Key points are not available for this paper at this time.
Theoreticians and experimentalists should work together more closely to establish reliable rankings and benchmarks for quantum chemical methods. Comparison to carefully designed experimental benchmark data should be a priority. Guidelines to improve the situation for experiments and calculations are proposed. Over the last years a subtle but profound change has taken place in chemical research. Electronic structure calculations have become ubiquitous, with much of the work published in the field today making use of theoretical results. The chemistry curricula have also been accompanying this change, incorporating a growing number of courses on quantum chemistry and computational chemistry, slowly but surely. The reasons behind this quiet revolution are all well known. The computing hardware has undergone significant developments, such that even a mobile phone surpasses the performance of computing clusters from 30 years back. This allows us to compute more and in less time, translating into larger systems, longer timescales. Furthermore, the computer algorithms have improved significantly, expanding the realm of application and the quality of the calculations performed. We find ourselves at a perceived turning point where quantum chemical calculations are believed by many to be on par with experimental methods. Even if the reader may not share this view, it is not difficult to find reasons why others would. Already in the 1960s we observe examples of theory competing with experiment in terms of accuracy. A seminal example is the H2 adiabatic dissociation energy computed by Kolos and Wolniewicz.1 The authors carried out a variational calculation and found their value (36 117.4 cm−1) to exceed the best experimental estimate at the time (36 113.6±0.6 cm−1).2 Given that the theoretical estimate would necessarily give an upper bound to the true energy of H2, and thus a lower bound to its dissociation energy, the experimental value was questioned. The episode was only concluded after a further experiment confirmed the theoretical result.3 Since then, there have been several cases where theory has made predictions which were only later confirmed by experiment,4-9 and many more where theory played an essential role in the interpretation of experimental results. Already the early H2 story points to a fascinating aspect, in particular now that the agreement between theory and experiment has progressed to eight significant digits.10 In contrast to other fields of science, where a model is mostly used as a framework to break down and understand complex systems, electronic structure methods are regarded as quantitative-data providers. This is a completely different scenario from that faced in biology or economics, just to name two examples. In many fields, the ruling equations are still unknown, or the conditions are ill-defined. An exact monthly weather forecast is unthinkable, but we chemists dare to believe that our theoretical models will be able to provide us with a reaction mechanism involving hundreds of atoms. This line of thought is fertile ground for illusions of grandeur, and calls for a critical analysis of the relation between experiment and theory in chemistry (Figure 1). A false change of paradigm. Experience has taught us that one should be critical on both ends. Since we are not philosophers of science (nor pretend to be), we will focus on a pragmatic view and only briefly touch on fundamental issues. In the language of Thomas Kuhn, quantum chemistry is largely in a rather unspectacular state of normal science.11 Apart from subtle issues such as parity violation,12 the Schrödinger equation and its relativistic variants are assumed to be essentially correct and complete for all practical purposes.13 Currently the cutting edge of quantum chemistry is simply modeling larger and larger systems in an increasingly quantitative way. Our concern is thus evolutionary rather than revolutionary. Beyond the usual method validations within the theoretical community, we address the relation between quantum chemical calculations and experiments. In order to guarantee a fruitful interplay, we need to define how to experimentally evaluate and benchmark the numerous methods in the best way. We argue that this subject is frequently neglected, and that this neglect leads to a slower development of quantum chemistry and the field of chemistry as a whole. While we are unable to provide unique and final recipes, we would like to share some of our thoughts and discussions with the community and conclude with some recommendations. Without prejudice against any of the methodologies covered under the general designation of theoretical/computational chemistry,14 this essay deals exclusively with electronic-structure methods, given their particular appeal and promise towards a first-principles description of chemical systems. Mechanical force fields and topological methods, for example, have a more pragmatic relation to experiment, and do not warrant the same type of discussion. The dynamics and statistics needed to connect the electronic-structure information to the real world would also be an important and critical topic of its own. This is beyond our scope, just like conceptual models and other qualitative theories. One permanent issue in quantum chemistry is to establish a practical overview of the many different approximations to the Schrödinger equation. Literally hundreds of electronic-structure methods are available at our fingertips, and it can be a cumbersome task to compare them or even to select one for a particular application. One way to deal with this question has been to establish a hierarchy of models. In the case of wave-function methods, the process can be rather straightforward. For single-reference correlated methods there are widely accepted orderings. The latter can be better understood through the use of diagrammatic representations. Generally, the method that includes the greater number of terms lies higher in the ranking. However, this is not necessarily true, since the nature of each term is different along with the impact on the overall performance. While a ranking for coupled cluster (CC) methods such as CCSD<CCSDT<CCSDTQ<CCSDTQP is rather straightforward, the ordering within the Møller–Plesset perturbation series (MP2, MP3, MP4, …) will depend on the system being studied, since it is not necessarily convergent.15 And although CCSDT is a more expensive approach, the CCSD(T) method is for most application purposes a more robust method than the full triples variant. This results from a quite favorable cancellation of error between the overestimated triples contribution and the neglect of quadruple excitations.16 Also in the case of density functional theory (DFT) there have been attempts at establishing similar hierarchies. The most well-known example is the DFT Jacob's Ladder, proposed by Perdew,17 which defines a set of steps featuring the different levels of approximation in the exchange-correlation kernels. Even for artificial molecules, such hierarchies have been shown to shine through.18 It is, however, relatively easy to find individual examples where a functional from a lower rank may be able to outperform methods higher up in the ladder19and truth be told, this was not the intended purpose of Perdew's proposal. Although such hierarchies can be questionable, they provide a very welcome order amidst the chaos of modern-day quantum chemistry toolboxes. They have also created the possibility to carry out theory benchmarks with theory as reference. The quality of a model is thereby no longer measured through any relation to experiment, but purely to the similarity to another model. This practice has become so popular that many manuscripts dedicated to quantum chemistry benchmarks do not feature a single experimental result.20 Sometimes, the word “experiment” is missing from the entire manuscript, or may at best be found once in the outlook. The current most complete set of benchmark sets, the GMTKN30,21 includes only a small amount of experimental reference data (Figure 2). The acceptance of CCSD(T) as a “gold standard”22 has been a particular encouragement to the practice, from DFT benchmarking23, 24 up to theory-only blind-test challenges.25 As further evidence, if one looks again at the GMTKN30 database, 14/30 sets use as reference data estimated CCSD(T)/CBS limits. Even improvements beyond CCSD(T) are sometimes judged without any reference to experiment.26 This is certainly related to the fact that most such calculations refer to a fixed geometry, although the importance of structural relaxation at least for selected degrees of freedom has been emphasized.27 Reference points in the GMTKN30 database, classified according to their origin (pure theoretical values or experiment-based). There are good reasons for theory-only benchmarking. Even weather forecasters can develop a reasonable sense for their simulation error by just comparing the results of disparate numerical models for the same starting conditions. Another obvious application is the comparison of different codes and numerical strategies within a model family.28 For the identification of isolated experimental database errors, sometimes even superficial comparisons between different computational predictions suffice.29 One may also be interested in a single quantity which is ill-defined or hard/impossible to measure experimentally (e.g., harmonic spectra,30 or transition-state structures31, 32) so that the best (or only) reference is the result of another calculation. Or there may simply be no satisfactory experimental data available or in reach for a specific class of compounds.33 Current state-of-the-art quantum chemical methods depend on a wide variety of approximations. They range from more technical issues such as density fitting34, 35 and numerical grids,36 up to developments such as explicit correlation,37 trying to reach the complete basis set limit. All these require the comparison to internal standards. However, the literature is filled with examples where the comparison to experiment is completely disregarded for no particularly good reason. There are many causes for this fault. On one hand, we have the focused education of theoretical chemists, who dedicate several years to equations and little time to lab experience. In other cases, one may find it difficult to establish relations between what has been measured in experiment and the results of calculations, or even in understanding the approximations the experiment itself involves. More trivially, insufficient time is spent in searching the literature due to the all too well known publication pressure.38 All of these issues can result in misguided comparisons (oranges and apples) leading to frustration or flawed benchmarking, or to a complete neglect of the published experimental data. There are also cases where the appropriate experimental literature is cited but not really digested by the theory/theory benchmark. As an example, we pick the ethanol dimer, which was shown semi-experimentally to revert the conformational preference of its monomer (trans) to a homochiral double-gauche arrangement in the lowest-energy pair structure.39, 40 The driving force is obviously a compact, dispersion-optimizing packing. A recent computational study claiming accurate (1 kJ mol−1) ab initio approximations for longer-chain alcohol clusters41 uses ethanol dimer as a stepping stone for the validation of their methods, which is not an unreasonable approach. Based on the results of a series of calculations of perceived increasing accuracy, the authors argue that in contrast to all the cited experimental and high-level computational evidence, the double-trans dimer is systematically the most stable dimer. This is due to a range of misconceptions, not the least of which is the overlooked difference between homo- and heterochiral pairings and the resulting apparently decreased importance of dispersion corrections. Even if some of the targeted cluster quantities happen to be close to older thermodynamic data, this is clearly no match for the right reason. Careful comparison to available spectroscopic data would immediately have revealed the flaws in this laborious study. We end this section with a little story told by Coulson42 about an exhibition of quantum chemistry in Paris after World War II. There were lovely diagrams of resonance structures of benzene and excellent numerical illustrations of the lowering in energy produced by them. But Linus Pauling, as he went round that exhibition and came to these diagrams, said, “Why don't you put a bottle of the stuff by the side of the diagrams?” At the end of the day, any weather-forecasting model has to be tested against reality, and success for one season or region is no guarantee for further successful predictions. Chemists also want reliable forecasts for exotic conditions,43 not only for well-trod paths. And even for those, there may be occasional surprises. In the end, there is no way around some experimental benchmarking. The benefits of this practice are obvious, when one looks back at the development of electronic-structure methods. Even though quantum chemists will often find shortcomings in their methods when performing theory/theory benchmarks, it is usually the hard test against experiment that brings about a change in practice. A recent example of this is the description of dispersion forces. Although conventional DFT functionals were well known to fail in the description of weakly bound van der Waals complexes,44 until about ten years ago a large community regarded this as a minor inconvenience for computational studies on “realistic” systems. Looking at the of was not and it was these small would out or A real change only came about when a series of benchmark studies with increasing system by experimental evidence, that the DFT functionals in use at the time were systematically and that the were in fact experiments to benchmark theory need to be carefully and the right have to be For the of a quantum chemistry may at the not be more quantitative than models in biology or is provide to popular from the theory side (Figure too it down to the to down the for a into the for a reliable of for the apparently of and such as relativistic or be by the can also a role in in different chemical systems and the best conditions for benchmarking. There is less exotic which is hard for experimentalists to energy or quantum of the an of the approximation which is so fruitful in electronic structure theory (Figure At least for the study of where theory/theory is very this can the system more than one would is needed from both in this and it is to have on how to best A between theory and the a state of The is to any in the experimental to test quantum chemical methods. This of with theory but also of as The is that you not you are the to the experimental result is known to the or the theoretical known to the there is a of the may the Another is chemists are very good at A single match between theory and experiment is it be the result of error A disparate data set is usually and this a of work at the it has become to density functionals on such benchmark data sets mostly theoretical we there is a for reliable experiments. While the of the theoretical is the of the benchmark should be fundamental The of a or the of a may be quantities to which quantum but or not a particular quantum chemical method the can often be most by simply at and of the points on the energy This why spectroscopic methods a role in benchmarking. can be for the correct if the structure or is experimentalists sometimes an to from their and system conditions and to address more In this the for error are A to the usually more at the This also allows much higher levels of to be used and these as for theory/theory In chemistry, are often less important than conformational experiments on such energy rankings of different structures for a given system may also be for it is for the comparison to quantum chemistry to and the For the case of may be although one should that the structure for is a completely different can be by experiments as a of chemical It out the between and of is not by the the functional this is for a good can be found out by the data examples of benchmark comparisons to experiment may be found in the literature and only a recent examples can be The different methods for adiabatic electronic on a large experimental database by can be rather to experiment, at least for small molecules, and their to a of the experimental to theoretical and also experimental for very large attempts to experiment and theory together can be in a by other theoretical leading to a between different A from such is the need for more accurate reference data for complex data in the provide a particularly and frequently for On the other end, to through experimental is about change and it would be too for to at the of structures and It is the to the H2 energy in reaction dynamics and the since the of its reaction it has been the subject of experimental experiments and to provide the most reference data for chemical dynamics and thus the Schrödinger equation. This also brings us to another are not only made through benchmark comparisons between theory and experiment for selected are also they have the to the of examples and even It to such are to other chemical systems. In this the role of experiments is than it may be that a out to be an experimental or even this quantum chemists have their with model or experimentalists have simply been too in their The latter is more often the case for experimental reference which are shown to be with state-of-the-art quantum chemical experimental they are publication through with which is of an excellent practice beyond benchmarking. Sometimes, there is a very fundamental and why theory and experiment such as for the between electronic of weakly bound More the situation is and theory with each other for a specific but for no good two or more A recent example is the from to as a of The force field is within experimental but only it the conformational energy and the dispersion between the two by similar can be in for variants of CCSD(T) calculations also with but now for much better It is quite that these calculations are more accurate than the experiment, which an estimated error of This a possibility spectroscopic and the in the But we will only for once a improved experiment some of on which have to be by with more application models have to be carefully by approximations in quantum In view of such important and the more and more between theory and experiment to be This change for the good of science, both from We that it would be very if quantum chemists and experimentalists were for the of in and forecasters should to the real weather and from a understanding of the models. In the end, we are all but in order to provide the best conditions for quantum chemistry to some be of Based on our and the recent we would like to Theoreticians should focus on a between the computed quantity and the experimental data. They should in touch with experimental and In the case of theory/theory benchmarks, should also a role in experimental data are can by an experimental should theory and work to establish They should that theoretical results and methods can and should be experimentally tested and in some cases even Theoreticians should published experimental results the results of their calculations are in although they have reasons to believe that they should not The can only both should be on selected experimental benchmark results. in all the there should be an to establish data sets for use in the quantum chemistry community, rather than only the popular data This should be by chemical of should an role in the latter and critical in from different of may as examples. Apart from the benchmark and experimentalists also by some of their so as to lower Theoreticians not only the methods, but also the to address the that in particular with wide acceptance in the community should be to this type of their work should provide all needed to and use of the This has to be by the of There is no current on what a theoretical study should provide in the and there is little of the in what should be A for energy and would be provide and on error and approximations in the calculations carried should to the quantum chemistry community results that are particularly to and to the quantum chemistry community experimental results where theory (or in a or on theoretical chemistry …) to the points experimental data by the available publication of critical This the to in quantum the growing in several to focus only on immediately and have to for the freedom to the fundamental points between their methods. And on this we with the that success will not us very This essay was and by a critical and of benchmark experiments with excellent to quantum We dedicate it to on the of We also our and for as well as the der for The authors no of has been a at the of since in the field of theoretical of to methods, and in some of are the and of computational from small to complex chemical systems. to experiment to theory in the field of to in and in quantum chemical Over the last two has and with to the and of and van der Waals is a of the of
Mata et al. (Fri,) studied this question.