Chembricks Blog
Complete results: six architectures on one learning curve
This page carries every table behind A head start from physics beats a bigger AI model. It is generated from one committed data file, so the figures in the article and the numbers here come from the same source.
The article calls the two setups learning from scratch and working with a physics head start. This page uses the names the data files carry, direct and delta. All errors are mean absolute error in kcal/mol, on the identical held-out test set of 49,235 molecules. The training pool holds 998,112 molecules with zero overlap against that test set.
Learning curves
Direct target, or learning from scratch
The wB97M-V total electronic energy minus a per-element composition fit. With no learning at all the reference sits at 22.6 to 23.3 kcal/mol.
| training molecules | chembricks-phy-AI-2 | SpookyNet | SchNet | MACE | chemprop | chembricks-phy-AI-1 |
|---|---|---|---|---|---|---|
| 1,000 | 9.453 | 20.151 | 15.479 | 18.409 | 15.266 | 23.034 |
| 2,000 | 7.360 | 12.767 | 11.599 | 20.862 | 11.090 | 19.728 |
| 5,000 | 6.681 | 8.193 | 10.244 | 12.268 | 9.429 | 19.578 |
| 10,000 | 5.277 | 6.226 | 8.823 | 7.069 | 10.469 | |
| 25,000 | 4.223 | 4.665 | 7.382 | 6.193 | 8.326 | |
| 50,000 | 3.690 | 6.820 | 5.748 | 7.204 | ||
| 100,000 | 4.465 | |||||
| 250,000 | 4.085 | |||||
| 500,000 | 4.000 | |||||
| 993,112 | 4.060 |
Delta target, or the physics head start
The same quantity after a g-xTB baseline has been subtracted as well. The reference alone sits at 2.74 to 2.99 kcal/mol.
| training molecules | chembricks-phy-AI-2 | SpookyNet | SchNet | MACE | chemprop | chembricks-phy-AI-1 |
|---|---|---|---|---|---|---|
| 1,000 | 1.682 | 2.264 | 1.842 | 2.345 | 2.489 | 2.623 |
| 2,000 | 1.356 | 1.879 | 1.555 | 1.640 | 1.853 | 2.047 |
| 5,000 | 1.142 | 1.461 | 1.311 | 1.226 | 1.521 | 1.668 |
| 10,000 | 1.038 | 1.290 | 1.047 | 1.298 | 1.368 | |
| 25,000 | 0.961 | 0.975 | 0.880 | 1.213 | 1.101 | |
| 50,000 | 0.829 | 0.775 | 1.168 | 0.906 | ||
| 100,000 | 1.019 | |||||
| 250,000 | 0.960 | |||||
| 500,000 | 0.940 | |||||
| 993,112 | 0.897 |
Cost of one training run
Core-hours for a single training run against the accuracy it bought. The two targets use different accounting, deliberately: on the direct target chembricks-phy-AI-2 alone is charged its g-xTB featurisation, because it is the only model that needs g-xTB there. On the delta target nobody is charged, because all six need it.
Direct target, learning from scratch: all 36 points
| model | training molecules | MAE kcal/mol | core-hours |
|---|---|---|---|
| chembricks-phy-AI-2 | 1,000 | 9.453 | 3.0238 |
| chembricks-phy-AI-2 | 2,000 | 7.360 | 3.1135 |
| chembricks-phy-AI-2 | 5,000 | 6.681 | 3.4049 |
| chembricks-phy-AI-2 | 10,000 | 5.277 | 4.0324 |
| chembricks-phy-AI-2 | 25,000 | 4.223 | 8.1840 |
| SpookyNet | 1,000 | 20.151 | 4.4598 |
| SpookyNet | 2,000 | 12.767 | 9.4164 |
| SpookyNet | 5,000 | 8.193 | 15.2559 |
| SpookyNet | 10,000 | 6.226 | 32.4914 |
| SpookyNet | 25,000 | 4.665 | 62.9665 |
| SpookyNet | 50,000 | 3.690 | 148.8932 |
| SchNet | 1,000 | 15.479 | 4.9327 |
| SchNet | 2,000 | 11.599 | 9.8724 |
| SchNet | 5,000 | 10.244 | 18.0934 |
| SchNet | 10,000 | 8.823 | 22.7771 |
| SchNet | 25,000 | 7.382 | 41.3035 |
| SchNet | 50,000 | 6.820 | 67.9368 |
| MACE | 1,000 | 18.409 | 130.0350 |
| MACE | 2,000 | 20.862 | 228.9185 |
| MACE | 5,000 | 12.268 | 472.6781 |
| chemprop | 1,000 | 15.266 | 0.2739 |
| chemprop | 2,000 | 11.090 | 0.2530 |
| chemprop | 5,000 | 9.429 | 0.3207 |
| chemprop | 10,000 | 7.069 | 0.7453 |
| chemprop | 25,000 | 6.193 | 1.5298 |
| chemprop | 50,000 | 5.748 | 2.5960 |
| chemprop | 100,000 | 4.465 | 6.4722 |
| chemprop | 250,000 | 4.085 | 7.7561 |
| chemprop | 500,000 | 4.000 | 13.5020 |
| chemprop | 993,112 | 4.060 | 31.1857 |
| chembricks-phy-AI-1 | 1,000 | 23.034 | 0.8870 |
| chembricks-phy-AI-1 | 2,000 | 19.728 | 1.7863 |
| chembricks-phy-AI-1 | 5,000 | 19.578 | 1.5731 |
| chembricks-phy-AI-1 | 10,000 | 10.469 | 7.3234 |
| chembricks-phy-AI-1 | 25,000 | 8.326 | 11.5733 |
| chembricks-phy-AI-1 | 50,000 | 7.204 | 17.9863 |
Frontier
| step | model | training molecules | core-hours | MAE kcal/mol |
|---|---|---|---|---|
| 1 | chemprop | 2,000 | 0.2530 | 11.090 |
| 2 | chemprop | 5,000 | 0.3207 | 9.429 |
| 3 | chemprop | 10,000 | 0.7453 | 7.069 |
| 4 | chemprop | 25,000 | 1.5298 | 6.193 |
| 5 | chemprop | 50,000 | 2.5960 | 5.748 |
| 6 | chembricks-phy-AI-2 | 10,000 | 4.0324 | 5.277 |
| 7 | chemprop | 100,000 | 6.4722 | 4.465 |
| 8 | chemprop | 250,000 | 7.7561 | 4.085 |
| 9 | chemprop | 500,000 | 13.5020 | 4.000 |
| 10 | SpookyNet | 50,000 | 148.8932 | 3.690 |
Delta target, the head start: all 36 points
| model | training molecules | MAE kcal/mol | core-hours |
|---|---|---|---|
| chembricks-phy-AI-2 | 1,000 | 1.682 | 0.0528 |
| chembricks-phy-AI-2 | 2,000 | 1.356 | 0.0790 |
| chembricks-phy-AI-2 | 5,000 | 1.142 | 0.1910 |
| chembricks-phy-AI-2 | 10,000 | 1.038 | 0.5100 |
| chembricks-phy-AI-2 | 25,000 | 0.961 | 3.8873 |
| SpookyNet | 1,000 | 2.264 | 4.9712 |
| SpookyNet | 2,000 | 1.879 | 8.1333 |
| SpookyNet | 5,000 | 1.461 | 13.8068 |
| SpookyNet | 10,000 | 1.290 | 16.5356 |
| SpookyNet | 25,000 | 0.975 | 39.2608 |
| SpookyNet | 50,000 | 0.829 | 77.4899 |
| SchNet | 1,000 | 1.842 | 4.7972 |
| SchNet | 2,000 | 1.555 | 9.8994 |
| SchNet | 5,000 | 1.311 | 16.4893 |
| SchNet | 10,000 | 1.047 | 26.1646 |
| SchNet | 25,000 | 0.880 | 50.1114 |
| SchNet | 50,000 | 0.775 | 102.5321 |
| MACE | 1,000 | 2.345 | 129.7634 |
| MACE | 2,000 | 1.640 | 228.9533 |
| MACE | 5,000 | 1.226 | 473.7951 |
| chemprop | 1,000 | 2.489 | 0.0947 |
| chemprop | 2,000 | 1.853 | 0.1573 |
| chemprop | 5,000 | 1.521 | 0.2854 |
| chemprop | 10,000 | 1.298 | 0.3580 |
| chemprop | 25,000 | 1.213 | 1.1496 |
| chemprop | 50,000 | 1.168 | 1.8961 |
| chemprop | 100,000 | 1.019 | 4.3929 |
| chemprop | 250,000 | 0.960 | 7.3272 |
| chemprop | 500,000 | 0.940 | 14.7699 |
| chemprop | 993,112 | 0.897 | 21.0877 |
| chembricks-phy-AI-1 | 1,000 | 2.623 | 1.1507 |
| chembricks-phy-AI-1 | 2,000 | 2.047 | 2.0597 |
| chembricks-phy-AI-1 | 5,000 | 1.668 | 3.4364 |
| chembricks-phy-AI-1 | 10,000 | 1.368 | 4.7581 |
| chembricks-phy-AI-1 | 25,000 | 1.101 | 13.7566 |
| chembricks-phy-AI-1 | 50,000 | 0.906 | 28.3505 |
Frontier
| step | model | training molecules | core-hours | MAE kcal/mol |
|---|---|---|---|---|
| 1 | chembricks-phy-AI-2 | 1,000 | 0.0528 | 1.682 |
| 2 | chembricks-phy-AI-2 | 2,000 | 0.0790 | 1.356 |
| 3 | chembricks-phy-AI-2 | 5,000 | 0.1910 | 1.142 |
| 4 | chembricks-phy-AI-2 | 10,000 | 0.5100 | 1.038 |
| 5 | chembricks-phy-AI-2 | 25,000 | 3.8873 | 0.961 |
| 6 | chemprop | 250,000 | 7.3272 | 0.960 |
| 7 | chemprop | 500,000 | 14.7699 | 0.940 |
| 8 | chemprop | 993,112 | 21.0877 | 0.897 |
| 9 | SchNet | 25,000 | 50.1114 | 0.880 |
| 10 | SpookyNet | 50,000 | 77.4899 | 0.829 |
| 11 | SchNet | 50,000 | 102.5321 | 0.775 |
Measured study cost
Process CPU accounting on production-length runs, not wall-clock estimates. MACE never early-stopped, so part of its total is protocol rather than architecture.
| model | runs | wall process-hours | measured core-hours | cores busy | utilisation |
|---|---|---|---|---|---|
| MACE | 18 | 643.35 | 4,992.4 | 7.76 of 8 | 97% |
| SpookyNet | 28 | 87.24 | 697.0 | 7.99 of 8 | 100% |
| SchNet | 28 | 80.63 | 643.4 | 7.98 of 8 | 100% |
| chemprop | 48 | 15.29 | 161.4 | 6.26 to 10.92 of 12 | 52 to 91% |
| chembricks-phy-AI-1 | 28 | 19.28 | 153.8 | 7.98 of 8 | 100% |
| chembricks-phy-AI-2 | 30 | 1.56 | 35.6 | 24.4 to 25.2 of 32 | 77% |
Interpolation against extrapolation
The training pool is built from small fragments, so molecules above 13 heavy atoms are outside the training size distribution entirely.
| model | target | training molecules | 13 heavy atoms or fewer | more than 13 heavy atoms | extrapolation better by |
|---|---|---|---|---|---|
| SpookyNet | direct | 50,000 | 4.228 | 3.160 | +25% |
| SchNet | direct | 50,000 | 6.495 | 7.901 | -22% |
| chembricks-phy-AI-2 | direct | 25,000 | 5.414 | 3.052 | +44% |
| SpookyNet | delta | 50,000 | 0.960 | 0.699 | +27% |
| chembricks-phy-AI-2 | delta | 25,000 | 1.130 | 0.796 | +30% |
| MACE | delta | 5,000 | 1.448 | 1.009 | +30% |
Methods
This is the detail the article leaves out. It is here so that every claim on that page can be checked rather than taken on trust.
The two targets, precisely
Both targets are residuals after a linear per-element composition fit, fitted on the training draw only. The direct target is the wB97M-V energy minus that fit. The delta target subtracts a g-xTB single point as well.
The composition term is not cosmetic on the delta arm. g-xTB is all-electron while this DFT uses ccECP pseudopotentials, so the raw difference between the two is dominated by core-electron energy the DFT never computed, and has a wider spread than the DFT energy itself. For HBr the g-xTB energy is minus 2,574.60 Hartree against a DFT value of minus 13.93. The per-element fit absorbs that offset before any learning starts.
The split, and the leak in the first version of it
The split originally intended for this study had complete test leakage: the file meant for testing was a strict subset of the file meant for training. The partition was rebuilt on data origin instead.
| set | molecules |
|---|---|
| training pool | 998,112 |
| test set as partitioned | 49,290 |
| dropped, present in both origins | 4,391 |
| excluded, elements occurring only in the test set | 55 |
| test set scored by every model | 49,235 |
Zero overlap between training and test is asserted when the split is built and re-asserted on every load. A further 55 conformers containing arsenic, selenium or tellurium occur only in the test set, where a composition reference cannot place them, so they are excluded and all six models score the identical 49,235 molecules. The training pool is drawn as nested prefixes of one shuffle per seed, so a curve moves with data rather than with luck, and a fixed 5,000-molecule validation set is held out of every draw.
Shared protocol
Five of the six models run through one harness with identical split, reference, targets, training loop and metrics. chemprop and chembricks-phy-AI-2 use their own harnesses on the same split and the same test set, verified molecule for molecule.
| setting | value |
|---|---|
| cutoff | 6.0 Å |
| optimiser | Adam |
| learning rate | 3e-4 |
| batch size | 64 |
| max epochs | 200 |
| patience | 20 |
| gradient clip | 10.0 |
| LR schedule | plateau, patience 5, factor 0.5 |
| weight averaging | EMA, decay 0.999 |
| loss normalisation | normalise_by_size=True (per-site) |
| target unit | kcal/mol |
Per-site loss normalisation matters here: without it a 76-atom structure would dominate a 3-atom one, and this corpus spans exactly that range. Molecules are split for reporting at 13 heavy atoms, since the training pool is built from small fragments and anything above that is genuinely outside the training size distribution.
The six models
Two of the six are ours. chembricks-phy-AI-1 is the message-passing network the study calls cmbx-nn. chembricks-phy-AI-2 is the kernel method it calls SORF, structured orthogonal random features over the FCHL19 representation with g-xTB charges, solved as ridge regression. Those internal identifiers are what the source data files use, so every number here can be traced back to them.
| model | type | parameters | uses 3-D geometry | gets g-xTB charges and bond orders |
|---|---|---|---|---|
| chemprop 2.3.1 | 2-D D-MPNN | ~1.4 M | no | no |
| chembricks-phy-AI-1 | site-wise MPNN | ~1.1 M | yes | yes |
| SchNet (PyG) | continuous-filter CNN | 455,809 | yes | no |
| MACE | equivariant (e3nn) | 936,848 / 984,336 | yes | no |
| SpookyNet | MPNN + nonlocal attention | 498,351 | yes | no |
| chembricks-phy-AI-2 | random features + ridge | 65,536 coefficients | yes | yes, both arms |
Physics terms in SpookyNet (repulsion, electrostatics and dispersion) were disabled, because they contribute on the absolute energy scale while the target here is a kcal/mol residual. Its nonlocal attention, the more distinctive part of the architecture, is untouched. MACE ran at its own defaults, with atomic energies zeroed because the harness already subtracts the composition reference.
The one protocol deviation, and its control
SpookyNet does not train under the shared protocol. It zero-initialises its nuclear embedding, so its output starts at exactly zero, and the shared exponential moving average keeps dragging the evaluated weights back to that state. At 1,000 molecules on the delta target it returned a 0.3% gain over the reference; with weight averaging off and a higher learning rate it returned 24.9%.
Since SpookyNet then took the best direct result and was the only model given any hyperparameter attention, the obvious worry is that its win is tuning rather than architecture. The control was to run SchNet at 5,000 molecules on the direct target, same seed, under SpookyNet's exact settings.
| run | MAE kcal/mol | epochs |
|---|---|---|
| shared protocol (lr 3e-4, EMA decay 0.999) | 11.049 | 193 |
| SpookyNet's protocol (lr 1e-3, no weight averaging) | 11.405 | 59 |
SpookyNet's settings make SchNet 3.2% worse, so the deviation is not a general advantage. Its result stands as a finding about the architecture.
Back to the article: A head start from physics beats a bigger AI model. A subset of the underlying data is public under an MIT license at chembricks/WB96MV-ORGANIC.
Get new articles by email
Notes on physics validation for AI chemistry, straight from the team. No spam.
Double opt-in. Unsubscribe anytime.