Chembricks Blog

Complete results: six architectures on one learning curve

This page carries every table behind A head start from physics beats a bigger AI model. It is generated from one committed data file, so the figures in the article and the numbers here come from the same source.

The article calls the two setups learning from scratch and working with a physics head start. This page uses the names the data files carry, direct and delta. All errors are mean absolute error in kcal/mol, on the identical held-out test set of 49,235 molecules. The training pool holds 998,112 molecules with zero overlap against that test set.

Learning curves

Direct target, or learning from scratch

The wB97M-V total electronic energy minus a per-element composition fit. With no learning at all the reference sits at 22.6 to 23.3 kcal/mol.

training molecules chembricks-phy-AI-2 SpookyNet SchNet MACE chemprop chembricks-phy-AI-1
1,000 9.453 20.151 15.479 18.409 15.266 23.034
2,000 7.360 12.767 11.599 20.862 11.090 19.728
5,000 6.681 8.193 10.244 12.268 9.429 19.578
10,000 5.277 6.226 8.823 7.069 10.469
25,000 4.223 4.665 7.382 6.193 8.326
50,000 3.690 6.820 5.748 7.204
100,000 4.465
250,000 4.085
500,000 4.000
993,112 4.060

Delta target, or the physics head start

The same quantity after a g-xTB baseline has been subtracted as well. The reference alone sits at 2.74 to 2.99 kcal/mol.

training molecules chembricks-phy-AI-2 SpookyNet SchNet MACE chemprop chembricks-phy-AI-1
1,000 1.682 2.264 1.842 2.345 2.489 2.623
2,000 1.356 1.879 1.555 1.640 1.853 2.047
5,000 1.142 1.461 1.311 1.226 1.521 1.668
10,000 1.038 1.290 1.047 1.298 1.368
25,000 0.961 0.975 0.880 1.213 1.101
50,000 0.829 0.775 1.168 0.906
100,000 1.019
250,000 0.960
500,000 0.940
993,112 0.897

Cost of one training run

Core-hours for a single training run against the accuracy it bought. The two targets use different accounting, deliberately: on the direct target chembricks-phy-AI-2 alone is charged its g-xTB featurisation, because it is the only model that needs g-xTB there. On the delta target nobody is charged, because all six need it.

Direct target, learning from scratch: all 36 points

model training molecules MAE kcal/mol core-hours
chembricks-phy-AI-2 1,000 9.453 3.0238
chembricks-phy-AI-2 2,000 7.360 3.1135
chembricks-phy-AI-2 5,000 6.681 3.4049
chembricks-phy-AI-2 10,000 5.277 4.0324
chembricks-phy-AI-2 25,000 4.223 8.1840
SpookyNet 1,000 20.151 4.4598
SpookyNet 2,000 12.767 9.4164
SpookyNet 5,000 8.193 15.2559
SpookyNet 10,000 6.226 32.4914
SpookyNet 25,000 4.665 62.9665
SpookyNet 50,000 3.690 148.8932
SchNet 1,000 15.479 4.9327
SchNet 2,000 11.599 9.8724
SchNet 5,000 10.244 18.0934
SchNet 10,000 8.823 22.7771
SchNet 25,000 7.382 41.3035
SchNet 50,000 6.820 67.9368
MACE 1,000 18.409 130.0350
MACE 2,000 20.862 228.9185
MACE 5,000 12.268 472.6781
chemprop 1,000 15.266 0.2739
chemprop 2,000 11.090 0.2530
chemprop 5,000 9.429 0.3207
chemprop 10,000 7.069 0.7453
chemprop 25,000 6.193 1.5298
chemprop 50,000 5.748 2.5960
chemprop 100,000 4.465 6.4722
chemprop 250,000 4.085 7.7561
chemprop 500,000 4.000 13.5020
chemprop 993,112 4.060 31.1857
chembricks-phy-AI-1 1,000 23.034 0.8870
chembricks-phy-AI-1 2,000 19.728 1.7863
chembricks-phy-AI-1 5,000 19.578 1.5731
chembricks-phy-AI-1 10,000 10.469 7.3234
chembricks-phy-AI-1 25,000 8.326 11.5733
chembricks-phy-AI-1 50,000 7.204 17.9863

Frontier

step model training molecules core-hours MAE kcal/mol
1 chemprop 2,000 0.2530 11.090
2 chemprop 5,000 0.3207 9.429
3 chemprop 10,000 0.7453 7.069
4 chemprop 25,000 1.5298 6.193
5 chemprop 50,000 2.5960 5.748
6 chembricks-phy-AI-2 10,000 4.0324 5.277
7 chemprop 100,000 6.4722 4.465
8 chemprop 250,000 7.7561 4.085
9 chemprop 500,000 13.5020 4.000
10 SpookyNet 50,000 148.8932 3.690

Delta target, the head start: all 36 points

model training molecules MAE kcal/mol core-hours
chembricks-phy-AI-2 1,000 1.682 0.0528
chembricks-phy-AI-2 2,000 1.356 0.0790
chembricks-phy-AI-2 5,000 1.142 0.1910
chembricks-phy-AI-2 10,000 1.038 0.5100
chembricks-phy-AI-2 25,000 0.961 3.8873
SpookyNet 1,000 2.264 4.9712
SpookyNet 2,000 1.879 8.1333
SpookyNet 5,000 1.461 13.8068
SpookyNet 10,000 1.290 16.5356
SpookyNet 25,000 0.975 39.2608
SpookyNet 50,000 0.829 77.4899
SchNet 1,000 1.842 4.7972
SchNet 2,000 1.555 9.8994
SchNet 5,000 1.311 16.4893
SchNet 10,000 1.047 26.1646
SchNet 25,000 0.880 50.1114
SchNet 50,000 0.775 102.5321
MACE 1,000 2.345 129.7634
MACE 2,000 1.640 228.9533
MACE 5,000 1.226 473.7951
chemprop 1,000 2.489 0.0947
chemprop 2,000 1.853 0.1573
chemprop 5,000 1.521 0.2854
chemprop 10,000 1.298 0.3580
chemprop 25,000 1.213 1.1496
chemprop 50,000 1.168 1.8961
chemprop 100,000 1.019 4.3929
chemprop 250,000 0.960 7.3272
chemprop 500,000 0.940 14.7699
chemprop 993,112 0.897 21.0877
chembricks-phy-AI-1 1,000 2.623 1.1507
chembricks-phy-AI-1 2,000 2.047 2.0597
chembricks-phy-AI-1 5,000 1.668 3.4364
chembricks-phy-AI-1 10,000 1.368 4.7581
chembricks-phy-AI-1 25,000 1.101 13.7566
chembricks-phy-AI-1 50,000 0.906 28.3505

Frontier

step model training molecules core-hours MAE kcal/mol
1 chembricks-phy-AI-2 1,000 0.0528 1.682
2 chembricks-phy-AI-2 2,000 0.0790 1.356
3 chembricks-phy-AI-2 5,000 0.1910 1.142
4 chembricks-phy-AI-2 10,000 0.5100 1.038
5 chembricks-phy-AI-2 25,000 3.8873 0.961
6 chemprop 250,000 7.3272 0.960
7 chemprop 500,000 14.7699 0.940
8 chemprop 993,112 21.0877 0.897
9 SchNet 25,000 50.1114 0.880
10 SpookyNet 50,000 77.4899 0.829
11 SchNet 50,000 102.5321 0.775

Measured study cost

Process CPU accounting on production-length runs, not wall-clock estimates. MACE never early-stopped, so part of its total is protocol rather than architecture.

model runs wall process-hours measured core-hours cores busy utilisation
MACE 18 643.35 4,992.4 7.76 of 8 97%
SpookyNet 28 87.24 697.0 7.99 of 8 100%
SchNet 28 80.63 643.4 7.98 of 8 100%
chemprop 48 15.29 161.4 6.26 to 10.92 of 12 52 to 91%
chembricks-phy-AI-1 28 19.28 153.8 7.98 of 8 100%
chembricks-phy-AI-2 30 1.56 35.6 24.4 to 25.2 of 32 77%

Interpolation against extrapolation

The training pool is built from small fragments, so molecules above 13 heavy atoms are outside the training size distribution entirely.

model target training molecules 13 heavy atoms or fewer more than 13 heavy atoms extrapolation better by
SpookyNet direct 50,000 4.228 3.160 +25%
SchNet direct 50,000 6.495 7.901 -22%
chembricks-phy-AI-2 direct 25,000 5.414 3.052 +44%
SpookyNet delta 50,000 0.960 0.699 +27%
chembricks-phy-AI-2 delta 25,000 1.130 0.796 +30%
MACE delta 5,000 1.448 1.009 +30%

Methods

This is the detail the article leaves out. It is here so that every claim on that page can be checked rather than taken on trust.

The two targets, precisely

Both targets are residuals after a linear per-element composition fit, fitted on the training draw only. The direct target is the wB97M-V energy minus that fit. The delta target subtracts a g-xTB single point as well.

The composition term is not cosmetic on the delta arm. g-xTB is all-electron while this DFT uses ccECP pseudopotentials, so the raw difference between the two is dominated by core-electron energy the DFT never computed, and has a wider spread than the DFT energy itself. For HBr the g-xTB energy is minus 2,574.60 Hartree against a DFT value of minus 13.93. The per-element fit absorbs that offset before any learning starts.

The split, and the leak in the first version of it

The split originally intended for this study had complete test leakage: the file meant for testing was a strict subset of the file meant for training. The partition was rebuilt on data origin instead.

set molecules
training pool 998,112
test set as partitioned 49,290
dropped, present in both origins 4,391
excluded, elements occurring only in the test set 55
test set scored by every model 49,235

Zero overlap between training and test is asserted when the split is built and re-asserted on every load. A further 55 conformers containing arsenic, selenium or tellurium occur only in the test set, where a composition reference cannot place them, so they are excluded and all six models score the identical 49,235 molecules. The training pool is drawn as nested prefixes of one shuffle per seed, so a curve moves with data rather than with luck, and a fixed 5,000-molecule validation set is held out of every draw.

Shared protocol

Five of the six models run through one harness with identical split, reference, targets, training loop and metrics. chemprop and chembricks-phy-AI-2 use their own harnesses on the same split and the same test set, verified molecule for molecule.

setting value
cutoff 6.0 Å
optimiser Adam
learning rate 3e-4
batch size 64
max epochs 200
patience 20
gradient clip 10.0
LR schedule plateau, patience 5, factor 0.5
weight averaging EMA, decay 0.999
loss normalisation normalise_by_size=True (per-site)
target unit kcal/mol

Per-site loss normalisation matters here: without it a 76-atom structure would dominate a 3-atom one, and this corpus spans exactly that range. Molecules are split for reporting at 13 heavy atoms, since the training pool is built from small fragments and anything above that is genuinely outside the training size distribution.

The six models

Two of the six are ours. chembricks-phy-AI-1 is the message-passing network the study calls cmbx-nn. chembricks-phy-AI-2 is the kernel method it calls SORF, structured orthogonal random features over the FCHL19 representation with g-xTB charges, solved as ridge regression. Those internal identifiers are what the source data files use, so every number here can be traced back to them.

model type parameters uses 3-D geometry gets g-xTB charges and bond orders
chemprop 2.3.1 2-D D-MPNN ~1.4 M no no
chembricks-phy-AI-1 site-wise MPNN ~1.1 M yes yes
SchNet (PyG) continuous-filter CNN 455,809 yes no
MACE equivariant (e3nn) 936,848 / 984,336 yes no
SpookyNet MPNN + nonlocal attention 498,351 yes no
chembricks-phy-AI-2 random features + ridge 65,536 coefficients yes yes, both arms

Physics terms in SpookyNet (repulsion, electrostatics and dispersion) were disabled, because they contribute on the absolute energy scale while the target here is a kcal/mol residual. Its nonlocal attention, the more distinctive part of the architecture, is untouched. MACE ran at its own defaults, with atomic energies zeroed because the harness already subtracts the composition reference.

The one protocol deviation, and its control

SpookyNet does not train under the shared protocol. It zero-initialises its nuclear embedding, so its output starts at exactly zero, and the shared exponential moving average keeps dragging the evaluated weights back to that state. At 1,000 molecules on the delta target it returned a 0.3% gain over the reference; with weight averaging off and a higher learning rate it returned 24.9%.

Since SpookyNet then took the best direct result and was the only model given any hyperparameter attention, the obvious worry is that its win is tuning rather than architecture. The control was to run SchNet at 5,000 molecules on the direct target, same seed, under SpookyNet's exact settings.

run MAE kcal/mol epochs
shared protocol (lr 3e-4, EMA decay 0.999) 11.049 193
SpookyNet's protocol (lr 1e-3, no weight averaging) 11.405 59

SpookyNet's settings make SchNet 3.2% worse, so the deviation is not a general advantage. Its result stands as a finding about the architecture.


Back to the article: A head start from physics beats a bigger AI model. A subset of the underlying data is public under an MIT license at chembricks/WB96MV-ORGANIC.