Introducing Chembricks

From proprietary quantum-chemistry data to physics-gated molecular design

Chembricks: physics-gated molecular design

Welcome to the Chembricks blog

Before the technical story, a quick word on why we keep a blog at all.

Chembricks builds tools for doing chemistry with AI agents, but doing it properly: with real physics underneath and an honest account of what has and hasn't been validated. That's a space full of hype, and we'd rather show our work than add to the noise. This blog is where we'll do that: walk through how the platform works, publish worked examples end-to-end, be candid about where our models are strong and where they aren't, and share what we're learning as we go.

We'll start where it matters most: with the idea the whole company is built on.

The problem with "AI for chemistry"

A language model can produce a confident-sounding answer about almost any molecule. The trouble is that a plausible sentence and a physically correct one look identical on the page. Generative chemistry has a habit of proposing molecules that read beautifully and then fall apart the moment anyone checks the underlying thermodynamics, usually after they've reached a lab bench.

Chembricks is not a language model answering chemistry questions from memory. It is an agentic environment in which an AI plans the work and named scientific tools do the calculating: proprietary Chembricks physics models and open, peer-reviewed engines, coordinated over large molecular datasets, with an explicit reliability layer that decides what may actually be claimed.

The design philosophy is a physics-gated loop: cognition proposes, physics validates, survivors advance. A reasoning agent frames a problem and proposes candidates; the platform's physics tools calculate the properties that decide whether each candidate can survive; only the candidates that pass feed the next cycle. It inverts the usual failure mode of generative chemistry.

The physics-gated decision loop
The physics-gated decision loop: frame, propose, calculate, challenge, learn, and then decide.

What makes the toolbox trustworthy is the quality of the numbers behind it. Every model under the cmbx_ namespace is built on quantum-chemistry reference data we compute in-house with property-specific protocols. The emphasis is on having the right data at high accuracy, rather than throwing large volumes of low-quality data at a model. That efficient data-generation methodology is our protected core asset.

A few numbers to set the scale: the production catalog holds 155 callable tools across 22 scientific categories, every result is tagged with one of 5 evidence classes, and answers export to 4 formats. But the count isn't the point. The point is what happens between the tools.

How it works: the agent is a workflow controller

The platform treats molecular work as an iterative decision loop, not a one-shot answer. Each result can change the next calculation, eliminate a candidate, trigger a cross-check, or escalate the method. The agent can fan out parallel calculations, manage asynchronous jobs, inspect intermediate results, revise a candidate set, fit a surrogate, request the next informative batch, and update a live Pareto front, all within one brief.

A validation layer runs underneath at three levels:

  • Input integrity. Structures are canonicalized, charge and spin are tracked, atoms are balanced, and unsupported inputs are rejected rather than silently coerced.
  • Method fitness. A property-specific cmbx_ model or external engine is selected for the task; vacuum energies and implicit-solvent energies are never mixed inside one thermodynamic cycle.
  • Decision fitness. Candidates are ranked across several objectives at once, constraint violations stay visible, and the system states where higher-level theory or experiment becomes mandatory.

The model foundry: why a small dataset generalizes

The proprietary layer is a foundry: Chembricks-computed quantum-chemistry reference data becomes deployed cmbx_ model services, while open engines remain a distinct, attributed layer the agent composes on top.

The Chembricks model foundry
The model foundry: proprietary quantum-chemistry data becomes trained cmbx_ services, with open engines as a separate, attributed layer.

The deployed cmbx_ models are trained surrogates that reproduce quantum-chemical reference quantities in milliseconds where the underlying DFT calculation would take minutes to hours. Throughout, "near-DFT" means a model reproduces its trained reference method (wB97M-V or M06-2X) to near that method's own accuracy; agreement with experiment is bounded by the reference method's error, not implied by the phrase. A few of them:

  • cmbx_internal_energy_v2: near-DFT internal and atomization energies, trained on order 10⁵ molecules with FCHL19 / aSLATM representations.
  • cmbx_solvation_logp_v2: vacuum, water and octanol solvation free energies and logP, computed at the M06-2X / cc-pVDZ level with the SMD solvation model.
  • cmbx_redox: gas-phase ionization potential and electron affinity, a delta-ML correction trained to reproduce wB97M-V references from a g-xTB baseline.

Chemistry as building blocks

Here's the part we're most proud of, and the principle we're named for. The cmbx_ models generalize because of how they represent a molecule. Rather than memorizing whole structures, the learning approach decomposes chemistry into transferable building blocks: local atomic environments whose behaviour recurs across organic chemistry. High-accuracy reference data computed on a compact, deliberately chosen set of these motifs then carries over to larger, previously unseen molecules assembled from the same blocks.

The transfer is bounded, not universal: it is weakest for environments absent from training and for strongly non-local or long-range electronic effects, so an applicability-domain check travels with every prediction rather than being assumed. This is why a small, carefully generated dataset can outperform a large, noisy one. The decomposition-and-generalization method is patent-pending; this post describes the principle, not the recipe.

Open science, and giving back to it

The platform is deliberately two-layer. The open layer is a curated set of peer-reviewed academic engines: RDKit, the xTB family, CREST, RMG, AiZynthFinder, Morfeus, CASCADE, BoTorch and more, each keeping its own method identity in the evidence ledger. The proprietary layer is what open tools can't give: models trained on data we compute ourselves. Neither suffices alone. We stand on decades of open academic work and aim to contribute back, naming and attributing every engine we build on, and preparing an openly documented benchmark suite the community can use and check.

Capability atlas
Eight client-facing capability domains span the 22 functional categories of the production catalog.

Case study 1: designing a reversible hydrogen carrier

Liquid organic hydrogen carriers (LOHCs) store hydrogen in reversible chemical bonds and release it by catalytic dehydrogenation. The design problem is inherently multi-objective: storage capacity, release energy, phase behaviour, stability, catalyst compatibility, safety, cost and availability all pull in different directions.

Reversible hydrogenation scheme
A carrier cycles between a hydrogen-lean and a hydrogen-rich form; every candidate must be scored with the same energy method, balanced atoms and a matched reference state.

Two quantities drive the first pass. Gravimetric capacity is a structural fact; release energy per H₂ is a thermochemical claim. And this is a thermodynamic screen only: release energy sets an equilibrium floor, but the temperature a carrier actually dehydrogenates at (typically 170 to 350 °C over Pt/Pd) is fixed by kinetics and catalyst, and long-term viability turns on cycling stability the screen doesn't address. A favourable release energy is necessary, not sufficient.

The agent enumerated eleven carrier chemistries as balanced lean/rich pairs, scored both forms with cmbx_internal_energy_v2, fixed the hydrogen reference against a single benzene/cyclohexane anchor, then predicted every other pair blind and ranked capacity against release energy on a live Pareto front.

LOHC design space
The live design space: each point is a lean/rich pair; N-heterocycles cluster at lower release energy, easier to unload.

Crucially, the screen is checked against experiment, not just plotted. Across the held-out pairs the mean absolute error is ≈0.16 kcal/mol H₂ (a small set, and mostly close analogues of the anchor; we're honest about that). And against a raw general-purpose GFN2-xTB baseline, the trained model's real advantage shows on chemically distant carriers, where a single global energy shift can't reach.

LOHC validation and baseline comparison
Held-out validation, and the trained model versus a raw semiempirical baseline.

The ranking is chemically legible: carbazole tops the trade-off but melts near 246 °C, which is exactly why N-ethylcarbazole and dibenzyltoluene are the pairs a real program carries forward. The screen reduces the search space; it does not replace expert phase, catalyst and safety judgement.

Case study 2: selecting a copper-over-iron ligand

Separating copper from iron is a selectivity problem: the goal is a ligand that binds Cu(II) while leaving Fe(III) behind. We use it here as a clean model problem for that kind of reasoning.

Copper over iron selectivity
Selectivity margin toward copper for two model chelators, from physics-based metal docking.

The result reproduces textbook hard-soft acid-base and charge-density chemistry directly from computed thermodynamics: the neutral N-donor is copper-selective (Cu binds, Fe(III) is rejected: a clean sign flip), while the anionic carboxylate is not copper-selective (both bind strongly; the raw ordering is dominated by a charge-neutralisation term and isn't a genuine iron preference). The robust output is the sign, not the magnitude.

Cu-en stability ladder
Stepwise Cu(II)-ethylenediamine stability constants in the model's uncalibrated reference frame; the values and step count are model artifacts.

We're deliberately careful here. A monotonic ladder is generic to any successive-addition model; the model also binds ethylenediamine through one nitrogen rather than resolving the chelate ring, so the computed margin is best read as a donor-type and charge signal, and the workflow presents this as a ranked selectivity signal with higher-level calculation and bench extraction tests as the mandatory next steps.

Evidence you can challenge

Two things we won't do: publish flattering domain-wide accuracy numbers from small internal sets, or hide a capability that underperforms. Instead the platform reports which route to trust for which property and where its limits lie, and we're building the openly documented benchmark suite the field currently lacks, and we intend to hold our own cmbx_ models to it in public, with error bars and stated applicability domains.

And every answer is an inspectable evidence package, not just a paragraph. In a field crowded with black-box predictions, every number is traceable to the exact method that produced it and tagged by evidence class, auditable enough to defend in a regulatory, safety or investment review.

Evidence package flow
Every claim carries an evidence card, a raw tool ledger and a structured export, with five evidence classes kept visually distinct.

Built for chemists, not just programmers

None of this comes with a programming tax. A chemist with no coding or AI background drives the whole platform by typing a question in plain English: no SMILES to write, no scripts, no environment to set up. The physics tools, model selection and provenance ledger all run underneath.

Plain question to readable answer with evidence
One plain-language question, answered on the production system, with the structure and computed numbers exposed for anyone who wants to check them.

The same conversational surface that runs the two case studies above also answers a first-year student asking which of two familiar molecules is more water-soluble, quietly running the real solubility and property tools and then returning an answer pitched at the reader while still exposing the evidence. The accessibility is in the interface, not in a weaker calculation.

Two more live examples

Because a screenshot beats a claim, here are two more runs, executed end-to-end on the production system at cmbx.ai/chat.

Reactivity map of quinoline. Asked in one sentence to compute Fukui indices for quinoline and explain its reactive sites, the agent canonicalized the SMILES, ran an xtb_fukui_indices calculation, mapped the values onto numbered atoms, and identified the ring nitrogen N1 as both the most nucleophilic site (f⁻ = 0.182, the pyridine lone pair) and the most electrophilic (f⁺ = 0.118, then C4 and C8 on the π* system), consistent with textbook quinoline chemistry.

IP and EA across the acenes. A second run computed the ionization potential and electron affinity of naphthalene, anthracene and tetracene with the proprietary cmbx_redox model and explained the trend for electron-transport materials: electron affinity rises with ring count (from ≈0.16 eV for naphthalene toward ≈1.35 eV for tetracene) while the ionization potential falls, lowering the barrier to electron injection for n-type transport, with anthracene near the ambipolar boundary and air-stability flagged as the practical trade-off. The absolute values are model outputs; the ordering and trend are the decision signal.

Try it. The fastest way to understand any of this is to run your own molecule.

Request access Explore the platform

Screening and simulation results support decisions; they do not replace experimental, regulatory or safety validation. We'll keep publishing worked examples, and their honest boundaries, here.

The Chembricks team