Components from open-source chemistry
FormulaIQ Layer 1 is fed by a multi-source open materials ETL — not a single vendor dump. Components stay CAS-keyed into inverse formulation.
Formula · open data
Open chemistry sources
Technology stack
Formula components from open-source chemistry data, ML/DL models we train, LLM evaluation, regulation RAG, workflow simulation & digital twins, process control commissioned from plant data, and Chemical Engineering Agents — a growing fleet for chemistry.
prediction = first-principles baseline + ML residual ± uncertainty
One inspectable stack. Each layer is purpose-built for chemistry — not a generic LLM wrapper.
Component-level properties and structures from multiple open chemistry databases — QM9, Materials Project, OQMD, AFLOW, NOMAD, PoLyInfo, Polymer Genome, PubChem, ZINC — normalised onto a CAS-keyed backbone for FormulaIQ Layer 1.
In-house machine learning and deep learning: molecular descriptors & GNNs, mixture models, physics-informed residuals, inverse search (Bayesian / multi-objective). Trained by Lavoisier — not rented property APIs.
Chemistry LLMs are evaluated on formulation, process, and regulation tasks before they orchestrate agents — grounded answers, cited sources, and failure modes measured — not vibes.
Local LLM + RAG over REACH, RoHS, market rules, and customer policy docs. Compliance answers cite clauses; inference can stay private.
Studio Pro workflow tools: full-metric process simulation, thermodynamic packages (Ideal, NRTL, Peng-Robinson, UNIQUAC), Monte Carlo risk, Bayesian optimization, and digital twin profiles ready for scale-up.
Models of a live process, built from its own operating data on top of the physics that governs it — with a verdict on whether the data was ever enough to support the model. From one model we commission the controller the loop needs, classical or model predictive, into the control system already running it.
Agents answer across formulation, process, and plant context under chemistry constraints, incomplete data, and quantified uncertainty — citing the run or the clause behind each answer. Advisory, with humans in the loop.
A fleet of chemistry AI agents — plant advisory today; formulation, regulation, simulation, and scale-up roles expanding on one backbone.
FormulaIQ starts with real components drawn from open-source chemistry corpora, then runs ML/DL forward models and inverse search so recipes stay physically plausible. LLM evaluation gates what agents are allowed to say; regulation RAG constrains what they are allowed to recommend.
See how FormulaIQ uses this stack.
Explore FormulaIQFirst-principles baseline plus ML residual — confidence and out-of-domain flags on every answer.
Data flywheel
Validated experiments
25→200
Studio Pro runs the workflow: simulate the process, extract a digital twin profile, then commission the controller against the real plant. Agents read the same record and advise — plant advisory today, a growing fleet ahead.
Formula confidence
+7.8%80.5%
Twin stability
+4.5%94%
Model readiness
72
Eval score
Chemical Engineering Agents
Open components, ML/DL + LLM evaluation, digital twin workflows, and chemistry agents — the stack investors and PhDs both ask about.
Open components, ML/DL, LLM evaluation, RAG, digital twins, process control, and chemistry agents — in one technical session.