Components from open-source chemistry
FormulaIQ Layer 1 is fed by a multi-source open materials ETL — not a single vendor dump. Components stay CAS-keyed into inverse formulation.
Formula · open data
Open chemistry sources
Technology stack
Formula components from open-source chemistry data, ML/DL models we train, LLM evaluation, regulation RAG, workflow simulation & digital twins, and Chemical Engineering Agents — a growing fleet for chemistry.
prediction = first-principles baseline + ML residual ± uncertainty
One inspectable stack. Each layer is purpose-built for chemistry — not a generic LLM wrapper.
Component-level properties and structures from multiple open chemistry databases — QM9, Materials Project, OQMD, AFLOW, NOMAD, PoLyInfo, Polymer Genome, PubChem, ZINC — normalised onto a CAS-keyed backbone for FormulaIQ Layer 1.
In-house machine learning and deep learning: molecular descriptors & GNNs, mixture models, physics-informed residuals, inverse search (Bayesian / multi-objective). Trained by Lavoisier — not rented property APIs.
Chemistry LLMs are evaluated on formulation, process, and regulation tasks before they orchestrate agents — grounded answers, cited sources, and failure modes measured — not vibes.
Local LLM + RAG over REACH, RoHS, market rules, and customer policy docs. Compliance answers cite clauses; inference can stay private.
Studio Pro workflow tools: full-metric process simulation, thermodynamic packages (Ideal, NRTL, Peng-Robinson, UNIQUAC), Monte Carlo risk, Bayesian optimization, and digital twin profiles ready for scale-up.
The same twin moves into operations. Agents keep plants on-spec under chemistry constraints, incomplete data, and quantified uncertainty — with humans in the loop, and more agent roles ahead.
A fleet of chemistry AI agents — plant ops today; formulation, regulation, simulation, and scale-up roles expanding on one backbone.
FormulaIQ starts with real components drawn from open-source chemistry corpora, then runs ML/DL forward models and inverse search so recipes stay physically plausible. LLM evaluation gates what agents are allowed to say; regulation RAG constrains what they are allowed to recommend.
See how FormulaIQ uses this stack.
Explore FormulaIQFirst-principles baseline plus ML residual — confidence and out-of-domain flags on every answer.
Data flywheel
Validated experiments
25→200
Studio Pro runs the workflow: simulate the process, extract a digital twin profile, then hand the same twin to Chemical Engineering Agents — plant ops today, a growing fleet tomorrow.
Formula confidence
+7.8%80.5%
Twin stability
+4.5%94%
Model readiness
72
Eval score
Chemical Engineering Agents
Open components, ML/DL + LLM evaluation, digital twin workflows, and chemistry agents — the stack investors and PhDs both ask about.
Open components, ML/DL, LLM evaluation, RAG, digital twins, process control, and chemistry agents — in one technical session.