World's first foundational model for toxicology

Get early signalsabout your compoundGet early signalsabout your compound

Review endpoint-specific hazards, exposure, uncertainty, and analogue evidence before deciding what to make, test, advance, or stop.

A molecular landscape connects early safety evidence to an inspectable scientific decision workflow.
15visible hazard outputs14 standard plus exploratory nephrotoxicity
72,827committed rowsprovenance-aware training records
48,480canonical structuresunique chemistry after normalization
4public benchmark tasksofficial TDC splits, five seeded runs

Small-molecule evidence at compound and series scale

Review individual compounds and related chemical series while keeping the endpoint evidence behind each pattern visible.

MLTox Small Molecule Safety

Decide what to make, test, redesign, advance, or stop

Review 15 visible hazard outputs, direct and tiered potency estimates, 15 ADME/PK models, physicochemistry, read-across, approved-drug context, and OECD-format model and prediction records in one report. Exploratory nephrotoxicity remains clearly labeled.

Explore Small Molecule Safety
Current MLTox small-molecule portal showing structure input, recent analyses, overall risk, canonical structure, applicability-domain status, and model version
Portal workflow example from an earlier build. A submitted structure stays connected to overall risk, molecular identity, applicability-domain status, and the exact model version. Current verified counts appear above.

SeriesLens workflow

Compare chemical series before the field narrows

SeriesLens is part of MLTox Small Molecule Safety. Compare predicted liabilities across user-defined series, then inspect the compounds and local changes that drive the pattern.

Explore the SeriesLens workflow
Four related chemical series pass through shared evidence filters and narrow to representative analogues
Concept visualization. SeriesLens compares patterns across related chemistry; it does not label a series safe.

World's first foundational model built for endpoint toxicity

Fifteen visible hazard outputs do not come from one universal model. Each endpoint is independently trained, selected, and evaluated against its own data, performance metrics, reliability evidence, and applicability evidence. Exploratory nephrotoxicity and the retained immunotoxicity champion keep their lower-confidence caveats visible.

  • One fine-tuned head per endpoint

    Classification, calibration, and model selection are performed for the biological question and data available at each endpoint.

  • One visible metric

    Measured endpoint performance stays beside the result. Lower-confidence outputs retain their metrics, caveats, and maturity labels for scientific interpretation.

  • One applicability boundary

    A structure outside the modeled chemical neighborhood is flagged as low-confidence instead of receiving a polished extrapolation.

  • One uncertainty record

    A 90% prediction set can contain both labels. That state is presented as unresolved evidence, not compressed into a false binary answer.

  • Experimental confirmation

    High-value predictions can be routed into appropriate assays. Returned outcomes remain connected to their source; eligible small-molecule endpoint models may use them in a controlled learning loop.

  • Grounded report assistance

    A separate language-model assistant is grounded in a report-derived fact sheet, instructed to label general background, and to say when candidate-specific evidence is absent. It does not compute the underlying predictions.

Held-out benchmark evidence

State-of-the-art model with highest aggregate score

Each task was rebuilt from scratch on its official Therapeutics Data Commons train fold and evaluated against public leaderboard methods.

Explore methods and full comparisons
hERGROC-AUC · higher is betterTDC rank #1outright lead
0.889 ± 0.002

First on the board. Built on MapLight + GNN, not an MLTox architecture — see Evidence for the caveats.

AmesROC-AUC · higher is betterTDC rank #1tied at the top
0.873 ± 0.003

Highest mean on the board — a statistical tie, not a win.

LD50MAE · lower is betterTDC rank #2tied at the top
0.558 ± 0.007

Second on the board — the leader’s edge is inside the noise.

DILIROC-AUC · higher is betterTDC rank #2leader ahead
0.931 ± 0.006

Second on the board; MiniMol keeps a separable lead.

Mean ± standard deviation across five seeds. Each task uses its own scale; do not compare interval lengths across cards.

Illustrative compound series routed to a selective assay or expert review based on supported, ambiguous, and unsupported evidence states
Illustrative workflow. Pilot conclusions remain tied to your compounds, decisions, and agreed experimental evidence.

A bounded first step

Run one blinded Safety Decision Pilot

Start with a fixed compound set and a decision your team already understands. Agree on success before MLTox reviews the blinded outcomes.

Input
A representative compound series, known outcomes held back, and the decision context.
Success
Pre-agreed scientific and workflow criteria, including where uncertainty must be visible.
Decision
Adopt, refine, or stop based on the unblinded review. No open-ended platform commitment.
Plan a pilot

Separate hazard from potency and drug disposition

A hazard flag does not say at what dose an effect may occur or whether a compound has a workable ADME profile. The report keeps those questions separate.

01 15 visible hazard outputs

Could the structure carry an intrinsic toxicity signal?

Hazard

Calibrated endpoint probabilities, conformal sets, measured AUC, structural alerts, and interpretable drivers.

02 3 direct + 2 tiered

At roughly what concentration or dose might the signal matter?

Potency

LD50, hERG IC50, ER AC50 as direct estimates, plus a tiered systemic POD and a derived cardiac IVIVE margin, with broad ranges where precision is limited.

03 15 ADME/PK models

Can the compound reach, persist, and clear in a workable way?

Drug disposition

Absorption, permeability, distribution, binding, clearance, half-life, and CYP substrate predictions, each with its own metric and domain gate.

One evidence layer supports all three judgments.

Nearest analogues, approved-drug percentiles, physicochemical rules, model and domain evidence, PDF/JSON, and OECD-format QMRF/QPRF records stay attached to the result.

Interpretation boundary: potency and ADME outputs are structure-based estimates for prioritization. They do not establish a safe human dose, exposure profile, or regulatory conclusion.

MLTox Proteins & Biologics

A separate workflow for sequence decisions

Protein and biologics work uses different evidence from small-molecule safety. Review selectable HLA panels, humanness, junctional neoepitopes, sequence liabilities, intrinsic solubility, and conservative de-immunization candidates.

Explore Proteins & Biologics
Illustrative biologics workspace with aligned sequences, hotspot patterns, liability summaries, protein surfaces, and experimental plans
Prototype artwork with illustrative data.

A decision, not a demo

Test MLTox on compounds you already understand

Define one blinded comparison, make the success criteria explicit, and decide whether the evidence improves your next compound review.

Plan a pilot