Scientific evidence

Evidence you can inspect

Inspect measured performance, applicability, limits, tool fit, and example reports before choosing the next experiment.

Scientific report pages and evidence records arranged across a white canvas.

Evidence paths

Inspect the claim from four directions

Scientific credibility, decision value, controlled examples, and enterprise operation require different evidence.

Validation

Methods and failure modes

Review intended use, leakage controls, performance, calibration, applicability, unsupported cases, limitations, and change history.

Model and data cards

Versioned capability facts

Inspect data sources, labels, counts, splits, methods, known gaps, versions, and owners.

Example reports

Controlled work product

Review approved synthetic, public-domain, or explicitly authorized outputs. Open examples

Security and deployment

Operating evidence

Review data flow, identity, change control, monitoring, support, and responsibility. Review deployment

Current small-molecule evidence

Start with the measured facts

The current source package combines real-assay and curated-outcome training data, scaffold-aware evaluation, endpoint-level calibration, and explicit reliability and applicability gates.

Model card

v20260803181840-hyb102633

MLTox hazard + ADME/PK engine

mltox-engine · schema v2 · built Aug 3, 2026

  • 72,827Training records
  • 15Hazard endpoints
  • 15ADME / PK models
  • 4Dose readouts
  • 0Lab feedback

Hazard endpoints

Scaffold-split CV AUC · bar 0.5→1.0
  • Ocular toxicity0.90
  • CYP450 inhibition (DDI)0.87
  • Mutagenicity (Ames)0.85
  • Acute systemic toxicity0.83
  • Cardiotoxicity (hERG)0.81
  • Respiratory toxicity0.79
  • Endocrine disruption0.78
  • Hepatotoxicity0.76
  • Immunotoxicity0.73
  • Carcinogenicity0.72
  • Skin sensitization0.72
  • Neurotoxicity0.71
  • Developmental toxicity0.64
  • Reproductive toxicity0.61
  • Nephrotoxicity0.61

ADME / PK — classification

CV AUC · bar 0.5→1.0
  • Human intestinal absorptionAUC 0.97
  • P-gp inhibitionAUC 0.93
  • Blood-brain-barrier penetrationAUC 0.90
  • CYP2D6 substrateAUC 0.77
  • PAMPA permeabilityAUC 0.76
  • Oral bioavailability (F≥20%)AUC 0.69
  • CYP3A4 substrateAUC 0.67
  • CYP2C9 substrateAUC 0.66

ADME / PK — regression

CV R²
  • Aqueous solubilityR² 0.80
  • Caco-2 permeabilityR² 0.67
  • Lipophilicity (logD7.4)R² 0.65
  • Volume of distribution (VDss)R² 0.49
  • Plasma protein bindingR² 0.44
  • Hepatocyte clearanceR² 0.20
  • Half-lifeR² 0.19

Dose / potency

CV R²
  • Estrogen receptor AC50R² 0.56
  • hERG IC50R² 0.43
  • Acute oral LD50 (rat)R² 0.40
  • Systemic POD (in vivo)R² 0.28

The same panel the portal shows for the deployed bundle. Every figure is the model’s own scaffold-split cross-validation result on its training corpus, not a held-out public benchmark — for those, see the four TDC tasks below.

Current portal model card. Rendered from deployed bundle v20260803181840-hyb102633, built August 3, 2026.
Leakage control

Scaffold-grouped evaluation

Bemis-Murcko scaffold groups keep close structural families from straddling training and validation folds.

Uncertainty

Calibration, conformal sets, and abstention

Eligible hazard classifiers carry isotonic calibration and split-conformal prediction sets. Out-of-domain estimates can be withheld.

Reliability

Lower-confidence outputs keep their limits visible

Each endpoint retains its measured performance, caveats, maturity label, and applicability evidence. Exploratory nephrotoxicity is currently below the 0.70 development gate and should receive less interpretive weight. The retained immunotoxicity champion predates the outcome-anchored trusted-subset metric, so its displayed confidence also warrants caution.

ADME / PK

Every endpoint carries its own metric

Reported performance ranges from HIA AUC near 0.97 and aqueous-solubility R² near 0.80 to directional half-life R² near 0.19 and clearance R² near 0.20.

Model generations

A mixed champion and challenger bundle

Nine endpoints adopted the v2 schema; six kept their v1 champion because the per-endpoint gate refused v2 — three below the trusted-AUC floor, two on calibration, one on a sentinel check. Adopted v2 endpoints use 371 named features plus a 1,024-bit fingerprint.

Interpret these as model-development results, not universal performance guarantees. AUC, R², calibration, and coverage are endpoint- and dataset-specific; prospective performance depends on the chemistry and decision context presented to the model.

hERGROC-AUC · higher is betterTDC rank #1outright lead
0.889 ± 0.002
.790higher →.895
  1. MLTox.889 ± .002
  2. MapLight + GNN.880 ± .002
  3. CFA.875 ± .014
  4. SimGCN.874 ± .014
  5. MapLight.871 ± .004
  6. ZairaChem.856 ± .009
  7. MiniMol.846 ± .016
  8. RDKit2D + MLP.841 ± .020
  9. Chemprop-RDKit.840 ± .007

First on the board and separably ahead of MapLight + GNN (Welch p < .001) — the only task where we lead outright. Two caveats carry equal weight. The method is adopted, not ours: MapLight’s feature union and CatBoost recipe plus the Hu et al. pretrained GIN — together, that combination is the MapLight + GNN entry below us — with a 3D pharmacophore block added. And this test fold is burned: roughly 15 scored configurations went through its 132 compounds, so the ±.002 measures training reproducibility on a fixed public fold, not generalization. The Hanley–McNeil standard error at this AUC and fold size is closer to ±.028.

AmesROC-AUC · higher is betterTDC rank #1tied at the top
0.873 ± 0.003
.825higher →.882
  1. MLTox.873 ± .003
  2. ZairaChem.871 ± .002
  3. ADMETrix.870 ± .006
  4. MapLight + GNN.869 ± .002
  5. MapLight.868 ± .002
  6. CFA.852 ± .005
  7. Chemprop-RDKit.850 ± .004
  8. MiniMol.849 ± .004
  9. DeepMol (AutoML).847 ± .007

MLTox now posts the highest mean on the board, but the .002 gap to ZairaChem sits inside the noise (Welch p = .198) — a tie at the top that sorts first, not a win. ADMETrix is also inseparable (p = .359). MapLight + GNN (p = .016) and MapLight (p = .004) are now separable, which they were not before.

LD50MAE · lower is betterTDC rank #2tied at the top
0.558 ± 0.007
.530← lower.675
  1. BaseBoosting.552 ± .009
  2. MLTox.558 ± .007
  3. ADMETrix.573 ± .010
  4. MiniMol.585 ± .008
  5. MACCS keys + autoML.588 ± .005
  6. Chemprop.606 ± .024
  7. DeepMol (AutoML).614 ± .004
  8. MapLight.621 ± .003
  9. QuGIN.622 ± .015

Lower is better. Matching the training loss to the scored metric and pretraining on external rat oral LD50 data cut the error from .591 to .558 — fifth to second. BaseBoosting’s .006 edge is inside the noise (p = .256), so this is a tie at the top; every method below is separable.

DILIROC-AUC · higher is betterTDC rank #2leader ahead
0.931 ± 0.006
.860higher →.980
  1. MiniMol.956 ± .006
  2. MLTox.931 ± .006
  3. ZairaChem.925 ± .005
  4. AttrMasking.919 ± .008
  5. CFA.919 ± .014
  6. MapLight + GNN.917 ± .005
  7. SimGCN.909 ± .011
  8. ADMETrix.906 ± .016
  9. Chemprop.899 ± .008

A six-member blend and a v2 GBM arm lifted the mean from .924 to .931, moving MLTox to second. MiniMol keeps a real, separable lead (p < .001). ZairaChem (p = .067) and CFA (p = .129) now sit below us but inside the noise; the rest of the board is separable.

MLTox meanPublished-method mean± 1 SDMLTox interval Five-seed mean ± standard deviation on the official TDC held-out split. Each task uses its own scale. Public leaderboard checked 30 July 2026.

Benchmark evidence

Four public endpoint benchmarks

Each chart shows MLTox against the eight leading public methods for an official TDC task. Each task was rebuilt from scratch on its train fold; these are not scores from the deployed artifact. The hERG entry adopts a published method rather than an MLTox architecture — its note sets out what is and is not ours.

  • 01Five seeded runs report variability instead of a single favorable run.
  • 02MLTox reached .873 Ames, .931 DILI, .889 hERG, and .558 LD50 MAE.
  • 03Whiskers show one standard deviation. ROC-AUC is higher-is-better; MAE is lower-is-better.

Evidence standard

Different claims require different proof

Model accuracy, usefulness in a real scientific decision, biologics design hypotheses, and enterprise operation are separate claims. Each needs evidence matched to the decision, the failure modes, and the people responsible for acting.

01

Predictive claim

Model performance

Locked datasets, predefined metrics, leakage controls, relevant splits, subgroup and domain analysis, and versioned methods.

02

Workflow claim

Decision value

Representative workflows, current-process comparators, useful disagreement, expert adjudication, unsupported cases, and prospective decision measures.

03

Design claim

Biologics design

Retrospective method evaluation plus relevant experimental follow-up for presentation, expression, function, binding, biophysics, developability, and immune evidence.

04

Operational claim

Enterprise operation

Architecture, data flow, identity, roles, responsibility, change control, monitoring, incident response, support, and rollback evidence.

Validation and release

Model evidence is implemented; release governance is the next layer

The repository documents intended use, scaffold-aware evaluation, calibration, applicability, and known limitations. Named approval, monitored promotion, and a formal rollback decision are additional controls to establish for governed enterprise deployment.

Five validation checkpoints separating intended use, leakage control, model characterization, decision value, and controlled release
Current scientific validation occupies the first three checkpoints. Decision studies and controlled release add the evidence needed for a specific enterprise use.
  1. 01

    Define intended use

    What input, user, stage, endpoint or liability, and decision define the use?

  2. 02

    Prevent information leakage

    How were identities, labels, related records, scaffolds, sequences, programs, and time handled?

  3. 03

    Characterize model behavior

    How does the method perform, calibrate, expose ambiguity, and identify unsupported cases?

  4. 04

    Measure decision value

    Does the result improve the real decision compared with the current process?

  5. 05

    Control the release

    Which version was evaluated, what changed, who approved it, and how can it be monitored or rolled back?

Honest limits

Limits belong beside the result

MLTox distinguishes model evidence from biological interpretation and regulatory judgment. Lower-performing reproductive, developmental, carcinogenicity, and other endpoints remain visible with their measured metrics. Potency and difficult ADME regressions are approximate. Biologics outputs are sequence- and structure-derived hypotheses, not clinical ADA, measured biophysics, or proof that a suggested edit preserves function. No report establishes that a candidate is safe.

Capability matrix

A practical view of the current field

This is a workflow-level comparison, not a claim that every capability is equivalent in depth, validation, or intended use.

Product Small-molecule
safety
Biologics
engineering
ADME/PK and
physicochemistry
Visible uncertainty
and applicability
Read-across or
reference context
Regulatory-style
documentation
Self-hosted or
on-premises
Batch and
programmatic access
mltox Full Full Full Full Full Full Full Full
ADMETlab 3.0
ProTox 3.0
ADMET-AI
Toxometris.ai
Optibrium StarDrop
ADMET Predictor
Schrodinger LiveDesign
Lhasa Derek Nexus
EpiVax ISPRI
OPIG SAbPred / TAP

Legend: full means the capability is a stated product function; partial means limited, indirect, module-dependent, or not presented with the same workflow depth; none means it was not identified in the reviewed public product material.

Choose the evidence system that fits the decision

Use these practical guides to identify when MLTox fits, when another category fits better, and when both belong in the workflow.

Broad ADMET suitesWhen you need breadth across physicochemistry, disposition, and early risk
Choose MLTox when

The review centers on endpoint-specific toxicology evidence, visible uncertainty, applicability, analogues, and an inspectable decision record.

Choose the suite when

Your primary need is broad property coverage, mature enterprise workflows, or a familiar all-purpose discovery environment.

Use both when

The suite supports broad optimization while MLTox provides a focused safety review before the next compound decision.

Free or internal modelsWhen flexibility, code access, or a low-cost first screen matters most
Choose MLTox when

Your team needs maintained models, calibrated outputs, applicability evidence, analogue context, exports, and a consistent reviewer-facing workflow.

Choose free or internal tools when

You have the expertise to validate, maintain, document, and operate the models for your chemistry and intended use.

Use both when

Internal models remain the custom baseline and MLTox supplies an independent, structured safety evidence layer.

Regulatory specialistsWhen expert rule systems and submission-oriented precedent are central
Choose MLTox when

The immediate goal is preclinical prioritization across multiple endpoints with probability, uncertainty, applicability, and read-across context.

Choose the specialist when

You need a long-established expert system, a specific regulatory workflow, or human expert interpretation tied to submission practice.

Use both when

MLTox informs early portfolio decisions and the specialist supports later expert or regulatory review.

Molecular-design platformsWhen multiparameter design and team-wide chemistry collaboration lead
Choose MLTox when

You need a focused toxicology evidence record and experimental next-step guidance without adopting a broader design environment.

Choose the platform when

Your priority is collaborative design, enumeration, simulation, or multiparameter optimization across an integrated discovery stack.

Use both when

The design platform manages the series and MLTox adds a dedicated safety decision checkpoint.

Biologics specialistsWhen protein modality depth matters more than a shared product surface
Choose MLTox when

You want proteins and biologics screening alongside small-molecule safety under one operating and deployment model.

Choose the specialist when

You need deeper modality-specific capabilities, services, or experimentally anchored expertise beyond the current MLTox product boundary.

Use both when

MLTox supplies the repeatable early screen and a specialist addresses the highest-value modality questions.

Example outputs

See the result before the demo

Review two synthetic teaching examples before sharing proprietary data. They show the scientific-review structure and separate a model result, its supporting evidence, and the next experimental action.

Interactive report

Portal evidence view

Hazard, potency, ADME/PK, physicochemical rules, nearest analogues, approved-drug percentiles, and a grounded report assistant.

Portable report

PDF and JSON

A readable report for scientific circulation and structured output for programmatic use or downstream transformation.

Prediction record

OECD-format QPRF

A per-molecule or per-endpoint prediction record with applicability-domain context, available as JSON, PDF, or editable DOCX.

Model record

OECD-format QMRF

An endpoint-level model record covering the OECD validation principles, available as JSON, PDF, or editable DOCX.

Small-molecule demo report
Page 1 / -- Open full small-molecule PDF

Loading report preview

Small-molecule example

A compound series with endpoint evidence and an explicit next action

Click through a ten-page synthetic concept report that keeps compounds comparable, opens endpoint evidence, separates supported, ambiguous, and unsupported cases, and records a proposed next action. It predates the complete 15-endpoint ADME/PK panel, approved-drug percentile radar, surfaced read-across, and QMRF/QPRF controls.

  • Controlled structure identity and example provenance.
  • Legacy 13-endpoint concept layout; the current product exposes 15 visible hazard outputs.
  • Call, probability, prediction set, applicability, and local evidence.
  • Reviewer rationale and next evidence request.
Biologics demo report
Page 1 / -- Open full biologics PDF

Loading report preview

Biologics example

A candidate comparison with liabilities, constraints, and a laboratory handoff

Click through a ten-page synthetic candidate-review concept. It illustrates how immunogenicity, intrinsic developability, optional predicted structure, constrained design ideas, and an experimental handoff should remain distinct. It is not a validated clinical ADA or measured biophysics report.

  • Sequence identity, construct context, and version.
  • HLA assumptions, residue hotspots, developability, and structural exposure.
  • Protected regions, proposed variants, and explicit tradeoffs.
  • Candidate-specific experimental handoff.

Decision proof

The comparison that matters happens on your molecules

A useful evaluation uses representative chemistry or sequences, agreed evidence, and a decision made before the next program milestone.

  1. Representative inputs

    Use the molecular classes, novelty, edge cases, and file formats your team actually handles.

  2. Declared outputs

    Agree which endpoints, confidence states, reports, and exports must be usable.

  3. Useful disagreements

    Inspect where methods differ, why they differ, and what experiment would resolve the uncertainty.

  4. Operating fit

    Test batch flow, API integration, deployment, identity, retention, versioning, and scientific review.

Compare MLTox on one real decision

Bring a representative molecule or sequence, the current workflow, and the evidence your team needs to act.