Validation
Methods and failure modes
Review intended use, leakage controls, performance, calibration, applicability, unsupported cases, limitations, and change history.
Evidence paths
Scientific credibility, decision value, controlled examples, and enterprise operation require different evidence.
Validation
Review intended use, leakage controls, performance, calibration, applicability, unsupported cases, limitations, and change history.
Model and data cards
Inspect data sources, labels, counts, splits, methods, known gaps, versions, and owners.
Example reports
Review approved synthetic, public-domain, or explicitly authorized outputs. Open examples
Security and deployment
Review data flow, identity, change control, monitoring, support, and responsibility. Review deployment
Current small-molecule evidence
The current source package combines real-assay and curated-outcome training data, scaffold-aware evaluation, endpoint-level calibration, and explicit reliability and applicability gates.
MLTox hazard + ADME/PK engine
The same panel the portal shows for the deployed bundle. Every figure is the model’s own scaffold-split cross-validation result on its training corpus, not a held-out public benchmark — for those, see the four TDC tasks below.
v20260803181840-hyb102633, built August 3, 2026.Bemis-Murcko scaffold groups keep close structural families from straddling training and validation folds.
Eligible hazard classifiers carry isotonic calibration and split-conformal prediction sets. Out-of-domain estimates can be withheld.
Each endpoint retains its measured performance, caveats, maturity label, and applicability evidence. Exploratory nephrotoxicity is currently below the 0.70 development gate and should receive less interpretive weight. The retained immunotoxicity champion predates the outcome-anchored trusted-subset metric, so its displayed confidence also warrants caution.
Reported performance ranges from HIA AUC near 0.97 and aqueous-solubility R² near 0.80 to directional half-life R² near 0.19 and clearance R² near 0.20.
Nine endpoints adopted the v2 schema; six kept their v1 champion because the per-endpoint gate refused v2 — three below the trusted-AUC floor, two on calibration, one on a sentinel check. Adopted v2 endpoints use 371 named features plus a 1,024-bit fingerprint.
Interpret these as model-development results, not universal performance guarantees. AUC, R², calibration, and coverage are endpoint- and dataset-specific; prospective performance depends on the chemistry and decision context presented to the model.
First on the board and separably ahead of MapLight + GNN (Welch p < .001) — the only task where we lead outright. Two caveats carry equal weight. The method is adopted, not ours: MapLight’s feature union and CatBoost recipe plus the Hu et al. pretrained GIN — together, that combination is the MapLight + GNN entry below us — with a 3D pharmacophore block added. And this test fold is burned: roughly 15 scored configurations went through its 132 compounds, so the ±.002 measures training reproducibility on a fixed public fold, not generalization. The Hanley–McNeil standard error at this AUC and fold size is closer to ±.028.
MLTox now posts the highest mean on the board, but the .002 gap to ZairaChem sits inside the noise (Welch p = .198) — a tie at the top that sorts first, not a win. ADMETrix is also inseparable (p = .359). MapLight + GNN (p = .016) and MapLight (p = .004) are now separable, which they were not before.
Lower is better. Matching the training loss to the scored metric and pretraining on external rat oral LD50 data cut the error from .591 to .558 — fifth to second. BaseBoosting’s .006 edge is inside the noise (p = .256), so this is a tie at the top; every method below is separable.
A six-member blend and a v2 GBM arm lifted the mean from .924 to .931, moving MLTox to second. MiniMol keeps a real, separable lead (p < .001). ZairaChem (p = .067) and CFA (p = .129) now sit below us but inside the noise; the rest of the board is separable.
Benchmark evidence
Each chart shows MLTox against the eight leading public methods for an official TDC task. Each task was rebuilt from scratch on its train fold; these are not scores from the deployed artifact. The hERG entry adopts a published method rather than an MLTox architecture — its note sets out what is and is not ours.
Evidence standard
Model accuracy, usefulness in a real scientific decision, biologics design hypotheses, and enterprise operation are separate claims. Each needs evidence matched to the decision, the failure modes, and the people responsible for acting.
Predictive claim
Locked datasets, predefined metrics, leakage controls, relevant splits, subgroup and domain analysis, and versioned methods.
Workflow claim
Representative workflows, current-process comparators, useful disagreement, expert adjudication, unsupported cases, and prospective decision measures.
Design claim
Retrospective method evaluation plus relevant experimental follow-up for presentation, expression, function, binding, biophysics, developability, and immune evidence.
Operational claim
Architecture, data flow, identity, roles, responsibility, change control, monitoring, incident response, support, and rollback evidence.
Validation and release
The repository documents intended use, scaffold-aware evaluation, calibration, applicability, and known limitations. Named approval, monitored promotion, and a formal rollback decision are additional controls to establish for governed enterprise deployment.
What input, user, stage, endpoint or liability, and decision define the use?
How were identities, labels, related records, scaffolds, sequences, programs, and time handled?
How does the method perform, calibrate, expose ambiguity, and identify unsupported cases?
Does the result improve the real decision compared with the current process?
Which version was evaluated, what changed, who approved it, and how can it be monitored or rolled back?
Honest limits
MLTox distinguishes model evidence from biological interpretation and regulatory judgment. Lower-performing reproductive, developmental, carcinogenicity, and other endpoints remain visible with their measured metrics. Potency and difficult ADME regressions are approximate. Biologics outputs are sequence- and structure-derived hypotheses, not clinical ADA, measured biophysics, or proof that a suggested edit preserves function. No report establishes that a candidate is safe.
Capability matrix
This is a workflow-level comparison, not a claim that every capability is equivalent in depth, validation, or intended use.
| Product | Small-molecule safety |
Biologics engineering |
ADME/PK and physicochemistry |
Visible uncertainty and applicability |
Read-across or reference context |
Regulatory-style documentation |
Self-hosted or on-premises |
Batch and programmatic access |
|---|---|---|---|---|---|---|---|---|
| mltox | Full | Full | Full | Full | Full | Full | Full | Full |
| ADMETlab 3.0 | ||||||||
| ProTox 3.0 | ||||||||
| ADMET-AI | ||||||||
| Toxometris.ai | ||||||||
| Optibrium StarDrop | ||||||||
| ADMET Predictor | ||||||||
| Schrodinger LiveDesign | ||||||||
| Lhasa Derek Nexus | ||||||||
| EpiVax ISPRI | ||||||||
| OPIG SAbPred / TAP |
Legend: full means the capability is a stated product function; partial means limited, indirect, module-dependent, or not presented with the same workflow depth; none means it was not identified in the reviewed public product material.
Use these practical guides to identify when MLTox fits, when another category fits better, and when both belong in the workflow.
The review centers on endpoint-specific toxicology evidence, visible uncertainty, applicability, analogues, and an inspectable decision record.
Your primary need is broad property coverage, mature enterprise workflows, or a familiar all-purpose discovery environment.
The suite supports broad optimization while MLTox provides a focused safety review before the next compound decision.
Your team needs maintained models, calibrated outputs, applicability evidence, analogue context, exports, and a consistent reviewer-facing workflow.
You have the expertise to validate, maintain, document, and operate the models for your chemistry and intended use.
Internal models remain the custom baseline and MLTox supplies an independent, structured safety evidence layer.
The immediate goal is preclinical prioritization across multiple endpoints with probability, uncertainty, applicability, and read-across context.
You need a long-established expert system, a specific regulatory workflow, or human expert interpretation tied to submission practice.
MLTox informs early portfolio decisions and the specialist supports later expert or regulatory review.
You need a focused toxicology evidence record and experimental next-step guidance without adopting a broader design environment.
Your priority is collaborative design, enumeration, simulation, or multiparameter optimization across an integrated discovery stack.
The design platform manages the series and MLTox adds a dedicated safety decision checkpoint.
You want proteins and biologics screening alongside small-molecule safety under one operating and deployment model.
You need deeper modality-specific capabilities, services, or experimentally anchored expertise beyond the current MLTox product boundary.
MLTox supplies the repeatable early screen and a specialist addresses the highest-value modality questions.
Example outputs
Review two synthetic teaching examples before sharing proprietary data. They show the scientific-review structure and separate a model result, its supporting evidence, and the next experimental action.
Hazard, potency, ADME/PK, physicochemical rules, nearest analogues, approved-drug percentiles, and a grounded report assistant.
A readable report for scientific circulation and structured output for programmatic use or downstream transformation.
A per-molecule or per-endpoint prediction record with applicability-domain context, available as JSON, PDF, or editable DOCX.
An endpoint-level model record covering the OECD validation principles, available as JSON, PDF, or editable DOCX.
Loading report preview
Small-molecule example
Click through a ten-page synthetic concept report that keeps compounds comparable, opens endpoint evidence, separates supported, ambiguous, and unsupported cases, and records a proposed next action. It predates the complete 15-endpoint ADME/PK panel, approved-drug percentile radar, surfaced read-across, and QMRF/QPRF controls.
Loading report preview
Biologics example
Click through a ten-page synthetic candidate-review concept. It illustrates how immunogenicity, intrinsic developability, optional predicted structure, constrained design ideas, and an experimental handoff should remain distinct. It is not a validated clinical ADA or measured biophysics report.
Decision proof
A useful evaluation uses representative chemistry or sequences, agreed evidence, and a decision made before the next program milestone.
Use the molecular classes, novelty, edge cases, and file formats your team actually handles.
Agree which endpoints, confidence states, reports, and exports must be usable.
Inspect where methods differ, why they differ, and what experiment would resolve the uncertainty.
Test batch flow, API integration, deployment, identity, retention, versioning, and scientific review.
Bring a representative molecule or sequence, the current workflow, and the evidence your team needs to act.