Scientific evidence

See what supports the result and where it can fail

Review what a model was built to do, how it was evaluated, where it applies, how uncertainty behaves, which cases remain unsupported, and what evidence a qualified reviewer should seek next.

Scientific evidence records, validation panels, calibration curves, and review states arranged across a white canvas.

Evidence paths

Inspect the claim from four directions

Scientific credibility, decision value, controlled examples, and enterprise operation require different evidence.

Validation

Methods and failure modes

Review intended use, leakage controls, performance, calibration, applicability, unsupported cases, limitations, and change history.

Model and data cards

Versioned capability facts

Inspect data sources, labels, counts, splits, methods, known gaps, versions, and owners.

Example reports

Controlled work product

Review approved synthetic, public-domain, or explicitly authorized outputs. Open examples

Security and deployment

Operating evidence

Review data flow, identity, change control, monitoring, support, and responsibility. Review deployment

Current small-molecule evidence

Start with the measured facts

The current model card reports real-assay and curated-outcome training data, scaffold-aware evaluation, endpoint-level calibration, and explicit reliability and applicability gates.

Current MLTox model card showing the version, training-record count, hazard and ADME model coverage, scaffold-split metrics, and endpoint performance
Current versioned model card. The displayed counts and validation metrics belong to this specific model build.
  • 58,063training records
  • 12hazard endpoints
  • 15ADME/PK models
  • 3direct dose readouts
  • 1lab feedback record in this model version
Leakage control

Scaffold-grouped evaluation

Bemis-Murcko scaffold groups keep close structural families from straddling training and validation folds.

Uncertainty

Calibration, conformal sets, and abstention

Isotonic calibration and split-conformal prediction sets accompany the 12 hazard classifiers. Out-of-domain estimates can be withheld.

Reliability

Weak endpoints do not silently dominate

Endpoints below 0.70 CV-AUC remain visible with their metric but are excluded from escalation of the overall hazard band.

ADME / PK

Every endpoint carries its own metric

Reported performance ranges from HIA AUC near 0.96 and LogS R² near 0.77 to directional half-life and clearance R² near 0.20.

Interpret these as model-development results, not universal performance guarantees. AUC, R², calibration, and coverage are endpoint- and dataset-specific; prospective performance depends on the chemistry and decision context presented to the model.

Evidence standard

Different claims require different proof

Model accuracy, usefulness in a real scientific decision, biologics design hypotheses, and enterprise operation are separate claims. Each needs evidence matched to the decision, the failure modes, and the people responsible for acting.

01

Predictive claim

Model performance

Locked datasets, predefined metrics, leakage controls, relevant splits, subgroup and domain analysis, and versioned methods.

02

Workflow claim

Decision value

Representative workflows, current-process comparators, useful disagreement, expert adjudication, unsupported cases, and prospective decision measures.

03

Design claim

Biologics design

Retrospective method evaluation plus relevant experimental follow-up for presentation, expression, function, binding, biophysics, developability, and immune evidence.

04

Operational claim

Enterprise operation

Architecture, data flow, identity, roles, responsibility, change control, monitoring, incident response, support, and rollback evidence.

Validation and release

Model evidence is implemented; release governance is the next layer

The repository documents intended use, scaffold-aware evaluation, calibration, applicability, and known limitations. Named approval, monitored promotion, and a formal rollback decision are additional controls to establish for governed enterprise deployment.

Five validation checkpoints separating intended use, leakage control, model characterization, decision value, and controlled release
Current scientific validation occupies the first three checkpoints. Decision studies and controlled release add the evidence needed for a specific enterprise use.
  1. 01

    Define intended use

    What input, user, stage, endpoint or liability, and decision define the use?

  2. 02

    Prevent information leakage

    How were identities, labels, related records, scaffolds, sequences, programs, and time handled?

  3. 03

    Characterize model behavior

    How does the method perform, calibrate, expose ambiguity, and identify unsupported cases?

  4. 04

    Measure decision value

    Does the result improve the real decision compared with the current process?

  5. 05

    Control the release

    Which version was evaluated, what changed, who approved it, and how can it be monitored or rolled back?

Honest limits

Limits belong beside the result

mltox distinguishes model evidence from biological interpretation and regulatory judgment. Lower-performing reproductive, developmental, neurotoxicity, and other endpoints remain visible with their measured metrics. Potency and difficult ADME regressions are approximate. Biologics outputs are sequence- and structure-derived hypotheses, not clinical ADA, measured biophysics, or proof that a suggested edit preserves function. No report establishes that a candidate is safe.