Validation
Methods and failure modes
Review intended use, leakage controls, performance, calibration, applicability, unsupported cases, limitations, and change history.
Evidence paths
Scientific credibility, decision value, controlled examples, and enterprise operation require different evidence.
Validation
Review intended use, leakage controls, performance, calibration, applicability, unsupported cases, limitations, and change history.
Model and data cards
Inspect data sources, labels, counts, splits, methods, known gaps, versions, and owners.
Example reports
Review approved synthetic, public-domain, or explicitly authorized outputs. Open examples
Security and deployment
Review data flow, identity, change control, monitoring, support, and responsibility. Review deployment
Current small-molecule evidence
The current model card reports real-assay and curated-outcome training data, scaffold-aware evaluation, endpoint-level calibration, and explicit reliability and applicability gates.
Bemis-Murcko scaffold groups keep close structural families from straddling training and validation folds.
Isotonic calibration and split-conformal prediction sets accompany the 12 hazard classifiers. Out-of-domain estimates can be withheld.
Endpoints below 0.70 CV-AUC remain visible with their metric but are excluded from escalation of the overall hazard band.
Reported performance ranges from HIA AUC near 0.96 and LogS R² near 0.77 to directional half-life and clearance R² near 0.20.
Interpret these as model-development results, not universal performance guarantees. AUC, R², calibration, and coverage are endpoint- and dataset-specific; prospective performance depends on the chemistry and decision context presented to the model.
Evidence standard
Model accuracy, usefulness in a real scientific decision, biologics design hypotheses, and enterprise operation are separate claims. Each needs evidence matched to the decision, the failure modes, and the people responsible for acting.
Predictive claim
Locked datasets, predefined metrics, leakage controls, relevant splits, subgroup and domain analysis, and versioned methods.
Workflow claim
Representative workflows, current-process comparators, useful disagreement, expert adjudication, unsupported cases, and prospective decision measures.
Design claim
Retrospective method evaluation plus relevant experimental follow-up for presentation, expression, function, binding, biophysics, developability, and immune evidence.
Operational claim
Architecture, data flow, identity, roles, responsibility, change control, monitoring, incident response, support, and rollback evidence.
Validation and release
The repository documents intended use, scaffold-aware evaluation, calibration, applicability, and known limitations. Named approval, monitored promotion, and a formal rollback decision are additional controls to establish for governed enterprise deployment.
What input, user, stage, endpoint or liability, and decision define the use?
How were identities, labels, related records, scaffolds, sequences, programs, and time handled?
How does the method perform, calibrate, expose ambiguity, and identify unsupported cases?
Does the result improve the real decision compared with the current process?
Which version was evaluated, what changed, who approved it, and how can it be monitored or rolled back?
Honest limits
mltox distinguishes model evidence from biological interpretation and regulatory judgment. Lower-performing reproductive, developmental, neurotoxicity, and other endpoints remain visible with their measured metrics. Potency and difficult ADME regressions are approximate. Biologics outputs are sequence- and structure-derived hypotheses, not clinical ADA, measured biophysics, or proof that a suggested edit preserves function. No report establishes that a candidate is safe.