In-silico safety screening

Predictive safety for
molecules & biologics.

Submit a small molecule or a protein and get a calibrated safety report in seconds — QSAR toxicology across twelve endpoints, or learned immunogenicity and developability screening for biologics. The models fine-tune themselves as your clinical and lab results come back.

  • 12tox endpoints
  • 44kreal-assay compounds
  • 90%conformal coverage

Research tool. Every output is an in-silico hypothesis on an unvalidated model — a prioritization aid, not a regulatory or clinical determination.

Built on established science
  • Real-assay data: TOXRIC · TDC · ChEMBL · EPA ToxRefDB
  • RDKit featurization
  • Bemis–Murcko scaffold split
  • IEDB-trained MHC models
  • CamSol-style solubility
  • Isotonic + conformal calibration
How it works

One input. The right analysis. A loop that improves.

mltox auto-detects what you submit and routes it to the appropriate model — a protein is never scored by a small-molecule QSAR. Real-world outcomes feed back in.

  1. 01

    Submit

    Paste a SMILES, InChI, MOL/SDF, or a FASTA/amino-acid sequence. No account, no setup.

  2. 02

    Detect & featurize

    Format is auto-detected. Small molecules are featurized (RDKit or a pure-Python fallback); biologics get a sequence-liability pipeline.

  3. 03

    Model predicts

    Per-endpoint gradient-boosted classifiers, or learned per-allele epitope models — each with calibrated confidence and an applicability-domain gate.

  4. 04

    Report & feedback

    A structured report with rationale. Submit clinical/lab outcomes; weighted retraining folds them back into the models.

SELF-HOSTED · ISOLATED · YOUR ENVIRONMENT PREDICTION LEARNING LOOP YOUR EXTERNAL DATA REFINES THE ENDPOINT MODELS YOU RUN THE EXPERIMENT Input SMILES · MOL · FASTA Detect & featurize format router Models QSAR · biologic analysis Safety report calibrated · conformal Weighted retrain seed + all feedback Feedback store your local DB · source-weighted Clinical & lab outcomes trial · in-vivo · assay SELF-HOSTED · ISOLATED PREDICTION Input SMILES · MOL · FASTA Detect & featurize format router Models QSAR · biologic analysis Safety report calibrated · conformal YOU RUN THE EXPERIMENT LEARNING LOOP Clinical & lab outcomes trial · in-vivo · assay Feedback store your local DB · source-weighted Weighted retrain seed + all feedback REFINES THE ENDPOINT MODELS

How the loop closes. Every report is a prediction. When a real result comes back — a clinical readout, an in-vivo study, a bench assay — you submit it as feedback. Each outcome is stored with a trust weight based on its source (a clinical trial counts for more than a literature note), then folded into a retrain over the seed data plus all accumulated feedback. That updated model scores your next molecule. The more real-world data you feed it, the more it is tuned to your chemistry.

Everything happens inside your boundary. mltox is self-hosted: the models, the feedback store, and retraining all run on infrastructure you control (a container or host you own). Scoring is fully local — RDKit + scikit-learn, and ESMFold on your own GPU — so a submitted molecule, sequence, or clinical outcome is never sent to a third-party service to be scored. Feedback lives in your own local database and is used only to retrain your own models; nothing is pooled across tenants.

Scope note: the fine-tune loop updates the small-molecule endpoint models. The biologics epitope models are trained separately on curated IEDB data and are not retrained on submitted outcomes. The only outbound fetch anywhere is the one-time download of the public ESMFold model weights — model data in, never your data out.

What it screens

Two deliberately distinct engines

The small-molecule models don't apply to a 300-residue protein — so proteins get a real sequence-based analysis instead of a fake QSAR score.

Small molecules

SMILES · InChI · MOL/SDF → QSAR toxicology report

  • Twelve hazard endpoints with calibrated probability
  • Structural-alert (toxicophore) matches with rationale
  • 90% split-conformal prediction sets
  • Applicability-domain novelty gate + 2D depiction
  • Scaffold-split evaluation (no analog leakage)

Proteins & biologics

FASTA · sequence → immunogenicity + developability

  • Per-allele MHC-II T-cell epitope models (IEDB-trained)
  • Humanness axis separates self-tolerated from non-human risk
  • Deamidation, isomerization, oxidation & glycosylation liabilities
  • CamSol-style continuous solubility profile
  • GPU structure prediction with a 3D viewer — surface-aware re-ranking
  • Designs, not just scores — de-immunization candidates
Coverage

Twelve toxicology endpoints

Each is an independent binary hazard model, trained on real assays and evaluated by scaffold-split CV (0.58–0.95 AUC by endpoint). Severity weights how much each drives the overall risk band.

Carcinogenicitytumor formation on chronic exposure
Acute systemic toxicitylow LD50 after single exposure
Cardiotoxicity (hERG)QT prolongation / arrhythmia
Developmental toxicityteratogenicity
Reproductive toxicityfertility / reproductive harm
Hepatotoxicitydrug-induced liver injury
Mutagenicity (Ames)bacterial reverse mutation
NeurotoxicityCNS / peripheral effects
Endocrine disruptionnuclear-receptor hormone pathways
Skin sensitizationallergic contact dermatitis
Ocular toxicityeye irritation / corrosion
CYP450 inhibition (DDI)drug–drug-interaction liability
Biologics engineering

It doesn't just score — it proposes redesigns

A suite of sequence-aware capabilities for antibody and protein engineers. Each is an in-silico hypothesis to prioritize wet-lab work, never a verified fix.

Junctional-neoepitope scanning

In fusions and multispecifics, flags T-cell epitopes created at the seam — present in neither parent domain — across auto-detected linkers.

De-immunization suggester

Proposes conservative BLOSUM62 substitutions the model predicts remove an epitope — anchors first, Cys/Pro/glyco fixed, with a greedy best-pair fallback.

Antibody CDR-awareness

Detects VH/VL, annotates CDR1/2/3 vs framework, and tags a risk as a redesign target (CDR) or a humanization gap (framework).

Expanded HLA panels

Selectable epitope panels — HLA-DR, DR/DQ/DP, and a supplementary class-I CD8 panel — each per-allele CV-AUC gated (0.80–0.93).

CamSol-style solubility

A continuous per-residue intrinsic-solubility profile from hydrophobicity, charge and β-propensity — flags aggregation-prone regions.

Calibration + conformal

Every endpoint is isotonic-calibrated and carries a 90%-coverage conformal set — a two-way {clean, hazard} means “can’t distinguish.”

GPU structure prediction

Folds the sequence with ESMFold on GPU, then uses per-residue solvent exposure to down-weight buried liabilities and up-weight exposed ones. Runs as a queued job (queued → running → done, with your place in line) so long folds never time out — results are cached and viewable in 3D.

Long multi-domain chains

ESMFold’s memory grows with the square of length, so a full-size protein won’t fit on one pass. Long chains are cut at flexible linker-like regions — which tend to fall between domains — and each segment is folded on its own, then recombined into one full-length exposure profile. The split is a heuristic, and how the segments sit relative to each other is not predicted.

Plain-language rationale

Every report explains why — the drivers behind the overall band and the developability call, in language a reviewer can act on.

Honest by design

Calibrated confidence beats a confident guess.

A safety tool that overstates certainty is worse than none. mltox is built to tell you when it doesn’t know — and to keep its numbers honest about novel chemistry.

Try it on your molecule →
  • Scaffold-split evaluation

    AUCs are measured across Bemis–Murcko scaffold groups on real assays, so the numbers reflect novel chemistry — not analog leakage from the training set.

  • Reliability gate on weak endpoints

    Endpoints below 0.70 CV-AUC (the hard, data-scarce in-vivo ones) are shown with their score but held out of the overall-risk escalation, so a weak model can’t inflate a verdict.

  • Applicability-domain gate

    Structures unlike anything seen in training are flagged low-confidence instead of extrapolating a confident-looking number.

  • Conformal prediction sets

    90%-coverage sets quantify ambiguity: a set containing both labels is an explicit “can’t distinguish,” not a silent coin flip.

  • Unvalidated-hypothesis framing

    Every biologics output is labeled an in-silico hypothesis on an unvalidated model — a capability to prioritize experiments, not a verified answer.

Screen your first molecule in under a minute.

No signup. Paste a structure or a sequence and read the report.