Share
Last year, insitro and Eli Lilly and Company (Lilly) teamed up to build AI models that predict in vivo properties of small molecules, powered by the corpus of proprietary data that Lilly has accumulated over decades of R&D. Together, we built the in vivo PK model available in Lilly TuneLab as of October 6, to predict in vivo pharmacokinetics (PK) across four preclinical species: mouse, rat, dog, and monkey.
Designing a clinic-ready small molecule drug requires optimizing the pharmacokinetic profile of that molecule—ensuring the drug reaches the right tissues at appropriate concentrations for the required length of time to have a therapeutic effect. Existing ML models that predict the PK profile of small molecules have focused on predicting in vitro ADME values due to limited in vivo PK data availability [1]. Lilly’s corpus of historical in vitro ADME and in vivo PK data lets us go one step further: training models to directly predict the in vivo PK behavior of a molecule. Our models are designed to outperform today’s industry standard method based on extrapolation from in vitro assays, potentially enabling prioritization of compounds entering expensive and lengthy animal studies.
The gap between a dish and an animal
Medicinal chemists designing a safe and effective small molecule have to cross a critical translation gap between what happens in a dish (in vitro) and what happens in the body (in vivo). The first critical checkpoint on that path is a preclinical PK study. Animal models have translation challenges, but they are where a candidate molecule has to survive the complexity of an intact physiological system capable of performing all of the aspects of a drug’s fate in the body.
A single IV dose study in animals produces three PK data points that are important signals about whether the candidate molecule has the potential to be a drug:
- Plasma clearance (CL): How efficiently the body breaks the molecule down and/or excretes it and its metabolites.
- Volume of distribution at steady state (Vdss): How widely the molecule spreads into tissues from the bloodstream.
- Mean residence time (MRT): The average time a molecule remains intact in the body before it’s eliminated.
Getting this data is slow, expensive, and requires animal testing. Anything that predicts these outcomes reliably ahead of time means running fewer studies on molecules that fail, reducing animal testing and allowing better molecules to move to the clinic faster and at lower cost.
The most common way to evaluate whether a design will have an acceptable in vivo PK profile is through a series of in vitro assays that act as fast, cheap proxies. For example, with plasma clearance, in vitro you can measure intrinsic clearance by incubating the molecule with liver microsomes or hepatocytes and tracking decreasing compound levels over time. Medicinal chemists use these data, via in vitro to in vivo extrapolation (IVIVE), to predict in vivo PK behavior and steer the next round of designs. IVIVE applies scaling equations to that measurement to produce a predicted in vivo clearance.
IVIVE is mechanistic and interpretable, but it uses a static characterization of a molecule in a dish to approximate complex behavior in an organism. As a result, its predictions can be off by several-fold, systematically under or over-predicting clearance, providing weak guidance for what happens in an animal.
Our work demonstrates that AI models trained on large, well curated datasets can now deliver more accurate predictions without requiring in vitro assay results at all.
The data that made this approach possible
The corpus of historical Lilly data allows us to overcome common challenges in using machine learning to predict in vivo PK:
- Dataset size: Models addressing this problem have relied on smaller datasets from publicly available sources such as ChEMBL [2] and TDC [3]. Lilly’s extensive internal dataset provides a far greater volume of in vitro and in vivo PK data than public repositories. The plot below shows the relative sizes of the combined in vitro and in vivo PK datasets from Lilly vs. TDC. Information about exact dataset sizes per assay are available to partners in the TuneLab ecosystem.
- Heterogenous conditions: Unlike public sources that compile disjointed data across varying experimental conditions, Lilly’s data corpus captures consistent, high-quality measurements using standardized assay and PK study protocols. This consistency reduces noise and supports reliable model training across diverse endpoints.
- Lack of in vitro → in vivo translation: Because the Lilly corpus contains overlapping in vitro and in vivo data, models can learn fundamental ADME and physicochemical features from larger, well-populated datasets, using this richer molecular understanding to better predict in vivo outcomes on smaller datasets.
This dataset has enabled us to predict in vivo PK with accuracy that exceeds the industry standard method of IVIVE modeling.

Data processing and splitting
In order to test the model’s ability to generalize to new chemical series, splits were generated to minimize the molecular similarity between the train and test set molecules, while maintaining a roughly even distribution of data from each endpoint in each split (train, validation, test).
Morgan fingerprints were computed for all unique molecules, and structural clusters were generated using sphere-exclusion clustering via RDKit’s LeaderPicker [4] algorithm. Each non-centroid molecule was assigned to the cluster of its nearest centroid by Tanimoto similarity. The training, validation, and test sets were created simultaneously across all datasets by assigning entire clusters to splits, ensuring that structurally similar compounds remain in the same partition and preventing information leakage between splits. Cluster-to-split assignment was performed using a greedy deficit-reduction algorithm that minimizes the sum of squared deviations from a target 75%/15%/10% (train/val/test) ratio across all endpoints simultaneously. Importantly, we assigned data from every endpoint associated with a molecule to the same split, i.e. if a molecule is in the test set, all in vitro and in vivo data points for that molecule are included in the test set.
Models were trained with the validation set used for early stopping and best-checkpoint selection. The test set was used only to compute final metrics using the best model checkpoint, reported below and in the model cards available to users of the TuneLab platform.
Model architecture
Our models use a Chemprop [5] multitask graph-based architecture, where each in vitro/in vivo endpoint has its own multi-layer prediction head. This allows us to use the combined data from in vitro and in vivo assay endpoints across the entire Lilly data corpus. Because there are more compounds with in vitro assay endpoint data than with full IV PK study results, the encoder builds a representation of chemistry that captures the physical properties driving all of them. The in vivo heads then inherit a representation that would be challenging to learn from a few thousand PK studies on their own.

Model Performance
We evaluated the model’s ability to predict the three core outputs of an IV PK study — plasma clearance, Vdss, and mean residence time — across mouse, rat, dog, and monkey, on the held-out test set described above. Across all twelve species-endpoint combinations, the model achieves Pearson correlations between predicted and measured values of roughly 0.6–0.8. Depending on species and endpoint, up to 75% of predictions fall within twofold of the measured value, and 69-91% fall within threefold. Predictive power is particularly robust in non-rodent species (dog and monkey), with steady-state volume of distribution reaching 90% or better within a threefold margin.
The performance is consistent across species and endpoints. Vdss and MRT (endpoints that in vitro assays don’t directly measure, because they depend on whole-body tissue distribution) match clearance in predictability. These results also hold up under challenging test conditions. By assigning entire structural clusters to splits, we maximize the structural distance between training and test molecules within our dataset. The true test of generalization will be performance on new chemistry.

To assess whether the model actually improves on the methods scientists commonly use to estimate in vivo PK, we benchmarked it directly against standard proxies for in vivo clearance:
- Measured in vitro intrinsic clearance (hepatocytes)
- in vivo clearance predicted from measured in vitro clearance via IVIVE

Our model’s predictions correlate more closely with observed outcomes than traditional proxies based on in vitro data. On mouse and rat plasma clearance, raw in vitro intrinsic clearance correlates with in vivo outcomes at Pearson r of approximately 0.39–0.45. Applying IVIVE scaling via the well-stirred model helps, lifting correlations to approximately 0.5. Note that in vitro and IVIVE values don’t use any model predictions and correlations are computed across the entire dataset. The multitask model reaches approximately 0.54–0.62, outperforming both experimentally derived proxies.
An in vitro clearance assay captures only hepatic metabolism. IVIVE using the well-stirred model scales that measurement, but doesn’t capture renal excretion, transporter effects, and everything else an intact organism does to a molecule. Trained across the full Lilly corpus of in vitro ADME, physicochemical, and in vivo PK endpoints, the model learns how these properties jointly determine the PK properties of a molecule.
Please note, performance metrics are based on internal testing and may vary with novel chemical space. Users should validate model predictions with experimental data.
Enabling the Ecosystem
The Lilly TuneLab platform is built on a federated learning infrastructure and hosted by a third-party provider, where each participant’s proprietary molecular structures are not shared with Lilly or other participants.
In typical drug discovery programs, the only reliable way to know what a molecule does in an animal is to test in an animal. That dictates how many compounds a program can evaluate, how long each cycle takes, how many resources go to exploring dead ends. Predicting these in vivo endpoints from structure loosens that constraint. Programs can synthesize far fewer compounds, run fewer animal studies, and put more effort into molecules that have the best chance of becoming safe and effective drugs.
References
[1] Tomin, Lucille, et al. “Gaps in AI-Driven Pharmacokinetic Property Prediction for Early Drug Development: A Scoping Review.” Journal of Chemical Information and Modeling 66.16 (2026): 9761-9783.
[2] Zdrazil, Barbara, et al. “The ChEMBL Database in 2023: a drug discovery platform spanning multiple bioactivity data types and time periods.” Nucleic Acids Research 52.D1 (2024): D1180–D1192.
[3] Huang, Kexin, et al. “Artificial intelligence foundation for therapeutic science.” Nature Chemical Biology 18.10 (2022): 1033–1036.
[4] Greg Landrum. RDKit: Open-source cheminformatics software. https://www.rdkit.org
[5] Heid, Esther, et al. “Chemprop: a machine learning package for chemical property prediction.” Journal of Chemical Information and Modeling 64.1 (2024): 9-17.