deployed for in silico prediction of the cytotoxicity of drug-like compounds against human cancer and normal cell lines. The service is built on classification SAR (Structure–Activity Relationship) models developed using the PASS (Prediction of Activity Spectra for Substances) algorithm. These models were trained on a dataset of 59,882 unique chemical structures extracted from the ChEMBL database (version 23), covering cytotoxicity data for 943 human cell lines. The final validated models cover 278 cancer cell lines and 27 normal cell lines, achieving mean prediction accuracies (AUC) of 0.930 and 0.948, respectively, as assessed by leave-one-out and 20-fold cross-validation procedures.
The PASS algorithm represents each molecule using Multilevel Neighbourhoods of Atoms (MNA) sub-structural descriptors derived from its 2D structure, then applies a Naive Bayes-based classification to estimate cytotoxicity probabilities. For each cell line, the service outputs two probabilities: Pa (probability to be active) and Pi (probability to be inactive), which can be used to rank and filter predictions according to the user's desired confidence threshold.
You can submit a chemical structure in three ways: by entering a SMILES string, uploading a mol file, or drawing the structure directly in the integrated Marvin JS molecular editor. The prediction results are presented in two sortable tables: one for tumor cell lines and one for normal cell lines. Each table lists Pa/Pi values, the short and full names of the cell lines, the tissue type, and the tumor classification. Results can be exported in SDF, CSV, or PDF formats, and each cell line name links to the corresponding ChEMBL entry with experimental data.
CLC-Pred covers cell lines spanning 27 different tissue/organ types, including breast, lung, colon, liver, and melanoma, among others. It is particularly valuable for drug repositioning and virtual anticancer screening, allowing researchers to identify candidate compounds with selective cytotoxicity across diverse tumor types before costly experimental testing. Importantly, the inclusion of normal cell line predictions enables early safety assessment of drug candidates, helping to flag compounds likely to be toxic to healthy tissues. Since its launch in 2016, the service has been used by independent research groups worldwide for the assessment of cytotoxicity of natural and synthetic compounds.
To further develop the CLC-Pred web service, in 2023 we have launched a new version called CLC-Pred 2.0. CLC-Pred 2.0 represents a substantially expanded version of the service. The training set more than doubled, growing to 128,545 structures (ChEMBL + PubChem), and coverage was expanded to include 391 tumor cell lines and 47 normal cell lines.A key innovation is the addition of two new prediction modes: cytotoxicity against the NCI60 panel at three activity thresholds (GI50: 1, 10, and 100 nM), and prediction of 2,170 molecular mechanisms of action (trained on 656,011 structures, AUC 0.979), enabling users to simultaneously obtain both phenotypic and mechanistic information about a compound.
Key differences between CLC-Pred and CLC-Pred 2.0
| Parameter | CLC-Pred (2018) | CLC-Pred 2.0 (2023) |
|---|---|---|
| Training set | 59,882 structures ChEMBL v23 |
128,545 structures ChEMBL v29 + PubChem |
| Tumor cell lines | 278 | 391 |
| Normal cell lines | 27 | 47 |
| NCI60 panel | — | Yes New 3 GI50 thresholds: 1, 10, 100 nM |
| Mechanisms of action | — | 2,170 MOA New 656,011 structures, AUC 0.979 |
| Mean AUC accuracy | 0.930 (tumor) 0.948 (normal) |
0.925 (tumor) 0.923 (normal) |
| Data source | ChEMBL only | ChEMBL + PubChem |
| URL | way2drug.com/cell-line/ | way2drug.com/clc-pred/ |
Early-stage anticancer screening - CLC-Pred allows you to rapidly prioritize compounds with predicted cytotoxic activity across hundreds of cancer cell lines before committing to time-consuming and expensive wet-lab experiments, making it an efficient first-pass virtual screening tool.
Selective toxicity profiling - by simultaneously predicting activity against both tumor and normal cell lines, the service helps you identify compounds that are selectively cytotoxic to cancer cells while sparing healthy tissues — a critical criterion in preclinical drug candidate selection.
Drug repositioning and scaffold exploration - whether you are working with approved drugs, natural products, or novel synthetic scaffolds, CLC-Pred can reveal unexpected cytotoxic potential across diverse tissue types, opening new directions for drug repurposing strategies.
Alexey A. Lagunin et al. (2018)
CLC-Pred: A freely available web-service for in silico prediction of human cell line cytotoxicity for drug-like compounds.
International Journal of Molecular Sciences, 24(2), 1689.
If you need to evaluate a large dataset, or if you need to maintain confidentiality of structural formulas transmitted via unsecured data channels, you can contact us to discuss the licensing opportunities.