Way2Drug Projects DIGEP-Pred 2.0
Way2Drug helps to understand in silico in Drug Discovery

DIGEP-Pred 2.0

(drug-induced gene expression profiles prediction 2.0) was developed for predicting drug-induced gene expression profiles and likely protein targets directly from the structural formula of a drug-like compound. It was developed to help researchers estimate molecular effects when experimental transcriptomic data are unavailable.

The web service is based on a combination of structure-activity relationship (SAR) modeling and network analysis. SAR models were built using PASS (Prediction of Activity Spectra for Substances) algorithm, trained on data from three major sources: the Comparative Toxicogenomics Database (CTD) and the Connectivity Map (CMap) for gene expression prediction, and PubChem and ChEMBL for predicting molecular mechanisms of action. Using leave-one-out cross-validation, the mean prediction accuracy was calculated to be 86.5% for 13,377 genes and 94.8% for 2,932 proteins from CTD data. Meanwhile, the MoA prediction accuracy was found to be 97.9% for 2,170 MoA types. Additional CMap-based models, with a mean accuracy of 87.5%, cover three cancer cell lines: MCF7, PC3, and HL60, at multiple fold-change thresholds.

How It Works?

You can submit a compound through a structural representation interface, including SMILES input and drug name entry. From this input, the system returns predicted gene expression changes, probable direct targets, and downstream biological interpretation linked to the compound.

A key strength of DIGEP-Pred 2.0 is that it does not stop at listing affected genes, but also provides enrichment analysis for pathways from KEGG and Reactome, biological processes from Gene Ontology, diseases from DisGeNet, and target-master regulator estimation based on OmniPath data, which helps connect predicted transcriptomic changes to mechanisms of action and possible therapeutic or adverse effects.

service details

Practical Use

DIGEP-Pred 2.0 can be applied across multiple stages of the drug discovery and development pipeline. Researchers can use it to elucidate the molecular mechanism of action of a newly synthesized or repurposed compound by identifying the genes and signaling pathways it is likely to modulate. This service can also be used for safety assessments. Predicted off-target gene expression changes can identify potential adverse effects early on, before costly in vitro or in vivo experiments are initiated. In phytochemistry and natural product research, DIGEP-Pred 2.0 enables rapid in silico profiling of bioactive secondary metabolites, helping prioritize candidates for further biological evaluation.

Other services that might interest you:

Why DIGEP-Pred 2.0 might be useful for you?

No experimental data required - the service generates a comprehensive gene expression and target profile from nothing more than a structural formula (SMILES), making it immediately accessible at the earliest stages of compound design or screening.

Multi-level biological interpretation - results go beyond a simple gene list, linking predicted expression changes to pathways (KEGG, Reactome), biological processes (Gene Ontology), and diseases (DisGeNet), enabling a systems-level view of a compound's potential effects.

High predictive accuracy - models trained on large, curated datasets from CTD, CMap, PubChem, and ChEMBL achieve mean accuracies above 86–98%, providing a reliable computational basis for research decisions.

Which publication describes this service and how should it be cited?

Sergey M. Ivanov (2024)

DIGEP-Pred 2.0: A web application for predicting drug-induced cell signaling and gene expression changes.

Molecular Informatics, 43, e202400032.

doi: 10.1002/minf.202400032

What to do if I have a large dataset?

If you need to evaluate a large dataset, or if you need to maintain confidentiality of structural formulas transmitted via unsecured data channels, you can contact us to discuss the licensing opportunities.