Machine learning in drug discovery refers to the use of computational algorithms that learn from previously accumulated chemical, pharmacological, toxicological, and biological data in order to predict the properties and behavior of new compounds. In practical terms, these approaches help researchers infer biological activity, probable targets, metabolism, toxicity, and other critical endpoints directly from molecular structure, thereby reducing the need for exhaustive experimental screening at the earliest stages of drug discovery. For Way2Drug platform, machine learning serves as the methodological foundation that transforms structural formulae into biologically meaningful hypotheses and supports faster, more rational decision-making in medicinal chemistry and translational research.
Modern drug discovery is an extraordinarily complex process, traditionally requiring years of laboratory work, enormous financial investment, and substantial experimental resources. Machine learning has fundamentally transformed the field by enabling researchers to quickly and cost-effectively extract actionable biological predictions from molecular structures on a large scale. The Way2Drug platform embodies this paradigm, offering a comprehensive ecosystem of machine learning driven services that guide scientists from a simple chemical formula to deep biological insight.
At the heart of the platform is PASS Online, a flagship tool that predicts more than 4,000 types of biological activity for molecules based solely on their structural formula from pharmacological effects to mechanisms of action and adverse effects. It is built on a machine learning method using a Naive Bayes classifier operating on 2D structural descriptors. The average prediction accuracy is approximately 95% by leave-one-out cross-validation. Thanks to its high computational speed (1,000 compounds in ~10 seconds), PASS is effectively used for screening large chemical libraries and corporate compound databases. The platform uses ML models to identify new uses for already approved drugs, a process known as drug repurposing. This is supported by numerous publications. Complementing it, PASS Targets maps likely molecular target interactions, while MetaPASS extends this analysis to metabolic products, predicting how biotransformation alters the activity profile of a parent compound.
Infectious disease research is richly served by specialized machine learning modules. AntiHIV-Pred estimates activity against HIV, and HVR focuses on HIV–host virus resistance, while AntiBac-Pred targets bacterial pathogens. AntiCOVID-19 addresses SARS-CoV-2 activity prediction, representing the platform's rapid response capacity for emerging global threats. The geroprotective dimension is covered by PASS Gero, which predicts potential anti-aging biological activities.
Oncology and cytotoxicity prediction form another major cluster. CLC Pred 2.0 estimate compound cytotoxicity across a broad panel of cancer cell lines, and BC CLC-Pred narrows the focus specifically to breast cancer models. DIGEP-Pred 2.0 predicts drug-induced gene expression changes, linking chemical structure with transcriptomic responses.
Drug-drug interactions and pharmacokinetic liabilities are addressed by DDI-Pred, which models interaction potential between co-administered compounds. The metabolic toolkit includes SOMP for sites-of-metabolism prediction, MDM-Pred for microbiota-mediated drug metabolism, P450-analyzer for cytochrome P450 interaction profiling, Metabolic Stability for in vitro half-life estimation, and RA for reactivity assessment of metabolites.
The platform provides machine learning tools for predicting metabolism (Phase I/II enzymes, P-glycoprotein transport), acute toxicity, cytotoxicity, cardiotoxicity, and hepatotoxicity. These models allow filtering out undesirable compounds at early development stages, before laboratory experiments begin. Safety and toxicity are covered by a robust suite. GUSAR Online predicts acute rat toxicity, antitarget liabilities, and ecotoxicity using QSAR models. Ames Mutagenicity Predictor provides mutagenicity assessment, hERG-Pred evaluates cardiotoxicity risk via hERG channel blockade, MetaTox 2.0 integrates metabolism and toxicity prediction, ADVERPred focuses on adverse drug reaction profiles, and ROSC-Pred assesses the risk of organ-specific carcinogenicity.
Research in the field of immunology and sequence-based analysis is represented by tools designed for highly specific biomedical tasks. TCR-Pred is intended for predicting T-cell receptor specificity to antigenic epitopes and major histocompatibility complex molecules, supporting research at the interface of immunoinformatics and antigen recognition. SAV-pred is aimed at predicting the clinical effect of single amino acid substitutions in proteins associated with monogenic hereditary diseases included in newborn screening panels. TIP is likewise focused on the prediction of the clinical significance of single amino acid substitutions in proteins linked to monogenic inherited disorders relevant to newborn screening. Proteochemometrics provides a framework for analyzing ligand–target relationships using combined descriptions of chemical compounds and biological targets, while SprOS serves as a specialized resource within the platform for structure-based analysis in a protein- and sequence-related biomedical context.
Rounding out the ecosystem, MNA-PSS-Pred enables prediction of protein secondary structures. CNER is more accurately described as a knowledge-based resource designed to support the exploration of relationships among chemicals, genes, proteins, diseases, and biological pathways, thereby helping users interpret compound action in a broader systems-biology and network-pharmacology context. CNER supports identification and analysis of chemical entities in scientific text. Its outputs can assist literature mining and the organization of chemical information for downstream biomedical research.In this role, CNER complements the predictive services of the platform by connecting individual molecular predictions with wider biological and biomedical associations.
![]()
We don’t know a single drug discovered solely in silico, but we don’t know a single drug discovered without in silico methods.
![]()
Taken together, the services of Way2Drug platform form an integrated computational environment for early-stage drug discovery and biomedical research. Rather than providing isolated predictions, the platform enables users to move from molecular structure to a connected interpretation of biological activity, target interactions, metabolism, toxicity, gene expression effects, disease relevance, and sequence-associated variation. This combination of complementary machine learning models supports a more systematic evaluation of compounds, metabolites, and biomolecular features within a single research framework. As a result, Way2Drug helps researchers generate hypotheses faster, prioritize experiments more rationally, and reduce the cost and time associated with exploratory screening. In this sense, the platform represents not only a collection of predictive services, but also a practical platform for transforming chemical and biological data into actionable scientific insight.
Benchmarking new machine learning models in drug discovery on real-world problems.
Enriching training datasets with predicted activity labels.
Demonstrating an end-to-end pipeline: molecule -> descriptors -> machine learning prediction -> interpretation.
Use in educational courses on cheminformatics and machine learning in biomedicine.
Validation of third-party machine learning models through independent Way2Drug predictions.