The quality of in silico predictions is directly determined by the quality of the data on which models are trained. This fundamental limitation of any machine learning and (Q)SAR modeling is captured by the "garbage in, garbage out" principle: even a mathematically sound model trained on erroneous, incomplete, or heterogeneous data will reproduce and amplify those errors in its predictions.
Public databases assembled through automated literature parsing or aggregation from disparate sources inevitably contain duplicates, conflicting activity values, errors in compound structures, and incomparable experimental conditions. Different laboratories measure, for example, IC50 values in non-comparable cell systems, at different compound concentrations, and with different incubation protocols. However, automated algorithms cannot evaluate these nuances. The result is models with high internal statistics but poor generalization ability.
An expert curator performing manual data collection carries out several critically important operations that are beyond the reach of automation:
Standardization of conditions - selecting only comparable experimental protocols to ensure homogeneity of the training set;
Structure verification - checking the correspondence between a compound's chemical structure and its assigned activity, and excluding artefactual compounds (reactive groups, aggregators);
Conflict resolution - analyzing discrepancies between sources and assigning reliable consensus values;
Class annotation - correctly defining "active/inactive" thresholds in light of biological context.
Manual curation is precisely what underpins the reliability of all key databases on the platform. WWAD was assembled through verification of records from national drug registers, not by simply merging them. RHIVDB contains clinically verified HIV sequences linked to real patient treatment histories. Phyto4Health includes only compounds with confirmed chemical structures from plants of the Russian Pharmacopoeia. HGMMX classifies compounds as "metabolized/not metabolized" on the basis of experimentally established data, rather than automated assumptions.
WWAD (World Wide Approved Drugs) is the most comprehensive database of approved therapeutic small molecules in the world, compiled from national drug registries of numerous countries. The structures of drugs in WWAD are supplemented with predictions of additional biological activity calculated using the PASS system (2022 version): for each compound, probability estimates of "being active" (Pa) and "being inactive" (Pi) are stored across thousands of types of biological activity. As a result, WWAD serves not only as a reference system for approved drugs, but also as a training dataset for our (Q)SAR models, and as a tool for generating drug repurposing hypotheses and analyzing safety profiles. The accompanying publication on "big data" from national registries analyzes the opportunities and limitations of aggregating information from heterogeneous sources.
RHIVDB is a freely accessible database of amino acid sequences of HIV proteins (reverse transcriptase, protease, integrase, envelope protein), collected primarily in the Russian Federation, supplemented with antiretroviral therapy history, CD4⁺ lymphocyte counts, and patients' viral loads. The resource is used by clinical specialists, biologists, and bioinformaticians to analyze treatment efficacy, assess HIV drug resistance, and study viral sequence variability in relation to treatment regimens. Based on RHIVDB, a web service for predicting HIV drug resistance (HVR) has been developed, analyzing amino acid substitutions in key targets. Additionally, patient data accumulated in the database, especially from those with long-term controlled viral loads, can be used to create personalized viral control models and develop individualized therapy regimens.
Phyto4Health contains information on the phytochemical composition of 268 medicinal plants from the Russian Pharmacopoeia: for each of the 3,128 phytocomponents, the database provides the structure, physicochemical properties, data on interactions with human molecular targets, and computational predictions of over 5,000 pharmacological effects and mechanisms of action. Searching the database is possible via both text queries (plant or compound name) and molecular structural similarity, making it well suited for cheminformatics research. Practical application areas include identifying promising phytocomponents as prototypes for new drugs, studying the mechanisms of action of traditional herbal preparations, and expanding training datasets for biological activity prediction models.
HGMMX (Host Gut Microbiota Metabolism Xenobiotics Database) accumulates data on 368 structures of drug-like compounds metabolized by the human gut microbiota and 310 compounds that are not subject to such metabolism. The relevance of the database stems from the fact that the metagenome of gut bacteria is almost 150 times larger than the host genome, and microbiota enzymes can produce metabolites with fundamentally different activity, reducing the efficacy of pharmacotherapy or causing adverse side effects. HGMMX is intended primarily for developing computational models to predict the biotransformation of compounds by the gut microbiota - a task previously constrained by a critical shortage of structured training data. Thus, the database fills a critical gap in ADMET analysis and serves as a foundation for Way2Drug web services such as Metabolic Stability and MDM-Pred, which assess microbiome-dependent drug metabolism, among other factors.
The combined set of Way2Drug databases and web applications forms a closed-loop in silico research cycle: from searching for known approved drugs (WWAD) and natural compounds (Phyto4Health) - through activity prediction (PASS Online, AntiHIV-Pred, BC CLC-Pred, CLC-Pred 2.0, DIGEP-Pred 2.0) - to safety assessment and pharmacovigilance data analysis. All resources are freely available, consistent with the principles of open science, and ensuring wide adoption within the international scientific community.