Modern drug discovery requires the analysis of vast collections of scientific texts. The Way2Drug.com platform integrates bioinformatics and text mining methods for the automated extraction of structured data from scientific literature, enabling continuous enrichment of its databases and enhancement of predictive models.
A key tool is the automated Chemical Named Entity Recognition (CNER) applied to the texts of scientific publications. We have demonstrated that a Naive Bayes-based classifier achieves high accuracy in identifying chemical compound names directly from abstracts and full-text articles. This approach enables the automatic construction of chemical structure dictionaries for subsequent activity and metabolism prediction within the Way2Drug services.
In particular, we developed an automated process for extracting parent compound–metabolite pairs from scientific abstracts. This process formed the basis for creating the XenoMet corpus, which is specialized for extracting metabolite data automatically. This approach has the potential to enrich the databases of metabolism prediction services.
A distinct research direction involves text mining for studying virus–host cell interactions. We applied text mining approaches to identify molecular mechanisms underlying HIV infection progression, and also automatically extracted data on virus–human interactions and antiviral compounds from large text collections. The resulting data were used, among other applications, to populate the informational sections of AntiCOVID-19 on our platform.
Text mining methods are also applied to the systematic discovery of molecular targets. Using text mining approaches, we identified proteins and genes of the Hedgehog signaling pathway involved in oncogenesis, drawing on data from open-access resources. We also performed a systematic mapping of scientific text fragments from Open Targets to identify potential therapeutic targets at the level of DNA/mRNA, proteins, and metabolites.
The integration of text mining with machine learning methods expands the capabilities of Way2Drug predictive models. In collaboration with co-authors from the Pontifical Catholic University of Rio Grande do Sul, we reviewed the scoring function space in the context of computational models for drug discovery, emphasizing the role of literature-derived data in optimizing docking and assessing ligand–protein binding affinity. We established that the "one size fits all" principle is fundamentally limited — scoring functions developed for a specific target (targeted scoring functions) demonstrate significantly superior predictive performance.
The accumulated arsenal of text mining methods serves three key functions of the platform:
Database enrichment - automated extraction of data on metabolites, antiviral compounds, and molecular targets from millions of publications;
Validation of predictive models - verified text corpora serve as reference datasets for (Q)SAR models;
Scientific trend monitoring - regular scanning of the literature enables timely updates of Way2Drug databases across emerging therapeutic areas, including oncology and HIV infection.