(Q)SAR modeling originally focused on building mathematical relationships between chemical structure and biological activity, typically using regression methods. These models were developed by analyzing training sets containing experimentally measured activity values for compounds tested in the same assay and belonging to a specific chemical series. Once established, these (Q)SAR models could be used to predict the activity of new compounds based on their structural features. However, applicability of such (Q)SAR models was generally limited to compounds that were structurally similar to those in the original chemical series used for model development.
It is now well established that most pharmaceutical agents belong to diverse chemical classes, interact with multiple molecular targets, and exhibit a range of pharmacological effects identified across different experimental test-systems. For example, ChEMBL (version 37) contains approximately 24.5 million activity records for 2.9 million compounds tested in 1.9 million bioassays. A search for beta-2 adrenergic receptor agonists alone yields more than 245,000 records, covering about 3,200 compounds and roughly 425,000 bioassay results.
Because such data are highly heterogeneous, accurate quantitative prediction of biological activity for new compounds can be challenging. As a result, (Q)SAR models are often developed using classification approaches rather than relying solely on precise activity value estimation. In this approach, compounds in the training set are classified as either "active" or "inactive" based on their biological activity values (such as IC50, Ki, etc.). A predefined threshold is used, for example, 10 µM, where compounds with activity values below this cutoff are considered active, while those above it are classified as inactive. Therefore, application of the obtained (Q)SAR model for estimation of a new compound's activity allows to classify it either as active or inactive, which implicitly suggests that the activity value for this compound is below or above the established threshold.
The core (Q)SAR/(Q)SPR models implemented at the Way2Drug platform are based on the approach realized in the computer program PASS (Prediction of Activity Spectra for Substances), which is offered on the platform through the freely accessible PASS Online web service, the first freely available web service capable of predicting biological activity from a compound's structural formula alone, used by more than 90,000 researchers from over 90 countries. PASS operates by the concept of a biological activity spectrum of a drug-like compound, defined as the set of different activities reflecting its interactions with different biological systems. The biological activity spectrum is considered an intrinsic characteristic of the compound, determined solely by its molecular structure. Based on this concept, it is possible to aggregate large, heterogeneous datasets from multiple sources, since any single publication typically reports only a subset of a compound’s biological effects; on the platform, this aggregated training foundation currently comprises structure–activity data for more than 300,000 drug-like biologically active compounds.
We adhere to the "Presumption of Innocence" principle: in PASS, it is assumed that a compound lacks those types of biological activity that are not listed in its spectrum. Certainly, some information about biological activity may be missing (an effect not yet reported or not yet assayed), but this approximation does not significantly affect structure–activity analysis or the resulting predictions because the PASS calculation method is statistically robust.
The PASS-predicted activity spectrum, now covering more than 4,000 types of activity on the platform, encompasses pharmacological effects (e.g., antihypertensive, antiinflammatory etc.), molecular mechanisms of action (e.g., 5-HT1A agonist, cyclooxygenase 1 inhibitor, adenosine uptake inhibitor, etc.), specific toxicities and adverse actions (e.g., mutagenicity, carcinogenicity, etc.), action on antitargets (e.g., HERG channel blocker, etc.), interaction with metabolic enzymes & transporters (e.g., CYP1A substrate, CYP3A4 inhibitor, dopamine transporter antagonist, sodium/bile acid cotransporter inhibitor, etc.), and influence on gene expression (e.g., APOA1 expression enhancer, NOS2 expression inhibitor, etc.), that implemented as a dedicated service, DIGEP-Pred, which predicts drug-induced changes in gene expression profiles using training data derived from the Comparative Toxicogenomics Database.
To describe the structures of drug-like compounds in PASS, we developed descriptors called Multilevel Neighborhoods of Atoms (MNA). Introduced in 1999, MNA descriptors are 2D atom-centered structural fingerprints generated recursively from a molecule's connection table, capturing important physicochemical features that influence how a molecule interacts with biological targets by encoding nested local atomic environments at growing radii. They are calculated directly from a molecule's structural formula, which means they can be used at very early stages of research — even for compounds that have only been designed computationally and not yet synthesized. Several extensions of this descriptor family now underpin distinct Way2Drug services: LMNA (introduced 2014) adds functional-group and charge-state labeling for biotransformation and site-of-metabolism prediction, and PoSMNA encodes molecular pairs as the Cartesian product of two MNA descriptor sets, forming the structural basis of the DDI-Pred web-service for drug–drug interaction prediction, which achieves an average IAP (Invariant Accuracy of Prediction, numerically equivalent to AUC-ROC) of approximately 0.92 across seven major cytochrome P450 isoforms.
The mathematical method underlying PASS is based on an original implementation of a Bayesian approach (specifically, a Bayesian classifier operating on 2D structural descriptors). This design ensures robust structure–activity relationship modeling, delivering reliable accuracy and predictive performance even when training data are incomplete. The method is specifically adapted to handle highly imbalanced datasets, enabling the development of (Q)SAR models that maintain meaningful predictive power under realistic data constraints. The average prediction accuracy of PASS, evaluated by leave-one-out cross-validation, is approximately 95%, a figure now consistently cited across the platform's flagship documentation.
The PASS prediction result is a ranked activity profile. For the compound under analysis, you obtain an ordered list of activities, each annotated with two probabilities: Pa - probability of belonging to the class of actives, and Pi - probability of belonging to the class of inactives. The list is sorted in descending order of Pa-Pi, so activities with the strongest positive evidence appear at the top. By default, the profile includes only activities for which Pa exceeds Pi, though users may apply a different cutoff. It is important to understand what Pa actually measures. It primarily captures how closely a compound's structural signature matches active prototypes in the training set. Accordingly, Pa should not be interpreted as potency or any other quantitative activity parameter. A genuinely active compound with an atypical scaffold for a given endpoint may still receive a low Pa estimate, sometimes even Pa < Pi because, by design of the appropriate calibration, the held-out scores for known actives and inactives are distributed across their full ranges. If, for instance, the Pa value equals 0.9, then for 90% of "actives" from the training set the values are less than for this compound, and only for 10% of "actives" is this value higher. If we decline the suggestion that this compound is active, we will make a wrong decision with probability 0.9.
In case the Pa value is less than 0.5, but Pa > Pi, for more than half of "actives" from the training set the values are higher than for this compound. If we decline the suggestion that this compound is active, we will make a wrong decision with probability less than 0.5. In such a case, the probability of confirming this kind of activity in the experiment is small, but there is more than a 50% chance that this structure has high novelty, a scenario directly exploited in the platform's Drug Repositioning workflow, where activities with Pa in the 0.5–0.7 range outside a drug's registered indication are flagged as repositioning candidates. If the predicted biological activity spectrum is wide, the structure of the compound is quite simple and does not contain peculiarities responsible for the selectivity of its biological action. If it appears that the structure under prediction contains a few new MNA descriptors (in comparison with the descriptors from the compounds of the training set), then the structure has low similarity with any structure from the training set, and the results of prediction should be considered as very rough estimates.
Advantages of the developed method have been demonstrated in comparative computational experiments using thousands of different chemical descriptors and dozens of various mathematical methods. Thanks to its high computational speed (roughly 1,000 compounds screened in 10 seconds), PASS is also effectively used for virtual screening of large chemical libraries and corporate compound collections. This computational speed, combined with an average cross-validated accuracy near 95%, is why PASS-based (Q)SAR forms the methodological backbone not just of PASS Online itself but of derivative services like PASS Targets (2,507 protein targets, ROC AUC 96–97%), DIGEP-Pred, CLC-Pred 2.0, and MetaPASS, all inheriting the same classification logic described for the core PASS approach.
Based on predicted biological activity spectra, the following tasks can be addressed: (1) identifying the most promising directions for experimental testing of a specific compound; and (2) conducting virtual screening to identify compounds likely to exhibit the required biological activity. In both cases, the general recommendation is to test activities (or compounds) sequentially in descending order of the Pa or (Pa−Pi) values reported by the predictions. This approach maximizes the probability of success. Downstream analysis of these ranked profiles is supported by the PharmaExpert expert system, which cross-references PASS output against a formalized knowledge base of drug interactions, protein targets, and signaling pathways to identify compounds satisfying a desired combination of activities while excluding those with defined toxic properties.
It should be emphasized that any method of objective classification of drug-like organic compounds can be used for prediction with PASS. If the corresponding classes are indeed determined by features of molecular structure, then predicting assignment to these classes can be quite successful. That is why the PASS approach is successfully applied for the creation of (Q)SAR/(Q)SPR models for a wide range of activities and properties, which accounts for its versatility, reflected in all of the platform’s services, including PASS Targets (interactions with over 2,500 human and non-human protein targets, ROC AUC 96–97%), CLC-Pred 2.0 (quantitative cytotoxicity, IC50/GI50, across 18 cell lines), Ames Mutagenicity Predictor (mean IAP ≈0.944 across 69 Salmonella typhimurium strains), AntiBac-Pred (growth inhibition across 353 bacterial strains), and AntiHIV-Pred/HVR (HIV activity and drug-resistance prediction). For strictly quantitative endpoints beyond binary classification, the platform complements PASS with GUSAR (General Unrestricted Structure–Activity Relationships), which builds regression-based QSAR/QSPR models from user-supplied numerical training data and underlies the GUSAR Online service for acute rat toxicity and anti-target interaction potency, using QNA quantum-chemical descriptors rather than MNA.