Way2Drug Projects PASS Targets
Way2Drug helps to understand in silico in Drug Discovery

PASS Targets

is designed to predict which protein targets a given drug-like organic compound may interact with using its 2D structure. The PASS Targets model was trained on public compound–target interaction data extracted from ChEMBL and additionally filtered/curated by us to reduce noise and contradictions in database records. As a result, the system leverages SAR knowledge for 589,107 distinct compounds and can predict interactions with 2,507 protein targets from different organisms. Prediction accuracy, validated by leave-one-out and 20-fold cross-validation, reaches an average AUC ROC of ~96–97% on the training set and ~90% on an external test set of known drugs across 206 targets, confirming the model's reliability for real-world applications.

An important evolution of the methodology was presented in the 2019 study, which addressed "sample selection bias" in the experimental testing of chemical compounds. The authors demonstrated that public database records are heavily biased because compounds are usually selected for testing based on prior knowledge rather than at random.

Using the PASS algorithm, researchers were able to accurately predict whether a given compound had ever been tested against a specific target. Crucially, if the algorithm "recognizes" a compound as one that would likely be tested (meaning it falls within the model's applicability domain), the prediction accuracy for its actual activity is significantly higher (an average ROC AUC of ~0.87 compared to ~0.75 for compounds outside this domain).

How it works?

In PASS Targets, a compound’s structure is encoded as a set of 2D descriptors called MNA (Multilevel Neighbourhoods of Atom), and a Naive Bayes–type approach (as implemented in PASS) is used to compute probabilistic interaction estimates for each target in the training set. A key distinction of PASS Targets versus “classic” PASS is that it predicts the fact of interaction with a protein target rather than a specific mechanism/type of action (e.g., inhibition), which enables broader use of heterogeneous public data and reduces bias from assay-condition differences. You submit the compound’s structural formula in the web interface, and the service returns a predicted interaction spectrum: a list of protein targets with probabilistic scores for each target. In PASS terminology, results are typically interpreted via probabilities of being “active” (Pa) and “inactive” (Pi) for each target; in practice, users often rank targets by Pa–Pi and/or focus on top targets with the highest confidence.

service details

Practical use

PASS Targets is designed to support several concrete scenarios in early-stage drug research. First, it can identify the most probable molecular targets for a newly synthesized or repurposed compound, allowing researchers to narrow down which in vitro assays to run rather than testing blindly across thousands of proteins. Second, the tool can be used in reverse. It can start with a known protein target and search a library of compounds to find those predicted to interact with it. This effectively acts as a virtual screening filter. Third, predicted off-target interaction profiles can flag potential side effects or toxicity risks before costly preclinical experiments: for example, the we demonstrated this by predicting non-kinase off-targets of 28,000 submicromolar kinase inhibitors, revealing candidates relevant to pathogens such as Plasmodium falciparum and Mycobacterium tuberculosis. Finally, PASS Targets supports drug repurposing workflows, where the full predicted interaction spectrum for a known drug can uncover new therapeutic applications that were not originally intended.

The algorithm's predictive performance is highly robust: the average ROC AUC is roughly 96-97% in cross-validation procedures, and around 90% when evaluated on independent external test sets of known drugs.

Other services that might interest you:

Why PASS Targets might be useful for you?

Rationalizing experimental screening: It significantly narrows down the experimental search space, allowing you to prioritize specific in vitro or in vivo tests and saving valuable time and laboratory resources.

Discovering new applications: By predicting the full spectrum of interacting protein targets, it helps uncover novel therapeutic potentials for your synthesized chemical compounds or existing drugs (drug repurposing).

Early prediction of side effects: Identifying unintended off-target interactions during the virtual screening stage allows you to flag potential adverse effects or toxicity risks long before moving to expensive preclinical trials.

Which publication describes this service and how should it be cited?

Pogodin P.V. et al. (2015)

PASS Targets: Ligand-based multi-target computational system based on a public data and naïve Bayes approach

SAR and QSAR in Environmental Research, 26(10):783-93.

doi: 10.1080/1062936X.2015.1078407

Pogodin P.V. et al. (2019)

Improving (Q)SAR predictions by examining bias in the selection of compounds for experimental testing

SAR and QSAR in Environmental Research, 30(10):759-773.

doi: 10.1080/1062936X.2019.1665580

What to do if I have a large dataset?

If you need to evaluate a large dataset, or if you need to maintain confidentiality of structural formulas transmitted via unsecured data channels, you can contact us to discuss the licensing opportunities.