Way2Drug Projects Proteochemometrics
Way2Drug helps to understand in silico in Drug Discovery

Proteochemometrics

is developed for predicting protein–ligand interactions using a combination of amino acid sequence analysis and ligand structural features.

The service implements the proteochemometrics approach, which explores the combined space of proteins and small-molecule ligands to build large-scale predictive models without requiring solved 3D protein structures. Its core algorithm computes positional similarity scores through segment-to-segment comparison of amino acid sequences, using fuzzy classification coefficients derived from the PASS (Prediction of Activity Spectra for Substances) algorithm to estimate interaction probabilities.

How It Works?

The service offers three distinct prediction scenarios, covering the full range of drug discovery situations:

Ligand-based prediction - predicts target proteins for a query ligand by comparing it against chemical structures of compounds with known target spectra ;

Sequence-based prediction - searches for small-molecule ligands for a test protein by comparing amino acid sequences, without considering ligand structures;

Proteochemometrics prediction - handles the most complex case where neither the target nor the ligand has a known interaction spectrum, requiring both a protein sequence and a ligand structure as input.

Proteochemometrics Apps 1
Proteochemometrics Apps 2
Proteochemometrics Apps 3
service details

Validation and Accuracy

The underlying models were trained on data from the ChEMBL25 database and validated across five major drug target families: GPCRs, protein kinases, ligand-gated ion channels, voltage-gated ion channels, and nuclear receptors. Predictive accuracy (AUC) reached an average of 0.96 for ligand-to-target prediction, 0.95 for target-to-ligand prediction, and values between 0.89 and 0.99 for the fully uncharacterized protein–ligand scenario. Training datasets used by the service are also publicly available.

Practical Use

Proteochemometrics web service fits naturally into a computational drug discovery pipeline at multiple stages. At the hit identification stage, researchers can submit a novel compound's structure and instantly retrieve a ranked list of the most probable protein targets across five major families, narrowing experimental assay priorities. For target deconvolution, when a biologically active compound has an unexplained mechanism of action, the service can suggest plausible off-targets based solely on structural similarity to known ligands. In reverse pharmacology, a newly sequenced or mutant protein variant can be queried against thousands of known ligands to identify candidate binders without performing any docking calculations. Finally, the service can be used to estimate the selectivity of a designed compound, or whether it is likely to interact with unintended members of a protein family, such as closely related kinase subtypes.

Other services that might interest you:

Why Proteochemometrics might be useful for you?

No 3D structure required - the method operates entirely on amino acid sequences and 2D ligand descriptors, making it applicable to proteins with no available crystal structure or homology model, which is common for newly discovered or mutant targets.

Handles truly novel entities - unlike conventional QSAR tools that require a training compound with known activity at the target of interest, this service can evaluate interaction likelihood even when both the protein and the ligand are outside any previously characterized interaction space, using fuzzy probabilistic coefficients.

Broad target family coverage - with validated models for GPCRs, kinases, ion channels, and nuclear receptors, the service covers the most pharmacologically relevant and druggable protein superfamilies in a single unified interface.

Which publication describes this service and how should it be cited?

Dmitry A. Karasev et al. (2020)

Prediction of protein-ligand interaction based on sequence similarity and ligand structural features.

International Journal of Molecular Sciences, 21, 8152.

doi: 10.3390/ijms21218152

What to do if I have a large dataset?

If you need to evaluate a large dataset, or if you need to maintain confidentiality of structural formulas transmitted via unsecured data channels, you can contact us to discuss the licensing opportunities.