is web application for in silico prediction of hERG potassium channel blocking activity, combining both qualitative (SAR) and quantitative (QSAR) models within a single interface.
The hERG channel is one of the most critical antitargets in drug development. Blocking it disrupts myocardial repolarization, prolonging the QT interval and potentially causing life-threatening arrhythmias such as Torsades de Pointes and sudden cardiac death. Because experimental hERG testing is mandatory for any new drug candidate, computational pre-screening tools like hERG-Pred help reduce time and cost at early development stages.
The service is built on five (Q)SAR models developed with GUSAR software, trained on data from ChEMBL database (ver. 24):
Two QSAR models for quantitative prediction of pIC50 and pKi values (trained on 4,903 and 1,153 compounds with exact values, respectively);
Three SAR models for qualitative classification (active/inactive hERG blocker) based on IC50, Ki, and inhibition % endpoints;
Structural similarity search using the JS molecular editor, supporting SMILES, MOL, and SDF inputs - similarity is computed via MNA (Multilevel Neighborhoods of Atoms) and QNA (Quantitative Neighborhoods of Atoms) descriptors with Tanimoto and Todeschini coefficients.
A key innovation of hERG-Pred is its use of both exact and inexact experimental values (those reported as >, <, ≥, ≤ a threshold) in SAR model training. This increased the IC50 training set by 50% and improved both prediction accuracy and applicability domain coverage.
Predictive Performance
| Model type | Endpoint | Key metric (5-fold CV) | Applicability domain |
|---|---|---|---|
| QSAR | pIC50 | R2 = 0.551, RMSE = 0.602 |
97.9% |
| QSAR | pKi | R2 = 0.574, RMSE = 0.608 |
97.9% |
| SAR | IC50 (exact+inexact) | BA = 0.816 | 99.9% |
| SAR | Ki (exact+inexact) | BA = 0.787 | 100% |
| SAR | Inhibition % | BA = 0.771 | 100% |
All models were validated by 5-fold cross-validation, and more than 97% of compounds fell within the applicability domain.
The web service accepts compound structures in multiple formats: SMILES strings, drug names, MOL/SDF files, or structures drawn interactively via the embedded JSME Applet. Results include:
Qualitative classification (blocker / non-blocker) from all three SAR models;
Quantitative predicted pIC50 and pKi values with applicability domain flags;
Downloadable output in PDF, CSV, and Excel formats, with clipboard copy support.
hERG-Pred is designed for use at early stages of drug discovery, when experimental cardiac safety data is not yet available. A researcher can submit a single compound or a small library of drug candidates as SMILES strings, drawn structures, or molfiles, and receive an immediate multi-model assessment of hERG blocking risk. The simultaneous output of qualitative classification (active/inactive across three endpoints) and quantitative potency estimates (pIC50, pKi) enables a tiered decision: compounds flagged as active by multiple SAR models and showing pIC50 or pKi > 6 should be prioritized for experimental hERG patch-clamp assays, while compounds predicted inactive across all models with high applicability domain coverage can proceed with greater confidence.
Early cardiac risk flagging - screen novel drug candidates for potential QT-prolongation liability before committing resources to synthesis or in vitro assays, reducing late-stage attrition caused by cardiotoxicity.
Multi-endpoint coverage - unlike most freely available tools that rely solely on IC50-based classification, hERG-Pred simultaneously provides two quantitative (pIC50, pKi) and three qualitative (IC50, Ki, Inhibition %) predictions, offering a more complete picture of hERG interaction risk.
Large, diverse applicability domain - models trained on over 10,000 curated ChEMBL structures with both exact and inexact experimental values achieve near-complete applicability domain coverage (up to 100%), meaning the tool is applicable to a broad range of organic scaffolds encountered in medicinal chemistry programs.
In the process of publication...
If you need to evaluate a large dataset, or if you need to maintain confidentiality of structural formulas transmitted via unsecured data channels, you can contact us to discuss the licensing opportunities.