Specificity Projection On Sequence is designed to analyze protein amino acid sequences and predict positions that determine their functional specificity toward ligands. The key premise of the method is that even a single amino acid substitution in a protein can substantially alter its molecular recognition specificity. Many protein families are divided into functional groups based on specificity toward recognized ligands, and predicting "group-discriminating" residues is critically important for theoretical studies, protein engineering, and drug design.
Unlike most existing methods that rely on multiple sequence alignment (MSA), SPrOS applies local pairwise comparison of segments from the test and training sequences. For each position of the test protein, a specificity score is calculated based on the similarity of the local neighborhood of that position to corresponding regions of training proteins, which the user divides into functional groups. This approach enables the identification of specificity-determining positions even when sequence divergence within a family is high — a task that alignment-based methods do not always handle well.
The method was tested on both simulated and real protein families. On model sequences, SPrOS detected specific positions that alignment-based methods missed. For bacterial transcription factors of the LacI/GalR family, the predicted specific residues matched published experimental data. In the more complex case of protein kinases classified by inhibitor specificity, significant positions were found, as expected, in ligand-binding regions. To eliminate bias associated with evolutionary proximity, close homologs of the test protein can be excluded from the training set.
The user submits an amino acid sequence of the test protein, defines groups of training proteins by ligand specificity, and receives a ranked list of positions most significant for distinguishing those groups. The service is also used in proteochemometrics scenarios to search for small-molecule ligands of a target protein by comparing its sequence to a protein-ligand database.
SPrOS addresses a broad range of research and development tasks in structural biology and drug discovery. The service can be applied to identify key residues responsible for substrate or inhibitor selectivity within enzyme families, guide site-directed mutagenesis experiments by narrowing down candidate positions for functional studies, and support the rational design of selective enzyme inhibitors by highlighting positions that differentiate closely related protein subfamilies. Beyond drug discovery, SPrOS can be used for protein engineering tasks, such as redesigning binding pockets to alter or enhance ligand specificity, and for comparative genomics studies, where understanding the determinants of functional divergence is essential.
No alignment required - SPrOS uses local pairwise sequence comparison instead of multiple sequence alignment, making it applicable to divergent protein families where classical MSA-based tools struggle or produce unreliable results.
Actionable, position-level output - the service delivers a ranked list of specific amino acid positions most responsible for functional group discrimination, giving researchers a direct, experimentally testable hypothesis rather than a generic sequence similarity score.
Flexible training set configuration - users can fully customize the composition of training protein groups, assign sequences to functional classes based on any ligand specificity criterion, and explicitly exclude close homologs of the test protein to avoid evolutionary bias, ensuring that the predicted specificity-determining positions reflect genuine functional differences rather than phylogenetic relatedness.
Dmitry A. Karasev et al. (2016)
Prediction of amino acid positions specific for functional groups in a protein family based on local sequence similarity.
International Journal of Molecular Sciences, 24(3), 2463.
doi: 10.1002/jmr.2515
If you need to evaluate a large dataset, or if you need to maintain confidentiality of structural formulas transmitted via unsecured data channels, you can contact us to discuss the licensing opportunities.