Way2Drug Projects RHIVDB
Way2Drug helps to understand in silico in Drug Discovery

RHIVDB

is a freely available web database containing 1,653 amino acid sequences of HIV proteins including reverse transcriptase (RT), protease (PR), integrase (IN), and envelope protein (ENV), treatment history information, and CD4+ cell count and viral load data available by the user’s query.

What the Database Contains?

RHIVDB stores three interconnected categories of data for each patient record, all collected from clinical samples across all federal districts of the Russian Federation:

HIV-1 amino acid sequences of four viral proteins: reverse transcriptase (RT), protease (PR), integrase (IN), and envelope protein (ENV), sequenced using Sanger sequencing via ViroSeq HIV-1 Genotyping System or AmpliSens HIV-Resist-Seq;

Antiretroviral treatment history: specific drug combinations (covering protease inhibitors, NRTIs, NNRTIs, and integrase inhibitors) and the time periods during which they were taken, including flags indicating therapy changes;

Clinical blood parameters: CD4+ lymphocyte cell count and plasma HIV RNA viral load (copies/ml), along with patient age, gender, and year of diagnosis;

The database hold 1,653 RT/PR sequence records from 1,094 unique patients, 281 IN sequences, 276 ENV sequences, and 434 drug combination records. Diagnosis dates span from 1997 to 2019, and blood sampling dates range from January 2014 to December 2019. No personally identifiable patient information is stored.

How to Use It?

Registration is not required to access the content. The data can be accessed and explored in the following ways:

Keyword search via the "Search" tab for quick lookups;

Complex Boolean filtering (AND/OR operators) to combine multiple query criteria, including ranges for CD4+ cell count or viral load;

Registered users can submit data, including new sequence and treatment records, which are added after expert verification in the field of HIV epidemiology.

service details

Database uniqueness

The primary distinguishing feature of RHIVDB is the integration of three data types within a single record: HIV-1 amino acid sequences, full antiretroviral treatment history, and longitudinal clinical outcome parameters (CD4+ cell count and viral load). While databases such as the LANL HIV Sequence Database primarily focus on sequence or genotype–phenotype relationships, RHIVDB uniquely links each sequence to the specific drug combinations a patient was taking and to measurable clinical indicators of therapy success or failure.

A second distinguishing aspect is its geographic origin: the database is built exclusively from clinical samples collected across all federal districts of the Russian Federation, representing a patient cohort and HIV-1 subtypes that are underrepresented in most globally oriented HIV databases. Diagnosis dates span over two decades (1997–2019), giving the dataset a meaningful longitudinal depth.

Third, RHIVDB records therapy dynamics — not just a single treatment snapshot, but sequential drug regimens per patient, with on average two therapy schemas per person and up to 14 regimens recorded for a single individual. This makes it particularly valuable for studying the development of drug resistance over successive lines of treatment.

Applications

The database is designed to support both clinical and computational research:

Therapy effectiveness analysis - comparing CD4+ cell count and viral load dynamics across different drug regimens to identify the most and least effective combinations;

Resistance mutation profiling - retrieving sequences associated with specific drugs (e.g., abacavir, zidovudine) to quantify amino acid substitution frequencies at key RT positions such as 65K/R and 74V/L;

Predictive modeling - serving as a curated training dataset for QSAR and machine learning models to enable the prediction of drug exposure, resistance emergence, and therapeutic outcomes based on sequence-derived and clinical features;

Viremic control research - future expansion is planned to include data on patients who do not develop high viral loads over time (elite controllers), supporting personalized HIV treatment strategies.

Other services that might interest you:

Why RHIVDB might be useful for you?

For clinical researchers and HIV specialists, RHIVDB supports retrospective research on associations between antiretroviral regimens, CD4+ cell counts, viral load and resistance-associated sequence variation.

For virologists and molecular biologists: The database provides ready-to-use sets of HIV-1 amino acid sequences linked to known drug exposures, enabling systematic analysis of resistance-associated substitution patterns in RT, PR, IN, and ENV proteins without the need to manually cross-reference multiple data sources.

For bioinformaticians and computational researchers: RHIVDB serves as a curated training and validation dataset for building QSAR or machine learning models that predict drug resistance, drug exposure, or therapeutic outcome from sequence and clinical features - a task that requires precisely the combination of sequence, treatment, and clinical data that the database provides in a single exportable file.

Which publication describes this service and how should it be cited?

Olga A. Tarasova et al. (2021)

RHIVDB: A Freely Accessible Database of HIV Amino Acid Sequences and Clinical Data of Infected Patients.

Frontiers in Genetics, 10:12:679029.

doi: 10.3389/fgene.2021.679029

How can I get this database for my personal use?

If you need to use the complete dataset presented in RHIVDB in your own studies, please contact us to discuss licensing opportunities.

Notice!

The database does not store user search queries and collects no personal data, handling all queries within the user session.