A multi-method bioinformatics investigation to characterize an unannotated bacteriophage protein — integrating sequence homology search, remote homology detection, structural prediction, and literature synthesis to build a case for functional assignment from first principles.
Bacteriophages — viruses that infect bacteria — are the most abundant biological entities on Earth, yet the majority of their encoded proteins remain functionally uncharacterized. Genomic sequencing has far outpaced functional annotation: databases are filled with entries labeled simply "hypothetical protein," with no information about what they actually do.
PRR1 is a cystovirus — a double-stranded RNA bacteriophage that infects Pseudomonas aeruginosa, a clinically significant pathogen responsible for serious infections in immunocompromised patients. Understanding the complete functional repertoire of PRR1's genome matters both for basic phage biology and for potential therapeutic applications.
The target for this project: gp1, an open reading frame in the PRR1 genome with no experimentally confirmed function. When standard BLAST searches return only "hypothetical protein" hits, the question becomes: how do you determine what a protein does when no one has characterized it before?
The challenge: Standard sequence similarity searches often fail for divergent phage proteins because they evolve rapidly and share little sequence identity with well-characterized proteins — even when their structures and functions are conserved. This project used a convergent multi-method approach to work around that limitation.
When sequence identity alone isn't enough, convergent evidence from multiple independent methods is the most reliable path to functional inference. Each tool probes a different aspect of protein structure and evolution.
PSI-BLAST was run iteratively against the UniProtKB database. Hits are shown by sequence identity (%) and E-value significance (-log₁₀ scale — higher bars indicate stronger statistical support). Related cystovirus proteins from φ6 and φ8 provided the most significant alignments, pointing toward conserved functional roles within this phage family.
Confidence in functional annotation increases when independent methods point to the same conclusion.
Amino acid composition analysis reveals the physicochemical character of the protein. Residue properties — hydrophobic, polar, charged, and structurally special — inform predictions about how the protein folds and interacts with other molecules, providing additional context for functional inference.
What this project demonstrates about working with biological data under uncertainty