Module 18: Sequence Bioinformatics & Evolution
Move from a biological database record to defensible alignments, similarity searches, sequence profiles, evolutionary relationships, and CADD decisions without confusing computational output with biological proof.
Learning outcomes
- Navigate sequence, protein, domain, genome, structure, and pathway records with provenance.
- Choose and interpret pairwise alignment, substitution matrices, gaps, and BLAST searches.
- Build and inspect multiple alignments, sequence profiles, and phylogenetic trees.
- Use conservation and genomic context to guide templates, selectivity, resistance, and validation.
Interactive interpretation challenge
The best sequence hit depends on the CADD question
Switch tasks and compare four realistic-looking hits. The smallest E-value or highest identity is not automatically the best biological choice.
Recommended interpretation
Hit B is preferred because its ligand-bound state and strong pocket coverage outweigh its lower global identity. Confirm local geometry and retrospective docking before prospective use.
Hit B: rank 1 for this task
- The hit fits the selected task, but the computational ranking still requires structural or experimental validation.
Model boundary: The task scores are transparent teaching weights, not BLAST outputs or validated probabilities. Real decisions require the alignment itself, domain boundaries, sequence quality, structural metadata, species context, and experimental evidence.
1. Navigate biological databases with provenance
Sequence bioinformatics starts with records, not algorithms. A useful record links molecular identity to evidence, but curated statements, computational predictions, and author-submitted annotations do not have the same reliability.
| Resource class | Typical content | Questions to ask |
|---|---|---|
| Sequence archives | Submitted DNA and RNA sequences | What sample, construct, assembly, and version produced this sequence? |
| Curated protein records | Protein sequence, function, isoforms, features, and cross-references | Which statements are reviewed, experimentally supported, or predicted? |
| Genome browsers | Genes, transcripts, variants, conservation, and regulatory tracks | Which genome build, transcript, strand, and coordinate system are in use? |
| Families and domains | Motifs, profiles, domain boundaries, and classifications | Does the match cover the functional domain and exceed the model-specific threshold? |
| Structure archives | Experimental coordinates, assemblies, ligands, and metadata | Does the structure represent the correct species, construct, state, and biological assembly? |
| Ontologies and pathways | Controlled terms, relationships, reactions, and evidence links | Which evidence code, organism, context, and database version support the annotation? |
2. Pairwise alignment answers a scoped question
An alignment proposes which residues correspond between two sequences. The correct method depends on whether the whole sequences are comparable or only a local region, such as a catalytic domain, is expected to match.
Dot plot
A visual comparison where matching regions form diagonals. Breaks, parallel diagonals, and repeated patterns can reveal insertions, deletions, repeats, or domain rearrangements.
Global alignment
Needleman–Wunsch aligns sequences end to end. It is most appropriate for sequences of similar length with comparable full-length architecture.
Local alignment
Smith–Waterman finds the best-scoring local regions. It is useful for shared domains or motifs within otherwise different sequences.
Dynamic programming
A score matrix evaluates match, substitution, and gap choices to obtain an optimal alignment under the selected scoring model. Optimal does not mean biologically correct.
3. Scoring substitutions and gaps
| Choice | Meaning | Practical interpretation |
|---|---|---|
| Log-odds score | A substitution is compared with its chance expectation. | Positive scores favor substitutions observed more often than expected by chance. |
| BLOSUM matrix | Built from conserved local blocks at a stated clustering threshold. | Higher BLOSUM numbers generally target closer, more conserved relationships. |
| PAM matrix | Derived from an evolutionary substitution model and extrapolated over distance. | Higher PAM numbers represent greater evolutionary divergence. |
| Affine gap penalty | Opening a gap costs more than extending it. | This models one insertion or deletion event more plausibly than many single gaps. |
| Alignment summary | Identity, similarity, length, gaps, and coverage describe different properties. | Report them together and inspect the domain or pocket region directly. |
4. Interpret BLAST as a heuristic search
BLAST accelerates database searching by extending promising sequence seeds into high-scoring segment pairs. It is fast and useful, but it is not guaranteed to return the mathematically optimal alignment.
- 1
Clean and define the query
Confirm sequence type and orientation; remove vector contamination; identify low-complexity, signal-peptide, and transmembrane regions that may dominate a search.
- 2
Choose the search space
Match program, database, organism scope, and filtering settings to the biological question. A larger database changes the expected number of chance hits.
- 3
Inspect more than the top hit
Compare E-value, score, query coverage, identity, gaps, conserved motifs, and domain architecture. Promiscuous domains can produce convincing but incomplete matches.
- 4
Validate the biological claim
Check reciprocal relationships, curated annotations, genomic context, structure, and experimentally supported function before transferring a name or mechanism.
| Program | Query | Database | Typical use |
|---|---|---|---|
| BLASTP | Protein | Protein | Find protein homologs, domains, and related families. |
| BLASTN | Nucleotide | Nucleotide | Find closely related DNA or RNA sequences. |
| BLASTX | Translated nucleotide | Protein | Detect coding regions or protein similarity in an unannotated nucleotide query. |
| TBLASTN | Protein | Translated nucleotide | Search genomes or transcript collections for sequences encoding related proteins. |
5. Multiple alignment and sequence profiles
Exact simultaneous alignment becomes computationally impractical as the number of sequences grows, so common tools use heuristics. Progressive alignment is useful, but early errors can propagate through later steps.
- 1
Curate representative sequences
Remove obvious fragments, duplicates, unrelated domain architectures, and poor-quality records while preserving meaningful diversity.
- 2
Estimate pairwise similarity
Initial pairwise comparisons provide distances used to organize the order of alignment.
- 3
Build a guide tree
The guide tree controls progressive alignment order. It is an algorithmic scaffold, not a validated evolutionary tree.
- 4
Align progressively and inspect
Add sequences or profiles in guide-tree order, then inspect conserved motifs, secondary structure, insertions, low-complexity regions, and domain boundaries.
- 5
Use the profile carefully
A position-specific scoring matrix captures residue preferences by alignment position. Iterative profile searches such as PSI-BLAST can detect distant relatives but can also drift if false positives enter the profile.
6. Read phylogenetic trees as inferences
Root and branch meaning
Rooting gives direction to the tree. Branch lengths may represent inferred change or simply layout, depending on whether the figure is a phylogram or cladogram.
Gene tree versus species tree
A gene family can have duplications, losses, horizontal transfer, or incomplete lineage sorting, so its tree need not match the organismal history.
Orthologs and paralogs
Orthologs diverge through speciation; paralogs through gene duplication. Both can retain or change function, so the label guides rather than replaces functional analysis.
Support and sensitivity
Bootstrap support measures stability under resampling. Test sensitivity to sequence sampling, alignment, trimming, evolutionary model, and rooting choice.
7. Infer function from genomic context
Sequence similarity is only one clue to function. Comparative genomics can add evidence when the protein is poorly annotated or belongs to a broad family.
Phylogenetic profiles
Genes that are repeatedly gained or lost together across genomes may participate in the same process, although shared ecology and genome quality can confound the pattern.
Gene neighborhood and fusion
Conserved proximity or fusion into one protein can support a functional relationship, especially in microbial genomes.
Conserved noncoding regions
Conservation outside coding sequence can suggest regulatory elements, but functional interpretation requires cell-type and experimental evidence.
Annotation boundaries
Gene prediction and automatic annotation can miss exons, split genes, or propagate an old error. Inspect the underlying sequence and evidence before modeling a target.
8. Translate sequence evidence into CADD decisions
- 1
Confirm the target sequence
Record species, isoform, mutations, signal peptide, transmembrane segments, domains, construct boundaries, and database version.
- 2
Choose structural templates
Evaluate full and local coverage, pocket-residue conservation, insertions, oligomeric state, ligand state, and experimental quality rather than percent identity alone.
- 3
Anticipate selectivity and resistance
Map conserved and variable residues across paralogs, species, pathogens, and known resistant variants onto the binding site.
- 4
Design validation
Test predicted function, binding, selectivity, and mechanism with orthogonal assays. Keep alignment, search, database, and tree settings with the project record.