Skip to course content
Learn CADD
Applied extension

Module 18: Sequence Bioinformatics & Evolution

Move from a biological database record to defensible alignments, similarity searches, sequence profiles, evolutionary relationships, and CADD decisions without confusing computational output with biological proof.

Learning outcomes

  • Navigate sequence, protein, domain, genome, structure, and pathway records with provenance.
  • Choose and interpret pairwise alignment, substitution matrices, gaps, and BLAST searches.
  • Build and inspect multiple alignments, sequence profiles, and phylogenetic trees.
  • Use conservation and genomic context to guide templates, selectivity, resistance, and validation.

Interactive interpretation challenge

The best sequence hit depends on the CADD question

Switch tasks and compare four realistic-looking hits. The smallest E-value or highest identity is not automatically the best biological choice.

Decision rule: Prioritize local pocket conservation, coverage, structural state, and architecture rather than identity alone.

Recommended interpretation

Hit B is preferred because its ligand-bound state and strong pocket coverage outweigh its lower global identity. Confirm local geometry and retrospective docking before prospective use.

Hit B: rank 1 for this task

  • The hit fits the selected task, but the computational ranking still requires structural or experimental validation.

Model boundary: The task scores are transparent teaching weights, not BLAST outputs or validated probabilities. Real decisions require the alignment itself, domain boundaries, sequence quality, structural metadata, species context, and experimental evidence.

1. Navigate biological databases with provenance

Sequence bioinformatics starts with records, not algorithms. A useful record links molecular identity to evidence, but curated statements, computational predictions, and author-submitted annotations do not have the same reliability.

Resource classTypical contentQuestions to ask
Sequence archivesSubmitted DNA and RNA sequencesWhat sample, construct, assembly, and version produced this sequence?
Curated protein recordsProtein sequence, function, isoforms, features, and cross-referencesWhich statements are reviewed, experimentally supported, or predicted?
Genome browsersGenes, transcripts, variants, conservation, and regulatory tracksWhich genome build, transcript, strand, and coordinate system are in use?
Families and domainsMotifs, profiles, domain boundaries, and classificationsDoes the match cover the functional domain and exceed the model-specific threshold?
Structure archivesExperimental coordinates, assemblies, ligands, and metadataDoes the structure represent the correct species, construct, state, and biological assembly?
Ontologies and pathwaysControlled terms, relationships, reactions, and evidence linksWhich evidence code, organism, context, and database version support the annotation?

2. Pairwise alignment answers a scoped question

An alignment proposes which residues correspond between two sequences. The correct method depends on whether the whole sequences are comparable or only a local region, such as a catalytic domain, is expected to match.

Dot plot

A visual comparison where matching regions form diagonals. Breaks, parallel diagonals, and repeated patterns can reveal insertions, deletions, repeats, or domain rearrangements.

Global alignment

Needleman–Wunsch aligns sequences end to end. It is most appropriate for sequences of similar length with comparable full-length architecture.

Local alignment

Smith–Waterman finds the best-scoring local regions. It is useful for shared domains or motifs within otherwise different sequences.

Dynamic programming

A score matrix evaluates match, substitution, and gap choices to obtain an optimal alignment under the selected scoring model. Optimal does not mean biologically correct.

3. Scoring substitutions and gaps

ChoiceMeaningPractical interpretation
Log-odds scoreA substitution is compared with its chance expectation.Positive scores favor substitutions observed more often than expected by chance.
BLOSUM matrixBuilt from conserved local blocks at a stated clustering threshold.Higher BLOSUM numbers generally target closer, more conserved relationships.
PAM matrixDerived from an evolutionary substitution model and extrapolated over distance.Higher PAM numbers represent greater evolutionary divergence.
Affine gap penaltyOpening a gap costs more than extending it.This models one insertion or deletion event more plausibly than many single gaps.
Alignment summaryIdentity, similarity, length, gaps, and coverage describe different properties.Report them together and inspect the domain or pocket region directly.

4. Interpret BLAST as a heuristic search

BLAST accelerates database searching by extending promising sequence seeds into high-scoring segment pairs. It is fast and useful, but it is not guaranteed to return the mathematically optimal alignment.

  1. 1

    Clean and define the query

    Confirm sequence type and orientation; remove vector contamination; identify low-complexity, signal-peptide, and transmembrane regions that may dominate a search.

  2. 2

    Choose the search space

    Match program, database, organism scope, and filtering settings to the biological question. A larger database changes the expected number of chance hits.

  3. 3

    Inspect more than the top hit

    Compare E-value, score, query coverage, identity, gaps, conserved motifs, and domain architecture. Promiscuous domains can produce convincing but incomplete matches.

  4. 4

    Validate the biological claim

    Check reciprocal relationships, curated annotations, genomic context, structure, and experimentally supported function before transferring a name or mechanism.

ProgramQueryDatabaseTypical use
BLASTPProteinProteinFind protein homologs, domains, and related families.
BLASTNNucleotideNucleotideFind closely related DNA or RNA sequences.
BLASTXTranslated nucleotideProteinDetect coding regions or protein similarity in an unannotated nucleotide query.
TBLASTNProteinTranslated nucleotideSearch genomes or transcript collections for sequences encoding related proteins.

5. Multiple alignment and sequence profiles

Exact simultaneous alignment becomes computationally impractical as the number of sequences grows, so common tools use heuristics. Progressive alignment is useful, but early errors can propagate through later steps.

  1. 1

    Curate representative sequences

    Remove obvious fragments, duplicates, unrelated domain architectures, and poor-quality records while preserving meaningful diversity.

  2. 2

    Estimate pairwise similarity

    Initial pairwise comparisons provide distances used to organize the order of alignment.

  3. 3

    Build a guide tree

    The guide tree controls progressive alignment order. It is an algorithmic scaffold, not a validated evolutionary tree.

  4. 4

    Align progressively and inspect

    Add sequences or profiles in guide-tree order, then inspect conserved motifs, secondary structure, insertions, low-complexity regions, and domain boundaries.

  5. 5

    Use the profile carefully

    A position-specific scoring matrix captures residue preferences by alignment position. Iterative profile searches such as PSI-BLAST can detect distant relatives but can also drift if false positives enter the profile.

6. Read phylogenetic trees as inferences

Root and branch meaning

Rooting gives direction to the tree. Branch lengths may represent inferred change or simply layout, depending on whether the figure is a phylogram or cladogram.

Gene tree versus species tree

A gene family can have duplications, losses, horizontal transfer, or incomplete lineage sorting, so its tree need not match the organismal history.

Orthologs and paralogs

Orthologs diverge through speciation; paralogs through gene duplication. Both can retain or change function, so the label guides rather than replaces functional analysis.

Support and sensitivity

Bootstrap support measures stability under resampling. Test sensitivity to sequence sampling, alignment, trimming, evolutionary model, and rooting choice.

7. Infer function from genomic context

Sequence similarity is only one clue to function. Comparative genomics can add evidence when the protein is poorly annotated or belongs to a broad family.

Phylogenetic profiles

Genes that are repeatedly gained or lost together across genomes may participate in the same process, although shared ecology and genome quality can confound the pattern.

Gene neighborhood and fusion

Conserved proximity or fusion into one protein can support a functional relationship, especially in microbial genomes.

Conserved noncoding regions

Conservation outside coding sequence can suggest regulatory elements, but functional interpretation requires cell-type and experimental evidence.

Annotation boundaries

Gene prediction and automatic annotation can miss exons, split genes, or propagate an old error. Inspect the underlying sequence and evidence before modeling a target.

8. Translate sequence evidence into CADD decisions

  1. 1

    Confirm the target sequence

    Record species, isoform, mutations, signal peptide, transmembrane segments, domains, construct boundaries, and database version.

  2. 2

    Choose structural templates

    Evaluate full and local coverage, pocket-residue conservation, insertions, oligomeric state, ligand state, and experimental quality rather than percent identity alone.

  3. 3

    Anticipate selectivity and resistance

    Map conserved and variable residues across paralogs, species, pathogens, and known resistant variants onto the binding site.

  4. 4

    Design validation

    Test predicted function, binding, selectivity, and mechanism with orthogonal assays. Keep alignment, search, database, and tree settings with the project record.

Knowledge check

Self-Assessment ChallengeQuestion 1 of 6

When is a local alignment more appropriate than a global alignment?