Skip to course content
Learn CADD
Applied extension

Module 13: Structural Bioinformatics & Molecular Visualization

Move from a database entry to a chemically defensible receptor model, then use PyMOL or Chimera to inspect, compare, communicate, and validate structural hypotheses.

Learning outcomes

  • Relate protein sequence, secondary structure, domains, and assemblies to a CADD question.
  • Select an experimental or predicted structure for a specific CADD task.
  • Interpret PDB/mmCIF records and classify ligands, waters, ions, and cofactors.
  • Prepare and visually validate a receptor without erasing mechanistic chemistry.
  • Build and assess a comparative model when no suitable experimental structure is available.

Interactive decision sandbox

Is this structure ready for the task?

Change the intended use and local evidence. The same coordinate model can be suitable for one question and unsafe for another.

90%
95%

Readiness

Ready for validation

95/100

Next defensible actions

Record preparation provenance, validate retrospectively where possible, and keep an alternative receptor model for sensitivity testing.

Model boundary: This score is an educational checklist, not a structure-validation metric. Inspect the experimental map or prediction confidence, chemistry, biological assembly, and task-specific retrospective performance before using a receptor prospectively.

1. From a biological question to a trustworthy structure

A structure is not simply a coordinate file. It is an experimental or predicted model with a specific biological assembly, sequence construct, resolution, missing regions, ligands, ions, waters, and uncertainty. Those details decide whether the structure is suitable for docking, pharmacophore modeling, or molecular dynamics.

Experimental structures

X-ray crystallography, cryo-EM, and NMR provide coordinates supported by experimental data, but each method has characteristic uncertainty and preparation needs.

  • Check method, resolution or map quality, R-values, and local confidence.
  • Inspect mutations, tags, missing loops, alternate locations, and crystal contacts.
  • Use the biologically relevant assembly rather than assuming the asymmetric unit is correct.

Predicted structures

AlphaFold-class models can fill structural gaps, but confidence is spatially variable and a predicted apo conformation may not reproduce a ligand-ready pocket.

  • Use per-residue predicted Local Distance Difference Test (pLDDT): values below 70 usually indicate low local confidence, so do not treat nearby pocket coordinates as docking-ready without additional evidence or remodeling.
  • Inspect Predicted Aligned Error (PAE) for uncertainty in the relative placement of domains or chains; high PAE can warn against interface or allosteric-site interpretation even when local pLDDT is high.
  • Remember that high pLDDT supports local fold confidence but does not guarantee correct pocket rotamers, protonation, ligand-induced conformations, or docking performance.
  • Prefer an experimental holo template when accurate ligand geometry is essential.

2. Protein structure from sequence to assembly

Before inspecting a binding pocket, connect what the coordinate model shows to the levels of protein organization. Each level answers a different CADD question, from residue chemistry to domain motion and complex formation.

Primary structure

The amino-acid sequence and its covalent connectivity define residue identity, numbering, variants, termini, disulfides, and the construct that was studied.

Secondary structure

Backbone phi and psi angles favor recurring helices, sheets, and turns. Ramachandran analysis flags unusual geometry, but functional strained residues require context rather than automatic deletion.

Tertiary structure and domains

The three-dimensional fold brings distant sequence positions together. Domains can fold or move semi-independently, changing allosteric sites and pocket accessibility.

Quaternary structure

Biological assemblies place chains, cofactors, nucleic acids, or partners together. Interfaces can create the binding site or stabilize the modeled state.

3. Reading PDB and mmCIF records

Coordinate formats connect atom names and residue identities to Cartesian coordinates. Legacy PDB files are column-based; mmCIF is the modern, extensible archive format. Both must be interpreted together with the entry metadata.

Record or conceptWhat it describesWhy it matters in CADD
ATOMStandard polymer atomsDefines the protein or nucleic-acid receptor coordinates.
HETATMLigands, cofactors, ions, modified residues, and often solventSeparates the reference ligand and essential cofactors from removable crystallization components.
Chain and residue IDsPolymer membership and residue positionPrevents accidental deletion of interfaces or docking against the wrong construct.
Occupancy and B-factorAlternate occupancy and positional disorderFlags uncertain atoms and flexible pocket regions.
ConnectivityCovalent links and chemical componentsImportant for metals, covalent inhibitors, disulfides, and non-standard residues.

4. A defensible receptor-preparation workflow

  1. 1

    Select the right entry

    Match species, construct, binding state, pocket conformation, experimental quality, and ligand similarity to the scientific question.

  2. 2

    Resolve structural issues

    Choose alternate conformers, repair missing side chains or loops when justified, assign bond orders, and check unusual residues and covalent links.

  3. 3

    Define chemistry

    Assign protonation and tautomer states at the intended pH, orient Asn/Gln/His where needed, retain mechanistic waters and metals, and add hydrogens consistently.

  4. 4

    Minimize conservatively

    Relax clashes while restraining experimentally supported heavy atoms. Large unvalidated movements can destroy the information supplied by the structure.

  5. 5

    Record provenance

    Keep the source identifier, assembly, preparation software, parameters, removed components, retained waters, protonation decisions, and final checks with the model.

5. PyMOL and Chimera for molecular interpretation

Molecular viewers are analytical tools, not only illustration software. Use representations deliberately: cartoons for fold and topology, sticks for chemistry, surfaces for pocket shape, and distance objects for candidate interactions.

Selection logic

Name the ligand, protein, waters, metals, chains, and neighboring residues as reusable selections. Distance-based selections should be expanded by residue so complete side chains are displayed.

Interaction inspection

Check donor-acceptor geometry, salt bridges, aromatic contacts, hydrophobic enclosure, solvent exposure, steric clashes, and whether a contact is mediated by water or metal.

Structure comparison

Align apo and holo structures, homologs, mutants, or trajectory clusters. Interpret global RMSD together with local pocket rearrangements and domain motion.

Publication output

Use saved scenes, consistent colors, orthographic views when appropriate, ray tracing, transparent surfaces, and labels that identify rather than decorate.

PyMOL pocket-inspection recipe

6. Comparative model construction and validation

Comparative modeling transfers structural information from one or more templates to a homologous target. The hardest decisions are template choice and alignment around insertions, deletions, active-site residues, and domain boundaries.

Scope boundary: Module 18 teaches sequence search and template-alignment evidence; this section owns coordinate construction and structural validation; Module 6 applies the finished receptor to docking-specific retrospective tests.

  1. 1

    Find and rank templates

    Combine sequence identity and coverage with structure quality, oligomeric state, ligand state, and functional relevance. A slightly less similar holo template may be more useful than a high-identity apo structure.

  2. 2

    Build a structure-aware alignment

    Keep catalytic motifs and secondary-structure elements aligned; place gaps preferentially in solvent-exposed loops rather than conserved cores.

  3. 3

    Generate an ensemble

    Model alternative loop and side-chain conformations instead of treating one output as uniquely correct, especially around the binding site.

  4. 4

    Validate before use

    Check stereochemistry, Ramachandran outliers, clashes, packing, residue environments, pocket geometry, and performance in retrospective docking when known ligands are available.

Knowledge check

Self-Assessment ChallengeQuestion 1 of 5

Why should domain orientation be inspected separately from whole-protein RMSD?