Module 13: Structural Bioinformatics & Molecular Visualization
Move from a database entry to a chemically defensible receptor model, then use PyMOL or Chimera to inspect, compare, communicate, and validate structural hypotheses.
Learning outcomes
- Relate protein sequence, secondary structure, domains, and assemblies to a CADD question.
- Select an experimental or predicted structure for a specific CADD task.
- Interpret PDB/mmCIF records and classify ligands, waters, ions, and cofactors.
- Prepare and visually validate a receptor without erasing mechanistic chemistry.
- Build and assess a comparative model when no suitable experimental structure is available.
Interactive decision sandbox
Is this structure ready for the task?
Change the intended use and local evidence. The same coordinate model can be suitable for one question and unsafe for another.
Readiness
Ready for validation
95/100
Next defensible actions
Record preparation provenance, validate retrospectively where possible, and keep an alternative receptor model for sensitivity testing.
Model boundary: This score is an educational checklist, not a structure-validation metric. Inspect the experimental map or prediction confidence, chemistry, biological assembly, and task-specific retrospective performance before using a receptor prospectively.
1. From a biological question to a trustworthy structure
A structure is not simply a coordinate file. It is an experimental or predicted model with a specific biological assembly, sequence construct, resolution, missing regions, ligands, ions, waters, and uncertainty. Those details decide whether the structure is suitable for docking, pharmacophore modeling, or molecular dynamics.
Experimental structures
X-ray crystallography, cryo-EM, and NMR provide coordinates supported by experimental data, but each method has characteristic uncertainty and preparation needs.
- Check method, resolution or map quality, R-values, and local confidence.
- Inspect mutations, tags, missing loops, alternate locations, and crystal contacts.
- Use the biologically relevant assembly rather than assuming the asymmetric unit is correct.
Predicted structures
AlphaFold-class models can fill structural gaps, but confidence is spatially variable and a predicted apo conformation may not reproduce a ligand-ready pocket.
- Use per-residue predicted Local Distance Difference Test (pLDDT): values below 70 usually indicate low local confidence, so do not treat nearby pocket coordinates as docking-ready without additional evidence or remodeling.
- Inspect Predicted Aligned Error (PAE) for uncertainty in the relative placement of domains or chains; high PAE can warn against interface or allosteric-site interpretation even when local pLDDT is high.
- Remember that high pLDDT supports local fold confidence but does not guarantee correct pocket rotamers, protonation, ligand-induced conformations, or docking performance.
- Prefer an experimental holo template when accurate ligand geometry is essential.
2. Protein structure from sequence to assembly
Before inspecting a binding pocket, connect what the coordinate model shows to the levels of protein organization. Each level answers a different CADD question, from residue chemistry to domain motion and complex formation.
Primary structure
The amino-acid sequence and its covalent connectivity define residue identity, numbering, variants, termini, disulfides, and the construct that was studied.
Secondary structure
Backbone phi and psi angles favor recurring helices, sheets, and turns. Ramachandran analysis flags unusual geometry, but functional strained residues require context rather than automatic deletion.
Tertiary structure and domains
The three-dimensional fold brings distant sequence positions together. Domains can fold or move semi-independently, changing allosteric sites and pocket accessibility.
Quaternary structure
Biological assemblies place chains, cofactors, nucleic acids, or partners together. Interfaces can create the binding site or stabilize the modeled state.
3. Reading PDB and mmCIF records
Coordinate formats connect atom names and residue identities to Cartesian coordinates. Legacy PDB files are column-based; mmCIF is the modern, extensible archive format. Both must be interpreted together with the entry metadata.
| Record or concept | What it describes | Why it matters in CADD |
|---|---|---|
| ATOM | Standard polymer atoms | Defines the protein or nucleic-acid receptor coordinates. |
| HETATM | Ligands, cofactors, ions, modified residues, and often solvent | Separates the reference ligand and essential cofactors from removable crystallization components. |
| Chain and residue IDs | Polymer membership and residue position | Prevents accidental deletion of interfaces or docking against the wrong construct. |
| Occupancy and B-factor | Alternate occupancy and positional disorder | Flags uncertain atoms and flexible pocket regions. |
| Connectivity | Covalent links and chemical components | Important for metals, covalent inhibitors, disulfides, and non-standard residues. |
4. A defensible receptor-preparation workflow
- 1
Select the right entry
Match species, construct, binding state, pocket conformation, experimental quality, and ligand similarity to the scientific question.
- 2
Resolve structural issues
Choose alternate conformers, repair missing side chains or loops when justified, assign bond orders, and check unusual residues and covalent links.
- 3
Define chemistry
Assign protonation and tautomer states at the intended pH, orient Asn/Gln/His where needed, retain mechanistic waters and metals, and add hydrogens consistently.
- 4
Minimize conservatively
Relax clashes while restraining experimentally supported heavy atoms. Large unvalidated movements can destroy the information supplied by the structure.
- 5
Record provenance
Keep the source identifier, assembly, preparation software, parameters, removed components, retained waters, protonation decisions, and final checks with the model.
5. PyMOL and Chimera for molecular interpretation
Molecular viewers are analytical tools, not only illustration software. Use representations deliberately: cartoons for fold and topology, sticks for chemistry, surfaces for pocket shape, and distance objects for candidate interactions.
Selection logic
Name the ligand, protein, waters, metals, chains, and neighboring residues as reusable selections. Distance-based selections should be expanded by residue so complete side chains are displayed.
Interaction inspection
Check donor-acceptor geometry, salt bridges, aromatic contacts, hydrophobic enclosure, solvent exposure, steric clashes, and whether a contact is mediated by water or metal.
Structure comparison
Align apo and holo structures, homologs, mutants, or trajectory clusters. Interpret global RMSD together with local pocket rearrangements and domain motion.
Publication output
Use saved scenes, consistent colors, orthographic views when appropriate, ray tracing, transparent surfaces, and labels that identify rather than decorate.
6. Comparative model construction and validation
Comparative modeling transfers structural information from one or more templates to a homologous target. The hardest decisions are template choice and alignment around insertions, deletions, active-site residues, and domain boundaries.
Scope boundary: Module 18 teaches sequence search and template-alignment evidence; this section owns coordinate construction and structural validation; Module 6 applies the finished receptor to docking-specific retrospective tests.
- 1
Find and rank templates
Combine sequence identity and coverage with structure quality, oligomeric state, ligand state, and functional relevance. A slightly less similar holo template may be more useful than a high-identity apo structure.
- 2
Build a structure-aware alignment
Keep catalytic motifs and secondary-structure elements aligned; place gaps preferentially in solvent-exposed loops rather than conserved cores.
- 3
Generate an ensemble
Model alternative loop and side-chain conformations instead of treating one output as uniquely correct, especially around the binding site.
- 4
Validate before use
Check stereochemistry, Ramachandran outliers, clashes, packing, residue environments, pocket geometry, and performance in retrospective docking when known ligands are available.