Module 17: Bioinformatics & Systems Biology Foundations
Connect omics, interaction, pathway, perturbation, and chemical-biology evidence so each CADD project begins with the right molecular identity and a defensible disease hypothesis.
Learning outcomes
- Trace genetics, expression, chemistry, structure, and pathway evidence across databases.
- Interpret omics results with attention to study design, uncertainty, and biological context.
- Distinguish physical, functional, regulatory, and correlative interaction evidence.
- Turn network and pathway hypotheses into falsifiable CADD milestones.
Interactive evidence exercise
Would you advance this target?
Grade each independent evidence stream. Watch how a chemically tractable target can still fail when disease causality or safety remains weak.
Do variants support the target and direction of modulation?
Do genetic or chemical perturbations change the disease phenotype?
Is the target present and altered in the relevant cell type and condition?
Do selective ligands or convincing chemical starting points exist?
Is a relevant pocket, interface, or ligandable state supported?
Do tissue, phenotype, and selectivity data support a usable safety window?
Evidence readiness
Safety uncertainty
61/100
The target may be biologically and chemically promising, but the therapeutic window is not yet supported.
Highest-value next experiments
- 1Safety margin evidence: Profile essential tissues, paralogs, liabilities, and on-target phenotypes before lead expansion.
- 2Human genetics: Resolve the causal gene and modulation direction with fine-mapping or functional variant studies.
Triangulation check
- Disease causality
- Moderate
- Chemical or structural access
- Moderate
- Biological context
- Moderate
- Safety margin
- Indirect
Model boundary: The weighted score is a teaching aid, not a probability of clinical success. Record the source, independence, quality, direction, and disease context of each evidence claim, then define experiments that could falsify the target hypothesis.
1. The evidence layers behind a CADD project
Computer-aided drug design begins before molecular modeling. Genetics, expression, pathways, phenotypes, chemical biology, and structural evidence determine which target, molecular state, and assay should be modeled.
| Evidence layer | Typical evidence | CADD decision supported |
|---|---|---|
| Molecular identity | Curated sequence, isoform, domain, and structure records | Target construct, biological assembly, binding site, and model boundaries. |
| Chemistry and activity | Standardized compounds, assay protocols, endpoints, and selectivity panels | Known ligands, activity definition, training data, and off-target risks. |
| Genetics and disease | Human variants, association studies, functional screens, and model systems | Causal support, desired modulation direction, patient group, and safety. |
| Expression and cell state | Bulk, single-cell, spatial, and proteomic measurements | Relevant tissue, cell type, disease state, compensatory response, and biomarkers. |
| Networks and pathways | Physical interactions, regulation, metabolism, perturbations, and phenotypes | Mechanism, bypass routes, combinations, and system-level consequences. |
2. What each omics layer contributes
| Layer | What is measured | Drug-discovery use | Common limitation |
|---|---|---|---|
| Genome | DNA sequence and variation | Causal variants, target direction, resistance, and pharmacogenomics. | A variant association does not by itself identify the causal gene or mechanism. |
| Transcriptome | RNA abundance and isoforms | Cell-state expression, response signatures, and compensatory pathways. | RNA abundance may not predict protein amount or activity. |
| Proteome | Protein abundance, interactions, and modifications | Target presence, complex formation, signaling state, and target engagement. | Coverage and dynamic range vary by protein class and method. |
| Metabolome | Small-molecule products and intermediates | Pathway consequences, mechanism, efficacy markers, and safety signals. | Metabolites can have several biological and technical sources. |
| Metagenome | Microbial genes, pathways, and taxa | Microbiome targets, xenobiotic metabolism, and resistance reservoirs. | Composition and function are strongly affected by environment and sampling. |
3. Reading expression evidence without overclaiming
Expression studies compare conditions, but the biological conclusion depends on study design. A large fold change from a confounded or poorly replicated experiment is not strong target evidence.
Design before statistics
Define the biological contrast, sample unit, replicates, covariates, and batch structure before testing differential expression.
- Avoid treating repeated measurements as independent samples.
- Balance batches across conditions whenever possible.
Effect size and uncertainty
Interpret normalized abundance, fold change, confidence intervals, and multiple-testing-adjusted significance together.
- A small precise effect may be real but not useful.
- A large uncertain effect needs replication.
Cellular resolution
Bulk tissue averages cell types. Single-cell and spatial data can localize a signal, but sparsity, cell annotation, and compositional shifts introduce new uncertainty.
Association is a starting point
Coexpression and clustering generate hypotheses. Genetic or chemical perturbation and orthogonal protein-level evidence are needed to test causality.
4. Interaction evidence is method-dependent
An interactome is assembled from experiments with different meanings. Low overlap between datasets can reflect limited coverage, biological context, and measurement noise rather than a single correct map.
| Evidence type | What it supports | Interpretation checkpoint |
|---|---|---|
| Binary interaction assay | Two proteins can associate in the assay system. | Confirm localization, construct quality, directionality, and an orthogonal assay. |
| Affinity purification–mass spectrometry | Proteins occur in the same captured complex. | Complex membership does not prove a direct pairwise contact. |
| Genetic or CRISPR interaction | The combined perturbation changes fitness or phenotype. | A functional relationship may be indirect and context-specific. |
| Coexpression and colocalization | Molecules vary together or occupy compatible compartments. | Compatibility is useful support, not proof of physical binding. |
| Regulatory occupancy | A factor is enriched near a genomic region. | Pair occupancy with expression or perturbation evidence to infer regulation. |
5. Reconstructing and contextualizing pathways
Pathway databases provide curated consensus models. A disease-, tissue-, or organism-specific pathway is a contextual hypothesis built from those references and the available measurements.
- 1
Define the graph
Represent genes, proteins, metabolites, or reactions as nodes and label each regulatory, physical, or biochemical edge with direction and evidence.
- 2
Anchor the question
Start from disease genes, perturbed proteins, metabolites, or pathway seeds rather than searching the entire network without a hypothesis.
- 3
Build the relevant subnetwork
Connect seeds with plausible paths, known reactions, and compartment constraints. Keep alternative routes instead of forcing one neat story.
- 4
Add biological context
Overlay expression, abundance, localization, variants, and perturbation responses for the relevant tissue, cell type, and condition.
- 5
Validate the model
Test predicted nodes or edges experimentally and compare against held-out evidence. Database agreement is not independent validation.
6. Networks in systems pharmacology
Topology
Communities and centrality can reveal influential nodes, but highly connected essential proteins may also carry greater safety risk.
Pathway enrichment
Test whether a defined gene set is overrepresented or coordinately shifted. Correct for multiple testing and recognize overlapping gene sets.
Polypharmacology
Map intended and off-target activities onto disease and safety networks. Multiple targets can explain efficacy, toxicity, or resistance.
Perturbation profiles
Genetic, chemical, and phenotypic perturbations connect a molecular intervention to system response and can support target deconvolution.
7. An integrated target-to-model evidence workflow
- 1
Frame the disease mechanism
Specify the tissue, cell type, molecular phenotype, desired modulation direction, and evidence that the mechanism is causal rather than correlated.
- 2
Resolve the molecular identity
Select organism, gene, isoform, domains, sequence construct, modifications, complexes, and relevant disease variants.
- 3
Connect chemical and biological evidence
Curate ligands and assays, standardize endpoints, distinguish biochemical from cellular activity, and inspect selectivity and phenotypic signatures.
- 4
Choose the structural hypothesis
Select or model the relevant conformation and assembly; map variants, interfaces, regulatory sites, and pathway context onto the structure.
- 5
Define falsifiable milestones
State which experiment will validate target engagement, mechanism, selectivity, and downstream phenotype before optimizing computational scores.