Skip to course content
Learn CADD
Applied extension

Module 17: Bioinformatics & Systems Biology Foundations

Connect omics, interaction, pathway, perturbation, and chemical-biology evidence so each CADD project begins with the right molecular identity and a defensible disease hypothesis.

Learning outcomes

  • Trace genetics, expression, chemistry, structure, and pathway evidence across databases.
  • Interpret omics results with attention to study design, uncertainty, and biological context.
  • Distinguish physical, functional, regulatory, and correlative interaction evidence.
  • Turn network and pathway hypotheses into falsifiable CADD milestones.

Interactive evidence exercise

Would you advance this target?

Grade each independent evidence stream. Watch how a chemically tractable target can still fail when disease causality or safety remains weak.

Do variants support the target and direction of modulation?

Moderate

Do genetic or chemical perturbations change the disease phenotype?

Moderate

Is the target present and altered in the relevant cell type and condition?

Moderate

Do selective ligands or convincing chemical starting points exist?

Moderate

Is a relevant pocket, interface, or ligandable state supported?

Moderate

Do tissue, phenotype, and selectivity data support a usable safety window?

Indirect

Evidence readiness

Safety uncertainty

61/100

The target may be biologically and chemically promising, but the therapeutic window is not yet supported.

Highest-value next experiments

  1. 1Safety margin evidence: Profile essential tissues, paralogs, liabilities, and on-target phenotypes before lead expansion.
  2. 2Human genetics: Resolve the causal gene and modulation direction with fine-mapping or functional variant studies.

Triangulation check

Disease causality
Moderate
Chemical or structural access
Moderate
Biological context
Moderate
Safety margin
Indirect

Model boundary: The weighted score is a teaching aid, not a probability of clinical success. Record the source, independence, quality, direction, and disease context of each evidence claim, then define experiments that could falsify the target hypothesis.

1. The evidence layers behind a CADD project

Computer-aided drug design begins before molecular modeling. Genetics, expression, pathways, phenotypes, chemical biology, and structural evidence determine which target, molecular state, and assay should be modeled.

Evidence layerTypical evidenceCADD decision supported
Molecular identityCurated sequence, isoform, domain, and structure recordsTarget construct, biological assembly, binding site, and model boundaries.
Chemistry and activityStandardized compounds, assay protocols, endpoints, and selectivity panelsKnown ligands, activity definition, training data, and off-target risks.
Genetics and diseaseHuman variants, association studies, functional screens, and model systemsCausal support, desired modulation direction, patient group, and safety.
Expression and cell stateBulk, single-cell, spatial, and proteomic measurementsRelevant tissue, cell type, disease state, compensatory response, and biomarkers.
Networks and pathwaysPhysical interactions, regulation, metabolism, perturbations, and phenotypesMechanism, bypass routes, combinations, and system-level consequences.

2. What each omics layer contributes

LayerWhat is measuredDrug-discovery useCommon limitation
GenomeDNA sequence and variationCausal variants, target direction, resistance, and pharmacogenomics.A variant association does not by itself identify the causal gene or mechanism.
TranscriptomeRNA abundance and isoformsCell-state expression, response signatures, and compensatory pathways.RNA abundance may not predict protein amount or activity.
ProteomeProtein abundance, interactions, and modificationsTarget presence, complex formation, signaling state, and target engagement.Coverage and dynamic range vary by protein class and method.
MetabolomeSmall-molecule products and intermediatesPathway consequences, mechanism, efficacy markers, and safety signals.Metabolites can have several biological and technical sources.
MetagenomeMicrobial genes, pathways, and taxaMicrobiome targets, xenobiotic metabolism, and resistance reservoirs.Composition and function are strongly affected by environment and sampling.

3. Reading expression evidence without overclaiming

Expression studies compare conditions, but the biological conclusion depends on study design. A large fold change from a confounded or poorly replicated experiment is not strong target evidence.

Design before statistics

Define the biological contrast, sample unit, replicates, covariates, and batch structure before testing differential expression.

  • Avoid treating repeated measurements as independent samples.
  • Balance batches across conditions whenever possible.

Effect size and uncertainty

Interpret normalized abundance, fold change, confidence intervals, and multiple-testing-adjusted significance together.

  • A small precise effect may be real but not useful.
  • A large uncertain effect needs replication.

Cellular resolution

Bulk tissue averages cell types. Single-cell and spatial data can localize a signal, but sparsity, cell annotation, and compositional shifts introduce new uncertainty.

Association is a starting point

Coexpression and clustering generate hypotheses. Genetic or chemical perturbation and orthogonal protein-level evidence are needed to test causality.

4. Interaction evidence is method-dependent

An interactome is assembled from experiments with different meanings. Low overlap between datasets can reflect limited coverage, biological context, and measurement noise rather than a single correct map.

Evidence typeWhat it supportsInterpretation checkpoint
Binary interaction assayTwo proteins can associate in the assay system.Confirm localization, construct quality, directionality, and an orthogonal assay.
Affinity purification–mass spectrometryProteins occur in the same captured complex.Complex membership does not prove a direct pairwise contact.
Genetic or CRISPR interactionThe combined perturbation changes fitness or phenotype.A functional relationship may be indirect and context-specific.
Coexpression and colocalizationMolecules vary together or occupy compatible compartments.Compatibility is useful support, not proof of physical binding.
Regulatory occupancyA factor is enriched near a genomic region.Pair occupancy with expression or perturbation evidence to infer regulation.

5. Reconstructing and contextualizing pathways

Pathway databases provide curated consensus models. A disease-, tissue-, or organism-specific pathway is a contextual hypothesis built from those references and the available measurements.

  1. 1

    Define the graph

    Represent genes, proteins, metabolites, or reactions as nodes and label each regulatory, physical, or biochemical edge with direction and evidence.

  2. 2

    Anchor the question

    Start from disease genes, perturbed proteins, metabolites, or pathway seeds rather than searching the entire network without a hypothesis.

  3. 3

    Build the relevant subnetwork

    Connect seeds with plausible paths, known reactions, and compartment constraints. Keep alternative routes instead of forcing one neat story.

  4. 4

    Add biological context

    Overlay expression, abundance, localization, variants, and perturbation responses for the relevant tissue, cell type, and condition.

  5. 5

    Validate the model

    Test predicted nodes or edges experimentally and compare against held-out evidence. Database agreement is not independent validation.

6. Networks in systems pharmacology

Topology

Communities and centrality can reveal influential nodes, but highly connected essential proteins may also carry greater safety risk.

Pathway enrichment

Test whether a defined gene set is overrepresented or coordinately shifted. Correct for multiple testing and recognize overlapping gene sets.

Polypharmacology

Map intended and off-target activities onto disease and safety networks. Multiple targets can explain efficacy, toxicity, or resistance.

Perturbation profiles

Genetic, chemical, and phenotypic perturbations connect a molecular intervention to system response and can support target deconvolution.

7. An integrated target-to-model evidence workflow

  1. 1

    Frame the disease mechanism

    Specify the tissue, cell type, molecular phenotype, desired modulation direction, and evidence that the mechanism is causal rather than correlated.

  2. 2

    Resolve the molecular identity

    Select organism, gene, isoform, domains, sequence construct, modifications, complexes, and relevant disease variants.

  3. 3

    Connect chemical and biological evidence

    Curate ligands and assays, standardize endpoints, distinguish biochemical from cellular activity, and inspect selectivity and phenotypic signatures.

  4. 4

    Choose the structural hypothesis

    Select or model the relevant conformation and assembly; map variants, interfaces, regulatory sites, and pathway context onto the structure.

  5. 5

    Define falsifiable milestones

    State which experiment will validate target engagement, mechanism, selectivity, and downstream phenotype before optimizing computational scores.

Knowledge check

Self-Assessment ChallengeQuestion 1 of 5

Why is differential expression alone insufficient to establish a therapeutic target?