Skip to course content
Learn CADD

Module 1: Introduction to Drug Discovery

Explore the stages of the drug discovery pipeline, learn to distinguish key compound classes, and understand how Computer-Aided Drug Design (CADD) accelerates the transition from biological idea to medicine.


Learning outcomes

  • Trace a compound through the discovery pipeline and explain the attrition at each stage.
  • Distinguish a hit, a lead, and a drug candidate by the evidence each one requires.
  • Choose between target-based and phenotypic hit finding, and state what each costs you.
  • Explain why lack of efficacy — not potency — dominates late-stage failure.
  • Place any module in this course at the pipeline stage it serves.

1. The Drug Discovery and Development Pipeline

Bringing a new drug to the market is a highly complex, multi-stage, interdisciplinary process. Historically, it requires 10–12 years and upwards of $2.6 billion, with an extremely high rate of attrition. For every 10,000 compounds screened at the outset, typically only one receives regulatory approval.

The pipeline acts as a funnel, filtering compounds through successive hurdles of affinity, selectivity, pharmacokinetics, and safety. Computational chemistry and biology (CADD) have become vital tools to cut costs and time by early filtering and rational design.

Interactive Playground: R&D Funnel Simulation

Click "Advance Pipeline" to simulate compound screening through the five major stages of the drug discovery pipeline. Notice the scale of attrition at each barrier.

Stage 1
Stage 2
Stage 3
Stage 4
Stage 5
Remaining Pool

10,000 compounds

Ready to begin the R&D pipeline simulations

2. What CADD Actually Does

Computer-aided drug design is the use of computational models to decide which molecules are worth making and testing. That is the whole of it. CADD does not discover drugs; it changes the order in which you spend your experimental budget, and at the scale of the funnel above, ordering is worth an enormous amount.

Everything in this course divides into two families, and the split depends on one question: do you have a three-dimensional structure of the target?

Structure-Based (SBDD)

You have the receptor — from crystallography, cryo-EM, NMR, or a predicted model. You can reason directly about the pocket: its shape, its charges, which contacts a ligand could make. Docking, structure-based pharmacophores, and free-energy calculations all live here.

Fails when: the structure is wrong, the wrong conformation, or missing the loop that forms the pocket.

Ligand-Based (LBDD)

You have no receptor, but you do have molecules known to work. You infer what the target wants from what its ligands have in common. Similarity searching, ligand-based pharmacophores, and QSAR all live here.

Fails when: the known actives are too few, too similar to each other, or measured in inconsistent assays.

Real projects use both, and the boundary is softer than the labels suggest — a structure-based screen still needs ligand-based filters to stay tractable, and a ligand-based model is far easier to trust once a structure explains why it works.

The honest framing

A computed number is a hypothesis, not a measurement. A docking score is not a binding affinity, a QSAR prediction is not an assay result, and a favourable ADMET profile is not safety. Used well, CADD raises the fraction of your experiments that are worth running. Used badly — trusted instead of tested — it produces confident, well-formatted, and entirely wrong answers. Every method in this course comes with a section on how it fails, and those sections matter more than the equations.

3. Key Definitions in the Screening Funnel

Stage A

Hit Compound

A molecule that shows reproducible, verified activity in a bioassay. It must possess validated structure/purity, novelty, and chemical tractability.

Stage B

Lead Compound

An optimized hit showing activity in vivo, clear Structure-Activity Relationships (SAR), no reactive groups, and clean cardiotoxicity markers.

Stage C

Drug Candidate

A fully optimized lead structure with robust preclinical safety profiles, ready for Investigational New Drug application and clinical trials.

4. Strategies for Identifying Active Hits

Before choosing a screening technology, a project makes a more fundamental choice: what counts as a hit?

Target-based

Pick a protein, assay compounds against it, and call a hit anything that binds or inhibits. The mechanism is known from the first day, so SAR, structural work, and everything in this course apply immediately — but only if the target was the right one. A beautifully optimized inhibitor of a protein that does not drive the disease is a wasted programme.

Phenotypic

Assay compounds against cells or an organism and call a hit anything that produces the desired biological outcome — regardless of how. You are guaranteed relevance in a living system, but you may not know what the compound hits, which makes optimization far harder and demands a separate effort to deconvolute the target.

It is tempting to assume the rational, target-based route dominates. It does not. In a well-known survey of the 259 agents approved by the FDA between 1999 and 2008, 75 were first-in-class — and of the 50 first-in-class small molecules, 28 came from phenotypic screening against 17 from target-based approaches (Swinney & Anthony, Nat Rev Drug Discov 2011). Genuinely novel mechanisms have historically been found more often by asking biology an open question than by interrogating a protein we had already nominated.

The lesson is not that computational, target-based work is misguided — the same survey shows target-based approaches dominating follower drugs, where the mechanism is already established, and CADD has grown enormously since 2008. The lesson is that the target hypothesis is the single riskiest assumption in the project, which is why Module 2 is devoted entirely to interrogating it before any modelling begins.

1

High-Throughput Screening (HTS)

Automated robotic testing of chemical libraries containing millions of synthesized compounds. Highly robust and unbiased, but extremely costly to configure.

2

Exploitation of Biological Information

Repurposing existing drugs based on unexpected clinical observation of side effects (e.g. sildenafil) or traditional medicine extracts.

3

Rational Drug Design

Using structural knowledge of the target protein (structure-based) or active ligands (ligand-based) to construct compounds atom-by-atom.

4

Fragment-Based Screening

Screen very small compounds (< 300 Da) that bind weakly (mM–µM) but with high ligand efficiency. Hits are then grown, linked, or merged into leads. Because fragments are small, a library of a few thousand samples chemical space far more efficiently than a million-compound HTS deck.

5

DNA-Encoded Libraries (DELs)

Each compound is built by split-and-pool synthesis and tagged with a DNA barcode recording its synthetic history. Billions of compounds can then be screened in a single tube: the pool is washed over immobilized target, non-binders are rinsed away, and the surviving barcodes are read by DNA sequencing. The scale is unmatched — but the readout is enrichment of a barcode, not a clean affinity, and hits must be re-synthesized without their DNA tag to be confirmed. DEL selection data has become a major training set for the machine learning models in Module 9.

5. Why Compounds Fail

The funnel says roughly one candidate in ten that enters Phase I reaches approval. The more useful question is why the other nine die — because each cause of death corresponds to something a computational method is trying to prevent.

Of the Phase II and Phase III failures reported between 2013 and 2015 with a stated reason, 52% failed for lack of efficacy and 24% for safety, with the remainder falling to commercial and strategic decisions (Harrison, Nat Rev Drug Discov 2016).

Lack of efficacy — the dominant killer

The compound did what it was designed to do and the patient did not improve. Note what this usually is not: a potency failure. It is most often a target failure — the biological hypothesis was wrong, or the target mattered less in humans than in the model system. No amount of computational optimization rescues this, which is the entire argument for spending real effort on target validation (Module 2) and systems-level evidence (Module 17) before optimizing anything.

Safety and toxicity

Off-target activity, reactive metabolites, cardiac liabilities such as hERG block, or liver injury. These are the failures computational toxicology attacks most directly, because many are predictable from structure — and because a liability caught at the design stage costs a chemist an afternoon, while the same liability caught in Phase II costs years. Modules 12 and 15 cover this ground.

Pharmacokinetics — the quiet one

A compound that never reaches its target at sufficient concentration cannot work, however potent it is in vitro. PK-driven attrition has fallen sharply since the 1990s, precisely because the field learned to screen for absorption and metabolism early rather than late. It is the clearest historical evidence that moving a filter earlier in the funnel actually works.

Read the funnel backwards and you have the design of this course. Every module exists because something kills compounds at that stage, and the whole enterprise rests on one economic asymmetry: the cost of a mistake grows by orders of magnitude the later you find it. Computation is cheap; a failed Phase III is not.

6. Targeting the "Undruggable" Proteome

For decades, drug discovery focused on target-based design against deep, well-defined active pockets (e.g. enzyme ATP-binding clefts). However, over 80% of disease-driving proteins lack such cavities, including transcription factors, intrinsically disordered proteins (IDPs), and flat protein-protein interaction (PPI) interfaces.

Once considered "undruggable," breakthroughs in biotechnology and CADD are opening these targets to therapeutic intervention via novel modalities:

Targeted Degradation (PROTACs)

Bifunctional molecules that bind the target protein on one end and recruit an E3 ubiquitin ligase on the other, tagging the target for destruction by the proteasome rather than merely inhibiting it.

PPI Inhibitors & glues

Targeting flat, solvent-exposed protein-protein interfaces. Drugs like Venetoclax target the BCL-2 interface, while molecular glues stabilize target complexes to drive degradation.

Drug Repurposing

Finding new clinical indications for FDA-approved drugs (e.g., sildenafil, aspirin). This bypasses phase I safety barriers, accounting for nearly one-third of recent approvals.

7. How the Rest of This Course Maps onto the Funnel

The remaining modules are not a list of techniques; they follow the pipeline you just simulated. If you ever lose the thread, come back to this table and ask which stage you are standing in.

Pipeline stageThe questionModules
Target selectionIs this protein worth attacking, and can anything bind it?2, 17
FoundationsWhy do molecules stick together, and how do we compute that?3, 4
RepresentationHow do we store and compare molecules so a machine can reason about them?5, 13, 18
Hit findingWhich molecules, out of millions, deserve an assay?6, 7, 8
Lead optimizationHow do we make this hit better without breaking something else?9, 10, 11
DevelopabilityWill it be safe, absorbed, and eliminated sensibly?12, 15
Beyond small moleculesWhat if the modality is a protein rather than a drug-like ligand?16
Doing it crediblyCould someone else — including you, in a year — reproduce this?14

Modules 1–12 form the core curriculum and are meant to be read in order. Modules 13–18 are shorter applied extensions covering the structural, computational, pharmacokinetic, biologics, and systems-level material a real project needs; take them in any order once the core is behind you.

Knowledge check

Self-Assessment ChallengeQuestion 1 of 6

What distinguishes a validated hit from an initial screening signal?