MODULE 04 · Discovery

Drug discovery engine

Solve one measurable prediction task before building a broad discovery engine. Make data quality, baseline comparisons and out-of-distribution uncertainty visible from the first experiment.

THE FIRST USEFUL PRODUCT

Start with a bounded problem.

Who it is for
Computational researchers and small teams learning or benchmarking molecular prediction.
First product
A reproducible molecular-property benchmark with a simple baseline and a transparent ranking report.
Next decision
A reviewer agrees the labels and metric support the claimed task.
Product boundary

Predicted properties and candidate rankings are research hypotheses. Computational performance alone does not establish efficacy, toxicity or clinical suitability.

PRODUCT 0 → 1.0

The Phase 0–5 strategy.

Each phase has a concrete deliverable and an acceptance gate. These are proposed development plans, not completed scientific or product validations.

PHASE 0 / 0% MILESTONECurrent planning

Select a prediction task

Choose one target with usable labels and a meaningful baseline.

Work to do

  • Define the intended user and scientific question.
  • Inspect assay context, label quality and likely confounders.
  • Set the evaluation metric before model selection.

Acceptance gate

A reviewer agrees the labels and metric support the claimed task.

Publish: Task brief and baseline specification.

PHASE 1 / 20% MILESTONEPlanned

Prepare trustworthy data

Prevent data leakage from making results look better than they are.

Work to do

  • Pin the dataset and record licensing and provenance.
  • Standardise identifiers and detect duplicate or conflicting records.
  • Define scaffold or temporal splits appropriate to the question.

Acceptance gate

The split and leakage checks are documented and reproducible.

Publish: Curated dataset manifest and split files.

PHASE 2 / 40% MILESTONEPlanned

Train the first baseline

Produce one complete experiment anyone can inspect.

Work to do

  • Start with a simple descriptor-based baseline.
  • Compare one established model under the same split.
  • Record configuration, seeds, environment and compute use.

Acceptance gate

A clean environment can rerun the experiment and recover the stated metrics.

Publish: Baseline notebook and experiment report.

PHASE 3 / 60% MILESTONEPlanned

Stress-test generalisation

Understand where predictions stop being reliable.

Work to do

  • Evaluate held-out chemical groups or later records.
  • Measure calibration and sensitivity to alternative splits.
  • Publish failure cases and uncertainty rather than only best scores.

Acceptance gate

The predefined benchmark is met without leakage; limitations are accepted by a reviewer.

Publish: Generalisation report and model card.

PHASE 4 / 80% MILESTONEPlanned

Pilot candidate review

Check whether ranked outputs help a real research workflow.

Work to do

  • Agree a bounded review task with a research partner.
  • Have experts assess candidates without relying on model rank alone.
  • Record useful leads, rejected outputs and missing evidence.

Acceptance gate

The partner can explain and reproduce each selected hypothesis and its limitations.

Publish: Pilot review and prioritisation feedback.

PHASE 5 / 100% MILESTONEPlanned

Release reproducible screening

Package the evaluated task as a stable research tool.

Work to do

  • Version models, datasets and experiment bundles together.
  • Document supported inputs, uncertainty and compute requirements.
  • Add regressions and upstream update monitoring.

Acceptance gate

Independent reruns pass and the release has a bounded scientific claim.

Publish: Product 1.0 discovery workspace.

Progress stays at 0% until the starting scope is accepted and subsequent milestone evidence is published. Phase 1–5 targets are 20%, 40%, 60%, 80% and 100%. These percentages track development, not treatment effectiveness.

WHAT THIS DEPENDS ON

Build with the right foundations.

  • Datasets & models library: ChEMBL or PubChem inputs and pinned tool versions.
  • RDKit and DeepChem: candidate software foundations, subject to task-level evaluation.
  • Scientific reviewer: assay interpretation and evaluation design.

Partner roles above are requirements. No partner participation or endorsement is claimed.

FUTURE INTERACTIVE PRODUCT

Discovery experiment workspace

Open after a baseline experiment can be independently reproduced within an agreed tolerance.

  1. Select a versioned benchmark.
  2. Inspect data quality and train-test separation.
  3. Run a baseline and compare a candidate model.
  4. Export rankings with uncertainty and experiment metadata.
View the portal specification
PUBLIC CHANGELOG

What changed. What is still planned.

Drug discovery engine

Phase 0–5 strategy and portal plan published

Published the module strategy, concrete phase deliverables, acceptance criteria and future portal workflow. This is a planning update; product completion remains 0%.

Entries are published with site updates. Planned work is labelled separately from completed product milestones.