eduKateSG · Why Science?
Estimate how perturbation direction changes across cell types and doses—without turning latent interpolation into a safety or efficacy claim
Reading routes
Science learning becomes useful when a familiar object or observation is turned into a system of quantities, mechanisms and claim limits. This guide owns one applied evidence-reading job inside eduKateSG’s wider Science estate. It connects naturally to Why Science Scgen Latent Space Arithmetic Cross Condition Perturbation Prediction; Why Science Cpa Compositional Perturbations Dose Response Prediction; Why Science Prescient Potential Landscapes Physical Time Cell Trajectories; Why Science Scnode Neural Odes Missing Timepoint Single Cell Prediction; Education Hub; Singapore Secondary School Directory; Career Adulthood Hub; Science Learning Hub. It also keeps current school and public claims traceable to visible primary sources: scVIDR primary study; scVIDR PubMed record; Official scVIDR repository; 2026 Singapore–Cambridge O-Level Biology syllabus; MOE G2/G3 Lower Secondary Science syllabus; 2026 MOE G2 Computing syllabus. The sources describe the scientific scope; this article translates that scope into a calm route for Primary Science, PSLE Science, Secondary Science, O-Level Science, STEM exploration, school choices and career pathways without inventing admission or employment outcomes.
scVIDR, presented in Patterns in 2023, uses a variational autoencoder and a regression model to estimate cell-type-specific perturbation directions, including responses for a held-out cell type. Its multi-dose workflow interpolates along estimated latent perturbation vectors, and the paper evaluated settings including mouse-liver TCDD doses, sci-Plex chemical responses and cross-species LPS transfer. This is a lively science-process-skills example of dose, controls and prediction; it does not turn gene-expression forecasts into recommended exposure, treatment, safety or efficacy.
Inside this guide
1–12 · Foundations and models
- 1. Begin with cell-type-specific response
- 2. Encode expression
- 3. Calculate perturbation vectors
- 4. Regress across cell types
- 5. Model multiple doses
- 6. Understand pseudo-dose
- 7. Read the 2023 evidence in scope
- 8. Define the dose question
- 9. Build an exposure manifest
- 10. Plot design coverage
- 11. Hold out the target response
- 12. Protect specimens and plates
13–24 · Evidence, testing and applications
- 13. Choose genes transparently
- 14. Predeclare interpolation
- 15. Practise with a fictional scVIDR table
- 16. Separate latent and gene curves
- 17. Audit regression fit
- 18. Audit the dose function
- 19. Check target support
- 20. Run seed ensembles
- 21. Use strong baselines
- 22. Audit rare subtypes
- 23. Construct dose-aware nulls
- 24. Test cross-species transfer carefully
25–36 · Learning, decisions and pathways
- 25. Compare with CPA
- 26. Build a dose-response evidence card
- 27. Keep toxicology separate
- 28. Design the next dose experiment
- 29. Connect scVIDR to school science
- 30. Try a dose-curve activity
- 31. Connect to mathematics
- 32. Connect to computing and AI literacy
- 33. Connect to school and career choices
- 34. Create a family dose habit
- 35. Use precise scVIDR language
- 36. Finish with a vector-to-dose bundle
Section 1 of 36
1. Begin with cell-type-specific response
Different cell types may react differently to the same chemical. scVIDR estimates how perturbation direction varies across known types and predicts the missing direction for a held-out type.
scVIDR uses a variational representation and regression across known cell-type perturbation vectors to estimate the direction for a held-out type.
Archive cell types, compounds, dose units, times, specimens, preprocessing, regression settings, seeds, checkpoints and repository version.
Section 2 of 36
2. Encode expression
A variational autoencoder compresses single-cell expression into a latent distribution. The decoder converts predicted latent states back into gene space.
Multiple-dose prediction interpolates along an estimated latent vector, so functional form and dose coverage become part of every result.
Hold out the complete treated target cell type or drug–cell combination and never use it for feature selection or tuning.
Section 3 of 36
3. Calculate perturbation vectors
Within known cell types, treated and control latent means define response directions. Those vectors become examples for a regression model.
A pseudo-dose orders cells by projected sensitivity in the fitted latent model; it is not administered exposure or a clinical measurement.
Compare no change, highest-dose vector scaling, scGen-style shared vector and simple dose regression on the same splits.
Section 4 of 36
4. Regress across cell types
Cell-type latent positions help predict a response vector for the unseen type. This extends one global vector with target-specific structure.
The 2023 paper evaluated TCDD liver data, IFN-beta, cross-study and cross-species settings plus sci-Plex drug-response tasks.
Report gene-level curves, distribution distances, dose-wise errors, subgroup errors, seed spread and extrapolation warnings.
Section 5 of 36
5. Model multiple doses
For dose-response tasks, the method interpolates along estimated latent perturbation vectors, using a log-linear dose relationship in the published workflow.
Response prediction across cell types is easiest when shared biology exists; target-specific pathways and sparse doses remain hard cases.
Build a dose card with observed range, interpolation status, target support, baselines, uncertainty and independent validation.
Section 6 of 36
6. Understand pseudo-dose
Projecting cells onto the perturbation direction can order them by modelled sensitivity. The resulting value is a latent summary, not literal exposure.
Dose–response expression forecasts can guide experiments but cannot establish safe exposure, therapeutic efficacy or human dosing.
Say predicted expression response and model-derived pseudo-dose; never call either safe dose, effective dose or measured exposure.
Section 7 of 36
7. Read the 2023 evidence in scope
The primary paper evaluated single- and multiple-dose tasks, cross-study and cross-species transfer using several published single-cell datasets.
scVIDR uses a variational representation and regression across known cell-type perturbation vectors to estimate the direction for a held-out type.
Archive cell types, compounds, dose units, times, specimens, preprocessing, regression settings, seeds, checkpoints and repository version.
Section 8 of 36
8. Define the dose question
Name chemical, target cell type, administered dose, units, duration and predicted genes. Mark which dimensions are held out.
Multiple-dose prediction interpolates along an estimated latent vector, so functional form and dose coverage become part of every result.
Hold out the complete treated target cell type or drug–cell combination and never use it for feature selection or tuning.
Section 9 of 36
9. Build an exposure manifest
Record specimen, cell type, dose, vehicle, route, time, batch, assay and cell count. Harmonise units and controls.
A pseudo-dose orders cells by projected sensitivity in the fitted latent model; it is not administered exposure or a clinical measurement.
Compare no change, highest-dose vector scaling, scGen-style shared vector and simple dose regression on the same splits.
Section 10 of 36
10. Plot design coverage
Show every observed cell-type-by-dose combination. A dense cell count cannot repair a missing biological replicate.
The 2023 paper evaluated TCDD liver data, IFN-beta, cross-study and cross-species settings plus sci-Plex drug-response tasks.
Report gene-level curves, distribution distances, dose-wise errors, subgroup errors, seed spread and extrapolation warnings.
Section 11 of 36
11. Hold out the target response
Keep treated target cells completely sealed during training and tuning. Their controls may be used only as declared by the task.
Response prediction across cell types is easiest when shared biology exists; target-specific pathways and sparse doses remain hard cases.
Build a dose card with observed range, interpolation status, target support, baselines, uncertainty and independent validation.
Section 12 of 36
12. Protect specimens and plates
Split animals, donors, cultures or plates, not random cells. Technical neighbours can exaggerate performance.
Dose–response expression forecasts can guide experiments but cannot establish safe exposure, therapeutic efficacy or human dosing.
Say predicted expression response and model-derived pseudo-dose; never call either safe dose, effective dose or measured exposure.
Section 13 of 36
13. Choose genes transparently
Document normalisation, highly variable genes and differential-gene rules. Evaluate biologically important genes outside the optimisation set.
scVIDR uses a variational representation and regression across known cell-type perturbation vectors to estimate the direction for a held-out type.
Archive cell types, compounds, dose units, times, specimens, preprocessing, regression settings, seeds, checkpoints and repository version.
Section 14 of 36
14. Predeclare interpolation
A dose between measured endpoints differs from one beyond the range. Mark every prediction as interpolation or extrapolation.
Multiple-dose prediction interpolates along an estimated latent vector, so functional form and dose coverage become part of every result.
Hold out the complete treated target cell type or drug–cell combination and never use it for feature selection or tuning.
Section 15 of 36
15. Practise with a fictional scVIDR table
This classroom table is invented and is not output from the scVIDR paper.
| Fictional dose | Prediction error | scGen error | Status | Next check |
|---|---|---|---|---|
| 0.3 units | 0.17 | 0.24 | Interpolation | New animal |
| 3 units | 0.22 | 0.26 | Interpolation | Marker assay |
| 60 units | 0.71 | 0.66 | Extrapolation | Do not infer |
A pseudo-dose orders cells by projected sensitivity in the fitted latent model; it is not administered exposure or a clinical measurement.
Compare no change, highest-dose vector scaling, scGen-style shared vector and simple dose regression on the same splits.
Section 16 of 36
16. Separate latent and gene curves
A smooth movement in latent space can decode into inaccurate or non-monotonic gene responses. Inspect both levels.
The 2023 paper evaluated TCDD liver data, IFN-beta, cross-study and cross-species settings plus sci-Plex drug-response tasks.
Report gene-level curves, distribution distances, dose-wise errors, subgroup errors, seed spread and extrapolation warnings.
Section 17 of 36
17. Audit regression fit
Plot known cell-type vectors and leave each type out in turn. This tests whether the regression captures a reusable relationship.
Response prediction across cell types is easiest when shared biology exists; target-specific pathways and sparse doses remain hard cases.
Build a dose card with observed range, interpolation status, target support, baselines, uncertainty and independent validation.
Section 18 of 36
18. Audit the dose function
Compare log-linear interpolation with plausible alternatives using training evidence only. Do not choose a curve after seeing final targets.
Dose–response expression forecasts can guide experiments but cannot establish safe exposure, therapeutic efficacy or human dosing.
Say predicted expression response and model-derived pseudo-dose; never call either safe dose, effective dose or measured exposure.
Section 19 of 36
19. Check target support
Measure how close the held-out cell type is to training types in the representation. Distant targets deserve wider uncertainty.
scVIDR uses a variational representation and regression across known cell-type perturbation vectors to estimate the direction for a held-out type.
Archive cell types, compounds, dose units, times, specimens, preprocessing, regression settings, seeds, checkpoints and repository version.
Section 20 of 36
20. Run seed ensembles
Retrain autoencoders and regressions. Report how vector angle, gene curves and pseudo-dose order change.
Multiple-dose prediction interpolates along an estimated latent vector, so functional form and dose coverage become part of every result.
Hold out the complete treated target cell type or drug–cell combination and never use it for feature selection or tuning.
Section 21 of 36
21. Use strong baselines
Compare no change, shared mean vector, scGen-style arithmetic and simple dose regression. Complexity must improve untouched evidence.
A pseudo-dose orders cells by projected sensitivity in the fitted latent model; it is not administered exposure or a clinical measurement.
Compare no change, highest-dose vector scaling, scGen-style shared vector and simple dose regression on the same splits.
Section 22 of 36
22. Audit rare subtypes
A cell-type average can hide responder and non-responder groups. Compare distributions and subtype proportions.
The 2023 paper evaluated TCDD liver data, IFN-beta, cross-study and cross-species settings plus sci-Plex drug-response tasks.
Report gene-level curves, distribution distances, dose-wise errors, subgroup errors, seed spread and extrapolation warnings.
Section 23 of 36
23. Construct dose-aware nulls
Shuffle dose within specimens or simulate flat responses with matching noise. The workflow should not invent a graded curve.
Response prediction across cell types is easiest when shared biology exists; target-specific pathways and sparse doses remain hard cases.
Build a dose card with observed range, interpolation status, target support, baselines, uncertainty and independent validation.
Section 24 of 36
24. Test cross-species transfer carefully
Orthologue mapping and physiology differ. Report mapping coverage and failures rather than calling species interchangeable.
Dose–response expression forecasts can guide experiments but cannot establish safe exposure, therapeutic efficacy or human dosing.
Say predicted expression response and model-derived pseudo-dose; never call either safe dose, effective dose or measured exposure.
Section 25 of 36
25. Compare with CPA
CPA learns compositional perturbation and dose embeddings; scVIDR regresses cell-type-specific vectors and interpolates doses. Use distinct tests.
scVIDR uses a variational representation and regression across known cell-type perturbation vectors to estimate the direction for a held-out type.
Archive cell types, compounds, dose units, times, specimens, preprocessing, regression settings, seeds, checkpoints and repository version.
Section 26 of 36
26. Build a dose-response evidence card
Record target, observed dose range, regression support, holdout unit, baselines, gene curves, seed spread and validation.
Multiple-dose prediction interpolates along an estimated latent vector, so functional form and dose coverage become part of every result.
Hold out the complete treated target cell type or drug–cell combination and never use it for feature selection or tuning.
Section 27 of 36
27. Keep toxicology separate
Expression change may accompany stress or toxicity. Viability, pathology and organism-level outcomes require dedicated measurements.
A pseudo-dose orders cells by projected sensitivity in the fitted latent model; it is not administered exposure or a clinical measurement.
Compare no change, highest-dose vector scaling, scGen-style shared vector and simple dose regression on the same splits.
Section 28 of 36
28. Design the next dose experiment
Measure doses where models disagree or curves bend, with matched controls, replication and orthogonal outcomes.
The 2023 paper evaluated TCDD liver data, IFN-beta, cross-study and cross-species settings plus sci-Plex drug-response tasks.
Report gene-level curves, distribution distances, dose-wise errors, subgroup errors, seed spread and extrapolation warnings.
Section 29 of 36
29. Connect scVIDR to school science
Dose, response, controlled variables and graph reading connect naturally to PSLE Science and secondary science at different depths.
Response prediction across cell types is easiest when shared biology exists; target-specific pathways and sparse doses remain hard cases.
Build a dose card with observed range, interpolation status, target support, baselines, uncertainty and independent validation.
Section 30 of 36
30. Try a dose-curve activity
Use a safe classroom system, measure several levels, hide one, fit alternative curves and reveal which prediction was justified.
Dose–response expression forecasts can guide experiments but cannot establish safe exposure, therapeutic efficacy or human dosing.
Say predicted expression response and model-derived pseudo-dose; never call either safe dose, effective dose or measured exposure.
Section 31 of 36
31. Connect to mathematics
Vectors, regression, logarithms, interpolation and error metrics turn response patterns into testable quantities.
scVIDR uses a variational representation and regression across known cell-type perturbation vectors to estimate the direction for a held-out type.
Archive cell types, compounds, dose units, times, specimens, preprocessing, regression settings, seeds, checkpoints and repository version.
Section 32 of 36
32. Connect to computing and AI literacy
A VAE and regression layer can fail differently. Honest holdouts show whether representation and transfer both work.
Multiple-dose prediction interpolates along an estimated latent vector, so functional form and dose coverage become part of every result.
Hold out the complete treated target cell type or drug–cell combination and never use it for feature selection or tuning.
Section 33 of 36
33. Connect to school and career choices
Learners interested in dose and data can explore biology, chemistry, computing, statistics, toxicology or pharmacology.
A pseudo-dose orders cells by projected sensitivity in the fitted latent model; it is not administered exposure or a clinical measurement.
Compare no change, highest-dose vector scaling, scGen-style shared vector and simple dose regression on the same splits.
Section 34 of 36
34. Create a family dose habit
When AI predicts a response, ask the measured dose range, units, target cell type and whether the requested point is extrapolated.
The 2023 paper evaluated TCDD liver data, IFN-beta, cross-study and cross-species settings plus sci-Plex drug-response tasks.
Report gene-level curves, distribution distances, dose-wise errors, subgroup errors, seed spread and extrapolation warnings.
Section 35 of 36
35. Use precise scVIDR language
Say cell-type perturbation vector, latent interpolation and model-derived pseudo-dose. Avoid safe, effective or prescribed dose.
Response prediction across cell types is easiest when shared biology exists; target-specific pathways and sparse doses remain hard cases.
Build a dose card with observed range, interpolation status, target support, baselines, uncertainty and independent validation.
Section 36 of 36
36. Finish with a vector-to-dose bundle
Deliver manifests, held-out targets, regression audits, dose-function sensitivity, baselines, seed ensembles, cards and new measurements.
Dose–response expression forecasts can guide experiments but cannot establish safe exposure, therapeutic efficacy or human dosing.
Say predicted expression response and model-derived pseudo-dose; never call either safe dose, effective dose or measured exposure.
Did you know? A mathematically smooth dose curve can pass through regions with no measured biology. Plot every observed dose and mark interpolation versus extrapolation.
A pseudo-dose is useful for ordering modelled sensitivity but should never be presented as exposure received by a cell, animal or person.
The best follow-up samples doses where curve families disagree, because that evidence can reject assumptions instead of merely adding points.
Report the number of independent animals, donors, cultures or plates at each dose. Thousands of cells from one biological unit cannot establish a general dose response.
Where the task concerns human health, link expression predictions only to appropriate experimental planning; never turn a latent curve into personal advice.
Contents · Previous section · Continue to the Science Learning Hub
