eduKateSG · Why Science?
Let neighbouring tissue locations inform cell-type composition estimates—while keeping measured spots, predicted proportions and refined resolution clearly separated
Reading routes
Science learning becomes useful when a familiar object or observation is turned into a system of quantities, mechanisms and claim limits. This guide owns one applied evidence-reading job inside eduKateSG’s wider Science estate. It connects naturally to Why Science Spatial Transcriptomics Tissue Coordinates Gene Expression Evidence; Why Science Rctd Reference Cell Type Decomposition Spatial Transcriptomics Evidence; Why Science Spotlight Topic Modelling Cell Type Deconvolution Spatial Transcriptomics Evidence; Why Science Spagcn Graph Convolution Spatial Domains Histology Evidence; Education Hub; Singapore Secondary School Directory; Career Adulthood Hub. It also keeps current school and public claims traceable to visible primary sources: CARD primary study; CARD PubMed record; CARD full text; Official CARD repository; Official CARD overview; 2026 Singapore–Cambridge O-Level Biology syllabus; MOE G2/G3 Lower Secondary Science syllabus; 2026 MOE G2 Computing syllabus. The sources describe the scientific scope; this article translates that scope into a calm route for Primary Science, PSLE Science, Secondary Science, O-Level Science, STEM exploration, school choices and career pathways without inventing admission or employment outcomes.
CARD is a statistical method for estimating spatially varying cell-type composition from spatial transcriptomics. Its 2022 Nature Biotechnology paper uses a conditional autoregressive structure to share information between nearby locations, can operate with or without a matched single-cell reference, and can generate a higher-resolution predicted composition map. Spatial borrowing can stabilise noisy estimates, but it can also smooth real boundaries when its assumptions are wrong. Reference coverage, neighbourhood construction, spot quality, tissue edges and orthogonal evidence therefore remain central.
Inside this guide
1–12 · Foundations and models
- 1. Start with nearby locations that may be related
- 2. Frame composition as an estimation problem
- 3. Use a conditional autoregressive structure
- 4. Borrow information without erasing boundaries
- 5. Use a reference when it fits
- 6. Understand the reference-free option
- 7. Label refined maps as predictions
- 8. Build a tissue and reference manifest
- 9. Audit reference cell types
- 10. Draw the neighbourhood graph
- 11. Map quality before composition
- 12. Harmonise genes carefully
13–24 · Evidence, testing and applications
- 13. Estimate capacity before refining space
- 14. Save every spatial parameter
- 15. Practise with an invented CARD audit table
- 16. Read each proportion with context
- 17. Compare measured and refined layers
- 18. Inspect boundaries as a stress test
- 19. Treat rare populations cautiously
- 20. Keep specimen-level replication visible
- 21. Distinguish tissue absence from reference absence
- 22. Audit spatial-weight sensitivity
- 23. Audit reference sensitivity
- 24. Audit resolution refinement
25–36 · Learning, decisions and pathways
- 25. Validate outside the fitted signal
- 26. Compare non-spatial and alternative methods
- 27. Publish uncertainty and failure regions
- 28. Use disagreement to plan measurement
- 29. Build a classroom model before code
- 30. Start with Primary Science habits
- 31. Use PSLE Science to practise claim limits
- 32. Connect Secondary and O-Level Science
- 33. Let Computing support the science
- 34. Choose school opportunities by fit
- 35. See the connected career families
- 36. Finish with a bounded scientific claim
Section 1 of 36
1. Start with nearby locations that may be related
Tissue is organised, so neighbouring spatial locations often have similar cell-type mixtures. CARD uses that spatial structure while estimating composition. The opportunity is to stabilise noisy spots; the risk is smoothing a real edge or dispersed population. Spatial borrowing is therefore an assumption to test, not a guarantee of anatomical truth.
Section 2 of 36
2. Frame composition as an estimation problem
Observed spot counts combine signals from multiple cell types. CARD explains them with cell-type-specific expression profiles and spatially varying proportions. The proportions are inferred summaries. They do not identify exact cells or prove that a cell type caused a regional phenotype.
Section 3 of 36
3. Use a conditional autoregressive structure
A conditional autoregressive component links each location to its neighbours. This encourages related composition where the graph says locations are close. The graph’s definition, tissue holes and boundary handling therefore enter the biological interpretation.
Section 4 of 36
4. Borrow information without erasing boundaries
Spatial regularisation can reduce noise when nearby mixtures truly resemble one another. It can also blur sharp interfaces, infiltrating cells or small niches. Inspect original counts and independent images at boundaries rather than assuming the smoother map is more realistic.
Section 5 of 36
5. Use a reference when it fits
In reference-based mode, single-cell expression profiles inform the expected signal for each type. Tissue, condition, platform and annotation quality determine transferability. A reference missing an important state can shift its expression into the nearest available type.
Section 6 of 36
6. Understand the reference-free option
CARD can also analyse data without an external single-cell reference under its published framework. That flexibility reduces dependency on a matched reference but does not remove identifiability limits. Resulting components still require markers, anatomy and external validation.
Section 7 of 36
7. Label refined maps as predictions
CARD can generate a finer predicted composition surface than the measured spot grid. This is interpolation or model-based refinement, not new sequencing. Every refined pixel should remain traceable to the measured spots and assumptions that support it.
Section 8 of 36
8. Build a tissue and reference manifest
Record specimens, conditions, sections, platforms, coordinates, images, spot size, count layers, gene identifiers, reference donors and annotation versions. State which factors are confounded. Composition comparisons inherit every difference in this manifest.
Section 9 of 36
9. Audit reference cell types
Check marker coherence, incompatible markers, doublets, donor balance and rare types. Combine categories the spatial assay cannot separate. Publish the mapping from original labels to analysis labels so readers can reconstruct the decision.
Section 10 of 36
10. Draw the neighbourhood graph
Plot all graph edges on the tissue image, including near holes, tears and disconnected pieces. Report distance units, neighbour rule and boundary treatment. A spatial model cannot know that a short coordinate distance crosses empty tissue unless the graph forbids it.
Section 11 of 36
11. Map quality before composition
Display depth, detected genes, background, tissue coverage and image artefacts. Low-quality regions can look depleted for every type, while edge effects can masquerade as composition gradients. Quality covariates should be considered before biological naming.
Section 12 of 36
12. Harmonise genes carefully
Document gene intersections, symbol conversions, duplicates, filtering and normalisation across reference and spatial datasets. Hold out informative genes for evaluation when feasible. Do not validate a cell type only with the markers used to define its reference profile.
Section 13 of 36
13. Estimate capacity before refining space
A refined map implies a spatial support and cell-density assumption. Connect its grid to measured spot dimensions and plausible tissue occupancy. More pixels do not mean more observed cells.
Section 14 of 36
14. Save every spatial parameter
Archive neighbourhoods, priors, initialisation, optimisation settings, resolution parameters, seeds, software version and input hashes. Keep measured and refined coordinates in separate fields. A future audit should be able to reproduce both the fit and the display.
Section 15 of 36
15. Practise with an invented CARD audit table
This fictional table teaches reasoning, not benchmark performance.
| Region | Composition stability | Graph sensitivity | Measured markers | Refined-map support | First reading |
|---|---|---|---|---|---|
| Core | high | low | agrees | strong | supported |
| Edge | medium | high | mixed | weak | smoothing risk |
| Niche | low | medium | agrees | local | inspect reference |
| Gap | low | high | absent | apparent | artefact risk |
Section 16 of 36
16. Read each proportion with context
A proportion map is meaningful relative to other types, spots and specimens. Check compositional constraints: increasing one share necessarily affects others. Report uncertainty or sensitivity and avoid ranking tiny differences that do not survive plausible settings.
Section 17 of 36
17. Compare measured and refined layers
Place measured spot estimates beside the refined prediction using aligned scales. Mark areas where the refined layer adds detail not directly observed. Readers should never have to guess which pixels were sequenced.
Section 18 of 36
18. Inspect boundaries as a stress test
Tissue interfaces are where conditional autoregression is most challenged. Quantify gradient width, compare alternative neighbour strengths and use morphology. A persistent sharp boundary is stronger evidence than one appearing only after a convenient setting.
Section 19 of 36
19. Treat rare populations cautiously
Rare cell types may have weak, overlapping signals. Spatial smoothing can spread a small estimate beyond plausible niches. Confirm rare-type maps with specific held-out markers or imaging and report detection limits.
Section 20 of 36
20. Keep specimen-level replication visible
Summarise composition within each biological specimen before pooling. Spot-level sample size does not replace donor-level replication. Show when a pattern is shared, absent or reversed across specimens.
Section 21 of 36
21. Distinguish tissue absence from reference absence
A zero estimate may mean true absence, insufficient marker counts, a mismatched profile or competition with another type. Use positive controls and spike-in simulations to understand what zero can mean under the data quality.
Section 22 of 36
22. Audit spatial-weight sensitivity
Repeat fits over a prespecified range of spatial strengths and graph definitions. Track stable cores, moving boundaries and dispersed populations. Interpret only the level of detail that survives reasonable alternatives.
Section 23 of 36
23. Audit reference sensitivity
Use alternative donor subsets, annotation schemes and, where appropriate, the reference-free route. Compare which cell-type patterns persist. A conclusion that disappears with one donor removal is conditional evidence.
Section 24 of 36
24. Audit resolution refinement
Change the refinement grid within biologically sensible limits and compare integrated abundance, boundary placement and held-out agreement. Fine predictions should not create more total biological material or unsupported islands.
Section 25 of 36
25. Validate outside the fitted signal
Use reserved genes, protein stains, morphology or expert annotations. Include obvious positives, clear negatives and difficult borders. Orthogonal validation should test where the spatial prior could be most misleading.
Section 26 of 36
26. Compare non-spatial and alternative methods
Fit a matched non-spatial composition baseline and another published spatial method. CARD adds value when spatial information improves held-out evidence without erasing known boundaries or creating implausible spread.
Section 27 of 36
27. Publish uncertainty and failure regions
Show unsupported types, graph-sensitive edges, poor-quality spots and unstable refined pixels. A transparent map can remain useful even when some regions are unresolved.
Section 28 of 36
28. Use disagreement to plan measurement
Target graph-sensitive boundaries or rare-type niches with extra markers, imaging or higher resolution. The best next experiment is often where spatial and non-spatial estimates disagree for a scientifically interpretable reason.
Section 29 of 36
29. Build a classroom model before code
Let students estimate coloured-bead mixtures in neighbouring cups, first independently and then with a rule that neighbours should resemble one another. Discuss when the rule helps and when it hides a real boundary. Keep observations, transformations, assumptions and conclusions in separate columns. The aim is not to imitate specialist software but to make each evidence hand-off visible and testable.
Section 30 of 36
30. Start with Primary Science habits
Primary Science already teaches the habits beneath CARD: observe carefully, compare fairly, record consistently and keep conclusions within the experiment. A map of plants, light or water can show why location changes interpretation without advanced mathematics.
Section 31 of 36
31. Use PSLE Science to practise claim limits
PSLE Science connects observations to processes and explanations. CARD adds a modern reminder that a statistical pattern is not automatically a cause. Ask what changed, what was measured, what remained uncontrolled and which new observation would separate competing explanations.
Section 32 of 36
32. Connect Secondary and O-Level Science
Secondary Science and O-Level Biology develop cells, organisation, variation, experimental design and evaluation. CARD is a contemporary case where molecular measurements, tissue position and computing meet. The learning goal is disciplined reasoning, not memorising software.
Section 33 of 36
33. Let Computing support the science
Computing contributes data structures, algorithms, optimisation, visualisation and reproducibility. Biology supplies the specimen, mechanism and independent validation. CARD shows why correct code is necessary but insufficient: an algorithm can run perfectly on mislabelled, confounded or incomplete data.
Section 34 of 36
34. Choose school opportunities by fit
Families should verify Biology, Computing, mathematics, statistics projects and research mentorship on current official school and MOE pages. A mention of genomics, AI or data science does not guarantee programme depth, admission or a career outcome. Look for sustained inquiry, careful teaching, accessible mentoring and time to explain evidence.
Section 35 of 36
35. See the connected career families
Reasoning used in CARD appears in spatial statistics, genomics, epidemiology, pathology, computational biology and biomedical data science. Routes can pass through polytechnic, junior college, university or continuing education with different blends of biology, mathematics, statistics, computing and communication. Check current course requirements directly; one project cannot guarantee entry or employment.
Section 36 of 36
36. Finish with a bounded scientific claim
A defensible conclusion states exactly what CARD estimated and under which inputs, settings and specimens. CARD estimates spatial cell-type composition and may predict refined maps; it does not directly observe individual cells, guarantee a complete reference or turn interpolated pixels into new measurements. Report uncertainty, sensitivity, replication, negative results and orthogonal validation together. That boundary is not weakness; it makes the result testable and reusable.
A rigorous CARD project begins by drawing the neighbourhood graph over the untouched tissue image. Colour edges by length and mark edges crossing holes, tears, folds or disconnected tissue. Compute degree and neighbour distance by region. This one diagnostic often reveals whether spatial borrowing means biological proximity or merely coordinate convenience. Keep the graph as a first-class research object throughout the analysis.
Reference-based and reference-free analyses answer related but not identical questions. In reference-based work, labels and profiles anchor cell types, so audit donor balance and transferability. In reference-free work, components require post-hoc interpretation and may not correspond one-to-one with canonical types. Do not present agreement between these routes as automatic; define a matching rule and show unmatched components.
Composition lives on a simplex: shares are relative and sum to a total. If one estimated type rises, at least one other share must fall even when its absolute abundance does not. Pair proportions with tissue area, cell density or imaging counts when the biological question concerns numbers. Use language such as ‘higher estimated share’ unless absolute abundance has independent support.
Boundary validation is central because spatial regularisation is strongest where local information is shared. Select known sharp interfaces, smooth gradients, small islands and dispersed populations before tuning. Measure boundary width and location across spatial strengths. The best setting should not merely maximise smoothness; it should preserve the different geometries expected by independent tissue evidence.
Refined-resolution maps need an explicit provenance overlay. Show the measured spot centres, the prediction grid and a distance-to-measurement map. A fine pixel far from an informative spot is supported differently from a fine pixel surrounded by consistent measurements. If the refinement predicts sub-spot islands, require especially strong independent validation before describing them as tissue niches.
Construct simulations that resemble the real tissue rather than only ideal circles. Include gradients, abrupt borders, rare types, variable depth, missing reference types and spatially structured noise. Measure abundance error, boundary displacement and false islands. Simulation cannot prove truth in the specimen, but it can expose which visual patterns the chosen graph and settings manufacture under known conditions.
Use leave-one-specimen-out analysis when enough specimens exist. Fit or define reference profiles without one specimen, then evaluate the held-out tissue using a declared procedure. Compare major regions, rare populations and edge behaviour. This is a stronger test of transportability than pooling all spots and randomly withholding a few locations from every specimen.
Negative controls should attack the spatial assumption. Shuffle coordinates within a tissue mask, permute neighbour links while preserving degree, or use a spatially implausible graph as appropriate. If the same polished composition map survives a broken geography, the result may be driven mainly by reference markers or global abundance rather than the intended spatial information.
Compare measured-marker evidence with the full multigene estimate. A cell type with a specific marker may provide a useful sanity check, but markers can vary by state and platform. Use several positive markers, incompatible markers and morphology. Reserve at least some evidence that did not determine the reference profiles or tune the spatial strength.
Did you know? Spatial smoothing can increase numerical stability while reducing biological accuracy at a narrow boundary. Stability and truth are related only when the model’s neighbourhood assumptions match the tissue. That is why edge-focused validation is as important as average reconstruction error.
Publish a refined-map audit table containing grid size, source spots, distance to measured locations, local uncertainty or sensitivity, boundary stability and orthogonal support. Include regions where refinement adds no reliable information. A map that says ‘unsupported here’ is more useful than one that fills every pixel with a confident colour.
Reproducibility materials should contain raw counts, coordinates, images, masks, reference metadata, annotations, shared genes, normalisation, graph edges, model settings, abundance tables, refined grids, seeds, versions, sensitivity results, validation images and checksums. Keep measured and predicted coordinates in separate files or clearly labelled columns so downstream users cannot silently confuse them.
A bounded conclusion might say: ‘CARD estimated specimen-reproducible cell-type composition patterns that remained stable across plausible reference and spatial-graph choices and agreed with held-out markers and imaging; a refined map was used as a prediction layer.’ It should not call the refined pixels new measurements, claim exact cell locations or treat spatial smoothness as proof of anatomy.
Finally, keep a claim ledger for every named tissue region. Record the measured counts, inferred proportions, graph settings, reference cells, refined predictions, orthogonal evidence, specimens represented and strongest alternative explanation. Review the ledger before drafting captions. If a claim depends on one spatial weight or one reference donor, narrow the wording. This small discipline prevents the polished refined map from becoming more authoritative than the data supporting it.
Contents · Previous section · Continue to the Science Learning Hub
