eduKateSG · Why Science?
Use neighbouring expression to discover spatial clusters and computationally enhance capture spots into subspots—without confusing Bayesian imputation with molecules or cells directly measured at finer resolution
Reading routes
Science learning becomes useful when a familiar object or observation is turned into a system of quantities, mechanisms and claim limits. This guide owns one applied evidence-reading job inside eduKateSG’s wider Science estate. It connects naturally to Why Science Spatial Transcriptomics Tissue Coordinates Gene Expression Evidence; Why Science Single Cell Rna Sequencing Barcodes Transcriptome Heterogeneity Evidence; Why Science Spagcn Graph Convolution Spatial Domains Histology Evidence; Why Science Squidpy Spatial Neighbourhoods Omics Image Analysis Evidence; Education Hub; Singapore Secondary School Directory; Career Adulthood Hub. It also keeps current school and public claims traceable to visible primary sources: BayesSpace primary study; BayesSpace PubMed record; BayesSpace full text; Bioconductor BayesSpace package; Official BayesSpace repository; 2026 Singapore–Cambridge O-Level Biology syllabus; 2026 MOE G2 Computing syllabus. The sources describe the scientific scope; this article translates that scope into a calm route for Primary Science, PSLE Science, Secondary Science, O-Level Science, STEM exploration, school choices and career pathways without inventing admission or employment outcomes.
BayesSpace is a Bayesian method for spatial clustering and resolution enhancement in spatial transcriptomics. The 2021 Nature Biotechnology paper models a low-dimensional representation of expression while using neighbourhood information to encourage spatially coherent clusters. It can then computationally enhance spots into subspots and impute expression or composition at that finer grid. This is a useful statistical reconstruction, not a new high-resolution measurement. Cluster number, neighbourhood structure, spatial smoothing, Markov chain Monte Carlo convergence, edge effects and orthogonal tissue evidence define how far the interpretation may travel.
Inside this guide
1–12 · Foundations and models
- 1. Separate measured spots from enhanced subspots
- 2. Represent expression in fewer dimensions
- 3. Use spatial neighbours as prior information
- 4. Cluster within a Bayesian model
- 5. Sample with Markov chain Monte Carlo
- 6. Enhance resolution computationally
- 7. Keep the primary demonstrations specific
- 8. Quality-control the expression matrix
- 9. Verify coordinates and neighbourhoods
- 10. Choose informative features
- 11. Choose the number of clusters
- 12. Specify spatial priors
13–24 · Evidence, testing and applications
- 13. Run multiple chains and seeds
- 14. Map uncertainty, not only labels
- 15. Practise with an invented clustering table
- 16. Interpret spatial clusters cautiously
- 17. Interpret enhanced subspots cautiously
- 18. Impute expression with clear labels
- 19. Estimate composition only with supporting models
- 20. Replicate domains across specimens
- 21. Validate at the claimed resolution
- 22. Challenge over-smoothing
- 23. Challenge cluster-number dependence
- 24. Challenge edges and tissue gaps
25–36 · Learning, decisions and pathways
- 25. Challenge MCMC convergence
- 26. State that imputation is not measurement
- 27. Compare BayesSpace with SpaGCN
- 28. Keep negative spatial claims bounded
- 29. Try a classroom model
- 30. Start the habit in Primary Science
- 31. Strengthen PSLE Science reasoning
- 32. Connect Secondary and O-Level Science
- 33. Use Computing as a partner
- 34. Choose schools by fit, not a single buzzword
- 35. See the career pathways
- 36. Finish with a bounded claim
Section 1 of 36
1. Separate measured spots from enhanced subspots
Spatial assays measure molecules at physical capture spots. BayesSpace can subdivide those spots computationally and estimate finer patterns, but the subspots are not newly measured locations. This distinction is the article’s anchor: enhanced resolution is inference supported by neighbours and expression, not additional molecules collected from tissue.
Section 2 of 36
2. Represent expression in fewer dimensions
The method works with a low-dimensional representation, commonly principal components, that summarises gene-expression variation. This reduces noise and computation. It also means feature selection, normalisation and component number determine which biological signals remain available for clustering.
Section 3 of 36
3. Use spatial neighbours as prior information
Adjacent spots often share tissue structure. BayesSpace encodes that expectation so neighbouring labels tend to be coherent. The prior can recover weak boundaries, yet it can also smooth across real sharp transitions or small rare regions if given too much influence.
Section 4 of 36
4. Cluster within a Bayesian model
Bayesian spatial clustering combines expression evidence with neighbourhood structure and uncertainty. The fitted labels describe domains under the chosen model and number of clusters. They are not automatically cell types, anatomical diagnoses or developmental lineages.
Section 5 of 36
5. Sample with Markov chain Monte Carlo
BayesSpace uses Markov chain Monte Carlo to explore model states. Chains require enough iterations, burn-in and diagnostic checking. A returned label map is not proof of convergence; repeated chains and stable summaries are essential before interpreting subtle boundaries.
Section 6 of 36
6. Enhance resolution computationally
For platforms such as Visium, the method can divide spots into subspots and infer finer patterns from the surrounding evidence. Enhancement can sharpen known structure, but several fine arrangements may explain the same coarse measurements. Validation must match the claimed spatial scale.
Section 7 of 36
7. Keep the primary demonstrations specific
The 2021 Nature Biotechnology paper introduced Bayesian spatial clustering and resolution enhancement and evaluated the method on spatial transcriptomics datasets. Those results show useful examples of coherent domains and enhanced patterns. Every new tissue still needs cluster, prior, convergence and validation checks.
Section 8 of 36
8. Quality-control the expression matrix
Inspect counts, detected genes, mitochondrial signal, tissue coverage, background and batch. Technical gradients can become spatially smooth clusters. Plot quality metrics beside inferred domains and prespecify exclusions so damaged tissue does not define biology.
Section 9 of 36
9. Verify coordinates and neighbourhoods
Check coordinate units, grid geometry, tissue holes, edges, rotation and missing spots. Neighbour relations should respect the physical array and tissue mask. Connecting across a tear or empty lumen can pull unrelated regions together.
Section 10 of 36
10. Choose informative features
Select genes and principal components with a documented rationale. Too few components may erase subtle biology; too many can carry noise and batch. Compare plausible choices and report which boundaries persist. Feature stability is part of evidence, not merely preprocessing.
Section 11 of 36
11. Choose the number of clusters
The requested cluster count strongly shapes the map. Examine a defensible range, hierarchy, stability and external landmarks. A larger number always creates more regions, but those extra colours do not automatically represent new tissue compartments.
Section 12 of 36
12. Specify spatial priors
Report neighbourhood definition and parameters controlling spatial coherence. Perform sensitivity analyses with weaker and stronger smoothing. If a region appears only under a narrow prior, it should remain provisional until supported by independent markers or morphology.
Section 13 of 36
13. Run multiple chains and seeds
Use several chains or runs, inspect traces or relevant diagnostics and compare label agreement. Save configurations and seeds. Consensus can identify stable cores, while disagreement may reveal poor mixing, weak data or a genuine transition zone.
Section 14 of 36
14. Map uncertainty, not only labels
Posterior summaries or repeated-run stability should accompany the hard cluster map. Edge spots and low-count regions often carry greater uncertainty. A confidence layer helps readers distinguish a stable domain centre from an attractive but fragile boundary.
Section 15 of 36
15. Practise with an invented clustering table
This fictional table is for interpretation only.
| Region | Expression evidence | Neighbour agreement | Run stability | First reading |
|---|---|---|---|---|
| R1 | strong | 94% | 96% | stable core |
| R2 | strong | 51% | 68% | possible edge |
| R3 | weak | 89% | 54% | prior-driven |
| R4 | weak | 38% | 31% | unresolved |
Section 16 of 36
16. Interpret spatial clusters cautiously
A cluster can correspond to a tissue layer, tumour compartment, compositional zone or technical region. Describe marker genes, morphology and sample context before naming it. Avoid using cell-type language unless cell composition is independently established.
Section 17 of 36
17. Interpret enhanced subspots cautiously
Subspot patterns can suggest where fine structure lies inside a capture spot. They should be presented with the original spot measurements and sensitivity maps. A sharply coloured subspot is a model allocation, not direct subcellular or single-cell observation.
Section 18 of 36
18. Impute expression with clear labels
BayesSpace can estimate gene expression at enhanced resolution. Label these values as imputed or enhanced in figures and files. Validate predictions for held-out or orthogonally measured genes before using them to support a novel boundary.
Section 19 of 36
19. Estimate composition only with supporting models
Cell-type composition can be analysed alongside enhanced structure, but each extra model adds assumptions. Propagate uncertainty rather than treating one inferred matrix as raw truth for the next step. Validate key populations with independent markers.
Section 20 of 36
20. Replicate domains across specimens
A domain recurring across independent sections and subjects is stronger than a cluster with thousands of correlated spots in one sample. Align only comparable anatomy, show all specimens and use specimen-level statistics for generalisation.
Section 21 of 36
21. Validate at the claimed resolution
Broad domains can be checked with histology and regional markers. Fine subspot claims require higher-resolution imaging, segmentation or another spatial assay. Validation at the original spot scale cannot prove a within-spot arrangement.
Section 22 of 36
22. Challenge over-smoothing
Insert synthetic small regions or examine known sharp boundaries to see whether the prior erases them. Compare weaker spatial influence and expression-only clustering. A smooth map can be visually appealing while missing biologically important islands.
Section 23 of 36
23. Challenge cluster-number dependence
Track how domains split or merge across candidate cluster counts. Report stable parent regions and uncertain subdivisions. Selecting the count that best matches a desired picture without independent evidence is circular.
Section 24 of 36
24. Challenge edges and tissue gaps
Spots at edges have fewer neighbours, and holes can distort local support. Map degree, uncertainty and enhanced assignments near gaps. Use tissue masks so the model does not borrow evidence through empty space.
Section 25 of 36
25. Challenge MCMC convergence
Longer runs, multiple chains and different starting states should give compatible broad conclusions. If labels switch or traces drift, increase sampling or simplify the claim. Computation ending successfully is not the same as statistical convergence.
Section 26 of 36
26. State that imputation is not measurement
Enhanced values may be useful for hypothesis generation and visualisation, but they do not increase the number of captured molecules. Preserve original counts, disclose which values were imputed and avoid treating subspot sample size as independent experimental replication.
Section 27 of 36
27. Compare BayesSpace with SpaGCN
BayesSpace uses Bayesian spatial clustering and MCMC, whereas SpaGCN combines expression, coordinates and histology in a graph-convolutional framework. Agreement can support broad domains. Differences may expose image influence, prior smoothing, feature choice or cluster resolution.
Section 28 of 36
28. Keep negative spatial claims bounded
Failure to see a small region may reflect coarse spots, low counts or smoothing. Simulate regions of different size and contrast to estimate the detection envelope. State what broad structures were excluded and which fine alternatives remain unresolved.
Section 29 of 36
29. Try a classroom model
Give students a coarse grid with noisy numbers and ask them to label regions using both values and neighbours. Then add one tiny island to test whether smoothing preserves or erases it. Ask students to state what was observed, what was inferred and what information remains unresolved. That three-column habit transfers directly to laboratory reports and prevents a plausible calculation from being mistaken for a new measurement.
Section 30 of 36
30. Start the habit in Primary Science
Primary Science builds careful observation, comparison and fair testing. A child can learn the foundation of BayesSpace without advanced algebra by sorting evidence from explanation, checking repeated measurements and asking whether a model has enough information. Those habits matter in PSLE Science and much later in computational biology.
Section 31 of 36
31. Strengthen PSLE Science reasoning
For PSLE Science, the useful link is not memorising software names. It is answering with evidence: identify the variable, read the table, connect cause and effect, and avoid a conclusion that exceeds the results. BayesSpace makes that familiar discipline visible in a modern research setting.
Section 32 of 36
32. Connect Secondary and O-Level Science
Secondary Science and O-Level Biology develop cells, organisation, variation, experimental design, graphs and evaluation. BayesSpace adds a contemporary example in which gene counts, tissue position and computation meet. Students can practise interpreting axes, controls, uncertainty and alternative explanations while keeping syllabus fundamentals central.
Section 33 of 36
33. Use Computing as a partner
Computing contributes data structures, algorithms, iteration, visualisation and reproducibility. Science contributes the question, samples, controls and biological interpretation. The productive boundary is clear: code can organise and test evidence, but it cannot rescue a weak reference or turn an assumption into an observation.
Section 34 of 36
34. Choose schools by fit, not a single buzzword
Families comparing Singapore schools should verify current subject combinations, laboratories, applied learning, CCAs, mentoring and timetable fit on official school and MOE pages. A fashionable term such as BayesSpace is not evidence of programme quality. Look for sustained inquiry, sound teaching and opportunities to explain results.
Section 35 of 36
35. See the career pathways
The reasoning behind BayesSpace appears in spatial genomics, biostatistics, medical imaging, computational pathology and scientific computing. Pathways can include polytechnic, junior college, university and continuing education, with different mixtures of biology, statistics, computing, engineering and communication. Careers depend on qualifications and experience; one school project does not promise an outcome.
Section 36 of 36
36. Finish with a bounded claim
A careful conclusion keeps the central distinction visible: BayesSpace computationally clusters and enhances measured spots; subspots and imputed expression are not new physical measurements. Report the data, model, uncertainty, sensitivity and validation together. That measured language is not timid—it is how science remains useful when the result moves from a colourful figure toward a decision.
A publication-ready BayesSpace analysis starts with a spatial data contract. Record platform, spot geometry, physical dimensions, tissue mask, coordinate system, specimens, processing batches and image registration. The neighbourhood graph must respect the array and actual tissue. Display edges around holes, tears and borders so readers can see where spatial support is sparse or potentially misleading.
Pre-register quality filters, normalisation, feature genes, principal components, candidate cluster counts, spatial parameters, MCMC iterations, burn-in, chains, seeds, primary domain comparison and validation. Resolution enhancement adds another layer: specify subspot geometry, enhanced outputs and the assay that will test fine-scale claims. Separate exploratory maps from locked confirmation.
Publish a funnel from raw spots to tissue spots, quality-controlled counts, selected genes, principal components, modelled spots and enhanced subspots. Map every exclusion and count metric. Technical gradients are often spatially coherent; a Bayesian prior can reinforce them unless raw quality and batch are inspected beside the cluster map.
Feature sensitivity deserves formal analysis. Repeat clustering across defensible gene and principal-component choices. Track which broad domains remain stable and where boundaries move. Principal components compress biology and noise together, so the chosen representation is part of the evidence. Save component loadings and explained variation, not only the final labels.
Cluster number should be reported as a hierarchy. Examine several plausible values, quantify agreement and compare with independent tissue landmarks. More clusters inevitably produce more colours. A fine subdivision is credible only when its marker, morphology and replication evidence also becomes more specific, not merely because the algorithm can partition the data.
Spatial-prior sensitivity belongs in the main results. Compare weaker and stronger coherence and, where useful, expression-only clustering. Stable domain cores may persist while transition zones move. Report that layered outcome. Selecting the smoothest map risks erasing rare islands or sharp biological interfaces that do not conform to neighbouring majority labels.
MCMC diagnostics are not optional. Run multiple chains or independent starts, retain enough post-burn-in samples and inspect relevant traces or stability summaries. Label switching and poor mixing can produce changing maps even when code completes. Archive samples or sufficient summaries so uncertainty can be re-examined rather than reduced to one hard label.
Resolution enhancement must remain visually tethered to original measurements. Pair every subspot panel with the measured spot grid, counts and uncertainty. Use a different legend or explicit label for imputed expression. A fine-scale prediction can guide targeted imaging, but it should not be counted as new molecules, new independent observations or direct single-cell localisation.
Validate at matching scale. Broad clusters can be checked with anatomy, histology and regional markers. A within-spot arrangement needs higher-resolution imaging, segmentation or another spatial assay. Validation performed only at spot scale cannot establish which subspot held the signal. Include expected negatives and challenging edges, not only stable centres.
Biological replication remains specimen-based. Neighbouring spots, subspots and MCMC samples are not independent animals or patients. Show all specimens, balance processing and use specimen-level summaries for general claims. Adjacent sections help test technical repeatability, while independent subjects test biological generalisation.
Reproducibility materials include raw and processed matrices, coordinates, tissue masks, quality maps, features, component loadings, cluster candidates, neighbourhood settings, priors, chains, seeds, diagnostics, posterior summaries, enhanced outputs, validation data, package version, environment, code and checksums. Preserve original and imputed values in separate objects to prevent accidental mixing downstream.
Negative results need size-and-contrast calibration. Simulate spatial islands and gradients at different spot scales, expression contrasts and count depths, then estimate which structures survive smoothing and enhancement. Failure to detect a tiny region does not prove absence. It may show that the coarse assay and spatial prior could not resolve that configuration.
A careful final sentence might say: ‘Prespecified BayesSpace analyses identified broad domains reproducible across independent sections, while computational subspot enhancement generated testable fine-scale hypotheses supported for selected targets by higher-resolution assays.’ It should not describe imputed subspots as additional measured locations or let smoothness stand in for truth.
The most useful enhanced map therefore behaves like a hypothesis generator with a visible audit trail. Readers should be able to move backward from each subspot claim to its neighbouring measured spots, feature representation, prior, chain diagnostics and independent validation. That traceability is what turns computational detail into scientific evidence rather than decorative precision.
Contents · Previous section · Continue to the Science Learning Hub
