eduKateSG · Why Science?
Learn recurring spatial ligand–receptor patterns and multi-hop relay candidates with attention—without turning an attention score into observed molecular transmission
Reading routes
Science learning becomes useful when a familiar object or observation is turned into a system of quantities, mechanisms and claim limits. This guide owns one applied evidence-reading job inside eduKateSG’s wider Science estate. It connects naturally to Why Science Spatalk Spatial Ligand Receptor Target Networks Knowledge Graph Evidence; Why Science Commot Collective Optimal Transport Cell Communication Evidence; Why Science Sctensor Many To Many Cell Interactions Hypergraph Evidence; Why Science Cellchat Communication Networks Signalling Pathways Pattern Evidence; Education Hub; Singapore Secondary School Directory; Career Adulthood Hub. It also keeps current school and public claims traceable to visible primary sources: CellNEST primary study; CellNEST PubMed record; CellNEST primary full text; Official CellNEST repository; CellNEST paper figure repository; 2026 Singapore–Cambridge O-Level Biology syllabus; MOE G2/G3 Lower Secondary Science syllabus; 2026 MOE G2 Computing syllabus. The sources describe the scientific scope; this article translates that scope into a calm route for Primary Science, PSLE Science, Secondary Science, O-Level Science, STEM exploration, school choices and career pathways without inventing admission or employment outcomes.
CellNEST is an open-source spatial cell–cell communication method published in Nature Methods in 2025. It represents cells or spots as vertices and potential ligand–receptor relations as directed spatial edges, then uses a graph-attention encoder trained with Deep Graph Infomax contrastive learning to prioritise recurring communication patterns. It also extracts putative ligand–receptor–ligand–receptor relay networks and can score their intracellular plausibility. The paper explicitly notes that spatial transcriptomics is a snapshot and cannot determine whether receptor activation caused the next ligand’s expression. Attention and relay scores therefore rank hypotheses rather than proving transmission.
Inside this guide
1–12 · Foundations and models
- 1. Start with a spatial multigraph
- 2. Keep the 2025 paper boundary
- 3. Activate genes by threshold
- 4. Define spatial neighbours
- 5. Encode edge features
- 6. Learn with graph attention
- 7. Train without labelled ground truth
- 8. Create a spatial graph manifest
- 9. Audit cell and spot identity
- 10. Freeze the pair database
- 11. Protect coordinate quality
- 12. Archive model training
13–24 · Evidence, testing and applications
- 13. Protect specimen independence
- 14. Plan realistic controls
- 15. Practise with an invented attention table
- 16. Read attention as learned importance
- 17. Read connected components
- 18. Read pair-frequency histograms
- 19. Understand two-hop relays
- 20. Inspect intracellular relay confidence
- 21. Compare specimens and regions
- 22. Audit expression percentiles
- 23. Audit neighbourhood distance
- 24. Audit top-edge filtering
25–36 · Learning, decisions and pathways
- 25. Audit seeds and convergence
- 26. Add protein and temporal evidence
- 27. Design a relay perturbation
- 28. Write a bounded CellNEST claim
- 29. Build a classroom relay graph
- 30. Begin with Primary Science location
- 31. Use PSLE Science for sequences
- 32. Connect Secondary and O-Level Science
- 33. Let Computing explain attention
- 34. Choose opportunities by current evidence
- 35. See connected careers
- 36. Finish with attention and causation separate
Section 1 of 36
1. Start with a spatial multigraph
CellNEST represents each cell or spot as a vertex and each eligible ligand–receptor relationship as a directed edge. Multiple molecular edges can connect the same vertices. The input graph describes possible neighbourhood relations, not confirmed communication.
Section 2 of 36
2. Keep the 2025 paper boundary
The Nature Methods study evaluated synthetic arrangements and spatial datasets across technologies, species and disease contexts. Reported performance belongs to those benchmarks and settings. A new tissue still needs its own controls and validation.
Section 3 of 36
3. Activate genes by threshold
Ligands and receptors above a chosen expression percentile are treated as active for graph construction. This is a modelling rule. Repeat key results across reasonable thresholds because a small expression change can add or remove many edges.
Section 4 of 36
4. Define spatial neighbours
Edges are considered only within a supplied neighbourhood distance, with defaults differing for spot and cell data. The correct range depends on platform and mediator biology. Proximity supplies opportunity, not transfer.
Section 5 of 36
5. Encode edge features
Each candidate edge includes physical distance, ligand–receptor co-expression and pair identity. These features let the model distinguish molecular signals and spatial context. They remain transcript-derived and database-derived attributes.
Section 6 of 36
6. Learn with graph attention
A graph-attention encoder learns vertex representations and assigns attention scores to edges. Attention indicates importance for reconstructing recurring graph patterns under the model. It is not a probability of receptor binding or a causal weight.
Section 7 of 36
7. Train without labelled ground truth
CellNEST uses Deep Graph Infomax contrastive learning because comprehensive labelled communication edges are unavailable. The model distinguishes observed graph structure from corrupted alternatives. Unsupervised learning reduces dependence on labels but does not create biological truth.
Section 8 of 36
8. Create a spatial graph manifest
Record platform, resolution, 2D or 3D coordinates, specimen, preprocessing, expression percentile, neighbourhood distance, ligand–receptor database, optional pathways, model version, seed and hardware. Spatial and deep-learning settings are part of the result.
Section 9 of 36
9. Audit cell and spot identity
Single-cell technologies and multicellular spots provide different units. Preserve whether vertices are measured cells, segmented cells or spots. Cell-type annotations may help interpretation but should not be confused with the graph’s spatial vertices.
Section 10 of 36
10. Freeze the pair database
The paper’s default combines CellChat and NicheNet resources into 12,605 pairs. Treat that as the reported resource, not a permanent universal total. Archive the exact table, species, direction and identifier mapping used.
Section 11 of 36
11. Protect coordinate quality
Check tissue orientation, units, masks, registration and segmentation before building edges. A coordinate error can manufacture recurring patterns that attention then learns faithfully.
Section 12 of 36
12. Archive model training
Retain graph features, corrupted graphs, learned embeddings, attention scores, thresholds, seeds, convergence diagnostics and filtered edges. A final interactive network cannot reconstruct training.
Section 13 of 36
13. Protect specimen independence
Train or evaluate findings with independent sections, patients or animals in mind. Millions of edges from one section do not constitute millions of biological replicates.
Section 14 of 36
14. Plan realistic controls
Use anatomical negatives, abundance-matched pairs, coordinate-preserving shuffles and held-out tissues where possible. A corruption that destroys every spatial and molecular property is easy to distinguish and may overstate confidence.
Section 15 of 36
15. Practise with an invented attention table
This fictional table is not CellNEST output. Decide which model layer requires checking.
| Fictional pattern | Attention stability | Relay support | Main next check |
|---|---|---|---|
| P1 | Stable | Two-hop recurrent | Intracellular assay |
| P2 | Threshold-sensitive | Frequent | Expression sensitivity |
| P3 | Stable in one section | Strong | New specimens |
| P4 | Stable | Database-sensitive | Pair provenance |
Section 16 of 36
16. Read attention as learned importance
Show attention distributions, ranks and stability across seeds. A score of one is not one hundred per cent biological certainty. It means the edge was highly weighted by the trained representation under the declared inputs.
Section 17 of 36
17. Read connected components
High-scoring edges can form communicating regions in the output graph. Compare those regions with independent anatomy rather than naming them from the model alone. A component may reflect density, expression or tissue boundaries.
Section 18 of 36
18. Read pair-frequency histograms
Frequent ligand–receptor identities among top edges reveal repeated motifs. Frequency can reflect widespread biology or abundant transcripts. Compare with eligible-edge denominators and spatially matched controls.
Section 19 of 36
19. Understand two-hop relays
A relay contains a first ligand–receptor edge into an intermediate cell or spot and a second ligand–receptor edge leaving it. Recurrence supports a pattern. It does not establish that the first receptor induced production of the second ligand.
Section 20 of 36
20. Inspect intracellular relay confidence
CellNEST can search receptor-to-transcriptional-activator paths that might connect the first input to the second ligand. Publish path provenance and expression support. This confidence layer is prior-informed, not direct biochemical measurement.
Section 21 of 36
21. Compare specimens and regions
Ask whether the same pair or relay recurs in independent samples and matching anatomical regions. Report prevalence and heterogeneity. A disease-subtype result needs patient-level support, not only many local edges.
Section 22 of 36
22. Audit expression percentiles
Repeat graph construction across planned thresholds. Record how many vertices and edges are eligible and whether headline signals persist. A stable pattern is stronger than one created by one percentile.
Section 23 of 36
23. Audit neighbourhood distance
Test biologically plausible distances for contact, paracrine and self-signalling modes. A single distance can blur mechanism classes. Report which spatial assumptions support each claim.
Section 24 of 36
24. Audit top-edge filtering
The paper uses a default top fraction for retaining high-attention edges. Show continuous scores and sensitivity to the retained fraction. A route should not become important only because the display hides competitors.
Section 25 of 36
25. Audit seeds and convergence
Deep models can vary across training runs. Repeat seeds, align edges and relay patterns, and publish stability. A one-run attention map is exploratory evidence.
Section 26 of 36
26. Add protein and temporal evidence
Confirm ligands and receptors in place, then test an early receiver event and the proposed relay ligand over time. Spatial transcriptomics is a snapshot and cannot establish relay order.
Section 27 of 36
27. Design a relay perturbation
Block the first ligand or receptor and measure the second ligand before blocking the second receptor and measuring the final response. This staged design directly tests the relay interpretation rather than two unrelated edges.
Section 28 of 36
28. Write a bounded CellNEST claim
Prefer: ‘CellNEST prioritised a recurring two-hop spatial pattern with stable attention and intracellular-path support.’ Avoid saying the tissue executed a proven relay. The bounded claim is specific and testable.
Section 29 of 36
29. Build a classroom relay graph
Place cell cards on a map, add two ligand–receptor arrows and ask whether the middle cell truly relays the message. Students identify the missing time and perturbation evidence.
Section 30 of 36
30. Begin with Primary Science location
Nearness, direction and repeated patterns are accessible ideas. Learners can observe a map and distinguish an adjacency from an action.
Section 31 of 36
31. Use PSLE Science for sequences
A relay is an ordered hypothesis: first signal, intermediate response, second signal. Students can predict what blocking the first step should change downstream.
Section 32 of 36
32. Connect Secondary and O-Level Science
Receptors, coordination, gene expression and experimental design explain the biology. CellNEST adds spatial graphs and model-based evidence while keeping causal limits clear.
Section 33 of 36
33. Let Computing explain attention
Computing contributes graph features, embeddings, attention, contrastive learning, search and interactive visualisation. Biology decides whether learned patterns correspond to plausible molecular routes.
Section 34 of 36
34. Choose opportunities by current evidence
Verify official programme details and look for sustained work in biology, spatial data, computing and research. Avoid assuming entry or outcomes from a modern-sounding programme label.
Section 35 of 36
35. See connected careers
The work connects to spatial biology, pathology, machine learning, bioinformatics, imaging, software engineering and drug discovery. Requirements evolve, so consult current official sources.
Section 36 of 36
36. Finish with attention and causation separate
A defensible CellNEST conclusion names the coordinates, thresholds, database, graph, training, filtering, relay search, specimens and stability, then separates learned spatial patterns from protein, temporal and perturbation validation. Attention guides experiments; it does not replace them.
CellNEST represents spatial transcriptomic data as a directed multigraph. Cells or spots become vertices, and eligible ligand–receptor relationships between nearby vertices become edges. Multiple molecular pairs can connect the same two locations. A graph-attention encoder then learns which edge patterns are useful under its training objective, and the workflow can assemble high-ranking edges into candidate relay networks.
Did you know? The CellNEST paper’s default ligand–receptor collection combined CellChat and NicheNet resources into a reported snapshot of 12,605 pairs. That number belongs to the publication’s resource version. Database updates, species mapping and complex rules can change the eligible edge universe, so reproduce a study with its frozen resource.
The spatial graph is the first model. Record coordinate units, image transformations, neighbourhood distance and whether spots or cells are the vertices. Plot edge-length distributions and check tissue boundaries. A single distance threshold is not equally plausible for contact molecules and diffusible signals; use sensitivity analyses or biologically informed classes.
Gene activation thresholds are another model choice. CellNEST can define active ligands and receptors using expression percentiles. Report the threshold, the reference population used to calculate it and the number of active genes per cell. A global percentile and a cell-type-specific percentile answer different questions and can reshape the graph.
Each edge can carry features such as distance, ligand–receptor co-expression and pair identity. Preserve those components alongside the final attention score. A high score may arise from strong expression, favourable geometry or a repeatedly learned pair pattern. Component visibility keeps interpretation grounded.
Deep Graph Infomax supplies an unsupervised contrastive objective that encourages informative local representations relative to broader graph structure. This is a learning signal, not a biological assay. The model does not watch a ligand move or a receptor activate. Attention indicates importance inside the fitted computation; it is not a probability of communication or a causal effect.
The paper used a default top-20-percent attention-edge filter in its analyses. Treat that as a reported analytical choice, not a law of nature. Show sensitivity across several cutoffs, report the number of retained edges and examine whether headline pathways persist. A beautiful network that exists only at one percentile is fragile.
Relay extraction is especially exciting. A putative ligand–receptor–ligand–receptor sequence can suggest how signalling might propagate across locations. CellNEST’s default relay depth can be extended, and intracellular evidence can assess whether an upstream receptor plausibly connects to the next ligand’s activators. Longer paths also create more chances for coincidental linkage.
The publication states an important limitation: a spatial transcriptomic snapshot cannot determine whether activation of one receptor caused expression of the next ligand. Preserve that sentence’s meaning in every relay interpretation. Use words such as candidate, putative and prioritised until temporal or perturbational evidence is available.
Evaluate recurrence across specimens rather than counting edges as independent replicates. Bootstrap cells within specimen, rerun training across seeds and match high-attention patterns. Summarise which pairs, locations and relay motifs recur. A pattern supported by several donors and seeds is more useful than one driven by a dense region in a single section.
Negative controls should challenge both geometry and biology. Shuffle coordinates within anatomical compartments, permute ligand identities while preserving prevalence, test receptor-negative targets and compare distance-matched random pairs. Also compare a simpler co-expression-plus-distance baseline. Attention adds value when it improves a fixed, testable prediction.
Validation should retain the relay structure. Confirm proteins and cellular compartments for each ligand and receptor, then assay a proximal response at each hop. Use time-resolved perturbations when possible: block the first receptor and test whether the proposed intermediate ligand and downstream receiver response change in sequence. One validated edge does not validate an entire relay.
For learners, CellNEST connects graphs with living tissue. Primary Science begins with patterns and observation; PSLE Science builds fair comparisons; Secondary and O-Level Biology add receptors and coordination; Computing adds graph edges, attention and contrastive learning. The durable lesson is that an algorithmic weight describes a model’s calculation, while an experiment tests the biological event.
A clear report includes the full eligible-edge denominator, threshold sensitivity, attention stability, specimen prevalence and intracellular support for every headline relay. Mark observed expression separately from learned attention and database-derived paths. Include a failed or unstable relay so readers can see the quality gate.
Archive the exact code revision, environment, interaction database, thresholds, graph, seeds, edge features and relay settings. The official repository supports implementation; the paper and its figure repository anchor the published analyses. If current code changes defaults, state that boundary explicitly.
The bounded conclusion is powerful enough: CellNEST can prioritise recurring spatial ligand–receptor edges and candidate multi-hop relay networks from spatial transcriptomic data. It does not directly measure ligand secretion, binding, receptor activation, temporal propagation or causal induction of a later ligand. Those gaps are an invitation to design the next experiment.
Before drawing a relay, verify edge orientation and cellular identity at every hop. A receptor and the next ligand must belong to the intended intermediate vertex, not merely to neighbouring spots that share transcripts. When resolution is spot-level, describe the relay at that resolution and avoid implying single-cell precision.
Audit attention against simple quantities. Plot attention beside distance, pair expression and prevalence. If the ranking is almost perfectly explained by one input feature, say so; if it adds a stable pattern beyond those features, show the comparison. Interpretation improves when readers can see what the learned model contributes.
Finally, design a relay-specific experiment rather than validating only the first pair. Time-resolved blockade of the first receptor, measurement of the intermediate ligand and observation of the downstream receiver can test the proposed order. An alternative branch or receptor-negative control helps distinguish propagation from a shared environmental response. That is how an attention-derived hypothesis can mature into biological evidence.
Show informative absences in the spatial graph. An edge can be absent because cells are too far apart, the ligand is below threshold, the receptor is below threshold, the pair is missing from the database or the learned score falls below a display cutoff. Encode those reasons separately. A blank space in a figure is otherwise impossible to interpret.
For condition comparisons, rebuild eligible graphs under harmonised rules and compare patterns across specimens, not raw edge counts alone. Changes in cell density, tissue coverage or sequencing sensitivity can change the number of possible edges. Use opportunity-aware denominators and uncertainty. A claimed increase in communication should mean more support relative to comparable opportunity, not simply a denser section.
Contents · Previous section · Continue to the Science Learning Hub
