eduKateSG · Why Science?
Turn expression and curated interactions into testable cell-communication networks—without mistaking a modelled probability or colourful pathway map for observed molecular conversation
Reading routes
Science learning becomes useful when a familiar object or observation is turned into a system of quantities, mechanisms and claim limits. This guide owns one applied evidence-reading job inside eduKateSG’s wider Science estate. It connects naturally to Why Science Commot Collective Optimal Transport Cell Communication Evidence; Why Science Ncem Spatial Cell Graphs Contextual Expression Communication Evidence; Why Science Misty Multiview Spatial Contexts Marker Dependency Evidence; Why Science Spatial Transcriptomics Tissue Coordinates Gene Expression Evidence; Education Hub; Singapore Secondary School Directory; Career Adulthood Hub. It also keeps current school and public claims traceable to visible primary sources: CellChat primary study; CellChat primary PubMed record; CellChat v2 protocol PubMed record; Official CellChat repository; 2026 Singapore–Cambridge O-Level Biology syllabus; MOE G2/G3 Lower Secondary Science syllabus; 2026 MOE G2 Computing syllabus. The sources describe the scientific scope; this article translates that scope into a calm route for Primary Science, PSLE Science, Secondary Science, O-Level Science, STEM exploration, school choices and career pathways without inventing admission or employment outcomes.
CellChat is an open-source R toolkit for inferring and analysing cell–cell communication from single-cell and spatially resolved transcriptomic data. Its published framework combines expression with a curated interaction database, estimates communication probabilities, aggregates interactions into signalling pathways and uses network analysis to study sender, receiver, mediator and influencer roles. Those outputs are hypotheses about possible signalling under declared data and modelling assumptions. They do not directly observe ligand secretion, receptor binding or a causal response.
Inside this guide
1–12 · Foundations and models
- 1. Start with a network question
- 2. Define senders and receivers
- 3. Bring in the interaction database
- 4. Estimate communication probability
- 5. Aggregate pairs into pathways
- 6. Study roles with network analysis
- 7. Search for communication patterns
- 8. Create a specimen-first manifest
- 9. Audit cell identity evidence
- 10. Control expression quality
- 11. Respect abundance and sampling
- 12. Keep complexes and cofactors explicit
13–24 · Evidence, testing and applications
- 13. Track the database version
- 14. Archive the complete configuration
- 15. Practise with an invented network audit table
- 16. Read probabilities as conditional evidence
- 17. Return from pathways to pairs
- 18. Compare centrality across matched networks
- 19. Interpret outgoing patterns carefully
- 20. Interpret incoming patterns carefully
- 21. Compare conditions at specimen level
- 22. Audit label sensitivity
- 23. Audit database sensitivity
- 24. Audit modelling choices
25–36 · Learning, decisions and pathways
- 25. Add spatial and molecular evidence
- 26. Compare with transparent baselines
- 27. Publish failures and negative controls
- 28. Turn one edge into an experiment
- 29. Build a classroom model before code
- 30. Begin with Primary Science habits
- 31. Use PSLE Science to practise claim limits
- 32. Connect Secondary and O-Level Science
- 33. Let Computing support the biology
- 34. Choose opportunities by verified fit
- 35. See connected career families
- 36. Finish with a bounded scientific claim
Section 1 of 36
1. Start with a network question
A tissue is not merely a list of cell types; it is a coordinated system in which populations may send, receive and relay signals. CellChat asks which curated ligand–receptor routes are compatible with the observed expression and how those candidate routes organise into networks.
Section 2 of 36
2. Define senders and receivers
Every inferred edge begins with a declared source group and target group. Those identities usually come from clustering and annotation, so uncertainty in cell labels flows directly into apparent communication. Report ambiguous or mixed populations instead of forcing every cell into a neat category.
Section 3 of 36
3. Bring in the interaction database
CellChatDB supplies curated ligands, receptors, cofactors and pathway membership. The database is prior knowledge, not a measurement from the sample. Record species, release, evidence scope and any subset because changing this catalogue changes which communications are even possible.
Section 4 of 36
4. Estimate communication probability
The model combines group-level expression and interaction rules to assign communication probabilities. A probability here is a model-derived score conditioned on inputs and assumptions; it is not a measured fraction of cells that exchanged molecules and should not be narrated as such.
Section 5 of 36
5. Aggregate pairs into pathways
Individual ligand–receptor interactions can be grouped into named signalling pathways. Aggregation creates a useful biological overview while hiding pair-level variation. Preserve the interactions that support every pathway so a strong label cannot conceal one fragile or ubiquitous component.
Section 6 of 36
6. Study roles with network analysis
Centrality measures can describe groups as prominent senders, receivers, mediators or influencers within an inferred network. These are graph roles, not permanent cellular identities. A cell type may change role by pathway, condition, database choice or preprocessing decision.
Section 7 of 36
7. Search for communication patterns
CellChat can identify coordinated outgoing or incoming signalling patterns across cell groups and pathways. Patterns help compress a complex network, but the chosen number and stability of patterns matter. Treat them as summaries to test, not hidden biological programmes automatically discovered.
Section 8 of 36
8. Create a specimen-first manifest
Record donor or animal, tissue region, condition, batch, cell count, annotation method and analysis unit before computing any edge. Many cells from one specimen improve resolution but do not create independent biological replicates.
Section 9 of 36
9. Audit cell identity evidence
Use canonical markers, reference mapping, differential features and expert review where appropriate. Check whether a communication claim disappears when uncertain subclusters are merged. A result that depends on a disputed label should be described as label-sensitive.
Section 10 of 36
10. Control expression quality
Inspect library size, mitochondrial signal, ambient RNA, doublets, dissociation stress and cell-cycle effects. Secreted genes can be sparse or induced by handling. Quality covariates should be mapped onto source and receiver groups, not buried in a methods appendix.
Section 11 of 36
11. Respect abundance and sampling
Rare populations can vanish after filtering, while abundant groups can dominate summaries. Compare cell counts per specimen and condition, and avoid interpreting sampling imbalance as changed signalling. Use prespecified minimum group sizes and report excluded groups.
Section 12 of 36
12. Keep complexes and cofactors explicit
A receptor may require multiple subunits, agonists or antagonists. State how the chosen CellChatDB version represents these components and how absent subunits are handled. Family names should not be treated as proof that a functional complex exists at the membrane.
Section 13 of 36
13. Track the database version
CellChat’s knowledge base and software can evolve. Freeze package version, database object, species conversion and any custom interactions. An analysis that cannot identify its knowledge source cannot explain why an edge appeared or disappeared.
Section 14 of 36
14. Archive the complete configuration
Save input matrices, metadata, group labels, filters, database, preprocessing choices, probability settings, pathway aggregation, centrality calculations, pattern settings, seeds and exported tables. A figure should be rebuildable without reconstructing choices from memory.
Section 15 of 36
15. Practise with an invented network audit table
This fictional table teaches interpretation, not biological performance.
| Pathway | Specimens supporting | Pair stability | Label sensitivity | Orthogonal check | First reading |
|---|---|---|---|---|---|
| Amber | 5/6 | high | low | protein agrees | supported candidate |
| Birch | 2/6 | medium | high | none | exploratory |
| Coral | 6/6 | low | low | mixed | database-sensitive |
| Delta | 0/6 controls | none | low | negative | useful null |
Section 16 of 36
16. Read probabilities as conditional evidence
Ranked probabilities are meaningful within the defined model and input. Compare them across planned groups only when preprocessing and population definitions are compatible. Avoid turning score differences into fold changes or physical signal strength without a justified calibration.
Section 17 of 36
17. Return from pathways to pairs
A pathway-level edge can combine several ligand–receptor pairs with different quality. Show the dominant and supporting pairs, their expression distributions and complex requirements. If one pair owns the pathway result, the conclusion should name that narrow dependency.
Section 18 of 36
18. Compare centrality across matched networks
Network roles depend on graph density and weighting. Before declaring a new ‘hub’, check whether networks were built from matched cells, thresholds and interaction sets. Normalise comparisons transparently and retain per-specimen role estimates.
Section 19 of 36
19. Interpret outgoing patterns carefully
An outgoing pattern summarises pathways that source groups tend to share. It can generate a hypothesis about coordinated secretion, but cannot establish common regulation. Examine whether the pattern is stable across seeds and whether one abundant pathway dominates it.
Section 20 of 36
20. Interpret incoming patterns carefully
An incoming pattern groups candidate signals received by target populations. Receptor transcripts, cofactors and downstream response need separate validation. A colourful incoming module is not evidence that the receiver activated all pathways shown.
Section 21 of 36
21. Compare conditions at specimen level
Case–control contrasts should begin with matched biological replicates, not pooled cells. Quantify within-condition variability and interaction prevalence across specimens. A condition difference driven by one donor should be narrowed to that observation.
Section 22 of 36
22. Audit label sensitivity
Repeat central claims with nearby annotation resolutions and justified merged groups. Track whether sender, receiver or pathway identity changes. Stable conclusions survive plausible labels; unstable ones identify where better reference data or imaging is needed.
Section 23 of 36
23. Audit database sensitivity
Compare a declared CellChatDB subset or version with a defensible alternative resource or high-confidence subset. Separate missing interactions from altered rankings. Database disagreement is information about knowledge uncertainty, not a reason to hide one result.
Section 24 of 36
24. Audit modelling choices
Vary robust mean, population-size adjustment, minimum cells and other defensible settings according to a prespecified plan. Report the stable core and the turnover. Selecting the parameter combination that produces the busiest network is not validation.
Section 25 of 36
25. Add spatial and molecular evidence
Use spatial transcriptomics, imaging, protein localisation or tissue architecture to test whether proposed partners can plausibly meet. Then use targeted assays for secretion, binding, activation or response. Proximity supports opportunity, not mechanism.
Section 26 of 36
26. Compare with transparent baselines
A simple ligand–receptor expression product and a second method under identical preprocessing reveal what CellChat adds. Compare results at pair, pathway and cell-group levels. Agreement strengthens prioritisation; disagreement should be traced to assumptions.
Section 27 of 36
27. Publish failures and negative controls
Permute labels within appropriate specimen strata, test incompatible interactions and show pathways that fail replication. Negative results establish the background against which a communication network is judged. They also prevent every tissue from appearing densely conversational.
Section 28 of 36
28. Turn one edge into an experiment
Choose one replicated source–ligand–receptor–receiver claim and specify the perturbation, readout, timing and result that would overturn it. The best network is not the most elaborate one; it is the one that helps design a decisive test.
Section 29 of 36
29. Build a classroom model before code
Give groups coloured ligand, receptor and cofactor cards. They may draw an arrow only when every required component is present, then must remove arrows after receiving a new quality-control clue. Label observations, transformations, assumptions and conclusions separately. The goal is not to imitate specialist software; it is to make every evidence hand-off visible enough for a classmate to question.
Section 30 of 36
30. Begin with Primary Science habits
Primary Science already supplies the foundation for CellChat: careful observation, fair comparison, consistent records and conclusions that fit the evidence. A simple plant, light or water investigation can show why compatible parts do not prove that an interaction actually occurred.
Section 31 of 36
31. Use PSLE Science to practise claim limits
PSLE Science asks students to connect evidence to process without leaping beyond an experiment. With CellChat, ask what was measured, what came from a database, what the software estimated and which new observation could separate two explanations.
Section 32 of 36
32. Connect Secondary and O-Level Science
Secondary Science and O-Level Biology develop cells, organisation, molecular transport, experimental design and evaluation. CellChat turns those ideas into a contemporary data problem. The useful lesson is disciplined interpretation, not memorising a software menu.
Section 33 of 36
33. Let Computing support the biology
Computing contributes tables, graphs, algorithms, ranking, version control and reproducibility. Biology supplies specimens, mechanisms and validation. CellChat shows why correct code is necessary but insufficient: a flawless program can still analyse confounded samples or unsuitable labels.
Section 34 of 36
34. Choose opportunities by verified fit
For school choices and science enrichment, verify current programmes on official school and MOE pages. Look for careful inquiry, reproducible computing and opportunities to explain evidence, rather than prestige labels alone. A fashionable mention of genomics or AI does not guarantee sustained teaching, admission, mentorship or a particular career outcome.
Section 35 of 36
35. See connected career families
The reasoning in CellChat appears in single-cell genomics, immunology, systems biology, bioinformatics, statistics, network science and experimental design. Routes may pass through polytechnic, junior college, university or continuing education, with different blends of biology, statistics, computing and communication. Requirements change, so check the responsible institution directly.
Section 36 of 36
36. Finish with a bounded scientific claim
A defensible conclusion states exactly what CellChat prioritised under which data, database, settings and specimens. CellChat infers expression- and database-supported communication probabilities and network summaries; it does not directly measure secretion, binding, receptor activation or causal response. Report uncertainty, sensitivity, replication, negative results and independent validation together; those boundaries make a claim more useful.
A reliable CellChat study keeps four layers apart: measured transcript counts, assigned cell identities, curated interaction knowledge and modelled communication networks. Use a diagram or data dictionary that names each hand-off. Readers should be able to trace a pathway arrow to the source cells, ligand, receptor or complex, database entry, probability settings, supporting specimens and independent checks.
Start with the experimental unit. If six specimens were sequenced, preserve six specimen-level views even when the software is run on a combined object for exploration. Show cell counts and detection for every source and receiver in each specimen. A dense network produced mainly by one large sample is not replicated communication.
Annotation deserves its own sensitivity study. Repeat central conclusions after merging uncertain neighbouring subtypes and after removing likely doublets or stressed cells. Compare whether pathway identity, direction and sender–receiver roles survive. This turns label uncertainty into a measurable property rather than a vague limitation.
CellChat’s pathway aggregation is attractive because it turns many pairs into a navigable network. Yet an attractive summary can hide a single dominant interaction, incomplete complex or ubiquitous ligand. Publish a drill-down table for every headline pathway, including cofactors, agonists, antagonists and the contribution of each pair.
Graph centrality is relational. A group becomes a prominent sender partly because of which other groups and edges are present. Removing a rare population, changing a threshold or restricting the database can alter roles. Compare centrality only on declared, compatible networks and report uncertainty across specimens or resamples.
Pattern discovery adds another compression layer. Predefine how the number of patterns is chosen and examine stability across seeds or nearby choices. Label a pattern from its actual pathway loadings, not from the most exciting member. If a pattern mixes unrelated pathways, retain a neutral name until biology supplies evidence.
Did you know? A model can infer the same broad pathway in two conditions while the supporting ligand–receptor pairs change. This is why pathway-level agreement and pair-level agreement answer different questions. Show both, especially when comparing treatments or developmental stages.
Expression preprocessing can own the outcome. Pseudocounts, normalisation, variable-feature selection and robust averages affect sparse ligand and receptor genes. Run planned checks with defensible alternatives, preserve the full result set and explain which conclusion remains stable. Avoid using parameter exploration as a hunt for the richest graph.
Population-size adjustment also changes interpretation. It may reflect the opportunity created by abundant cell populations, but abundance itself can be biological or technical. Provide results with and without the adjustment when it materially changes conclusions, and connect both versions to measured cell composition.
Spatial information should restrict possibility, not decorate a dissociated analysis after the fact. If spatial data exist, test whether senders and receivers share a niche at a biologically relevant scale. If only dissociated data exist, say that proximity is unknown and avoid directional tissue language.
Negative controls need structure. Randomly scramble cell labels within specimens, use biologically incompatible pairs or replace the interaction resource with matched random pairs. An informative analysis should distinguish the real arrangement from controls. If many pathways survive, general expression or group size may be driving the network.
Orthogonal validation should target the weakest inference step. Protein imaging tests localisation, secretion assays test ligand availability, phospho-protein measurements test receptor-linked activation and perturbations test necessity or sufficiency. One assay rarely validates the entire source-to-response chain, so state what each check addresses.
For condition comparisons, display prevalence and effect together. A pathway detected in five of six treated specimens and one of six controls tells a different story from a pathway whose pooled score is higher because one treated specimen dominates. Avoid cell-level p-values that ignore specimen-level independence.
Reproducibility includes the knowledge base. Archive the precise CellChatDB object, custom additions, software session, random seeds and exported tables. Screenshots and circular network plots are outputs, not sufficient records. Future readers should be able to recover why a named edge was eligible.
A publication-ready sentence might read: “CellChat prioritised a pathway-level communication hypothesis from source group A to receiver group B that was supported by specified ligand–receptor pairs in five of six specimens and remained stable under declared label, database and parameter checks.” It should not claim that cells sent or received a signal unless direct evidence supports that verb.
Finish by asking what observation would make the favourite edge disappear from the story. Perhaps the protein is absent, the receptor complex is incomplete, the populations never meet or the receiver response persists after blockade. Writing the falsifier before validation protects the experiment from becoming a search for confirmation.
Contents · Previous section · Continue to the Science Learning Hub
