VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

Why Science? | Citrus, Stratifying Signatures and Single-Cell Biomarkers

Three students sit around open books and worksheets at a classroom table, reading, writing and discussing the work together.

eduKateSG · Why Science?

Let the endpoint guide a search across many cell subsets—then make prediction, association and mechanism stay in their proper lanes

Full section index · Science Learning Hub

Science learning becomes useful when a familiar object or observation is turned into a system of quantities, mechanisms and claim limits. This guide owns one applied evidence-reading job inside eduKateSG’s wider Science estate. It connects naturally to Why Science Mass Cytometry Metal Isotope Antibodies Single Cell Protein Evidence; Why Science Cydar Hypersphere Counts Spatial Fdr Mass Cytometry; Why Science Milo Neighbourhood Differential Abundance Single Cell Population Shifts; Why Science Sccoda Compositional Models Credible Cell Type Abundance Changes; Education Hub; Singapore Secondary School Directory; Career Adulthood Hub; Science Learning Hub. It also keeps current school and public claims traceable to visible primary sources: Citrus primary study; Citrus PubMed record; Official Citrus repository; 2026 Singapore–Cambridge O-Level Biology syllabus; MOE G2/G3 Lower Secondary Science syllabus; 2026 MOE G2 Computing syllabus. The sources describe the scientific scope; this article translates that scope into a calm route for Primary Science, PSLE Science, Secondary Science, O-Level Science, STEM exploration, school choices and career pathways without inventing admission or employment outcomes.

Citrus—cluster identification, characterization and regression—was introduced by Bruggner and colleagues in Proceedings of the National Academy of Sciences in 2014. It hierarchically clusters pooled cytometry cells, calculates abundance or functional features for each sample, and uses supervised models to find cell-subset features that stratify an experimental or clinical endpoint. Citrus can organise a powerful biomarker search, but a selected signature remains conditional on the cohort, panel, preprocessing, model and validation design; it does not by itself establish mechanism, diagnosis or clinical utility.

Section 1 of 36

1. Begin with a hidden subset

Two groups may have similar broad immune-cell percentages while differing in a narrow functional or phenotypic subset. Manual gates can miss that subset or find it only after an analyst already knows where to look.

Citrus joins unsupervised hierarchical clustering to supervised endpoint modelling: cell subsets are discovered broadly, while endpoint labels are used to select informative per-sample features.

Archive sample manifests, preprocessing, feature roles, model formulae, random seeds, software versions and exclusions so another analyst can reconstruct the decision path.

Contents · Next section

Section 2 of 36

2. Define the endpoint

State whether the target is stimulation, treatment response, disease class, survival or another sample-level outcome. The endpoint scale and timing determine which model and validation are sensible.

The original workflow can summarise cluster abundance and median functional-marker expression, so the scientific question must determine which feature family is interpretable.

Keep the biological sample as the unit of replication. Many cells improve resolution, but they do not turn one donor, animal or culture into many independent experiments.

Contents · Previous section · Next section

Section 3 of 36

3. Know what Citrus means

Citrus expands to cluster identification, characterization and regression. The name describes a pipeline: find nested clusters, characterise them per sample and relate those features to an endpoint.

Hierarchical clusters are nested and correlated. A selected parent and child may describe one biological signal rather than several independent biomarkers.

Use negative controls, label permutations or design-matched null analyses where valid. A trustworthy signal should exceed artefacts that the same workflow can manufacture.

Contents · Previous section · Next section

Section 4 of 36

4. Pool cells with care

Cells from training samples are pooled for hierarchical clustering. Equal sampling rules help prevent a high-yield specimen from defining the hierarchy simply because it contributed more events.

Cross-validation and regularisation control selection inside the training data; truly independent samples are still needed to estimate generalisation and clinical usefulness.

Report effect direction, uncertainty, sample support and sensitivity—not only a colourful embedding, selected feature or adjusted value.

Contents · Previous section · Next section

Section 5 of 36

5. Separate discovery from confirmation

Use one cohort to discover and tune the signature, then an untouched cohort or rigorously nested resampling procedure to assess performance. Reusing held-out labels during tuning breaks the test.

A stratifying feature predicts or separates the recorded endpoint under the analysed design. It does not automatically explain why the endpoint differs.

Create an evidence card naming the contrast, measured modality, statistical unit, model, result, validation status, unresolved alternatives and narrowest defensible claim.

Contents · Previous section · Next section

Section 6 of 36

6. Choose abundance or function

Abundance features ask how much of a sample occupies a cluster. Functional features ask whether a marker such as a signalling readout shifts within that cluster.

Citrus was developed for multidimensional cytometry, where panel design, compensation or normalisation, batch balance and viable-cell quality set the ceiling for inference.

Write association, stratification, differential abundance or differential state when that is what was tested. Reserve cause, diagnosis and benefit for evidence that directly supports them.

Contents · Previous section · Next section

Section 7 of 36

7. Read the 2014 evidence in scope

Bruggner and colleagues demonstrated Citrus on stimulated peripheral-blood mononuclear-cell mass-cytometry data and public datasets, framing it as data-driven discovery of stratifying cellular subpopulations.

Citrus joins unsupervised hierarchical clustering to supervised endpoint modelling: cell subsets are discovered broadly, while endpoint labels are used to select informative per-sample features.

Archive sample manifests, preprocessing, feature roles, model formulae, random seeds, software versions and exclusions so another analyst can reconstruct the decision path.

Contents · Previous section · Next section

Section 8 of 36

8. Build a sample manifest

Record donor, group, outcome, batch, instrument run, stimulation, replicate structure, cell count and exclusions before clustering. Hidden grouping variables can mimic a biomarker.

The original workflow can summarise cluster abundance and median functional-marker expression, so the scientific question must determine which feature family is interpretable.

Keep the biological sample as the unit of replication. Many cells improve resolution, but they do not turn one donor, animal or culture into many independent experiments.

Contents · Previous section · Next section

Section 9 of 36

9. Quality-control the files

Inspect acquisition stability, spillover or normalisation, beads, debris, doublets, dead cells and marker distributions. Automated discovery cannot distinguish a biological subset from an unremoved technical population by intention alone.

Hierarchical clusters are nested and correlated. A selected parent and child may describe one biological signal rather than several independent biomarkers.

Use negative controls, label permutations or design-matched null analyses where valid. A trustworthy signal should exceed artefacts that the same workflow can manufacture.

Contents · Previous section · Next section

Section 10 of 36

10. Choose clustering markers

Lineage and phenotype markers usually define cellular neighbourhoods. Functional markers used as outcomes should not quietly reshape the cluster geometry unless that is a deliberate, explained choice.

Cross-validation and regularisation control selection inside the training data; truly independent samples are still needed to estimate generalisation and clinical usefulness.

Report effect direction, uncertainty, sample support and sensitivity—not only a colourful embedding, selected feature or adjusted value.

Contents · Previous section · Next section

Section 11 of 36

11. Subsample reproducibly

If each file contributes a subset of cells, declare the count and seed. Compare whether a candidate cluster survives reasonable alternative subsamples.

A stratifying feature predicts or separates the recorded endpoint under the analysed design. It does not automatically explain why the endpoint differs.

Create an evidence card naming the contrast, measured modality, statistical unit, model, result, validation status, unresolved alternatives and narrowest defensible claim.

Contents · Previous section · Next section

Section 12 of 36

12. Build the hierarchy

Hierarchical clustering creates nested populations from broad branches to smaller leaves. The tree supports multi-resolution exploration but also creates correlated candidates.

Citrus was developed for multidimensional cytometry, where panel design, compensation or normalisation, batch balance and viable-cell quality set the ceiling for inference.

Write association, stratification, differential abundance or differential state when that is what was tested. Reserve cause, diagnosis and benefit for evidence that directly supports them.

Contents · Previous section · Next section

Section 13 of 36

13. Calculate per-sample features

For every cluster, summarise abundance or functional-marker behaviour separately for each sample. The model should never receive pooled cell values as though they were independent people.

Citrus joins unsupervised hierarchical clustering to supervised endpoint modelling: cell subsets are discovered broadly, while endpoint labels are used to select informative per-sample features.

Archive sample manifests, preprocessing, feature roles, model formulae, random seeds, software versions and exclusions so another analyst can reconstruct the decision path.

Contents · Previous section · Next section

Section 14 of 36

14. Filter unstable features

Very rare clusters or nearly constant features may carry little reliable information. Declare minimum cluster size and prevalence rules before outcome-driven model selection.

The original workflow can summarise cluster abundance and median functional-marker expression, so the scientific question must determine which feature family is interpretable.

Keep the biological sample as the unit of replication. Many cells improve resolution, but they do not turn one donor, animal or culture into many independent experiments.

Contents · Previous section · Next section

Section 15 of 36

15. Practise with a fictional Citrus table

This teaching example is invented and is not a result from the Citrus paper.

Fictional clusterFeatureTraining associationHeld-out directionInterpretation
C14AbundanceHigher in group BSameCandidate subset
C27Marker-P medianLower after stimulusMixedNeeds replication
C41AbundanceStrongReversedLikely unstable
Invented classroom data for comparison practice; not an operational, product-certification or safety dataset.

Hierarchical clusters are nested and correlated. A selected parent and child may describe one biological signal rather than several independent biomarkers.

Use negative controls, label permutations or design-matched null analyses where valid. A trustworthy signal should exceed artefacts that the same workflow can manufacture.

Contents · Previous section · Next section

Section 16 of 36

16. Select a supervised model

The original software supports regularised prediction and significance-analysis approaches. Model choice should match classification, regression or survival goals and the available sample size.

Cross-validation and regularisation control selection inside the training data; truly independent samples are still needed to estimate generalisation and clinical usefulness.

Report effect direction, uncertainty, sample support and sensitivity—not only a colourful embedding, selected feature or adjusted value.

Contents · Previous section · Next section

Section 17 of 36

17. Use nested cross-validation

All feature selection and hyperparameter tuning belong inside each training fold. Selecting a cluster on the full dataset before cross-validation leaks outcome information.

A stratifying feature predicts or separates the recorded endpoint under the analysed design. It does not automatically explain why the endpoint differs.

Create an evidence card naming the contrast, measured modality, statistical unit, model, result, validation status, unresolved alternatives and narrowest defensible claim.

Contents · Previous section · Next section

Section 18 of 36

18. Read regularisation paths

A feature that enters only at a permissive penalty may be fragile. Compare selection frequency across folds and nearby regularisation values.

Citrus was developed for multidimensional cytometry, where panel design, compensation or normalisation, batch balance and viable-cell quality set the ceiling for inference.

Write association, stratification, differential abundance or differential state when that is what was tested. Reserve cause, diagnosis and benefit for evidence that directly supports them.

Contents · Previous section · Next section

Section 19 of 36

19. Estimate error honestly

Report held-out performance with uncertainty and class balance. Accuracy alone can look impressive when one group dominates; sensitivity, specificity or calibration may be more useful.

Citrus joins unsupervised hierarchical clustering to supervised endpoint modelling: cell subsets are discovered broadly, while endpoint labels are used to select informative per-sample features.

Archive sample manifests, preprocessing, feature roles, model formulae, random seeds, software versions and exclusions so another analyst can reconstruct the decision path.

Contents · Previous section · Next section

Section 20 of 36

20. Inspect cluster redundancy

Parent, child and sibling clusters can carry similar information. Collapse the biological story to the smallest coherent signature instead of counting every selected node as a discovery.

The original workflow can summarise cluster abundance and median functional-marker expression, so the scientific question must determine which feature family is interpretable.

Keep the biological sample as the unit of replication. Many cells improve resolution, but they do not turn one donor, animal or culture into many independent experiments.

Contents · Previous section · Next section

Section 21 of 36

21. Visualise markers and samples

Show marker distributions, sample-level feature values and cell counts beside the model score. A predictive coefficient without cellular context is hard to validate.

Hierarchical clusters are nested and correlated. A selected parent and child may describe one biological signal rather than several independent biomarkers.

Use negative controls, label permutations or design-matched null analyses where valid. A trustworthy signal should exceed artefacts that the same workflow can manufacture.

Contents · Previous section · Next section

Section 22 of 36

22. Test batch predictability

Attempt to predict batch or acquisition run using the same features. If technical labels are easier to predict than biology, rebalance or redesign before making biological claims.

Cross-validation and regularisation control selection inside the training data; truly independent samples are still needed to estimate generalisation and clinical usefulness.

Report effect direction, uncertainty, sample support and sensitivity—not only a colourful embedding, selected feature or adjusted value.

Contents · Previous section · Next section

Section 23 of 36

23. Compare manual gates

Expert gating supplies a familiar benchmark. Agreement can validate interpretation; disagreement should trigger marker-level inspection rather than automatic dismissal of either approach.

A stratifying feature predicts or separates the recorded endpoint under the analysed design. It does not automatically explain why the endpoint differs.

Create an evidence card naming the contrast, measured modality, statistical unit, model, result, validation status, unresolved alternatives and narrowest defensible claim.

Contents · Previous section · Next section

Section 24 of 36

24. Compare with diffcyt

diffcyt tests differential abundance or state across high-resolution clusters using explicit statistical designs. Citrus searches for a supervised multifeature signature that stratifies an endpoint.

Citrus was developed for multidimensional cytometry, where panel design, compensation or normalisation, batch balance and viable-cell quality set the ceiling for inference.

Write association, stratification, differential abundance or differential state when that is what was tested. Reserve cause, diagnosis and benefit for evidence that directly supports them.

Contents · Previous section · Next section

Section 25 of 36

25. Compare with cydar

cydar tests local hypersphere counts with spatial false-discovery control. Citrus uses nested clusters and supervised selection, so its inferential target and multiplicity structure differ.

Citrus joins unsupervised hierarchical clustering to supervised endpoint modelling: cell subsets are discovered broadly, while endpoint labels are used to select informative per-sample features.

Archive sample manifests, preprocessing, feature roles, model formulae, random seeds, software versions and exclusions so another analyst can reconstruct the decision path.

Contents · Previous section · Next section

Section 26 of 36

26. Build a signature evidence card

Place cluster path, markers, feature type, coefficient or score, fold stability, held-out performance, sample support and intended use on one record.

The original workflow can summarise cluster abundance and median functional-marker expression, so the scientific question must determine which feature family is interpretable.

Keep the biological sample as the unit of replication. Many cells improve resolution, but they do not turn one donor, animal or culture into many independent experiments.

Contents · Previous section · Next section

Section 27 of 36

27. Design orthogonal validation

Translate the selected subset into a simpler flow panel, targeted assay or sorted-cell experiment. Validation should test the same phenotype in new samples, not merely rerun Citrus.

Hierarchical clusters are nested and correlated. A selected parent and child may describe one biological signal rather than several independent biomarkers.

Use negative controls, label permutations or design-matched null analyses where valid. A trustworthy signal should exceed artefacts that the same workflow can manufacture.

Contents · Previous section · Next section

Section 28 of 36

28. Separate a biomarker from a mechanism

A feature may forecast an endpoint without mediating it. Perturbation, longitudinal sampling or functional experiments are required before causal language becomes credible.

Cross-validation and regularisation control selection inside the training data; truly independent samples are still needed to estimate generalisation and clinical usefulness.

Report effect direction, uncertainty, sample support and sensitivity—not only a colourful embedding, selected feature or adjusted value.

Contents · Previous section · Next section

Section 29 of 36

29. Connect Citrus to school science

The workflow turns classification, controls and fair testing into a real example: a rule that fits known examples must still face unseen examples.

A stratifying feature predicts or separates the recorded endpoint under the analysed design. It does not automatically explain why the endpoint differs.

Create an evidence card naming the contrast, measured modality, statistical unit, model, result, validation status, unresolved alternatives and narrowest defensible claim.

Contents · Previous section · Next section

Section 30 of 36

30. Try a nested-groups activity

Give students coloured objects with several measured properties, build broad-to-narrow groups and test which group feature predicts a hidden label in a separate set.

Citrus was developed for multidimensional cytometry, where panel design, compensation or normalisation, batch balance and viable-cell quality set the ceiling for inference.

Write association, stratification, differential abundance or differential state when that is what was tested. Reserve cause, diagnosis and benefit for evidence that directly supports them.

Contents · Previous section · Next section

Section 31 of 36

31. Connect to mathematics

Hierarchies, proportions, medians, regression penalties and cross-validation show how measurement becomes a prediction with quantifiable uncertainty.

Citrus joins unsupervised hierarchical clustering to supervised endpoint modelling: cell subsets are discovered broadly, while endpoint labels are used to select informative per-sample features.

Archive sample manifests, preprocessing, feature roles, model formulae, random seeds, software versions and exclusions so another analyst can reconstruct the decision path.

Contents · Previous section · Next section

Section 32 of 36

32. Connect to computing and AI literacy

Citrus is automated pattern discovery, yet the algorithm inherits every human decision about labels, features, quality filters and evaluation.

The original workflow can summarise cluster abundance and median functional-marker expression, so the scientific question must determine which feature family is interpretable.

Keep the biological sample as the unit of replication. Many cells improve resolution, but they do not turn one donor, animal or culture into many independent experiments.

Contents · Previous section · Next section

Section 33 of 36

33. Connect to school and career pathways

Students who enjoy cells and prediction can explore biology, chemistry, statistics, computing, biomedical engineering, laboratory medicine or bioinformatics through different education routes.

Hierarchical clusters are nested and correlated. A selected parent and child may describe one biological signal rather than several independent biomarkers.

Use negative controls, label permutations or design-matched null analyses where valid. A trustworthy signal should exceed artefacts that the same workflow can manufacture.

Contents · Previous section · Next section

Section 34 of 36

34. Create a family evidence habit

When a report calls a cell subset a biomarker, ask whether it was discovered and tested on separate people and whether the assay can be repeated reliably.

Cross-validation and regularisation control selection inside the training data; truly independent samples are still needed to estimate generalisation and clinical usefulness.

Report effect direction, uncertainty, sample support and sensitivity—not only a colourful embedding, selected feature or adjusted value.

Contents · Previous section · Next section

Section 35 of 36

35. Use precise Citrus language

Say selected cluster feature, stratifying signature and held-out performance. Avoid diagnostic, prognostic or mechanistic claims unless the study design validates that use.

A stratifying feature predicts or separates the recorded endpoint under the analysed design. It does not automatically explain why the endpoint differs.

Create an evidence card naming the contrast, measured modality, statistical unit, model, result, validation status, unresolved alternatives and narrowest defensible claim.

Contents · Previous section · Next section

Section 36 of 36

36. Finish with a discovery-to-validation bundle

Deliver manifests, quality plots, hierarchy settings, feature tables, nested resampling, selection stability, independent validation, evidence cards and a predeclared next experiment.

Citrus was developed for multidimensional cytometry, where panel design, compensation or normalisation, batch balance and viable-cell quality set the ceiling for inference.

Write association, stratification, differential abundance or differential state when that is what was tested. Reserve cause, diagnosis and benefit for evidence that directly supports them.

## Source date and scope note Primary papers and official software documentation were checked for this article on 10 October 2026. Software interfaces and recommendations can change, so readers should consult the linked current documentation before reproducing an analysis.

## Final reader checklist Before accepting a claim, confirm the biological sample count, condition definition, measured modality, preprocessing, statistical unit, model formula, uncertainty, software version, negative controls, sensitivity checks, independent validation and the exact boundary between association and causation.

Contents · Previous section · Continue to the Science Learning Hub

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading