eduKateSG · Why Science?
Let each log-count observation carry an evidence weight—then use flexible linear models without pretending RNA-seq variance is constant
Reading routes
Science learning becomes useful when a familiar object or observation is turned into a system of quantities, mechanisms and claim limits. This guide owns one applied evidence-reading job inside eduKateSG’s wider Science estate. It connects naturally to Why Science Single Cell Rna Sequencing Barcodes Transcriptome Heterogeneity Evidence; Why Science Dreamlet Pseudobulk Mixed Models Complex Single Cell Cohorts; Why Science Muscat Multi Sample Multi Group Differential State Analysis; Why Science Distinct Full Distributions Multi Sample Single Cell Evidence; Education Hub; Singapore Secondary School Directory; Career Adulthood Hub; Science Learning Hub. It also keeps current school and public claims traceable to visible primary sources: voom primary study; Official limma User's Guide; Official Bioconductor limma package; 2026 Singapore–Cambridge O-Level Biology syllabus; MOE G2/G3 Lower Secondary Science syllabus; 2026 MOE G2 Computing syllabus. The sources describe the scientific scope; this article translates that scope into a calm route for Primary Science, PSLE Science, Secondary Science, O-Level Science, STEM exploration, school choices and career pathways without inventing admission or employment outcomes.
The voom method, presented by Law and colleagues in Genome Biology in 2014, estimates the mean–variance relationship in log-counts per million and assigns observation-level precision weights before fitting limma linear models. Limma then provides flexible contrasts and empirical-Bayes moderation across genes. The 2026 user guide distinguishes voom, limma-trend, quality weights and specialised extensions. A weighted model remains conditional on library preparation, filtering, normalisation, design rank, replication and the exact contrast being interpreted.
Inside this guide
1–12 · Foundations and models
- 1. Begin with unequal precision
- 2. Define counts per million
- 3. Understand observation weights
- 4. Meet the linear model
- 5. Understand empirical Bayes moderation
- 6. Choose the right voom family
- 7. Read the 2014 voom evidence in scope
- 8. Start from documented counts
- 9. Filter weakly expressed features
- 10. Normalise effective library sizes
- 11. Construct the design matrix
- 12. Run voom with diagnostic plotting
13–24 · Evidence, testing and applications
- 13. Inspect the weight distribution
- 14. Consider sample-quality weights
- 15. Practise with a fictional voom table
- 16. Fit weighted gene-wise models
- 17. Specify contrasts visibly
- 18. Apply empirical Bayes moderation
- 19. Control false discoveries
- 20. Rank with effect and precision
- 21. Inspect residuals by sample
- 22. Check the mean–variance trend again
- 23. Audit low-count sensitivity
- 24. Audit sample influence
25–36 · Learning, decisions and pathways
- 25. Model repeated measures appropriately
- 26. Separate batch from condition
- 27. Compare with DESeq2
- 28. Build a voom evidence card
- 29. Connect voom to school science
- 30. Try a weighted-measurement activity
- 31. Connect to mathematics
- 32. Connect to computing and AI literacy
- 33. Connect to school and career pathways
- 34. Create a family evidence habit
- 35. Use precise limma-voom language
- 36. Finish with a weights-to-claim bundle
Section 1 of 36
1. Begin with unequal precision
RNA-seq log-counts do not have constant variance across abundance. Voom begins by estimating how measurement precision changes with the fitted expression level.
Voom converts counts to log-counts per million and estimates a mean–variance trend, producing observation-level precision weights for linear modelling.
Archive the sample manifest, count provenance, feature identifiers, formula, contrast, filters, package version, random seeds and exclusions so another analyst can reconstruct the decision path.
Section 2 of 36
2. Define counts per million
Counts per million adjust counts for library size on a relative scale. The logarithm improves visual and modelling behaviour, while offsets and small-count handling still matter.
The weights quantify estimated reliability on the transformed scale. They are not probabilities that a gene is true or important.
Keep the biological sample as the unit of replication. More reads or cells can improve measurement, but they do not create more independent people, animals or cultures.
Section 3 of 36
3. Understand observation weights
Each gene–sample observation receives an estimated precision weight. Larger weights mean greater model influence, not greater biological importance.
Limma fits gene-wise linear models and moderates variance estimates across genes with empirical Bayes methods, improving stability in small replicated studies.
Pair adjusted evidence with effect direction, uncertainty, sample-level plots and sensitivity checks; a short ranked table is not a complete scientific result.
Section 4 of 36
4. Meet the linear model
The design matrix describes expected log expression as a combination of coefficients. Linear modelling makes complex comparisons possible when the design is full rank.
Contrasts are algebraic questions asked of fitted coefficients. A correct contrast must match the design coding and the biological comparison.
Use negative controls, simulated nulls or label permutations only when their assumptions match the design, and name the particular false signal each check could reveal.
Section 5 of 36
5. Understand empirical Bayes moderation
Limma shares information across genes to stabilise residual variance estimates. Moderation helps small studies but cannot substitute for missing experimental units.
Quality weights and observation weights solve different problems: one can downweight a sample, while the other follows the mean-dependent precision pattern.
Create an evidence card naming the question, measured material, statistical unit, model, comparison, result, validation status, alternatives and narrowest defensible claim.
Section 6 of 36
6. Choose the right voom family
Standard voom, limma-trend, sample-quality weights and specialised voomLmFit routes address different mean–variance or sample-quality structures. Follow current documentation for the design at hand.
Weighted residuals, library-level plots and design-rank checks reveal failures that a smooth mean–variance curve alone can miss.
Write association or differential expression when that is what was tested. Reserve cause, mechanism, diagnosis and benefit for designs with stronger supporting evidence.
Section 7 of 36
7. Read the 2014 voom evidence in scope
Law and colleagues showed how precision weights unlock limma’s linear-model tools for RNA-seq counts. The paper evaluates statistical performance, not biological truth for every detected gene.
Voom converts counts to log-counts per million and estimates a mean–variance trend, producing observation-level precision weights for linear modelling.
Archive the sample manifest, count provenance, feature identifiers, formula, contrast, filters, package version, random seeds and exclusions so another analyst can reconstruct the decision path.
Section 8 of 36
8. Start from documented counts
Record quantification, annotation, transcript-to-gene summarisation and matrix orientation. Do not feed already normalised or transformed expression into a workflow expecting counts.
The weights quantify estimated reliability on the transformed scale. They are not probabilities that a gene is true or important.
Keep the biological sample as the unit of replication. More reads or cells can improve measurement, but they do not create more independent people, animals or cultures.
Section 9 of 36
9. Filter weakly expressed features
Retain features with enough expression across relevant libraries to support a stable variance estimate. Base the rule on the design rather than the outcome labels alone.
Limma fits gene-wise linear models and moderates variance estimates across genes with empirical Bayes methods, improving stability in small replicated studies.
Pair adjusted evidence with effect direction, uncertainty, sample-level plots and sensitivity checks; a short ranked table is not a complete scientific result.
Section 10 of 36
10. Normalise effective library sizes
A common workflow uses edgeR objects and composition factors before voom. Record these factors because voom uses effective, not merely raw, library sizes.
Contrasts are algebraic questions asked of fitted coefficients. A correct contrast must match the design coding and the biological comparison.
Use negative controls, simulated nulls or label permutations only when their assumptions match the design, and name the particular false signal each check could reveal.
Section 11 of 36
11. Construct the design matrix
Represent conditions, batches, donors, time points and interactions only when supported by the samples. Inspect column names and rank before fitting.
Quality weights and observation weights solve different problems: one can downweight a sample, while the other follows the mean-dependent precision pattern.
Create an evidence card naming the question, measured material, statistical unit, model, comparison, result, validation status, alternatives and narrowest defensible claim.
Section 12 of 36
12. Run voom with diagnostic plotting
Estimate the mean–variance trend and inspect the curve, points and residual standard-deviation pattern. A smooth line is a model component, not automatic validation.
Weighted residuals, library-level plots and design-rank checks reveal failures that a smooth mean–variance curve alone can miss.
Write association or differential expression when that is what was tested. Reserve cause, mechanism, diagnosis and benefit for designs with stronger supporting evidence.
Section 13 of 36
13. Inspect the weight distribution
Compare weights across abundance, samples and groups. Systematically low weights in one library may signal quality problems requiring investigation.
Voom converts counts to log-counts per million and estimates a mean–variance trend, producing observation-level precision weights for linear modelling.
Archive the sample manifest, count provenance, feature identifiers, formula, contrast, filters, package version, random seeds and exclusions so another analyst can reconstruct the decision path.
Section 14 of 36
14. Consider sample-quality weights
When whole libraries differ in reliability, quality weights may complement observation weights. Use them because diagnostics justify the model, not because they improve a preferred ranking.
The weights quantify estimated reliability on the transformed scale. They are not probabilities that a gene is true or important.
Keep the biological sample as the unit of replication. More reads or cells can improve measurement, but they do not create more independent people, animals or cultures.
Section 15 of 36
15. Practise with a fictional voom table
This classroom table is invented and is not a result from the voom paper.
| Fictional gene | logCPM | Precision weight | Moderated t | Reading |
|---|---|---|---|---|
| GENE-V | 7.2 | 1.8 | 5.1 | Precise difference |
| GENE-W | 1.0 | 0.3 | 1.2 | Noisy low count |
| GENE-X | 5.4 | 0.9 | -3.7 | Follow direction |
Limma fits gene-wise linear models and moderates variance estimates across genes with empirical Bayes methods, improving stability in small replicated studies.
Pair adjusted evidence with effect direction, uncertainty, sample-level plots and sensitivity checks; a short ranked table is not a complete scientific result.
Section 16 of 36
16. Fit weighted gene-wise models
lmFit uses the design and weights to estimate coefficients for every gene. Preserve the fitted object and the exact commands so contrasts can be audited.
Contrasts are algebraic questions asked of fitted coefficients. A correct contrast must match the design coding and the biological comparison.
Use negative controls, simulated nulls or label permutations only when their assumptions match the design, and name the particular false signal each check could reveal.
Section 17 of 36
17. Specify contrasts visibly
Use makeContrasts or direct coefficient tests only after translating the comparison into plain language. Confirm reference levels and sign direction.
Quality weights and observation weights solve different problems: one can downweight a sample, while the other follows the mean-dependent precision pattern.
Create an evidence card naming the question, measured material, statistical unit, model, comparison, result, validation status, alternatives and narrowest defensible claim.
Section 18 of 36
18. Apply empirical Bayes moderation
eBayes or treat moderates gene-wise variance information. treat asks whether effects exceed a chosen fold-change threshold rather than merely differing from zero.
Weighted residuals, library-level plots and design-rank checks reveal failures that a smooth mean–variance curve alone can miss.
Write association or differential expression when that is what was tested. Reserve cause, mechanism, diagnosis and benefit for designs with stronger supporting evidence.
Section 19 of 36
19. Control false discoveries
Adjust across the declared gene set and comparisons. The adjusted value answers a testing-family question, not the probability that a gene is biologically useful.
Voom converts counts to log-counts per million and estimates a mean–variance trend, producing observation-level precision weights for linear modelling.
Archive the sample manifest, count provenance, feature identifiers, formula, contrast, filters, package version, random seeds and exclusions so another analyst can reconstruct the decision path.
Section 20 of 36
20. Rank with effect and precision
Examine log fold change, average expression, moderated statistics, uncertainty and sample plots together. A strong statistic can accompany a modest effect when precision is high.
The weights quantify estimated reliability on the transformed scale. They are not probabilities that a gene is true or important.
Keep the biological sample as the unit of replication. More reads or cells can improve measurement, but they do not create more independent people, animals or cultures.
Section 21 of 36
21. Inspect residuals by sample
Plot residual summaries and sample relationships after fitting. Batch structure or nonlinearity left in residuals can invalidate a simple contrast.
Limma fits gene-wise linear models and moderates variance estimates across genes with empirical Bayes methods, improving stability in small replicated studies.
Pair adjusted evidence with effect direction, uncertainty, sample-level plots and sensitivity checks; a short ranked table is not a complete scientific result.
Section 22 of 36
22. Check the mean–variance trend again
Review whether weighted residual variation is approximately stabilised. Persistent abundance-dependent patterns may call for a different voom or trend setting.
Contrasts are algebraic questions asked of fitted coefficients. A correct contrast must match the design coding and the biological comparison.
Use negative controls, simulated nulls or label permutations only when their assumptions match the design, and name the particular false signal each check could reveal.
Section 23 of 36
23. Audit low-count sensitivity
Repeat headline analyses after reasonable filtering changes. Very low-count genes should not dominate scientific interpretation.
Quality weights and observation weights solve different problems: one can downweight a sample, while the other follows the mean-dependent precision pattern.
Create an evidence card naming the question, measured material, statistical unit, model, comparison, result, validation status, alternatives and narrowest defensible claim.
Section 24 of 36
24. Audit sample influence
Use leave-one-library checks or sample weights to see whether one observation creates the effect. Report influential samples instead of silently removing them.
Weighted residuals, library-level plots and design-rank checks reveal failures that a smooth mean–variance curve alone can miss.
Write association or differential expression when that is what was tested. Reserve cause, mechanism, diagnosis and benefit for designs with stronger supporting evidence.
Section 25 of 36
25. Model repeated measures appropriately
For repeated observations, use supported correlation or mixed-model extensions and enough subjects. Repeated measurements are not independent replication.
Voom converts counts to log-counts per million and estimates a mean–variance trend, producing observation-level precision weights for linear modelling.
Archive the sample manifest, count provenance, feature identifiers, formula, contrast, filters, package version, random seeds and exclusions so another analyst can reconstruct the decision path.
Section 26 of 36
26. Separate batch from condition
A design cannot recover a condition effect when batch and condition are perfectly aligned. Narrow the claim or redesign the experiment.
The weights quantify estimated reliability on the transformed scale. They are not probabilities that a gene is true or important.
Keep the biological sample as the unit of replication. More reads or cells can improve measurement, but they do not create more independent people, animals or cultures.
Section 27 of 36
27. Compare with DESeq2
DESeq2 models counts directly with negative-binomial GLMs; voom uses precision-weighted log-count models. Compare them only with aligned inputs, design and contrasts.
Limma fits gene-wise linear models and moderates variance estimates across genes with empirical Bayes methods, improving stability in small replicated studies.
Pair adjusted evidence with effect direction, uncertainty, sample-level plots and sensitivity checks; a short ranked table is not a complete scientific result.
Section 28 of 36
28. Build a voom evidence card
Record counts, filters, scaling factors, design, voom variant, weight diagnostics, contrasts, moderation, testing family, sensitivity and biological validation.
Contrasts are algebraic questions asked of fitted coefficients. A correct contrast must match the design coding and the biological comparison.
Use negative controls, simulated nulls or label permutations only when their assumptions match the design, and name the particular false signal each check could reveal.
Section 29 of 36
29. Connect voom to school science
Repeated measurements do not all deserve equal confidence. Precision weighting formalises the familiar idea that reliable measurements should influence a conclusion more.
Quality weights and observation weights solve different problems: one can downweight a sample, while the other follows the mean-dependent precision pattern.
Create an evidence card naming the question, measured material, statistical unit, model, comparison, result, validation status, alternatives and narrowest defensible claim.
Section 30 of 36
30. Try a weighted-measurement activity
Combine thermometer readings with different stated uncertainties. Compare an ordinary average with a reliability-weighted summary and discuss when weighting can mislead.
Weighted residuals, library-level plots and design-rank checks reveal failures that a smooth mean–variance curve alone can miss.
Write association or differential expression when that is what was tested. Reserve cause, mechanism, diagnosis and benefit for designs with stronger supporting evidence.
Section 31 of 36
31. Connect to mathematics
Logarithms, regression, residuals, variance trends, weights and moderated t-statistics show how mathematical models adapt to changing precision.
Voom converts counts to log-counts per million and estimates a mean–variance trend, producing observation-level precision weights for linear modelling.
Archive the sample manifest, count provenance, feature identifiers, formula, contrast, filters, package version, random seeds and exclusions so another analyst can reconstruct the decision path.
Section 32 of 36
32. Connect to computing and AI literacy
A smooth algorithmic result rests on encoded assumptions. Inspect weights and design columns the way one would inspect features and labels in an AI system.
The weights quantify estimated reliability on the transformed scale. They are not probabilities that a gene is true or important.
Keep the biological sample as the unit of replication. More reads or cells can improve measurement, but they do not create more independent people, animals or cultures.
Section 33 of 36
33. Connect to school and career pathways
Precision modelling links biology, mathematics, statistics, computing, medicine and data science through multiple education choices.
Limma fits gene-wise linear models and moderates variance estimates across genes with empirical Bayes methods, improving stability in small replicated studies.
Pair adjusted evidence with effect direction, uncertainty, sample-level plots and sensitivity checks; a short ranked table is not a complete scientific result.
Section 34 of 36
34. Create a family evidence habit
When a gene-expression plot looks decisive, ask how precision was estimated, how many samples contributed and whether one library was downweighted.
Contrasts are algebraic questions asked of fitted coefficients. A correct contrast must match the design coding and the biological comparison.
Use negative controls, simulated nulls or label permutations only when their assumptions match the design, and name the particular false signal each check could reveal.
Section 35 of 36
35. Use precise limma-voom language
Say the weighted linear model supports a differential-expression estimate under the stated contrast. Avoid claiming a direct regulatory mechanism without further experiments.
Quality weights and observation weights solve different problems: one can downweight a sample, while the other follows the mean-dependent precision pattern.
Create an evidence card naming the question, measured material, statistical unit, model, comparison, result, validation status, alternatives and narrowest defensible claim.
Section 36 of 36
36. Finish with a weights-to-claim bundle
Deliver counts, metadata, scaling, code, design, voom plots, weights, fits, contrasts, complete tables, sample diagnostics, sensitivity and validation.
Weighted residuals, library-level plots and design-rank checks reveal failures that a smooth mean–variance curve alone can miss.
Write association or differential expression when that is what was tested. Reserve cause, mechanism, diagnosis and benefit for designs with stronger supporting evidence.
## Source date and scope note Primary studies and current official software documentation were checked for this article on 11 October 2026. Software interfaces and recommendations can change, so readers should consult the linked current documentation before reproducing an analysis.
## Final reader checklist Before accepting a claim, confirm the biological sample count, condition definition, measured material, preprocessing, statistical unit, design matrix, exact contrast, effect size, uncertainty, software version, negative controls, sensitivity checks, independent validation and the boundary between association and causation.
Contents · Previous section · Continue to the Science Learning Hub
