eduKateSG · Why Science?
Read accessible chromatin and RNA from the same cell, connect distal sites with genes and test regulatory timing—while keeping sparsity, barcode collisions and correlation visible
Reading routes
Science learning becomes useful when a familiar object or observation is turned into a system of quantities, mechanisms and claim limits. This guide owns one applied evidence-reading job inside eduKateSG’s wider Science estate. It connects naturally to Why Science Single Cell Rna Sequencing Barcodes Transcriptome Heterogeneity Evidence; Why Science Atac Seq Transposase Accessible Chromatin Evidence; Why Science Cite Seq Oligonucleotide Antibody Tags Rna Protein Evidence; Why Science Spatial Transcriptomics Tissue Coordinates Gene Expression Evidence; Education Hub; Singapore Secondary School Directory; Career Adulthood Hub. It also keeps current school and public claims traceable to visible primary sources: Foundational SHARE-seq primary study; Foundational SHARE-seq PubMed record; Foundational SHARE-seq DOI; 2026 Singapore–Cambridge O-Level Chemistry syllabus; 2026 Singapore–Cambridge O-Level Biology syllabus. The sources describe the scientific scope; this article translates that scope into a calm route for Primary Science, PSLE Science, Secondary Science, O-Level Science, STEM exploration, school choices and career pathways without inventing admission or employment outcomes.
SHARE-seq—Simultaneous High-throughput ATAC and RNA Expression with sequencing—profiles chromatin accessibility and gene expression from the same cell by repeatedly splitting, barcoding and pooling fixed cells or nuclei. The foundational study applied the method to mouse skin and reported 34,774 joint profiles, using distal elements co-accessible with promoters to define domains of regulatory chromatin, or DORCs. Some accessibility changes preceded expression along differentiation trajectories. That is valuable temporal evidence, not automatic proof that an accessible site caused a gene to change. RNA capture, ATAC fragments, barcode collisions, doublets, trajectory assumptions, cell-state composition and validation determine the strength of each claim.
Inside this guide
1–12 · Foundations and models
- 1. Begin with a regulatory timing question
- 2. Did you know one cell can carry two molecular diaries?
- 3. Define accessibility carefully
- 4. Define expression as sampled RNA
- 5. Why same-cell pairing matters
- 6. Keep association below mechanism
- 7. Plan specimens before barcodes
- 8. Preserve cells or nuclei for joint chemistry
- 9. Tag accessible DNA with transposase
- 10. Copy RNA before repeated pooling
- 11. Split, barcode, pool and repeat
- 12. Distinguish collisions from doublets
13–24 · Evidence, testing and applications
- 13. Separate the two libraries transparently
- 14. Sequence for the biological question
- 15. Practise with an invented SHARE-seq table
- 16. Demultiplex without wishful rescue
- 17. Call accessible regions without circularity
- 18. Normalise each modality on its own terms
- 19. Build joint cell states, then test them
- 20. Link distal elements to genes cautiously
- 21. Understand DORCs as an operational definition
- 22. Test whether accessibility precedes expression
- 23. Challenge the trajectory
- 24. Challenge sparse zeros
25–36 · Learning, decisions and pathways
- 25. Challenge barcode mixtures
- 26. Challenge tissue dissociation
- 27. Validate the decisive regulatory link
- 28. Build Primary Science habits
- 29. Prepare for PSLE Science data questions
- 30. Deepen Secondary Science reasoning
- 31. Connect to O-Level Science
- 32. Use school choices as fit questions
- 33. See the career pathway ecosystem
- 34. Know when science tuition helps
- 35. Did you know accessibility can lead without causing?
- 36. Finish with an evidence ladder
Section 1 of 36
1. Begin with a regulatory timing question
Genes are not simply switched on because nearby DNA is open. SHARE-seq measures accessible chromatin and RNA in the same cell so researchers can ask whether regulatory regions become accessible before, with or after transcriptional change. The design converts a vague story about control into two linked, testable measurements.
Section 2 of 36
2. Did you know one cell can carry two molecular diaries?
An ATAC-derived library records transposase-accessible DNA fragments, while an RNA-derived library records captured transcripts. Shared cellular barcodes connect both diaries to the same cell. This removes uncertain matching between separate cells, but it does not make either diary complete; both remain sparse samples.
Section 3 of 36
3. Define accessibility carefully
Accessible chromatin is DNA that the transposase can reach under the assay conditions. It often marks promoters, enhancers and other regulatory regions, yet accessibility is not equivalent to activity, factor binding or causation. Nucleosome position, fixation, cell quality and sequence bias can all change the recovered fragments.
Section 4 of 36
4. Define expression as sampled RNA
RNA counts reflect captured and sequenced molecules, not every transcript present in the cell. Dropout is common, especially for low-abundance genes. Transcript abundance also reflects production and degradation. A zero count therefore means not detected, which is different from proving that a gene was not expressed.
Section 5 of 36
5. Why same-cell pairing matters
When two measurements come from one barcode, researchers can compare accessibility and expression without first guessing which ATAC-only cell matches which RNA-only cell. That is a major evidential improvement. It still requires proof that the barcode belongs to one cell and that both modalities passed quality control.
Section 6 of 36
6. Keep association below mechanism
A distal site whose accessibility covaries with a gene is a candidate regulatory element. The relationship may arise from shared cell state, developmental time or a third regulator. Language should progress from associated, to temporally ordered, to functionally tested. Only perturbation can directly test whether changing the site changes the gene.
Section 7 of 36
7. Plan specimens before barcodes
Specify tissues, independent animals, conditions, expected cell states, fixation batches and target cell numbers before processing. Randomise biological groups across plates where possible. Thousands of barcoded cells from one specimen remain one biological replicate for specimen-level conclusions, however impressive the cellular total appears.
Section 8 of 36
8. Preserve cells or nuclei for joint chemistry
SHARE-seq begins by fixing and permeabilising material so chromatin and RNA can survive repeated reactions. Too much fixation can reduce enzyme access; too little can leak RNA or mix molecules. Record times, concentrations, temperatures, washes and microscopic integrity for every preparation.
Section 9 of 36
9. Tag accessible DNA with transposase
A loaded transposase inserts adaptor sequences into accessible chromatin. Fragment yield, transcription-start-site enrichment, mitochondrial fraction and nucleosomal pattern help assess the result. The reaction is a biochemical sampling step, so enzyme lot, input concentration and incubation belong in the evidence record.
Section 10 of 36
10. Copy RNA before repeated pooling
Reverse transcription converts captured RNA into cDNA that can receive the cellular indexing history. RNA quality, priming and reverse-transcriptase efficiency shape which genes appear. Include RNA-focused controls and remember that excellent chromatin data cannot rescue a failed transcriptome from the same barcode.
Section 11 of 36
11. Split, barcode, pool and repeat
Cells or nuclei move through successive wells, acquiring an ordered series of oligonucleotide barcodes. The combination creates a large address space without a separate tube for every cell. Accurate plate maps, balanced wells and barcode sequences with useful edit distance are essential to preserve identity.
Section 12 of 36
12. Distinguish collisions from doublets
A collision assigns unrelated material to the same barcode combination; a doublet carries two cells through the workflow together. Both can create apparently hybrid regulatory states. Species-mixing controls, unusually high counts, incompatible markers and barcode-occupancy models detect different parts of the problem and should be reported together.
Section 13 of 36
13. Separate the two libraries transparently
After cellular identities are encoded, DNA- and RNA-derived products are amplified and prepared for sequencing. The read structure should show where modality tags, cellular indexes and molecular sequences occur. Publish the demultiplexing rules and an evidence funnel from raw reads to accepted cells in each modality.
Section 14 of 36
14. Sequence for the biological question
A pilot estimates unique ATAC fragments, RNA molecules, genes, duplicates and saturation. Rare-state discovery may need more cells; promoter-level linkage may need more information per cell. Additional reads help only while the libraries retain unseen molecules. Depth, cell count and biological replication solve different limitations.
Section 15 of 36
15. Practise with an invented SHARE-seq table
These fictional values illustrate joint quality review; they are not performance benchmarks.
| Cell barcode | ATAC fragments | RNA molecules | Mixed markers | First reading |
|---|---|---|---|---|
| S-041 | 18,400 | 7,200 | no | retain |
| S-042 | 1,050 | 6,900 | no | weak ATAC |
| S-043 | 20,100 | 420 | no | weak RNA |
| S-044 | 39,700 | 14,600 | yes | inspect doublet |
Section 16 of 36
16. Demultiplex without wishful rescue
Specify exact-match and mismatch policies for every index. Retain index qualities and unused combinations. Permissive correction can move reads between plausible cells, especially when one barcode is abundant. Repeat important conclusions under strict demultiplexing so rescued molecules do not manufacture a rare joint state.
Section 17 of 36
17. Call accessible regions without circularity
Pool suitable cells or use a declared peak set, then quantify cell-level fragments against it. If clusters define peaks and the same peaks prove the clusters, validation becomes circular. A held-out specimen, consensus reference or alternative feature set shows whether the cell states generalise.
Section 18 of 36
18. Normalise each modality on its own terms
ATAC fragments and RNA molecules have different distributions, biases and sparsity. Process them with modality-appropriate quality metrics before integration. A single blended score can hide an excellent RNA profile paired with unusable chromatin, or the reverse. Display the two-dimensional quality landscape.
Section 19 of 36
19. Build joint cell states, then test them
A joint embedding can sharpen cell-state separation because chromatin and RNA contribute complementary information. Colour it by specimen, batch, read depth and cell cycle before biological interpretation. Confirm markers in raw counts and ask whether clusters remain when either modality or one specimen is held out.
Section 20 of 36
20. Link distal elements to genes cautiously
Co-accessibility and correlation can identify distal sites whose accessibility moves with a promoter or transcript. Genomic distance, cell composition and shared trajectories can create similar patterns. Report the candidate rule, null model, multiple-testing control and independent evidence rather than labelling every linked site an enhancer.
Section 21 of 36
21. Understand DORCs as an operational definition
The foundational study described domains of regulatory chromatin, or DORCs, around genes with many associated accessibility sites. DORCs can highlight regulatory complexity and overlap known regulatory landscapes. They are analysis-defined domains, not physical containers, and their boundaries depend on features, thresholds and available cells.
Section 22 of 36
22. Test whether accessibility precedes expression
Order cells along a defensible differentiation trajectory, estimate when accessibility and RNA change, and quantify uncertainty. Temporal precedence is more informative than simultaneous correlation, but pseudotime is not a clock. Repeat across independent specimens and observed stages before proposing that chromatin primes later transcription.
Section 23 of 36
23. Challenge the trajectory
Different roots, neighbours, smoothing settings and branch assignments can change apparent lead–lag order. Compare RNA-only, ATAC-only and joint trajectories. Check cell cycle, batch and specimen composition. A timing conclusion is stronger when it survives several reasonable models and matches independently observed developmental stages.
Section 24 of 36
24. Challenge sparse zeros
A missing ATAC fragment or transcript often reflects sampling. Aggregate only after defining biologically coherent groups and report the number of cells, specimens and molecules behind each curve. Downsample high-depth groups. Claims that disappear after coverage matching may describe measurement opportunity rather than regulatory biology.
Section 25 of 36
25. Challenge barcode mixtures
Cells with unusually high counts or incompatible lineage markers can dominate accessibility–RNA correlations. Estimate doublets and collisions, then repeat headline analyses after removing high-risk barcodes. State how many cells and which biological groups are lost. Filtering should not quietly erase a real transitional population.
Section 26 of 36
26. Challenge tissue dissociation
Enzymatic and mechanical handling can induce stress RNA and change representation of fragile cell types. Chromatin may appear relatively stable while RNA changes rapidly. Shorten processing, compare nuclei with whole cells where relevant, inspect stress programmes and balance batches before treating a joint stress state as normal development.
Section 27 of 36
27. Validate the decisive regulatory link
Use targeted accessibility assays, reporter tests, chromatin conformation, imaging or CRISPR perturbation according to the claim. A candidate element becomes much more convincing if its deletion changes the predicted gene in the relevant cell state. Validation should target the weakest inference, not merely repeat the easiest measurement.
Section 28 of 36
28. Build Primary Science habits
Children can sort picture cards into ‘observed now’ and ‘possible cause later’. That simple distinction mirrors the difference between accessible DNA, measured RNA and a regulatory explanation. Asking what was measured, what changed and what remains a guess develops careful observation without requiring molecular vocabulary.
Section 29 of 36
29. Prepare for PSLE Science data questions
Use two simple graphs for the same fictional cells: one for an accessibility signal and one for RNA. Ask which pattern rises first, whether every cell follows it and what extra test is needed. This practises trend reading, fair comparison and the difference between association and proof.
Section 30 of 36
30. Deepen Secondary Science reasoning
Secondary Science students can map the workflow as specimen, chemistry, barcode, sequence, table and claim. At each arrow, identify a possible error and a control. This turns an advanced method into familiar ideas about variables, reliability, replication and the limits of an instrument.
Section 31 of 36
31. Connect to O-Level Science
The official 2026 Singapore–Cambridge Biology and Chemistry syllabuses reward evidence-based explanations, experimental planning and evaluation. SHARE-seq provides a modern context for enzymes, nucleic acids, chemical conditions, measurement uncertainty and valid conclusions. Students should apply syllabus concepts rather than memorise a brand name.
Section 32 of 36
32. Use school choices as fit questions
Families comparing schools can ask how students encounter inquiry: practical work, data interpretation, research exposure, computing, communication and mentoring. Confirm current opportunities on official school pages and at open houses. One sophisticated method does not define a school; sustained habits and student fit matter more.
Section 33 of 36
33. See the career pathway ecosystem
Related work spans molecular biology, genomics, bioinformatics, statistics, software engineering, laboratory operations, instrument support, clinical research, data governance and science communication. Progress can begin through junior college, polytechnic or other routes. Careers develop through skills, projects and further training, not one predetermined subject combination.
Section 34 of 36
34. Know when science tuition helps
Science tuition can be useful when it diagnoses a precise gap: weak graph reading, imprecise explanations, poor experimental evaluation or fragile prerequisite knowledge. It should not replace sleep, school feedback or independent practice. A good tutor asks learners to justify evidence and transfer the reasoning to unfamiliar contexts.
Section 35 of 36
35. Did you know accessibility can lead without causing?
In the foundational SHARE-seq analysis, some chromatin-accessibility changes preceded gene-expression changes along differentiation. That makes regulatory priming plausible and testable. It does not mean every early-opening site caused the later transcript. The happy lesson is that a good dataset can reveal exactly which experiment should come next.
Section 36 of 36
36. Finish with an evidence ladder
A careful SHARE-seq conclusion climbs from valid barcodes, to usable ATAC and RNA libraries, to reproducible cell states, to robust element–gene associations, to timing across specimens, and finally to targeted functional tests. The method matters because it joins two molecular views while making their remaining uncertainties visible.
A publication-ready SHARE-seq study needs an audit trail that begins before cells enter a well. List independent specimens, tissue regions, collection times, dissociation or nuclei-isolation conditions, fixation batches, target cell states and planned comparisons. Randomise biological groups across processing units when possible and preserve specimen identity in every downstream table. The statistical unit for an animal- or treatment-level claim is the independent specimen, not every barcode recovered from it. Report both because they answer different questions: cells reveal heterogeneity, while specimens support generalisation.
The molecular record should make each modality independently reproducible. For accessible chromatin, provide transposase conditions, adaptor sequences, fragment-quality rules, mapping settings, duplicate handling, transcription-start-site enrichment and accepted fragments. For RNA, provide priming, reverse-transcription conditions, read structure, transcript assignment, unique-molecule rules and detected-gene counts. For the cellular link, provide the complete split-pool index design, whitelist, mismatch policy, index-quality distributions and unused combinations. A reader should be able to trace a raw read to its modality, cell and final feature without relying on software names alone.
Barcode validation deserves a full analysis rather than one doublet score. A species-mixing experiment tests some cross-cell assignments, but same-species collisions may be harder to recognise. Combine occupancy calculations, count outliers, mutually exclusive lineage markers, synthetic mixtures and unused-index leakage. Report how many cells each rule removes and whether one state or specimen is disproportionately affected. Re-run headline element–gene links after excluding high-risk cells. A regulatory relationship that survives stricter identity rules is much more credible than one driven by a few unusually rich barcodes.
Accessibility and RNA quality should remain visibly two-dimensional. Plot unique ATAC fragments against RNA molecules, colour by specimen and show accepted thresholds. Then repeat for enrichment, genes, duplicates and cell-cycle or stress scores. Avoid a composite score that lets an excellent transcriptome compensate for chromatin failure. When the biological question concerns a DORC or regulatory lead–lag pattern, specify the minimum usable information in both channels. Include sensitivity results under more and less stringent thresholds so readers can see whether filtering determines the conclusion.
Peak and feature selection can introduce circularity. If cell clusters define accessible regions and those same regions validate the clusters, the evidence has looped back on itself. Use a consensus peak set, a held-out specimen, external annotations or cross-validation. For distal element–gene links, state genomic windows, correlation measure, covariates, matched nulls and multiple-testing procedure. Compare results after controlling for total fragments, RNA depth, cell state and pseudotime. A link should be described as a candidate even when it recurs, unless direct functional evidence tests the regulatory relationship.
DORC analyses need their operational choices on the page. Report how candidate distal sites were selected, how promoters were defined, which association threshold was used and how many cells supported each domain. Compare DORCs across alternative peak sets and subsampled specimens. Overlap with known super-enhancers or regulatory annotations can strengthen interpretation, but it is not independent proof of function if those annotations were created from related tissues or signals. Display the raw accessibility and RNA evidence beneath summary scores so the analytical label never replaces the underlying measurements.
Lead–lag claims require a developmental model and a timing audit. State the trajectory root, neighbours, branch assignments, smoothing, transition-point estimator and uncertainty. Compare RNA-only, ATAC-only and joint orderings, then anchor them to observed stages where available. Repeat the estimate within each independent specimen and after balancing cell numbers and depth. Chromatin opening before RNA is biologically intriguing, yet it remains temporal precedence rather than causation. The strongest next test perturbs the candidate element or regulator and asks whether the predicted expression transition changes.
Figures should progress from materials to inference. Begin with specimen balance, cell integrity, barcode occupancy, collisions and the two-channel quality landscape. Continue with modality-specific clusters, their joint alignment and raw marker evidence. Only then show DORCs, candidate links and timing. Label raw versus smoothed values, measured stages versus pseudotime and direct observations versus inferred connections. Every integrated locus should include cell numbers, specimen numbers and coverage. This ordering prevents an elegant final model from obscuring the evidence that made it possible.
Stewardship includes consent and provenance for human material, animal approvals where relevant, sample identifiers, plate maps, oligonucleotide sequences, reagent lots, raw reads, index assignments, fragment files, RNA matrices, cell exclusions, code and software environments. Single-cell genomic data can contain sensitive sequence information even when the study focuses on regulation. Apply proportionate access controls. Laboratory safety covers biological material, fixatives, enzymes, heat, sharps and amplified DNA, with clean separation of pre- and post-amplification areas and institution-approved procedures.
For learners, SHARE-seq is a vivid example of how science turns a story into a sequence of claims. First, establish that two measurements truly came from one cell. Next, show that both measurements are usable. Then compare states, nominate regulatory links and test timing. Finally, design a perturbation that could contradict the model. This ladder supports Primary Science observation, PSLE Science data use, Secondary Science variables and O-Level Science evaluation. It also gives families a practical way to judge science tuition: does teaching strengthen the reasoning ladder, or merely add vocabulary?
A useful final supplement is a sensitivity matrix. Vary barcode mismatch allowance, doublet thresholds, ATAC and RNA quality cut-offs, peak sets, DORC thresholds, trajectory roots, neighbour counts, smoothing and specimen inclusion. Mark which findings persist: cell identities, candidate links, DORCs and accessibility-before-expression patterns. Robust conclusions should survive several defensible pipelines and recur across specimens. Findings that depend on one setting are still valuable as hypotheses, because they point directly to the measurement or validation that should improve next.
Contents · Previous section · Continue to the Science Learning Hub
