eduKateSG · Why Science?
Translate carefully described phenotypes into a transparent candidate-gene shortlist—then reconnect every gene to variants, inheritance, evidence and uncertainty
Choose the closest route, then return to the full index whenever you need the wider evidence chain.
Science learning becomes powerful when students can tell a measured pattern from a model, a statistical decision from a mechanism, and a useful lead from a finished conclusion. Phen2Gene is an open-source phenotype-driven gene-prioritisation tool from Wang Genomics Lab. Its current website and repository accept Human Phenotype Ontology identifiers; the web service also describes input from clinical notes and Phenopackets. Phen2Gene uses a precomputed HPO-to-gene knowledgebase and term weighting to return a rapid ranked list of candidate genes. That list narrows attention; it does not examine a patient variant by itself, establish inheritance, confirm causality or replace diagnostic review. This guide supports Primary Science, PSLE Science, Secondary Science, O-Level Science and STEM education as a longform evidence-reading route; it is not medical advice, a laboratory protocol, an admission promise or a career guarantee.
Related eduKate reading: Human Phenotype Ontology and structured clinical phenotypes; Gene2Phenotype and structured gene–disease models; Matchmaker Exchange and federated rare-disease case matching; Science Learning Hub; Education Hub; STEM Education; Careers by Subject Capability.
Primary and official sources checked on 12 October 2026: Official Phen2Gene documentation observed 12 October 2026; Official Phen2Gene repository; Phen2Gene primary study; 2026 Singapore–Cambridge O-Level Biology syllabus; MOE G2/G3 Lower Secondary Science teaching and learning syllabus; 2026 MOE G2 Computing syllabus. Resource interfaces and recommendations can change, so preserve the release, access date and documentation used.
| Stage | Question | Evidence | Responsible output |
|---|---|---|---|
| Describe | Which verified patient phenotypes can be represented as HPO terms? | Clinical observations, mappings and term status | A traceable phenotype profile |
| Rank | Which genes are most associated with the combined profile? | Term weights and HPO2Gene knowledgebase | A ranked candidate-gene list |
| Integrate | Which variants and inheritance models fit those genes? | VCF, pedigree, gene–disease and variant evidence | Case-specific candidates |
| Validate | What independent evidence supports the leading explanation? | Segregation, function, literature and expert review | A bounded interpretation, not a gene-list diagnosis |
Inside this guide
1–12 · Foundations and design
- 1. Start with phenotype-to-gene translation
- 2. Understand the input
- 3. Map notes to HPO carefully
- 4. Prefer observed specificity
- 5. Keep absent and unknown distinct
- 6. Understand term weighting
- 7. Understand HPO2Gene
- 8. Understand precomputation
- 9. Read a rank as prioritisation
- 10. Read the full list
- 11. Inspect returned gene–disease context
- 12. Use the API reproducibly
13–24 · Models and estimation
- 13. Handle errors explicitly
- 14. Use Phenopackets as structured input
- 15. Separate gene ranking from variant ranking
- 16. Add inheritance after ranking
- 17. Add variant consequence
- 18. Add population evidence
- 19. Add gene–disease validity
- 20. Test term removal
- 21. Test redundant terms
- 22. Test phenotype expansion
- 23. Challenge annotation bias
- 24. Challenge benchmark leakage
25–36 · Diagnostics and interpretation
- 25. Compare tools fairly
- 26. Protect clinical notes
- 27. Use bounded clinical language
- 28. Connect to Primary Science
- 29. Connect to Secondary Science
- 30. Connect to mathematics
- 31. Connect to computing
- 32. Connect to pathways
- 33. Did you know?
- 34. Write a reader-ready result
- 35. Build a validation ladder
- 36. Finish with a reproducible bundle
Section 1 of 36
1. Start with phenotype-to-gene translation
Phen2Gene converts a structured phenotype profile into a ranked gene list so researchers can focus the next evidence search.
Working checkpoint — HPO term: Define the intended downstream use before submitting terms. Keep the source observations, HPO mappings, submitted request, raw response, knowledge date and software version together.
For a classroom model, use invented traits and gene cards rather than real clinical notes.
Section 2 of 36
2. Understand the input
The documented API accepts HPO identifiers, and the web interface describes clinical-note input and Phenopacket support.
Working checkpoint — candidate gene: Preserve the exact submitted terms or source text. Read each gene rank beside the phenotype terms and knowledge links that support it.
For future study, connect ontology work to biology, language, statistics and computing while checking official pathways separately.
Section 3 of 36
3. Map notes to HPO carefully
Free text can contain negation, uncertainty, family history and timing that a simple term extraction may miss.
Working checkpoint — term weighting: Review every generated term against the note. Preserve rejected terms and the complete list; a short favourite set is not the whole query result.
For scientific writing, call the output a candidate-gene ranking until downstream evidence supports more.
Section 4 of 36
4. Prefer observed specificity
A specific HPO term can discriminate more strongly than a broad ancestor, but only when the observation supports it.
Working checkpoint — HPO2Gene knowledgebase: Use the narrowest accurate term. Separate phenotype-to-gene ranking from variant interpretation, inheritance and diagnosis.
For students, ask which phenotype term contributed information and what evidence is still missing about the gene.
Section 5 of 36
5. Keep absent and unknown distinct
A missing term is not evidence that a feature is absent, and the documented ranking input is not a complete clinical assessment.
Working checkpoint — REST API: Store negated and unassessed findings separately. Treat HPO identifiers, mappings and knowledge versions as scientific data: a silent change can reorder the shortlist.
For parents, ask whether the tool ranked genes only or also evaluated a specific variant and family.
Section 6 of 36
6. Understand term weighting
Phen2Gene weights phenotype terms so more informative observations can contribute differently to the ranking.
Working checkpoint — ranked shortlist: Inspect how unusual and broad terms affect the result. Keep the source observations, HPO mappings, submitted request, raw response, knowledge date and software version together.
For a classroom model, use invented traits and gene cards rather than real clinical notes.
Section 7 of 36
7. Understand HPO2Gene
The precomputed knowledgebase links phenotype terms to genes through disease and gene knowledge used by the algorithm.
Working checkpoint — HPO term: Record the knowledgebase or software release. Read each gene rank beside the phenotype terms and knowledge links that support it.
For future study, connect ontology work to biology, language, statistics and computing while checking official pathways separately.
Section 8 of 36
8. Understand precomputation
Precomputed relationships make the response fast, while updates to source knowledge may change later rankings.
Working checkpoint — candidate gene: Keep retrieval date and version with every result. Preserve rejected terms and the complete list; a short favourite set is not the whole query result.
For scientific writing, call the output a candidate-gene ranking until downstream evidence supports more.
Section 9 of 36
9. Read a rank as prioritisation
Rank one means the gene scored highest among candidates under the submitted profile and current knowledgebase.
Working checkpoint — term weighting: Do not call the gene causal. Separate phenotype-to-gene ranking from variant interpretation, inheritance and diagnosis.
For students, ask which phenotype term contributed information and what evidence is still missing about the gene.
Section 10 of 36
10. Read the full list
Several genes can score similarly, and a low-ranked true gene may reflect incomplete phenotyping or knowledge.
Working checkpoint — HPO2Gene knowledgebase: Archive the complete ordered output. Treat HPO identifiers, mappings and knowledge versions as scientific data: a silent change can reorder the shortlist.
For parents, ask whether the tool ranked genes only or also evaluated a specific variant and family.
Section 11 of 36
11. Inspect returned gene–disease context
A gene can be linked to multiple diseases with different phenotypes, mechanisms and inheritance patterns.
Working checkpoint — REST API: Open the exact relationship relevant to the case. Keep the source observations, HPO mappings, submitted request, raw response, knowledge date and software version together.
For a classroom model, use invented traits and gene cards rather than real clinical notes.
Section 12 of 36
12. Use the API reproducibly
The REST service returns JSON results and errors; a reproducible call records the endpoint, parameters, response and date.
Working checkpoint — ranked shortlist: Save raw JSON before reformatting. Read each gene rank beside the phenotype terms and knowledge links that support it.
For future study, connect ontology work to biology, language, statistics and computing while checking official pathways separately.
Section 13 of 36
13. Handle errors explicitly
Unrecognised HPO identifiers, malformed input or service failures can silently shrink a profile if code ignores the errors list.
Working checkpoint — HPO term: Fail the workflow when decisive terms are rejected. Preserve rejected terms and the complete list; a short favourite set is not the whole query result.
For scientific writing, call the output a candidate-gene ranking until downstream evidence supports more.
Section 14 of 36
14. Use Phenopackets as structured input
Phenopackets can package case phenotypes and context, but schema and ontology versions remain part of the evidence.
Working checkpoint — candidate gene: Validate the packet before submission. Separate phenotype-to-gene ranking from variant interpretation, inheritance and diagnosis.
For students, ask which phenotype term contributed information and what evidence is still missing about the gene.
Section 15 of 36
15. Separate gene ranking from variant ranking
Phen2Gene prioritises genes from phenotype evidence; it does not by itself evaluate every variant in a VCF.
Working checkpoint — term weighting: Add variant annotation and filtering as a distinct step. Treat HPO identifiers, mappings and knowledge versions as scientific data: a silent change can reorder the shortlist.
For parents, ask whether the tool ranked genes only or also evaluated a specific variant and family.
Section 16 of 36
16. Add inheritance after ranking
A phenotype-associated gene may not fit the observed zygosity, family structure or allelic requirement.
Working checkpoint — HPO2Gene knowledgebase: Check the gene–disease inheritance model. Keep the source observations, HPO mappings, submitted request, raw response, knowledge date and software version together.
For a classroom model, use invented traits and gene cards rather than real clinical notes.
Section 17 of 36
17. Add variant consequence
A candidate gene can contain benign, uncertain and pathogenic variants with different transcript and molecular effects.
Working checkpoint — REST API: Interpret the exact allele independently. Read each gene rank beside the phenotype terms and knowledge links that support it.
For future study, connect ontology work to biology, language, statistics and computing while checking official pathways separately.
Section 18 of 36
18. Add population evidence
A variant frequency incompatible with a proposed rare high-penetrance model can weaken the candidate even when the gene ranks well.
Working checkpoint — ranked shortlist: Use ancestry-aware frequency data. Preserve rejected terms and the complete list; a short favourite set is not the whole query result.
For scientific writing, call the output a candidate-gene ranking until downstream evidence supports more.
Section 19 of 36
19. Add gene–disease validity
A phenotype association may be well established, disputed or emerging.
Working checkpoint — HPO term: Consult current expert-curated relationship evidence. Separate phenotype-to-gene ranking from variant interpretation, inheritance and diagnosis.
For students, ask which phenotype term contributed information and what evidence is still missing about the gene.
Section 20 of 36
20. Test term removal
A single highly weighted feature can dominate a ranking and may itself be uncertain.
Working checkpoint — candidate gene: Remove one term at a time in a planned influence check. Treat HPO identifiers, mappings and knowledge versions as scientific data: a silent change can reorder the shortlist.
For parents, ask whether the tool ranked genes only or also evaluated a specific variant and family.
Section 21 of 36
21. Test redundant terms
Closely related HPO terms can describe one feature and may make the profile appear richer than it is.
Working checkpoint — term weighting: Compare a deduplicated term set. Keep the source observations, HPO mappings, submitted request, raw response, knowledge date and software version together.
For a classroom model, use invented traits and gene cards rather than real clinical notes.
Section 22 of 36
22. Test phenotype expansion
Age-dependent or newly observed features can legitimately change the candidate list during reanalysis.
Working checkpoint — HPO2Gene knowledgebase: Version the phenotype profile rather than overwriting it. Read each gene rank beside the phenotype terms and knowledge links that support it.
For future study, connect ontology work to biology, language, statistics and computing while checking official pathways separately.
Section 23 of 36
23. Challenge annotation bias
Well-studied genes and diseases usually have more associations than recently discovered ones.
Working checkpoint — REST API: Discuss knowledge-depth bias in every shortlist. Preserve rejected terms and the complete list; a short favourite set is not the whole query result.
For scientific writing, call the output a candidate-gene ranking until downstream evidence supports more.
Section 24 of 36
24. Challenge benchmark leakage
A published case can be easy to rank when its phenotype–gene links already entered the knowledgebase.
Working checkpoint — ranked shortlist: Use time-aware evaluation when claiming performance. Separate phenotype-to-gene ranking from variant interpretation, inheritance and diagnosis.
For students, ask which phenotype term contributed information and what evidence is still missing about the gene.
Section 25 of 36
25. Compare tools fairly
Different tools rank genes, diseases or variants and may use overlapping data, so output length alone is not a fair comparison.
Working checkpoint — HPO term: Align inputs, releases and target task. Treat HPO identifiers, mappings and knowledge versions as scientific data: a silent change can reorder the shortlist.
For parents, ask whether the tool ranked genes only or also evaluated a specific variant and family.
Section 26 of 36
26. Protect clinical notes
Submitting free text can expose sensitive details that HPO-only input would omit.
Working checkpoint — candidate gene: Use approved systems and minimise text. Keep the source observations, HPO mappings, submitted request, raw response, knowledge date and software version together.
For a classroom model, use invented traits and gene cards rather than real clinical notes.
Section 27 of 36
27. Use bounded clinical language
Phen2Gene helps prioritise candidate genes; it does not diagnose, classify variants or recommend care.
Working checkpoint — term weighting: Write “ranked for review”. Read each gene rank beside the phenotype terms and knowledge links that support it.
For future study, connect ontology work to biology, language, statistics and computing while checking official pathways separately.
Section 28 of 36
28. Connect to Primary Science
A fictional feature-to-organism matching game shows how several clues can narrow a list without proving identity.
Working checkpoint — HPO2Gene knowledgebase: Use invented features and transparent scoring. Preserve rejected terms and the complete list; a short favourite set is not the whole query result.
For scientific writing, call the output a candidate-gene ranking until downstream evidence supports more.
Section 29 of 36
29. Connect to Secondary Science
Genes, phenotypes and inheritance show how biological observations guide a search through many candidates.
Working checkpoint — REST API: Add the missing variant step explicitly. Separate phenotype-to-gene ranking from variant interpretation, inheritance and diagnosis.
For students, ask which phenotype term contributed information and what evidence is still missing about the gene.
Section 30 of 36
30. Connect to mathematics
Weights, ranks, ties and sensitivity checks demonstrate why an ordered list is not a probability scale.
Working checkpoint — ranked shortlist: Re-rank a toy list after removing one term. Treat HPO identifiers, mappings and knowledge versions as scientific data: a silent change can reorder the shortlist.
For parents, ask whether the tool ranked genes only or also evaluated a specific variant and family.
Section 31 of 36
31. Connect to computing
Ontologies, APIs, JSON, identifiers and precomputed databases turn clinical concepts into reproducible queries.
Working checkpoint — HPO term: Validate every identifier before execution. Keep the source observations, HPO mappings, submitted request, raw response, knowledge date and software version together.
For a classroom model, use invented traits and gene cards rather than real clinical notes.
Section 32 of 36
32. Connect to pathways
Phenotype informatics connects biology, medicine, language technology, statistics and software and can illuminate STEM pathways.
Working checkpoint — candidate gene: Check official education and professional requirements separately. Read each gene rank beside the phenotype terms and knowledge links that support it.
For future study, connect ontology work to biology, language, statistics and computing while checking official pathways separately.
Section 33 of 36
33. Did you know?
Phen2Gene can return a candidate-gene ranking quickly because much of the HPO-to-gene knowledge is precomputed before the case query.
Working checkpoint — term weighting: Speed does not remove the need for validation. Preserve rejected terms and the complete list; a short favourite set is not the whole query result.
For scientific writing, call the output a candidate-gene ranking until downstream evidence supports more.
Section 34 of 36
34. Write a reader-ready result
Name the profile, mapping method, software and data date, leading genes, score context, rejected terms, sensitivity checks and next evidence step.
Working checkpoint — HPO2Gene knowledgebase: Keep ranking and causality separate. Separate phenotype-to-gene ranking from variant interpretation, inheritance and diagnosis.
For students, ask which phenotype term contributed information and what evidence is still missing about the gene.
Section 35 of 36
35. Build a validation ladder
Verify HPO mappings, test term influence, inspect gene–disease and inheritance evidence, evaluate variants, then seek segregation and functional confirmation.
Working checkpoint — REST API: Do not skip from phenotype to diagnosis. Treat HPO identifiers, mappings and knowledge versions as scientific data: a silent change can reorder the shortlist.
For parents, ask whether the tool ranked genes only or also evaluated a specific variant and family.
Section 36 of 36
36. Finish with a reproducible bundle
Deliver permitted source observations, HPO mappings, request, raw response, versions, complete ranking, sensitivity results and downstream review status.
Working checkpoint — ranked shortlist: Make the shortlist reconstructable. Keep the source observations, HPO mappings, submitted request, raw response, knowledge date and software version together.
For a classroom model, use invented traits and gene cards rather than real clinical notes.
Contents · Previous section · Continue to the Science Learning Hub
