VIEW THIS AS

Auto mode follows the Route Engine until you choose a viewpoint.

YOU ARE HERE

ROUTE CHECK

CONNECTED TO

WHAT NEXT

Use the canonical route for this room, or HELP if you are unsure.

Why Science? | GA4GH Data Connect and Federated Dataset Queries

Three students sit around open books and worksheets at a classroom table, reading, writing and discussing the work together.

eduKateSG · Why Science?

Ask one carefully bounded question across tabular data services—and keep schema, meaning, permissions and provenance visible

Choose the closest route, then return to the full index whenever you need the wider evidence chain.

Full section index · Science Learning Hub

Science learning becomes powerful when students can tell a measured pattern from a model, a statistical decision from a mechanism, and a useful lead from a finished conclusion. GA4GH Data Connect is a standard API for discovering and querying tabular datasets, including data held across federated services. GA4GH lists Data Connect as current and Support Ready, with v1.0 approved on 22 June 2021. Its primary container is a table whose rows are JSON objects described with JSON Schema. Core routes let clients list tables, inspect one table and page through its rows; an optional search route accepts SQL with positional parameters. The standard improves machine-readable discovery and query portability, but it does not harmonise scientific meaning, grant access, prove datasets comparable or turn a syntactically valid query into a valid study design. This guide supports Primary Science, PSLE Science, Secondary Science, O-Level Science and STEM education as a longform evidence-reading route; it is not medical advice, a laboratory protocol, an admission promise or a career guarantee.

Related eduKate reading: GA4GH Data Repository Service and portable data access; GA4GH Beacon v2 and federated genomic data discovery; GA4GH Phenopackets and computable case representation; Science Learning Hub; Education Hub; STEM Education; Careers by Subject Capability.

Primary and official sources checked on 12 October 2026: Official GA4GH Data Connect product page, current v1.0; Official GA4GH Data Connect specification; Official GA4GH Data Connect repository; 2026 Singapore–Cambridge O-Level Biology syllabus; MOE G2/G3 Lower Secondary Science teaching and learning syllabus; 2026 MOE G2 Computing syllabus. Standards, service interfaces and recommendations can change, so preserve the dated documentation and versions used.

A deeper reading route

Data Connect is easiest to understand as a shared grammar for tabular questions. A service can explain which tables it offers, describe the structure of a row and return a bounded page of data. That common shape matters when a researcher wants to write one client for several institutions. Yet a column named age, diagnosis or score still needs a definition, unit, coding rule and population context. Interoperable syntax is the beginning of a scientific conversation, not its conclusion.

The table model is deliberately broad. Any data that can be represented as arrays of JSON objects can fit, which allows implementations to sit over relational databases, data lakes or other stores. JSON Schema tells a client whether a property is a string, number, array or nested object and can link to semantic definitions. It cannot by itself guarantee that two sites collected a value in the same way or that missing values have the same cause.

Federation changes the location of computation more than it changes the logic of evidence. A coordinator can send related requests to several services and combine permitted responses, while each institution keeps control of its own system. The combined answer still inherits local eligibility rules, refresh schedules, coding practices and disclosure protections. A sound report therefore preserves the endpoint, table identity, schema version, query, parameters and response provenance for every contributing service.

The optional SQL search operation is powerful because filters, projections, joins and aggregations can express rich questions. It is also a boundary that demands care. Positional parameters should carry user-supplied values; services need validation, timeouts, row limits and authorisation; researchers need to inspect denominators and avoid queries that expose tiny groups. A result that is fast and technically correct can still be ethically unsafe or scientifically misleading.

Pagination and long-running queries make completeness a testable property. A single response may be only the first page, and a federated request may finish at different times across sites. Clients should follow paging information, detect repeated or missing pages, record partial failures and distinguish “zero matching rows” from “service unavailable”. Those habits turn a convenient dashboard number into an auditable result.

For students, Data Connect offers a cheerful bridge from classroom tables to modern science infrastructure. The familiar questions remain: What is one row? What does each column mean? Which records were included? What is the denominator? Computing adds schemas, APIs and federation, while science keeps the deeper responsibility of deciding whether the comparison is fair.

A four-step claim ladder for this method
StageQuestionEvidenceResponsible output
DiscoverWhich tables and structures are available?Table lists, identifiers and JSON SchemaA candidate data source
QueryWhich rows or summaries answer the question?Paged data routes or parameterised SQLA bounded response
FederateWhich authorised services contribute?Endpoint, schema and response provenanceA traceable multi-site result
InterpretAre values genuinely comparable?Definitions, units, cohorts and missingnessA qualified scientific claim

Section 1 of 36

1. Start with the scientific question

A federated query should begin with a population, variable, comparison and intended interpretation. A broad technical search can return many rows while answering no coherent scientific question.

Working checkpoint — Data Connect: Write the claim and denominator before selecting a table. Keep the endpoint, table and schema versions, query, parameters, pages and response provenance together.

For a classroom model, query fictional plant measurements rather than personal or health records.

Contents · Next section

Section 2 of 36

2. Understand the current standard

GA4GH lists Data Connect as a current, Support Ready product and identifies v1.0, approved on 22 June 2021, as the last approved version.

Working checkpoint — table: Record the standard, implementation and access date. Read every summary beside site-specific definitions, units, inclusion rules, missingness and denominators.

For future study, connect tables to biology, statistics and data engineering while checking official pathways separately.

Contents · Previous section · Next section

Section 3 of 36

3. Treat a table as the main container

The specification models a table as the primary container for data serialisable as rows of JSON objects. A table may sit over many kinds of storage.

Working checkpoint — JSON Schema: Name the logical table and its service endpoint together. Preserve partial failures and excluded services; a combined answer is not evidence that every node responded.

For scientific writing, cite Data Connect, endpoints, schemas, query parameters and denominators.

Contents · Previous section · Next section

Section 4 of 36

4. Read rows as JSON objects

Each row is a JSON object whose properties carry the values exposed by that table. Nested values and arrays need explicit parsing rather than spreadsheet assumptions.

Working checkpoint — search query: Validate a small page against the declared schema. Separate syntactic interoperability and query success from harmonisation, privacy approval and scientific comparability.

For students, ask what one row represents, what each column means and which records were filtered.

Contents · Previous section · Next section

Section 5 of 36

5. Use JSON Schema for structure

JSON Schema can describe property names, types, required fields and nested shapes. It makes machine validation possible without claiming that every value is scientifically equivalent.

Working checkpoint — federation: Archive the exact schema returned for the analysis. Treat table, field, namespace and release identifiers as scientific inputs: a silent mismatch can alter the cohort.

For parents, ask whether federated data stay governed and whether comparisons share definitions.

Contents · Previous section · Next section

Section 6 of 36

6. Separate structure from semantics

Two columns can share a numeric type while measuring different constructs, units or time points. Structural compatibility does not establish scientific comparability.

Working checkpoint — data provenance: Attach definitions, units and collection context to every analysed field. Keep the endpoint, table and schema versions, query, parameters, pages and response provenance together.

For a classroom model, query fictional plant measurements rather than personal or health records.

Contents · Previous section · Next section

Section 7 of 36

7. Discover the table catalogue

The table-list operation lets a client learn which logical tables a service exposes instead of relying on undocumented database knowledge.

Working checkpoint — Data Connect: Save the complete catalogue response and paging state. Read every summary beside site-specific definitions, units, inclusion rules, missingness and denominators.

For future study, connect tables to biology, statistics and data engineering while checking official pathways separately.

Contents · Previous section · Next section

Section 8 of 36

8. Inspect one table before querying

The table-information operation supplies the table description and schema needed to build a valid client request.

Working checkpoint — table: Compare expected fields with the live table description. Preserve partial failures and excluded services; a combined answer is not evidence that every node responded.

For scientific writing, cite Data Connect, endpoints, schemas, query parameters and denominators.

Contents · Previous section · Next section

Section 9 of 36

9. Browse rows through the data route

The table-data operation can return rows without requiring the optional SQL search interface. It is useful for small, bounded inspection and simple clients.

Working checkpoint — JSON Schema: Request a tiny page before attempting full retrieval. Separate syntactic interoperability and query success from harmonisation, privacy approval and scientific comparability.

For students, ask what one row represents, what each column means and which records were filtered.

Contents · Previous section · Next section

Section 10 of 36

10. Treat search as optional

Data Connect defines an optional search operation; a conforming service may expose table discovery and row access without supporting SQL search.

Working checkpoint — search query: Discover capability instead of assuming every endpoint accepts SQL. Treat table, field, namespace and release identifiers as scientific inputs: a silent mismatch can alter the cohort.

For parents, ask whether federated data stay governed and whether comparisons share definitions.

Contents · Previous section · Next section

Section 11 of 36

11. Use positional query parameters

The search design supports positional parameters so values can be separated from the SQL text. This reduces ambiguity and helps implementations enforce safer query handling.

Working checkpoint — federation: Bind values as parameters and never construct SQL by string concatenation. Keep the endpoint, table and schema versions, query, parameters, pages and response provenance together.

For a classroom model, query fictional plant measurements rather than personal or health records.

Contents · Previous section · Next section

Section 12 of 36

12. Keep the SQL dialect visible

SQL capabilities can vary across engines and implementations. A query that succeeds at one service may use a function or type conversion unavailable at another.

Working checkpoint — data provenance: Document the supported subset and test it at every target. Read every summary beside site-specific definitions, units, inclusion rules, missingness and denominators.

For future study, connect tables to biology, statistics and data engineering while checking official pathways separately.

Contents · Previous section · Next section

Section 13 of 36

13. Use schema references carefully

A JSON Schema can refer to shared or semantic definitions, helping related tables express more than primitive types. References still require stable resolution and versioning.

Working checkpoint — Data Connect: Preserve referenced identifiers and the content they resolved to. Preserve partial failures and excluded services; a combined answer is not evidence that every node responded.

For scientific writing, cite Data Connect, endpoints, schemas, query parameters and denominators.

Contents · Previous section · Next section

Section 14 of 36

14. Name identifiers and namespaces

An identifier is meaningful only with its namespace and issuing context. The same character string can refer to different entities in different catalogues.

Working checkpoint — table: Store namespace, version and service beside every identifier. Separate syntactic interoperability and query success from harmonisation, privacy approval and scientific comparability.

For students, ask what one row represents, what each column means and which records were filtered.

Contents · Previous section · Next section

Section 15 of 36

15. Connect logical data to DRS objects

Data Connect can expose tabular discovery while DRS identifies retrievable data objects. A row reference and a file object are related but not interchangeable.

Working checkpoint — JSON Schema: Record both logical row context and object identity where used. Treat table, field, namespace and release identifiers as scientific inputs: a silent mismatch can alter the cohort.

For parents, ask whether federated data stay governed and whether comparisons share definitions.

Contents · Previous section · Next section

Section 16 of 36

16. Traverse pagination completely

Large responses can be split across pages. Reading only the first page changes the denominator and can bias summaries without producing an obvious error.

Working checkpoint — search query: Follow paging links or tokens and reconcile row counts. Keep the endpoint, table and schema versions, query, parameters, pages and response provenance together.

For a classroom model, query fictional plant measurements rather than personal or health records.

Contents · Previous section · Next section

Section 17 of 36

17. Plan for long-running queries

A federated or computationally heavy search may not finish within one short request. Services and clients need bounded, observable handling of delayed results.

Working checkpoint — federation: Record submission, polling and completion states without duplicating requests. Read every summary beside site-specific definitions, units, inclusion rules, missingness and denominators.

For future study, connect tables to biology, statistics and data engineering while checking official pathways separately.

Contents · Previous section · Next section

Section 18 of 36

18. Draw the federation tree

A coordinating service may query other Data Connect services, which can themselves expose local tables. The final answer should retain the path of contribution.

Working checkpoint — data provenance: Map every endpoint and transformation involved in aggregation. Preserve partial failures and excluded services; a combined answer is not evidence that every node responded.

For scientific writing, cite Data Connect, endpoints, schemas, query parameters and denominators.

Contents · Previous section · Next section

Section 19 of 36

19. Respect heterogeneous stores

A common API can front relational systems, object stores or analytical engines. Backend differences affect performance, refresh timing and supported operations.

Working checkpoint — Data Connect: Test behaviour rather than inferring it from the shared interface. Separate syntactic interoperability and query success from harmonisation, privacy approval and scientific comparability.

For students, ask what one row represents, what each column means and which records were filtered.

Contents · Previous section · Next section

Section 20 of 36

20. Version schemas and data releases

Adding, renaming or retyping a property can break clients or change interpretation. Data updates can also alter counts while the endpoint remains the same.

Working checkpoint — table: Pin schema and dataset release identifiers in the report. Treat table, field, namespace and release identifiers as scientific inputs: a silent mismatch can alter the cohort.

For parents, ask whether federated data stay governed and whether comparisons share definitions.

Contents · Previous section · Next section

Section 21 of 36

21. Preserve units and denominators

A value needs a unit, reference population and aggregation rule. Counts from two services may use different inclusion windows or privacy thresholds.

Working checkpoint — JSON Schema: Define every numerator and denominator before combining results. Keep the endpoint, table and schema versions, query, parameters, pages and response provenance together.

For a classroom model, query fictional plant measurements rather than personal or health records.

Contents · Previous section · Next section

Section 22 of 36

22. Keep missingness explicit

Null, absent, suppressed and not-collected values are different scientific states. Flattening them into zero creates false observations.

Working checkpoint — search query: Maintain a missingness codebook and report site-specific rates. Read every summary beside site-specific definitions, units, inclusion rules, missingness and denominators.

For future study, connect tables to biology, statistics and data engineering while checking official pathways separately.

Contents · Previous section · Next section

Section 23 of 36

23. Audit joins before believing them

Joining tables can duplicate rows, drop unmatched records or create many-to-many expansion. Valid SQL does not guarantee a valid scientific unit.

Working checkpoint — federation: Count keys and rows before and after each join. Preserve partial failures and excluded services; a combined answer is not evidence that every node responded.

For scientific writing, cite Data Connect, endpoints, schemas, query parameters and denominators.

Contents · Previous section · Next section

Section 24 of 36

24. Audit aggregations and small groups

Group counts and averages can hide unequal coverage and may expose sensitive small cells. Federation does not remove disclosure risk.

Working checkpoint — data provenance: Apply approved disclosure controls and display contributing denominators. Separate syntactic interoperability and query success from harmonisation, privacy approval and scientific comparability.

For students, ask what one row represents, what each column means and which records were filtered.

Contents · Previous section · Next section

Section 25 of 36

25. Separate authentication from meaning

Credentials decide whether a caller may reach a service; they do not certify that the chosen variables or cohort answer the question.

Working checkpoint — Data Connect: Review permissions and study design as separate gates. Treat table, field, namespace and release identifiers as scientific inputs: a silent mismatch can alter the cohort.

For parents, ask whether federated data stay governed and whether comparisons share definitions.

Contents · Previous section · Next section

Section 26 of 36

26. Defend the query boundary

Search interfaces need input validation, resource limits, timeouts and logging. Parameterisation reduces injection risk but does not prevent expensive or disclosive queries.

Working checkpoint — table: Threat-model syntax, compute cost and returned information. Keep the endpoint, table and schema versions, query, parameters, pages and response provenance together.

For a classroom model, query fictional plant measurements rather than personal or health records.

Contents · Previous section · Next section

Section 27 of 36

27. Preserve provenance for every response

A federated result should identify services, table and schema versions, query text, parameters, time, paging and transformations.

Working checkpoint — JSON Schema: Create a machine-readable manifest alongside the result. Read every summary beside site-specific definitions, units, inclusion rules, missingness and denominators.

For future study, connect tables to biology, statistics and data engineering while checking official pathways separately.

Contents · Previous section · Next section

Section 28 of 36

28. Connect to Primary Science

A class table of plant heights shows that rows are observations and columns are named measurements. Filtering changes which observations support a conclusion.

Working checkpoint — search query: Use invented, non-personal data and keep the original table. Preserve partial failures and excluded services; a combined answer is not evidence that every node responded.

For scientific writing, cite Data Connect, endpoints, schemas, query parameters and denominators.

Contents · Previous section · Next section

Section 29 of 36

29. Connect to Secondary Science

Fair comparisons depend on consistent measurement, sampling and variables. A common column name cannot repair different experimental designs.

Working checkpoint — federation: Compare two fictional schemas and identify hidden differences. Separate syntactic interoperability and query success from harmonisation, privacy approval and scientific comparability.

For students, ask what one row represents, what each column means and which records were filtered.

Contents · Previous section · Next section

Section 30 of 36

30. Connect to mathematics

Sets, filters, joins, proportions and denominators connect federated data work to logic, algebra and statistics.

Working checkpoint — data provenance: Calculate how a missing page changes a reported percentage. Treat table, field, namespace and release identifiers as scientific inputs: a silent mismatch can alter the cohort.

For parents, ask whether federated data stay governed and whether comparisons share definitions.

Contents · Previous section · Next section

Section 31 of 36

31. Connect to computing

REST routes, JSON objects, schemas, SQL, pagination and federation show how software interfaces become scientific infrastructure.

Working checkpoint — Data Connect: Test invalid types, missing fields and interrupted pages. Keep the endpoint, table and schema versions, query, parameters, pages and response provenance together.

For a classroom model, query fictional plant measurements rather than personal or health records.

Contents · Previous section · Next section

Section 32 of 36

32. Connect to pathways

Data federation joins biology, statistics, data engineering, security and research governance and can illuminate STEM pathways.

Working checkpoint — table: Check current official course and professional requirements separately. Read every summary beside site-specific definitions, units, inclusion rules, missingness and denominators.

For future study, connect tables to biology, statistics and data engineering while checking official pathways separately.

Contents · Previous section · Next section

Section 33 of 36

33. Did you know?

Data Connect can describe a table with JSON Schema and optionally search it with SQL, allowing a client to discover structure before composing a query.

Working checkpoint — JSON Schema: Treat discovery as a safety and meaning check, not a formality. Preserve partial failures and excluded services; a combined answer is not evidence that every node responded.

For scientific writing, cite Data Connect, endpoints, schemas, query parameters and denominators.

Contents · Previous section · Next section

Section 34 of 36

34. Write a reader-ready query report

Name Data Connect and implementation versions, endpoints, tables, schemas, query, parameters, pages, timestamps, permissions and transformations.

Working checkpoint — search query: State which services failed or contributed zero rows. Separate syntactic interoperability and query success from harmonisation, privacy approval and scientific comparability.

For students, ask what one row represents, what each column means and which records were filtered.

Contents · Previous section · Next section

Section 35 of 36

35. Build a validation ladder

Inspect the catalogue and schema, query a tiny fixture, verify types and paging, test missing services, review privacy, then run the full federation.

Working checkpoint — federation: Resolve the earliest failing layer before scaling. Treat table, field, namespace and release identifiers as scientific inputs: a silent mismatch can alter the cohort.

For parents, ask whether federated data stay governed and whether comparisons share definitions.

Contents · Previous section · Next section

Section 36 of 36

36. Finish with a reproducible bundle

Deliver schemas, semantic references, query and parameters, endpoint manifest, responses or hashes, page log, provenance, permissions and analysis outputs.

Working checkpoint — data provenance: Make every result traceable to a defined row and service. Keep the endpoint, table and schema versions, query, parameters, pages and response provenance together.

For a classroom model, query fictional plant measurements rather than personal or health records.

Contents · Previous section · Continue to the Science Learning Hub

Discover more from eduKate Singapore

Subscribe now to keep reading and get access to the full archive.

Continue reading