Lowers barriers to entry
Shared infrastructure reduces the burden of navigating separate access processes, agreements, and analysis environments for each cohort.
The Extended Twin, Adoption, and Family Data Commons is an early-stage global academic initiative working toward a secure, controlled-access resource that integrates genetically informative family cohorts, harmonized phenotypes, and genomic data to support more powerful, transparent, and reproducible research.
The ETAF Data Commons remains in the planning stage. Here is a brief summary of where things stand.
A working group has met monthly during 2025–2026 and convened in person in Boulder, Colorado on June 1–4, 2026. The group has developed the scientific rationale, initial governance thinking, and a funding strategy.
An NIH PF5 application was submitted in May 2026. If funded, the seed phase would support initial infrastructure development and a small number of pilot cohorts. Additional seed-funding opportunities are under consideration.
A manuscript describing the scientific rationale and proposed implementation has been drafted and is expected to be finalized in 2026. An ETAF Data Commons Protocol has also been drafted to define a future path for cohort participation.
Broader rollout, including a call for cohort participation, is expected only after seed funding is obtained and initial infrastructure has been tested. Calls to include new cohorts will begin once preliminary funding and infrastructure are established.
A community survey of the twin- and family-register world suggests that the phenotypic and genotypic resources for global expansion already exist.
Twin, adoption, and extended-family cohorts have delivered decades of insight into genetic and environmental influences on human behavior, health, and development. Yet these datasets remain scattered across institutions, governed independently, and difficult to combine at scale.
Shared infrastructure reduces the burden of navigating separate access processes, agreements, and analysis environments for each cohort.
Harmonization pursued where scientifically appropriate; cohort-specific measures preserved when harmonization would reduce scientific value.
Combining independently collected cohorts enables more robust, reproducible findings that no single cohort can support alone.
Within-family GWAS, indirect genetic effects models, and extended pedigree analyses require large samples that no single cohort can provide.
Learn more about the science, how to get involved, and how the initiative is being built.
Learn how twin, adoption, and family study leaders can contribute to a shared scientific resource — and what participation might offer in terms of scientific reach, harmonization support, and governance input.
Cohort participationDiscover the scientific questions the resource is designed to support — from within-family genomics and assortative mating to developmental change and cross-cohort replication. Future data access will be controlled and project-based.
Research opportunitiesThe scientific case for shared infrastructure, the data flow pipeline, and research use cases.
Learn moreHow the resource is designed around controlled access, consent, and responsible data use.
Our approachLeadership, working groups, and the people building the initiative's scientific and governance foundations.
Meet the teamQuestions about cohort participation, funding, data access, or collaboration? Get in touch.
Get in touchThe initiative is in early development and welcomes conversations with cohort leaders, researchers, funders, and collaborators.