A public, consent-based reference dataset for underrepresented populations.
Overview
Population genomics references are heavily skewed toward a handful of ancestries. This project builds a consent-based, public reference dataset for South Asian populations.
Objectives
- Recruit consenting participants across four regions
- Sequence and publish a de-identified reference panel
- Document consent design decisions publicly
Research questions
What reference-panel gaps most affect variant-calling accuracy for South Asian populations specifically?
Methodology
Whole-genome sequencing with a tiered, revocable consent model developed with regional ethics boards.
Current progress
Two of four planned regions recruited; sequencing underway for the first cohort.
Future work
Complete recruitment, sequence remaining cohorts, publish the reference panel.
Acknowledgements
With thanks to our regional ethics review partners.