Indian data layer7 of 10 integrated
The moat
The edge is not the foundation models — they are commoditised. The edge is the combination of Indian population data, graph-constrained reasoning that cannot hallucinate entities, and clinical validation in Indian cohorts. This page declares the data layer state honestly, including what is pending.
Source: Build & Launch Reference §1 (strategic principle) + §7 (Indian Data Layer).
🇮🇳
INDIGEN COMMERCIAL LICENSING · IN ACTIVE CONVERSATION
CSIR-IGIB IndiGen (1,029 genomes · 55M variants)
IndiGen is the single most important Indian population dataset in PetriDish — it's where the "23% CYP2C19*2 in South Asians" comes from. Research-tier use is free. Commercial licensing with CSIR-IGIB Business Development is in active conversation and will be resolved before paid product launch. We disclose this here because customer due-diligence will surface it — better proactively than late.
Reference: Build & Launch Reference §7 + §10 landmine #1.
CSIR-IGIB
1,029 whole genomes · 55M variants · 32% India-unique
Population-specific allele frequencies; PGx panel; LD reference for India PRS calibration
Research free · Commercial license required
Indian Genome Variation Database
900K+ annotated Indian variants
Variant annotation with India-specific frequency; identifies variants absent from dbSNP/gnomAD
Open
International Genome Project
~500 samples · BEB/GIH/ITU/PJL/STU · 30× WGS
Population structure analysis; LD reference complement; ancestry inference
Open
GenomeAsia 100K
GenomeAsia Consortium
~2,600 South Asian WGS
Largest available South Asian WGS for rare variant frequency
Controlled access · application required
DATRI
Donor Awareness Trust of India
600,000+ HLA-typed bone marrow donors
Most comprehensive Indian HLA frequency dataset
Application required for research access
EMBL-EBI
~15,000 Indian HLA-typed individuals (DATRI/MDRI contributions)
HLA allele frequency reference
Open
ICMR-NCD Registry
Indian Council of Medical Research
5,000+ Indian autoimmune cases with HLA typing (target)
Clinical + partial genomic data for validation cohorts
Application required
Published Indian GWAS
Various (AIIMS, CMC, PGIMER, ICMR consortia)
RA · SLE · AS · T1D · IBD · psoriasis (~6 published India studies)
India-specific effect sizes; supplementary tables liftover to GRCh38
Literature mining (PMID-cited)
IMPPAT 2.0
IIT-Madras (curated subset)
17,967 Indian phytochemicals · 100 currently in graph
Ayurvedic compound knowledge layer — pathway-target mapping for AyurBridge / AutoCure
Academic / CC-BY (curated subset)
Stanford / NIH
61 PGx variants ingested · CPIC 2024 guideline corpus
Pharmacogenomic guideline + evidence grading for IndoPGx
Free access · commercial redistribution requires verification
Roadmap
- Resolve IndiGen commercial licensing with CSIR-IGIB BD (Build & Launch §9 Phase 2)
- Apply for DATRI HLA dataset for India HLA Atlas validation
- Apply for GenomeAsia 100K for rare-variant fine-mapping
- Apply for ICMR-NCD Registry access for autoimmune validation cohorts
- Close ACTREC partnership for IMPPAT 17,967-compound expansion (currently 100)