PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 3, 2026Nature1 citationsOpen Access

The 1000 Chinese Pangenome empowers medical and population genetics

YWYifei WangZDZhongqu DuanDCD. Lu Chen

Key Points

  • To construct a comprehensive pangenome that characterizes complex genetic variations in a large sample of Chinese genomes.
  • Generated 1,116 diploid genome assemblies including 55 de novo and 1,061 pangenome-informed.
  • Catalogued various genetic variations such as small variants and structural variants.
  • Conducted pan-variant expression quantitative trait locus mapping to identify eQTLs.
  • Constructed a pangenome containing 405.3 million base pairs of unique sequences.
  • Identified 3,256 eQTLs related to complex genetic variants.
  • Developed a pan-variant imputation reference panel for future genetic studies.

Abstract

Pangenomes are revolutionizing our ability to resolve genomic regions with complex variations1. However, existing human pangenomes2,3, constrained by small sample sizes, provide limited utility for medical and population genetic applications. Here we generated 1,116 diploid genome assemblies (55 de novo and 1,061 pangenome-informed) with an average size of 2.98 Gb and a mean quality value of 46 as part of the 1000 Chinese Pangenome (1KCP) project. On the basis of these assemblies, we constructed a pangenome comprising 405.3 million base pairs of sequences absent from the current references GRCh38 and CHM13, including 26.2 million base pairs of functional genic and predicted regulatory elements. We catalogued a full spectrum of genetic variation, including 35.4 million small variants, 110,530 structural variants (SVs), 485,575 tandem repeats (TRs) and 0.86 million nested variants embedded in non-reference sequences. This extensive dataset enabled detailed characterization of multiscale genic variations relevant to medical genetics, including gene-altering SVs, TR expansions, gene cluster variations and HLA gene haplotypes. Coupled with the 1KCP gene expression data, we conducted pan-variant expression quantitative trait locus (eQTL) mapping to analyse diverse variant types. We identified 3,256 eQTLs involving complex variants (SVs, TRs and nested variants) and elucidated their regulatory complexity. Finally, we developed a 1KCP pan-variant imputation reference panel, which provides multitype genetic markers to enhance the resolution of future association studies. This resource advances our understanding of complex variants and their functional implications to provide new insights into human health. Development of the pangenome-informed genome assembly (PIGA) workflow enabled the generation of 1,116 diploid genome assemblies (55 de novo and 1,061 pangenome-informed), representing an extensive resource of medically relevant genic variations.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Wang et al. (2026) studied this question.

synapsesocial.com/papers/69cf5e3d5a333a821460c6fdhttps://doi.org/10.1038/s41586-026-10315-y
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Using probabilistic estimation of expression residuals (PEER) to obtain increased power and interpretability of gene expression analyses2012 · 1,283 citations
  2. 2Mash: fast genome and metagenome distance estimation using MinHash2016 · 3,680 citations
  3. 3CPC2: a fast and accurate coding potential calculator based on sequence intrinsic features2017 · 1,837 citations
  4. 4High-coverage whole-genome sequencing of the expanded 1000 Genomes Project cohort including 602 trios2022 · 1,088 citations
  5. 5Structural variation in 1,019 diverse humans based on long-read sequencing2025 · 65 citations