Preprint introduces summary-statistics method to detect ascertainment bias in biobanks
Researchers describe a new approach that estimates sampling bias in large genetic studies using only summary data, without requiring access to individual-level records.
A preprint posted to bioRxiv on 7 August 2026 presents a summary-statistics-based method for detecting and quantifying ascertainment bias in large-scale genetic studies such as biobanks. Non-random participation — where individuals who enrol differ systematically from the general population — can distort associations between genetic variants and health or trait outcomes, yet most existing correction methods require individual-level data that are rarely shared across institutions.
The authors introduce an estimator, denoted theta, that captures the deviation between the mean polygenic score of an ascertained sample and its expectation in a non-ascertained or differentially ascertained reference population. Simulation experiments and applications to real biobank data are described, though full details are not available from the lede text alone.
The method is of direct relevance to researchers who use polygenic scores derived from biobank GWAS, as uncorrected ascertainment can lead to overestimation or underestimation of variant effects and miscalibrated scores. The work also has implications for meta-analyses that aggregate summary statistics across multiple cohorts with differing recruitment strategies. As a preprint, the findings have not yet been peer-reviewed.
Sources
Read the original reporting — these are the public sources this summary draws from.
-
Primary sourcePreprint bioRxiv (Cold Spring Harbor Laboratory) · 2026-08-07Using summary data to detect and quantify ascertainment in biobanks