Preprint proposes corrected liability-scale variance explained for polygenic scores derived from logistic regression
A bioRxiv preprint identifies a systematic discrepancy in how polygenic score accuracy is calculated when logistic rather than linear regression is used, and derives corrected expressions adjusted for case-control ascertainment.
A preprint deposited on bioRxiv addresses a technical problem in the quantification of polygenic score (PGS) prediction accuracy for binary disease outcomes. Standard practice converts Nagelkerke's R² from logistic regression on the observed 0–1 scale to a liability-scale R² adjusted for case-control ascertainment, enabling comparisons across studies with different disease prevalences and sampling designs. However, previous derivations were developed primarily with linear regression in mind. The preprint argues that applying these transformations to logistic regression outputs introduces systematic error, and derives corrected expressions appropriate for the logistic setting.
The work is technical in scope and directed at researchers who construct, evaluate, or compare polygenic scores for disease traits — a large and methodologically active community given the proliferation of PGS analyses in biobank-scale datasets. Incorrect liability-scale R² estimates have downstream consequences for published comparisons of model performance, heritability benchmarking, and meta-analyses of polygenic prediction.
The preprint does not present new biological findings or polygenic scores for specific diseases; its contribution is methodological. Researchers in statistical genetics, genetic epidemiology, and bioinformatics are the primary intended audience. The work has not yet undergone peer review, and independent verification of the derived expressions will be an important validation step.
Sources
Read the original reporting — these are the public sources this summary draws from.
-
Primary sourcePreprint bioRxiv (Cold Spring Harbor Laboratory) · 2026-09-17From Nagelkerke's R2 to Liability-Scale Variance Explained for Polygenic Scores