UniLasso sparse regression framework scales polygenic risk score computation to biobank data
A team including researchers from Stanford University has adapted a two-stage penalised regression method, uniLasso, to compute polygenic risk scores efficiently in high-dimensional genomic datasets.
Researchers including Joshua Richland, Tuomo Kiiskinen, William Wang, Wenhui Sophia Lu, Balasubramanian Narasimhan, Trevor Hastie, Manuel Rivas, and Robert Tibshirani have published a study in *PLOS Genetics* introducing a scalable framework for computing polygenic risk scores (PRS) in biobank-scale, high-dimensional genomic settings.
The framework is built around Univariate-Guided Sparse Regression (uniLasso), a two-stage penalised regression procedure. In the first stage, univariate association coefficients and their magnitudes are used to guide feature selection; in the second, a sparse predictive model is fitted using this informed prior. The authors argue this approach stabilises variable selection in settings where the number of genetic variants vastly exceeds sample size — a persistent challenge in PRS methodology.
The paper provides both theoretical justification and empirical benchmarking, reporting improved predictive performance relative to established PRS methods across several traits. The authors frame the contribution primarily as a methodological advance for statistical genetics and population genomics research.
The study complements a growing body of work on PRS methodology published recently, including efforts to improve score transferability across ancestries and to correct for biases introduced by standard logistic regression frameworks — topics that have featured in several recent preprints and PLOS Genetics papers.
Sources
Read the original reporting — these are the public sources this summary draws from.
-
Primary source Public Library of Science · 2026-09-23Univariate-guided sparse regression for Biobank-scale high-dimensional omics data