Preprint · not peer-reviewed Researchers Educators Students

Preprint uses large language models to audit population descriptor use in genetics research

A preprint from Cold Spring Harbor Laboratory applies LLM-based bibliometric analysis to assess how race, ethnicity, and ancestry are described across human genetics literature, benchmarking practice against 2023 NASEM recommendations.

Published · AI-drafted summary based on 1 public source
Illustration for ancestry story
Illustrative image — not from the source article.
Share

A preprint posted to bioRxiv applies large language model-based bibliometric methods to evaluate how population descriptors — including race, ethnicity, and ancestry — are used in human genetics and genomics publications. The work responds directly to the 2023 report from the National Academies of Sciences, Engineering, and Medicine (NASEM), titled *Using Population Descriptors in Genetics and Genomics Research: A New Framework for an Evolving Field*, which issued eight actionable recommendations for researchers.

The authors used language model-assisted classification to systematically assess whether and how published genetics studies conform to those NASEM recommendations, characterising patterns across a large corpus of literature. The bibliometric approach is designed to scale the kind of qualitative audit that would be impractical to perform manually across thousands of papers.

The preprint has not yet undergone peer review. Its findings are relevant to researchers thinking about how to describe study populations accurately and ethically, to educators teaching research methods and responsible conduct, and to students beginning to engage with primary genetics literature. The work also bears on ongoing debates in the field about the relationship between social categories and biological ancestry, and how that relationship should be communicated in scientific writing.

No specific institutional affiliation or lead author is named in the feed lede; full authorship details are available at the bioRxiv record.

Sources

Read the original reporting — these are the public sources this summary draws from.

  1. Primary sourcePreprint bioRxiv (Cold Spring Harbor Laboratory) · 2026-09-28
    Large language model-based bibliometric evaluation of population descriptors in human genetics

Tags

population-descriptors research-ethics bibliometrics nasem ancestry race-ethnicity large-language-models
Share

About Genetic Current

Educational summaries of public genetics news

Genetic Current is the news section of Evagene, an academic, research, and educational pedigree-modelling platform. Stories are AI-drafted summaries of items from trusted public sources, written for researchers, clinicians, educators, students, genealogists, and patients with an interest in genetics. Summaries are for educational and research purposes only and are not medical advice.

Join the Evagene Alpha Waiting List