Preprint benchmarks five machine-learning classifiers for CRISPR-Cas9 off-target prediction from GUIDE-seq data
A bioRxiv preprint evaluates logistic regression, random forest, gradient boosting, CNN, and ensemble approaches on a class-imbalanced GUIDE-seq dataset, exposing methodological pitfalls relevant to therapeutic genome editing safety.
A preprint posted to bioRxiv (Cold Spring Harbor Laboratory) benchmarks five machine-learning classifiers — logistic regression on mismatch-count features, random forest, gradient boosting, a one-dimensional convolutional neural network (CNN), and a gradient-boosting/CNN ensemble — on their ability to classify CRISPR-Cas9 off-target cleavage sites from published GUIDE-seq data (Kleinstiver et al., 2016, *Nature*).
Off-target double-strand breaks are a central safety concern for therapeutic CRISPR applications, and predicting which candidate genomic sites will be cleaved by a given single-guide RNA (sgRNA) remains a difficult problem in part because genuine off-target sites are rare relative to the vast number of possible candidate sequences — a pronounced class-imbalance problem. The preprint frames this as a machine-learning benchmark, making explicit how different modelling choices and evaluation metrics perform under realistic imbalance conditions.
The work is relevant to researchers developing or assessing CRISPR-based therapeutics and to those building computational pipelines for genome-editing safety evaluation. It also provides a methodological reference for computational biologists working on similar classification problems in genomics.
This is a preprint and has not yet undergone peer review. Results and conclusions should be treated as preliminary.
Sources
Read the original reporting — these are the public sources this summary draws from.
-
Primary sourcePreprint bioRxiv (Cold Spring Harbor Laboratory) · 2026-08-23Classifying CRISPR-Cas9 Off-Target Cleavage Sites from GUIDE-seq Data: A Class-Imbalanced Machine Learning Benchmark