Prediction of crossover recombination using parental genomes

Feb 16, 2023·
Mauricio Peñuela
,
Camila Riccio-Rengifo
,
Jorge Finke
,
Camilo Rocha
Anestis Gkanogiannis
Anestis Gkanogiannis
,
Rod A. Wing
,
Mathias Lorieux
· 0 min read
Abstract
Meiotic recombination is a crucial cellular process, being one of the major drivers of evolution and adaptation of species. In plant breeding, crossing is used to introduce genetic variation among individuals and populations. While different approaches to predict recombination rates for different species have been developed, they fail to estimate the outcome of crossings between two specific accessions. This paper builds on the hypothesis that chromosomal recombination correlates positively to a measure of sequence identity. It presents a model that uses sequence identity, combined with other features derived from a genome alignment (including the number of variants, inversions, absent bases, and CentO sequences) to predict local chromosomal recombination in rice. Model performance is validated in an inter-subspecific indica x japonica cross, using 212 recombinant inbred lines. Across chromosomes, an average correlation of about 0.8 between experimental and prediction rates is achieved. The proposed model, a characterization of the variation of the recombination rates along the chromosomes, can enable breeding programs to increase the chances of creating novel allele combinations and, more generally, to introduce new varieties with a collection of desirable traits. It can be part of a modern panel of tools that breeders can use to reduce costs and execution times of crossing experiments.
Type
Publication
PLOS ONE 2023, 18(2), 1-21;, 18(2), 1-21. Public Library of Science
publication
Anestis Gkanogiannis
Authors
Senior AI/ML and genomics practitioner

Senior AI/ML and genomics practitioner with ~15 years building open-source, production-grade tools for large-scale biological data. Maintainer of multiple Bioconductor packages (fastreeR, metabinR, jvecfor), and author of agentic, LLM-driven tooling that runs reproducible bioinformatics workflows from natural-language requests.

Broad multi-omics background spanning genome assembly and annotation, population genomics, large-scale NGS and functional-genomics analysis, and metagenomics, backed by reproducible HPC software and end-to-end program leadership. Currently focused on bringing modern AI — embeddings, deep learning, and LLM-based agents — to making complex omics datasets faster and easier to interrogate.