Genome Annotation and Other Post-Assembly Workflows for the Tree of Life

Jan 1, 2024·
Tom Brown
,
Kathleen Collier
,
Fernando Cruz
Anestis Gkanogiannis
Anestis Gkanogiannis
,
Sagane Joye-Dind
,
Yannis Nevers
,
Stepan Saenko
,
Tyler Alioto
,
Anthony Bretaudeau
,
Michael Charleston
,
Phuong Duy Doan
,
Christoph Hahn
,
Thomas Harrop
,
Katie Herron
,
Fredrick Kebaso
,
Romane Libouban
,
Locedi Mansueto
,
Shivakumara Manu
,
Asime Oba
,
David Swarbreck
,
Anna Syme
,
Fabio Zanarello
,
Jean-Marc Aury
,
Jèssica Gómez-Garrido
,
Alice Dennis
· 0 min read
Abstract
Rapid advances in genome sequencing technologies have resulted in an explosion of reference-quality genome assemblies across the tree of life. While these resources will be invaluable toward goals of species and biodiversity conservation, their application is limited when they lack accurate annotations of functional elements. The European Reference Genome Atlas (ERGA), the European node of the Earth Biogenome Project (EBP), aims to share resources and knowledge to create fully annotated reference genomes in a distributed manner, bringing together researchers worldwide with common goals and understandings. In the BioHackathon Europe 2023, we constructed and tested tools, pipelines, and workflows for annotating protein-coding regions in assembled genomes, evaluating performance across diverse non-model organisms and the usability of pipelines for newcomers. This required deploying tools across multiple compute environments, sharing genomic resources and expertise across institutes, and assessing workflow performance and input data requirements to achieve high-quality genome annotation. Here we present results from over 20 researchers in 8 time zones working toward robust genome-annotation workflows in eukaryotes.
Type
Publication
BioHackrXiv
publication
Anestis Gkanogiannis
Authors
Senior AI/ML and genomics practitioner

Senior AI/ML and genomics practitioner with ~15 years building open-source, production-grade tools for large-scale biological data. Maintainer of multiple Bioconductor packages (fastreeR, metabinR, jvecfor), and author of agentic, LLM-driven tooling that runs reproducible bioinformatics workflows from natural-language requests.

Broad multi-omics background spanning genome assembly and annotation, population genomics, large-scale NGS and functional-genomics analysis, and metagenomics, backed by reproducible HPC software and end-to-end program leadership. Currently focused on bringing modern AI — embeddings, deep learning, and LLM-based agents — to making complex omics datasets faster and easier to interrogate.