u/0MGEM0

Genome annotation folks: what do you wish current pipelines did better?

Hi everyone! A collaborator and I are in the early stages of building a new open-source pipeline for whole-genome gene prediction and annotation. Before we get too far into solidifying what it should do, I would really like to hear from people who actually annotate genomes.

What do current tools make harder than it needs to be? What still takes too much manual work or too many custom scripts? What results are difficult to trust? Are there useful tools, types of evidence, or separate parts of your workflow that you wish worked together better?

I am interested in both structural annotation, meaning predicting and refining gene models, and functional annotation. Feedback from any organism or project size is welcome, especially from people working with non-model organisms.

This could include installation and portability, combining gene predictors, incorporating RNA or protein evidence, GFF/GTF wrangling, comparing different annotations, QC, choosing which gene models to keep, manual review, functional annotation, HPC use, reproducibility, or anything else I have not thought of.

If you can, it would be helpful to mention:

  • what organism or taxonomic group you work on
  • what software or workflow you use now
  • what part causes the most frustration or uncertainty
  • what feature or integration would genuinely improve your work

No need to answer every bullet. Anecdotes, wish lists, horror stories, and “please just make X talk to Y” answers are all welcome.

The eventual goal is a pipeline that can take over after genome assembly and help get from “I have an assembly” to “I have an annotation I can trust and actually use.” We do not yet have a finished tool to promote, but we have a skeleton. We are trying to learn what the community needs before we build ourselves into a corner.

If you could change, add, or better connect one thing in genome annotation software, what would it be?

Edit: Got some comments echoing concern about lack of background research on my part, so I thought I would include an edit. I should have made this clearer in the original post.

I have read papers on existing tools and their issue trackers, used and modified existing tools, and built several annotation pipelines both independently and with collaborators. This project is growing out of those experiences, not an assumption that we can start from scratch and solve everything.

I am asking here because issue trackers do not always capture the workarounds people have learned to live with, why they abandoned a tool, or needs that never became a formal issue. I wanted broader, more organic feedback before we lock in the design.

I am approaching this in good faith and with plenty of humility about what I do not know. If you have a specific failure mode or design mistake you think we should avoid, I would genuinely value the input.

reddit.com
u/0MGEM0 — 23 hours ago