Skip to content

EviAnn

Introduction

EviAnn is an evidence-based gene annotation system for eukaryotic genomes. Unlike ab initio gene finders, which predict genes from a statistical model trained on a particular species, EviAnn builds gene structures directly from evidence: RNA-seq or assembled transcript alignments combined with protein homology from related species.

This approach has two practical consequences. Repeats do not need to be soft-masked beforehand, and no species-specific training is required. EviAnn annotates protein-coding genes and long non-coding RNAs, and handles genomes up to 32 Gbp as well as organelles.

The annotation is produced as GFF3 that passes NCBI submission checks, optionally accompanied by the predicted protein and transcript sequences.

This functionality can be found under Genome Analysis → Gene Finding → Eukaryotic Gene Finding.

Please cite EviAnn as:

Zimin AV, Puiu D, Pertea M, Yorke JA, Salzberg SL (2026). "Efficient evidence-based genome annotation with EviAnn." Nature Methods, 23(8), 1521-1527.

Run EviAnn annotation

EviAnn has no mode selector. The way it annotates follows from the evidence provided, so the choice of inputs is the main decision.

Genome

A genome FASTA file is always required. Repeats do not need to be masked.

All sequence inputs may be supplied either uncompressed or gzip-compressed, in any combination. Compressed files are uploaded as they are and expanded in the cloud, so there is no need to decompress anything beforehand.

Transcript evidence

EviAnn requires transcript evidence, and there are two ways to supply it. At least one must be enabled.

Use RNA-Seq Reads. Sequencing reads from the species being annotated. This is the primary mode and generally produces the most complete annotation, because the evidence comes from the organism itself. Illumina single-end and paired-end reads, PacBio Iso-seq long reads, and already-aligned BAM files are all accepted, and different sample types can be combined in one run.

Combine sequencing runs so that each file, pair, or triplet corresponds to a single experiment. EviAnn assembles transcripts per sample and merges the results afterwards.

Use Transcripts From a Related Species. Assembled transcripts from a close relative, used when no RNA-seq data is available for the target organism. The relative should be more than 95 % identical at the DNA level for this to work well.

Both options can be enabled together, in which case EviAnn uses all the evidence provided.

Protein evidence

A protein FASTA file from related species guides the gene models. Concatenate proteins from several relatives into a single file. As a rule of thumb, provide 5 to 10 times as many proteins as the number of genes expected: roughly 100,000 to 200,000 for an insect, or roughly 500,000 for a typical plant or mammalian genome.

Note

If no protein file is provided, EviAnn downloads the UniProt database during the run. This requires internet access from the analysis and makes the result depend on UniProt availability. Providing proteins from close relatives is strongly recommended.

Additional evidence

Three optional inputs refine the annotation. Each is enabled by its own checkbox.

Use Existing CDS Annotations. A GFF file with CDS features in the coordinates of the genome being annotated. Each CDS must carry gene, transcript or mRNA, exon, and CDS attributes.

Add Extra Features. An external GFF file whose features are merged into the final annotation. Features must have gene records. Any that overlap existing annotations are ignored.

Specify Mitochondrial Contigs. A plain text file listing the contigs to annotate with the mitochondrial genetic code, in which AGA, AGG, TAA, and TAG are stop codons.

Configuration

Maximum Intron Size. Maximum intron length in base pairs. Leave this at 0 to let EviAnn derive the value from the genome size and ploidy. The derived value is usually a better choice than a fixed one.

Ploidy. Ploidy of the genome. This is used only when estimating the maximum intron size, so it has no effect when that value is set explicitly.

Include Partial CDS Transcripts. Include transcripts whose coding sequence is missing a start or stop codon. This increases sensitivity at the cost of more incomplete gene models.

Minimum lncRNA TPM. Minimum expression, in transcripts per million, for a non-coding transcript to be annotated as a long non-coding RNA. Raising this value yields fewer but better-supported lncRNAs.

Minimum Protein Length. Minimum protein length in amino acids for ab initio open reading frame detection where no homology evidence is available.

Output

Annotation GFF3. Destination for the resulting annotation.

Also Produce Protein and Transcript Sequences. When enabled, the predicted protein and transcript sequences are delivered as FASTA files to a chosen folder.

Results

Summary report

An HTML report opens when the analysis finishes. It contains the effective analysis parameters, a table of annotation statistics, and the EviAnn citation.

The statistics are the counts reported by EviAnn itself, so they always agree with the annotation produced:

Statistic Meaning
Number of genes All gene features
Number of protein coding genes Genes tagged as protein coding, excluding pseudogenes
Number of processed pseudo gene transcripts Transcripts marked as pseudogenes
Number of processed pseudo genes Genes marked as pseudogenes
Number of transcripts All mRNA features
Number of long non-coding RNAs Annotated lncRNA features
Number of predicted protein sequences Records in the protein FASTA file
Number of predicted transcript sequences Records in the transcript FASTA file

Annotation

The annotation opens in OmicsBox as a browsable object, with gene, mRNA, exon, CDS, and lncRNA features. The same annotation is written to the chosen destination as GFF3.

When protein and transcript sequences are requested, they are written to the selected folder as FASTA files.

OmicsBox Engine

This tool can be run from the command line via the OmicsBox Engine.

Command: omicsbox eviann [options]

Inputs

Flag Type Required Description
--i-genome-file file Yes Genome FASTA
--i-rna-seq-reads-single-end file (multiple) No Single-End
--i-rna-seq-reads-paired-end file (multiple) No Paired-End
--i-rna-seq-reads-pacbio-hifi file (multiple) No PacBio Iso-seq
--i-rna-seq-reads-alignments file (multiple) No Aligned BAM
--i-transcripts-file file No Transcripts FASTA
--i-proteins-file file No Protein Sequences FASTA
--i-cds-gff-file file No CDS GFF File
--i-extra-gff-file file No Extra Features GFF File
--i-mito-contigs-file file No Mitochondrial Contigs List

Parameters

Flag Type Default Range / Candidates Required Description
--use-rna-seq boolean true No Use RNA-Seq Reads
--upstream-pattern string _1 No Upstream Files Pattern
--downstream-pattern string _2 No Downstream Files Pattern
--use-transcripts boolean false No Use Transcripts From a Related Species
--use-cds-gff boolean false No Use Existing CDS Annotations
--use-extra-gff boolean false No Add Extra Features
--use-mito-contigs boolean false No Specify Mitochondrial Contigs
--max-intron-size integer 0 0 – 10000000 No Maximum Intron Size
--ploidy integer 2 1 – 20 No Ploidy
--include-partial boolean false No Include Partial CDS Transcripts
--lnc-rna-min-tpm double 0.5 0.0 – 1000.0 No Minimum lncRNA TPM
--min-protein-length integer 75 1 – 10000 No Minimum Protein Length
--provide-sequences boolean true No Also Produce Protein and Transcript Sequences
--o-annotation-file string No Annotation GFF3
--o-sequences-folder string No Sequences Output Folder

Parameter relationships

Flag When Effect Affected flags
--use-rna-seq true enables --i-rna-seq-reads-single-end, --i-rna-seq-reads-paired-end, --i-rna-seq-reads-pacbio-hifi, --i-rna-seq-reads-alignments, --upstream-pattern, --downstream-pattern
--use-transcripts true enables --i-transcripts-file
--use-cds-gff true enables --i-cds-gff-file
--use-extra-gff true enables --i-extra-gff-file
--use-mito-contigs true enables --i-mito-contigs-file
--provide-sequences true enables --o-sequences-folder

Warning

Either --use-rna-seq or --use-transcripts must be enabled, together with the corresponding input. Enabling neither leaves EviAnn without transcript evidence and the analysis stops.

Global options (--local-folder, --cloud-folder, --output-format, --config, --detach, --verbose, …) are shared by every Engine tool and are not repeated here — see the OmicsBox Engine reference.