Skip to content

Functional Annotation with EggNOG Mapper

Introduction

EggNOG Mapper annotates novel sequences (genes or proteins) using precomputed eggNOG orthology assignments. Typical uses include annotating genomes, transcriptomes, and metagenomic gene catalogs. Orthology-based transfer is often more specific than simple homology, because annotations are not taken from paralogs that may have diverged in function.

The methodology and the underlying database are described on the eggNOG methods page.

EggNOG Mapper produces its own results object, separate from a functional annotation project: a table where each row is a query sequence with the functional information transferred from its eggNOG orthologs. These annotations can later be merged into a functional annotation project (see Merge EggNOG GOs to a Functional Annotation Project).

Run

EggNOG Mapper can be launched from:

When the job is started from a loaded project, the sequences come from the project and the Input page is skipped. When it is started from the menu without a project, the input files are selected on the Input page first.

Input

The Input page is shown only when no project is loaded.

  • Genes or Proteins: One or more multi-FASTA files containing gene or protein sequences (extensions such as .fasta, .faa, .fna, .fa, .fn, .ffn).

  • Input Type: How sequences are searched against eggNOG (default CDS):

  • Proteins: Amino acid sequences are searched with BLASTp.

  • CDS: Coding sequences are translated to proteins, then searched with BLASTp.
  • Genes/Contigs: Nucleotide sequences are searched in all six reading frames with BLASTx.

OmicsBox can detect whether files contain protein or nucleotide sequences and warns if the chosen Input Type does not match. Sequences longer than 2.5 Mb are filtered automatically and skipped.

Configuration

The Configuration page sets the search and annotation options (Figure 1).

Search filters

  • Minimum Hit e-Value: Minimum e-value for seed ortholog hits in the homology search phase (default 1.0E-3). Lower values are more stringent; higher values are more permissive.

Annotation options

  • Taxonomic Scope: Restrict annotation to orthologs from a chosen clade. Default Adjust Automatically picks the best scope per query sequence. Select a specific taxon when the origin of the data is known.

  • Target Orthologs: Which ortholog relationships are used to transfer function (default All):

  • All: All ortholog types (recommended).

  • One to One: Single ortholog per pair of species (most stringent).
  • One to Many: One-to-one and one-to-many orthologs.
  • Many to One: One-to-one and many-to-one orthologs.
  • Many to Many: All types including many-to-many (least stringent, most comprehensive).

  • GO Evidence: Which Gene Ontology (GO) terms are transferred (default Non-Electronic):

  • Experimental: Only terms with experimental evidence codes (fewer, high-confidence terms).

  • Non-Electronic: All evidence types except electronic inference only (recommended balance of quality and coverage).

Figure 1. EggNOG Mapper wizard, Configuration page.

Results

The result is an EggNOG Mapper results object: a table with one row per query sequence holding the functional information transferred from its eggNOG orthologs (Figure 2). This is a standalone object, not a functional annotation project; its GO terms and enzyme codes can be added to a project afterwards (see Merge EggNOG GOs to a Functional Annotation Project). The run also produces a Summary Report with totals for GO terms, COG categories, and orthologous groups.

Results table

The table can be sorted and filtered. The columns shown by default are:

  • Type. The orthologous group type the annotation is transferred from: COG (prokaryotic Clusters of Orthologous Groups), KOG (eukaryotic orthologous groups), or ENOG (eggNOG non-supervised orthologous groups).
  • Query ID. The query sequence identifier.
  • Gene Name. The predicted preferred gene name.
  • EggNOG Description. The functional description of the matched orthologous group.
  • E-Value. The e-value of the seed ortholog hit.
  • Bit-Score. The bit-score of the seed ortholog hit.
  • Best Tax-Level. The taxonomic level of the orthologous group from which the annotation was transferred.
  • EC Codes. The Enzyme Commission numbers assigned.
  • #GO. The number of GO terms transferred.
  • GOs. The GO term identifiers.
  • GO Names. The names of the GO terms.
  • KEGG KO. The KEGG Orthology (KO) identifiers.
  • KEGG Pathway. The KEGG pathways.

The following columns are hidden by default and can be shown from the column selector:

  • EggNOG Protein. The seed ortholog protein matched in eggNOG.
  • Taxa Scope. The taxonomic scope used for the annotation.
  • KEGG Module, KEGG Reaction, KEGG RCLASS, and KEGG TC. Additional KEGG cross-references: modules, reactions, reaction classes, and transporter classification.
  • Brite. KEGG BRITE functional hierarchies.
  • CAZy. Carbohydrate-Active enZymes families.
  • BiGG. BiGG metabolic reaction identifiers.
  • Matching OGs. The orthologous groups matched by the query.
  • COG Categories. The COG functional category classification.

Figure 2. EggNOG Mapper results table.

Side Panel

Actions

Export

  • Export Table. Export the results table to a tab-separated text file.

Context Menu

Right-click a row to open its context menu. Apart from the generic Context Menu options, the EggNOG results have:

  • Show GOs. Open the GO terms assigned to the selected sequence.
  • Show Annotation Details. Open a detailed report of the eggNOG annotation for the sequence, including link-outs and GO information (Figure 3).

Figure 3. EggNOG Mapper annotation details.

Merge EggNOG GOs to a Functional Annotation Project

The EggNOG Mapper results are a standalone object. To incorporate their annotations into a functional annotation project, open the project and, in its Side Panel → EggNOG → Merge EggNOG GOs, select the EggNOG Mapper results to merge.

The GO terms and enzyme codes (EC) from the EggNOG results are added to the project's annotations; when a sequence already has annotations, the new ones are added to the existing set. The annotations to merge can be filtered by e-value or bit-score. When the merge finishes, a bar chart reports the total number of GO terms and enzyme codes added to the project.

OmicsBox Engine

This tool can be run from the command line via the OmicsBox Engine.

Command: omicsbox eggnog-mapper [options]

Inputs

Flag Type Required Description
--i-sequences file (multiple) No Genes or Proteins
--i-type-options enum No Input Type
--i-local-project file No Sequence Project (.box)

Parameters

Flag Type Default Range / Candidates Required Description
--taxonomic-scope enum auto auto
84992
225057
373 more57723
204432
422676
201174
85005
622450
7898
186827
135624
311790
5338
155619
355688
506
186823
1283313
28211
72275
256005
554915
150247
5794
200783
290174
2157
183980
178469
34384
6656
4890
71274
255475
8782
539002
2836
91061
1386
2
815
976
1100069
200643
772
5204
213481
45404
28216
423358
85004
33213
39782
572511
41294
3699
85019
118882
32199
119060
830
1016
33554
186828
28883
204458
85016
104264
10
314294
91561
35718
451870
9397
204428
1090
200795
32061
3041
7711
119089
135613
59732
5878
544
34397
186801
31979
538999
5796
267889
80864
84998
1653
28889
246874
1117
167375
43988
768503
766764
200930
301297
1297
68525
28221
85020
145357
85018
213118
213115
213113
69541
204037
7147
326319
189330
147541
451867
7214
35237
35325
308865
547
81852
29547
551
526524
561
314146
186806
2759
5042
147545
28890
91835
1239
117743
237
85013
4751
112252
10474
32066
1236
142182
129337
244698
1028384
85014
85026
74385
5819
53433
45667
183963
10404
548681
9604
119069
7399
45401
69657
5129
5125
267893
10860
50557
85021
5653
2063
1164882
1506553
33958
1357
283735
118969
191028
147548
7088
81850
11989
1511857
4447
10477
186820
245186
400634
639021
40674
252356
558415
9263
33208
183925
183939
224756
119045
135618
31993
85023
1268
85008
10841
11157
468
1762
10662
76831
29
110618
909932
206351
7148
6231
76804
206350
40117
85025
363408
1161
252301
182709
135619
33183
34383
5151
33154
414999
265975
1150
216572
75682
186822
53335
265
33342
135625
122277
186807
1570339
186804
4776
302485
69277
464095
5863
203682
186818
92860
52604
38820
10744
52959
289201
171551
9443
1212
85017
85009
1224
583
586
267888
46205
136841
136843
136845
136846
136849
85010
83612
267894
29000
121069
34037
35268
11632
6236
82115
1060
119043
206389
204441
34008
766
171550
9989
93682
2433
74030
84995
157897
97050
541000
4893
4891
590
5809
613
267890
10699
5148
5139
147550
117747
204457
203691
186821
29258
35301
35278
439488
90964
1189
228398
671232
119603
68892
28037
1303
1313
1305
1307
629295
35493
85012
60136
995019
1129
1142
508458
213462
68298
451866
82986
10656
544448
8459
651137
186824
68295
183968
200940
189775
183967
200918
285107
72273
5234
675063
52018
82117
2323
119066
119065
186813
191675
33867
61432
118884
186928
58840
326457
452284
74201
203494
7742
135623
84406
33090
10239
335928
135614
629
No Taxonomic Scope
--target-orthologs enum all all
one2one
many2one
one2many
many2many
No Target Orthologs
--go-evidence enum non-electronic experimental
non-electronic
No GO Evidence
--blast-expect-value enum 1.0E-3 1000
10
5
10 more1
0.1
1.0E-3
1.0E-5
1.0E-10
1.0E-15
1.0E-25
1.0E-50
1.0E-75
1.0E-100
No Minimum Hit e-Value

Global options (--local-folder, --cloud-folder, --output-format, --config, --detach, --verbose, …) are shared by every Engine tool and are not repeated here — see the OmicsBox Engine reference.

References

Huerta-Cepas J et al. (2019). eggNOG 5.0: a hierarchical, functionally and phylogenetically annotated orthology resource based on 5090 organisms and 2502 viruses. Nucleic Acids Research, 47(D1), D309-D314.