1 Introduction

Gene identifier conversion and annotation is a common and critical task in bioinformatics research. Existing databases and tools use different naming conventions for genes or provide only partial annotations, making it challenging to integrate data from multiple sources. geneslator addresses this problem by providing a unified interface for genome annotation across different databases in several model organisms.

Key Features:

  • Multiple database integration: Integrates data from cross-organism databases (NCBI, Ensembl, UniProt, Alliance of Genome Resources, GO, KEGG, Reactome, Wikipathways) and organism-specific resources (HGNC, MGI, RGD, SGD, WormBase, Flybase, ZFIN, TAIR);
  • Archive search: Supports searching using both current and archived gene identifiers in NCBI and Ensembl databases;
  • Alias resolution: Supports automatic disambiguation between symbols and aliases in annotations involving gene symbols;
  • Multi-organism support: Integrates annotation data from several model organisms into a unique R package.

geneslator currently supports 14 different organisms, including Human, Mouse, Rat, Yeast, Worm, Fly, Zebrafish, Arabidopsis and other plant species. In the future releases of the package, the number of supported species will increase to include more model organisms.

2 Installation

if (!require("BiocManager", quietly = TRUE)) {
  install.packages("BiocManager")
}
BiocManager::install("geneslator")

3 Load the package

library(geneslator)

4 Import annotation databases

geneslator provides species-specific annotation databases for several organisms. Annotation databases are stored as SQLite files in different versions of a Zenodo record at https://doi.org/10.5281/zenodo.20448208. Each release refers to a specific version of the databases. Versions are tagged as year.month, where year and month denote the year and the month of the publication of the release (e.g. ‘2026.03’ for March 2026). Databases are updated on a monthly basis.

Type availableDatabases() to retrieve the list of available databases and supported species in the most recent release.

# List organisms annotated in geneslator
availableDatabases()
#>                     Name                 Organism  TaxID
#> 16     org.Amellifera.db           Apis mellifera   7460
#> 8       org.Athaliana.db     Arabidopsis thaliana   3702
#> 10         org.Bnapus.db           Brassica napus   3708
#> 9       org.Boleracea.db        Brassica oleracea   3712
#> 6        org.Celegans.db   Caenorhabditis elegans   6239
#> 4          org.Drerio.db              Danio rerio   7955
#> 5   org.Dmelanogaster.db  Drosophila melanogaster   7227
#> 1        org.Hsapiens.db             Homo sapiens   9606
#> 13 org.Langustifolius.db    Lupinus angustifolius   3871
#> 15       org.Mmulatta.db           Macaca mulatta   9544
#> 2       org.Mmusculus.db             Mus musculus  10090
#> 18        org.Osativa.db             Oryza sativa  39947
#> 14      org.Pvulgaris.db       Phaseolus vulgaris   3885
#> 3     org.Rnorvegicus.db        Rattus norvegicus  10116
#> 7     org.Scerevisiae.db Saccharomyces cerevisiae 559292
#> 11  org.Slycopersicum.db     Solanum lycopersicum   4081
#> 12      org.Vvinifera.db           Vitis vinifera  29760
#> 17        org.Xlaevis.db           Xenopus laevis   8355
#> 19          org.Zmays.db                 Zea mays   4577
#>                                 MD5 Version                     DOI
#> 16 b8f86715dfeaf5c80d3203c9a7eed3c0 2026.08 10.5281/zenodo.22066305
#> 8  37a6d4797d8d2f5fc86f5cc5b663f3ff 2026.08 10.5281/zenodo.22066305
#> 10 3e2a212913f4e3c65b27ef6b3d6cb50e 2026.08 10.5281/zenodo.22066305
#> 9  014de35fcc6bafe0e88cc80ab6baeaf1 2026.08 10.5281/zenodo.22066305
#> 6  6b5d35e4309230d7007aee52a9cc6124 2026.08 10.5281/zenodo.22066305
#> 4  a302d3e5a80437e62849b68210226245 2026.08 10.5281/zenodo.22066305
#> 5  212985069f50559df9abe6075e43f9b9 2026.08 10.5281/zenodo.22066305
#> 1  d01b9ffa7fef66798d602466b0a50059 2026.08 10.5281/zenodo.22066305
#> 13 bd36b7a7b3db6a85ccd18e0d6430cb69 2026.08 10.5281/zenodo.22066305
#> 15 efd33ae5455717513fa28540f40946ba 2026.08 10.5281/zenodo.22066305
#> 2  38726d1b781159fe9bfd00dab8690aaa 2026.08 10.5281/zenodo.22066305
#> 18 e51d45ff6afb0bfc98284fcb999ee563 2026.08 10.5281/zenodo.22066305
#> 14 1501c039beb58818bd378e5fb9ad4b23 2026.08 10.5281/zenodo.22066305
#> 3  772b932eaa7fcdc2a577b1f9c75e963a 2026.08 10.5281/zenodo.22066305
#> 7  96b8319f446922bd38afb51a0cdf7962 2026.08 10.5281/zenodo.22066305
#> 11 29d62767d69463bfd0cd66cda5bb3078 2026.08 10.5281/zenodo.22066305
#> 12 174eec4749a2629f17aefd72559c9308 2026.08 10.5281/zenodo.22066305
#> 17 c23f1758d13dcfc02bcba156372a209f 2026.08 10.5281/zenodo.22066305
#> 19 4154960b8b59c5208fca883980642921 2026.08 10.5281/zenodo.22066305

The parameter release.version can be used to retrieve the list of all available databases in an older release.

# List organisms annotated in geneslator (release December 2025)
availableDatabases(release.version = "2025.12")
#>                   Name                 Organism  TaxID
#> 1     org.Athaliana.db     Arabidopsis thaliana   3702
#> 6      org.Celegans.db   Caenorhabditis elegans   6239
#> 8        org.Drerio.db              Danio rerio   7955
#> 2 org.Dmelanogaster.db  Drosophila melanogaster   7227
#> 3      org.Hsapiens.db             Homo sapiens   9606
#> 4     org.Mmusculus.db             Mus musculus  10090
#> 5   org.Rnorvegicus.db        Rattus norvegicus  10116
#> 7   org.Scerevisiae.db Saccharomyces cerevisiae 559292
#>                                MD5 Version                     DOI
#> 1 a292153eee87600c5d8c27977fe7ea45 2025.12 10.5281/zenodo.20448209
#> 6 fb4f03098e379712c17196a1f2b6c6a4 2025.12 10.5281/zenodo.20448209
#> 8 e292dcb2cca5c038d9c369b31ca16d8c 2025.12 10.5281/zenodo.20448209
#> 2 06031138af0a7e44af7d9f938f8f4239 2025.12 10.5281/zenodo.20448209
#> 3 6b6ffd437724b029e3ec5f24ab866d97 2025.12 10.5281/zenodo.20448209
#> 4 1f5af73caf5e89f65e7bcf31669f62d0 2025.12 10.5281/zenodo.20448209
#> 5 7cb6dbed9441b0b142032a5206b66126 2025.12 10.5281/zenodo.20448209
#> 7 7b36f0be0eecce6e12bf05f32d5d8779 2025.12 10.5281/zenodo.20448209

A complete list of all available release versions can be obtained with availableVersions().

# List available versions of geneslator annotation databases
availableVersions()
#> [1] "2025.12" "2026.03" "2026.04" "2026.05" "2026.06" "2026.07" "2026.08"

To query a database for a specific organism org, you first need to import it, by using the GeneslatorDb function. org can be either the scientific name of the organism (e.g. “Homo sapiens”) or its Taxonomy ID (e.g. “10090” for Mouse). The function creates a new GeneslatorDb object for the requested database, which is then exported to the global environment of the user as a variable having the same name of the SQLite annotation database (e.g. org.Hsapiens.db for Human, org.Mmusculus.db for Mouse).

# Import human annotation db (after downloading it from remote repository)
GeneslatorDb("Homo sapiens")
# Info about the imported human annotation database object
org.Hsapiens.db
#> GeneslatorDb object
#> Organism: Homo sapiens 
#> Columns: ALIAS, ENSEMBL, ENSEMBLOLD, ENTREZID, ENTREZIDOLD, GENENAME, GENETYPE, GID, GO, GOEVIDENCE, GONAME, GOTYPE, HGNC, KEGGPATH, KEGGPATHNAME, LOCUS, ORTHOFLY, ORTHOMOUSE, ORTHORAT, ORTHOWORM, ORTHOYEAST, ORTHOZEBRAFISH, REACTOMEPATH, REACTOMEPATHNAME, SYMBOL, UNIPROT, WIKIPATH, WIKIPATHNAME
# Import mouse annotation database using its Taxonomy ID
GeneslatorDb("10090")
# Info about the imported human annotation database object
org.Mmusculus.db
#> GeneslatorDb object
#> Organism: Mus musculus 
#> Columns: ALIAS, ENSEMBL, ENSEMBLOLD, ENTREZID, ENTREZIDOLD, GENENAME, GENETYPE, GID, GO, GOEVIDENCE, GONAME, GOTYPE, KEGGPATH, KEGGPATHNAME, MGI, ORTHOFLY, ORTHOHUMAN, ORTHORAT, ORTHOWORM, ORTHOYEAST, ORTHOZEBRAFISH, REACTOMEPATH, REACTOMEPATHNAME, SYMBOL, UNIPROT, WIKIPATH, WIKIPATHNAME

When called for the first time on a specific organism, GeneslatorDb function downloads the annotation database from the remote repository, stores a local copy into your R cache folder and finally imports the database. Future calls to GeneslatorDb function will simply import the database from your cache, unless a new version of the database is present in the remote repository. In the latter case, you will be notified about that and you will be able to choose whether or not updating your local copy in the R cache, before importing the database.

# Import human db again. Now cache data will be used to import db
GeneslatorDb("Homo sapiens")

By default, GeneslatorDb queries the latest release. To retrieve an older version of the database, you can set the release.version parameter to the desired release version. Again, a local copy of the database (distinct from the latest release) will be stored into your R cache folder, so that future calls to the same database will simply import it from your cache.

# Import yeast annotation db from release 2025.12 (December 2025)
GeneslatorDb("Saccharomyces cerevisiae", release.version = "2025.12")
# Info about the imported human annotation database object
org.Scerevisiae.db
#> GeneslatorDb object
#> Organism: Saccharomyces cerevisiae 
#> Columns: ALIAS, ENSEMBL, ENSEMBLOLD, ENTREZID, ENTREZIDOLD, GENENAME, GENETYPE, GID, GO, GOEVIDENCE, GONAME, GOTYPE, KEGGPATH, KEGGPATHNAME, LOCUS, ORTHOFLY, ORTHOHUMAN, ORTHOMOUSE, ORTHORAT, ORTHOWORM, ORTHOZEBRAFISH, REACTOMEPATH, REACTOMEPATHNAME, SGD, SYMBOL, UNIPROT, WIKIPATH, WIKIPATHNAME

5 Columns and values of annotation databases

Annotation databases are internally represented as collections of R dataframes that can be queried through functions that map a set of values of an input column (the key) of a dataframe to the corresponding values of one or more output columns of the same or a different dataframe.

Function keytypes() lists all columns that can be used as keys.

# Get all columns that can be used as keys in mouse annotation db
geneslator::keytypes(org.Mmusculus.db)
#>  [1] "ALIAS"          "ENSEMBL"        "ENSEMBLOLD"     "ENTREZID"      
#>  [5] "ENTREZIDOLD"    "GENENAME"       "GENETYPE"       "GO"            
#>  [9] "KEGGPATH"       "MGI"            "ORTHOFLY"       "ORTHOHUMAN"    
#> [13] "ORTHORAT"       "ORTHOWORM"      "ORTHOYEAST"     "ORTHOZEBRAFISH"
#> [17] "REACTOMEPATH"   "SYMBOL"         "UNIPROT"        "WIKIPATH"

Similarly, function columns() lists all possible output columns.

# Get all available types of output values in mouse annotation db
geneslator::columns(org.Mmusculus.db)
#>  [1] "ALIAS"            "ENSEMBL"          "ENSEMBLOLD"       "ENTREZID"        
#>  [5] "ENTREZIDOLD"      "GENENAME"         "GENETYPE"         "GO"              
#>  [9] "GOEVIDENCE"       "GONAME"           "GOTYPE"           "KEGGPATH"        
#> [13] "KEGGPATHNAME"     "MGI"              "ORTHOFLY"         "ORTHOHUMAN"      
#> [17] "ORTHORAT"         "ORTHOWORM"        "ORTHOYEAST"       "ORTHOZEBRAFISH"  
#> [21] "REACTOMEPATH"     "REACTOMEPATHNAME" "SYMBOL"           "UNIPROT"         
#> [25] "WIKIPATH"         "WIKIPATHNAME"

Note that the output of the two functions is different, because only identifier columns can be used as keys, while any column can be an output column. Type help("columns","geneslator") to see the complete list of columns available in the annotation databases of geneslator, together with their description.

Function keys() is used to retrieve all values of a column in an annotation database.

# Get the first 10 Entrez IDs in mouse annotation db
head(geneslator::keys(org.Mmusculus.db, keytype = "ENTREZID"), 10)
#>  [1] "100008564" "100008567" "100009600" "100009609" "100009614" "100009664"
#>  [7] "100009698" "100010"    "100012"    "100014"

6 Query the annotation databases

Columns of the annotation databases can be queried using properly re-defined versions of the well-known query functions select() and mapIds() of AnnotationDbi R package.

The select() function allows you to query an input key column of the annotation database (keytype argument) and retrieve related information across one or more other columns (columns argument).

The output of select() is a dataframe with all columns specified by keytype and columns arguments and one row for each mapping found between input and output values.

# Map NCBI Gene IDs to gene symbols and Ensembl IDs in Human
genes <- c("1", "2", "9")
result <- geneslator::select(org.Hsapiens.db,
  keys = genes,
  columns = c("SYMBOL", "ENSEMBL"), keytype = "ENTREZID"
)
result
#>   ENTREZID SYMBOL         ENSEMBL
#> 1        1   A1BG ENSG00000121410
#> 2        2    A2M ENSG00000175899
#> 3        9   NAT1 ENSG00000171428

Unlike select(), mapIds() maps an input key column (argument keytype) to a single output column (argument column).

# Convert gene symbols to ENTREZ IDs (first match only)
genes <- c("TP53", "BRCA1", "EGFR")
entrez_ids <- geneslator::mapIds(org.Hsapiens.db,
  keys = genes,
  column = "ENTREZID", keytype = "SYMBOL"
)
entrez_ids
#>   TP53  BRCA1   EGFR 
#> "7157"  "672" "1956"

By default, the return type is a named vector, where each value is the first mapping found (if any) for a given key, even if multiple output values map to that key. However, this behaviour can be changed through the multiVals parameter, which also controls the shape of the output result. For example, multiVals="list" produces a list object with all matches found for each input.

# Get all possible mappings as a list
entrez_list <- geneslator::mapIds(org.Hsapiens.db,
  keys = genes,
  column = "ENTREZID", keytype = "SYMBOL", multiVals = "list"
)
entrez_list
#> $TP53
#> [1] "7157"
#> 
#> $BRCA1
#> [1] "672"
#> 
#> $EGFR
#> [1] "1956"

7 Search options

7.1 Search using aliases

In select() and mapIds() functions, by default, queries of annotation databases involving gene symbols are performed by first looking at column “SYMBOL” and, if no mapping is found using “SYMBOL”, the query is performed using the “ALIAS” column. This is helpful when users unknowingly start from a list of names that is actually a mix of official gene symbols and aliases.

This behaviour of select() and mapIds() can be controlled through the boolean parameter search.aliases, whose default value is TRUE.

In the following example, “BRCAI” is actually an alias of BRCA1 gene, while “PTEN” is the official symbol of the PTEN gene. When mapping these two keys (treated as SYMBOL) to ENTREZID by using select(), BRCAI is correctly viewed as an alias of BRCA1 gene and mapped to the NCBI gene id of BRCA1.

# Map gene symbols to their NCBI gene ids, querying also the ALIAS column
# if needed
result <- geneslator::select(org.Hsapiens.db,
  keys = c("BRCAI", "PTEN"),
  columns = "ENTREZID", keytype = "SYMBOL"
)
result
#>   SYMBOL ENTREZID
#> 1  BRCAI      672
#> 2   PTEN     5728

Whenever ALIAS column is used in place of SYMBOL column (as in this example), a warning message is sent to the user. If we repeat the same query with search.aliases=FALSE, select() is unable to map BRCAI to the correct NCBI gene id.

# Map gene symbols to their NCBI gene ids, querying only the SYMBOL column
result <- geneslator::select(org.Hsapiens.db,
  keys = c("BRCAI", "PTEN"),
  columns = "ENTREZID", keytype = "SYMBOL", search.aliases = FALSE
)
result
#>   SYMBOL ENTREZID
#> 1  BRCAI     <NA>
#> 2   PTEN     5728

7.2 Search using archived identifiers

Gene identifiers and symbols can change over time or become deprecated, as a result of periodic updates of databases such as NCBI or Ensembl. This could be troublesome in annotation tasks, especially when user starts from an old set of identifiers or symbols. To overcome this, annotation databases in geneslator contain columns “ENTREZIDOLD” and “ENSEMBLOLD”, which collect old gene identifiers of NCBI Gene and Ensembl databases. By default, these columns are queried by select() and mapIds() methods whenever a gene cannot be annotated using current identifiers. This behaviour can be controlled through the boolean parameter search.archives, whose default value is TRUE.

For example, in the following query key “3” corresponds to the old NCBI Gene identifier of gene “A2MP1”. By using archived data, select() is able to correctly map NCBI Gene ID “3” to gene symbol “A2MP1”.

# Map NCBI gene id 3 to gene symbol, using both current and old identifiers
result <- geneslator::select(org.Hsapiens.db,
  keys = "3", columns = "SYMBOL",
  keytype = "ENTREZID"
)
result
#>   ENTREZID SYMBOL
#> 1        3  PZP2P

Whenever archived identifiers are used to solve a query (as in this example), a warning message is sent to the user. If we set search.archives=FALSE, select() is unable to map the identifier to the correct symbol.

# Map NCBI gene id 3 to gene symbol, using only current identifiers
result <- geneslator::select(org.Hsapiens.db,
  keys = "3", columns = "SYMBOL",
  keytype = "ENTREZID", search.archives = FALSE
)
result
#>   ENTREZID SYMBOL
#> 1        3     NA

7.3 Orthologs mapping

In queries involving orthologs mapping, by default, select() returns all possible ortholog mappings. This behavior is controlled by parameter orthologs.mapping, whose default value is “multiple”.

# Get orthologs of yeast genes CHC1 and NMA2 in worm and fly
result <- geneslator::select(org.Hsapiens.db,
  keys = c("CHC1", "SCAMP5"),
  columns = c("ORTHOWORM", "ORTHOFLY"), keytype = "SYMBOL"
)
result
#>   SYMBOL ORTHOWORM ORTHOFLY
#> 1   CHC1     ran-3  CG33288
#> 2   CHC1     ran-3   CG7420
#> 3   CHC1     ran-3     Rcc1
#> 4 SCAMP5     scm-1    Scamp

To get only the first ortholog, set orthologs.mapping="single":

result <- geneslator::select(org.Hsapiens.db,
  keys = c("CHC1", "SCAMP5"),
  columns = c("ORTHOWORM", "ORTHOFLY"), keytype = "SYMBOL",
  orthologs.mapping = "single"
)
result
#>   SYMBOL ORTHOWORM ORTHOFLY
#> 1   CHC1     ran-3  CG33288
#> 2 SCAMP5     scm-1    Scamp

For mapIds() function, the option orthologs.mapping is absent, because the number of mapped orthologs can be directly controlled through parameter multiVals.

8 Session Information

sessionInfo()
#> R version 4.6.1 (2026-06-24)
#> Platform: x86_64-pc-linux-gnu
#> Running under: Ubuntu 24.04.4 LTS
#> 
#> Matrix products: default
#> BLAS:   /home/biocbuild/bbs-3.24-bioc/R/lib/libRblas.so 
#> LAPACK: /usr/lib/x86_64-linux-gnu/lapack/liblapack.so.3.12.0  LAPACK version 3.12.0
#> 
#> locale:
#>  [1] LC_CTYPE=en_US.UTF-8       LC_NUMERIC=C              
#>  [3] LC_TIME=en_GB              LC_COLLATE=C              
#>  [5] LC_MONETARY=en_US.UTF-8    LC_MESSAGES=en_US.UTF-8   
#>  [7] LC_PAPER=en_US.UTF-8       LC_NAME=C                 
#>  [9] LC_ADDRESS=C               LC_TELEPHONE=C            
#> [11] LC_MEASUREMENT=en_US.UTF-8 LC_IDENTIFICATION=C       
#> 
#> time zone: America/New_York
#> tzcode source: system (glibc)
#> 
#> attached base packages:
#> [1] stats4    stats     graphics  grDevices utils     datasets  methods  
#> [8] base     
#> 
#> other attached packages:
#> [1] AnnotationDbi_1.75.2 IRanges_2.47.5       S4Vectors_0.51.9    
#> [4] Biobase_2.73.2       BiocGenerics_0.59.12 generics_0.1.4      
#> [7] geneslator_0.99.14   BiocStyle_2.41.0    
#> 
#> loaded via a namespace (and not attached):
#>  [1] utf8_1.2.6           sass_0.4.10          xml2_1.6.0          
#>  [4] zen4R_0.10.6         RSQLite_3.53.3       digest_0.6.39       
#>  [7] magrittr_2.0.5       evaluate_1.0.5       bookdown_0.48       
#> [10] fastmap_1.2.0        blob_1.3.0           plyr_1.8.9          
#> [13] jsonlite_2.0.0       DBI_1.3.0            BiocManager_1.30.27 
#> [16] httr_1.4.9           purrr_1.2.2          XML_3.99-0.24       
#> [19] Biostrings_2.81.9    httr2_1.3.0          jquerylib_0.1.4     
#> [22] cli_3.6.6            rlang_1.3.0          crayon_1.5.3        
#> [25] dbplyr_2.6.0         XVector_0.53.0       bit64_4.8.6         
#> [28] withr_3.0.3          cachem_1.1.0         yaml_2.3.12         
#> [31] otel_0.2.0           BiocBaseUtils_1.15.1 tools_4.6.1         
#> [34] memoise_2.0.1        dplyr_1.2.1          filelock_1.0.3      
#> [37] curl_8.0.0           vctrs_0.7.3          R6_2.6.1            
#> [40] png_0.1-9            lifecycle_1.0.5      BiocFileCache_3.3.0 
#> [43] KEGGREST_1.53.6      Seqinfo_1.3.2        bit_4.6.0           
#> [46] pkgconfig_2.0.3      bslib_0.12.0         pillar_1.11.1       
#> [49] Rcpp_1.1.2           glue_1.8.1           xfun_0.60           
#> [52] tibble_3.3.1         tidyselect_1.2.1     keyring_1.4.1       
#> [55] knitr_1.52           htmltools_0.5.9      rmarkdown_2.32      
#> [58] compiler_4.6.1

9 References

Appendix