Updates to FTP site of the new Ensembl website

The FTP site for the new Ensembl website has a new structure. In the updated structure, files are now organised under GCA/ and GCF/ subdirectories, for example: https://ftp.ebi.ac.uk/pub/ensemblorganisms/GCA/. This structure is available for all releases on the site, including the latest partial release 2026-06-07. It will be the standard structure for all subsequent releases. 

The structure organised by species name at  https://ftp.ebi.ac.uk/pub/ensemblorganisms/ will be retired in August 2026. Please be aware that this change will affect workflows and pipelines pointing to this site. This structure has been maintained for partial release 2026-06-07, and it will remain for three more partial releases in July 2026. We recommend adapting pipelines or workflows to the updated structure as soon as possible.** When it is retired, data will be deleted, and URL links to these will fail. 

This update is intended to make the ever-increasing number of species assembly datasets easier to organise and maintain. To achieve this, we have updated the site structure to reflect naming conventions of the assembly accession number for each assembly. Accession identifiers are split into groups of three digits, with the version forming the final segment.

Examples of the updated structure:

1) Human assembly, annotated by Ensembl 

The human GRCh38 assembly accession is  GCA_000001405.29, the corresponding url for data related to this assembly will be:

https://ftp.ebi.ac.uk/pub/ensemblorganisms/GCA/000/001/405/29

Genome annotation data is then grouped into provider directories based on the source of genome annotation, (in this case – Ensembl), and subsequently by date of the annotation:

https://ftp.ebi.ac.uk/pub/ensemblorganisms/GCA/000/001/405/29/ensembl/2023_03

2) Wheat assembly annotated by IWGSC

The wheat IWGSC assembly accession is GCA_900519105.1, the corresponding url for data related to this assembly will be:

https://ftp.ebi.ac.uk/pub/ensemblorganisms/GCA/900/519/105/1

Genome annotation data is then grouped into provider directories based on the source of genome annotation, (in this case labelled “community” from the collaborative source – IWGSC), and subsequently by date of the annotation:

https://ftp.ebi.ac.uk/pub/ensemblorganisms/GCA/900/519/105/1/community/2018_04

This format is consistent with other file transfer sites such as those [hosted by NCBI](https://ftp.ncbi.nlm.nih.gov/genomes/all/GCA/).

Assembly identifiers e.g. GCA_000001405.29 are described in the ENA accession documentation

For the most up to date information on this FTP structure, please see the README file available in https://ftp.ebi.ac.uk/pub/ensemblorganisms/ 

To facilitate searching we have included a JSON manifest file which lists genome annotation data for each species. 

This structure is different from that of the Ensembl FTP site serving data up to version 116 (https://ftp.ebi.ac.uk/ensemblorg/), Ensembl Genomes up to version 63 (https://ftp.ebi.ac.uk/ensemblgenomes/), and the frozen Rapid Release FTP site (https://ftp.ensembl.org/pub/rapid-release/). Any automatic downloading configurations would need to be updated accordingly to use on the new Ensembl FTP site.

If you have any questions please contact the Ensembl Helpdesk

** authors note: there is a known bug from June 2026 affecting release date directories for homology and variation data. Please take note of this while adapting your pipelines.

Authors: Jorge Batista da Rocha, Daniel Poppleton, Natalie Willhoft