We are pleased to announce the public availability of a second Ensembl mirror in the USA. This can be found at http://useast.ensembl.org/ . This provides a fully functional Ensembl website, but there are some things to note which I’ve listed below.

Redirection

We don’t automatically redirect users to the new mirror, although we have plans for this in future. So for now you’ll need to explicitly visit http://useast.ensembl.org/ to access it.

User logins

If you use the login functionality, your existing login will work on http://useast.ensembl.org/ , although configuration changes will not be reflected between sites. We plan to support shared logins very soon.

Other services

We don’t yet offer the Biomart or BLAST/BLAT services on the new mirror; these will come in the near future. We currently have no plans to offer a US-based MySQL mirror, so you should continue use ensembldb.ensembl.org for MySQL queries.

We’re very keen to hear your experiences with this new mirror, particularly from our US users, please use the Helpdesk in the first instance, or contact me directly. We also have advanced plans for mirrors in other parts of the world, so stay tuned!

Owing to problems with BLOB data in some tables, we have had to change the way these tables are dumped, which in turn means a change to the mysqlimport parameters.

It has been necessary to escape the table fields when dumping the MySQL text files. Hence to import successfully, a field escape parameter needs to be specified when using mysqlimport, e.g.

mysqlimport -h your_host -u write_user -p write_pass -P 3306 –fields_escaped_by=\\ database_name /full/path/to/table_name.txt

Tables affected in Release 59 are:

* homo_sapiens_funcgen_59_37d.result_feature
* mus_musculus_funcgen_59_37l.result_feature
* ensembl_compara_59.conservation_scores

and the meta_conf__xml__dm and meta_template__xml__dm tables in all mart databases.

From release 60, all Ensembl mysql dumps will be done this way, so you will need to use the above parameters on all imports.


People following the declarations of intentions for the next release (these are sent to announce@ensembl.org) may have noticed that we are releasing LASTZ pairwise alignments instead of BLASTZ ones. LASTZ is written by Bob Harris from the Penn State University as a replacement of BLASTZ, as BLASTZ is now considered obsolete (read the announcement).

This is the first release where we use LASTZ for the new alignments. We will update the previous alignments in the following releases.

Ensembl is already working on the forthcoming release (e60!). The declaration of intentions have been published and include the new Giant Panda (Ailuropoda melanoleuca) genome.

While I was looking at the alignments we are planing to release, I saw the H.sap-A.mel LASTZ-net alignments. Hey! I thought we banned Apis mellifera from Ensembl long time ago.

Then I realised that the scientific names of both the honey bee and the giant panda start with the same letters. So, don’t get confused, one of them eats flowers and makes honey, while the other one eats bamboo and makes very nice teddies!

Since release 58, the Ensembl Regulatory Build has been cell type specific. Regulatory Features are defined as sites of open chromatin which are potentially involved in gene regulation. These are built using data from different cell types, resulting in differing structures, attributes and classifications across the various cell types.

Release 59 also boasts greater coverage due to the incorporation of more data sets (see previous blog post), and a new ‘projection’ methodology. Projection allows Regulatory Features to be built on cell lines with sparse data, and also consolidates existing higher quality builds. The result of these changes is an increase in the number of Regulatory Features per cell type, as well as an improvement in the number of features which are assigned a classification e.g. Gene Associated or Promoter Associated etc.

The ‘Regulation’ panel has also been updated to reflect the new cell type specific nature of the build e.g. ENSR00000515919. The ‘Details by cell line’ view is now split into several sections, each showing data for a specific cell line. The uppermost section details a species ‘MultiCell’ cell line, showing data used to define the core regions of the given Regulatory Feature. More information on this view can be found here. We are always looking at ways to improve our Regulatory Build process, some areas we are currently considering are listed in our development road map.

Over the last 3 releases we have significantly increased the content of the functional genomics databases, by including data from large public projects such as ENCODE and The Epigenomics Roadmap. In release 59 we have over 200 human data sets representing 10 cell types, 41 histone modifications and 14 transcription factors. These numbers will steadily increase in the forthcoming releases as more data is incorporated.

These data sets are now available in ‘Region In Detail’. Cell type tracks can be turned on via the ‘Functional genomics’ menu of the configuration panel – click on ‘configure this page’ on the left to access it. These are split into ‘Core’ and ‘Other’ evidence types, reflecting how we deal with these data within the Regulatory Build. Display options include a peak track, with the underlying raw data is available as a ‘multi-wiggle’ track. Further configuration is available via the ‘Cell/Tissue’ tab, where individual feature types can be turned on or off.

Work is ongoing to improve the flexibility of these displays.

Ensembl release 59 includes the first human assembly patches released by the Genome Reference Consortium (GRC).

The goal of the GRC is to ensure that the human reference assembly is biologically relevant by closing gaps, fixing errors and representing complex variation. Their ongoing efforts are made available to the community via minor releases called patches. The patches do not change the chromosome coordinate system but do provide either a new alternate haplotype (novel patch) or provide a preview of the chromosome tiling path for that region in the next major release (fix patch).

The patched update GRCh37.p1 affects only 2 regions of the reference assembly.The patch update GRCh37.p1 includes:

Fix Patch: HG79_PATCH (GL339450.1) on chromosome 9, correction for the ABO gene. Fix patches are coloured green on the Chromosome Summary page. Click here for an example region.

Novel Patch: HSCHR5_1_CTG1 (GL339449.1) on chromosome 5. This patch provides an alternative region (haplotype). Novel patches are coloured red on the Chromosome Summary page. Click here for an example region.

The two patched regions have undergone preliminary gene annotation. Human cDNAs with their annotated ORFs were aligned to the genome using the Exonerate cdna2genome model to generate coding transcripts.

We expect future patch releases on a quarterly basis.

In the near future, we will be changing the way that the ensembldev and ensembl-announce mailing lists are managed.

The move to new hardware and list management software will increase responsiveness on the lists (ensembldev in particular has suffered from slow deliveries for some time), and in time will also allow us to provide a searchable archive of past posts.

The list addresses will change with the move:

  • ensembl-dev@ebi.ac.uk becomes dev@ensembl.org
  • ensembl-announce@ebi.ac.uk becomes announce@ensembl.org

The new addresses are active now.

All subscribers to the old lists have been moved to the new lists. Posts to the old list addresses will be automatically forwarded for a short time, but please update your email address books and any spam filters to reflect the new address list.

Information about the new lists, including details on how to subscribe and unsubscribe, can be found at http://lists.ensembl.org/.

As before, dev@ensembl.org is an open list which any subscriber can post to. Posting to announce@ensembl.org is restricted to Ensembl people.


A new section has been added to the variation page in Ensembl 58 that allows you to find other variations in strong linkage disequilibrium with the SNP you are viewing.

Clicking on “Linked variations” from the menu on the left hand side of the variation page takes you to a view like this one for rs1333049. Linkage disequilibrium values are calcuated on the fly and presented in one table for each population.

The table shows both r2 and D' values, along with the distance between the linked and current variations, any overlapping genes and any phenotypes associated with the linked variations. The table can be sorted by any of these columns by clicking on the column header (see previous post). The view is extensively configurable – clicking on Configure this page allows you to select populations to be displayed, change the distance over which linked variations are looked for, and filter the variations returned.

This view is currently only available for Ensembl Human, and is limited to variations with enough associated genotypes to calculate linkage disequilibrium values.

It is now possible to get the GERP constrained elements via the DAS protocol. For instance the DAS command to get all the GERP elements on the BRCA2 gene (Human chr 13: 32889611-32973347) is:
http://www.ensembl.org/das/Homo_sapiens.GRCh37.constrained_element/features?segment=13:32889611,32973347

By default, you obtain both the constrained derived from our 16-way amniote alignments and the 33-way placental mammals ones (these include all the low-coverage genomes). You can filter the elements you want by using the argument type:
http://www.ensembl.org/das/Homo_sapiens.GRCh37.constrained_element/features?segment=13:32889611,32973347;type=33_eutherian_mammals

Read more on DAS or on the multiple alignments and constrained elements.