Comparing and assessing the quality of whole-genome multiple alignments is a difficult task. In the protein world, many mathematical models are available. They are based on synonymous and non-synonymous substitutions and the physicochemical similarities among the aminoacids. None of this can be applied to non-coding sequences, the 99% of the human genome.

There are two main trends for whole-genome alignments. Authors have either used genomic features like ancestral repeats or have developed phylogenetic models to generate synthetic sequences for which the “real” alignment is known.

Two articles have been published recently, one proposing a new method based on artificial sequences (Kim & Sinha, BMC Bioinformatics 2010, 11:54) and the other one looking at the coverage, agreement and accuracy of the alignments in the ENCODE pilot regions (Chen & Tompa, Nature Biotechnology 2010, doi:10.1038/nbt.1637).

According to both studies, Pecan is the strongest contender, showing the clear advantage of using a consistency-based approach (see Paten et al., Genome Res. 2008, 18:1814-28) to align the sequences.

Release 57 saw the release of a 5-way EPO alignment across the telost fish – Zebrafish, Stickleback, Medaka, Tetraodon and Fugu. Just recently I’ve spent some time browsing through them. They are very interesting, with the ancestral duplication in Fish showing more complex homology relationships than in mammals. Here’s a nice, clean example

Simple Fish multiple alignments

Here is a far more complex region, when ENREDO has clearly picked up the ancestral duplication but is struggling to make it colinear across the entire region

Complex Fish multiple alignments

One thing which I don’t think most people appreciate is the incredible phylogenetic depth in the telost linage. In terms of “millions of years of evolution” or “sequence divergence” actually the deepest splits in the telosts – such as ZebraFish to Stickleback – are almost as deep as telosts to mammals – certainly deeper than birds to mammals. So this is asking alot to find good, clearly co linear stretches, in particular when you think of the draft nature of these genomes.

Chatting to Javier, it might be much better to also look at a 4-way EPO on the “Stickleback” side of the telost linage, in other words, Medaka/Stickleback/Fugu/Tetraodon. This I think will come together better (in fact most of the “nice” regions in the 5-way EPO are actually regions without ZebraFish) and we might be able to look at taking that ancestral chromosome ordering and perhaps sequence in comparisons to Zebrafish.

This week, Albert Vilella and myself participated in the Xfam consortium meeting. The meeting focussed on protein, domains and ncRNA classification, and on the new developments of the HMMER package.

Although Ensembl is not part of Xfam, we share many interests. We are getting increasingly interested in the use HMMER models, especially since the release of HMMER3.0. Also, in the forthcoming release (version 58), Ensembl will provide gene trees for ncRNAs. Most of these ncRNA genes are annotated using Rfam models.

Stay tuned for more!

In collaboration with the Neandertal Genome Project, we have created an Ensembl-style browser of the Neandertal data available at http://projects.ensembl.org/neandertal. A draft sequence of the Neandertal genome was published in the May 7 issue of Science.

The Neandertal browser includes the ability to visualise the Neandertal data using the new Resembl code developed in collaboration with Illumina. The Resembl code will be introduced in the 1000 Genomes browser later this month and in Ensembl over the summer.

Data include:
– Neandertal sequencing reads from all 6 Neandertal fossils
– Neandertal contigs/consensus from all individuals combined
– Modern human sequencing reads to put the divergence of the Neandertal genomes into perspective
– Selective sweep scan to detect positive selection in early modern humans
– A catalog of changes consisting of Neandertal alleles for positions of non-synonymous difference between human and chimpanzee

Full details of the data types and instructions for using our new display tools are available on the data information page.

Links are also provided from the Neandertal Browser home page to the raw sequence data stored at the EBI for the Neandertal genome project and the modern human genome data.

Further information about the project is available from the project page at Max Planck Institute for Evolutionary Anthropology, from the genome paper and from other companion papers in the same issue of science.

We thank Janet Kelso, Ed Green and Udo Stenzel at the MPI for assistance and Eugene Kulesha at the EBI for work to create the Neandertal browser.

This could be come a new word in English is it becomes more frequent… well whilst most of Europe was grounded (and some colleagues from Ensembl stranded throughout the world) we only had to postpone one workshop… And our tour of Australia continued (we only had domestic flights and these weren’t disrupted), coming to an end this week (I’m flying to Canberra where the ANU is hosting the last of the series).

Great feedback (thanks for filling those surveys, guys!), and our SNP Effect Predictor tool is becoming very useful, this trip also gave us a a chance to test (first hand) the speed of access to the browser and API from Australia.

This coming month, as Bert has already posted, we are attending a couple of conferences in France, if you are attending the Paris workshop (it may not be too late to register if you plan to be in town) you will have to clear security (it’s hosted at the UNESCO at the end of the day), so come early (with your passport or ID card), we are also running a workshop in Montepellier (thanks to HUGO and specially Cathy for this one).

Back to Europe soon,

ǝsoX

As part of the EBI Roadshow training programme, Ensembl teamed up with ArrayExpress to run workshops for students, postdocs, and professors at ITESM, UNAM, and CIBNOR in these bioinformatic tools. The response was very positive. Feedback from 86 participants includes comments such as:

“I am an undergraduate student, and know little about bioinformatics. In the future, I will be able to use EBI as my primary resource.”

“It is a really good opportunity to now get all these tools, to help facilitate understanding and analysis of scientific data”

“An excellent course, and very useful tools!”

Ensembl and ArrayExpress were ranked by 99% of participants as being useful to their work. Not only are people made more aware of individual projects through these workshops, EBI resources are publicised. 31% of our participants were unaware of EBI resources before the workshop, which contrasts to 95% responding that after this workshop, they would most likely use EBI resources. 88% of participants would like more training in these resources and others; specifically mentioned were ontologies, proteomics, and genome sequencing as topics to learn more about. This reflects a need for bioinformatics courses in the life sciences in that part of the world, if not in all the world.

And finally, an after-effect of the workshops was to prove that there is a lot of interest in bioinformatics. This from our host at CIBNOR:

“CIBNOR is in a growing stage, we have a project for an Innovation and Technology Park and I am trying to convince people about the need for a Bioinformatics Unit. I am sure that things like this course will help us a lot.”

We greatly enjoyed training in Mexico, because of all the keen interest, energy, and the evenings on the sand dunes. We took the course into the field, discovering a pufferfish and spine on the beach, in honor of vertebrate genomes!

Ensembl announces the release of http://ncbi36.ensembl.org. This Ensembl site is for users who still need access to the NCBI36 human assembly. It is actually a complete copy of the Ensembl 54 release which was the last Ensembl release containing NCBI36.

Although access was already possible through the Ensembl archive sites, the new ncbi36.ensembl.org site will provide better performance because it is running on separate hardware. Also ncbi36.ensembl.org provides Blast/Blat search support which the archives do not.

The main reason we have provided a dedicated site for NCBI36 is for two large projects (Encode and 1000 Genomes) which have some of their data aligned on this assembly. ncbi36.ensembl.org will only be up for as long as there is significant need for it. We will be reviewing usage in Spring 2010 and currently plan to remove the site by Summer 2010. After that time users will still be access the NCBI36 assembly via the archive sites, there just won’t be a dedicated site for it anymore.

In addition to Release 2 of Ensembl Genomes early this month, EBI-EMBL would also like to announce the new arrival of Ensembl Fungi (beta).

Release includes:

  • 2 yeast genomes: Saccharomyces cerevisiae and Schizosaccharomyces pombe.
  • 7 Aspergillus genomes A.clavatus, A.flavus, A.fumigatus, A.niger, A.oryzae, A.terreus and Neosartorya fischeri.

If you have any comments or feedback please do not hesitate to contact us at helpdesk@ensemblgenomes.org.

This year we’ve invested in our own mirror – maintained by us – on the west coast of the US. This was mainly because assessing the web return time for our users showed a consistent additional 3 to 4 seconds if you were lucky enough to live out on the west coast (worse still if you are in Australia!). Although we did alot last year to improve the general response time of our web pages (for example, compressing our CSS and Javascript down to single files for the whole site, so these are only loaded once and then cach’ed locally), the Ensembl site delivers alot of dynamic content – and nothing but getting closer to the users can help this.

You can reach the site directly at uswest.ensembl.org or alternatively there is a little “world” icon on the top right of the page which switches to the star-and-stripes when you’re on the west coast. Having the mirror not only helps our users who are on the west coast but also provides resilience when our main site goes down. As we’re responsibile for provisioning it in-sync with our main site (its part of our release process) this mirror will stay current with the main site.

In some sense the mirror should be a low cost “per user” for us having the mirror – if users go to the mirror, it means less load on the main site, and so it’s really how we distribute the “web farm” that sits behind Ensembl geographically. However, there are overheads from hiring rack space in the US to making our own release cycle more complex. This means we will need to assess whether running a US mirror makes sense in the long term. Our instinct is yes, but we need hard data on this.

These things need time to pick up, but already we’d be interested in feedback on this – for US users, is this site faster for you – in particular for East coast people who we think are probably still best off on the main site. Does it change with time of day? For Pacific rim users – Japan, Singapore, Korea, Australia – is the west coast site snappier for you? We’ll be putting in place our own monitoring schemes, but user feedback is always good…

Ensembl is pleased to announce the release of its West Coast US mirror (uswest.ensembl.org). This is a full mirror of the current Ensembl 54 release. We are providing this mirror to improve performance for users in the US, particularly on the West coast. It includes full search, BioMart and BLAST support (BLAST searching is actually run at Sanger with results passed back to the mirror).

This mirror is managed directly by the Ensembl web team, and we will aim to update it along with the main site, to keep it current. Credit for gettting this mirror up goes to James Smith and Eugene Bragin from the web team, with support from the Sanger systems team, particularly Peter Clapham, John Nicholson and Dave Holland.

Future plans: We will improve the mirror in the near future by allowing users to switch between the main and mirror site. Currently, we do not suggest logging in to the mirror. All user data must be retrieved by the main site at the Wellcome Trust Genome Campus. Speed is optimal if login is not used, however this will be improved in the future.

Many thanks,
The Ensembl Team