Ensembl archive sites allow you to browse older versions of our website (and the underlying data).  These are maintained going back three years from the live version.

Versions 50-52 of the website (2008) will be retired tomorrow (30 March).   Data and code for these releases are still available using MySQL or the Perl API.  No major genome assembly changes are included in these versions.

Archives are especially useful in citing our webpages in publications.  As Ensembl has frequent updates (every two months) it is useful to use the ‘permalink’ found at the bottom left of Ensembl views.

 

The url will follow the format of this example:

http://Feb2012.archive.ensembl.org

and will bring the reader to that specific Ensembl version, rather than the live site (at www.ensembl.org).

Cheers,

The Ensembl Team

 

 

Pre! sites have been released for two species: Southern Platyfish (Xiphophorus maculatus) and Budgerigar (Melopsittacus undulatus).

Xiphophorus maculatus

The playtfish assembly Xiphophorus maculatus 4.4.2 (GCA_000241075.1) was submitted by the The Genome Institute, Washington University at St. Louis. The Southern Platyfish is of interest for development of malignant melanoma when crossed with other members of the Xiphophorus family, and the family is used as a model for the study of hereditary cancer. This assembly is composed of 20,640 unplaced scaffolds. Click here to go to the platyfish Pre! site, where you can view platyfish protein and EST alignments, as well as alignments of the Ensembl release 65 zebrafish and stickleback translations. This assembly will undergo full automatic gene annotation in due course.

 

Melopsittacus undulatus

The budgerigar assembly Melopsittacus_undulatus_6.3 (GCA_000238935.1) was submitted by the The Genome Institute, Washington University at St. Louis. The budgerigar is important as a model species in the study of hereditary mutations which cause color variation in the feathers. This assembly is composed of 25211 unplaced scaffolds. Click here to go to the budgerigar Pre! site, where you can view budgerigar protein alignments, as well as alignments of the Ensembl release 65 human, chicken and zebra finch translations. This assembly will undergo full automatic gene annotation in due course.

The mouse assembly GRCm38 (GCA_000001635.2) was submitted by the Genome Reference Consortium (GRC). The whole assembly comprises 65 toplevel sequences: 19 autosomes, X and Y, 22 unlocalized scaffolds and 22 unplaced scaffolds. These toplevel sequences are assembled from 20,423 contigs with a N50 value of 191kb. The N50 size is the length such that 50% of the assembled genome lies in blocks of the N50 size or longer. Click here to go to the mouse Pre! site, where you can view mouse protein, cDNA and EST alignments, as well as alignments of the Ensembl release 66 mouse translations. This assembly will undergo full automatic gene annotation in due course.

Pre! sites have been released for five species: cat (Felis catus), chicken (Gallus gallus), dog (Canis lupus familiaris), squirrel monkey (Saimiri boliviensis) and thirteen-lined ground squirrel (Spermophilus tridecemlineatus).

Felis catus

The cat assembly Felis_catus-6.2 (GCA_000181335.2) was submitted by the International Cat Genome Sequencing Consortium. The domestic cat is important as a model organism for human infectious disease and for the conservation of endangered cat species. This assembly is composed of 19 chromosomes, 749 unlocalized scaffolds and 4731 unplaced scaffolds. Click here to go to the cat Pre! site, where you can view cat protein, cDNA and EST alignments, as well as alignments of the Ensembl release 66 human and cat translations. This assembly will undergo full automatic gene annotation in due course.

 

Gallus gallus

The chicken assembly Gallus_gallus-4.0 (GCA_000002315.2) was submitted by the International Chicken Genome Consortium. The chicken is important not only as a food source but also for studies in vertebrate embryology and as a model for other bird species. This assembly is composed of 33 chromosomes, 1805 unlocalized and 14093 unplaced scaffolds. Click here to go to the chicken Pre! site, where you can view chicken protein alignments, as well as alignments of the Ensembl release 65 human, chicken, turkey and zebra finch translations. This assembly will undergo full automatic gene annotation in due course.

 

Canis familiaris

The dog assembly CanFam3.1 (GCA_000002285.2) was submitted by the Dog Genome Sequencing Consortium. The dog is an important model organism for the study of human disease including cancer, heart disease and obsessive-compulsive disorder. This assembly is composed of 39 chromosomes and 3228 unplaced scaffolds. Click here to go to the dog Pre! site, where you can view dog protein, cDNA and EST alignments, as well as alignments of the Ensembl release 65 human and dog translations. Full automatic gene annotation is in progress and will include RNA-seq data.

 

Spermophilus tridecemlineatus

The squirrel assembly SpeTri2.0 (GCA_000236235.1) was submitted by the Broad Institute. The squirrel is a model for mammal hibernation. This assembly is composed of 12483 unplaced scaffolds. Click here to go to the squirrel Pre! site, where you can view alignments of the Ensembl release 66 squirrel, mouse and human translations. Full automatic gene annotation is in progress and will include RNA-seq data.

 

Saimiri boliviensis

The squirrel monkey assembly SaiBol1.0 (GCA_000235385.1) was submitted by the Broad Institute. The squirrel monkey is a model organism for infectious disease, behaviour and reproduction. This assembly is composed of 2685 unplaced scaffolds. Click here to go to the squirrel monkey Pre! site, where you can view alignments of the Ensembl release 65 human translations. This assembly will undergo automatic gene annotation in due course.

Starting from this release (65), we provide further details on the internal aspects of the GeneTrees. Our trees are built using a several phylogenetic reconstruction methods, each of them resulting in a different tree. These are combined by TreeBeST to produce final tree (Read more on the pipeline). Each node of the final tree is supported by at least one of the original trees.

The new functionality is not quite like an X-ray plate, but you can click on any node to find out about which methods support this particular node. The 5 methods we use for protein-coding gene trees are:

  • phyml_aa: maximum likelihood (ML) tree based on the protein alignment with the WAG model
  • phyml_nt: ML tree based on the codon alignment with the HKY model
  • nj_mm: neighbour-joining (NJ) tree based on the codon alignment using p-distance
  • nj_dn: NJ tree based on the codon alignment using dN distance
  • nj_dn: NJ tree based on the codon alignment using dS distance

New in this release (version 65 of Ensembl) is the ability to save your selected tracks as a configuration.  You can load in a set of tracks you chose with a simple click of the ‘Load configuration’ button.  A simple example follows.

From the Region in detail view, click on Configure this page.  Choose some tracks, for example, in the dbSNP menu: ‘Sequence variants (dbSNP and all other sources)’, in the EST alignments menu: ‘Human EST’, and in the RNA-Seq models menu: ‘blood intron’ and ‘blood’.  Now click the Save as button at the bottom left of the menu.

Name the configuration as My Tracks 1.  You can add a description as well, for example I selected SNPs, ESTs, and RNA-Seq blood models.  Click Save.

Now you have a configuration named My Tracks 1!  Even if you deselect the tracks and turn other data on, you can always go back to this selection by clicking on Configure this page and then Load configuration.

Try it on other views like Gene summary!

Clicking on Configure this page allows acess to the Manage configurations tab.  Here, you can make a configuration set.  A configuration set allows tracks across many different Ensembl views to be switched on all at once.

We hope you find this useful!  Keep your eyes out for future sharing of configuration, making it easy to send ‘your Ensembl view’ to a lab mate or colleague.

 

 

 

Ensembl 65 brought a major change to our core data model; we decided to merge the stable id tables with their parent tables. The relationship between a stable id record and its parent record was 1:1 making these tables an unnecessary step of normalisation & increasing the number of joins the API and MySQL had to perform. If you are using the Perl API this change will be transparent. However if you use direct SQL then views have been provided to replicate the stable id tables and allow your SQL to remain compatible. These views will be removed in Ensembl release 67. To support the new schema queries should be performed against the parent table for example:

  -- Original SQL
  select g.seq_region_start, g.seq_region_end
  from gene g join gene_stable_id gsi using (gene_id)
  where gsi.stable_id = 'ENSG00000139618';

  -- Should now be
  select g.seq_region_start, g.seq_region_end
  from gene g where g.stable_id = 'ENSG00000139618';

If you have any other queries about the changes then please contact helpdesk or our dev mailing list.

Pre! sites have been released for three species: Comorese coelacanth (Latimeria chalumnae), pig (Sus scrofa) and Chinese hamster (Cricetulus griseus).

 

Latimeria chalumnae

Comorese coelacanth. Picture courtesy of Robbie Cada

The Comorese coelacanth assembly LatCha1, provided by the Broad Institute, is in the process of full automatic gene annotation. This interesting fish species is a member of the lobe-finned fishes, which were thought to be extinct since the Late Cretaceous period. The first living specimen was discovered off the east shores of South Africa in 1938. The coelacanth is an important outgroup to tetrapods. Click here to go to the coelacanth Pre site, where you can view coelacanth protein and cDNA alignments, as well as alignments from the Ensembl release 64 human, stickleback and zebrafish translations.

 

Sus scrofa

Pig

The pig assembly Sscrofa10.2, provided by the Swine Genome Sequencing Consortium, is in the process of full automatic gene annotation that will include RNAseq data. The pig is important not only for pork production but also as an important model organism as it enables research into human health issues such as cardiovascular disease, obesity and immunity. Click here to go to the pig Pre site, where you can view pig protein, cDNA and EST alignments as well as alignments from the Ensembl release 64 human and pig translations.

 

Cricetulus griseus

Chinese hamster

The Chinese hamster assembly CriGro_1.0 is provided by the BGI and published here. The Chinese hamster Ovary (CHO) K1 cell line is widely used for the production of biopharmaceutical proteins. There are currently no plans to progress to full genome annotation for this assembly. Click here to go to the Chinese hamster Pre site.

 

We have recently added a new table at the top of the Orthologues view. This table shows the number of species that have a 1:1, 1:many or many:many orthology relationship with the current gene. This table shows that information for the human BRCA2 gene:

The table also contains a ‘Show details’ column. This can be used to restrict the list of orthologues shown on the page to this or these species clades.

Please, refer to the Ensembl documentation at http://www.ensembl.org if you want to know more about how we infer orthologues.

DNA methylation is an epigenetic mark known to be important in many biological processes, namely in the regulation of gene expression.

In addition to MeDIP-Chip methylation datasets for 17 cell lines and tissues from Rakyan et al (2008), Ensembl now also provides, in the form of DAS tracks, RRBS (Reduced Representation Bisulfite Sequencing) data for 44 cell lines from ENCODE. This data provides genome-scale information at single-nucleotide resolution regarding the methylation state of the genome in each cell line.

To visualize this data, when in the Location View, go to Configure this page in the left panel, and select the appropriate DNA Methylation tracks within the Regulation section of the configuration panel.

The methylation state is indicated by a color gradient which goes from dark blue, indicating highly methylated areas, through green and then towards yellow which indicates low methylation.

 

We merged replicates by a weighted average of the percentage of methylation of the replicates, but only taking positions with at least 20 reads in the combined set.

More information regarding the ENCODE datasets can be seen here. We follow the ENCODE data policies regarding data usage.