Large plant protein clusters accessible via gene families view

In Ensembl release 116 (Ensembl Genomes 63), two large protein clusters in Plants have incomplete or missing homology data due to production constraints. To address this, we set about making these clusters accessible via the gene families view. Please read on if you’d like to know more about these clusters, both associated with the Panther subfamily ‘PTHR11439_SF127’.

Recent releases of Ensembl Plants have seen the protein-trees pipeline challenged in new and interesting ways.

As previously reported in Ensembl release 113, about 40,000 genes from Triticum aestivum cultivar Kariega were assigned to the Panther subfamily ‘PTHR11439_SF127’. Confronted with an impractical delay during inference of the 261 million ancient paralogies between these genes, we excluded their paralogy relationships from further comparative analyses.

While preparing Oat pangenome protein trees for Ensembl release 116, just over 200,000 genes were assigned to the same Panther subfamily, ‘PTHR11439_SF127’. With this cluster facing impractical delay already during the alignment and subclustering stage of the Oat protein-trees pipeline, no gene tree or homologies were inferred for it.

To mitigate the omission of gene-tree/homology data, both these clusters have been stored as families, and we’ve worked to make them accessible via the Ensembl Plants gene families view.

While viewing the webpage of a relevant gene, you will be able to access its family view by clicking on the “Gene families” link in the sidebar menu of genes in the Wheat or Oat ‘PTHR11439_SF127’ clusters (e.g. T. aestivum cv. Kariega gene TraesKAR1A01G0000500 A. sativa cv. Sang gene AVESA.00010b.r2.1AG0008120).

The gene families view offers a preview of the first thousand genes. Data can be downloaded for the full set of genes in FASTA (via the “Download protein sequences” link) or tabular format (via the “Download all … genes” link).

For users with the Ensembl Compara Perl modules installed, a multiple alignment file may be generated for each gene family that has alignment data, by using the script dumpFamilyAlignment.pl .