Phylogenetic profiling: how much input data is enough?

Details

Ressource 1Download: BIB_BCB3501714F5.P001.pdf (1461.61 [Ko])
State: Public
Version: Final published version
Serval ID
serval:BIB_BCB3501714F5
Type
Article: article from journal or magazin.
Collection
Publications
Institution
Title
Phylogenetic profiling: how much input data is enough?
Journal
Plos One
Author(s)
Skunca N., Dessimoz C.
ISSN
1932-6203 (Electronic)
ISSN-L
1932-6203
Publication state
Published
Issued date
2015
Peer-reviewed
Oui
Volume
10
Number
2
Pages
e0114701
Language
english
Notes
Publication types: Journal Article ; Research Support, Non-U.S. Gov't Publication Status: epublish
Abstract
Phylogenetic profiling is a well-established approach for predicting gene function based on patterns of gene presence and absence across species. Much of the recent developments have focused on methodological improvements, but relatively little is known about the effect of input data size on the quality of predictions. In this work, we ask: how many genomes and functional annotations need to be considered for phylogenetic profiling to be effective? Phylogenetic profiling generally benefits from an increased amount of input data. However, by decomposing this improvement in predictive accuracy in terms of the contribution of additional genomes and of additional annotations, we observed diminishing returns in adding more than ∼ 100 genomes, whereas increasing the number of annotations remained strongly beneficial throughout. We also observed that maximising phylogenetic diversity within a clade of interest improves predictive accuracy, but the effect is small compared to changes in the number of genomes under comparison. Finally, we show that these findings are supported in light of the Open World Assumption, which posits that functional annotation databases are inherently incomplete. All the tools and data used in this work are available for reuse from http://lab.dessimoz.org/14_phylprof. Scripts used to analyse the data are available on request from the authors.
Keywords
Genetic Variation, Genomics/methods, Molecular Sequence Annotation, Phylogeny
Pubmed
Open Access
Yes
Create date
17/02/2016 16:41
Last modification date
20/08/2019 15:30
Usage data