WorldWideScience

Sample records for accurate phylogenetic classification

  1. Accurate phylogenetic classification of DNA fragments based onsequence composition

    Energy Technology Data Exchange (ETDEWEB)

    McHardy, Alice C.; Garcia Martin, Hector; Tsirigos, Aristotelis; Hugenholtz, Philip; Rigoutsos, Isidore

    2006-05-01

    Metagenome studies have retrieved vast amounts of sequenceout of a variety of environments, leading to novel discoveries and greatinsights into the uncultured microbial world. Except for very simplecommunities, diversity makes sequence assembly and analysis a verychallenging problem. To understand the structure a 5 nd function ofmicrobial communities, a taxonomic characterization of the obtainedsequence fragments is highly desirable, yet currently limited mostly tothose sequences that contain phylogenetic marker genes. We show that forclades at the rank of domain down to genus, sequence composition allowsthe very accurate phylogenetic 10 characterization of genomic sequence.We developed a composition-based classifier, PhyloPythia, for de novophylogenetic sequence characterization and have trained it on adata setof 340 genomes. By extensive evaluation experiments we show that themethodis accurate across all taxonomic ranks considered, even forsequences that originate fromnovel organisms and are as short as 1kb.Application to two metagenome datasets 15 obtained from samples ofphosphorus-removing sludge showed that the method allows the accurateclassification at genus level of most sequence fragments from thedominant populations, while at the same time correctly characterizingeven larger parts of the samples at higher taxonomic levels.

  2. Concepts of Classification and Taxonomy. Phylogenetic Classification

    CERN Document Server

    Fraix-Burnet, Didier

    2016-01-01

    Phylogenetic approaches to classification have been heavily developed in biology by bioinformaticians. But these techniques have applications in other fields, in particular in linguistics. Their main characteristics is to search for relationships between the objects or species in study, instead of grouping them by similarity. They are thus rather well suited for any kind of evolutionary objects. For nearly fifteen years, astrocladistics has explored the use of Maximum Parsimony (or cladistics) for astronomical objects like galaxies or globular clusters. In this lesson we will learn how it works. 1 Why phylogenetic tools in astrophysics? 1.1 History of classification The need for classifying living organisms is very ancient, and the first classification system can be dated back to the Greeks. The goal was very practical since it was intended to distinguish between eatable and toxic aliments, or kind and dangerous animals. Simple resemblance was used and has been used for centuries. Basically, until the XVIIIth...

  3. Concepts of Classification and Taxonomy Phylogenetic Classification

    Science.gov (United States)

    Fraix-Burnet, D.

    2016-05-01

    Phylogenetic approaches to classification have been heavily developed in biology by bioinformaticians. But these techniques have applications in other fields, in particular in linguistics. Their main characteristics is to search for relationships between the objects or species in study, instead of grouping them by similarity. They are thus rather well suited for any kind of evolutionary objects. For nearly fifteen years, astrocladistics has explored the use of Maximum Parsimony (or cladistics) for astronomical objects like galaxies or globular clusters. In this lesson we will learn how it works.

  4. A higher-level phylogenetic classification of the Fungi

    NARCIS (Netherlands)

    Hibbett, D.S.; Binder, M.; Bischoff, J.F.; Blackwell, M.; Cannon, P.F.; Eriksson, O.E.; Huhndorf, S.; James, T.; Kirk, P.M.; Lücking, R.; Thorsten Lumbsch, H.; Lutzoni, F.; Brandon Matheny, P.; McLaughlin, D.J.; Powell, M.J.; Redhead, S.; Schoch, C.L.; Spatafora, J.W.; Stalpers, J.A.; Vilgalys, R.; Aime, M.C.; Aptroot, A.; Bauer, R.; Begerow, D.; Benny, G.L.; Castlebury, L.A.; Crous, P.W.; Dai, Y.C.; Gams, W.; Geiser, D.M.; Griffith, G.W.; Gueidan, C.; Hawksworth, D.L.; Hestmark, G.; Hosaka, K.; Humber, R.A.; Hyde, K.D.; Ironside, J.E.; Koljalg, U.; Kurtzman, C.P.; Larsson, K.H.; Lichtwardt, R.; Longcore, J.; Miadlikowska, J.; Miller, A.; Moncalvo, J.M.; Mozley-Standridge, S.; Oberwinkler, F.; Parmasto, E.; Reeb, V.; Rogers, J.D.; Roux, Le C.; Ryvarden, L.; Sampaio, J.P.; Schüssler, A.; Sugiyama, J.; Thorn, R.G.; Tibell, L.; Untereiner, W.A.; Walker, C.; Wang, Z.; Weir, A.; Weiss, M.; White, M.M.; Winka, K.; Yao, Y.J.; Zhang, N.

    2007-01-01

    A comprehensive phylogenetic classification of the kingdom Fungi is proposed, with reference to recent molecular phylogenetic analyses, and with input from diverse members of the fungal taxonomic community. The classification includes 195 taxa, down to the level of order, of which 16 are described o

  5. A higher-level phylogenetic classification of the Fungi

    NARCIS (Netherlands)

    Hibbett, David S; Binder, Manfred; Bischoff, Joseph F; Blackwell, Meredith; Cannon, Paul F; Eriksson, Ove E; Huhndorf, Sabine; James, Timothy; Kirk, Paul M; Lücking, Robert; Thorsten Lumbsch, H; Lutzoni, François; Matheny, P Brandon; McLaughlin, David J; Powell, Martha J; Redhead, Scott; Schoch, Conrad L; Spatafora, Joseph W; Stalpers, Joost A; Vilgalys, Rytas; Aime, M Catherine; Aptroot, André; Bauer, Robert; Begerow, Dominik; Benny, Gerald L; Castlebury, Lisa A; Crous, Pedro W; Dai, Yu-Cheng; Gams, Walter; Geiser, David M; Griffith, Gareth W; Gueidan, Cécile; Hawksworth, David L; Hestmark, Geir; Hosaka, Kentaro; Humber, Richard A; Hyde, Kevin D; Ironside, Joseph E; Kõljalg, Urmas; Kurtzman, Cletus P; Larsson, Karl-Henrik; Lichtwardt, Robert; Longcore, Joyce; Miadlikowska, Jolanta; Miller, Andrew; Moncalvo, Jean-Marc; Mozley-Standridge, Sharon; Oberwinkler, Franz; Parmasto, Erast; Reeb, Valérie; Rogers, Jack D; Roux, Claude; Ryvarden, Leif; Sampaio, José Paulo; Schüssler, Arthur; Sugiyama, Junta; Thorn, R Greg; Tibell, Leif; Untereiner, Wendy A; Walker, Christopher; Wang, Zheng; Weir, Alex; Weiss, Michael; White, Merlin M; Winka, Katarina; Yao, Yi-Jian; Zhang, Ning

    2007-01-01

    A comprehensive phylogenetic classification of the kingdom Fungi is proposed, with reference to recent molecular phylogenetic analyses, and with input from diverse members of the fungal taxonomic community. The classification includes 195 taxa, down to the level of order, of which 16 are described o

  6. DNACLUST: accurate and efficient clustering of phylogenetic marker genes

    Directory of Open Access Journals (Sweden)

    Liu Bo

    2011-06-01

    Full Text Available Abstract Background Clustering is a fundamental operation in the analysis of biological sequence data. New DNA sequencing technologies have dramatically increased the rate at which we can generate data, resulting in datasets that cannot be efficiently analyzed by traditional clustering methods. This is particularly true in the context of taxonomic profiling of microbial communities through direct sequencing of phylogenetic markers (e.g. 16S rRNA - the domain that motivated the work described in this paper. Many analysis approaches rely on an initial clustering step aimed at identifying sequences that belong to the same operational taxonomic unit (OTU. When defining OTUs (which have no universally accepted definition, scientists must balance a trade-off between computational efficiency and biological accuracy, as accurately estimating an environment's phylogenetic composition requires computationally-intensive analyses. We propose that efficient and mathematically well defined clustering methods can benefit existing taxonomic profiling approaches in two ways: (i the resulting clusters can be substituted for OTUs in certain applications; and (ii the clustering effectively reduces the size of the data-sets that need to be analyzed by complex phylogenetic pipelines (e.g., only one sequence per cluster needs to be provided to downstream analyses. Results To address the challenges outlined above, we developed DNACLUST, a fast clustering tool specifically designed for clustering highly-similar DNA sequences. Given a set of sequences and a sequence similarity threshold, DNACLUST creates clusters whose radius is guaranteed not to exceed the specified threshold. Underlying DNACLUST is a greedy clustering strategy that owes its performance to novel sequence alignment and k-mer based filtering algorithms. DNACLUST can also produce multiple sequence alignments for every cluster, allowing users to manually inspect clustering results, and enabling more

  7. Descriptive Statistics of the Genome: Phylogenetic Classification of Viruses.

    Science.gov (United States)

    Hernandez, Troy; Yang, Jie

    2016-10-01

    The typical process for classifying and submitting a newly sequenced virus to the NCBI database involves two steps. First, a BLAST search is performed to determine likely family candidates. That is followed by checking the candidate families with the pairwise sequence alignment tool for similar species. The submitter's judgment is then used to determine the most likely species classification. The aim of this article is to show that this process can be automated into a fast, accurate, one-step process using the proposed alignment-free method and properly implemented machine learning techniques. We present a new family of alignment-free vectorizations of the genome, the generalized vector, that maintains the speed of existing alignment-free methods while outperforming all available methods. This new alignment-free vectorization uses the frequency of genomic words (k-mers), as is done in the composition vector, and incorporates descriptive statistics of those k-mers' positional information, as inspired by the natural vector. We analyze five different characterizations of genome similarity using k-nearest neighbor classification and evaluate these on two collections of viruses totaling over 10,000 viruses. We show that our proposed method performs better than, or as well as, other methods at every level of the phylogenetic hierarchy. The data and R code is available upon request.

  8. Accurate stemming of Dutch for text classification

    NARCIS (Netherlands)

    Gaustad, T; Bouma, G; Theune, M; Nijholt, A; Hondorp, H

    2002-01-01

    This paper investigates the use of stemming for classification of Dutch (email) texts. We introduce a stemmer, which combines dictionary lookup (implemented efficiently as a finite state automaton) with a rule-based backup strategy and,how, that it outperforms the Dutch Porter stemmer in terms of ac

  9. Towards an integrated phylogenetic classification of the Tremellomycetes.

    Science.gov (United States)

    Liu, X-Z; Wang, Q-M; Göker, M; Groenewald, M; Kachalkin, A V; Lumbsch, H T; Millanes, A M; Wedin, M; Yurkov, A M; Boekhout, T; Bai, F-Y

    2015-06-01

    Families and genera assigned to Tremellomycetes have been mainly circumscribed by morphology and for the yeasts also by biochemical and physiological characteristics. This phenotype-based classification is largely in conflict with molecular phylogenetic analyses. Here a phylogenetic classification framework for the Tremellomycetes is proposed based on the results of phylogenetic analyses from a seven-genes dataset covering the majority of tremellomycetous yeasts and closely related filamentous taxa. Circumscriptions of the taxonomic units at the order, family and genus levels recognised were quantitatively assessed using the phylogenetic rank boundary optimisation (PRBO) and modified general mixed Yule coalescent (GMYC) tests. In addition, a comprehensive phylogenetic analysis on an expanded LSU rRNA (D1/D2 domains) gene sequence dataset covering as many as available teleomorphic and filamentous taxa within Tremellomycetes was performed to investigate the relationships between yeasts and filamentous taxa and to examine the stability of undersampled clades. Based on the results inferred from molecular data and morphological and physiochemical features, we propose an updated classification for the Tremellomycetes. We accept five orders, 17 families and 54 genera, including seven new families and 18 new genera. In addition, seven families and 17 genera are emended and one new species name and 185 new combinations are proposed. We propose to use the term pro tempore or pro tem. in abbreviation to indicate the species names that are temporarily maintained.

  10. Accurate reconstruction of insertion-deletion histories by statistical phylogenetics.

    Directory of Open Access Journals (Sweden)

    Oscar Westesson

    Full Text Available The Multiple Sequence Alignment (MSA is a computational abstraction that represents a partial summary either of indel history, or of structural similarity. Taking the former view (indel history, it is possible to use formal automata theory to generalize the phylogenetic likelihood framework for finite substitution models (Dayhoff's probability matrices and Felsenstein's pruning algorithm to arbitrary-length sequences. In this paper, we report results of a simulation-based benchmark of several methods for reconstruction of indel history. The methods tested include a relatively new algorithm for statistical marginalization of MSAs that sums over a stochastically-sampled ensemble of the most probable evolutionary histories. For mammalian evolutionary parameters on several different trees, the single most likely history sampled by our algorithm appears less biased than histories reconstructed by other MSA methods. The algorithm can also be used for alignment-free inference, where the MSA is explicitly summed out of the analysis. As an illustration of our method, we discuss reconstruction of the evolutionary histories of human protein-coding genes.

  11. Accurate molecular classification of cancer using simple rules

    Directory of Open Access Journals (Sweden)

    Gotoh Osamu

    2009-10-01

    Full Text Available Abstract Background One intractable problem with using microarray data analysis for cancer classification is how to reduce the extremely high-dimensionality gene feature data to remove the effects of noise. Feature selection is often used to address this problem by selecting informative genes from among thousands or tens of thousands of genes. However, most of the existing methods of microarray-based cancer classification utilize too many genes to achieve accurate classification, which often hampers the interpretability of the models. For a better understanding of the classification results, it is desirable to develop simpler rule-based models with as few marker genes as possible. Methods We screened a small number of informative single genes and gene pairs on the basis of their depended degrees proposed in rough sets. Applying the decision rules induced by the selected genes or gene pairs, we constructed cancer classifiers. We tested the efficacy of the classifiers by leave-one-out cross-validation (LOOCV of training sets and classification of independent test sets. Results We applied our methods to five cancerous gene expression datasets: leukemia (acute lymphoblastic leukemia [ALL] vs. acute myeloid leukemia [AML], lung cancer, prostate cancer, breast cancer, and leukemia (ALL vs. mixed-lineage leukemia [MLL] vs. AML. Accurate classification outcomes were obtained by utilizing just one or two genes. Some genes that correlated closely with the pathogenesis of relevant cancers were identified. In terms of both classification performance and algorithm simplicity, our approach outperformed or at least matched existing methods. Conclusion In cancerous gene expression datasets, a small number of genes, even one or two if selected correctly, is capable of achieving an ideal cancer classification effect. This finding also means that very simple rules may perform well for cancerous class prediction.

  12. Towards a phylogenetic classification of Leptothecata (Cnidaria, Hydrozoa).

    Science.gov (United States)

    Maronna, Maximiliano M; Miranda, Thaís P; Peña Cantero, Álvaro L; Barbeitos, Marcos S; Marques, Antonio C

    2016-01-29

    Leptothecata are hydrozoans whose hydranths are covered by perisarc and gonophores and whose medusae bear gonads on their radial canals. They develop complex polypoid colonies and exhibit considerable morphological variation among species with respect to growth, defensive structures and mode of development. For instance, several lineages within this order have lost the medusa stage. Depending on the author, traditional taxonomy in hydrozoans may be either polyp- or medusa-oriented. Therefore, the absence of the latter stage in some lineages may lead to very different classification schemes. Molecular data have proved useful in elucidating this taxonomic challenge. We analyzed a super matrix of new and published rRNA gene sequences (16S, 18S and 28S), employing newly proposed methods to measure branch support and improve phylogenetic signal. Our analysis recovered new clades not recognized by traditional taxonomy and corroborated some recently proposed taxa. We offer a thorough taxonomic revision of the Leptothecata, erecting new orders, suborders, infraorders and families. We also discuss the origination and diversification dynamics of the group from a macroevolutionary perspective.

  13. Redundancy-Free, Accurate Analytical Center Machine for Classification

    Institute of Scientific and Technical Information of China (English)

    ZHENGFanzi; QIUZhengding; LengYonggang; YueJianhai

    2005-01-01

    Analytical center machine (ACM) has remarkable generalization performance based on analytical center of version space and outperforms SVM. From the analysis of geometry of machine learning and principle of ACM, it is showed that some training patterns are redundant to the definition of version space. Redundant patterns push ACM classifier away from analytical center of the prime version space so that the generalization performance degrades, at the same time redundant patterns slow down the classifier and reduce the efficiency of storage. Thus, an incremental algorithm is proposed to remove redundant patterns and embed into the frame of ACM that yields a Redundancy free accurate-Analytical center machine (RFA-ACM) for classification. Experiments with Heart, Thyroid, Banana datasets demonstrate the validity of RFA-ACM.

  14. Automatic classification and accurate size measurement of blank mask defects

    Science.gov (United States)

    Bhamidipati, Samir; Paninjath, Sankaranarayanan; Pereira, Mark; Buck, Peter

    2015-07-01

    complexity of defects encountered. The variety arises due to factors such as defect nature, size, shape and composition; and the optical phenomena occurring around the defect. This paper focuses on preliminary characterization results, in terms of classification and size estimation, obtained by Calibre MDPAutoClassify tool on a variety of mask blank defects. It primarily highlights the challenges faced in achieving the results with reference to the variety of defects observed on blank mask substrates and the underlying complexities which make accurate defect size measurement an important and challenging task.

  15. Accurate mobile malware detection and classification in the cloud.

    Science.gov (United States)

    Wang, Xiaolei; Yang, Yuexiang; Zeng, Yingzhi

    2015-01-01

    As the dominator of the Smartphone operating system market, consequently android has attracted the attention of s malware authors and researcher alike. The number of types of android malware is increasing rapidly regardless of the considerable number of proposed malware analysis systems. In this paper, by taking advantages of low false-positive rate of misuse detection and the ability of anomaly detection to detect zero-day malware, we propose a novel hybrid detection system based on a new open-source framework CuckooDroid, which enables the use of Cuckoo Sandbox's features to analyze Android malware through dynamic and static analysis. Our proposed system mainly consists of two parts: anomaly detection engine performing abnormal apps detection through dynamic analysis; signature detection engine performing known malware detection and classification with the combination of static and dynamic analysis. We evaluate our system using 5560 malware samples and 6000 benign samples. Experiments show that our anomaly detection engine with dynamic analysis is capable of detecting zero-day malware with a low false negative rate (1.16 %) and acceptable false positive rate (1.30 %); it is worth noting that our signature detection engine with hybrid analysis can accurately classify malware samples with an average positive rate 98.94 %. Considering the intensive computing resources required by the static and dynamic analysis, our proposed detection system should be deployed off-device, such as in the Cloud. The app store markets and the ordinary users can access our detection system for malware detection through cloud service.

  16. How accurate are the European Union's classifications of chemical substances.

    Science.gov (United States)

    Rudén, Christina; Hansson, Sven Ove

    2003-09-30

    The European Commission has decided on harmonized classifications for a large number of individual chemicals according to its own directive for classification and labeling of dangerous substances. We have compared the harmonized classifications for acute oral toxicity to the acute oral toxicity data available in the RTECS database. Of the 992 substances eligible for this comparison, 15% were assigned a too low danger class and 8% a too high danger class according to the RTECS data. Due to insufficient transparency-scientific documentations of the classification decisions are not available-the causes of this discrepancy can only be hypothesized. We propose that the scientific motivations of future classifications be published and that the apparent over- and underclassifications in the present system be either explained or rectified, according to what are the facts in the matter.

  17. Short interspersed elements (SINEs) in plants: origin, classification, and use as phylogenetic markers.

    Science.gov (United States)

    Deragon, Jean-Marc; Zhang, Xiaoyu

    2006-12-01

    Short interspersed elements (SINEs) are a class of dispersed mobile sequences that use RNA as an intermediate in an amplification process called retroposition. The presence-absence of a SINE at a given locus has been used as a meaningful classification criterion to evaluate phylogenetic relations among species. We review here recent developments in the characterisation of plant SINEs and their use as molecular makers to retrace phylogenetic relations among wild and cultivated Oryza and Brassica species. In Brassicaceae, further use of SINE markers is limited by our partial knowledge of endogenous SINE families (their origin and evolution histories) and by the absence of a clear classification. To solve this problem, phylogenetic relations among all known Brassicaceae SINEs were analyzed and a new classification, grouping SINEs in 15 different families, is proposed. The relative age and size of each Brassicaceae SINE family was evaluated and new phylogenetically supported subfamilies were described. We also present evidence suggesting that new potentially active SINEs recently emerged in Brassica oleracea from the shuffling of preexisting SINE portions. Finally, the comparative evolution history of SINE families present in Arabidopsis thaliana and Brassica oleracea revealed that SINEs were in general more active in the Brassica lineage. The importance of these new data for the use of Brassicaceae SINEs as molecular markers in future applications is discussed.

  18. HIPPI: highly accurate protein family classification with ensembles of HMMs

    Directory of Open Access Journals (Sweden)

    Nam-phuong Nguyen

    2016-11-01

    Full Text Available Abstract Background Given a new biological sequence, detecting membership in a known family is a basic step in many bioinformatics analyses, with applications to protein structure and function prediction and metagenomic taxon identification and abundance profiling, among others. Yet family identification of sequences that are distantly related to sequences in public databases or that are fragmentary remains one of the more difficult analytical problems in bioinformatics. Results We present a new technique for family identification called HIPPI (Hierarchical Profile Hidden Markov Models for Protein family Identification. HIPPI uses a novel technique to represent a multiple sequence alignment for a given protein family or superfamily by an ensemble of profile hidden Markov models computed using HMMER. An evaluation of HIPPI on the Pfam database shows that HIPPI has better overall precision and recall than blastp, HMMER, and pipelines based on HHsearch, and maintains good accuracy even for fragmentary query sequences and for protein families with low average pairwise sequence identity, both conditions where other methods degrade in accuracy. Conclusion HIPPI provides accurate protein family identification and is robust to difficult model conditions. Our results, combined with observations from previous studies, show that ensembles of profile Hidden Markov models can better represent multiple sequence alignments than a single profile Hidden Markov model, and thus can improve downstream analyses for various bioinformatic tasks. Further research is needed to determine the best practices for building the ensemble of profile Hidden Markov models. HIPPI is available on GitHub at https://github.com/smirarab/sepp .

  19. The phagotrophic origin of eukaryotes and phylogenetic classification of Protozoa.

    Science.gov (United States)

    Cavalier-Smith, T

    2002-03-01

    ancestrally biciliate clade, named 'bikonts'. The apparently conflicting rRNA and protein trees can be reconciled with each other and this ultrastructural interpretation if long-branch distortions, some mechanistically explicable, are allowed for. Bikonts comprise two groups: corticoflagellates, with a younger anterior cilium, no centrosomal cone and ancestrally a semi-rigid cell cortex with a microtubular band on either side of the posterior mature centriole; and Rhizaria [a new infrakingdom comprising Cercozoa (now including Ascetosporea classis nov.), Retaria phylum nov., Heliozoa and Apusozoa phylum nov.], having a centrosomal cone or radiating microtubules and two microtubular roots and a soft surface, frequently with reticulopodia. Corticoflagellates comprise photokaryotes (Plantae and chromalveolates, both ancestrally with cortical alveoli) and Excavata (a new protozoan infrakingdom comprising Loukozoa, Discicristata and Archezoa, ancestrally with three microtubular roots). All basal eukaryotic radiations were of mitochondrial aerobes; hydrogenosomes evolved polyphyletically from mitochondria long afterwards, the persistence of their double envelope long after their genomes disappeared being a striking instance of membrane heredity. I discuss the relationship between the 13 protozoan phyla recognized here and revise higher protozoan classification by updating as subkingdoms Lankester's 1878 division of Protozoa into Corticata (Excavata, Alveolata; with prominent cortical microtubules and ancestrally localized cytostome--the Parabasalia probably secondarily internalized the cytoskeleton) and Gymnomyxa [infrakingdoms Sarcomastigota (Choanozoa, Amoebozoa) and Rhizaria; both ancestrally with a non-cortical cytoskeleton of radiating singlet microtubules and a relatively soft cell surface with diffused feeding]. As the eukaryote root almost certainly lies within Gymnomyxa, probably among the Sarcomastigota, Corticata are derived. Following the single symbiogenetic origin of

  20. Accurate Classification of RNA Structures Using Topological Fingerprints

    Science.gov (United States)

    Li, Kejie; Gribskov, Michael

    2016-01-01

    While RNAs are well known to possess complex structures, functionally similar RNAs often have little sequence similarity. While the exact size and spacing of base-paired regions vary, functionally similar RNAs have pronounced similarity in the arrangement, or topology, of base-paired stems. Furthermore, predicted RNA structures often lack pseudoknots (a crucial aspect of biological activity), and are only partially correct, or incomplete. A topological approach addresses all of these difficulties. In this work we describe each RNA structure as a graph that can be converted to a topological spectrum (RNA fingerprint). The set of subgraphs in an RNA structure, its RNA fingerprint, can be compared with the fingerprints of other RNA structures to identify and correctly classify functionally related RNAs. Topologically similar RNAs can be identified even when a large fraction, up to 30%, of the stems are omitted, indicating that highly accurate structures are not necessary. We investigate the performance of the RNA fingerprint approach on a set of eight highly curated RNA families, with diverse sizes and functions, containing pseudoknots, and with little sequence similarity–an especially difficult test set. In spite of the difficult test set, the RNA fingerprint approach is very successful (ROC AUC > 0.95). Due to the inclusion of pseudoknots, the RNA fingerprint approach both covers a wider range of possible structures than methods based only on secondary structure, and its tolerance for incomplete structures suggests that it can be applied even to predicted structures. Source code is freely available at https://github.rcac.purdue.edu/mgribsko/XIOS_RNA_fingerprint. PMID:27755571

  1. Phylogeny and phylogenetic classification of the antbirds, ovenbirds, woodcreepers, and allies (Aves: Passeriformes: Infraorder Furnariides)

    Science.gov (United States)

    Moyle, R.G.; Chesser, R.T.; Brumfield, R.T.; Tello, J.G.; Marchese, D.J.; Cracraft, J.

    2009-01-01

    The infraorder Furnariides is a diverse group of suboscine passerine birds comprising a substantial component of the Neotropical avifauna. The included species encompass a broad array of morphologies and behaviours, making them appealing for evolutionary studies, but the size of the group (ca. 600 species) has limited well-sampled higher-level phylogenetic studies. Using DNA sequence data from the nuclear RAG-1 and RAG-2 exons, we undertook a phylogenetic analysis of the Furnariides sampling 124 (more than 88%) of the genera. Basal relationships among family-level taxa differed depending on phylogenetic method, but all topologies had little nodal support, mirroring the results from earlier studies in which discerning relationships at the base of the radiation was also difficult. In contrast, branch support for family-rank taxa and for many relationships within those clades was generally high. Our results support the Melanopareidae and Grallariidae as distinct from the Rhinocryptidae and Formicariidae, respectively. Within the Furnariides our data contradict some recent phylogenetic hypotheses and suggest that further study is needed to resolve these discrepancies. Of the few genera represented by multiple species, several were not monophyletic, indicating that additional systematic work remains within furnariine families and must include dense taxon sampling. We use this study as a basis for proposing a new phylogenetic classification for the group and in the process erect new family-group names for clades having high branch support across methods. ?? 2009 The Willi Hennig Society.

  2. A bootstrap based analysis pipeline for efficient classification of phylogenetically related animal miRNAs

    Directory of Open Access Journals (Sweden)

    Gu Xun

    2007-03-01

    Full Text Available Abstract Background Phylogenetically related miRNAs (miRNA families convey important information of the function and evolution of miRNAs. Due to the special sequence features of miRNAs, pair-wise sequence identity between miRNA precursors alone is often inadequate for unequivocally judging the phylogenetic relationships between miRNAs. Most of the current methods for miRNA classification rely heavily on manual inspection and lack measurements of the reliability of the results. Results In this study, we designed an analysis pipeline (the Phylogeny-Bootstrap-Cluster (PBC pipeline to identify miRNA families based on branch stability in the bootstrap trees derived from overlapping genome-wide miRNA sequence sets. We tested the PBC analysis pipeline with the miRNAs from six animal species, H. sapiens, M. musculus, G. gallus, D. rerio, D. melanogaster, and C. elegans. The resulting classification was compared with the miRNA families defined in miRBase. The two classifications were largely consistent. Conclusion The PBC analysis pipeline is an efficient method for classifying large numbers of heterogeneous miRNA sequences. It requires minimum human involvement and provides measurements of the reliability of the classification results.

  3. Using ESTs for phylogenomics: Can one accurately infer a phylogenetic tree from a gappy alignment?

    Directory of Open Access Journals (Sweden)

    Hartmann Stefanie

    2008-03-01

    Full Text Available Abstract Background While full genome sequences are still only available for a handful of taxa, large collections of partial gene sequences are available for many more. The alignment of partial gene sequences results in a multiple sequence alignment containing large gaps that are arranged in a staggered pattern. The consequences of this pattern of missing data on the accuracy of phylogenetic analysis are not well understood. We conducted a simulation study to determine the accuracy of phylogenetic trees obtained from gappy alignments using three commonly used phylogenetic reconstruction methods (Neighbor Joining, Maximum Parsimony, and Maximum Likelihood and studied ways to improve the accuracy of trees obtained from such datasets. Results We found that the pattern of gappiness in multiple sequence alignments derived from partial gene sequences substantially compromised phylogenetic accuracy even in the absence of alignment error. The decline in accuracy was beyond what would be expected based on the amount of missing data. The decline was particularly dramatic for Neighbor Joining and Maximum Parsimony, where the majority of gappy alignments contained 25% to 40% incorrect quartets. To improve the accuracy of the trees obtained from a gappy multiple sequence alignment, we examined two approaches. In the first approach, alignment masking, potentially problematic columns and input sequences are excluded from from the dataset. Even in the absence of alignment error, masking improved phylogenetic accuracy up to 100-fold. However, masking retained, on average, only 83% of the input sequences. In the second approach, alignment subdivision, the missing data is statistically modelled in order to retain as many sequences as possible in the phylogenetic analysis. Subdivision resulted in more modest improvements to alignment accuracy, but succeeded in including almost all of the input sequences. Conclusion These results demonstrate that partial gene

  4. Molecular phylogenetic perspectives for character classification and convergence: Framing some issues with nematode vulval appendages and telotylenchid tail termini

    Science.gov (United States)

    Characters flagged as convergent based on newer molecular phylogenetic trees inform both practical identification and more esoteric classification. Nematode morphological characters such as lateral lines, bullae and laciniae are quite independent structures from those similarly named in other organi...

  5. Classification of Phylogenetic Profiles for Protein Function Prediction: An SVM Approach

    Science.gov (United States)

    Kotaru, Appala Raju; Joshi, Ramesh C.

    Predicting the function of an uncharacterized protein is a major challenge in post-genomic era due to problems complexity and scale. Having knowledge of protein function is a crucial link in the development of new drugs, better crops, and even the development of biochemicals such as biofuels. Recently numerous high-throughput experimental procedures have been invented to investigate the mechanisms leading to the accomplishment of a protein’s function and Phylogenetic profile is one of them. Phylogenetic profile is a way of representing a protein which encodes evolutionary history of proteins. In this paper we proposed a method for classification of phylogenetic profiles using supervised machine learning method, support vector machine classification along with radial basis function as kernel for identifying functionally linked proteins. We experimentally evaluated the performance of the classifier with the linear kernel, polynomial kernel and compared the results with the existing tree kernel. In our study we have used proteins of the budding yeast saccharomyces cerevisiae genome. We generated the phylogenetic profiles of 2465 yeast genes and for our study we used the functional annotations that are available in the MIPS database. Our experiments show that the performance of the radial basis kernel is similar to polynomial kernel is some functional classes together are better than linear, tree kernel and over all radial basis kernel outperformed the polynomial kernel, linear kernel and tree kernel. In analyzing these results we show that it will be feasible to make use of SVM classifier with radial basis function as kernel to predict the gene functionality using phylogenetic profiles.

  6. Multigene phylogenetic analysis redefines dung beetles relationships and classification (Coleoptera: Scarabaeidae: Scarabaeinae).

    Science.gov (United States)

    Tarasov, Sergei; Dimitrov, Dimitar

    2016-11-29

    Dung beetles (subfamily Scarabaeinae) are popular model organisms in ecology and developmental biology, and for the last two decades they have experienced a systematics renaissance with the adoption of modern phylogenetic approaches. Within this period 16 key phylogenies and numerous additional studies with limited scope have been published, but higher-level relationships of this pivotal group of beetles remain contentious and current classifications contain many unnatural groupings. The present study provides a robust phylogenetic framework and a revised classification of dung beetles. We assembled the so far largest molecular dataset for dung beetles using sequences of 8 gene regions and 547 terminals including the outgroup taxa. This dataset was analyzed using Bayesian, maximum likelihood and parsimony approaches. In order to test the sensitivity of results to different analytical treatments, we evaluated alternative partitioning schemes based on secondary structure, domains and codon position. We assessed substitution models adequacy using Bayesian framework and used these results to exclude partitions where substitution models did not adequately depict the processes that generated the data. We show that exclusion of partitions that failed the model adequacy evaluation has a potential to improve phylogenetic inference, but efficient implementation of this approach on large datasets is problematic and awaits development of new computationally advanced software. In the class Insecta it is uncommon for the results of molecular phylogenetic analysis to lead to substantial changes in classification. However, the results presented here are congruent with recent morphological studies and support the largest change in dung beetle systematics for the last 50 years. Here we propose the revision of the concepts for the tribes Deltochilini (Canthonini), Dichotomiini and Coprini; additionally, we redefine the tribe Sisyphini. We provide and illustrate synapomorphies and

  7. Accurate and interpretable classification of microspectroscopy pixels using artificial neural networks.

    Science.gov (United States)

    Manescu, Petru; Jong Lee, Young; Camp, Charles; Cicerone, Marcus; Brady, Mary; Bajcsy, Peter

    2017-04-01

    This paper addresses the problem of classifying materials from microspectroscopy at a pixel level. The challenges lie in identifying discriminatory spectral features and obtaining accurate and interpretable models relating spectra and class labels. We approach the problem by designing a supervised classifier from a tandem of Artificial Neural Network (ANN) models that identify relevant features in raw spectra and achieve high classification accuracy. The tandem of ANN models is meshed with classification rule extraction methods to lower the model complexity and to achieve interpretability of the resulting model. The contribution of the work is in designing each ANN model based on the microspectroscopy hypothesis about a discriminatory feature of a certain target class being composed of a linear combination of spectra. The novelty lies in meshing ANN and decision rule models into a tandem configuration to achieve accurate and interpretable classification results. The proposed method was evaluated using a set of broadband coherent anti-Stokes Raman scattering (BCARS) microscopy cell images (600 000  pixel-level spectra) and a reference four-class rule-based model previously created by biochemical experts. The generated classification rule-based model was on average 85% accurate measured by the DICE pixel label similarity metric, and on average 96% similar to the reference rules measured by the vector cosine metric.

  8. Accurate crop classification using hierarchical genetic fuzzy rule-based systems

    Science.gov (United States)

    Topaloglou, Charalampos A.; Mylonas, Stelios K.; Stavrakoudis, Dimitris G.; Mastorocostas, Paris A.; Theocharis, John B.

    2014-10-01

    This paper investigates the effectiveness of an advanced classification system for accurate crop classification using very high resolution (VHR) satellite imagery. Specifically, a recently proposed genetic fuzzy rule-based classification system (GFRBCS) is employed, namely, the Hierarchical Rule-based Linguistic Classifier (HiRLiC). HiRLiC's model comprises a small set of simple IF-THEN fuzzy rules, easily interpretable by humans. One of its most important attributes is that its learning algorithm requires minimum user interaction, since the most important learning parameters affecting the classification accuracy are determined by the learning algorithm automatically. HiRLiC is applied in a challenging crop classification task, using a SPOT5 satellite image over an intensively cultivated area in a lake-wetland ecosystem in northern Greece. A rich set of higher-order spectral and textural features is derived from the initial bands of the (pan-sharpened) image, resulting in an input space comprising 119 features. The experimental analysis proves that HiRLiC compares favorably to other interpretable classifiers of the literature, both in terms of structural complexity and classification accuracy. Its testing accuracy was very close to that obtained by complex state-of-the-art classification systems, such as the support vector machines (SVM) and random forest (RF) classifiers. Nevertheless, visual inspection of the derived classification maps shows that HiRLiC is characterized by higher generalization properties, providing more homogeneous classifications that the competitors. Moreover, the runtime requirements for producing the thematic map was orders of magnitude lower than the respective for the competitors.

  9. Molecular phylogenetic evaluation of classification and scenarios of character evolution in calcareous sponges (Porifera, Class Calcarea.

    Directory of Open Access Journals (Sweden)

    Oliver Voigt

    Full Text Available Calcareous sponges (Phylum Porifera, Class Calcarea are known to be taxonomically difficult. Previous molecular studies have revealed many discrepancies between classically recognized taxa and the observed relationships at the order, family and genus levels; these inconsistencies question underlying hypotheses regarding the evolution of certain morphological characters. Therefore, we extended the available taxa and character set by sequencing the complete small subunit (SSU rDNA and the almost complete large subunit (LSU rDNA of additional key species and complemented this dataset by substantially increasing the length of available LSU sequences. Phylogenetic analyses provided new hypotheses about the relationships of Calcarea and about the evolution of certain morphological characters. We tested our phylogeny against competing phylogenetic hypotheses presented by previous classification systems. Our data reject the current order-level classification by again finding non-monophyletic Leucosolenida, Clathrinida and Murrayonida. In the subclass Calcinea, we recovered a clade that includes all species with a cortex, which is largely consistent with the previously proposed order Leucettida. Other orders that had been rejected in the current system were not found, but could not be rejected in our tests either. We found several additional families and genera polyphyletic: the families Leucascidae and Leucaltidae and the genus Leucetta in Calcinea, and in Calcaronea the family Amphoriscidae and the genus Ute. Our phylogeny also provided support for the vaguely suspected close relationship of several members of Grantiidae with giantortical diactines to members of Heteropiidae. Similarly, our analyses revealed several unexpected affinities, such as a sister group relationship between Leucettusa (Leucaltidae and Leucettidae and between Leucascandra (Jenkinidae and Sycon carteri (Sycettidae. According to our results, the taxonomy of Calcarea is in

  10. Molecular Phylogenetic Evaluation of Classification and Scenarios of Character Evolution in Calcareous Sponges (Porifera, Class Calcarea)

    Science.gov (United States)

    Voigt, Oliver; Wülfing, Eilika; Wörheide, Gert

    2012-01-01

    Calcareous sponges (Phylum Porifera, Class Calcarea) are known to be taxonomically difficult. Previous molecular studies have revealed many discrepancies between classically recognized taxa and the observed relationships at the order, family and genus levels; these inconsistencies question underlying hypotheses regarding the evolution of certain morphological characters. Therefore, we extended the available taxa and character set by sequencing the complete small subunit (SSU) rDNA and the almost complete large subunit (LSU) rDNA of additional key species and complemented this dataset by substantially increasing the length of available LSU sequences. Phylogenetic analyses provided new hypotheses about the relationships of Calcarea and about the evolution of certain morphological characters. We tested our phylogeny against competing phylogenetic hypotheses presented by previous classification systems. Our data reject the current order-level classification by again finding non-monophyletic Leucosolenida, Clathrinida and Murrayonida. In the subclass Calcinea, we recovered a clade that includes all species with a cortex, which is largely consistent with the previously proposed order Leucettida. Other orders that had been rejected in the current system were not found, but could not be rejected in our tests either. We found several additional families and genera polyphyletic: the families Leucascidae and Leucaltidae and the genus Leucetta in Calcinea, and in Calcaronea the family Amphoriscidae and the genus Ute. Our phylogeny also provided support for the vaguely suspected close relationship of several members of Grantiidae with giantortical diactines to members of Heteropiidae. Similarly, our analyses revealed several unexpected affinities, such as a sister group relationship between Leucettusa (Leucaltidae) and Leucettidae and between Leucascandra (Jenkinidae) and Sycon carteri (Sycettidae). According to our results, the taxonomy of Calcarea is in desperate need of a

  11. Molecular phylogenetic evaluation of classification and scenarios of character evolution in calcareous sponges (Porifera, Class Calcarea).

    Science.gov (United States)

    Voigt, Oliver; Wülfing, Eilika; Wörheide, Gert

    2012-01-01

    Calcareous sponges (Phylum Porifera, Class Calcarea) are known to be taxonomically difficult. Previous molecular studies have revealed many discrepancies between classically recognized taxa and the observed relationships at the order, family and genus levels; these inconsistencies question underlying hypotheses regarding the evolution of certain morphological characters. Therefore, we extended the available taxa and character set by sequencing the complete small subunit (SSU) rDNA and the almost complete large subunit (LSU) rDNA of additional key species and complemented this dataset by substantially increasing the length of available LSU sequences. Phylogenetic analyses provided new hypotheses about the relationships of Calcarea and about the evolution of certain morphological characters. We tested our phylogeny against competing phylogenetic hypotheses presented by previous classification systems. Our data reject the current order-level classification by again finding non-monophyletic Leucosolenida, Clathrinida and Murrayonida. In the subclass Calcinea, we recovered a clade that includes all species with a cortex, which is largely consistent with the previously proposed order Leucettida. Other orders that had been rejected in the current system were not found, but could not be rejected in our tests either. We found several additional families and genera polyphyletic: the families Leucascidae and Leucaltidae and the genus Leucetta in Calcinea, and in Calcaronea the family Amphoriscidae and the genus Ute. Our phylogeny also provided support for the vaguely suspected close relationship of several members of Grantiidae with giantortical diactines to members of Heteropiidae. Similarly, our analyses revealed several unexpected affinities, such as a sister group relationship between Leucettusa (Leucaltidae) and Leucettidae and between Leucascandra (Jenkinidae) and Sycon carteri (Sycettidae). According to our results, the taxonomy of Calcarea is in desperate need of a

  12. Phylogenetic Classification Of Bartonella Species By Comparing The Two-Component System Response Regulator Feup Sequences

    Directory of Open Access Journals (Sweden)

    Mhamad Abou-Hamdan

    2015-08-01

    Full Text Available Abstract The bacterial genus Bartonella is classified in the alpha-2 Proteobacteria on the basis of 16S rDNA sequence comparison. The Bartonella two-component system feuPQ is found in nearly all bacterial species. We investigated the usefulness of the response regulator feuP gene sequence in the classification of 18 well characterized Bartonella species. Phylogenetic relationships were inferred using parsimony neighbour-joining and maximum-likelihood methods. Reliable classifications of most of the studied species were obtained. Bartonella were divided into two supported clades containing two supported clusters each. These results were similar to our previous data obtained with groEL ftsZ and ribC genes sequences. The wide range of feuP DNA sequence similarity 78.6 to 96.5 among Bartonella species makes it a promising candidate for multi-locus sequence typing MLST of clinical isolates. This is the first report proving the usefulness of feuP sequences in bartonellae classification at the species level.

  13. Rapid phylogenetic and functional classification of short genomic fragments with signature peptides

    Directory of Open Access Journals (Sweden)

    Berendzen Joel

    2012-08-01

    Full Text Available Abstract Background Classification is difficult for shotgun metagenomics data from environments such as soils, where the diversity of sequences is high and where reference sequences from close relatives may not exist. Approaches based on sequence-similarity scores must deal with the confounding effects that inheritance and functional pressures exert on the relation between scores and phylogenetic distance, while approaches based on sequence alignment and tree-building are typically limited to a small fraction of gene families. We describe an approach based on finding one or more exact matches between a read and a precomputed set of peptide 10-mers. Results At even the largest phylogenetic distances, thousands of 10-mer peptide exact matches can be found between pairs of bacterial genomes. Genes that share one or more peptide 10-mers typically have high reciprocal BLAST scores. Among a set of 403 representative bacterial genomes, some 20 million 10-mer peptides were found to be shared. We assign each of these peptides as a signature of a particular node in a phylogenetic reference tree based on the RNA polymerase genes. We classify the phylogeny of a genomic fragment (e.g., read at the most specific node on the reference tree that is consistent with the phylogeny of observed signature peptides it contains. Using both synthetic data from four newly-sequenced soil-bacterium genomes and ten real soil metagenomics data sets, we demonstrate a sensitivity and specificity comparable to that of the MEGAN metagenomics analysis package using BLASTX against the NR database. Phylogenetic and functional similarity metrics applied to real metagenomics data indicates a signal-to-noise ratio of approximately 400 for distinguishing among environments. Our method assigns ~6.6 Gbp/hr on a single CPU, compared with 25 kbp/hr for methods based on BLASTX against the NR database. Conclusions Classification by exact matching against a precomputed list of signature

  14. HMM-FRAME: accurate protein domain classification for metagenomic sequences containing frameshift errors

    Directory of Open Access Journals (Sweden)

    Sun Yanni

    2011-05-01

    Full Text Available Abstract Background Protein domain classification is an important step in metagenomic annotation. The state-of-the-art method for protein domain classification is profile HMM-based alignment. However, the relatively high rates of insertions and deletions in homopolymer regions of pyrosequencing reads create frameshifts, causing conventional profile HMM alignment tools to generate alignments with marginal scores. This makes error-containing gene fragments unclassifiable with conventional tools. Thus, there is a need for an accurate domain classification tool that can detect and correct sequencing errors. Results We introduce HMM-FRAME, a protein domain classification tool based on an augmented Viterbi algorithm that can incorporate error models from different sequencing platforms. HMM-FRAME corrects sequencing errors and classifies putative gene fragments into domain families. It achieved high error detection sensitivity and specificity in a data set with annotated errors. We applied HMM-FRAME in Targeted Metagenomics and a published metagenomic data set. The results showed that our tool can correct frameshifts in error-containing sequences, generate much longer alignments with significantly smaller E-values, and classify more sequences into their native families. Conclusions HMM-FRAME provides a complementary protein domain classification tool to conventional profile HMM-based methods for data sets containing frameshifts. Its current implementation is best used for small-scale metagenomic data sets. The source code of HMM-FRAME can be downloaded at http://www.cse.msu.edu/~zhangy72/hmmframe/ and at https://sourceforge.net/projects/hmm-frame/.

  15. Accurate Classification of Protein Subcellular Localization from High-Throughput Microscopy Images Using Deep Learning

    Directory of Open Access Journals (Sweden)

    Tanel Pärnamaa

    2017-05-01

    Full Text Available High-throughput microscopy of many single cells generates high-dimensional data that are far from straightforward to analyze. One important problem is automatically detecting the cellular compartment where a fluorescently-tagged protein resides, a task relatively simple for an experienced human, but difficult to automate on a computer. Here, we train an 11-layer neural network on data from mapping thousands of yeast proteins, achieving per cell localization classification accuracy of 91%, and per protein accuracy of 99% on held-out images. We confirm that low-level network features correspond to basic image characteristics, while deeper layers separate localization classes. Using this network as a feature calculator, we train standard classifiers that assign proteins to previously unseen compartments after observing only a small number of training examples. Our results are the most accurate subcellular localization classifications to date, and demonstrate the usefulness of deep learning for high-throughput microscopy.

  16. Phylogenetic classification at generic level in the absence of distinct phylogenetic patterns of phenotypical variation: a case study in graphidaceae (ascomycota.

    Directory of Open Access Journals (Sweden)

    Sittiporn Parnmen

    Full Text Available Molecular phylogenies often reveal that taxa circumscribed by phenotypical characters are not monophyletic. While re-examination of phenotypical characters often identifies the presence of characters characterizing clades, there is a growing number of studies that fail to identify diagnostic characters, especially in organismal groups lacking complex morphologies. Taxonomists then can either merge the groups or split taxa into smaller entities. Due to the nature of binomial nomenclature, this decision is of special importance at the generic level. Here we propose a new approach to choose among classification alternatives using a combination of morphology-based phylogenetic binning and a multiresponse permutation procedure to test for morphological differences among clades. We illustrate the use of this method in the tribe Thelotremateae focusing on the genus Chapsa, a group of lichenized fungi in which our phylogenetic estimate is in conflict with traditional classification and the morphological and chemical characters do not show a clear phylogenetic pattern. We generated 75 new DNA sequences of mitochondrial SSU rDNA, nuclear LSU rDNA and the protein-coding RPB2. This data set was used to infer phylogenetic estimates using maximum likelihood and Bayesian approaches. The genus Chapsa was found to be polyphyletic, forming four well-supported clades, three of which clustering into one unsupported clade, and the other, supported clade forming two supported subclades. While these clades cannot be readily separated morphologically, the combined binning/multiresponse permutation procedure showed that accepting the four clades as different genera each reflects the phenotypical pattern significantly better than accepting two genera (or five genera if splitting the first clade. Another species within the Thelotremateae, Thelotrema petractoides, a unique taxon with carbonized excipulum resembling Schizotrema, was shown to fall outside Thelotrema

  17. Accurate and reliable cancer classification based on probabilistic inference of pathway activity.

    Directory of Open Access Journals (Sweden)

    Junjie Su

    Full Text Available With the advent of high-throughput technologies for measuring genome-wide expression profiles, a large number of methods have been proposed for discovering diagnostic markers that can accurately discriminate between different classes of a disease. However, factors such as the small sample size of typical clinical data, the inherent noise in high-throughput measurements, and the heterogeneity across different samples, often make it difficult to find reliable gene markers. To overcome this problem, several studies have proposed the use of pathway-based markers, instead of individual gene markers, for building the classifier. Given a set of known pathways, these methods estimate the activity level of each pathway by summarizing the expression values of its member genes, and use the pathway activities for classification. It has been shown that pathway-based classifiers typically yield more reliable results compared to traditional gene-based classifiers. In this paper, we propose a new classification method based on probabilistic inference of pathway activities. For a given sample, we compute the log-likelihood ratio between different disease phenotypes based on the expression level of each gene. The activity of a given pathway is then inferred by combining the log-likelihood ratios of the constituent genes. We apply the proposed method to the classification of breast cancer metastasis, and show that it achieves higher accuracy and identifies more reproducible pathway markers compared to several existing pathway activity inference methods.

  18. Towards a formal genealogical classification of the Lezgian languages (North Caucasus): testing various phylogenetic methods on lexical data.

    Science.gov (United States)

    Kassian, Alexei

    2015-01-01

    A lexicostatistical classification is proposed for 20 languages and dialects of the Lezgian group of the North Caucasian family, based on meticulously compiled 110-item wordlists, published as part of the Global Lexicostatistical Database project. The lexical data have been subsequently analyzed with the aid of the principal phylogenetic methods, both distance-based and character-based: Starling neighbor joining (StarlingNJ), Neighbor joining (NJ), Unweighted pair group method with arithmetic mean (UPGMA), Bayesian Markov chain Monte Carlo (MCMC), Unweighted maximum parsimony (UMP). Cognation indexes within the input matrix were marked by two different algorithms: traditional etymological approach and phonetic similarity, i.e., the automatic method of consonant classes (Levenshtein distances). Due to certain reasons (first of all, high lexicographic quality of the wordlists and a consensus about the Lezgian phylogeny among Caucasologists), the Lezgian database is a perfect testing area for appraisal of phylogenetic methods. For the etymology-based input matrix, all the phylogenetic methods, with the possible exception of UMP, have yielded trees that are sufficiently compatible with each other to generate a consensus phylogenetic tree of the Lezgian lects. The obtained consensus tree agrees with the traditional expert classification as well as some of the previously proposed formal classifications of this linguistic group. Contrary to theoretical expectations, the UMP method has suggested the least plausible tree of all. In the case of the phonetic similarity-based input matrix, the distance-based methods (StarlingNJ, NJ, UPGMA) have produced the trees that are rather close to the consensus etymology-based tree and the traditional expert classification, whereas the character-based methods (Bayesian MCMC, UMP) have yielded less likely topologies.

  19. Towards a formal genealogical classification of the Lezgian languages (North Caucasus: testing various phylogenetic methods on lexical data.

    Directory of Open Access Journals (Sweden)

    Alexei Kassian

    Full Text Available A lexicostatistical classification is proposed for 20 languages and dialects of the Lezgian group of the North Caucasian family, based on meticulously compiled 110-item wordlists, published as part of the Global Lexicostatistical Database project. The lexical data have been subsequently analyzed with the aid of the principal phylogenetic methods, both distance-based and character-based: Starling neighbor joining (StarlingNJ, Neighbor joining (NJ, Unweighted pair group method with arithmetic mean (UPGMA, Bayesian Markov chain Monte Carlo (MCMC, Unweighted maximum parsimony (UMP. Cognation indexes within the input matrix were marked by two different algorithms: traditional etymological approach and phonetic similarity, i.e., the automatic method of consonant classes (Levenshtein distances. Due to certain reasons (first of all, high lexicographic quality of the wordlists and a consensus about the Lezgian phylogeny among Caucasologists, the Lezgian database is a perfect testing area for appraisal of phylogenetic methods. For the etymology-based input matrix, all the phylogenetic methods, with the possible exception of UMP, have yielded trees that are sufficiently compatible with each other to generate a consensus phylogenetic tree of the Lezgian lects. The obtained consensus tree agrees with the traditional expert classification as well as some of the previously proposed formal classifications of this linguistic group. Contrary to theoretical expectations, the UMP method has suggested the least plausible tree of all. In the case of the phonetic similarity-based input matrix, the distance-based methods (StarlingNJ, NJ, UPGMA have produced the trees that are rather close to the consensus etymology-based tree and the traditional expert classification, whereas the character-based methods (Bayesian MCMC, UMP have yielded less likely topologies.

  20. Molecular and morphological data supporting phylogenetic reconstruction of the genus Goniothalamus (Annonaceae, including a reassessment of previous infrageneric classifications

    Directory of Open Access Journals (Sweden)

    Chin Cheung Tang

    2015-09-01

    Full Text Available Data is presented in support of a phylogenetic reconstruction of the species-rich early-divergent angiosperm genus Goniothalamus (Annonaceae (Tang et al., Mol. Phylogenetic Evol., 2015 [1], inferred using chloroplast DNA (cpDNA sequences. The data includes a list of primers for amplification and sequencing for nine cpDNA regions: atpB-rbcL, matK, ndhF, psbA-trnH, psbM-trnD, rbcL, trnL-F, trnS-G, and ycf1, the voucher information and molecular data (GenBank accession numbers of 67 ingroup Goniothalamus accessions and 14 outgroup accessions selected from across the tribe Annoneae, and aligned data matrices for each gene region. We also present our Bayesian phylogenetic reconstructions for Goniothalamus, with information on previous infrageneric classifications superimposed to enable an evaluation of monophyly, together with a taxon-character data matrix (with 15 morphological characters scored for 66 Goniothalamus species and seven other species from the tribe Annoneae that are shown to be phylogenetically correlated.

  1. Towards a phylogenetic classification of dendrocoelid freshwater planarians (Platyhelminthes): a morphological and eclectic approach

    NARCIS (Netherlands)

    Sluys, R.; Kawakatsu, M.

    2006-01-01

    We explore and review the taxonomic distribution of morphological features that may be used as supporting apomorphies for the monophyletic status of various taxa in future, more comprehensive phylogenetic analyses of the dendrocoelid freshwater planarians and their close relatives. Characters

  2. Photometric brown-dwarf classification. I. A method to identify and accurately classify large samples of brown dwarfs without spectroscopy

    Science.gov (United States)

    Skrzypek, N.; Warren, S. J.; Faherty, J. K.; Mortlock, D. J.; Burgasser, A. J.; Hewett, P. C.

    2015-02-01

    Aims: We present a method, named photo-type, to identify and accurately classify L and T dwarfs onto the standard spectral classification system using photometry alone. This enables the creation of large and deep homogeneous samples of these objects efficiently, without the need for spectroscopy. Methods: We created a catalogue of point sources with photometry in 8 bands, ranging from 0.75 to 4.6 μm, selected from an area of 3344 deg2, by combining SDSS, UKIDSS LAS, and WISE data. Sources with 13.0 0.8, were then classified by comparison against template colours of quasars, stars, and brown dwarfs. The L and T templates, spectral types L0 to T8, were created by identifying previously known sources with spectroscopic classifications, and fitting polynomial relations between colour and spectral type. Results: Of the 192 known L and T dwarfs with reliable photometry in the surveyed area and magnitude range, 189 are recovered by our selection and classification method. We have quantified the accuracy of the classification method both externally, with spectroscopy, and internally, by creating synthetic catalogues and accounting for the uncertainties. We find that, brighter than J = 17.5, photo-type classifications are accurate to one spectral sub-type, and are therefore competitive with spectroscopic classifications. The resultant catalogue of 1157 L and T dwarfs will be presented in a companion paper.

  3. Towards a phylogenetic classification of dendrocoelid freshwater planarians (Platyhelminthes): a morphological and eclectic approach

    NARCIS (Netherlands)

    Sluys, R.; Kawakatsu, M.

    2006-01-01

    We explore and review the taxonomic distribution of morphological features that may be used as supporting apomorphies for the monophyletic status of various taxa in future, more comprehensive phylogenetic analyses of the dendrocoelid freshwater planarians and their close relatives. Characters examin

  4. Beyond classification: gene-family phylogenies from shotgun metagenomic reads enable accurate community analysis.

    Science.gov (United States)

    Riesenfeld, Samantha J; Pollard, Katherine S

    2013-06-22

    Sequence-based phylogenetic trees are a well-established tool for characterizing diversity of both macroorganisms and microorganisms. Phylogenetic methods have recently been applied to shotgun metagenomic data from microbial communities, particularly with the aim of classifying reads. But the accuracy of gene-family phylogenies that characterize evolutionary relationships among short, non-overlapping sequencing reads has not been thoroughly evaluated. To quantify errors in metagenomic read trees, we developed MetaPASSAGE, a software pipeline to generate in silico bacterial communities, simulate a sample of shotgun reads from a gene family represented in the community, orient or translate reads, and produce a profile-based alignment of the reads from which a gene-family phylogenetic tree can be built. We applied MetaPASSAGE to a variety of RNA and protein-coding gene families, built trees using a range of different phylogenetic methods, and compared the resulting trees using topological and branch-length error metrics. We identified read length as one of the major sources of error. Because phylogenetic methods use a reference database of full-length sequences from the gene family to guide construction of alignments and trees, we found that error can also be substantially reduced through increasing the size and diversity of the reference database. Finally, UniFrac analysis, which compares metagenomic samples based on a summary statistic computed over all branches in a read tree, is very robust to the level of error we observe. Bacterial community diversity can be quantified using phylogenetic approaches applied to shotgun metagenomic data. As sequencing reads get longer and more genomes across the bacterial tree of life are sequenced, the accuracy of this approach will continue to improve, opening the door to more applications.

  5. Short notes and reviews Simplifying hydrozoan classification: inappropriateness of the group Hydroidomedusae in a phylogenetic context

    NARCIS (Netherlands)

    Marques, Antonio C.

    2001-01-01

    The systematics of Hydrozoa is considered from the viewpoint of logical consistency between phylogeny and classification. The validity of the nominal taxon Hydroidomedusae (including all groups of Hydrozoa except the Siphonophorae) is discussed with regard to its distinctness and inclusive relations

  6. Chemical classification of cattle. 2. Phylogenetic tree and specific status of the Zebu.

    Science.gov (United States)

    Manwell, C; Baker, C M

    1980-01-01

    Phylogenetic trees for the ten major breed groups of cattle were constructed by Farris's (1972) maximum parsimony method, or Fitch & Margoliash's (1967) method, which averages ou the deviation over the entire assemblage. Both techniques yield essentially identical trees. The phylogenetic tree for the ten major cattle breed groups can be superimposed on a map of Europe and western Asia, the root of the tree being close to the 'fertile crescent' in Asia Minor, believed to be a primary centre of bovine domestication. For some but not all protein variants there is a cline of gene frequencies as one proceeds from the British Isles and northwest Europe towards southeast Europe and Asia Minor, with the most extreme gene frequencies in the Zebu breeds of India. It is not clear to what extent the observed clines are primary or secondary, i.e., consequent to the initial migrations of cattle towards the end of the Pleistocene or consequent to the many migrations of man with his domesticated cattle. Such clines as exist are not in themselves sufficient to prove either selection versus genetic drift or to establish taxonomic ranking. Contrary to some suggestions in the literature, the biochemical evidence supports Linnaeus's original conclusions: Bos taurus and Bos indicus are distinct species.

  7. A revised phylogenetic classification of the ant subfamily Formicinae (Hymenoptera: Formicidae), with resurrection of the genera Colobopsis and Dinomyrmex.

    Science.gov (United States)

    Ward, Philip S; Blaimer, Bonnie B; Fisher, Brian L

    2016-02-02

    The classification of the ant subfamily Formicinae is revised to reflect findings from a recent molecular phylogenetic study and complementary morphological investigations. The existing classification is maintained as far as possible, but some tribes and genera are redefined to ensure monophyly. Eleven tribes are recognized, all of which are strongly supported as monophyletic groups: Camponotini, Formicini, Gesomyrmecini, Gigantiopini, Lasiini (= Prenolepidii syn. n.), Melophorini (= Myrmecorhynchini syn. n.; = Notostigmatini syn. n.), Myrmelachistini stat. rev. (= Brachymyrmicini syn. n.), Myrmoteratini, Oecophyllini, Plagiolepidini, and Santschiellini stat. rev. Most of the tribes remain similar in content, but the generic composition of Lasiini, Melophorini, and Plagiolepidini is changed substantially. Species that have been placed in the genus Camponotus belong to three separate lineages. To ensure monophyly of this large, cosmopolitan genus we institute the following changes: Colobopsis and Dinomyrmex, both former subgenera of Camponotus, are elevated to genus level (stat. rev.), and two former genera, Forelophilus and Phasmomyrmex, are demoted to subgenus status (stat. n. and stat. rev., respectively) under Camponotus; two erstwhile subgenera of Phasmomyrmex, Myrmorhachis and Myrmacantha, become junior synonyms (syn. n.) of Camponotus (Phasmomyrmex); and the Camponotus subgenus Myrmogonia becomes a junior synonym (syn. n.) of Colobopsis. Dinomyrmex, represented by a single species from southeast Asia, D. gigas, is quite distinctive, but Camponotus and Colobopsis exhibit more subtle differences, despite being well separated phylogenetically. We identify morphological features of the worker caste that are broadly useful for distinguishing these two genera. Colobopsis species on the islands of New Caledonia and Fiji-regions with few native Camponotus species-tend to exceed these diagnostic bounds, but in this case regionally applicable character differences can

  8. Towards a phylogenetic classification of reef corals: The Indo-Pacific genera Merulina, Goniastrea and Scapophyllia (Scleractinia, Merulinidae)

    KAUST Repository

    Huang, Danwei

    2014-06-03

    Recent advances in scleractinian systematics and taxonomy have been achieved through the integration of molecular and morphological data, as well as rigorous analysis using phylogenetic methods. In this study, we continue in our pursuit of a phylogenetic classification by examining the evolutionary relationships between the closely related reef coral genera Merulina, Goniastrea, Paraclavarina and Scapophyllia (Merulinidae). In particular, we address the extreme polyphyly of Favites and Goniastrea that was discovered a decade ago. We sampled 145 specimens belonging to 16 species from a wide geographic range in the Indo-Pacific, focusing especially on type localities, including the Red Sea, western Indian Ocean and central Pacific. Tree reconstructions based on both nuclear and mitochondrial markers reveal a novel lineage composed of three species previously placed in Favites and Goniastrea. Morphological analyses indicate that this clade, Paragoniastrea Huang, Benzoni & Budd, gen. n., has a unique combination of corallite and subcorallite features observable with scanning electron microscopy and thin sections. Molecular and morphological evidence furthermore indicates that the monotypic genus Paraclavarina is nested within Merulina, and the former is therefore synonymised. © 2014 Royal Swedish Academy of Sciences.

  9. Archaeal-eubacterial mergers in the origin of Eukarya: phylogenetic classification of life.

    OpenAIRE

    Margulis, L

    1996-01-01

    A symbiosis-based phylogeny leads to a consistent, useful classification system for all life. "Kingdoms" and "Domains" are replaced by biological names for the most inclusive taxa: Prokarya (bacteria) and Eukarya (symbiosis-derived nucleated organisms). The earliest Eukarya, anaerobic mastigotes, hypothetically originated from permanent whole-cell fusion between members of Archaea (e.g., Thermoplasma-like organisms) and of Eubacteria (e.g., Spirochaeta-like organisms). Molecular biology, life...

  10. Separation of Benign and Malicious Network Events for Accurate Malware Family Classification

    Science.gov (United States)

    2015-09-28

    Perdisci, D. Dagon, W. Lee, and N. Feamster, “Building a dynamic reputation system for dns ,” in USENIX Security , 2010. [24] M. Antonakakis, R. Perdisci...W. Lee, N. V. II, and D. Dagon, “Detecting malware domains at the upper dns hierarchy,” in USENIX Security , 2011. [25] R. Perdisci, A. Lanzi, and W...ABSTRACT 16. SECURITY CLASSIFICATION OF: Labeling malware samples with their appropriate malware family helps understand and track malware evolution and

  11. Accurate HEp-2 cell classification based on sparse bag of words coding.

    Science.gov (United States)

    Ensafi, Shahab; Lu, Shijian; Kassim, Ashraf A; Tan, Chew Lim

    2017-04-01

    Autoimmune diseases (AD) are the abnormal response of the immune system of the body to healthy tissues. ADs have generally been on the increase. Efficient computer aided diagnosis of ADs through classification of the human epithelial type 2 (HEp-2) cells become beneficial. These methods make lower diagnosis costs, faster response and better diagnosis repeatability. In this paper, we present an automated HEp-2 cell image classification technique that exploits the sparse coding of the visual features together with the Bag of Words model (SBoW). In particular, SURF (Speeded Up Robust Features) and SIFT (Scale-invariant feature transform) features are specially integrated to work in a complementary fashion. This method helps greatly improve the cell classification accuracy. Additionally, a hierarchical max-pooling method is proposed to aggregate the local sparse codes in different layers to provide final feature vector. Furthermore, various parameters of the dictionary learning including the dictionary size, the learning iteration number, and the pooling strategy is also investigated. Experiments conducted on publicly available datasets show that the proposed technique clearly outperforms state-of-the-art techniques in cell and specimen levels. Copyright © 2016 Elsevier Ltd. All rights reserved.

  12. Archaeal-eubacterial mergers in the origin of Eukarya: phylogenetic classification of life

    Science.gov (United States)

    Margulis, L.

    1996-01-01

    A symbiosis-based phylogeny leads to a consistent, useful classification system for all life. "Kingdoms" and "Domains" are replaced by biological names for the most inclusive taxa: Prokarya (bacteria) and Eukarya (symbiosis-derived nucleated organisms). The earliest Eukarya, anaerobic mastigotes, hypothetically originated from permanent whole-cell fusion between members of Archaea (e.g., Thermoplasma-like organisms) and of Eubacteria (e.g., Spirochaeta-like organisms). Molecular biology, life-history, and fossil record evidence support the reunification of bacteria as Prokarya while subdividing Eukarya into uniquely defined subtaxa: Protoctista, Animalia, Fungi, and Plantae.

  13. Classification algorithms with multi-modal data fusion could accurately distinguish neuromyelitis optica from multiple sclerosis.

    Science.gov (United States)

    Eshaghi, Arman; Riyahi-Alam, Sadjad; Saeedi, Roghayyeh; Roostaei, Tina; Nazeri, Arash; Aghsaei, Aida; Doosti, Rozita; Ganjgahi, Habib; Bodini, Benedetta; Shakourirad, Ali; Pakravan, Manijeh; Ghana'ati, Hossein; Firouznia, Kavous; Zarei, Mojtaba; Azimi, Amir Reza; Sahraian, Mohammad Ali

    2015-01-01

    Neuromyelitis optica (NMO) exhibits substantial similarities to multiple sclerosis (MS) in clinical manifestations and imaging results and has long been considered a variant of MS. With the advent of a specific biomarker in NMO, known as anti-aquaporin 4, this assumption has changed; however, the differential diagnosis remains challenging and it is still not clear whether a combination of neuroimaging and clinical data could be used to aid clinical decision-making. Computer-aided diagnosis is a rapidly evolving process that holds great promise to facilitate objective differential diagnoses of disorders that show similar presentations. In this study, we aimed to use a powerful method for multi-modal data fusion, known as a multi-kernel learning and performed automatic diagnosis of subjects. We included 30 patients with NMO, 25 patients with MS and 35 healthy volunteers and performed multi-modal imaging with T1-weighted high resolution scans, diffusion tensor imaging (DTI) and resting-state functional MRI (fMRI). In addition, subjects underwent clinical examinations and cognitive assessments. We included 18 a priori predictors from neuroimaging, clinical and cognitive measures in the initial model. We used 10-fold cross-validation to learn the importance of each modality, train and finally test the model performance. The mean accuracy in differentiating between MS and NMO was 88%, where visible white matter lesion load, normal appearing white matter (DTI) and functional connectivity had the most important contributions to the final classification. In a multi-class classification problem we distinguished between all of 3 groups (MS, NMO and healthy controls) with an average accuracy of 84%. In this classification, visible white matter lesion load, functional connectivity, and cognitive scores were the 3 most important modalities. Our work provides preliminary evidence that computational tools can be used to help make an objective differential diagnosis of NMO and MS.

  14. Classification algorithms with multi-modal data fusion could accurately distinguish neuromyelitis optica from multiple sclerosis

    Directory of Open Access Journals (Sweden)

    Arman Eshaghi

    2015-01-01

    Full Text Available Neuromyelitis optica (NMO exhibits substantial similarities to multiple sclerosis (MS in clinical manifestations and imaging results and has long been considered a variant of MS. With the advent of a specific biomarker in NMO, known as anti-aquaporin 4, this assumption has changed; however, the differential diagnosis remains challenging and it is still not clear whether a combination of neuroimaging and clinical data could be used to aid clinical decision-making. Computer-aided diagnosis is a rapidly evolving process that holds great promise to facilitate objective differential diagnoses of disorders that show similar presentations. In this study, we aimed to use a powerful method for multi-modal data fusion, known as a multi-kernel learning and performed automatic diagnosis of subjects. We included 30 patients with NMO, 25 patients with MS and 35 healthy volunteers and performed multi-modal imaging with T1-weighted high resolution scans, diffusion tensor imaging (DTI and resting-state functional MRI (fMRI. In addition, subjects underwent clinical examinations and cognitive assessments. We included 18 a priori predictors from neuroimaging, clinical and cognitive measures in the initial model. We used 10-fold cross-validation to learn the importance of each modality, train and finally test the model performance. The mean accuracy in differentiating between MS and NMO was 88%, where visible white matter lesion load, normal appearing white matter (DTI and functional connectivity had the most important contributions to the final classification. In a multi-class classification problem we distinguished between all of 3 groups (MS, NMO and healthy controls with an average accuracy of 84%. In this classification, visible white matter lesion load, functional connectivity, and cognitive scores were the 3 most important modalities. Our work provides preliminary evidence that computational tools can be used to help make an objective differential diagnosis

  15. Automatic phylogenetic classification of bacterial beta-lactamase sequences including structural and antibiotic substrate preference information.

    Science.gov (United States)

    Ma, Jianmin; Eisenhaber, Frank; Maurer-Stroh, Sebastian

    2013-12-01

    Beta lactams comprise the largest and still most effective group of antibiotics, but bacteria can gain resistance through different beta lactamases that can degrade these antibiotics. We developed a user friendly tree building web server that allows users to assign beta lactamase sequences to their respective molecular classes and subclasses. Further clinically relevant information includes if the gene is typically chromosomal or transferable through plasmids as well as listing the antibiotics which the most closely related reference sequences are known to target and cause resistance against. This web server can automatically build three phylogenetic trees: the first tree with closely related sequences from a Tachyon search against the NCBI nr database, the second tree with curated reference beta lactamase sequences, and the third tree built specifically from substrate binding pocket residues of the curated reference beta lactamase sequences. We show that the latter is better suited to recover antibiotic substrate assignments through nearest neighbor annotation transfer. The users can also choose to build a structural model for the query sequence and view the binding pocket residues of their query relative to other beta lactamases in the sequence alignment as well as in the 3D structure relative to bound antibiotics. This web server is freely available at http://blac.bii.a-star.edu.sg/.

  16. Accurate Medium-Term Wind Power Forecasting in a Censored Classification Framework

    DEFF Research Database (Denmark)

    Dahl, Christian M.; Croonenbroeck, Carsten

    2014-01-01

    We provide a wind power forecasting methodology that exploits many of the actual data's statistical features, in particular both-sided censoring. While other tools ignore many of the important “stylized facts” or provide forecasts for short-term horizons only, our approach focuses on medium......-term forecasts, which are especially necessary for practitioners in the forward electricity markets of many power trading places; for example, NASDAQ OMX Commodities (formerly Nord Pool OMX Commodities) in northern Europe. We show that our model produces turbine-specific forecasts that are significantly more...... accurate in comparison to established benchmark models and present an application that illustrates the financial impact of more accurate forecasts obtained using our methodology....

  17. Deceptive desmas: molecular phylogenetics suggests a new classification and uncovers convergent evolution of lithistid demosponges.

    Directory of Open Access Journals (Sweden)

    Astrid Schuster

    Full Text Available Reconciling the fossil record with molecular phylogenies to enhance the understanding of animal evolution is a challenging task, especially for taxa with a mostly poor fossil record, such as sponges (Porifera. 'Lithistida', a polyphyletic group of recent and fossil sponges, are an exception as they provide the richest fossil record among demosponges. Lithistids, currently encompassing 13 families, 41 genera and >300 recent species, are defined by the common possession of peculiar siliceous spicules (desmas that characteristically form rigid articulated skeletons. Their phylogenetic relationships are to a large extent unresolved and there has been no (taxonomically comprehensive analysis to formally reallocate lithistid taxa to their closest relatives. This study, based on the most comprehensive molecular and morphological investigation of 'lithistid' demosponges to date, corroborates some previous weakly-supported hypotheses, and provides novel insights into the evolutionary relationships of the previous 'order Lithistida'. Based on molecular data (partial mtDNA CO1 and 28S rDNA sequences, we show that 8 out of 13 'Lithistida' families belong to the order Astrophorida, whereas Scleritodermidae and Siphonidiidae form a separate monophyletic clade within Tetractinellida. Most lithistid astrophorids are dispersed between different clades of the Astrophorida and we propose to formally reallocate them, respectively. Corallistidae, Theonellidae and Phymatellidae are monophyletic, whereas the families Pleromidae and Scleritodermidae are polyphyletic. Family Desmanthidae is polyphyletic and groups within Halichondriidae--we formally propose a reallocation. The sister group relationship of the family Vetulinidae to Spongillida is confirmed and we propose here for the first time to include Vetulina into a new Order Sphaerocladina. Megascleres and microscleres possibly evolved and/or were lost several times independently in different 'lithistid' taxa, and

  18. The International Classification of Headache Disorders: accurate diagnosis of orofacial pain?

    Science.gov (United States)

    Benoliel, R; Birman, N; Eliav, E; Sharav, Y

    2008-07-01

    The aim was to apply diagnostic criteria, as published by the International Headache Society (IHS), to the diagnosis of orofacial pain. A total of 328 consecutive patients with orofacial pain were collected over a period of 2 years. The orofacial pain clinic routinely employs criteria published by the IHS, the American Academy of Orofacial Pain (AAOP) and the Research Diagnostic Criteria for Temporomandibular Disorders (RDCTMD). Employing IHS criteria, 184 patients were successfully diagnosed (56%), including 34 with persistent idiopathic facial pain. In the remaining 144 we applied AAOP/RDCTMD criteria and diagnosed 120 as masticatory myofascial pain (MMP) resulting in a diagnostic efficiency of 92.7% (304/328) when applying the three classifications (IHS, AAOP, RDCTMD). Employing further published criteria, 23 patients were diagnosed as neurovascular orofacial pain (NVOP, facial migraine) and one as a neuropathy secondary to connective tissue disease. All the patients were therefore allocated to predefined diagnoses. MMP is clearly defined by AAOP and the RDCTMD. However, NVOP is not defined by any of the above classification systems. The features of MMP and NVOP are presented and analysed with calculations for positive (PPV) and negative predictive values (NPV). In MMP the combination of facial pain aggravated by jaw movement, and the presence of three or more tender muscles resulted in a PPV = 0.82 and a NPV = 0.86. For NVOP the combination of facial pain, throbbing quality, autonomic and/or systemic features and attack duration of > 60 min gave a PPV = 0.71 and a NPV = 0.95. Expansion of the IHS system is needed so as to integrate more orofacial pain syndromes.

  19. Photometric brown-dwarf classification. I. A method to identify and accurately classify large samples of brown dwarfs without spectroscopy

    CERN Document Server

    Skrzypek, Nathalie; Faherty, Jacqueline K; Mortlock, Daniel J; Burgasser, Adam J; Hewett, Paul C

    2014-01-01

    Aims. We present a method, named photo-type, to identify and accurately classify L and T dwarfs onto the standard spectral classification system using photometry alone. This enables the creation of large and deep homogeneous samples of these objects efficiently, without the need for spectroscopy. Methods. We created a catalogue of point sources with photometry in 8 bands, ranging from 0.75 to 4.6 microns, selected from an area of 3344 deg^2, by combining SDSS, UKIDSS LAS, and WISE data. Sources with 13.0 0.8, were then classified by comparison against template colours of quasars, stars, and brown dwarfs. The L and T templates, spectral types L0 to T8, were created by identifying previously known sources with spectroscopic classifications, and fitting polynomial relations between colour and spectral type. Results. Of the 192 known L and T dwarfs with reliable photometry in the surveyed area and magnitude range, 189 are recovered by our selection and classification method. We have quantified the accuracy of th...

  20. Accurate classification of brain gliomas by discriminate dictionary learning based on projective dictionary pair learning of proton magnetic resonance spectra.

    Science.gov (United States)

    Adebileje, Sikiru Afolabi; Ghasemi, Keyvan; Aiyelabegan, Hammed Tanimowo; Saligheh Rad, Hamidreza

    2017-04-01

    Proton magnetic resonance spectroscopy is a powerful noninvasive technique that complements the structural images of cMRI, which aids biomedical and clinical researches, by identifying and visualizing the compositions of various metabolites within the tissues of interest. However, accurate classification of proton magnetic resonance spectroscopy is still a challenging issue in clinics due to low signal-to-noise ratio, overlapping peaks of metabolites, and the presence of background macromolecules. This paper evaluates the performance of a discriminate dictionary learning classifiers based on projective dictionary pair learning method for brain gliomas proton magnetic resonance spectroscopy spectra classification task, and the result were compared with the sub-dictionary learning methods. The proton magnetic resonance spectroscopy data contain a total of 150 spectra (74 healthy, 23 grade II, 23 grade III, and 30 grade IV) from two databases. The datasets from both databases were first coupled together, followed by column normalization. The Kennard-Stone algorithm was used to split the datasets into its training and test sets. Performance comparison based on the overall accuracy, sensitivity, specificity, and precision was conducted. Based on the overall accuracy of our classification scheme, the dictionary pair learning method was found to outperform the sub-dictionary learning methods 97.78% compared with 68.89%, respectively. Copyright © 2016 John Wiley & Sons, Ltd.

  1. Treatment response classification of liver metastatic disease evaluated on imaging. Are RECIST unidimensional measurements accurate?

    Science.gov (United States)

    Mantatzis, Michael; Kakolyris, Stylianos; Amarantidis, Kyriakos; Karayiannakis, Anastasios; Prassopoulos, Panos

    2009-07-01

    The purpose of this study was to evaluate the accuracy of unidimensional measurements (response evaluation criteria in solid tumors, RECIST) compared with volumetric measurements in patients with liver metastases undergoing chemotherapy. Forty-four patients with newly diagnosed liver lesions underwent three MRI examinations at treatment initiation, during chemotherapy, and immediately post-treatment. Measurements based on RECIST guidelines and volume calculations were performed on the "target" lesions (TLs). The two methods were in agreement in 64/77 of patients and 253/301 of individual lesions classification in response categories ("good" agreement, Cohen kappa = 0.735 and 0.741, respectively). In 16.88% of the comparisons the two methods stratified patients to a different response category; 27.6% of TLs did not follow the response category of the patient in whom lesions were located. The actual volume of TLs differs from the calculated volume of a sphere with the same diameter. Our study supports the use of volumetric techniques that may overcome certain disadvantages of unidimensional measurements.

  2. A Void Reference Sensor-Multiple Signal Classification Algorithm for More Accurate Direction of Arrival Estimation of Low Altitude Target

    Institute of Scientific and Technical Information of China (English)

    XIAO Hui; SUN Jin-cai; YUAN Jun; NIU Yi-long

    2007-01-01

    There exists MUSIC (multiple signal classification) algorithm for direction of arrival (DOA) estimation. This paper is to present a different MUSIC algorithm for more accurate estimation of low altitude target. The possibility of better performance is analyzed using a void reference sensor (VRS) in MUSIC algorithm. The following two topics are discussed: 1) the time delay formula and VRS-MUSIC algorithm with VRS located on the minus of z-axes; 2) the DOA estimation results of VRS-MUSIC and MUSIC algorithms. The simulation results show VRS-MUSIC algorithm has three advantages compared with MUSIC: 1 ) When the signal to noise ratio (SNR) is more than - 5 dB, the direction estimation error is 1/2 as much as that obtained by MUSIC; 2) The side lobe is more lower and the stability is better; 3) The size of array that the algorithm requires is smaller.

  3. Bilateral weighted radiographs are required for accurate classification of acromioclavicular separation: an observational study of 59 cases.

    Science.gov (United States)

    Ibrahim, E F; Forrest, N P; Forester, A

    2015-10-01

    Misinterpretation of the Rockwood classification system for acromioclavicular joint (ACJ) separations has resulted in a trend towards using unilateral radiographs for grading. Further, the use of weighted views to 'unmask' a grade III injury has fallen out of favour. Recent evidence suggests that many radiographic grade III injuries represent only a partial injury to the stabilising ligaments. This study aimed to determine (1) whether accurate classification is possible on unilateral radiographs and (2) the efficacy of weighted bilateral radiographs in unmasking higher-grade injuries. Complete bilateral non-weighted and weighted sets of radiographs for patients presenting with an acromioclavicular separation over a 10-year period were analysed retrospectively, and they were graded I-VI according to Rockwood's criteria. Comparison was made between grading based on (1) a single antero-posterior (AP) view of the injured side, (2) bilateral non-weighted views and (3) bilateral weighted views. Radiographic measurements for cases that changed grade after weighted views were statistically compared to see if this could have been predicted beforehand. Fifty-nine sets of radiographs on 59 patients (48 male, mean age of 33 years) were included. Compared with unilateral radiographs, non-weighted bilateral comparison films resulted in a grade change for 44 patients (74.5%). Twenty-eight of 56 patients initially graded as I, II or III were upgraded to grade V and two of three initial grade V patients were downgraded to grade III. The addition of a weighted view further upgraded 10 patients to grade V. No grade II injury was changed to grade III and no injury of any severity was downgraded by a weighted view. Grade III injuries upgraded on weighted views had a significantly greater baseline median percentage coracoclavicular distance increase than those that were not upgraded (80.7% vs. 55.4%, p=0.015). However, no cut-off point for this value could be identified to predict an

  4. Protein clustering and RNA phylogenetic reconstruction of the influenza A [corrected] virus NS1 protein allow an update in classification and identification of motif conservation.

    Directory of Open Access Journals (Sweden)

    Edgar E Sevilla-Reyes

    Full Text Available The non-structural protein 1 (NS1 of influenza A virus (IAV, coded by its third most diverse gene, interacts with multiple molecules within infected cells. NS1 is involved in host immune response regulation and is a potential contributor to the virus host range. Early phylogenetic analyses using 50 sequences led to the classification of NS1 gene variants into groups (alleles A and B. We reanalyzed NS1 diversity using 14,716 complete NS IAV sequences, downloaded from public databases, without host bias. Removal of sequence redundancy and further structured clustering at 96.8% amino acid similarity produced 415 clusters that enhanced our capability to detect distinct subgroups and lineages, which were assigned a numerical nomenclature. Maximum likelihood phylogenetic reconstruction using RNA sequences indicated the previously identified deep branching separating group A from group B, with five distinct subgroups within A as well as two and five lineages within the A4 and A5 subgroups, respectively. Our classification model proposes that sequence patterns in thirteen amino acid positions are sufficient to fit >99.9% of all currently available NS1 sequences into the A subgroups/lineages or the B group. This classification reduces host and virus bias through the prioritization of NS1 RNA phylogenetics over host or virus phenetics. We found significant sequence conservation within the subgroups and lineages with characteristic patterns of functional motifs, such as the differential binding of CPSF30 and crk/crkL or the availability of a C-terminal PDZ-binding motif. To understand selection pressures and evolution acting on NS1, it is necessary to organize the available data. This updated classification may help to clarify and organize the study of NS1 interactions and pathogenic differences and allow the drawing of further functional inferences on sequences in each group, subgroup and lineage rather than on a strain-by-strain basis.

  5. Revisiting the phylogeny of Bombacoideae (Malvaceae): Novel relationships, morphologically cohesive clades, and a new tribal classification based on multilocus phylogenetic analyses.

    Science.gov (United States)

    Carvalho-Sobrinho, Jefferson G; Alverson, William S; Alcantara, Suzana; Queiroz, Luciano P; Mota, Aline C; Baum, David A

    2016-08-01

    Bombacoideae (Malvaceae) is a clade of deciduous trees with a marked dominance in many forests, especially in the Neotropics. The historical lack of a well-resolved phylogenetic framework for Bombacoideae hinders studies in this ecologically important group. We reexamined phylogenetic relationships in this clade based on a matrix of 6465 nuclear (ETS, ITS) and plastid (matK, trnL-trnF, trnS-trnG) DNA characters. We used maximum parsimony, maximum likelihood, and Bayesian inference to infer relationships among 108 species (∼70% of the total number of known species). We analyzed the evolution of selected morphological traits: trunk or branch prickles, calyx shape, endocarp type, seed shape, and seed number per fruit, using ML reconstructions of their ancestral states to identify possible synapomorphies for major clades. Novel phylogenetic relationships emerged from our analyses, including three major lineages marked by fruit or seed traits: the winged-seed clade (Bernoullia, Gyranthera, and Huberodendron), the spongy endocarp clade (Adansonia, Aguiaria, Catostemma, Cavanillesia, and Scleronema), and the Kapok clade (Bombax, Ceiba, Eriotheca, Neobuchia, Pachira, Pseudobombax, Rhodognaphalon, and Spirotheca). The Kapok clade, the most diverse lineage of the subfamily, includes sister relationships (i) between Pseudobombax and "Pochota fendleri" a historically incertae sedis taxon, and (ii) between the Paleotropical genera Bombax and Rhodognaphalon, implying just two bombacoid dispersals to the Old World, the other one involving Adansonia. This new phylogenetic framework offers new insights and a promising avenue for further evolutionary studies. In view of this information, we present a new tribal classification of the subfamily, accompanied by an identification key.

  6. Comprehensive phylogenetic reconstructions of African swine fever virus: proposal for a new classification and molecular dating of the virus.

    Directory of Open Access Journals (Sweden)

    Vincent Michaud

    Full Text Available African swine fever (ASF is a highly lethal disease of domestic pigs caused by the only known DNA arbovirus. It was first described in Kenya in 1921 and since then many isolates have been collected worldwide. However, although several phylogenetic studies have been carried out to understand the relationships between the isolates, no molecular dating analyses have been achieved so far. In this paper, comprehensive phylogenetic reconstructions were made using newly generated, publicly available sequences of hundreds of ASFV isolates from the past 70 years. Analyses focused on B646L, CP204L, and E183L genes from 356, 251, and 123 isolates, respectively. Phylogenetic analyses were achieved using maximum likelihood and Bayesian coalescence methods. A new lineage-based nomenclature is proposed to designate 35 different clusters. In addition, dating of ASFV origin was carried out from the molecular data sets. To avoid bias, diversity due to positive selection or recombination events was neutralized. The molecular clock analyses revealed that ASFV strains currently circulating have evolved over 300 years, with a time to the most recent common ancestor (TMRCA in the early 18(th century.

  7. The molecular genetics and morphometry-based Endometrial Intraepithelial Neoplasia classification system predicts disease progression in Endometrial hyperplasia more accurately than the 1994 World Health Organization classification system

    NARCIS (Netherlands)

    Baak, JP; Mutter, GL; Robboy, S; van Diest, PJ; Uyterlinde, AM; Orbo, A; Palazzo, J; Fiane, B; Lovslett, K; Burger, C; Voorhorst, F; Verheijen, RH

    2005-01-01

    BACKGROUND. The objective of this study was to compare the accuracy of disease progression prediction of the molecular genetics and morphometry-based Endometrial Intraepithelial Neoplasia (EIN) and World Health Organization 1994 (WHO94) classification systems in patients with endometrial hyperplasia

  8. When proglottids and scoleces conflict: phylogenetic relationships and a family-level classification of the Lecanicephalidea (Platyhelminthes: Cestoda).

    Science.gov (United States)

    Jensen, Kirsten; Caira, Janine N; Cielocha, Joanna J; Littlewood, D Timothy J; Waeschenbach, Andrea

    2016-05-01

    This study presents the first comprehensive phylogenetic analysis of the interrelationships of the morphologically diverse elasmobranch-hosted tapeworm order Lecanicephalidea, based on molecular sequence data. With almost half of current generic diversity having been erected or resurrected within the last decade, an apparent conflict between scolex morphology and proglottid anatomy has hampered the assignment of many of these genera to families. Maximum likelihood and Bayesian analyses of two nuclear markers (D1-D3 of lsrDNA and complete ssrDNA) and two mitochondrial markers (partial rrnL and partial cox1) for 61 lecanicephalidean species representing 22 of the 25 valid genera were conducted; new sequence data were generated for 43 species and 11 genera, including three undescribed genera. The monophyly of the order was confirmed in all but the analyses based on cox1 data alone. Sesquipedalapex placed among species of Anteropora and was thus synonymized with the latter genus. Based on analyses of the concatenated dataset, eight major groups emerged which are herein formally recognised at the familial level. Existing family names (i.e., Lecanicephalidae, Polypocephalidae, Tetragonocephalidae, and Cephalobothriidae) are maintained for four of the eight clades, and new families are proposed for the remaining four groups (Aberrapecidae n. fam., Eniochobothriidae n. fam., Paraberrapecidae n. fam., and Zanobatocestidae n. fam.). The four new families and the Tetragonocephalidae are monogeneric, while the Cephalobothriidae, Lecanicephalidae and Polypocephalidae comprise seven, eight and four genera, respectively. As a result of their unusual morphologies, the three genera not included here (i.e., Corrugatocephalum, Healyum and Quadcuspibothrium) are considered incertae sedis within the order until their familial affinities can be examined in more detail. All eight families are newly circumscribed based on morphological features and a key to the families is provided

  9. Classification

    Science.gov (United States)

    Clary, Renee; Wandersee, James

    2013-01-01

    In this article, Renee Clary and James Wandersee describe the beginnings of "Classification," which lies at the very heart of science and depends upon pattern recognition. Clary and Wandersee approach patterns by first telling the story of the "Linnaean classification system," introduced by Carl Linnacus (1707-1778), who is…

  10. Classification

    DEFF Research Database (Denmark)

    Hjørland, Birger

    2017-01-01

    This article presents and discusses definitions of the term “classification” and the related concepts “Concept/conceptualization,”“categorization,” “ordering,” “taxonomy” and “typology.” It further presents and discusses theories of classification including the influences of Aristotle...... and Wittgenstein. It presents different views on forming classes, including logical division, numerical taxonomy, historical classification, hermeneutical and pragmatic/critical views. Finally, issues related to artificial versus natural classification and taxonomic monism versus taxonomic pluralism are briefly...

  11. Accurate white matter lesion segmentation by k nearest neighbor classification with tissue type priors (kNN-TTPs).

    Science.gov (United States)

    Steenwijk, Martijn D; Pouwels, Petra J W; Daams, Marita; van Dalen, Jan Willem; Caan, Matthan W A; Richard, Edo; Barkhof, Frederik; Vrenken, Hugo

    2013-01-01

    The segmentation and volumetric quantification of white matter (WM) lesions play an important role in monitoring and studying neurological diseases such as multiple sclerosis (MS) or cerebrovascular disease. This is often interactively done using 2D magnetic resonance images. Recent developments in acquisition techniques allow for 3D imaging with much thinner sections, but the large number of images per subject makes manual lesion outlining infeasible. This warrants the need for a reliable automated approach. Here we aimed to improve k nearest neighbor (kNN) classification of WM lesions by optimizing intensity normalization and using spatial tissue type priors (TTPs). The kNN-TTP method used kNN classification with 3.0 T 3DFLAIR and 3DT1 intensities as well as MNI-normalized spatial coordinates as features. Additionally, TTPs were computed by nonlinear registration of data from healthy controls. Intensity features were normalized using variance scaling, robust range normalization or histogram matching. The algorithm was then trained and evaluated using a leave-one-out experiment among 20 patients with MS against a reference segmentation that was created completely manually. The performance of each normalization method was evaluated both with and without TTPs in the feature set. Volumetric agreement was evaluated using intra-class coefficient (ICC), and voxelwise spatial agreement was evaluated using Dice similarity index (SI). Finally, the robustness of the method across different scanners and patient populations was evaluated using an independent sample of elderly subjects with hypertension. The intensity normalization method had a large influence on the segmentation performance, with average SI values ranging from 0.66 to 0.72 when no TTPs were used. Independent of the normalization method, the inclusion of TTPs as features increased performance particularly by reducing the lesion detection error. Best performance was achieved using variance scaled intensity

  12. Magnetic resonance imaging is comparable to computed tomography for determination of glenoid version but does not accurately distinguish between Walch B2 and C classifications.

    Science.gov (United States)

    Lowe, Jeremiah T; Testa, Edward J; Li, Xinning; Miller, Suzanne; DeAngelis, Joseph P; Jawa, Andrew

    2017-04-01

    Computed tomography (CT) scan is the standard for the preoperative assessment of glenoid version and morphology before total shoulder arthroplasty. However, the capacity of magnetic resonance imaging (MRI) to visualize bone morphology has improved with advancing technology. The purpose of this study was to compare the accuracy of MRI to CT for assessment of glenoid version and Walch classification. Three fellowship-trained shoulder surgeons assessed glenoid version and Walch classification of 30 patients with primary shoulder osteoarthritis who received both CT and MRI scans before total shoulder arthroplasty. Version measurements, Walch classification, and observer agreement were compared. Mean glenoid version was -15.5° and -18.6° by CT and MRI, respectively (P = .17). Interobserver reliability coefficients were good for both imaging modalities (CT, 0.73; MRI, 0.62). Intraobserver coefficients were good to excellent for CT (range, 0.76-0.87) and good for MRI (range, 0.75-0.79). For Walch classification, interobserver reliability for both modalities was merely fair, whereas intraobserver reliability was moderate to good. Although identification of type A1, A2, and B1 was nearly identical between CT and MRI, there was observer disagreement on type B2 (P = .001) and C glenoids (P = .03). Specifically, MRI underidentified type B2 and overidentified type C compared with CT. MRI is largely comparable to CT scan for evaluation of the glenoid, with similar measurements of version and identification of less extreme Walch glenoids. However, MRI is less accurate at distinguishing between type B2 and C glenoids. Copyright © 2017 Journal of Shoulder and Elbow Surgery Board of Trustees. Published by Elsevier Inc. All rights reserved.

  13. A species independent universal bio-detection microarray for pathogen forensics and phylogenetic classification of unknown microorganisms

    Directory of Open Access Journals (Sweden)

    McCormick John

    2011-06-01

    Full Text Available Abstract Background The ability to differentiate a bioterrorist attack or an accidental release of a research pathogen from a naturally occurring pandemic or disease event is crucial to the safety and security of this nation by enabling an appropriate and rapid response. It is critical in samples from an infected patient, the environment, or a laboratory to quickly and accurately identify the precise pathogen including natural or engineered variants and to classify new pathogens in relation to those that are known. Current approaches for pathogen detection rely on prior genomic sequence information. Given the enormous spectrum of genetic possibilities, a field deployable, robust technology, such as a universal (any species microarray has near-term potential to address these needs. Results A new and comprehensive sequence-independent array (Universal Bio-Signature Detection Array was designed with approximately 373,000 probes. The main feature of this array is that the probes are computationally derived and sequence independent. There is one probe for each possible 9-mer sequence, thus 49 (262,144 probes. Each genome hybridized on this array has a unique pattern of signal intensities corresponding to each of these probes. These signal intensities were used to generate an un-biased cluster analysis of signal intensity hybridization patterns that can easily distinguish species into accepted and known phylogenomic relationships. Within limits, the array is highly sensitive and is able to detect synthetically mixed pathogens. Examples of unique hybridization signal intensity patterns are presented for different Brucella species as well as relevant host species and other pathogens. These results demonstrate the utility of the UBDA array as a diagnostic tool in pathogen forensics. Conclusions This pathogen detection system is fast, accurate and can be applied to any species. Hybridization patterns are unique to a specific genome and these can be used

  14. MetaShot: an accurate workflow for taxon classification of host-associated microbiome from shotgun metagenomic data.

    Science.gov (United States)

    Fosso, B; Santamaria, M; D'Antonio, M; Lovero, D; Corrado, G; Vizza, E; Passaro, N; Garbuglia, A R; Capobianchi, M R; Crescenzi, M; Valiente, G; Pesole, G

    2017-06-01

    Shotgun metagenomics by high-throughput sequencing may allow deep and accurate characterization of host-associated total microbiomes, including bacteria, viruses, protists and fungi. However, the analysis of such sequencing data is still extremely challenging in terms of both overall accuracy and computational efficiency, and current methodologies show substantial variability in misclassification rate and resolution at lower taxonomic ranks or are limited to specific life domains (e.g. only bacteria). We present here MetaShot, a workflow for assessing the total microbiome composition from host-associated shotgun sequence data, and show its overall optimal accuracy performance by analyzing both simulated and real datasets. https://github.com/bfosso/MetaShot. graziano.pesole@uniba.it. Supplementary data are available at Bioinformatics online.

  15. Fast, Simple and Accurate Handwritten Digit Classification by Training Shallow Neural Network Classifiers with the 'Extreme Learning Machine' Algorithm.

    Science.gov (United States)

    McDonnell, Mark D; Tissera, Migel D; Vladusich, Tony; van Schaik, André; Tapson, Jonathan

    2015-01-01

    Recent advances in training deep (multi-layer) architectures have inspired a renaissance in neural network use. For example, deep convolutional networks are becoming the default option for difficult tasks on large datasets, such as image and speech recognition. However, here we show that error rates below 1% on the MNIST handwritten digit benchmark can be replicated with shallow non-convolutional neural networks. This is achieved by training such networks using the 'Extreme Learning Machine' (ELM) approach, which also enables a very rapid training time (∼ 10 minutes). Adding distortions, as is common practise for MNIST, reduces error rates even further. Our methods are also shown to be capable of achieving less than 5.5% error rates on the NORB image database. To achieve these results, we introduce several enhancements to the standard ELM algorithm, which individually and in combination can significantly improve performance. The main innovation is to ensure each hidden-unit operates only on a randomly sized and positioned patch of each image. This form of random 'receptive field' sampling of the input ensures the input weight matrix is sparse, with about 90% of weights equal to zero. Furthermore, combining our methods with a small number of iterations of a single-batch backpropagation method can significantly reduce the number of hidden-units required to achieve a particular performance. Our close to state-of-the-art results for MNIST and NORB suggest that the ease of use and accuracy of the ELM algorithm for designing a single-hidden-layer neural network classifier should cause it to be given greater consideration either as a standalone method for simpler problems, or as the final classification stage in deep neural networks applied to more difficult problems.

  16. Fast, Simple and Accurate Handwritten Digit Classification by Training Shallow Neural Network Classifiers with the 'Extreme Learning Machine' Algorithm.

    Directory of Open Access Journals (Sweden)

    Mark D McDonnell

    Full Text Available Recent advances in training deep (multi-layer architectures have inspired a renaissance in neural network use. For example, deep convolutional networks are becoming the default option for difficult tasks on large datasets, such as image and speech recognition. However, here we show that error rates below 1% on the MNIST handwritten digit benchmark can be replicated with shallow non-convolutional neural networks. This is achieved by training such networks using the 'Extreme Learning Machine' (ELM approach, which also enables a very rapid training time (∼ 10 minutes. Adding distortions, as is common practise for MNIST, reduces error rates even further. Our methods are also shown to be capable of achieving less than 5.5% error rates on the NORB image database. To achieve these results, we introduce several enhancements to the standard ELM algorithm, which individually and in combination can significantly improve performance. The main innovation is to ensure each hidden-unit operates only on a randomly sized and positioned patch of each image. This form of random 'receptive field' sampling of the input ensures the input weight matrix is sparse, with about 90% of weights equal to zero. Furthermore, combining our methods with a small number of iterations of a single-batch backpropagation method can significantly reduce the number of hidden-units required to achieve a particular performance. Our close to state-of-the-art results for MNIST and NORB suggest that the ease of use and accuracy of the ELM algorithm for designing a single-hidden-layer neural network classifier should cause it to be given greater consideration either as a standalone method for simpler problems, or as the final classification stage in deep neural networks applied to more difficult problems.

  17. Phylogenetic trees

    OpenAIRE

    Baños, Hector; Bushek, Nathaniel; Davidson, Ruth; Gross, Elizabeth; Harris, Pamela E.; Krone, Robert; Long, Colby; Stewart, Allen; WALKER, Robert

    2016-01-01

    We introduce the package PhylogeneticTrees for Macaulay2 which allows users to compute phylogenetic invariants for group-based tree models. We provide some background information on phylogenetic algebraic geometry and show how the package PhylogeneticTrees can be used to calculate a generating set for a phylogenetic ideal as well as a lower bound for its dimension. Finally, we show how methods within the package can be used to compute a generating set for the join of any two ideals.

  18. Phylogenetic Status of an Unrecorded Species of Curvularia, C. spicifera, Based on Current Classification System of Curvularia and Bipolaris Group Using Multi Loci.

    Science.gov (United States)

    Jeon, Sun Jeong; Nguyen, Thi Thuong Thuong; Lee, Hyang Burm

    2015-09-01

    A seed-borne fungus, Curvularia sp. EML-KWD01, was isolated from an indigenous wheat seed by standard blotter method. This fungus was characterized based on the morphological characteristics and molecular phylogenetic analysis. Phylogenetic status of the fungus was determined using sequences of three loci: rDNA internal transcribed spacer, large ribosomal subunit, and glyceraldehyde 3-phosphate dehydrogenase gene. Multi loci sequencing analysis revealed that this fungus was Curvularia spicifera within Curvularia group 2 of family Pleosporaceae.

  19. A High Resolution/Accurate Mass (HRAM) Data-Dependent MS3 Neutral Loss Screening, Classification, and Relative Quantitation Methodology for Carbonyl Compounds in Saliva

    Science.gov (United States)

    Dator, Romel; Carrà, Andrea; Maertens, Laura; Guidolin, Valeria; Villalta, Peter W.; Balbo, Silvia

    2016-10-01

    Reactive carbonyl compounds (RCCs) are ubiquitous in the environment and are generated endogenously as a result of various physiological and pathological processes. These compounds can react with biological molecules inducing deleterious processes believed to be at the basis of their toxic effects. Several of these compounds are implicated in neurotoxic processes, aging disorders, and cancer. Therefore, a method characterizing exposures to these chemicals will provide insights into how they may influence overall health and contribute to disease pathogenesis. Here, we have developed a high resolution accurate mass (HRAM) screening strategy allowing simultaneous identification and relative quantitation of DNPH-derivatized carbonyls in human biological fluids. The screening strategy involves the diagnostic neutral loss of hydroxyl radical triggering MS3 fragmentation, which is only observed in positive ionization mode of DNPH-derivatized carbonyls. Unique fragmentation pathways were used to develop a classification scheme for characterizing known and unanticipated/unknown carbonyl compounds present in saliva. Furthermore, a relative quantitation strategy was implemented to assess variations in the levels of carbonyl compounds before and after exposure using deuterated d 3 -DNPH. This relative quantitation method was tested on human samples before and after exposure to specific amounts of alcohol. The nano-electrospray ionization (nano-ESI) in positive mode afforded excellent sensitivity with detection limits on-column in the high-attomole levels. To the best of our knowledge, this is the first report of a method using HRAM neutral loss screening of carbonyl compounds. In addition, the method allows simultaneous characterization and relative quantitation of DNPH-derivatized compounds using nano-ESI in positive mode.

  20. A High Resolution/Accurate Mass (HRAM) Data-Dependent MS3 Neutral Loss Screening, Classification, and Relative Quantitation Methodology for Carbonyl Compounds in Saliva

    Science.gov (United States)

    Dator, Romel; Carrà, Andrea; Maertens, Laura; Guidolin, Valeria; Villalta, Peter W.; Balbo, Silvia

    2017-04-01

    Reactive carbonyl compounds (RCCs) are ubiquitous in the environment and are generated endogenously as a result of various physiological and pathological processes. These compounds can react with biological molecules inducing deleterious processes believed to be at the basis of their toxic effects. Several of these compounds are implicated in neurotoxic processes, aging disorders, and cancer. Therefore, a method characterizing exposures to these chemicals will provide insights into how they may influence overall health and contribute to disease pathogenesis. Here, we have developed a high resolution accurate mass (HRAM) screening strategy allowing simultaneous identification and relative quantitation of DNPH-derivatized carbonyls in human biological fluids. The screening strategy involves the diagnostic neutral loss of hydroxyl radical triggering MS3 fragmentation, which is only observed in positive ionization mode of DNPH-derivatized carbonyls. Unique fragmentation pathways were used to develop a classification scheme for characterizing known and unanticipated/unknown carbonyl compounds present in saliva. Furthermore, a relative quantitation strategy was implemented to assess variations in the levels of carbonyl compounds before and after exposure using deuterated d 3 -DNPH. This relative quantitation method was tested on human samples before and after exposure to specific amounts of alcohol. The nano-electrospray ionization (nano-ESI) in positive mode afforded excellent sensitivity with detection limits on-column in the high-attomole levels. To the best of our knowledge, this is the first report of a method using HRAM neutral loss screening of carbonyl compounds. In addition, the method allows simultaneous characterization and relative quantitation of DNPH-derivatized compounds using nano-ESI in positive mode.

  1. Didiscus verdensis spec. nov. (Porifera: Halichondrida) from the Cape Verde Islands, with a revision and phylogenetic classification of the genus Didiscus

    NARCIS (Netherlands)

    Hiemstra, F.; Soest, van R.W.M.

    1991-01-01

    A new species of the circumtropical/subtropical genus Didiscus Dendy, 1922 is described from the Cape Verde Islands. Based on a phylogenetic analysis of all known species of the genus, using morphological and microscopical (including SEM) characters, it was demonstrated that the new species is close

  2. Evolutionary Phylogenetic Networks: Models and Issues

    Science.gov (United States)

    Nakhleh, Luay

    Phylogenetic networks are special graphs that generalize phylogenetic trees to allow for modeling of non-treelike evolutionary histories. The ability to sequence multiple genetic markers from a set of organisms and the conflicting evolutionary signals that these markers provide in many cases, have propelled research and interest in phylogenetic networks to the forefront in computational phylogenetics. Nonetheless, the term 'phylogenetic network' has been generically used to refer to a class of models whose core shared property is tree generalization. Several excellent surveys of the different flavors of phylogenetic networks and methods for their reconstruction have been written recently. However, unlike these surveys, this chapte focuses specifically on one type of phylogenetic networks, namely evolutionary phylogenetic networks, which explicitly model reticulate evolutionary events. Further, this chapter focuses less on surveying existing tools, and addresses in more detail issues that are central to the accurate reconstruction of phylogenetic networks.

  3. Molecular phylogenetics of Alchemilla, Aphanes and Lachemilla (Rosaceae) inferred from plastid and nuclear intron and spacer DNA sequences, with comments on generic classification.

    Science.gov (United States)

    Gehrke, B; Bräuchler, C; Romoleroux, K; Lundberg, M; Heubl, G; Eriksson, T

    2008-06-01

    Alchemilla (the lady's mantles) is a well known but inconspicuous group in the Rosaceae, notable for its ornamental leaves and pharmaceutical properties. The systematics of Alchemilla has remained poorly understood, most likely due to confusion resulting from apomixis, polyploidisation and hybridisation, which are frequently observed in the group, and which have led to the description of a large number of (micro-) species. A molecular phylogeny of the genus, including all sections of Alchemilla and Lachemilla as well as five representatives of Aphanes, based on the analysis of the chloroplast trnL-trnF and the nuclear ITS regions is presented here. Gene phylogenies reconstructed from the nuclear and chloroplast sequence data were largely congruent. Limited conflict between the data partitions was observed with respect to a small number of taxa. This is likely to be the result of hybridisation/introgression or incomplete lineage sorting. Four distinct clades were resolved, corresponding to major geographical division and life forms: Eurasian Alchemilla, annual Aphanes, South American Lachemilla and African Alchemilla. We argue for a wider circumscription of the genus Alchemilla, including Lachemilla and Aphanes, based on the morphology and the phylogenetic relationships between the different clades.

  4. Redescription and phylogenetic position of Myxobolus aeglefini and Myxobolus platessae n. comb. (Myxosporea), parasites in the cartilage of some North Atlantic marine fishes, with notes on the phylogeny and classification of the Platysporina.

    Science.gov (United States)

    Karlsbakk, Egil; Kristmundsson, Árni; Albano, Marco; Brown, Paul; Freeman, Mark A

    2017-02-01

    Myxobolus 'aeglefini' Auerbach, 1906 was originally described from cranial cartilage of North sea haddock (Melanogrammus aeglefinus), but has subsequently been recorded from cartilaginous tissues of a range of other gadoid hosts, from pleuronectids and from lumpsucker (Cyclopterus lumpus) in the North Atlantic and from a zoarcid fish in the Japan Sea (Pacific). We obtained partial small-subunit rDNA sequences of Myxobolus 'aeglefini' from gadoids and pleuronectids from Norway and Iceland. The sequences from gadoids and pleuronectids represented two different genotypes, showing 98.2% identity. Morphometric studies on the spores from selected gadids and pleuronectids revealed slight but statistically significant differences in spore dimensions associated with the genotypes, the spores from pleuronectids were thicker and with larger polar capsules. We identify the morpho- and genotype from gadoids with Myxobolus 'aeglefini' sensu Auerbach, and the one from pleuronectids with Sphaerospora platessae Woodcock, 1904 as Myxobolus platessae n. comb. The latter species was originally described from Irish Sea plaice (Pleuronectes platessa). Myxobolus albi Picon et al., 2009 described from the common goby Pomatoschistus microps in Scotland is a synonym of M. 'aeglefini'. The Pacific Myxobolus 'aeglefini' represents a separate species, showing only 97.4-97.6% identity to the Atlantic species. In phylogenetic analyses based on SSU rDNA sequences, these and some related marine chondrotropic Myxobolus spp. form a distinct well supported group. This clusters with freshwater and marine myxobolids and Triangula and Cardimyxobolus species, in a basal clade in the phylogeny of the Platysporina. Members of family Myxobilatidae, Ortholinea spp. (currently Ortholineidae) and sequences of some other urinary system infecting myxosporeans form a well supported clade among members of the suborder Platysporina. Based on phylogenetic analyses, we propose the following changes to the

  5. Combining multiple hypothesis testing and affinity propagation clustering leads to accurate, robust and sample size independent classification on gene expression data

    Directory of Open Access Journals (Sweden)

    Sakellariou Argiris

    2012-10-01

    Full Text Available Abstract Background A feature selection method in microarray gene expression data should be independent of platform, disease and dataset size. Our hypothesis is that among the statistically significant ranked genes in a gene list, there should be clusters of genes that share similar biological functions related to the investigated disease. Thus, instead of keeping N top ranked genes, it would be more appropriate to define and keep a number of gene cluster exemplars. Results We propose a hybrid FS method (mAP-KL, which combines multiple hypothesis testing and affinity propagation (AP-clustering algorithm along with the Krzanowski & Lai cluster quality index, to select a small yet informative subset of genes. We applied mAP-KL on real microarray data, as well as on simulated data, and compared its performance against 13 other feature selection approaches. Across a variety of diseases and number of samples, mAP-KL presents competitive classification results, particularly in neuromuscular diseases, where its overall AUC score was 0.91. Furthermore, mAP-KL generates concise yet biologically relevant and informative N-gene expression signatures, which can serve as a valuable tool for diagnostic and prognostic purposes, as well as a source of potential disease biomarkers in a broad range of diseases. Conclusions mAP-KL is a data-driven and classifier-independent hybrid feature selection method, which applies to any disease classification problem based on microarray data, regardless of the available samples. Combining multiple hypothesis testing and AP leads to subsets of genes, which classify unknown samples from both, small and large patient cohorts with high accuracy.

  6. A preliminary phylogenetic analysis of the New World Helopini (Coleoptera, Tenebrionidae, Tenebrioninae indicates the need for profound rearrangements of the classification

    Directory of Open Access Journals (Sweden)

    Paulina Cifuentes-Ruiz

    2014-06-01

    Full Text Available Helopini is a diverse tribe in the subfamily Tenebrioninae with a worldwide distribution. The New World helopine species have not been reviewed recently and several doubts emerge regarding their generic assignment as well as the naturalness of the tribe and subordinate taxa. To assess these questions, a preliminary cladistic analysis was conducted with emphasis on sampling the genera distributed in the New World, but including representatives from other regions. The parsimony analysis includes 30 ingroup species from America, Europe and Asia of the subtribes Helopina and Cylindrinotina, plus three outgroups, and 67 morphological characters. Construction of the matrix resulted in the discovery of morphological character states not previously reported for the tribe, particularly from the genitalia of New World species. A consensus of the 12 most parsimonious trees supports the monophyly of the tribe based on a unique combination of characters, including one synapomorphy. None of the subtribes or the genera of the New World represented by more than one species (Helops Fabricius, Nautes Pascoe and Tarpela Bates were recovered as monophyletic. Helopina was recovered as paraphyletic in relation to Cylindrinotina. One Nearctic species of Helops and one Palearctic species of Tarpela (subtribe Helopina were more closely related to species of Cylindrinotina. A relatively derived clade, mainly composed by Neotropical species, was found; it includes seven species of Tarpela, seven species of Nautes, and three species of Helops, two Nearctic and one Neotropical. Our results reveal the need to deeply re-evaluate the current classification of the tribe and subordinated taxa, but a broader taxon sampling and further character exploration is needed in order to fully recognize monophyletic groups at different taxonomic levels (from subtribes to genera.

  7. Accurate age classification of 6 and 12 month-old infants based on resting-state functional connectivity magnetic resonance imaging data.

    Science.gov (United States)

    Pruett, John R; Kandala, Sridhar; Hoertel, Sarah; Snyder, Abraham Z; Elison, Jed T; Nishino, Tomoyuki; Feczko, Eric; Dosenbach, Nico U F; Nardos, Binyam; Power, Jonathan D; Adeyemo, Babatunde; Botteron, Kelly N; McKinstry, Robert C; Evans, Alan C; Hazlett, Heather C; Dager, Stephen R; Paterson, Sarah; Schultz, Robert T; Collins, D Louis; Fonov, Vladimir S; Styner, Martin; Gerig, Guido; Das, Samir; Kostopoulos, Penelope; Constantino, John N; Estes, Annette M; Petersen, Steven E; Schlaggar, Bradley L; Piven, Joseph

    2015-04-01

    Human large-scale functional brain networks are hypothesized to undergo significant changes over development. Little is known about these functional architectural changes, particularly during the second half of the first year of life. We used multivariate pattern classification of resting-state functional connectivity magnetic resonance imaging (fcMRI) data obtained in an on-going, multi-site, longitudinal study of brain and behavioral development to explore whether fcMRI data contained information sufficient to classify infant age. Analyses carefully account for the effects of fcMRI motion artifact. Support vector machines (SVMs) classified 6 versus 12 month-old infants (128 datasets) above chance based on fcMRI data alone. Results demonstrate significant changes in measures of brain functional organization that coincide with a special period of dramatic change in infant motor, cognitive, and social development. Explorations of the most different correlations used for SVM lead to two different interpretations about functional connections that support 6 versus 12-month age categorization. Copyright © 2015 The Authors. Published by Elsevier Ltd.. All rights reserved.

  8. Accurate age classification of 6 and 12 month-old infants based on resting-state functional connectivity magnetic resonance imaging data

    Directory of Open Access Journals (Sweden)

    John R. Pruett, Jr.

    2015-04-01

    Full Text Available Human large-scale functional brain networks are hypothesized to undergo significant changes over development. Little is known about these functional architectural changes, particularly during the second half of the first year of life. We used multivariate pattern classification of resting-state functional connectivity magnetic resonance imaging (fcMRI data obtained in an on-going, multi-site, longitudinal study of brain and behavioral development to explore whether fcMRI data contained information sufficient to classify infant age. Analyses carefully account for the effects of fcMRI motion artifact. Support vector machines (SVMs classified 6 versus 12 month-old infants (128 datasets above chance based on fcMRI data alone. Results demonstrate significant changes in measures of brain functional organization that coincide with a special period of dramatic change in infant motor, cognitive, and social development. Explorations of the most different correlations used for SVM lead to two different interpretations about functional connections that support 6 versus 12-month age categorization.

  9. Xenolog classification.

    Science.gov (United States)

    Darby, Charlotte A; Stolzer, Maureen; Ropp, Patrick J; Barker, Daniel; Durand, Dannie

    2017-03-01

    Orthology analysis is a fundamental tool in comparative genomics. Sophisticated methods have been developed to distinguish between orthologs and paralogs and to classify paralogs into subtypes depending on the duplication mechanism and timing, relative to speciation. However, no comparable framework exists for xenologs: gene pairs whose history, since their divergence, includes a horizontal transfer. Further, the diversity of gene pairs that meet this broad definition calls for classification of xenologs with similar properties into subtypes. We present a xenolog classification that uses phylogenetic reconciliation to assign each pair of genes to a class based on the event responsible for their divergence and the historical association between genes and species. Our classes distinguish between genes related through transfer alone and genes related through duplication and transfer. Further, they separate closely-related genes in distantly-related species from distantly-related genes in closely-related species. We present formal rules that assign gene pairs to specific xenolog classes, given a reconciled gene tree with an arbitrary number of duplications and transfers. These xenology classification rules have been implemented in software and tested on a collection of ∼13 000 prokaryotic gene families. In addition, we present a case study demonstrating the connection between xenolog classification and gene function prediction. The xenolog classification rules have been implemented in N otung 2.9, a freely available phylogenetic reconciliation software package. http://www.cs.cmu.edu/~durand/Notung . Gene trees are available at http://dx.doi.org/10.7488/ds/1503 . durand@cmu.edu. Supplementary data are available at Bioinformatics online.

  10. Stratification of co-evolving genomic groups using ranked phylogenetic profiles

    Directory of Open Access Journals (Sweden)

    Tsoka Sophia

    2009-10-01

    Full Text Available Abstract Background Previous methods of detecting the taxonomic origins of arbitrary sequence collections, with a significant impact to genome analysis and in particular metagenomics, have primarily focused on compositional features of genomes. The evolutionary patterns of phylogenetic distribution of genes or proteins, represented by phylogenetic profiles, provide an alternative approach for the detection of taxonomic origins, but typically suffer from low accuracy. Herein, we present rank-BLAST, a novel approach for the assignment of protein sequences into genomic groups of the same taxonomic origin, based on the ranking order of phylogenetic profiles of target genes or proteins across the reference database. Results The rank-BLAST approach is validated by computing the phylogenetic profiles of all sequences for five distinct microbial species of varying degrees of phylogenetic proximity, against a reference database of 243 fully sequenced genomes. The approach - a combination of sequence searches, statistical estimation and clustering - analyses the degree of sequence divergence between sets of protein sequences and allows the classification of protein sequences according to the species of origin with high accuracy, allowing taxonomic classification of 64% of the proteins studied. In most cases, a main cluster is detected, representing the corresponding species. Secondary, functionally distinct and species-specific clusters exhibit different patterns of phylogenetic distribution, thus flagging gene groups of interest. Detailed analyses of such cases are provided as examples. Conclusion Our results indicate that the rank-BLAST approach can capture the taxonomic origins of sequence collections in an accurate and efficient manner. The approach can be useful both for the analysis of genome evolution and the detection of species groups in metagenomics samples.

  11. Accurate Arabic Script Language/Dialect Classification

    Science.gov (United States)

    2014-01-01

    dialects. language identification, Arabic, dialect, natural language processing, machine learning 30 Stephen C. Tratz 301-394-1057Unclassified...Arabic, Farsi, Urdu), Cyrillic script (Bulgarian, Russian, Ukrainian), and Devanagari script ( Hindi , Marathi, Nepali). They use Mechanical Turk to...to 1, which can be a useful feature. The Java port of the LIBLINEAR (Fan et al., 2008) machine learning software package1 is used to train all our

  12. Phylogenetic relationships within and among Brassica species from ...

    African Journals Online (AJOL)

    Phylogenetic relationships within and among Brassica species from RAPD loci ... The genus Brassica comprises economically important oilseed and vegetable crops. ... genetic diversity for conservation, cultivar classification and molecular ...

  13. Phylogenetic and phytogeographical relationships in Maloideae (Rosaceae) based on morphological and anatomical characters

    NARCIS (Netherlands)

    Aldasoro, J.J.; Aedo, C.; Navarro, C.

    2005-01-01

    Phylogenetic relationships among 24 genera of Rosaceae subfam. Maloideae and Spiraeoideae are explored by means of a cladistic analysis; 16 morphological and anatomical characters were included in the analysis. Published suprageneric classifications and characters used in these classifications are

  14. Taxonomic update on proposed nomenclature and classification changes for bacteria of medical importance, 2016.

    Science.gov (United States)

    Janda, J Michael

    2017-02-13

    A key aspect of medical, public health, and diagnostic microbiology laboratories is the accurate identification and rapid reporting and communication to medical staff regarding patients with infectious agents of clinical importance. Microbial taxonomy in the age of molecular diagnostics and phylogenetics creates changes in taxonomy at a logarithmic rate further complicating this process. This update focuses on the description of new species and classification changes proposed in 2016.

  15. Advances in phylogenetic studies of Nematoda

    Institute of Scientific and Technical Information of China (English)

    2002-01-01

    Nematoda is a metazoan group with extremely high diversity only next to Insecta. Caenorhabditis elegans is now a favorable experimental model animal in modern developmental biology, genetics and genomics studies. However, the phylogeny of Nematoda and the phylogenetic position of the phylum within animal kingdom have long been in debate. Recent molecular phylogenetic studies gave great challenges to the traditional nematode classification. The new phylogenies not only placed the Nematoda in the Ecdysozoan and divided the phylum into five clades, but also provided new insights into animal molecular identification and phylogenetic biodiversity studies. The present paper reviews major progress and remaining problems in the current molecular phylogenetic studies of Nematoda, and prospects the developmental tendencies of this field.

  16. Accurate model selection of relaxed molecular clocks in bayesian phylogenetics.

    Science.gov (United States)

    Baele, Guy; Li, Wai Lok Sibon; Drummond, Alexei J; Suchard, Marc A; Lemey, Philippe

    2013-02-01

    Recent implementations of path sampling (PS) and stepping-stone sampling (SS) have been shown to outperform the harmonic mean estimator (HME) and a posterior simulation-based analog of Akaike's information criterion through Markov chain Monte Carlo (AICM), in bayesian model selection of demographic and molecular clock models. Almost simultaneously, a bayesian model averaging approach was developed that avoids conditioning on a single model but averages over a set of relaxed clock models. This approach returns estimates of the posterior probability of each clock model through which one can estimate the Bayes factor in favor of the maximum a posteriori (MAP) clock model; however, this Bayes factor estimate may suffer when the posterior probability of the MAP model approaches 1. Here, we compare these two recent developments with the HME, stabilized/smoothed HME (sHME), and AICM, using both synthetic and empirical data. Our comparison shows reassuringly that MAP identification and its Bayes factor provide similar performance to PS and SS and that these approaches considerably outperform HME, sHME, and AICM in selecting the correct underlying clock model. We also illustrate the importance of using proper priors on a large set of empirical data sets.

  17. Phylogenetic molecular function annotation

    Science.gov (United States)

    Engelhardt, Barbara E.; Jordan, Michael I.; Repo, Susanna T.; Brenner, Steven E.

    2009-07-01

    It is now easier to discover thousands of protein sequences in a new microbial genome than it is to biochemically characterize the specific activity of a single protein of unknown function. The molecular functions of protein sequences have typically been predicted using homology-based computational methods, which rely on the principle that homologous proteins share a similar function. However, some protein families include groups of proteins with different molecular functions. A phylogenetic approach for predicting molecular function (sometimes called "phylogenomics") is an effective means to predict protein molecular function. These methods incorporate functional evidence from all members of a family that have functional characterizations using the evolutionary history of the protein family to make robust predictions for the uncharacterized proteins. However, they are often difficult to apply on a genome-wide scale because of the time-consuming step of reconstructing the phylogenies of each protein to be annotated. Our automated approach for function annotation using phylogeny, the SIFTER (Statistical Inference of Function Through Evolutionary Relationships) methodology, uses a statistical graphical model to compute the probabilities of molecular functions for unannotated proteins. Our benchmark tests showed that SIFTER provides accurate functional predictions on various protein families, outperforming other available methods.

  18. [Analysis phylogenetic relationship of Gynostemma (Cucurbitaceae)].

    Science.gov (United States)

    Qin, Shuang-shuang; Li, Hai-tao; Wang, Zhou-yong; Cui, Zhan-hu; Yu, Li-ying

    2015-05-01

    The sequences of ITS, matK, rbcL and psbA-trnH of 9 Gynostemma species or variety including 38 samples were compared and analyzed by molecular phylogeny method. Hemsleya macrosperma was designated as outgroup. The MP and NJ phylogenetic tree of Gynostemma was built based on ITS sequence, the results of PAUP phylogenetic analysis showed the following results: (1) The eight individuals of G. pentaphyllum var. pentaphyllum were not supported as monophyletic in the strict consensus trees and NJ trees. (2) It is suspected whether G. longipes and G. laxum should be classified as the independent species. (3)The classification of subgenus units of Gynostemma plants is supported.

  19. Morphological phylogenetics of Bignoniaceae Juss.

    Directory of Open Access Journals (Sweden)

    Usama K. Abdel-Hameed

    2014-09-01

    Full Text Available The most recent classification of Bignoniaceae recognized seven tribes, Phylogenetic and monographic studies focusing on clades within Bignoniaceae had revised tribal and generic boundaries and species numbers for several groups, the portions of the family that remain most poorly known are the African and Asian groups. The goal of the present study is to identify the primary lineages of Bignoniaceae in Egypt based on macromorphological traits. A total of 25 species of Bignoniaceae in Egypt was included in this study (Table 1, along with Barleria cristata as outgroup. Parsimony analyses were conducted using the program NONA 1.6, preparation of data set matrices and phylogenetic tree editing were achieved in WinClada Software. The obtained cladogram showed that within the studied taxa of Bignoniaceae there was support for eight lineages. The present study revealed that the two studied species of Tabebuia showed a strong support for monophyly as well as Tecoma and Kigelia. It was revealed that Bignonia, Markhamia and Parmentiera are not monophyletic genera.

  20. Phylogenetic Trees From Sequences

    Science.gov (United States)

    Ryvkin, Paul; Wang, Li-San

    In this chapter, we review important concepts and approaches for phylogeny reconstruction from sequence data.We first cover some basic definitions and properties of phylogenetics, and briefly explain how scientists model sequence evolution and measure sequence divergence. We then discuss three major approaches for phylogenetic reconstruction: distance-based phylogenetic reconstruction, maximum parsimony, and maximum likelihood. In the third part of the chapter, we review how multiple phylogenies are compared by consensus methods and how to assess confidence using bootstrapping. At the end of the chapter are two sections that list popular software packages and additional reading.

  1. Taxonomic Identity Resolution of Highly Phylogenetically Related Strains and Selection of Phylogenetic Markers by Using Genome-Scale Methods: The Bacillus pumilus Group Case

    Science.gov (United States)

    Espariz, Martín; Zuljan, Federico A.; Esteban, Luis; Magni, Christian

    2016-01-01

    Bacillus pumilus group strains have been studied due their agronomic, biotechnological or pharmaceutical potential. Classifying strains of this taxonomic group at species level is a challenging procedure since it is composed of seven species that share among them over 99.5% of 16S rRNA gene identity. In this study, first, a whole-genome in silico approach was used to accurately demarcate B. pumilus group strains, as a case of highly phylogenetically related taxa, at the species level. In order to achieve that and consequently to validate or correct taxonomic identities of genomes in public databases, an average nucleotide identity correlation, a core-based phylogenomic and a gene function repertory analyses were performed. Eventually, more than 50% such genomes were found to be misclassified. Hierarchical clustering of gene functional repertoires was also used to infer ecotypes among B. pumilus group species. Furthermore, for the first time the machine-learning algorithm Random Forest was used to rank genes in order of their importance for species classification. We found that ybbP, a gene involved in the synthesis of cyclic di-AMP, was the most important gene for accurately predicting species identity among B. pumilus group strains. Finally, principal component analysis was used to classify strains based on the distances between their ybbP genes. The methodologies described could be utilized more broadly to identify other highly phylogenetically related species in metagenomic or epidemiological assessments. PMID:27658251

  2. Phylogenetic reconstruction of the wolf spiders (Araneae: Lycosidae) using sequences from the 12S rRNA, 28S rRNA, and NADH1 genes: implications for classification, biogeography, and the evolution of web building behavior.

    Science.gov (United States)

    Murphy, Nicholas P; Framenau, Volker W; Donnellan, Stephen C; Harvey, Mark S; Park, Yung-Chul; Austin, Andrew D

    2006-03-01

    Current knowledge of the evolutionary relationships amongst the wolf spiders (Araneae: Lycosidae) is based on assessment of morphological similarity or phylogenetic analysis of a small number of taxa. In order to enhance the current understanding of lycosid relationships, phylogenies of 70 lycosid species were reconstructed by parsimony and Bayesian methods using three molecular markers; the mitochondrial genes 12S rRNA, NADH1, and the nuclear gene 28S rRNA. The resultant trees from the mitochondrial markers were used to assess the current taxonomic status of the Lycosidae and to assess the evolutionary history of sheet-web construction in the group. The results suggest that a number of genera are not monophyletic, including Lycosa, Arctosa, Alopecosa, and Artoria. At the subfamilial level, the status of Pardosinae needs to be re-assessed, and the position of a number of genera within their respective subfamilies is in doubt (e.g., Hippasa and Arctosa in Lycosinae and Xerolycosa, Aulonia and Hygrolycosa in Venoniinae). In addition, a major clade of strictly Australasian taxa may require the creation of a new subfamily. The analysis of sheet-web building in Lycosidae revealed that the interpretation of this trait as an ancestral state relies on two factors: (1) an asymmetrical model favoring the loss of sheet-webs and (2) that the suspended silken tube of Pirata is directly descended from sheet-web building. Paralogous copies of the nuclear 28S rRNA gene were sequenced, confounding the interpretation of the phylogenetic analysis and suggesting that a cautionary approach should be taken to the further use of this gene for lycosid phylogenetic analysis.

  3. Phylogenetic lineages in Entomophthoromycota

    NARCIS (Netherlands)

    Gryganskyi, A.P.; Humber, R.A.; Smith, M.E.; Hodge, K.; Huang, B.; Voigt, K.; Vilgalys, R.

    2013-01-01

    Entomophthoromycota is one of six major phylogenetic lineages among the former phylum Zygomycota. These early terrestrial fungi share evolutionarily ancestral characters such as coenocytic mycelium and gametangiogamy as a sexual process resulting in zygospore formation. Previous molecular studies ha

  4. A new higher classification of planarian flatworms (Platyhelminthes, Tricladida)

    NARCIS (Netherlands)

    Sluys, R.; Kawakatsu, M .; Riutort, M.; Baguñà, J.

    2009-01-01

    This paper presents a revised classification for the higher taxa within the Tricladida. A historical sketch is provided of the higher classificatory systems of triclad flatworms. As far as possible, the new classification is based on published phylogenetic studies. A phylogenetic tree generalizing c

  5. Cyber infrastructure for Fusarium: three integrated platforms supporting strain identification, phylogenetics, comparative genomics and knowledge sharing.

    Science.gov (United States)

    Park, Bongsoo; Park, Jongsun; Cheong, Kyeong-Chae; Choi, Jaeyoung; Jung, Kyongyong; Kim, Donghan; Lee, Yong-Hwan; Ward, Todd J; O'Donnell, Kerry; Geiser, David M; Kang, Seogchan

    2011-01-01

    The fungal genus Fusarium includes many plant and/or animal pathogenic species and produces diverse toxins. Although accurate species identification is critical for managing such threats, it is difficult to identify Fusarium morphologically. Fortunately, extensive molecular phylogenetic studies, founded on well-preserved culture collections, have established a robust foundation for Fusarium classification. Genomes of four Fusarium species have been published with more being currently sequenced. The Cyber infrastructure for Fusarium (CiF; http://www.fusariumdb.org/) was built to support archiving and utilization of rapidly increasing data and knowledge and consists of Fusarium-ID, Fusarium Comparative Genomics Platform (FCGP) and Fusarium Community Platform (FCP). The Fusarium-ID archives phylogenetic marker sequences from most known species along with information associated with characterized isolates and supports strain identification and phylogenetic analyses. The FCGP currently archives five genomes from four species. Besides supporting genome browsing and analysis, the FCGP presents computed characteristics of multiple gene families and functional groups. The Cart/Favorite function allows users to collect sequences from Fusarium-ID and the FCGP and analyze them later using multiple tools without requiring repeated copying-and-pasting of sequences. The FCP is designed to serve as an online community forum for sharing and preserving accumulated experience and knowledge to support future research and education.

  6. Cyber infrastructure for Fusarium: three integrated platforms supporting strain identification, phylogenetics, comparative genomics and knowledge sharing

    Science.gov (United States)

    Park, Bongsoo; Park, Jongsun; Cheong, Kyeong-Chae; Choi, Jaeyoung; Jung, Kyongyong; Kim, Donghan; Lee, Yong-Hwan; Ward, Todd J.; O'Donnell, Kerry; Geiser, David M.; Kang, Seogchan

    2011-01-01

    The fungal genus Fusarium includes many plant and/or animal pathogenic species and produces diverse toxins. Although accurate species identification is critical for managing such threats, it is difficult to identify Fusarium morphologically. Fortunately, extensive molecular phylogenetic studies, founded on well-preserved culture collections, have established a robust foundation for Fusarium classification. Genomes of four Fusarium species have been published with more being currently sequenced. The Cyber infrastructure for Fusarium (CiF; http://www.fusariumdb.org/) was built to support archiving and utilization of rapidly increasing data and knowledge and consists of Fusarium-ID, Fusarium Comparative Genomics Platform (FCGP) and Fusarium Community Platform (FCP). The Fusarium-ID archives phylogenetic marker sequences from most known species along with information associated with characterized isolates and supports strain identification and phylogenetic analyses. The FCGP currently archives five genomes from four species. Besides supporting genome browsing and analysis, the FCGP presents computed characteristics of multiple gene families and functional groups. The Cart/Favorite function allows users to collect sequences from Fusarium-ID and the FCGP and analyze them later using multiple tools without requiring repeated copying-and-pasting of sequences. The FCP is designed to serve as an online community forum for sharing and preserving accumulated experience and knowledge to support future research and education. PMID:21087991

  7. Rounding up the usual suspects: a standard target-gene approach for resolving the interfamilial phylogenetic relationships of ecribellate orb-weaving spiders with a new family-rank classification (Araneae, Araneoidea)

    DEFF Research Database (Denmark)

    Dimitrov, Dimitar; Benevidas, Ligia R.; Arnedo, Miquel A.;

    2017-01-01

    We test the limits of the spider superfamily Araneoidea and reconstruct its interfamilial relationships using standard molecular markers. The taxon sample (363 terminals) comprises for the first time representatives of all araneoid families, including the first molecular data of the family...... Synaphridae. We use the resulting phylogenetic framework to study web evolution in araneoids. Araneoidea is monophyletic and sister to Nicodamoidea rank. n. Orbiculariae are not monophyletic and also include the RTA clade, Oecobiidae and Hersiliidae. Deinopoidea is paraphyletic with respect to a lineage...... holarchaeids but the family remains diphyletic even if Holarchaea is considered an anapid. The orb-web is ancient, having evolved by the early Jurassic; a single origin of the orb with multiple “losses” is implied by our analyses. By the late Jurassic, the orb-web had already been transformed into different...

  8. Phylogenetically resolving epidemiologic linkage.

    Science.gov (United States)

    Romero-Severson, Ethan O; Bulla, Ingo; Leitner, Thomas

    2016-03-08

    Although the use of phylogenetic trees in epidemiological investigations has become commonplace, their epidemiological interpretation has not been systematically evaluated. Here, we use an HIV-1 within-host coalescent model to probabilistically evaluate transmission histories of two epidemiologically linked hosts. Previous critique of phylogenetic reconstruction has claimed that direction of transmission is difficult to infer, and that the existence of unsampled intermediary links or common sources can never be excluded. The phylogenetic relationship between the HIV populations of epidemiologically linked hosts can be classified into six types of trees, based on cladistic relationships and whether the reconstruction is consistent with the true transmission history or not. We show that the direction of transmission and whether unsampled intermediary links or common sources existed make very different predictions about expected phylogenetic relationships: (i) Direction of transmission can often be established when paraphyly exists, (ii) intermediary links can be excluded when multiple lineages were transmitted, and (iii) when the sampled individuals' HIV populations both are monophyletic a common source was likely the origin. Inconsistent results, suggesting the wrong transmission direction, were generally rare. In addition, the expected tree topology also depends on the number of transmitted lineages, the sample size, the time of the sample relative to transmission, and how fast the diversity increases after infection. Typically, 20 or more sequences per subject give robust results. We confirm our theoretical evaluations with analyses of real transmission histories and discuss how our findings should aid in interpreting phylogenetic results.

  9. Clustering with phylogenetic tools in astrophysics

    CERN Document Server

    Fraix-Burnet, Didier

    2016-01-01

    Phylogenetic approaches are finding more and more applications outside the field of biology. Astrophysics is no exception since an overwhelming amount of multivariate data has appeared in the last twenty years or so. In particular, the diversification of galaxies throughout the evolution of the Universe quite naturally invokes phylogenetic approaches. We have demonstrated that Maximum Parsimony brings useful astrophysical results, and we now proceed toward the analyses of large datasets for galaxies. In this talk I present how we solve the major difficulties for this goal: the choice of the parameters, their discretization, and the analysis of a high number of objects with an unsupervised NP-hard classification technique like cladistics. 1. Introduction How do the galaxy form, and when? How did the galaxy evolve and transform themselves to create the diversity we observe? What are the progenitors to present-day galaxies? To answer these big questions, observations throughout the Universe and the physical mode...

  10. Host specificity and phylogenetic relationships of chicken and turkey parvoviruses

    Science.gov (United States)

    Previous reports indicate that the newly discovered chicken parvoviruses (ChPV) and turkey parvoviruses (TuPV) are very similar to each other, yet they represent different species within a new genus of Parvoviridae. Currently, strain classification is based on the phylogenetic analysis of a 561 bas...

  11. Phylogenetic placement of two species known only from resting spores

    DEFF Research Database (Denmark)

    Hajek, Ann E; Gryganskyi, Andrii; Bittner, Tonya;

    2016-01-01

    Molecular methods were used to determine the generic placement of two species of Entomophthorales known only from resting spores. Historically, these species would belong in the form-genus Tarichium, but this classification provides no information about phylogenetic relationships. Using DNA from...

  12. Charles Darwin, beetles and phylogenetics.

    Science.gov (United States)

    Beutel, Rolf G; Friedrich, Frank; Leschen, Richard A B

    2009-11-01

    Here, we review Charles Darwin's relation to beetles and developments in coleopteran systematics in the last two centuries. Darwin was an enthusiastic beetle collector. He used beetles to illustrate different evolutionary phenomena in his major works, and astonishingly, an entire sub-chapter is dedicated to beetles in "The Descent of Man". During his voyage on the Beagle, Darwin was impressed by the high diversity of beetles in the tropics, and he remarked that, to his surprise, the majority of species were small and inconspicuous. However, despite his obvious interest in the group, he did not get involved in beetle taxonomy, and his theoretical work had little immediate impact on beetle classification. The development of taxonomy and classification in the late nineteenth and earlier twentieth century was mainly characterised by the exploration of new character systems (e.g. larval features and wing venation). In the mid-twentieth century, Hennig's new methodology to group lineages by derived characters revolutionised systematics of Coleoptera and other organisms. As envisioned by Darwin and Ernst Haeckel, the new Hennigian approach enabled systematists to establish classifications truly reflecting evolution. Roy A. Crowson and Howard E. Hinton, who both made tremendous contributions to coleopterology, had an ambivalent attitude towards the Hennigian ideas. The Mickoleit school combined detailed anatomical work with a classical Hennigian character evaluation, with stepwise tree building, comparatively few characters and a priori polarity assessment without explicit use of the outgroup comparison method. The rise of cladistic methods in the 1970s had a strong impact on beetle systematics. Cladistic computer programs facilitated parsimony analyses of large data matrices, mostly morphological characters not requiring detailed anatomical investigations. Molecular studies on beetle phylogeny started in the 1990s with modest taxon sampling and limited DNA data. This has

  13. Charles Darwin, beetles and phylogenetics

    Science.gov (United States)

    Beutel, Rolf G.; Friedrich, Frank; Leschen, Richard A. B.

    2009-11-01

    Here, we review Charles Darwin’s relation to beetles and developments in coleopteran systematics in the last two centuries. Darwin was an enthusiastic beetle collector. He used beetles to illustrate different evolutionary phenomena in his major works, and astonishingly, an entire sub-chapter is dedicated to beetles in “The Descent of Man”. During his voyage on the Beagle, Darwin was impressed by the high diversity of beetles in the tropics, and he remarked that, to his surprise, the majority of species were small and inconspicuous. However, despite his obvious interest in the group, he did not get involved in beetle taxonomy, and his theoretical work had little immediate impact on beetle classification. The development of taxonomy and classification in the late nineteenth and earlier twentieth century was mainly characterised by the exploration of new character systems (e.g. larval features and wing venation). In the mid-twentieth century, Hennig’s new methodology to group lineages by derived characters revolutionised systematics of Coleoptera and other organisms. As envisioned by Darwin and Ernst Haeckel, the new Hennigian approach enabled systematists to establish classifications truly reflecting evolution. Roy A. Crowson and Howard E. Hinton, who both made tremendous contributions to coleopterology, had an ambivalent attitude towards the Hennigian ideas. The Mickoleit school combined detailed anatomical work with a classical Hennigian character evaluation, with stepwise tree building, comparatively few characters and a priori polarity assessment without explicit use of the outgroup comparison method. The rise of cladistic methods in the 1970s had a strong impact on beetle systematics. Cladistic computer programs facilitated parsimony analyses of large data matrices, mostly morphological characters not requiring detailed anatomical investigations. Molecular studies on beetle phylogeny started in the 1990s with modest taxon sampling and limited DNA data

  14. First phylogenetic analyses of galaxy evolution

    CERN Document Server

    Fraix-Burnet, D

    2004-01-01

    The Hubble tuning fork diagram, based on morphology, has always been the preferred scheme for classification of galaxies and is still the only one originally built from historical/evolutionary relationships. At the opposite, biologists have long taken into account the parenthood links of living entities for classification purposes. Assuming branching evolution of galaxies as a "descent with modification", we show that the concepts and tools of phylogenetic systematics widely used in biology can be heuristically transposed to the case of galaxies. This approach that we call "astrocladistics" has been first applied to Dwarf Galaxies of the Local Group and provides the first evolutionary galaxy tree. The cladogram is sufficiently solid to support the existence of a hierarchical organization in the diversity of galaxies, making it possible to track ancestral types of galaxies. We also find that morphology is a summary of more fundamental properties. Astrocladistics applied to cosmology simulated galaxies can, uns...

  15. Phylogenetic molecular function annotation

    OpenAIRE

    Engelhardt, Barbara E.; Jordan, Michael I.; Repo, Susanna T; Brenner, Steven E.

    2009-01-01

    It is now easier to discover thousands of protein sequences in a new microbial genome than it is to biochemically characterize the specific activity of a single protein of unknown function. The molecular functions of protein sequences have typically been predicted using homology-based computational methods, which rely on the principle that homologous proteins share a similar function. However, some protein families include groups of proteins with different molecular functions. A phylogenetic ...

  16. Statistical Methods in Phylogenetic and Evolutionary Inferences

    Directory of Open Access Journals (Sweden)

    Luigi Bertolotti

    2013-05-01

    Full Text Available Molecular instruments are the most accurate methods in organisms’identification and characterization. Biologists are often involved in studies where the main goal is to identify relationships among individuals. In this framework, it is very important to know and apply the most robust approaches to infer correctly these relationships, allowing the right conclusions about phylogeny. In this review, we will introduce the reader to the most used statistical methods in phylogenetic analyses, the Maximum Likelihood and the Bayesian approaches, considering for simplicity only analyses regardingDNA sequences. Several studieswill be showed as examples in order to demonstrate how the correct phylogenetic inference can lead the scientists to highlight very peculiar features in pathogens biology and evolution.

  17. ClassyFlu: classification of influenza A viruses with Discriminatively trained profile-HMMs.

    Directory of Open Access Journals (Sweden)

    Sandra Van der Auwera

    Full Text Available Accurate and rapid characterization of influenza A virus (IAV hemagglutinin (HA and neuraminidase (NA sequences with respect to subtype and clade is at the basis of extended diagnostic services and implicit to molecular epidemiologic studies. ClassyFlu is a new tool and web service for the classification of IAV sequences of the HA and NA gene into subtypes and phylogenetic clades using discriminatively trained profile hidden Markov models (HMMs, one for each subtype or clade. ClassyFlu merely requires as input unaligned, full-length or partial HA or NA DNA sequences. It enables rapid and highly accurate assignment of HA sequences to subtypes H1-H17 but particularly focusses on the finer grained assignment of sequences of highly pathogenic avian influenza viruses of subtype H5N1 according to the cladistics proposed by the H5N1 Evolution Working Group. NA sequences are classified into subtypes N1-N10. ClassyFlu was compared to semiautomatic classification approaches using BLAST and phylogenetics and additionally for H5 sequences to the new "Highly Pathogenic H5N1 Clade Classification Tool" (IRD-CT proposed by the Influenza Research Database. Our results show that both web tools (ClassyFlu and IRD-CT, although based on different methods, are nearly equivalent in performance and both are more accurate and faster than semiautomatic classification. A retraining of ClassyFlu to altered cladistics as well as an extension of ClassyFlu to other IAV genome segments or fragments thereof is undemanding. This is exemplified by unambiguous assignment to a distinct cluster within subtype H7 of sequences of H7N9 viruses which emerged in China early in 2013 and caused more than 130 human infections. http://bioinf.uni-greifswald.de/ClassyFlu is a free web service. For local execution, the ClassyFlu source code in PERL is freely available.

  18. ClassyFlu: classification of influenza A viruses with Discriminatively trained profile-HMMs.

    Science.gov (United States)

    Van der Auwera, Sandra; Bulla, Ingo; Ziller, Mario; Pohlmann, Anne; Harder, Timm; Stanke, Mario

    2014-01-01

    Accurate and rapid characterization of influenza A virus (IAV) hemagglutinin (HA) and neuraminidase (NA) sequences with respect to subtype and clade is at the basis of extended diagnostic services and implicit to molecular epidemiologic studies. ClassyFlu is a new tool and web service for the classification of IAV sequences of the HA and NA gene into subtypes and phylogenetic clades using discriminatively trained profile hidden Markov models (HMMs), one for each subtype or clade. ClassyFlu merely requires as input unaligned, full-length or partial HA or NA DNA sequences. It enables rapid and highly accurate assignment of HA sequences to subtypes H1-H17 but particularly focusses on the finer grained assignment of sequences of highly pathogenic avian influenza viruses of subtype H5N1 according to the cladistics proposed by the H5N1 Evolution Working Group. NA sequences are classified into subtypes N1-N10. ClassyFlu was compared to semiautomatic classification approaches using BLAST and phylogenetics and additionally for H5 sequences to the new "Highly Pathogenic H5N1 Clade Classification Tool" (IRD-CT) proposed by the Influenza Research Database. Our results show that both web tools (ClassyFlu and IRD-CT), although based on different methods, are nearly equivalent in performance and both are more accurate and faster than semiautomatic classification. A retraining of ClassyFlu to altered cladistics as well as an extension of ClassyFlu to other IAV genome segments or fragments thereof is undemanding. This is exemplified by unambiguous assignment to a distinct cluster within subtype H7 of sequences of H7N9 viruses which emerged in China early in 2013 and caused more than 130 human infections. http://bioinf.uni-greifswald.de/ClassyFlu is a free web service. For local execution, the ClassyFlu source code in PERL is freely available.

  19. Fast phylogenetic DNA barcoding

    DEFF Research Database (Denmark)

    Terkelsen, Kasper Munch; Boomsma, Wouter Krogh; Willerslev, Eske

    2008-01-01

    We present a heuristic approach to the DNA assignment problem based on phylogenetic inferences using constrained neighbour joining and non-parametric bootstrapping. We show that this method performs as well as the more computationally intensive full Bayesian approach in an analysis of 500 insect...... DNA sequences obtained from GenBank. We also analyse a previously published dataset of environmental DNA sequences from soil from New Zealand and Siberia, and use these data to illustrate the fact that statistical approaches to the DNA assignment problem allow for more appropriate criteria...... for determining the taxonomic level at which a particular DNA sequence can be assigned....

  20. Associations of Leaf Spectra with Genetic and Phylogenetic Variation in Oaks: Prospects for Remote Detection of Biodiversity

    Directory of Open Access Journals (Sweden)

    Jeannine Cavender-Bares

    2016-03-01

    Full Text Available Species and phylogenetic lineages have evolved to differ in the way that they acquire and deploy resources, with consequences for their physiological, chemical and structural attributes, many of which can be detected using spectral reflectance form leaves. Recent technological advances for assessing optical properties of plants offer opportunities to detect functional traits of organisms and differentiate levels of biological organization across the tree of life. Here, we connect leaf-level full range spectral data (400–2400 nm of leaves to the hierarchical organization of plant diversity within the oak genus (Quercus using field and greenhouse experiments in which environmental factors and plant age are controlled. We show that spectral data significantly differentiate populations within a species and that spectral similarity is significantly associated with phylogenetic similarity among species. We further show that hyperspectral information allows more accurate classification of taxa than spectrally-derived traits, which by definition are of lower dimensionality. Finally, model accuracy increases at higher levels in the hierarchical organization of plant diversity, such that we are able to better distinguish clades than species or populations. This pattern supports an evolutionary explanation for the degree of optical differentiation among plants and demonstrates potential for remote detection of genetic and phylogenetic diversity.

  1. CREST--classification resources for environmental sequence tags.

    Directory of Open Access Journals (Sweden)

    Anders Lanzén

    Full Text Available Sequencing of taxonomic or phylogenetic markers is becoming a fast and efficient method for studying environmental microbial communities. This has resulted in a steadily growing collection of marker sequences, most notably of the small-subunit (SSU ribosomal RNA gene, and an increased understanding of microbial phylogeny, diversity and community composition patterns. However, to utilize these large datasets together with new sequencing technologies, a reliable and flexible system for taxonomic classification is critical. We developed CREST (Classification Resources for Environmental Sequence Tags, a set of resources and tools for generating and utilizing custom taxonomies and reference datasets for classification of environmental sequences. CREST uses an alignment-based classification method with the lowest common ancestor algorithm. It also uses explicit rank similarity criteria to reduce false positives and identify novel taxa. We implemented this method in a web server, a command line tool and the graphical user interfaced program MEGAN. Further, we provide the SSU rRNA reference database and taxonomy SilvaMod, derived from the publicly available SILVA SSURef, for classification of sequences from bacteria, archaea and eukaryotes. Using cross-validation and environmental datasets, we compared the performance of CREST and SilvaMod to the RDP Classifier. We also utilized Greengenes as a reference database, both with CREST and the RDP Classifier. These analyses indicate that CREST performs better than alignment-free methods with higher recall rate (sensitivity as well as precision, and with the ability to accurately identify most sequences from novel taxa. Classification using SilvaMod performed better than with Greengenes, particularly when applied to environmental sequences. CREST is freely available under a GNU General Public License (v3 from http://apps.cbu.uib.no/crest and http://lcaclassifier.googlecode.com.

  2. Phylogenetic trees in bioinformatics

    Energy Technology Data Exchange (ETDEWEB)

    Burr, Tom L [Los Alamos National Laboratory

    2008-01-01

    Genetic data is often used to infer evolutionary relationships among a collection of viruses, bacteria, animal or plant species, or other operational taxonomic units (OTU). A phylogenetic tree depicts such relationships and provides a visual representation of the estimated branching order of the OTUs. Tree estimation is unique for several reasons, including: the types of data used to represent each OTU; the use ofprobabilistic nucleotide substitution models; the inference goals involving both tree topology and branch length, and the huge number of possible trees for a given sample of a very modest number of OTUs, which implies that fmding the best tree(s) to describe the genetic data for each OTU is computationally demanding. Bioinformatics is too large a field to review here. We focus on that aspect of bioinformatics that includes study of similarities in genetic data from multiple OTUs. Although research questions are diverse, a common underlying challenge is to estimate the evolutionary history of the OTUs. Therefore, this paper reviews the role of phylogenetic tree estimation in bioinformatics, available methods and software, and identifies areas for additional research and development.

  3. Entanglement, Invariants, and Phylogenetics

    Science.gov (United States)

    Sumner, J. G.

    2007-10-01

    This thesis develops and expands upon known techniques of mathematical physics relevant to the analysis of the popular Markov model of phylogenetic trees required in biology to reconstruct the evolutionary relationships of taxonomic units from biomolecular sequence data. The techniques of mathematical physics are plethora and have been developed for some time. The Markov model of phylogenetics and its analysis is a relatively new technique where most progress to date has been achieved by using discrete mathematics. This thesis takes a group theoretical approach to the problem by beginning with a remarkable mathematical parallel to the process of scattering in particle physics. This is shown to equate to branching events in the evolutionary history of molecular units. The major technical result of this thesis is the derivation of existence proofs and computational techniques for calculating polynomial group invariant functions on a multi-linear space where the group action is that relevant to a Markovian time evolution. The practical results of this thesis are an extended analysis of the use of invariant functions in distance based methods and the presentation of a new reconstruction technique for quartet trees which is consistent with the most general Markov model of sequence evolution.

  4. Update on diabetes classification.

    Science.gov (United States)

    Thomas, Celeste C; Philipson, Louis H

    2015-01-01

    This article highlights the difficulties in creating a definitive classification of diabetes mellitus in the absence of a complete understanding of the pathogenesis of the major forms. This brief review shows the evolving nature of the classification of diabetes mellitus. No classification scheme is ideal, and all have some overlap and inconsistencies. The only diabetes in which it is possible to accurately diagnose by DNA sequencing, monogenic diabetes, remains undiagnosed in more than 90% of the individuals who have diabetes caused by one of the known gene mutations. The point of classification, or taxonomy, of disease, should be to give insight into both pathogenesis and treatment. It remains a source of frustration that all schemes of diabetes mellitus continue to fall short of this goal.

  5. The Phylogenetic Diversity of Metagenomes

    Science.gov (United States)

    Kembel, Steven W.; Eisen, Jonathan A.; Pollard, Katherine S.; Green, Jessica L.

    2011-01-01

    Phylogenetic diversity—patterns of phylogenetic relatedness among organisms in ecological communities—provides important insights into the mechanisms underlying community assembly. Studies that measure phylogenetic diversity in microbial communities have primarily been limited to a single marker gene approach, using the small subunit of the rRNA gene (SSU-rRNA) to quantify phylogenetic relationships among microbial taxa. In this study, we present an approach for inferring phylogenetic relationships among microorganisms based on the random metagenomic sequencing of DNA fragments. To overcome challenges caused by the fragmentary nature of metagenomic data, we leveraged fully sequenced bacterial genomes as a scaffold to enable inference of phylogenetic relationships among metagenomic sequences from multiple phylogenetic marker gene families. The resulting metagenomic phylogeny can be used to quantify the phylogenetic diversity of microbial communities based on metagenomic data sets. We applied this method to understand patterns of microbial phylogenetic diversity and community assembly along an oceanic depth gradient, and compared our findings to previous studies of this gradient using SSU-rRNA gene and metagenomic analyses. Bacterial phylogenetic diversity was highest at intermediate depths beneath the ocean surface, whereas taxonomic diversity (diversity measured by binning sequences into taxonomically similar groups) showed no relationship with depth. Phylogenetic diversity estimates based on the SSU-rRNA gene and the multi-gene metagenomic phylogeny were broadly concordant, suggesting that our approach will be applicable to other metagenomic data sets for which corresponding SSU-rRNA gene sequences are unavailable. Our approach opens up the possibility of using metagenomic data to study microbial diversity in a phylogenetic context. PMID:21912589

  6. Angle′s Molar Classification Revisited

    Directory of Open Access Journals (Sweden)

    Devanshi Yadav

    2014-01-01

    Results: Of the 500 pretreatment study casts assessed 52.4% were definitive Class I, 23.6% were Class II, 2.6% were Class III and the ambiguous cases were 21%. These could be easily classified with our method of classification. Conclusion: This improvised classification technique will help orthodontists in making classification of malocclusion accurate and simple.

  7. Molecular Phylogenetic: Organism Taxonomy Method Based on Evolution History

    Directory of Open Access Journals (Sweden)

    N.L.P Indi Dharmayanti

    2011-03-01

    Full Text Available Phylogenetic is described as taxonomy classification of an organism based on its evolution history namely its phylogeny and as a part of systematic science that has objective to determine phylogeny of organism according to its characteristic. Phylogenetic analysis from amino acid and protein usually became important area in sequence analysis. Phylogenetic analysis can be used to follow the rapid change of a species such as virus. The phylogenetic evolution tree is a two dimensional of a species graphic that shows relationship among organisms or particularly among their gene sequences. The sequence separation are referred as taxa (singular taxon that is defined as phylogenetically distinct units on the tree. The tree consists of outer branches or leaves that represents taxa and nodes and branch represent correlation among taxa. When the nucleotide sequence from two different organism are similar, they were inferred to be descended from common ancestor. There were three methods which were used in phylogenetic, namely (1 Maximum parsimony, (2 Distance, and (3 Maximum likehoood. Those methods generally are applied to construct the evolutionary tree or the best tree for determine sequence variation in group. Every method is usually used for different analysis and data.

  8. Incompletely resolved phylogenetic trees inflate estimates of phylogenetic conservatism.

    Science.gov (United States)

    Davies, T Jonathan; Kraft, Nathan J B; Salamin, Nicolas; Wolkovich, Elizabeth M

    2012-02-01

    The tendency for more closely related species to share similar traits and ecological strategies can be explained by their longer shared evolutionary histories and represents phylogenetic conservatism. How strongly species traits co-vary with phylogeny can significantly impact how we analyze cross-species data and can influence our interpretation of assembly rules in the rapidly expanding field of community phylogenetics. Phylogenetic conservatism is typically quantified by analyzing the distribution of species values on the phylogenetic tree that connects them. Many phylogenetic approaches, however, assume a completely sampled phylogeny: while we have good estimates of deeper phylogenetic relationships for many species-rich groups, such as birds and flowering plants, we often lack information on more recent interspecific relationships (i.e., within a genus). A common solution has been to represent these relationships as polytomies on trees using taxonomy as a guide. Here we show that such trees can dramatically inflate estimates of phylogenetic conservatism quantified using S. P. Blomberg et al.'s K statistic. Using simulations, we show that even randomly generated traits can appear to be phylogenetically conserved on poorly resolved trees. We provide a simple rarefaction-based solution that can reliably retrieve unbiased estimates of K, and we illustrate our method using data on first flowering times from Thoreau's woods (Concord, Massachusetts, USA).

  9. On Nakhleh's metric for reduced phylogenetic networks.

    Science.gov (United States)

    Cardona, Gabriel; Llabrés, Mercè; Rosselló, Francesc; Valiente, Gabriel

    2009-01-01

    We prove that Nakhleh's metric for reduced phylogenetic networks is also a metric on the classes of tree-child phylogenetic networks, semibinary tree-sibling time consistent phylogenetic networks, and multilabeled phylogenetic trees. We also prove that it separates distinguishable phylogenetic networks. In this way, it becomes the strongest dissimilarity measure for phylogenetic networks available so far. Furthermore, we propose a generalization of that metric that separates arbitrary phylogenetic networks.

  10. Ultrafast Approximation for Phylogenetic Bootstrap

    NARCIS (Netherlands)

    Bui Quang Minh, [No Value; Nguyen, Thi; von Haeseler, Arndt

    2013-01-01

    Nonparametric bootstrap has been a widely used tool in phylogenetic analysis to assess the clade support of phylogenetic trees. However, with the rapidly growing amount of data, this task remains a computational bottleneck. Recently, approximation methods such as the RAxML rapid bootstrap (RBS) and

  11. Bayesian phylogenetic estimation of fossil ages

    Science.gov (United States)

    Drummond, Alexei J.; Stadler, Tanja

    2016-01-01

    Recent advances have allowed for both morphological fossil evidence and molecular sequences to be integrated into a single combined inference of divergence dates under the rule of Bayesian probability. In particular, the fossilized birth–death tree prior and the Lewis-Mk model of discrete morphological evolution allow for the estimation of both divergence times and phylogenetic relationships between fossil and extant taxa. We exploit this statistical framework to investigate the internal consistency of these models by producing phylogenetic estimates of the age of each fossil in turn, within two rich and well-characterized datasets of fossil and extant species (penguins and canids). We find that the estimation accuracy of fossil ages is generally high with credible intervals seldom excluding the true age and median relative error in the two datasets of 5.7% and 13.2%, respectively. The median relative standard error (RSD) was 9.2% and 7.2%, respectively, suggesting good precision, although with some outliers. In fact, in the two datasets we analyse, the phylogenetic estimate of fossil age is on average less than 2 Myr from the mid-point age of the geological strata from which it was excavated. The high level of internal consistency found in our analyses suggests that the Bayesian statistical model employed is an adequate fit for both the geological and morphological data, and provides evidence from real data that the framework used can accurately model the evolution of discrete morphological traits coded from fossil and extant taxa. We anticipate that this approach will have diverse applications beyond divergence time dating, including dating fossils that are temporally unconstrained, testing of the ‘morphological clock', and for uncovering potential model misspecification and/or data errors when controversial phylogenetic hypotheses are obtained based on combined divergence dating analyses. This article is part of the themed issue ‘Dating species divergences

  12. Dengue virus type 3 in Brazil: a phylogenetic perspective

    Directory of Open Access Journals (Sweden)

    Josélio Maria Galvão de Araújo

    2009-05-01

    Full Text Available Circulation of a new dengue virus (DENV-3 genotype was recently described in Brazil and Colombia, but the precise classification of this genotype has been controversial. Here we perform phylogenetic and nucleotide-distance analyses of the envelope gene, which support the subdivision of DENV-3 strains into five distinct genotypes (GI to GV and confirm the classification of the new South American genotype as GV. The extremely low genetic distances between Brazilian GV strains and the prototype Philippines/L11423 GV strain isolated in 1956 raise important questions regarding the origin of GV in South America.

  13. Quartets and unrooted phylogenetic networks.

    Science.gov (United States)

    Gambette, Philippe; Berry, Vincent; Paul, Christophe

    2012-08-01

    Phylogenetic networks were introduced to describe evolution in the presence of exchanges of genetic material between coexisting species or individuals. Split networks in particular were introduced as a special kind of abstract network to visualize conflicts between phylogenetic trees which may correspond to such exchanges. More recently, methods were designed to reconstruct explicit phylogenetic networks (whose vertices can be interpreted as biological events) from triplet data. In this article, we link abstract and explicit networks through their combinatorial properties, by introducing the unrooted analog of level-k networks. In particular, we give an equivalence theorem between circular split systems and unrooted level-1 networks. We also show how to adapt to quartets some existing results on triplets, in order to reconstruct unrooted level-k phylogenetic networks. These results give an interesting perspective on the combinatorics of phylogenetic networks and also raise algorithmic and combinatorial questions.

  14. A phylogenetic analysis of the myxobacteria: basis for their classification

    Science.gov (United States)

    Shimkets, L.; Woese, C. R.

    1992-01-01

    The primary sequence and secondary structural features of the 16S rRNA were compared for 12 different myxobacteria representing all the known cultivated genera. Analysis of these data show the myxobacteria to form a monophyletic grouping consisting of three distinct families, which lies within the delta subdivision of the purple bacterial phylum. The composition of the families is consistent with differences in cell and spore morphology, cell behavior, and pigment and secondary metabolite production but is not correlated with the morphological complexity of the fruiting bodies. The Nannocystis exedens lineage has evolved at an unusually rapid pace and its rRNA shows numerous primary and secondary structural idiosyncrasies.

  15. Three Domains, Not Five Kingdoms: A Phylogenetic Classification System.

    Science.gov (United States)

    Peirce, Susan K.

    1999-01-01

    Argues that the Woesian three domain view of life should replace the five kingdom taxonomic scheme presented in most general biology texts and courses. Presents evidence for employing the three domain scheme and a related activity for classroom use. Contains 11 references. (WRM)

  16. A phylogenetic analysis of the myxobacteria: basis for their classification

    Science.gov (United States)

    Shimkets, L.; Woese, C. R.

    1992-01-01

    The primary sequence and secondary structural features of the 16S rRNA were compared for 12 different myxobacteria representing all the known cultivated genera. Analysis of these data show the myxobacteria to form a monophyletic grouping consisting of three distinct families, which lies within the delta subdivision of the purple bacterial phylum. The composition of the families is consistent with differences in cell and spore morphology, cell behavior, and pigment and secondary metabolite production but is not correlated with the morphological complexity of the fruiting bodies. The Nannocystis exedens lineage has evolved at an unusually rapid pace and its rRNA shows numerous primary and secondary structural idiosyncrasies.

  17. Application of Data Mining in Protein Sequence Classification

    Directory of Open Access Journals (Sweden)

    Suprativ Saha

    2012-11-01

    Full Text Available Protein sequence classification involves feature selection for accurate classification. Popular protein sequence classification techniques involve extraction of specific features from the sequences. Researchers apply some well-known classification techniques like neural networks, Genetic algorithm, Fuzzy ARTMAP,Rough Set Classifier etc for accurate classification. This paper presents a review is with three different classification models such as neural network model, fuzzy ARTMAP model and Rough set classifier model.This is followed by a new technique for classifying protein sequences. The proposed model is typicallyimplemented with an own designed tool and tries to reduce the computational overheads encountered by earlier approaches and increase the accuracy of classification.

  18. Phylogenetic and biogeographic analysis of sphaerexochine trilobites.

    Directory of Open Access Journals (Sweden)

    Curtis R Congreve

    Full Text Available BACKGROUND: Sphaerexochinae is a speciose and widely distributed group of cheirurid trilobites. Their temporal range extends from the earliest Ordovician through the Silurian, and they survived the end Ordovician mass extinction event (the second largest mass extinction in Earth history. Prior to this study, the individual evolutionary relationships within the group had yet to be determined utilizing rigorous phylogenetic methods. Understanding these evolutionary relationships is important for producing a stable classification of the group, and will be useful in elucidating the effects the end Ordovician mass extinction had on the evolutionary and biogeographic history of the group. METHODOLOGY/PRINCIPAL FINDINGS: Cladistic parsimony analysis of cheirurid trilobites assigned to the subfamily Sphaerexochinae was conducted to evaluate phylogenetic patterns and produce a hypothesis of relationship for the group. This study utilized the program TNT, and the analysis included thirty-one taxa and thirty-nine characters. The results of this analysis were then used in a Lieberman-modified Brooks Parsimony Analysis to analyze biogeographic patterns during the Ordovician-Silurian. CONCLUSIONS/SIGNIFICANCE: The genus Sphaerexochus was found to be monophyletic, consisting of two smaller clades (one composed entirely of Ordovician species and another composed of Silurian and Ordovician species. By contrast, the genus Kawina was found to be paraphyletic. It is a basal grade that also contains taxa formerly assigned to Cydonocephalus. Phylogenetic patterns suggest Sphaerexochinae is a relatively distinctive trilobite clade because it appears to have been largely unaffected by the end Ordovician mass extinction. Finally, the biogeographic analysis yields two major conclusions about Sphaerexochus biogeography: Bohemia and Avalonia were close enough during the Silurian to exchange taxa; and during the Ordovician there was dispersal between Eastern Laurentia and

  19. High-resolution phylogenetic microbial community profiling

    Energy Technology Data Exchange (ETDEWEB)

    Singer, Esther; Coleman-Derr, Devin; Bowman, Brett; Schwientek, Patrick; Clum, Alicia; Copeland, Alex; Ciobanu, Doina; Cheng, Jan-Fang; Gies, Esther; Hallam, Steve; Tringe, Susannah; Woyke, Tanja

    2014-03-17

    The representation of bacterial and archaeal genome sequences is strongly biased towards cultivated organisms, which belong to merely four phylogenetic groups. Functional information and inter-phylum level relationships are still largely underexplored for candidate phyla, which are often referred to as microbial dark matter. Furthermore, a large portion of the 16S rRNA gene records in the GenBank database are labeled as environmental samples and unclassified, which is in part due to low read accuracy, potential chimeric sequences produced during PCR amplifications and the low resolution of short amplicons. In order to improve the phylogenetic classification of novel species and advance our knowledge of the ecosystem function of uncultivated microorganisms, high-throughput full length 16S rRNA gene sequencing methodologies with reduced biases are needed. We evaluated the performance of PacBio single-molecule real-time (SMRT) sequencing in high-resolution phylogenetic microbial community profiling. For this purpose, we compared PacBio and Illumina metagenomic shotgun and 16S rRNA gene sequencing of a mock community as well as of an environmental sample from Sakinaw Lake, British Columbia. Sakinaw Lake is known to contain a large age of microbial species from candidate phyla. Sequencing results show that community structure based on PacBio shotgun and 16S rRNA gene sequences is highly similar in both the mock and the environmental communities. Resolution power and community representation accuracy from SMRT sequencing data appeared to be independent of GC content of microbial genomes and was higher when compared to Illumina-based metagenome shotgun and 16S rRNA gene (iTag) sequences, e.g. full-length sequencing resolved all 23 OTUs in the mock community, while iTags did not resolve closely related species. SMRT sequencing hence offers various potential benefits when characterizing uncharted microbial communities.

  20. Speaking Fluently And Accurately

    Institute of Scientific and Technical Information of China (English)

    JosephDeVeto

    2004-01-01

    Even after many years of study,students make frequent mistakes in English. In addition, many students still need a long time to think of what they want to say. For some reason, in spite of all the studying, students are still not quite fluent.When I teach, I use one technique that helps students not only speak more accurately, but also more fluently. That technique is dictations.

  1. Phylogenetics and the human microbiome.

    Science.gov (United States)

    Matsen, Frederick A

    2015-01-01

    The human microbiome is the ensemble of genes in the microbes that live inside and on the surface of humans. Because microbial sequencing information is now much easier to come by than phenotypic information, there has been an explosion of sequencing and genetic analysis of microbiome samples. Much of the analytical work for these sequences involves phylogenetics, at least indirectly, but methodology has developed in a somewhat different direction than for other applications of phylogenetics. In this article, I review the field and its methods from the perspective of a phylogeneticist, as well as describing current challenges for phylogenetics coming from this type of work.

  2. [Foundations of the new phylogenetics].

    Science.gov (United States)

    Pavlinov, I Ia

    2004-01-01

    Evolutionary idea is the core of the modern biology. Due to this, phylogenetics dealing with historical reconstructions in biology takes a priority position among biological disciplines. The second half of the 20th century witnessed growth of a great interest to phylogenetic reconstructions at macrotaxonomic level which replaced microevolutionary studies dominating during the 30s-60s. This meant shift from population thinking to phylogenetic one but it was not revival of the classical phylogenetics; rather, a new approach emerged that was baptized The New Phylogenetics. It arose as a result of merging of three disciplines which were developing independently during 60s-70s, namely cladistics, numerical phyletics, and molecular phylogenetics (now basically genophyletics). Thus, the new phylogenetics could be defined as a branch of evolutionary biology aimed at elaboration of "parsimonious" cladistic hypotheses by means of numerical methods on the basis of mostly molecular data. Classical phylogenetics, as a historical predecessor of the new one, emerged on the basis of the naturphilosophical worldview which included a superorganismal idea of biota. Accordingly to that view, historical development (the phylogeny) was thought an analogy of individual one (the ontogeny) so its most basical features were progressive parallel developments of "parts" (taxa), supplemented with Darwinian concept of monophyly. Two predominating traditions were diverged within classical phylogenetics according to a particular interpretation of relation between these concepts. One of them (Cope, Severtzow) belittled monophyly and paid most attention to progressive parallel developments of morphological traits. Such an attitude turned this kind of phylogenetics to be rather the semogenetics dealing primarily with evolution of structures and not of taxa. Another tradition (Haeckel) considered both monophyletic and parallel origins of taxa jointly: in the middle of 20th century it was split into

  3. An Optimization-Based Sampling Scheme for Phylogenetic Trees

    Science.gov (United States)

    Misra, Navodit; Blelloch, Guy; Ravi, R.; Schwartz, Russell

    Much modern work in phylogenetics depends on statistical sampling approaches to phylogeny construction to estimate probability distributions of possible trees for any given input data set. Our theoretical understanding of sampling approaches to phylogenetics remains far less developed than that for optimization approaches, however, particularly with regard to the number of sampling steps needed to produce accurate samples of tree partition functions. Despite the many advantages in principle of being able to sample trees from sophisticated probabilistic models, we have little theoretical basis for concluding that the prevailing sampling approaches do in fact yield accurate samples from those models within realistic numbers of steps. We propose a novel approach to phylogenetic sampling intended to be both efficient in practice and more amenable to theoretical analysis than the prevailing methods. The method depends on replacing the standard tree rearrangement moves with an alternative Markov model in which one solves a theoretically hard but practically tractable optimization problem on each step of sampling. The resulting method can be applied to a broad range of standard probability models, yielding practical algorithms for efficient sampling and rigorous proofs of accurate sampling for some important special cases. We demonstrate the efficiency and versatility of the method in an analysis of uncertainty in tree inference over varying input sizes. In addition to providing a new practical method for phylogenetic sampling, the technique is likely to prove applicable to many similar problems involving sampling over combinatorial objects weighted by a likelihood model.

  4. Semiparametric Gaussian copula classification

    OpenAIRE

    Zhao, Yue; Wegkamp, Marten

    2014-01-01

    This paper studies the binary classification of two distributions with the same Gaussian copula in high dimensions. Under this semiparametric Gaussian copula setting, we derive an accurate semiparametric estimator of the log density ratio, which leads to our empirical decision rule and a bound on its associated excess risk. Our estimation procedure takes advantage of the potential sparsity as well as the low noise condition in the problem, which allows us to achieve faster convergence rate of...

  5. Fast and accurate methods for phylogenomic analyses

    Directory of Open Access Journals (Sweden)

    Warnow Tandy

    2011-10-01

    Full Text Available Abstract Background Species phylogenies are not estimated directly, but rather through phylogenetic analyses of different gene datasets. However, true gene trees can differ from the true species tree (and hence from one another due to biological processes such as horizontal gene transfer, incomplete lineage sorting, and gene duplication and loss, so that no single gene tree is a reliable estimate of the species tree. Several methods have been developed to estimate species trees from estimated gene trees, differing according to the specific algorithmic technique used and the biological model used to explain differences between species and gene trees. Relatively little is known about the relative performance of these methods. Results We report on a study evaluating several different methods for estimating species trees from sequence datasets, simulating sequence evolution under a complex model including indels (insertions and deletions, substitutions, and incomplete lineage sorting. The most important finding of our study is that some fast and simple methods are nearly as accurate as the most accurate methods, which employ sophisticated statistical methods and are computationally quite intensive. We also observe that methods that explicitly consider errors in the estimated gene trees produce more accurate trees than methods that assume the estimated gene trees are correct. Conclusions Our study shows that highly accurate estimations of species trees are achievable, even when gene trees differ from each other and from the species tree, and that these estimations can be obtained using fairly simple and computationally tractable methods.

  6. Skeletal Rigidity of Phylogenetic Trees

    CERN Document Server

    Cheng, Howard; Li, Brian; Risteski, Andrej

    2012-01-01

    Motivated by geometric origami and the straight skeleton construction, we outline a map between spaces of phylogenetic trees and spaces of planar polygons. The limitations of this map is studied through explicit examples, culminating in proving a structural rigidity result.

  7. Contextual classification of multispectral image data: Approximate algorithm

    Science.gov (United States)

    Tilton, J. C. (Principal Investigator)

    1980-01-01

    An approximation to a classification algorithm incorporating spatial context information in a general, statistical manner is presented which is computationally less intensive. Classifications that are nearly as accurate are produced.

  8. Accelerating metagenomic read classification on CUDA-enabled GPUs

    National Research Council Canada - National Science Library

    Kobus, Robin; Hundt, Christian; Müller, André; Schmidt, Bertil

    2017-01-01

    ... metagenomic read classification are urgently needed. Results We present cuCLARK, a read-level classifier for CUDA-enabled GPUs, based on the fast and accurate classification of metagenomic sequences using reduced k-mers (CLARK) method...

  9. Quantum Simulation of Phylogenetic Trees

    CERN Document Server

    Ellinas, Demosthenes

    2011-01-01

    Quantum simulations constructing probability tensors of biological multi-taxa in phylogenetic trees are proposed, in terms of positive trace preserving maps, describing evolving systems of quantum walks with multiple walkers. Basic phylogenetic models applying on trees of various topologies are simulated following appropriate decoherent quantum circuits. Quantum simulations of statistical inference for aligned sequences of biological characters are provided in terms of a quantum pruning map operating on likelihood operator observables, utilizing state-observable duality and measurement theory.

  10. Community Phylogenetics: Assessing Tree Reconstruction Methods and the Utility of DNA Barcodes.

    Science.gov (United States)

    Boyle, Elizabeth E; Adamowicz, Sarah J

    2015-01-01

    Studies examining phylogenetic community structure have become increasingly prevalent, yet little attention has been given to the influence of the input phylogeny on metrics that describe phylogenetic patterns of co-occurrence. Here, we examine the influence of branch length, tree reconstruction method, and amount of sequence data on measures of phylogenetic community structure, as well as the phylogenetic signal (Pagel's λ) in morphological traits, using Trichoptera larval communities from Churchill, Manitoba, Canada. We find that model-based tree reconstruction methods and the use of a backbone family-level phylogeny improve estimations of phylogenetic community structure. In addition, trees built using the barcode region of cytochrome c oxidase subunit I (COI) alone accurately predict metrics of phylogenetic community structure obtained from a multi-gene phylogeny. Input tree did not alter overall conclusions drawn for phylogenetic signal, as significant phylogenetic structure was detected in two body size traits across input trees. As the discipline of community phylogenetics continues to expand, it is important to investigate the best approaches to accurately estimate patterns. Our results suggest that emerging large datasets of DNA barcode sequences provide a vast resource for studying the structure of biological communities.

  11. Integrated classification of inflammatory myopathies.

    Science.gov (United States)

    Allenbach, Y; Benveniste, O; Goebel, H-H; Stenzel, W

    2017-02-01

    Inflammatory myopathies comprise a multitude of diverse diseases, most often occurring in complex clinical settings. To ensure accurate diagnosis, multidisciplinary expertise is required. Here, we propose a comprehensive myositis classification that incorporates clinical, morphological and molecular data as well as autoantibody profile. This review focuses on recent advances in myositis research, in particular, the correlation between autoantibodies and morphological or clinical phenotypes that can be used as the basis for an 'integrated' classification system.

  12. Phylogenetic incongruence in E. coli O104: understanding the evolutionary relationships of emerging pathogens in the face of homologous recombination.

    Directory of Open Access Journals (Sweden)

    Weilong Hao

    Full Text Available Escherichia coli O104:H4 was identified as an emerging pathogen during the spring and summer of 2011 and was responsible for a widespread outbreak that resulted in the deaths of 50 people and sickened over 4075. Traditional phenotypic and genotypic assays, such as serotyping, pulsed field gel electrophoresis (PFGE, and multilocus sequence typing (MLST, permit identification and classification of bacterial pathogens, but cannot accurately resolve relationships among genotypically similar but pathotypically different isolates. To understand the evolutionary origins of E. coli O104:H4, we sequenced two strains isolated in Ontario, Canada. One was epidemiologically linked to the 2011 outbreak, and the second, unrelated isolate, was obtained in 2010. MLST analysis indicated that both isolates are of the same sequence type (ST678, but whole-genome sequencing revealed differences in chromosomal and plasmid content. Through comprehensive phylogenetic analysis of five O104:H4 ST678 genomes, we identified 167 genes in three gene clusters that have undergone homologous recombination with distantly related E. coli strains. These recombination events have resulted in unexpectedly high sequence diversity within the same sequence type. Failure to recognize or adjust for homologous recombination can result in phylogenetic incongruence. Understanding the extent of homologous recombination among different strains of the same sequence type may explain the pathotypic differences between the ON2010 and ON2011 strains and help shed new light on the emergence of this new pathogen.

  13. apex: phylogenetics with multiple genes.

    Science.gov (United States)

    Jombart, Thibaut; Archer, Frederick; Schliep, Klaus; Kamvar, Zhian; Harris, Rebecca; Paradis, Emmanuel; Goudet, Jérome; Lapp, Hilmar

    2017-01-01

    Genetic sequences of multiple genes are becoming increasingly common for a wide range of organisms including viruses, bacteria and eukaryotes. While such data may sometimes be treated as a single locus, in practice, a number of biological and statistical phenomena can lead to phylogenetic incongruence. In such cases, different loci should, at least as a preliminary step, be examined and analysed separately. The r software has become a popular platform for phylogenetics, with several packages implementing distance-based, parsimony and likelihood-based phylogenetic reconstruction, and an even greater number of packages implementing phylogenetic comparative methods. Unfortunately, basic data structures and tools for analysing multiple genes have so far been lacking, thereby limiting potential for investigating phylogenetic incongruence. In this study, we introduce the new r package apex to fill this gap. apex implements new object classes, which extend existing standards for storing DNA and amino acid sequences, and provides a number of convenient tools for handling, visualizing and analysing these data. In this study, we introduce the main features of the package and illustrate its functionalities through the analysis of a simple data set.

  14. BIOACCESSIBILITY TESTS ACCURATELY ESTIMATE ...

    Science.gov (United States)

    Hazards of soil-borne Pb to wild birds may be more accurately quantified if the bioavailability of that Pb is known. To better understand the bioavailability of Pb to birds, we measured blood Pb concentrations in Japanese quail (Coturnix japonica) fed diets containing Pb-contaminated soils. Relative bioavailabilities were expressed by comparison with blood Pb concentrations in quail fed a Pb acetate reference diet. Diets containing soil from five Pb-contaminated Superfund sites had relative bioavailabilities from 33%-63%, with a mean of about 50%. Treatment of two of the soils with P significantly reduced the bioavailability of Pb. The bioaccessibility of the Pb in the test soils was then measured in six in vitro tests and regressed on bioavailability. They were: the “Relative Bioavailability Leaching Procedure” (RBALP) at pH 1.5, the same test conducted at pH 2.5, the “Ohio State University In vitro Gastrointestinal” method (OSU IVG), the “Urban Soil Bioaccessible Lead Test”, the modified “Physiologically Based Extraction Test” and the “Waterfowl Physiologically Based Extraction Test.” All regressions had positive slopes. Based on criteria of slope and coefficient of determination, the RBALP pH 2.5 and OSU IVG tests performed very well. Speciation by X-ray absorption spectroscopy demonstrated that, on average, most of the Pb in the sampled soils was sorbed to minerals (30%), bound to organic matter 24%, or present as Pb sulfate 18%. Ad

  15. Bosque: integrated phylogenetic analysis software.

    Science.gov (United States)

    Ramírez-Flandes, Salvador; Ulloa, Osvaldo

    2008-11-01

    Phylogenetic analyses today involve dealing with computer files in different formats and often several computer programs. Although some widely used applications have integrated important functionalities for such analyses, they still work with local resources only: input/output files (users have to manage them) and local computing (users have sometimes to leave their programs, on their desktop computers, running for extended periods of time). To address these problems we have developed 'Bosque', a multi-platform client-server software that performs standard phylogenetic tasks either locally or remotely on servers, and integrates the results on a local relational database. Bosque performs sequence alignments and graphical visualization and editing of trees, thus providing a powerful environment that integrates all the steps of phylogenetic analyses. http://bosque.udec.cl

  16. Factors that affect large subunit ribosomal DNA amplicon sequencing studies of fungal communities: classification method, primer choice, and error.

    Directory of Open Access Journals (Sweden)

    Teresita M Porter

    Full Text Available Nuclear large subunit ribosomal DNA is widely used in fungal phylogenetics and to an increasing extent also amplicon-based environmental sequencing. The relatively short reads produced by next-generation sequencing, however, makes primer choice and sequence error important variables for obtaining accurate taxonomic classifications. In this simulation study we tested the performance of three classification methods: 1 a similarity-based method (BLAST + Metagenomic Analyzer, MEGAN; 2 a composition-based method (Ribosomal Database Project naïve bayesian classifier, NBC; and, 3 a phylogeny-based method (Statistical Assignment Package, SAP. We also tested the effects of sequence length, primer choice, and sequence error on classification accuracy and perceived community composition. Using a leave-one-out cross validation approach, results for classifications to the genus rank were as follows: BLAST + MEGAN had the lowest error rate and was particularly robust to sequence error; SAP accuracy was highest when long LSU query sequences were classified; and, NBC runs significantly faster than the other tested methods. All methods performed poorly with the shortest 50-100 bp sequences. Increasing simulated sequence error reduced classification accuracy. Community shifts were detected due to sequence error and primer selection even though there was no change in the underlying community composition. Short read datasets from individual primers, as well as pooled datasets, appear to only approximate the true community composition. We hope this work informs investigators of some of the factors that affect the quality and interpretation of their environmental gene surveys.

  17. Absolute Pitch in Boreal Chickadees and Humans: Exceptions that Test a Phylogenetic Rule

    Science.gov (United States)

    Weisman, Ronald G.; Balkwill, Laura-Lee; Hoeschele, Marisa; Moscicki, Michele K.; Bloomfield, Laurie L.; Sturdy, Christopher B.

    2010-01-01

    This research examined generality of the phylogenetic rule that birds discriminate frequency ranges more accurately than mammals. Human absolute pitch chroma possessors accurately tracked transitions between frequency ranges. Independent tests showed that they used note naming (pitch chroma) to remap the tones into ranges; neither possessors nor…

  18. On the nature of global classification

    Science.gov (United States)

    Wheelis, M. L.; Kandler, O.; Woese, C. R.

    1992-01-01

    Molecular sequencing technology has brought biology into the era of global (universal) classification. Methodologically and philosophically, global classification differs significantly from traditional, local classification. The need for uniformity requires that higher level taxa be defined on the molecular level in terms of universally homologous functions. A global classification should reflect both principal dimensions of the evolutionary process: genealogical relationship and quality and extent of divergence within a group. The ultimate purpose of a global classification is not simply information storage and retrieval; such a system should also function as an heuristic representation of the evolutionary paradigm that exerts a directing influence on the course of biology. The global system envisioned allows paraphyletic taxa. To retain maximal phylogenetic information in these cases, minor notational amendments in existing taxonomic conventions should be adopted.

  19. HYBRID INTERNET TRAFFIC CLASSIFICATION TECHNIQUE1

    Institute of Scientific and Technical Information of China (English)

    Li Jun; Zhang Shunyi; Lu Yanqing; Yan Junrong

    2009-01-01

    Accurate and real-time classification of network traffic is significant to network operation and management such as QoS differentiation, traffic shaping and security surveillance. However, with many newly emerged P2P applications using dynamic port numbers, masquerading techniques, and payload encryption to avoid detection, traditional classification approaches turn to be ineffective. In this paper, we present a layered hybrid system to classify current Internet traffic, motivated by variety of network activities and their requirements of traffic classification. The proposed method could achieve fast and accurate traffic classification with low overheads and robustness to accommodate both known and unknown/encrypted applications. Furthermore, it is feasible to be used in the context of real-time traffic classification. Our experimental results show the distinct advantages of the proposed classification system, compared with the one-step Machine Learning (ML) approach.

  20. Interpreting the universal phylogenetic tree

    Science.gov (United States)

    Woese, C. R.

    2000-01-01

    The universal phylogenetic tree not only spans all extant life, but its root and earliest branchings represent stages in the evolutionary process before modern cell types had come into being. The evolution of the cell is an interplay between vertically derived and horizontally acquired variation. Primitive cellular entities were necessarily simpler and more modular in design than are modern cells. Consequently, horizontal gene transfer early on was pervasive, dominating the evolutionary dynamic. The root of the universal phylogenetic tree represents the first stage in cellular evolution when the evolving cell became sufficiently integrated and stable to the erosive effects of horizontal gene transfer that true organismal lineages could exist.

  1. Phylogenetic analysis of a spontaneous cocoa bean fermentation metagenome reveals new insights into its bacterial and fungal community diversity.

    Directory of Open Access Journals (Sweden)

    Koen Illeghems

    Full Text Available This is the first report on the phylogenetic analysis of the community diversity of a single spontaneous cocoa bean box fermentation sample through a metagenomic approach involving 454 pyrosequencing. Several sequence-based and composition-based taxonomic profiling tools were used and evaluated to avoid software-dependent results and their outcome was validated by comparison with previously obtained culture-dependent and culture-independent data. Overall, this approach revealed a wider bacterial (mainly γ-Proteobacteria and fungal diversity than previously found. Further, the use of a combination of different classification methods, in a software-independent way, helped to understand the actual composition of the microbial ecosystem under study. In addition, bacteriophage-related sequences were found. The bacterial diversity depended partially on the methods used, as composition-based methods predicted a wider diversity than sequence-based methods, and as classification methods based solely on phylogenetic marker genes predicted a more restricted diversity compared with methods that took all reads into account. The metagenomic sequencing analysis identified Hanseniaspora uvarum, Hanseniaspora opuntiae, Saccharomyces cerevisiae, Lactobacillus fermentum, and Acetobacter pasteurianus as the prevailing species. Also, the presence of occasional members of the cocoa bean fermentation process was revealed (such as Erwinia tasmaniensis, Lactobacillus brevis, Lactobacillus casei, Lactobacillus rhamnosus, Lactococcus lactis, Leuconostoc mesenteroides, and Oenococcus oeni. Furthermore, the sequence reads associated with viral communities were of a restricted diversity, dominated by Myoviridae and Siphoviridae, and reflecting Lactobacillus as the dominant host. To conclude, an accurate overview of all members of a cocoa bean fermentation process sample was revealed, indicating the superiority of metagenomic sequencing over previously used techniques.

  2. Functional Basis of Microorganism Classification.

    Directory of Open Access Journals (Sweden)

    Chengsheng Zhu

    2015-08-01

    Full Text Available Correctly identifying nearest "neighbors" of a given microorganism is important in industrial and clinical applications where close relationships imply similar treatment. Microbial classification based on similarity of physiological and genetic organism traits (polyphasic similarity is experimentally difficult and, arguably, subjective. Evolutionary relatedness, inferred from phylogenetic markers, facilitates classification but does not guarantee functional identity between members of the same taxon or lack of similarity between different taxa. Using over thirteen hundred sequenced bacterial genomes, we built a novel function-based microorganism classification scheme, functional-repertoire similarity-based organism network (FuSiON; flattened to fusion. Our scheme is phenetic, based on a network of quantitatively defined organism relationships across the known prokaryotic space. It correlates significantly with the current taxonomy, but the observed discrepancies reveal both (1 the inconsistency of functional diversity levels among different taxa and (2 an (unsurprising bias towards prioritizing, for classification purposes, relatively minor traits of particular interest to humans. Our dynamic network-based organism classification is independent of the arbitrary pairwise organism similarity cut-offs traditionally applied to establish taxonomic identity. Instead, it reveals natural, functionally defined organism groupings and is thus robust in handling organism diversity. Additionally, fusion can use organism meta-data to highlight the specific environmental factors that drive microbial diversification. Our approach provides a complementary view to cladistic assignments and holds important clues for further exploration of microbial lifestyles. Fusion is a more practical fit for biomedical, industrial, and ecological applications, as many of these rely on understanding the functional capabilities of the microbes in their environment and are less

  3. A comparison of phylogenetic network methods using computer simulation.

    Directory of Open Access Journals (Sweden)

    Steven M Woolley

    Full Text Available BACKGROUND: We present a series of simulation studies that explore the relative performance of several phylogenetic network approaches (statistical parsimony, split decomposition, union of maximum parsimony trees, neighbor-net, simulated history recombination upper bound, median-joining, reduced median joining and minimum spanning network compared to standard tree approaches, (neighbor-joining and maximum parsimony in the presence and absence of recombination. PRINCIPAL FINDINGS: In the absence of recombination, all methods recovered the correct topology and branch lengths nearly all of the time when the substitution rate was low, except for minimum spanning networks, which did considerably worse. At a higher substitution rate, maximum parsimony and union of maximum parsimony trees were the most accurate. With recombination, the ability to infer the correct topology was halved for all methods and no method could accurately estimate branch lengths. CONCLUSIONS: Our results highlight the need for more accurate phylogenetic network methods and the importance of detecting and accounting for recombination in phylogenetic studies. Furthermore, we provide useful information for choosing a network algorithm and a framework in which to evaluate improvements to existing methods and novel algorithms developed in the future.

  4. Phylogenetic relationships among Maloideae species

    Science.gov (United States)

    The Maloideae is a highly diverse sub-family of the Rosaceae containing several agronomically important species (Malus sp. and Pyrus sp.) and their wild relatives. Previous phylogenetic work within the group has revealed extensive intergeneric hybridization and polyploidization. In order to develop...

  5. Phylogenetics of neotropical Platymiscium (Leguminosae

    DEFF Research Database (Denmark)

    Saslis-Lagoudakis, C. Haris; Chase, Mark W; Robinson, Daniel N

    2008-01-01

    Platymiscium is a neotropical legume genus of forest trees in the Pterocarpus clade of the pantropical "dalbergioid" clade. It comprises 19 species (29 taxa), distributed from Mexico to southern Brazil. This study presents a molecular phylogenetic analysis of Platymiscium and allies inferred from...

  6. PoInTree: A Polar and Interactive Phylogenetic Tree

    Institute of Scientific and Technical Information of China (English)

    Carreras Marco; Gianti Eleonora; Sartori Luca; Plyte Simon Edward; Isacchi Antonella; Bosotti Roberta

    2005-01-01

    PoInTree (Polar and Innteractive Tree) is an application that allows to build, visualize, and customize phylogenetic trees in a polar, interactive, and highly flexible view. It takes as input a FASTA file or multiple alignment formats. Phylogenetic tree calculation is based on a sequence distance method and utilizes the Neighbor Joining (NJ) algorithm. It also allows displaying precalculated trees of the major protein families based on Pfam classification. In PoInTree, nodes can be dynamically opened and closed and distances between genes are graphically represented.Tree root can be centered on a selected leaf. Text search mechanism, color-coding and labeling display are integrated. The visualizer can be connected to an Oracle database containing information on sequences and other biological data, helping to guide their interpretation within a given protein family across multiple species.The application is written in Borland Delphi and based on VCL Teechart Pro 6 graphical component (Steema software).

  7. Phylogenetic biogeography and taxonomy of disjunctly distributed bryophytes

    Institute of Scientific and Technical Information of China (English)

    Jochen HEINRICHS; J(o)rn HENTSCHEL; Kathrin FELDBERG; Andrea BOMBOSCH; Harald SCHNEIDER

    2009-01-01

    More than 200 research papers on the molecular phylogeny and phylogenetic biogeography of bryophytes have been published since the beginning of this millenium. These papers corroborated assumptions of a complex ge-netic structure of morphologically circumscribed bryophytes, and raised reservations against many morphologically justified species concepts, especially within the mosses. However, many molecular studies allowed for corrections and modifications of morphological classification schemes. Several studies reported that the phylogenetic structure of disjunctly distributed bryophyte species reflects their geographical ranges rather than morphological disparities. Molecular data led to new appraisals of distribution ranges and allowed for the reconstruction of refugia and migra-tion routes. Intercontinental ranges of bryophytes are often caused by dispersal rather than geographical vicariance. Many distribution patterns of disjunct bryophytes are likely formed by processes such as short distance dispersal, rare long distance dispersal events, extinction, recolonization and diversification.

  8. Groundwater recharge: Accurately representing evapotranspiration

    CSIR Research Space (South Africa)

    Bugan, Richard DH

    2011-09-01

    Full Text Available Groundwater recharge is the basis for accurate estimation of groundwater resources, for determining the modes of water allocation and groundwater resource susceptibility to climate change. Accurate estimations of groundwater recharge with models...

  9. Tissue Classification

    DEFF Research Database (Denmark)

    Van Leemput, Koen; Puonti, Oula

    2015-01-01

    Computational methods for automatically segmenting magnetic resonance images of the brain have seen tremendous advances in recent years. So-called tissue classification techniques, aimed at extracting the three main brain tissue classes (white matter, gray matter, and cerebrospinal fluid), are now...... well established. In their simplest form, these methods classify voxels independently based on their intensity alone, although much more sophisticated models are typically used in practice. This article aims to give an overview of often-used computational techniques for brain tissue classification...

  10. Text Classification using Artificial Intelligence

    CERN Document Server

    Kamruzzaman, S M

    2010-01-01

    Text classification is the process of classifying documents into predefined categories based on their content. It is the automated assignment of natural language texts to predefined categories. Text classification is the primary requirement of text retrieval systems, which retrieve texts in response to a user query, and text understanding systems, which transform text in some way such as producing summaries, answering questions or extracting data. Existing supervised learning algorithms for classifying text need sufficient documents to learn accurately. This paper presents a new algorithm for text classification using artificial intelligence technique that requires fewer documents for training. Instead of using words, word relation i.e. association rules from these words is used to derive feature set from pre-classified text documents. The concept of na\\"ive Bayes classifier is then used on derived features and finally only a single concept of genetic algorithm has been added for final classification. A syste...

  11. Text Classification using Data Mining

    CERN Document Server

    Kamruzzaman, S M; Hasan, Ahmed Ryadh

    2010-01-01

    Text classification is the process of classifying documents into predefined categories based on their content. It is the automated assignment of natural language texts to predefined categories. Text classification is the primary requirement of text retrieval systems, which retrieve texts in response to a user query, and text understanding systems, which transform text in some way such as producing summaries, answering questions or extracting data. Existing supervised learning algorithms to automatically classify text need sufficient documents to learn accurately. This paper presents a new algorithm for text classification using data mining that requires fewer documents for training. Instead of using words, word relation i.e. association rules from these words is used to derive feature set from pre-classified text documents. The concept of Naive Bayes classifier is then used on derived features and finally only a single concept of Genetic Algorithm has been added for final classification. A system based on the...

  12. Cyber-infrastructure for Fusarium (CiF): Three integrated platforms supporting strain identification, phylogenetics, comparative genomics, and knowledge sharing

    Science.gov (United States)

    The fungal genus Fusarium includes many plant and/or animal pathogenic species and produces diverse toxins. Although accurate identification is critical for managing such threats, it is difficult to identify Fusarium morphologically. Fortunately, extensive molecular phylogenetic studies, founded on ...

  13. Transporter Classification Database (TCDB)

    Data.gov (United States)

    U.S. Department of Health & Human Services — The Transporter Classification Database details a comprehensive classification system for membrane transport proteins known as the Transporter Classification (TC)...

  14. Molecular identification and phylogenetic study of Demodex caprae.

    Science.gov (United States)

    Zhao, Ya-E; Cheng, Juan; Hu, Li; Ma, Jun-Xian

    2014-10-01

    The DNA barcode has been widely used in species identification and phylogenetic analysis since 2003, but there have been no reports in Demodex. In this study, to obtain an appropriate DNA barcode for Demodex, molecular identification of Demodex caprae based on mitochondrial cox1 was conducted. Firstly, individual adults and eggs of D. caprae were obtained for genomic DNA (gDNA) extraction; Secondly, mitochondrial cox1 fragment was amplified, cloned, and sequenced; Thirdly, cox1 fragments of D. caprae were aligned with those of other Demodex retrieved from GenBank; Finally, the intra- and inter-specific divergences were computed and the phylogenetic trees were reconstructed to analyze phylogenetic relationship in Demodex. Results obtained from seven 429-bp fragments of D. caprae showed that sequence identities were above 99.1% among three adults and four eggs. The intraspecific divergences in D. caprae, Demodex folliculorum, Demodex brevis, and Demodex canis were 0.0-0.9, 0.5-0.9, 0.0-0.2, and 0.0-0.5%, respectively, while the interspecific divergences between D. caprae and D. folliculorum, D. canis, and D. brevis were 20.3-20.9, 21.8-23.0, and 25.0-25.3, respectively. The interspecific divergences were 10 times higher than intraspecific ones, indicating considerable barcoding gap. Furthermore, the phylogenetic trees showed that four Demodex species gathered separately, representing independent species; and Demodex folliculorum gathered with canine Demodex, D. caprae, and D. brevis in sequence. In conclusion, the selected 429-bp mitochondrial cox1 gene is an appropriate DNA barcode for molecular classification, identification, and phylogenetic analysis of Demodex. D. caprae is an independent species and D. folliculorum is closer to D. canis than to D. caprae or D. brevis.

  15. Global patterns of amphibian phylogenetic diversity

    DEFF Research Database (Denmark)

    Fritz, Susanne; Rahbek, Carsten

    2012-01-01

    phylogeny (2792 species). We combined each tree with global species distributions to map four indices of phylogenetic diversity. To investigate congruence between global spatial patterns of amphibian species richness and phylogenetic diversity, we selected Faith’s phylogenetic diversity (PD) index...... successfully colonized these archipelagos. Areas with unusually high phylogenetic diversity were located around biogeographic contact zones in Central America and southern China, and seem to have experienced high immigration or in situ diversification rates, combined with local persistence of old lineages...

  16. Phyx: phylogenetic tools for unix.

    Science.gov (United States)

    Brown, Joseph W; Walker, Joseph F; Smith, Stephen A

    2017-02-08

    The ease with which phylogenomic data can be generated has drastically escalated the computational burden for even routine phylogenetic investigations. To address this, we present phyx : a collection of programs written in C ++ to explore, manipulate, analyze and simulate phylogenetic objects (alignments, trees and MCMC logs). Modelled after Unix/GNU/Linux command line tools, individual programs perform a single task and operate on standard I/O streams that can be piped to quickly and easily form complex analytical pipelines. Because of the stream-centric paradigm, memory requirements are minimized (often only a single tree or sequence in memory at any instance), and hence phyx is capable of efficiently processing very large datasets. phyx runs on POSIX-compliant operating systems. Source code, installation instructions, documentation and example files are freely available under the GNU General Public License at https://github.com/FePhyFoFum/phyx. eebsmith@umich.edu. Supplementary data are available at Bioinformatics online.

  17. Stochastic Models for Phylogenetic Trees on Higher-order Taxa

    CERN Document Server

    Aldous, David; Popovic, Lea

    2007-01-01

    Simple stochastic models for phylogenetic trees on species have been well studied. But much paleontology data concerns time series or trees on higher-order taxa, and any broad picture of relationships between extant groups requires use of higher-order taxa. A coherent model for trees on (say) genera should involve both a species-level model and a model for the classification scheme by which species are assigned to genera. We present a general framework for such models, and describe three alternate classification schemes. Combining with the species-level model of Aldous-Popovic (2005), one gets models for higher-order trees, and we initiate analytic study of such models. In particular we derive formulas for the lifetime of genera, for the distribution of number of species per genus, and for the offspring structure of the tree on genera.

  18. Mitochondrial coi in phylogenetic relationships of Laimaphelenchus belgradiensis (nematoda: Aphelenchoididae

    Directory of Open Access Journals (Sweden)

    Oro Violeta

    2015-01-01

    Full Text Available Nematodes of the genus Laimaphelenchus are small and tiny organisms. Some parts of their body are measured in nanometers. The identification and classification of such organisms is a complex task. Previously, the major source of classification was morphology based on anatomical characters and measurements. Nowadays, this approach is supplemented by: “nano-morphology” based on scanning electron microscopy and molecular data and phylogeny, resulting in molecular systematics. Laimaphelenchus belgradiensis was recently described species. Since cytochrome c oxidase subunit I gene was successful in DNA based species diagnosis, it was chosen as a molecular marker to infer phylogeny of the newly discovered species. Phylogenetic relationships were based on Bayesian inference, the pairwise distances and the content of nitrogenous bases. The great genetic diversity was observed among close and distant species. [Projekat Ministarstva nauke Republike Srbije, br. TR 31018 i br. III 46007

  19. Vestige: Maximum likelihood phylogenetic footprinting

    Directory of Open Access Journals (Sweden)

    Maxwell Peter

    2005-05-01

    Full Text Available Abstract Background Phylogenetic footprinting is the identification of functional regions of DNA by their evolutionary conservation. This is achieved by comparing orthologous regions from multiple species and identifying the DNA regions that have diverged less than neutral DNA. Vestige is a phylogenetic footprinting package built on the PyEvolve toolkit that uses probabilistic molecular evolutionary modelling to represent aspects of sequence evolution, including the conventional divergence measure employed by other footprinting approaches. In addition to measuring the divergence, Vestige allows the expansion of the definition of a phylogenetic footprint to include variation in the distribution of any molecular evolutionary processes. This is achieved by displaying the distribution of model parameters that represent partitions of molecular evolutionary substitutions. Examination of the spatial incidence of these effects across regions of the genome can identify DNA segments that differ in the nature of the evolutionary process. Results Vestige was applied to a reference dataset of the SCL locus from four species and provided clear identification of the known conserved regions in this dataset. To demonstrate the flexibility to use diverse models of molecular evolution and dissect the nature of the evolutionary process Vestige was used to footprint the Ka/Ks ratio in primate BRCA1 with a codon model of evolution. Two regions of putative adaptive evolution were identified illustrating the ability of Vestige to represent the spatial distribution of distinct molecular evolutionary processes. Conclusion Vestige provides a flexible, open platform for phylogenetic footprinting. Underpinned by the PyEvolve toolkit, Vestige provides a framework for visualising the signatures of evolutionary processes across the genome of numerous organisms simultaneously. By exploiting the maximum-likelihood statistical framework, the complex interplay between mutational

  20. Phylogenetic analysis of otospiralin protein

    Science.gov (United States)

    Torktaz, Ibrahim; Behjati, Mohaddeseh; Rostami, Amin

    2016-01-01

    Background: Fibrocyte-specific protein, otospiralin, is a small protein, widely expressed in the central nervous system as neuronal cell bodies and glia. The increased expression of otospiralin in reactive astrocytes implicates its role in signaling pathways and reparative mechanisms subsequent to injury. Indeed, otospiralin is considered to be essential for the survival of fibrocytes of the mesenchymal nonsensory regions of the cochlea. It seems that other functions of this protein are not yet completely understood. Materials and Methods: Amino acid sequences of otospiralin from 12 vertebrates were derived from National Center for Biotechnology Information database. Phylogenetic analysis and phylogeny estimation were performed using MEGA 5.0.5 program, and neighbor-joining tree was constructed by this software. Results: In this computational study, the phylogenetic tree of otospiralin has been investigated. Therefore, dendrograms of otospiralin were depicted. Alignment performed in MUSCLE method by UPGMB algorithm. Also, entropy plot determined for a better illustration of amino acid variations in this protein. Conclusion: In the present study, we used otospiralin sequence of 12 different species and by constructing phylogenetic tree, we suggested out group for some related species. PMID:27099854

  1. HIV classification using coalescent theory

    Energy Technology Data Exchange (ETDEWEB)

    Zhang, Ming [Los Alamos National Laboratory; Letiner, Thomas K [Los Alamos National Laboratory; Korber, Bette T [Los Alamos National Laboratory

    2008-01-01

    Algorithms for subtype classification and breakpoint detection of HIV-I sequences are based on a classification system of HIV-l. Hence, their quality highly depend on this system. Due to the history of creation of the current HIV-I nomenclature, the current one contains inconsistencies like: The phylogenetic distance between the subtype B and D is remarkably small compared with other pairs of subtypes. In fact, it is more like the distance of a pair of subsubtypes Robertson et al. (2000); Subtypes E and I do not exist any more since they were discovered to be composed of recombinants Robertson et al. (2000); It is currently discussed whether -- instead of CRF02 being a recombinant of subtype A and G -- subtype G should be designated as a circulating recombination form (CRF) nd CRF02 as a subtype Abecasis et al. (2007); There are 8 complete and over 400 partial HIV genomes in the LANL-database which belong neither to a subtype nor to a CRF (denoted by U). Moreover, the current classification system is somehow arbitrary like all complex classification systems that were created manually. To this end, it is desirable to deduce the classification system of HIV systematically by an algorithm. Of course, this problem is not restricted to HIV, but applies to all fast mutating and recombining viruses. Our work addresses the simpler subproblem to score classifications of given input sequences of some virus species (classification denotes a partition of the input sequences in several subtypes and CRFs). To this end, we reconstruct ancestral recombination graphs (ARG) of the input sequences under restrictions determined by the given classification. These restritions are imposed in order to ensure that the reconstructed ARGs do not contradict the classification under consideration. Then, we find the ARG with maximal probability by means of Markov Chain Monte Carlo methods. The probability of the most probable ARG is interpreted as a score for the classification. To our

  2. Classification and Analysis of Computer Network Traffic

    DEFF Research Database (Denmark)

    Bujlow, Tomasz

    2014-01-01

    various classification modes (decision trees, rulesets, boosting, softening thresholds) regarding the classification accuracy and the time required to create the classifier. We showed how to use our VBS tool to obtain per-flow, per-application, and per-content statistics of traffic in computer networks...... classification (as by using transport layer port numbers, Deep Packet Inspection (DPI), statistical classification) and assessed their usefulness in particular areas. We found that the classification techniques based on port numbers are not accurate anymore as most applications use dynamic port numbers, while...... DPI is relatively slow, requires a lot of processing power, and causes a lot of privacy concerns. Statistical classifiers based on Machine Learning Algorithms (MLAs) were shown to be fast and accurate. At the same time, they do not consume a lot of resources and do not cause privacy concerns. However...

  3. Reanalysis and simulation suggest a phylogenetic microarray does not accurately profile microbial communities.

    Directory of Open Access Journals (Sweden)

    David J Midgley

    Full Text Available The second generation (G2 PhyloChip is designed to detect over 8700 bacteria and archaeal and has been used over 50 publications and conference presentations. Many of those publications reveal that the PhyloChip measures of species richness greatly exceed statistical estimates of richness based on other methods. An examination of probes downloaded from Greengenes suggested that the system may have the potential to distort the observed community structure. This may be due to the sharing of probes by taxa; more than 21% of the taxa in that downloaded data have no unique probes. In-silico simulations using these data showed that a population of 64 taxa representing a typical anaerobic subterranean community returned 96 different taxa, including 15 families incorrectly called present and 19 families incorrectly called absent. A study of nasal and oropharyngeal microbial communities by Lemon et al (2010 found some 1325 taxa using the G2 PhyloChip, however, about 950 of these taxa have, in the downloaded data, no unique probes and cannot be definitively called present. Finally, data from Brodie et al (2007, when re-examined, indicate that the abundance of the majority of detected taxa, are highly correlated with one another, suggesting that many probe sets do not act independently. Based on our analyses of downloaded data, we conclude that outputs from the G2 PhyloChip should be treated with some caution, and that the presence of taxa represented solely by non-unique probes be independently verified.

  4. Neuromuscular disease classification system

    Science.gov (United States)

    Sáez, Aurora; Acha, Begoña; Montero-Sánchez, Adoración; Rivas, Eloy; Escudero, Luis M.; Serrano, Carmen

    2013-06-01

    Diagnosis of neuromuscular diseases is based on subjective visual assessment of biopsies from patients by the pathologist specialist. A system for objective analysis and classification of muscular dystrophies and neurogenic atrophies through muscle biopsy images of fluorescence microscopy is presented. The procedure starts with an accurate segmentation of the muscle fibers using mathematical morphology and a watershed transform. A feature extraction step is carried out in two parts: 24 features that pathologists take into account to diagnose the diseases and 58 structural features that the human eye cannot see, based on the assumption that the biopsy is considered as a graph, where the nodes are represented by each fiber, and two nodes are connected if two fibers are adjacent. A feature selection using sequential forward selection and sequential backward selection methods, a classification using a Fuzzy ARTMAP neural network, and a study of grading the severity are performed on these two sets of features. A database consisting of 91 images was used: 71 images for the training step and 20 as the test. A classification error of 0% was obtained. It is concluded that the addition of features undetectable by the human visual inspection improves the categorization of atrophic patterns.

  5. A phylogenetic re-analysis of groupers with applications for ciguatera fish poisoning.

    Directory of Open Access Journals (Sweden)

    Charlotte Schoelinck

    Full Text Available Ciguatera fish poisoning (CFP is a significant public health problem due to dinoflagellates. It is responsible for one of the highest reported incidence of seafood-borne illness and Groupers are commonly reported as a source of CFP due to their position in the food chain. With the role of recent climate change on harmful algal blooms, CFP cases might become more frequent and more geographically widespread. Since there is no appropriate treatment for CFP, the most efficient solution is to regulate fish consumption. Such a strategy can only work if the fish sold are correctly identified, and it has been repeatedly shown that misidentifications and species substitutions occur in fish markets.We provide here both a DNA-barcoding reference for groupers, and a new phylogenetic reconstruction based on five genes and a comprehensive taxonomical sampling. We analyse the correlation between geographic range of species and their susceptibility to ciguatera accumulation, and the co-occurrence of ciguatoxins in closely related species, using both character mapping and statistical methods.Misidentifications were encountered in public databases, precluding accurate species identifications. Epinephelinae now includes only twelve genera (vs. 15 previously. Comparisons with the ciguatera incidences show that in some genera most species are ciguateric, but statistical tests display only a moderate correlation with the phylogeny. Atlantic species were rarely contaminated, with ciguatera occurrences being restricted to the South Pacific.The recent changes in classification based on the reanalyses of the relationships within Epinephelidae have an impact on the interpretation of the ciguatera distribution in the genera. In this context and to improve the monitoring of fish trade and safety, we need to obtain extensive data on contamination at the species level. Accurate species identifications through DNA barcoding are thus an essential tool in controlling CFP since

  6. HoxPred: automated classification of Hox proteins using combinations of generalised profiles

    Directory of Open Access Journals (Sweden)

    Leyns Luc

    2007-07-01

    Full Text Available Abstract Background Correct identification of individual Hox proteins is an essential basis for their study in diverse research fields. Common methods to classify Hox proteins focus on the homeodomain that characterise homeobox transcription factors. Classification is hampered by the high conservation of this short domain. Phylogenetic tree reconstruction is a widely used but time-consuming classification method. Results We have developed an automated procedure, HoxPred, that classifies Hox proteins in their groups of homology. The method relies on a discriminant analysis that classifies Hox proteins according to their scores for a combination of protein generalised profiles. 54 generalised profiles dedicated to each Hox homology group were produced de novo from a curated dataset of vertebrate Hox proteins. Several classification methods were investigated to select the most accurate discriminant functions. These functions were then incorporated into the HoxPred program. Conclusion HoxPred shows a mean accuracy of 97%. Predictions on the recently-sequenced stickleback fish proteome identified 44 Hox proteins, including HoxC1a only found so far in zebrafish. Using the Uniprot databank, we demonstrate that HoxPred can efficiently contribute to large-scale automatic annotation of Hox proteins into their paralogous groups. As orthologous group predictions show a higher risk of misclassification, they should be corroborated by additional supporting evidence. HoxPred is accessible via SOAP and Web interface http://cege.vub.ac.be/hoxpred/. Complete datasets, results and source code are available at the same site.

  7. Comparison of tree-child phylogenetic networks.

    Science.gov (United States)

    Cardona, Gabriel; Rosselló, Francesc; Valiente, Gabriel

    2009-01-01

    Phylogenetic networks are a generalization of phylogenetic trees that allow for the representation of nontreelike evolutionary events, like recombination, hybridization, or lateral gene transfer. While much progress has been made to find practical algorithms for reconstructing a phylogenetic network from a set of sequences, all attempts to endorse a class of phylogenetic networks (strictly extending the class of phylogenetic trees) with a well-founded distance measure have, to the best of our knowledge and with the only exception of the bipartition distance on regular networks, failed so far. In this paper, we present and study a new meaningful class of phylogenetic networks, called tree-child phylogenetic networks, and we provide an injective representation of these networks as multisets of vectors of natural numbers, their path multiplicity vectors. We then use this representation to define a distance on this class that extends the well-known Robinson-Foulds distance for phylogenetic trees and to give an alignment method for pairs of networks in this class. Simple polynomial algorithms for reconstructing a tree-child phylogenetic network from its path multiplicity vectors, for computing the distance between two tree-child phylogenetic networks and for aligning a pair of tree-child phylogenetic networks, are provided. They have been implemented as a Perl package and a Java applet, which can be found at http://bioinfo.uib.es/~recerca/phylonetworks/mudistance/.

  8. Functional and phylogenetic ecology in R

    CERN Document Server

    Swenson, Nathan G

    2014-01-01

    Functional and Phylogenetic Ecology in R is designed to teach readers to use R for phylogenetic and functional trait analyses. Over the past decade, a dizzying array of tools and methods were generated to incorporate phylogenetic and functional information into traditional ecological analyses. Increasingly these tools are implemented in R, thus greatly expanding their impact. Researchers getting started in R can use this volume as a step-by-step entryway into phylogenetic and functional analyses for ecology in R. More advanced users will be able to use this volume as a quick reference to understand particular analyses. The volume begins with an introduction to the R environment and handling relevant data in R. Chapters then cover phylogenetic and functional metrics of biodiversity; null modeling and randomizations for phylogenetic and functional trait analyses; integrating phylogenetic and functional trait information; and interfacing the R environment with a popular C-based program. This book presents a uni...

  9. Molecular phylogenetics of New World searobins (Triglidae; Prionotinae).

    Science.gov (United States)

    Portnoy, David S; Willis, Stuart C; Hunt, Elizabeth; Swift, Dominic G; Gold, John R; Conway, Kevin W

    2017-02-01

    Phylogenetic relationships among members of the New World searobin genera Bellator and Prionotus (Family Triglidae, Subfamily Prionotinae) and among other searobins in the families Triglidae and Peristediidae were investigated using both mitochondrial and nuclear DNA sequences. Phylogenetic hypotheses derived from maximum likelihood and Bayesian methodologies supported a monophyletic Prionotinae that included four well resolved clades of uncertain relationship; three contained species in the genus Prionotus and one contained species in the genus Bellator. Bellator was always recovered within the genus Prionotus, a result supported by post hoc model testing. Two nominal species of Prionotus (P. alatus and P. paralatus) were not recovered as exclusive lineages, suggesting the two may comprise a single species. Phylogenetic hypotheses also supported a monophyletic Triglidae but only if armored searobins (Family Peristediidae) were included. A robust morphological assessment is needed to further characterize relationships and suggest classification of clades within Prionotinae; for the time being we recommend that Bellator be considered a synonym of Prionotus. Relationships between armored searobins (Family Peristediidae) and searobins (Family Triglidae) and relationships within Triglidae also warrant further study.

  10. Making Mosquito Taxonomy Useful: A Stable Classification of Tribe Aedini that Balances Utility with Current Knowledge of Evolutionary Relationships.

    Science.gov (United States)

    Wilkerson, Richard C; Linton, Yvonne-Marie; Fonseca, Dina M; Schultz, Ted R; Price, Dana C; Strickman, Daniel A

    2015-01-01

    The tribe Aedini (Family Culicidae) contains approximately one-quarter of the known species of mosquitoes, including vectors of deadly or debilitating disease agents. This tribe contains the genus Aedes, which is one of the three most familiar genera of mosquitoes. During the past decade, Aedini has been the focus of a series of extensive morphology-based phylogenetic studies published by Reinert, Harbach, and Kitching (RH&K). Those authors created 74 new, elevated or resurrected genera from what had been the single genus Aedes, almost tripling the number of genera in the entire family Culicidae. The proposed classification is based on subjective assessments of the "number and nature of the characters that support the branches" subtending particular monophyletic groups in the results of cladistic analyses of a large set of morphological characters of representative species. To gauge the stability of RH&K's generic groupings we reanalyzed their data with unweighted parsimony jackknife and maximum-parsimony analyses, with and without ordering 14 of the characters as in RH&K. We found that their phylogeny was largely weakly supported and their taxonomic rankings failed priority and other useful taxon-naming criteria. Consequently, we propose simplified aedine generic designations that 1) restore a classification system that is useful for the operational community; 2) enhance the ability of taxonomists to accurately place new species into genera; 3) maintain the progress toward a natural classification based on monophyletic groups of species; and 4) correct the current classification system that is subject to instability as new species are described and existing species more thoroughly defined. We do not challenge the phylogenetic hypotheses generated by the above-mentioned series of morphological studies. However, we reduce the ranks of the genera and subgenera of RH&K to subgenera or informal species groups, respectively, to preserve stability as new data become

  11. Making Mosquito Taxonomy Useful: A Stable Classification of Tribe Aedini that Balances Utility with Current Knowledge of Evolutionary Relationships

    Science.gov (United States)

    Wilkerson, Richard C.; Linton, Yvonne-Marie; Fonseca, Dina M.; Schultz, Ted R.; Price, Dana C.; Strickman, Daniel A.

    2015-01-01

    The tribe Aedini (Family Culicidae) contains approximately one-quarter of the known species of mosquitoes, including vectors of deadly or debilitating disease agents. This tribe contains the genus Aedes, which is one of the three most familiar genera of mosquitoes. During the past decade, Aedini has been the focus of a series of extensive morphology-based phylogenetic studies published by Reinert, Harbach, and Kitching (RH&K). Those authors created 74 new, elevated or resurrected genera from what had been the single genus Aedes, almost tripling the number of genera in the entire family Culicidae. The proposed classification is based on subjective assessments of the “number and nature of the characters that support the branches” subtending particular monophyletic groups in the results of cladistic analyses of a large set of morphological characters of representative species. To gauge the stability of RH&K’s generic groupings we reanalyzed their data with unweighted parsimony jackknife and maximum-parsimony analyses, with and without ordering 14 of the characters as in RH&K. We found that their phylogeny was largely weakly supported and their taxonomic rankings failed priority and other useful taxon-naming criteria. Consequently, we propose simplified aedine generic designations that 1) restore a classification system that is useful for the operational community; 2) enhance the ability of taxonomists to accurately place new species into genera; 3) maintain the progress toward a natural classification based on monophyletic groups of species; and 4) correct the current classification system that is subject to instability as new species are described and existing species more thoroughly defined. We do not challenge the phylogenetic hypotheses generated by the above-mentioned series of morphological studies. However, we reduce the ranks of the genera and subgenera of RH&K to subgenera or informal species groups, respectively, to preserve stability as new data

  12. Phylogenetic trees and Euclidean embeddings.

    Science.gov (United States)

    Layer, Mark; Rhodes, John A

    2017-01-01

    It was recently observed by de Vienne et al. (Syst Biol 60(6):826-832, 2011) that a simple square root transformation of distances between taxa on a phylogenetic tree allowed for an embedding of the taxa into Euclidean space. While the justification for this was based on a diffusion model of continuous character evolution along the tree, here we give a direct and elementary explanation for it that provides substantial additional insight. We use this embedding to reinterpret the differences between the NJ and BIONJ tree building algorithms, providing one illustration of how this embedding reflects tree structures in data.

  13. 类群取样与系统发育分析精确度之探索%Taxon sampling and the accuracy of phylogenetic analyses

    Institute of Scientific and Technical Information of China (English)

    Tracy A. HEATH; Shannon M. HEDTKE; David M. HILLIS

    2008-01-01

    Appropriate and extensive taxon sampling is one of the most important determinants of accurate phylogenetic estimation. In addition, accuracy of inferences about evolutionary processes obtained from phylogenetic analyses is improved significantly by thorough taxon sampling efforts. Many recent efforts to improve phylogenetic estimates have focused instead on increasing sequence length or the number of overall characters in the analysis, and this often does have a beneficial effect on the accuracy of phylogenetic analyses. However, phylogenetic analyses of few taxa (but each represented by many characters) can be subject to strong systematic biases, which in turn produce high measures of repeatability (such as bootstrap proportions) in support of incorrect or misleading phylogenetic results. Thus, it is important for phylogeneticists to consider both the sampling of taxa, as well as the sampling of characters, in designing phylogenetic studies. Taxon sampling also improves estimates of evolutionary parameters derived from phylogenetic trees, and is thus important for improved applications of phylogenetic analyses. Analysis of sensitivity to taxon inclusion, the possible effects of long-branch attraction, and sensitivity of parameter estimation for model-based methods should be a part of any careful and thorough phylogenetic analysis. Furthermore, recent improvements in phylogenetic algorithms and in computational power have removed many constraints on analyzing large, thoroughly sampled data sets. Thorough taxon sampling is thus one of the most practical ways to improve the accuracy of phylogenetic estimates, as well as the accuracy of biological inferences that are based on these phylogenetic trees.

  14. Experimental design in phylogenetics: testing predictions from expected information.

    Science.gov (United States)

    San Mauro, Diego; Gower, David J; Cotton, James A; Zardoya, Rafael; Wilkinson, Mark; Massingham, Tim

    2012-07-01

    Taxon and character sampling are central to phylogenetic experimental design; yet, we lack general rules. Goldman introduced a method to construct efficient sampling designs in phylogenetics, based on the calculation of expected Fisher information given a probabilistic model of sequence evolution. The considerable potential of this approach remains largely unexplored. In an earlier study, we applied Goldman's method to a problem in the phylogenetics of caecilian amphibians and made an a priori evaluation and testable predictions of which taxon additions would increase information about a particular weakly supported branch of the caecilian phylogeny by the greatest amount. We have now gathered mitogenomic and rag1 sequences (some newly determined for this study) from additional caecilian species and studied how information (both expected and observed) and bootstrap support vary as each new taxon is individually added to our previous data set. This provides the first empirical test of specific predictions made using Goldman's method for phylogenetic experimental design. Our results empirically validate the top 3 (more intuitive) taxon addition predictions made in our previous study, but only information results validate unambiguously the 4th (less intuitive) prediction. This highlights a complex relationship between information and support, reflecting that each measures different things: Information is related to the ability to estimate branch length accurately and support to the ability to estimate the tree topology accurately. Thus, an increase in information may be correlated with but does not necessitate an increase in support. Our results also provide the first empirical validation of the widely held intuition that additional taxa that join the tree proximal to poorly supported internal branches are more informative and enhance support more than additional taxa that join the tree more distally. Our work supports the view that adding more data for a single (well

  15. Phylogenetic assignment of Mycobacterium tuberculosis Beijing clinical isolates in Japan by maximum a posteriori estimation.

    Science.gov (United States)

    Seto, Junji; Wada, Takayuki; Iwamoto, Tomotada; Tamaru, Aki; Maeda, Shinji; Yamamoto, Kaori; Hase, Atsushi; Murakami, Koichi; Maeda, Eriko; Oishi, Akira; Migita, Yuji; Yamamoto, Taro; Ahiko, Tadayuki

    2015-10-01

    Intra-species phylogeny of Mycobacterium tuberculosis has been regarded as a clue to estimate its potential risk to develop drug-resistance and various epidemiological tendencies. Genotypic characterization of variable number of tandem repeats (VNTR), a standard tool to ascertain transmission routes, has been improving as a public health effort, but determining phylogenetic information from those efforts alone is difficult. We present a platform based on maximum a posteriori (MAP) estimation to estimate phylogenetic information for M. tuberculosis clinical isolates from individual profiles of VNTR types. This study used 1245 M. tuberculosis clinical isolates obtained throughout Japan for construction of an MAP estimation formula. Two MAP estimation formulae, classification of Beijing family and other lineages, and classification of five Beijing sublineages (ST11/26, STK, ST3, and ST25/19 belonging to the ancient Beijing subfamily and modern Beijing subfamily), were created based on 24 loci VNTR (24Beijing-VNTR) profiles and phylogenetic information of the isolates. Recursive estimation based on the formulae showed high concordance with their authentic phylogeny by multi-locus sequence typing (MLST) of the isolates. The formulae might further support phylogenetic estimation of the Beijing lineage M. tuberculosis from the VNTR genotype with various geographic backgrounds. These results suggest that MAP estimation can function as a reliable probabilistic process to append phylogenetic information to VNTR genotypes of M. tuberculosis independently, which might improve the usage of genotyping data for control, understanding, prevention, and treatment of TB.

  16. Multiple sequence alignment accuracy and phylogenetic inference.

    Science.gov (United States)

    Ogden, T Heath; Rosenberg, Michael S

    2006-04-01

    Phylogenies are often thought to be more dependent upon the specifics of the sequence alignment rather than on the method of reconstruction. Simulation of sequences containing insertion and deletion events was performed in order to determine the role that alignment accuracy plays during phylogenetic inference. Data sets were simulated for pectinate, balanced, and random tree shapes under different conditions (ultrametric equal branch length, ultrametric random branch length, nonultrametric random branch length). Comparisons between hypothesized alignments and true alignments enabled determination of two measures of alignment accuracy, that of the total data set and that of individual branches. In general, our results indicate that as alignment error increases, topological accuracy decreases. This trend was much more pronounced for data sets derived from more pectinate topologies. In contrast, for balanced, ultrametric, equal branch length tree shapes, alignment inaccuracy had little average effect on tree reconstruction. These conclusions are based on average trends of many analyses under different conditions, and any one specific analysis, independent of the alignment accuracy, may recover very accurate or inaccurate topologies. Maximum likelihood and Bayesian, in general, outperformed neighbor joining and maximum parsimony in terms of tree reconstruction accuracy. Results also indicated that as the length of the branch and of the neighboring branches increase, alignment accuracy decreases, and the length of the neighboring branches is the major factor in topological accuracy. Thus, multiple-sequence alignment can be an important factor in downstream effects on topological reconstruction.

  17. Phylogenetic Conservatism in Plant Phenology

    Science.gov (United States)

    Davies, T. Jonathan; Wolkovich, Elizabeth M.; Kraft, Nathan J. B.; Salamin, Nicolas; Allen, Jenica M.; Ault, Toby R.; Betancourt, Julio L.; Bolmgren, Kjell; Cleland, Elsa E.; Cook, Benjamin I.; Crimmins, Theresa M.; Mazer, Susan J.; McCabe, Gregory J.; Pau, Stephanie; Regetz, Jim; Schwartz, Mark D.; Travers, Steven E.

    2013-01-01

    Phenological events defined points in the life cycle of a plant or animal have been regarded as highly plastic traits, reflecting flexible responses to various environmental cues. The ability of a species to track, via shifts in phenological events, the abiotic environment through time might dictate its vulnerability to future climate change. Understanding the predictors and drivers of phenological change is therefore critical. Here, we evaluated evidence for phylogenetic conservatism the tendency for closely related species to share similar ecological and biological attributes in phenological traits across flowering plants. We aggregated published and unpublished data on timing of first flower and first leaf, encompassing 4000 species at 23 sites across the Northern Hemisphere. We reconstructed the phylogeny for the set of included species, first, using the software program Phylomatic, and second, from DNA data. We then quantified phylogenetic conservatism in plant phenology within and across sites. We show that more closely related species tend to flower and leaf at similar times. By contrasting mean flowering times within and across sites, however, we illustrate that it is not the time of year that is conserved, but rather the phenological responses to a common set of abiotic cues. Our findings suggest that species cannot be treated as statistically independent when modelling phenological responses.Closely related species tend to resemble each other in the timing of their life-history events, a likely product of evolutionarily conserved responses to environmental cues. The search for the underlying drivers of phenology must therefore account for species' shared evolutionary histories.

  18. NNLOPS accurate associated HW production

    CERN Document Server

    Astill, William; Re, Emanuele; Zanderighi, Giulia

    2016-01-01

    We present a next-to-next-to-leading order accurate description of associated HW production consistently matched to a parton shower. The method is based on reweighting events obtained with the HW plus one jet NLO accurate calculation implemented in POWHEG, extended with the MiNLO procedure, to reproduce NNLO accurate Born distributions. Since the Born kinematics is more complex than the cases treated before, we use a parametrization of the Collins-Soper angles to reduce the number of variables required for the reweighting. We present phenomenological results at 13 TeV, with cuts suggested by the Higgs Cross Section Working Group.

  19. Molecular Phylogenetics: Concepts for a Newcomer.

    Science.gov (United States)

    Ajawatanawong, Pravech

    2016-10-26

    Molecular phylogenetics is the study of evolutionary relationships among organisms using molecular sequence data. The aim of this review is to introduce the important terminology and general concepts of tree reconstruction to biologists who lack a strong background in the field of molecular evolution. Some modern phylogenetic programs are easy to use because of their user-friendly interfaces, but understanding the phylogenetic algorithms and substitution models, which are based on advanced statistics, is still important for the analysis and interpretation without a guide. Briefly, there are five general steps in carrying out a phylogenetic analysis: (1) sequence data preparation, (2) sequence alignment, (3) choosing a phylogenetic reconstruction method, (4) identification of the best tree, and (5) evaluating the tree. Concepts in this review enable biologists to grasp the basic ideas behind phylogenetic analysis and also help provide a sound basis for discussions with expert phylogeneticists.

  20. Tripartitions do not always discriminate phylogenetic networks.

    Science.gov (United States)

    Cardona, Gabriel; Rosselló, Francesc; Valiente, Gabriel

    2008-02-01

    Phylogenetic networks are a generalization of phylogenetic trees that allow for the representation of non-treelike evolutionary events, like recombination, hybridization, or lateral gene transfer. In a recent series of papers devoted to the study of reconstructibility of phylogenetic networks, Moret, Nakhleh, Warnow and collaborators introduced the so-called tripartition metric for phylogenetic networks. In this paper we show that, in fact, this tripartition metric does not satisfy the separation axiom of distances (zero distance means isomorphism, or, in a more relaxed version, zero distance means indistinguishability in some specific sense) in any of the subclasses of phylogenetic networks where it is claimed to do so. We also present a subclass of phylogenetic networks whose members can be singled out by means of their sets of tripartitions (or even clusters), and hence where the latter can be used to define a meaningful metric.

  1. Phylogenetic analysis of the spirochetes.

    Science.gov (United States)

    Paster, B J; Dewhirst, F E; Weisburg, W G; Tordoff, L A; Fraser, G J; Hespell, R B; Stanton, T B; Zablen, L; Mandelco, L; Woese, C R

    1991-10-01

    The 16S rRNA sequences were determined for species of Spirochaeta, Treponema, Borrelia, Leptospira, Leptonema, and Serpula, using a modified Sanger method of direct RNA sequencing. Analysis of aligned 16S rRNA sequences indicated that the spirochetes form a coherent taxon composed of six major clusters or groups. The first group, termed the treponemes, was divided into two subgroups. The first treponeme subgroup consisted of Treponema pallidum, Treponema phagedenis, Treponema denticola, a thermophilic spirochete strain, and two species of Spirochaeta, Spirochaeta zuelzerae and Spirochaeta stenostrepta, with an average interspecies similarity of 89.9%. The second treponeme subgroup contained Treponema bryantii, Treponema pectinovorum, Treponema saccharophilum, Treponema succinifaciens, and rumen strain CA, with an average interspecies similarity of 86.2%. The average interspecies similarity between the two treponeme subgroups was 84.2%. The division of the treponemes into two subgroups was verified by single-base signature analysis. The second spirochete group contained Spirochaeta aurantia, Spirochaeta halophila, Spirochaeta bajacaliforniensis, Spirochaeta litoralis, and Spirochaeta isovalerica, with an average similarity of 87.4%. The Spirochaeta group was related to the treponeme group, with an average similarity of 81.9%. The third spirochete group contained borrelias, including Borrelia burgdorferi, Borrelia anserina, Borrelia hermsii, and a rabbit tick strain. The borrelias formed a tight phylogenetic cluster, with average similarity of 97%. THe borrelia group shared a common branch with the Spirochaeta group and was closer to this group than to the treponemes. A single spirochete strain isolated fromt the shew constituted the fourth group. The fifth group was composed of strains of Serpula (Treponema) hyodysenteriae and Serpula (Treponema) innocens. The two species of this group were closely related, with a similarity of greater than 99%. Leptonema illini

  2. Phylogenetic diversity of Amazonian tree communities

    OpenAIRE

    Honorio Coronado, Eurídice N.; Dexter, Kyle G.; Pennington, R. Toby; Chave, Jérôme; Lewis, Simon L.; Alexiades, Miguel N.; Alvarez, Esteban; Alves de Oliveira, Atila; Amaral, Iêda L.; Araujo-Murakami, Alejandro; Arets, Eric J. M. M.; Aymard, Gerardo A.; Baraloto, Christopher; Bonal, Damien; Brienen, Roel

    2015-01-01

    Aim: To examine variation in the phylogenetic diversity (PD) of tree communities across geographical and environmental gradients in Amazonia. Location: Two hundred and eighty-three c. 1 ha forest inventory plots from across Amazonia. Methods: We evaluated PD as the total phylogenetic branch length across species in each plot (PDss), the mean pairwise phylogenetic distance between species (MPD), the mean nearest taxon distance (MNTD) and their equivalents standardized for species richness (ses...

  3. Classification of Bacteria and Archaea: past, present and future.

    Science.gov (United States)

    Schleifer, Karl Heinz

    2009-12-01

    The late 19th century was the beginning of bacterial taxonomy and bacteria were classified on the basis of phenotypic markers. The distinction of prokaryotes and eukaryotes was introduced in the 1960s. Numerical taxonomy improved phenotypic identification but provided little information on the phylogenetic relationships of prokaryotes. Later on, chemotaxonomic and genotypic methods were widely used for a more satisfactory classification. Archaea were first classified as a separate group of prokaryotes in 1977. The current classification of Bacteria and Archaea is based on an operational-based model, the so-called polyphasic approach, comprised of phenotypic, chemotaxonomic and genotypic data, as well as phylogenetic information. The provisional status Candidatus has been established for describing uncultured prokaryotic cells for which their phylogenetic relationship has been determined and their authenticity revealed by in situ probing. The ultimate goal is to achieve a theory-based classification system based on a phylogenetic/evolutionary concept. However, there are currently two contradictory opinions about the future classification of Bacteria and Archaea. A group of mostly molecular biologists posits that the yet-unclear effect of gene flow, in particular lateral gene transfer, makes the line of descent difficult, if not impossible, to describe. However, even in the face of genomic fluidity it seems that the typical geno- and phenotypic characteristics of a taxon are still maintained, and are sufficient for reliable classification and identification of Bacteria and Archaea. There are many well-defined genotypic clusters that are congruent with known species delineated by polyphasic approaches. Comparative sequence analysis of certain core genes, including rRNA genes, may be useful for the characterization of higher taxa, whereas various character genes may be suitable as phylogenetic markers for the delineation of lower taxa. Nevertheless, there may still be

  4. Phylogenetic positions of RH blood group-related genes in cyclostomes.

    Science.gov (United States)

    Suzuki, Akinori; Endo, Kouhei; Kitano, Takashi

    2014-06-10

    The RH gene family in vertebrates consists of four major genes (RH, RHAG, RHBG, and RHCG). They are thought to have emerged in the common ancestor of vertebrates after two rounds of whole genome duplication (2R-WGD). To analyze the detailed phylogenetic relationships within the RH gene family, we determined three types of cDNA sequence that belong to the RH gene family in lamprey (Lethenteron reissneri) and designated them as RHBG-like, RHCG-like1, and RHCG-like2. Phylogenetic analyses clearly showed that RHCG-like1 and RHCG-like2 genes, which were probably duplicated in the lamprey lineage, are orthologs of gnathostome RHCG. In contrast, the clear phylogenetic position of the RHBG-like gene could not be obtained. Probably some convergent events for cyclostome RHBG-like genes prevented the accurate identification of their phylogenetic positions. Copyright © 2014 Elsevier B.V. All rights reserved.

  5. Updated angiosperm family tree for analyzing phylogenetic diversity and community structure

    Directory of Open Access Journals (Sweden)

    Markus Gastauer

    Full Text Available ABSTRACT The computation of phylogenetic diversity and phylogenetic community structure demands an accurately calibrated, high-resolution phylogeny, which reflects current knowledge regarding diversification within the group of interest. Herein we present the angiosperm phylogeny R20160415.new, which is based on the topology proposed by the Angiosperm Phylogeny Group IV, a recently released compilation of angiosperm diversification. R20160415.new is calibratable by different sets of recently published estimates of mean node ages. Its application for the computation of phylogenetic diversity and/or phylogenetic community structure is straightforward and ensures the inclusion of up-to-date information in user specific applications, as long as users are familiar with the pitfalls of such hand-made supertrees.

  6. Relevant phylogenetic invariants of evolutionary models

    CERN Document Server

    Casanellas, Marta

    2009-01-01

    Recently there have been several attempts to provide a whole set of generators of the ideal of the algebraic variety associated to a phylogenetic tree evolving under an algebraic model. These algebraic varieties have been proven to be useful in phylogenetics. In this paper we prove that, for phylogenetic reconstruction purposes, it is enough to consider generators coming from the edges of the tree, the so-called edge invariants. This is the algebraic analogous to Buneman's Splits Equivalence Theorem. The interest of this result relies on its potential applications in phylogenetics for the widely used evolutionary models such as Jukes-Cantor, Kimura 2 and 3 parameters, and General Markov models.

  7. Automatic web services classification based on rough set theory

    Institute of Scientific and Technical Information of China (English)

    陈立; 张英; 宋自林; 苗壮

    2013-01-01

    With development of web services technology, the number of existing services in the internet is growing day by day. In order to achieve automatic and accurate services classification which can be beneficial for service related tasks, a rough set theory based method for services classification was proposed. First, the services descriptions were preprocessed and represented as vectors. Elicited by the discernibility matrices based attribute reduction in rough set theory and taking into account the characteristic of decision table of services classification, a method based on continuous discernibility matrices was proposed for dimensionality reduction. And finally, services classification was processed automatically. Through the experiment, the proposed method for services classification achieves approving classification result in all five testing categories. The experiment result shows that the proposed method is accurate and could be used in practical web services classification.

  8. Alu elements and hominid phylogenetics

    Science.gov (United States)

    Salem, Abdel-Halim; Ray, David A.; Xing, Jinchuan; Callinan, Pauline A.; Myers, Jeremy S.; Hedges, Dale J.; Garber, Randall K.; Witherspoon, David J.; Jorde, Lynn B.; Batzer, Mark A.

    2003-01-01

    Alu elements have inserted in primate genomes throughout the evolution of the order. One particular Alu lineage (Ye) began amplifying relatively early in hominid evolution and continued propagating at a low level as many of its members are found in a variety of hominid genomes. This study represents the first conclusive application of short interspersed elements, which are considered nearly homoplasy-free, to elucidate the phylogeny of hominids. Phylogenetic analysis of Alu Ye5 elements and elements from several other subfamilies reveals high levels of support for monophyly of Hominidae, tribe Hominini and subtribe Hominina. Here we present the strongest evidence reported to date for a sister relationship between humans and chimpanzees while clearly distinguishing the chimpanzee and human lineages. PMID:14561894

  9. Classification in context

    DEFF Research Database (Denmark)

    Mai, Jens Erik

    2004-01-01

    This paper surveys classification research literature, discusses various classification theories, and shows that the focus has traditionally been on establishing a scientific foundation for classification research. This paper argues that a shift has taken place, and suggests that contemporary...... classification research focus on contextual information as the guide for the design and construction of classification schemes....

  10. Classification in Australia.

    Science.gov (United States)

    McKinlay, John

    Despite some inroads by the Library of Congress Classification and short-lived experimentation with Universal Decimal Classification and Bliss Classification, Dewey Decimal Classification, with its ability in recent editions to be hospitable to local needs, remains the most widely used classification system in Australia. Although supplemented at…

  11. New Phylogenetic Groups of Torque Teno Virus Identified in Eastern Taiwan Indigenes.

    Directory of Open Access Journals (Sweden)

    Kuang-Liang Hsiao

    Full Text Available Torque teno virus (TTV is a single-stranded DNA virus highly prevalent in the world. It has been detected in eastern Taiwan indigenes with a low prevalence of 11% by using N22 region of which known to underestimate TTV prevalence excessively. In order to clarify their realistic epidemiology, we re-analyzed TTV prevalence with UTR region. One hundred and forty serum samples from eastern Taiwanese indigenous population were collected and TTV DNA was detected in 133 (95% samples. Direct sequencing revealed an extensive mix-infection of different TTV strains within the infected individual. Entire TTV open reading frame 1 was amplified and cloned from a TTV positive individual to distinguish mix-infected strains. Phylogenetic analysis showed eleven isolates were clustered into a monophyletic group that is distinct from all known groups. In addition, another our isolate was clustered with recently described Hebei-1 strain and formed an independent clade. Based on the distribution pattern of pairwise distances, both new clusters were placed at phylogenetic group level, designed as the 6th and 7th phylogenetic group. In present study, we showed a very high prevalence of TTV infection in eastern Taiwan indigenes and indentified new phylogenetic groups from the infected individual. Both intra- and inter-phylogenetic group mix-infections can be found from one healthy person. Our study has further broadened the field of human TTVs and proposed a robust criterion for classification of the major TTV phylogenetic groups.

  12. Efficient segmentation by sparse pixel classification

    DEFF Research Database (Denmark)

    Dam, Erik B; Loog, Marco

    2008-01-01

    Segmentation methods based on pixel classification are powerful but often slow. We introduce two general algorithms, based on sparse classification, for optimizing the computation while still obtaining accurate segmentations. The computational costs of the algorithms are derived......, and they are demonstrated on real 3-D magnetic resonance imaging and 2-D radiograph data. We show that each algorithm is optimal for specific tasks, and that both algorithms allow a speedup of one or more orders of magnitude on typical segmentation tasks....

  13. Progress, pitfalls and parallel universes: a history of insect phylogenetics.

    Science.gov (United States)

    Kjer, Karl M; Simon, Chris; Yavorskaya, Margarita; Beutel, Rolf G

    2016-08-01

    The phylogeny of insects has been both extensively studied and vigorously debated for over a century. A relatively accurate deep phylogeny had been produced by 1904. It was not substantially improved in topology until recently when phylogenomics settled many long-standing controversies. Intervening advances came instead through methodological improvement. Early molecular phylogenetic studies (1985-2005), dominated by a few genes, provided datasets that were too small to resolve controversial phylogenetic problems. Adding to the lack of consensus, this period was characterized by a polarization of philosophies, with individuals belonging to either parsimony or maximum-likelihood camps; each largely ignoring the insights of the other. The result was an unfortunate detour in which the few perceived phylogenetic revolutions published by both sides of the philosophical divide were probably erroneous. The size of datasets has been growing exponentially since the mid-1980s accompanied by a wave of confidence that all relationships will soon be known. However, large datasets create new challenges, and a large number of genes does not guarantee reliable results. If history is a guide, then the quality of conclusions will be determined by an improved understanding of both molecular and morphological evolution, and not simply the number of genes analysed.

  14. Phylogenetic signal dissection identifies the root of starfishes.

    Science.gov (United States)

    Feuda, Roberto; Smith, Andrew B

    2015-01-01

    Relationships within the class Asteroidea have remained controversial for almost 100 years and, despite many attempts to resolve this problem using molecular data, no consensus has yet emerged. Using two nuclear genes and a taxon sampling covering the major asteroid clades we show that non-phylogenetic signal created by three factors--Long Branch Attraction, compositional heterogeneity and the use of poorly fitting models of evolution--have confounded accurate estimation of phylogenetic relationships. To overcome the effect of this non-phylogenetic signal we analyse the data using non-homogeneous models, site stripping and the creation of subpartitions aimed to reduce or amplify the systematic error, and calculate Bayes Factor support for a selection of previously suggested topological arrangements of asteroid orders. We show that most of the previous alternative hypotheses are not supported in the most reliable data partitions, including the previously suggested placement of either Forcipulatida or Paxillosida as sister group to the other major branches. The best-supported solution places Velatida as the sister group to other asteroids, and the implications of this finding for the morphological evolution of asteroids are presented.

  15. Phylogenetic signal dissection identifies the root of starfishes.

    Directory of Open Access Journals (Sweden)

    Roberto Feuda

    Full Text Available Relationships within the class Asteroidea have remained controversial for almost 100 years and, despite many attempts to resolve this problem using molecular data, no consensus has yet emerged. Using two nuclear genes and a taxon sampling covering the major asteroid clades we show that non-phylogenetic signal created by three factors--Long Branch Attraction, compositional heterogeneity and the use of poorly fitting models of evolution--have confounded accurate estimation of phylogenetic relationships. To overcome the effect of this non-phylogenetic signal we analyse the data using non-homogeneous models, site stripping and the creation of subpartitions aimed to reduce or amplify the systematic error, and calculate Bayes Factor support for a selection of previously suggested topological arrangements of asteroid orders. We show that most of the previous alternative hypotheses are not supported in the most reliable data partitions, including the previously suggested placement of either Forcipulatida or Paxillosida as sister group to the other major branches. The best-supported solution places Velatida as the sister group to other asteroids, and the implications of this finding for the morphological evolution of asteroids are presented.

  16. Hyperspectral image classification using functional data analysis.

    Science.gov (United States)

    Li, Hong; Xiao, Guangrun; Xia, Tian; Tang, Y Y; Li, Luoqing

    2014-09-01

    The large number of spectral bands acquired by hyperspectral imaging sensors allows us to better distinguish many subtle objects and materials. Unlike other classical hyperspectral image classification methods in the multivariate analysis framework, in this paper, a novel method using functional data analysis (FDA) for accurate classification of hyperspectral images has been proposed. The central idea of FDA is to treat multivariate data as continuous functions. From this perspective, the spectral curve of each pixel in the hyperspectral images is naturally viewed as a function. This can be beneficial for making full use of the abundant spectral information. The relevance between adjacent pixel elements in the hyperspectral images can also be utilized reasonably. Functional principal component analysis is applied to solve the classification problem of these functions. Experimental results on three hyperspectral images show that the proposed method can achieve higher classification accuracies in comparison to some state-of-the-art hyperspectral image classification methods.

  17. Conflicting phylogenetic position of Schizosaccharomyces pombe

    NARCIS (Netherlands)

    Kuramae, Eiko E.; Robert, Vincent; Snel, Berend; Boekhout, Teun

    2006-01-01

    The phylogenetic position of the fission yeast Schizosaccharomyces pombe in the fungal Tree of Life is still controversial. Three alternative phylogenetic positions have been proposed in the literature, namely (1) a position basal to the Hemiascomycetes and Euascomycetes, (2) a position as a sister

  18. Efficient Computation of Popular Phylogenetic Tree Measures

    DEFF Research Database (Denmark)

    Tsirogiannis, Constantinos; Sandel, Brody Steven; Cheliotis, Dimitris

    2012-01-01

    Given a phylogenetic tree $\\mathcal{T}$ of n nodes, and a sample R of its tips (leaf nodes) a very common problem in ecological and evolutionary research is to evaluate a distance measure for the elements in R. Two of the most common measures of this kind are the Mean Pairwise Distance...... software package for processing phylogenetic trees....

  19. Insect phylogenetics in the digital age.

    Science.gov (United States)

    Dietrich, Christopher H; Dmitriev, Dmitry A

    2016-12-01

    Insect systematists have long used digital data management tools to facilitate phylogenetic research. Web-based platforms developed over the past several years support creation of comprehensive, openly accessible data repositories and analytical tools that support large-scale collaboration, accelerating efforts to document Earth's biota and reconstruct the Tree of Life. New digital tools have the potential to further enhance insect phylogenetics by providing efficient workflows for capturing and analyzing phylogenetically relevant data. Recent initiatives streamline various steps in phylogenetic studies and provide community access to supercomputing resources. In the near future, automated, web-based systems will enable researchers to complete a phylogenetic study from start to finish using resources linked together within a single portal and incorporate results into a global synthesis.

  20. Use of whole genome sequences to develop a molecular phylogenetic framework for Rhodococcus fascians and the Rhodococcus genus

    Directory of Open Access Journals (Sweden)

    Allison L. Creason

    2014-08-01

    Full Text Available The accurate diagnosis of diseases caused by pathogenic bacteria requires a stable species classification. Rhodococcus fascians is the only documented member of its ill-defined genus that is capable of causing disease on a wide range of agriculturally important plants. Comparisons of genome sequences generated from isolates of Rhodococcus associated with diseased plants revealed a level of genetic diversity consistent with them representing multiple species. To test this, we generated a tree based on more than 1700 homologous sequences from plant-associated isolates of Rhodococcus, and obtained support from additional approaches that measure and cluster based on genome similarities. Results were consistent in supporting the definition of new Rhodococcus species within clades containing phytopathogenic members. We also used the genome sequences, along with other rhodococcal genome sequences to construct a molecular phylogenetic tree as a framework for resolving the Rhodococcus genus. Results indicated that Rhodococcus has the potential for having 20 species and also confirmed a need to revisit the taxonomic groupings within Rhodococcus.

  1. Optimizing Phylogenetic Queries for Performance.

    Science.gov (United States)

    Jamil, Hasan M

    2017-08-24

    The vast majority of phylogenetic databases do not support declarative querying using which their contents can be flexibly and conveniently accessed and the template based query interfaces they support do not allow arbitrary speculative queries. They therefore also do not support query optimization leveraging unique phylogeny properties. While a small number of graph query languages such as XQuery, Cypher and GraphQL exist for computer savvy users, most are too general and complex to be useful for biologists, and too inefficient for large phylogeny querying. In this paper, we discuss a recently introduced visual query language, called PhyQL, that leverages phylogeny specific properties to support essential and powerful constructs for a large class of phylogentic queries. We develop a range of pruning aids, and propose a substantial set of query optimization strategies using these aids suitable for large phylogeny querying. A hybrid optimization technique that exploits a set of indices and ``graphlet" partitioning is discussed. A ``fail soonest" strategy is used to avoid hopeless processing and is shown to produce dividends. Possible novel optimization techniques yet to be explored are also discussed.

  2. Plasmid Classification in an Era of Whole-Genome Sequencing: Application in Studies of Antibiotic Resistance Epidemiology

    Science.gov (United States)

    Orlek, Alex; Stoesser, Nicole; Anjum, Muna F.; Doumith, Michel; Ellington, Matthew J.; Peto, Tim; Crook, Derrick; Woodford, Neil; Walker, A. Sarah; Phan, Hang; Sheppard, Anna E.

    2017-01-01

    Plasmids are extra-chromosomal genetic elements ubiquitous in bacteria, and commonly transmissible between host cells. Their genomes include variable repertoires of ‘accessory genes,’ such as antibiotic resistance genes, as well as ‘backbone’ loci which are largely conserved within plasmid families, and often involved in key plasmid-specific functions (e.g., replication, stable inheritance, mobility). Classifying plasmids into different types according to their phylogenetic relatedness provides insight into the epidemiology of plasmid-mediated antibiotic resistance. Current typing schemes exploit backbone loci associated with replication (replicon typing), or plasmid mobility (MOB typing). Conventional PCR-based methods for plasmid typing remain widely used. With the emergence of whole-genome sequencing (WGS), large datasets can be analyzed using in silico plasmid typing methods. However, short reads from popular high-throughput sequencers can be challenging to assemble, so complete plasmid sequences may not be accurately reconstructed. Therefore, localizing resistance genes to specific plasmids may be difficult, limiting epidemiological insight. Long-read sequencing will become increasingly popular as costs decline, especially when resolving accurate plasmid structures is the primary goal. This review discusses the application of plasmid classification in WGS-based studies of antibiotic resistance epidemiology; novel in silico plasmid analysis tools are highlighted. Due to the diverse and plastic nature of plasmid genomes, current typing schemes do not classify all plasmids, and identifying conserved, phylogenetically concordant genes for subtyping and phylogenetics is challenging. Analyzing plasmids as nodes in a network that represents gene-sharing relationships between plasmids provides a complementary way to assess plasmid diversity, and allows inferences about horizontal gene transfer to be made. PMID:28232822

  3. Classifying the bacterial gut microbiota of termites and cockroaches: A curated phylogenetic reference database (DictDb).

    Science.gov (United States)

    Mikaelyan, Aram; Köhler, Tim; Lampert, Niclas; Rohland, Jeffrey; Boga, Hamadi; Meuser, Katja; Brune, Andreas

    2015-10-01

    Recent developments in sequencing technology have given rise to a large number of studies that assess bacterial diversity and community structure in termite and cockroach guts based on large amplicon libraries of 16S rRNA genes. Although these studies have revealed important ecological and evolutionary patterns in the gut microbiota, classification of the short sequence reads is limited by the taxonomic depth and resolution of the reference databases used in the respective studies. Here, we present a curated reference database for accurate taxonomic analysis of the bacterial gut microbiota of dictyopteran insects. The Dictyopteran gut microbiota reference Database (DictDb) is based on the Silva database but was significantly expanded by the addition of clones from 11 mostly unexplored termite and cockroach groups, which increased the inventory of bacterial sequences from dictyopteran guts by 26%. The taxonomic depth and resolution of DictDb was significantly improved by a general revision of the taxonomic guide tree for all important lineages, including a detailed phylogenetic analysis of the Treponema and Alistipes complexes, the Fibrobacteres, and the TG3 phylum. The performance of this first documented version of DictDb (v. 3.0) using the revised taxonomic guide tree in the classification of short-read libraries obtained from termites and cockroaches was highly superior to that of the current Silva and RDP databases. DictDb uses an informative nomenclature that is consistent with the literature also for clades of uncultured bacteria and provides an invaluable tool for anyone exploring the gut community structure of termites and cockroaches.

  4. Phylogenetic analysis of Pectinidae (Bivalvia) based on the ribosomal DNA internal transcribed spacer region

    Institute of Scientific and Technical Information of China (English)

    2007-01-01

    The ribosomal DNA internal transcribed spacer (ITS) region is a useful genomic region for understanding evolutionary and genetic relationships. In the current study, the molecular phylogenetic analysis of Pectinidae (Mollusca: Bivalvia) was performed using the nucleotide sequences of the nuclear ITS region in nine species of this family. The sequences were obtained from the scallop species Argopecten irradians, Mizuhopecten yessoensis, Amusium pleuronectes and Mimachlamys nobilis, and compared with the published sequences of Aequipecten opercularis, Chlamys farreri, C. distorta, M. varia, Pecten maximus, and an outgroup species Perna viridis. The molecular phylogenetic tree was constructed by the neighbor-joining and maximum parsimony methods. Phylogenetic analysis based on ITS1, ITS2, or their combination always yielded trees of similar topology. The results support the morphological classifications of bivalve and are nearly consistent with classification of two subfamilies (Chlamydinae and Pectininae) formulated by Waller. However, A. irradians, together with A. opercularis made up of genera Amusium, evidences that they may belong to the subfamily Pectinidae. The data are incompatible with the conclusion of Waller who placed them in Chlamydinae by morphological characteristics. These results provide new insights into the evolutionary relationships among scallop species and contribute to the improvement of existing classification systems.

  5. Phylogenetic systematics of the Eucarida (Crustacea malacostraca

    Directory of Open Access Journals (Sweden)

    Martin L. Christoffersen

    1988-01-01

    Full Text Available Ninety-four morphological characters belonging to particular ontogenetic sequences within the Eucarida were used to produce a hierarchy of 128 evolutionary novelties (73 synapomorphies and 55 homoplasies and to delimit 15 monophyletic taxa. The following combined Recent-fossil sequenced phylogenetic classification is proposed: Superorder Eucarida; Order Euphausiacea; Family Bentheuphausiidae; Family Euphausiidae; Order Amphionidacea; Order Decapoda; Suborder Penaeidea; Suborder Pleocyemata; Infraorder Stenopodidea; Infraorder Reptantia; Infraorder Procarididea, Infraorder Caridea. The position of the Amphionidacea as the sister-group of the Decapoda is corroborated, while the Reptantia are proposed to be the sister-group of the Procarididea + Caridea for the first time. The fossil groups Uncina Quenstedt, 1850, and Palaeopalaemon Whitfield, 1880, are included as incertae sedis taxa within the Reptantia, which establishes the minimum ages of all the higher taxa of Eucarida except the Procarididea and Caridea in the Upper Devonian. The fossil group "Pygocephalomorpha" Beurlen, 1930, of uncertain status as a monophyletic taxon, is provisionally considered to belong to the "stem-group" of the Reptantia. Among the more important characters hypothesized to have evolved in the stem-lineage of each eucaridan monophyletic taxon are: (1 in Eucarida, attachement of post-zoeal carapace to all thoracic somites; (2 in Euphausiacea, reduction of endopod of eighth thoracopod; (3 in Bentheuphausiidae, compound eyes vestigial, associated with abyssal life; (4 in Euphausiidae, loss of endopod of eighth thoracopod and development of specialized luminescent organs; (5 in Amphionidacea + Decapoda, ambulatory ability of thoracic exopods reduced, scaphognathite, one pair of maxillipedes, pleurobranch gill series and carapace covering gills, associated with loss of pelagic life; (6 in Amphionidacea, unique thoracic brood pouch in females formed by inflated carapace and

  6. Phylogenetic analysis of the kinesin superfamily from Physcomitrella

    Directory of Open Access Journals (Sweden)

    Zhiyuan eShen

    2012-10-01

    Full Text Available Kinesins are an ancient superfamily of microtubule dependent motors. They participate in an ex-tensive and diverse list of essential cellular functions, including mitosis, cytokinesis, cell polari-zation, cell elongation, flagellar development, and intracellular transport. Based on phylogenetic relationships, the kinesin superfamily has been subdivided into 14 families, which are represented in most eukaryotic phyla. The functions of these families are sometimes conserved between species, but important variations in function across species have been observed. Plants possess most kinesin families including a few plant-specific families. With the availability of an ever in-creasing number of genome sequences from plants, it is important to document the complete complement of kinesins present in a given organism. This will help develop a molecular frame-work to explore the function of each family using genetics, biochemistry and cell biology. The moss Physcomitrella patens has emerged as a powerful model organism to study gene function in plants, which makes it a key candidate to explore complex gene families, such as the kinesin superfamily. Here we report a detailed phylogenetic characterization of the 71 kinesins of the kinesin superfamily in Physcomitrella. We found a remarkable conservation of families and sub-family classes with Arabidopsis, which is important for future comparative analysis of function. Some of the families, such as kinesins 14s are composed of fewer members in moss, while other families, such as the kinesin 12s are greatly expanded. To improve the comparison between spe-cies, and to simplify communication between research groups, we propose a classification of subfamilies based on our phylogenetic analysis.

  7. Phylogenetic placement of Hydra and relationships within Aplanulata (Cnidaria: Hydrozoa).

    Science.gov (United States)

    Nawrocki, Annalise M; Collins, Allen G; Hirano, Yayoi M; Schuchert, Peter; Cartwright, Paulyn

    2013-04-01

    The model organism Hydra belongs to the hydrozoan clade Aplanulata. Despite being a popular model system for development, little is known about the phylogenetic placement of this taxon or the relationships of its closest relatives. Previous studies have been conflicting regarding sister group relationships and have been unable to resolve deep nodes within the clade. In addition, there are several putative Aplanulata taxa that have never been sampled for molecular data or analyzed using multiple markers. Here, we combine the fast-evolving cytochrome oxidase 1 (CO1) mitochondrial marker with mitochondrial 16S, nuclear small ribosomal subunit (18S, SSU) and large ribosomal subunit (28S, LSU) sequences to examine relationships within the clade Aplanulata. We further discuss the relative contribution of four different molecular markers to resolving phylogenetic relationships within Aplanulata. Lastly, we report morphological synapomorphies for some of the major Aplanulata genera and families, and suggest new taxonomic classifications for two species of Aplanulata, Fukaurahydra anthoformis and Corymorpha intermedia, based on a preponderance of molecular and morphological data that justify the designation of these species to different genera. Copyright © 2012 Elsevier Inc. All rights reserved.

  8. Classification of the web

    DEFF Research Database (Denmark)

    Mai, Jens Erik

    2004-01-01

    This paper discusses the challenges faced by investigations into the classification of the Web and outlines inquiries that are needed to use principles for bibliographic classification to construct classifications of the Web. This paper suggests that the classification of the Web meets challenges...

  9. A statistical approach to root system classification.

    Directory of Open Access Journals (Sweden)

    Gernot eBodner

    2013-08-01

    Full Text Available Plant root systems have a key role in ecology and agronomy. In spite of fast increase in root studies, still there is no classification that allows distinguishing among distinctive characteristics within the diversity of rooting strategies. Our hypothesis is that a multivariate approach for plant functional type identification in ecology can be applied to the classification of root systems. We demonstrate that combining principal component and cluster analysis yields a meaningful classification of rooting types based on morphological traits. The classification method presented is based on a data-defined statistical procedure without a priori decision on the classifiers. Biplot inspection is used to determine key traits and to ensure stability in cluster based grouping. The classification method is exemplified with simulated root architectures and morphological field data. Simulated root architectures showed that morphological attributes with spatial distribution parameters capture most distinctive features within root system diversity. While developmental type (tap vs. shoot-borne systems is a strong, but coarse classifier, topological traits provide the most detailed differentiation among distinctive groups. Adequacy of commonly available morphologic traits for classification is supported by field data. Three rooting types emerged from measured data, distinguished by diameter/weight, density and spatial distribution respectively. Similarity of root systems within distinctive groups was the joint result of phylogenetic relation and environmental as well as human selection pressure. We concluded that the data-define classification is appropriate for integration of knowledge obtained with different root measurement methods and at various scales. Currently root morphology is the most promising basis for classification due to widely used common measurement protocols. To capture details of root diversity efforts in architectural measurement

  10. Use of phylogenetical analysis to predict susceptibility of pathogenic Candida spp. to antifungal drugs.

    Science.gov (United States)

    Maheux, Andrée F; Sellam, Adnane; Piché, Yves; Boissinot, Maurice; Pelletier, René; Boudreau, Dominique K; Picard, François J; Trépanier, Hélène; Boily, Marie-Josée; Ouellette, Marc; Roy, Paul H; Bergeron, Michel G

    2016-12-01

    Successful treatment of a Candida infection relies on 1) an accurate identification of the pathogenic fungus and 2) on its susceptibility to antifungal drugs. In the present study we investigated the level of correlation between phylogenetical evolution and susceptibility of pathogenic Candida spp. to antifungal drugs. For this, we compared a phylogenetic tree, assembled with the concatenated sequences (2475-bp) of the ATP2, TEF1, and TUF1 genes from 20 representative Candida species, with published minimal inhibitory concentrations (MIC) of the four principal antifungal drug classes commonly used in the treatment of candidiasis: polyenes, triazoles, nucleoside analogues, and echinocandins. The phylogenetic tree revealed three distinct phylogenetic clusters among Candida species. Species within a given phylogenetic cluster have generally similar susceptibility profiles to antifungal drugs and species within Clusters II and III were less sensitive to antifungal drugs than Cluster I species. These results showed that phylogenetical relationship between clusters and susceptibility to several antifungal drugs could be used to guide therapy when only species identification is available prior to information pertaining to its resistance profile. An extended study comprising a large panel of clinical samples should be conducted to confirm the efficiency of this approach in the treatment of candidiasis. Copyright © 2016. Published by Elsevier B.V.

  11. Testing for phylogenetic signal in biological traits: the ubiquity of cross-product statistics.

    Science.gov (United States)

    Pavoine, Sandrine; Ricotta, Carlo

    2013-03-01

    To evaluate rates of evolution, to establish tests of correlation between two traits, or to investigate to what degree the phylogeny of a species assemblage is predictive of a trait value so-called tests for phylogenetic signal are used. Being based on different approaches, these tests are generally thought to possess quite different statistical performances. In this article, we show that the Blomberg et al. K and K*, the Abouheif index, the Moran's I, and the Mantel correlation are all based on a cross-product statistic, and are thus all related to each other when they are associated to a permutation test of phylogenetic signal. What changes is only the way phylogenetic and trait similarities (or dissimilarities) among the tips of a phylogeny are computed. The definitions of the phylogenetic and trait-based (dis)similarities among tips thus determines the performance of the tests. We shortly discuss the biological and statistical consequences (in terms of power and type I error of the tests) of the observed relatedness among the statistics that allow tests for phylogenetic signal. Blomberg et al. K* statistic appears as one on the most efficient approaches to test for phylogenetic signal. When branch lengths are not available or not accurate, Abouheif's Cmean statistic is a powerful alternative to K*.

  12. On distances between phylogenetic trees

    Energy Technology Data Exchange (ETDEWEB)

    DasGupta, B. [Rutgers Univ., Camden, NJ (United States); He, X. [SUNY, Buffalo, NY (United States); Jiang, T. [McMaster Univ., Hamilton, Ontario (Canada)] [and others

    1997-06-01

    Different phylogenetic trees for the same group of species are often produced either by procedures that use diverse optimality criteria or from different genes in the study of molecular evolution. Comparing these trees to find their similarities and dissimilarities, i.e. distance, is thus an important issue in computational molecular biology. The nearest neighbor interchange distance and the subtree-transfer distance are two major distance metrics that have been proposed and extensively studied for different reasons. Despite their many appealing aspects such as simplicity and sensitivity to tree topologies, computing these distances has remained very challenging. This article studies the complexity and efficient approximation algorithms for computing the nni distance and a natural extension of the subtree-transfer distance, called the linear-cost subtree-transfer distance. The linear-cost subtree-transfer model is more logical than the subtree-transfer model and in fact coincides with the nni model under certain conditions. The following results have been obtained as part of our project of building a comprehensive software package for computing distances between phylogenies. (1) Computing the nni distance is NP-complete. This solves a 25 year old open question appearing again and again in, for example, under the complexity-theoretic assumption of P {ne} NP. We also answer an open question regarding the nni distance between unlabeled trees for which an erroneous proof appeared in. We give an algorithm to compute the optimal nni sequence in time O(n{sup 2} logn + n {circ} 2{sup O(d)}), where the nni distance is at most d. (2) Biological applications require us to extend the nni and linear-cost subtree-transfer models to weighted phylogenies, where edge weights indicate the length of evolution along each edge. We present a logarithmic ratio approximation algorithm for nni and a ratio 2 approximation algorithm for linear-cost subtree-transfer, on weighted trees.

  13. Molecular Phylogenetics: Mathematical Framework and Unsolved Problems

    Science.gov (United States)

    Xia, Xuhua

    Phylogenetic relationship is essential in dating evolutionary events, reconstructing ancestral genes, predicting sites that are important to natural selection, and, ultimately, understanding genomic evolution. Three categories of phylogenetic methods are currently used: the distance-based, the maximum parsimony, and the maximum likelihood method. Here, I present the mathematical framework of these methods and their rationales, provide computational details for each of them, illustrate analytically and numerically the potential biases inherent in these methods, and outline computational challenges and unresolved problems. This is followed by a brief discussion of the Bayesian approach that has been recently used in molecular phylogenetics.

  14. On Tree-Based Phylogenetic Networks.

    Science.gov (United States)

    Zhang, Louxin

    2016-07-01

    A large class of phylogenetic networks can be obtained from trees by the addition of horizontal edges between the tree edges. These networks are called tree-based networks. We present a simple necessary and sufficient condition for tree-based networks and prove that a universal tree-based network exists for any number of taxa that contains as its base every phylogenetic tree on the same set of taxa. This answers two problems posted by Francis and Steel recently. A byproduct is a computer program for generating random binary phylogenetic networks under the uniform distribution model.

  15. Classification issues related to neuropathic trigeminal pain.

    Science.gov (United States)

    Zakrzewska, Joanna M

    2004-01-01

    The goal of a classification system of medical conditions is to facilitate accurate communication, to ensure that each condition is described uniformly and universally and that all data banks for the storage and retrieval of research and clinical data related to the conditions are consistent. Classification entails deciding which kinds of diagnostic entities should be recognized and how to order them in a meaningful way. Currently there are 3 major pain classification systems of relevance to orofacial pain: The International Association for the Study of Pain classification system, the International Headache Society classification system, and the Research Diagnostic Criteria for Temporomandibular Disorders (RDC/TMD). All use different methodologies, and only the RDC/TMD take into account social and psychologic factors in the classification of conditions. Classification systems need to be reliable, valid, comprehensive, generalizable, and flexible, and they need to be tested using consensus views of experts as well as the available literature. There is an urgent need for a robust classification system for neuropathic trigeminal pain.

  16. Estudio filogenético de los géneros de Lithinini de Sudamérica Austral (Lepidoptera, Geometridae: una nueva clasificación Phylogenetic study of the genera of Lithinini (Lepidoptera, Geometridae from southern South America: a new classification

    Directory of Open Access Journals (Sweden)

    Luis E. Parra

    2010-03-01

    work we evaluate the taxonomy of the Lithinini of Austral South America based on a phylogenetic analysis. In our analysis we used outgroup Catophoenissa. Two approaches were used to evaluate phylogenetic relationships: 1 parsimony criterion, and 2 Bayesian inference. Parsimony analysis was conducted in PAUP software, and Bayesian analysis with Markov chains Monte Carlo using the BayesPhylogenies software. Our results based on the phylogenetic hypothesis suggest a new taxonomic order for Austral American Lithinini. The valid genera are: Asestra Warren, Acauro Rindge, Calta Rindge, Euclidiodes Warren, Franciscoia Orfila and Schajovskoy, Incalvertia Bartlett-Calvert, Lacaria Orfila and Schajovskoy, Laneco Rindge, Maeandrogonaria Butler, Martindoelloia Orfila and Schajovskoy, Nucara Rindge, Odontothera Butler, Proteopharmacis Warren, Psilaspilates Butler, Rhinoligia Warren and Tanagridia Butler. The main changes with respect to the previous taxonomic order are: 1 Yalpa Rindge is the synonymous junior of Odontothera; 2 the genus Rhinoligia Warren is incorporated into the Lithinini; 3 while our analysis reaffirms that Siopla Rindge is junior synonym of Asestra, Yapoma Rindge and Duraglia Rindge are synonymous of Euclidiodes Warren, while Callemo Rindge and Guara Rindge are synonymous of Tanagridia; 4 the genus Calta Rindge, Incalvertia Rindge, Odontothera Butler and Proteopharmacis Warren, synonymized by Pitkin, are redefined, revalidated and incorporated into the Lithinini tribe. A new species for the genus Franciscoia, F. ediliae Parra is described. A catalogue of the genera and species of the tribe in the region, and the figures of adults and genitalia of some species are included.

  17. DendroBLAST: approximate phylogenetic trees in the absence of multiple sequence alignments.

    Directory of Open Access Journals (Sweden)

    Steven Kelly

    Full Text Available The rapidly growing availability of genome information has created considerable demand for both fast and accurate phylogenetic inference algorithms. We present a novel method called DendroBLAST for reconstructing phylogenetic dendrograms/trees from protein sequences using BLAST. This method differs from other methods by incorporating a simple model of sequence evolution to test the effect of introducing sequence changes on the reliability of the bipartitions in the inferred tree. Using realistic simulated sequence data we demonstrate that this method produces phylogenetic trees that are more accurate than other commonly-used distance based methods though not as accurate as maximum likelihood methods from good quality multiple sequence alignments. In addition to tests on simulated data, we use DendroBLAST to generate input trees for a supertree reconstruction of the phylogeny of the Archaea. This independent analysis produces an approximate phylogeny of the Archaea that has both high precision and recall when compared to previously published analysis of the same dataset using conventional methods. Taken together these results demonstrate that approximate phylogenetic trees can be produced in the absence of multiple sequence alignments, and we propose that these trees will provide a platform for improving and informing downstream bioinformatic analysis. A web implementation of the DendroBLAST method is freely available for use at http://www.dendroblast.com/.

  18. Establishment and application of medication error classification standards in nursing care based on the International Classification of Patient Safety

    Directory of Open Access Journals (Sweden)

    Xiao-Ping Zhu

    2014-09-01

    Conclusion: Application of this classification system will help nursing administrators to accurately detect system- and process-related defects leading to medication errors, and enable the factors to be targeted to improve the level of patient safety management.

  19. Molecular systematics of the Amazonian genus Aldina, a phylogenetically enigmatic ectomycorrhizal lineage of papilionoid legumes.

    Science.gov (United States)

    Ramos, Gustavo; de Lima, Haroldo Cavalcante; Prenner, Gerhard; de Queiroz, Luciano Paganucci; Zartman, Charles E; Cardoso, Domingos

    2016-04-01

    Aldina (Leguminosae) is among the very few ecologically successful ectomycorrhizal lineages in a family largely marked by the evolution of nodulating symbiosis. The genus comprises 20 species predominantly distributed in Amazonia and has been traditionally classified in the tribe Swartzieae because of its radial flowers with an entire calyx and numerous free stamens. The taxonomy of Aldina is complicated due to its poor representation in herbaria and the lack of a robust phylogenetic hypothesis of relationship. Recent phylogenetic analyses of matK and trnL sequences confirmed the placement of Aldina in the 50-kb inversion clade, although the genus remained phylogenetically isolated or unresolved in the context of the evolutionary history of the main early-branching papilionoid lineages. We performed maximum likelihood and Bayesian analyses of combined chloroplast datasets (matK, rbcL, and trnL) and explored the effect of incomplete taxa or missing data in order to shed light on the enigmatic phylogenetic position of Aldina. Unexpectedly, a sister relationship of Aldina with the Andira clade (Andira and Hymenolobium) is revealed. We suggest that a new tribal phylogenetic classification of the papilionoid legumes should place Aldina along with Andira and Hymenolobium. These results highlight yet another example of the independent evolution of radial floral symmetry within the early-branching Papilionoideae, a large collection of florally heterogeneous lineages dominated by papilionate or bilaterally symmetric flower morphology.

  20. [Comparative leaf anatomy and phylogenetic relationships of 11 species of Laeliinae with emphasis on Brassavola (Orchidaceae)].

    Science.gov (United States)

    Noguera-Savelli, Eliana; Jáuregui, Damelis

    2011-09-01

    Brassavola inhabits a wide altitude range and habitat types from Northern Mexico to Northern Argentina. Classification schemes in plants have normally used vegetative and floral characters, but when species are very similar, as in this genus, conflicts arise in species delimitation, and alternative methods should be applied. In this study we explored the taxonomic and phylogenetic value of the anatomical structure of leaves in Brassavola; as ingroup, seven species of Brassavola were considered, and as an outgroup Guarianthe skinneri, Laelia anceps, Rhyncholaelia digbyana and Rhyncholaelia glauca were evaluated. Leaf anatomical characters were studied in freehand cross sections of the middle portion with a light microscope. Ten vegetative anatomical characters were selected and coded for the phylogenetic analysis. Phylogenetic reconstruction was carried out under maximum parsimony using the program NONA through WinClada. Overall, Brassavola species reveal a wide variety of anatomical characters, many of them associated with xeromorphic plants: thick cuticle, hypodermis and cells of the mesophyll with spiral thickenings in the secondary wall. Moreover, mesophyll is either homogeneous or heterogeneous, often with extravascular bundles of fibers near the epidermis at both terete and flat leaves. All vascular bundles are collateral, arranged in more than one row in the mesophyll. The phylogenetic analysis did not resolve internal relationships of the genus; we obtained a polytomy, indicating that the anatomical characters by themselves have little phylogenetic value in Brassavola. We concluded that few anatomical characters are phylogenetically important; however, they would provide more support to elucidate the phylogenetic relantionships in the Orchidaceae and other plant groups if they are used in conjunction with morphological and/or molecular characters.

  1. Phylogenetic mixture models can reduce node-density artifacts.

    Science.gov (United States)

    Venditti, Chris; Meade, Andrew; Pagel, Mark

    2008-04-01

    We investigate the performance of phylogenetic mixture models in reducing a well-known and pervasive artifact of phylogenetic inference known as the node-density effect, comparing them to partitioned analyses of the same data. The node-density effect refers to the tendency for the amount of evolutionary change in longer branches of phylogenies to be underestimated compared to that in regions of the tree where there are more nodes and thus branches are typically shorter. Mixture models allow more than one model of sequence evolution to describe the sites in an alignment without prior knowledge of the evolutionary processes that characterize the data or how they correspond to different sites. If multiple evolutionary patterns are common in sequence evolution, mixture models may be capable of reducing node-density effects by characterizing the evolutionary processes more accurately. In gene-sequence alignments simulated to have heterogeneous patterns of evolution, we find that mixture models can reduce node-density effects to negligible levels or remove them altogether, performing as well as partitioned analyses based on the known simulated patterns. The mixture models achieve this without knowledge of the patterns that generated the data and even in some cases without specifying the full or true model of sequence evolution known to underlie the data. The latter result is especially important in real applications, as the true model of evolution is seldom known. We find the same patterns of results for two real data sets with evidence of complex patterns of sequence evolution: mixture models substantially reduced node-density effects and returned better likelihoods compared to partitioning models specifically fitted to these data. We suggest that the presence of more than one pattern of evolution in the data is a common source of error in phylogenetic inference and that mixture models can often detect these patterns even without prior knowledge of their presence in the

  2. The disentangling number for phylogenetic mixtures

    CERN Document Server

    Sullivant, Seth

    2011-01-01

    We provide a logarithmic upper bound for the disentangling number on unordered lists of leaf labeled trees. This results is useful for analyzing phylogenetic mixture models. The proof depends on interpreting multisets of trees as high dimensional contingency tables.

  3. Efficient and accurate fragmentation methods.

    Science.gov (United States)

    Pruitt, Spencer R; Bertoni, Colleen; Brorsen, Kurt R; Gordon, Mark S

    2014-09-16

    Conspectus Three novel fragmentation methods that are available in the electronic structure program GAMESS (general atomic and molecular electronic structure system) are discussed in this Account. The fragment molecular orbital (FMO) method can be combined with any electronic structure method to perform accurate calculations on large molecular species with no reliance on capping atoms or empirical parameters. The FMO method is highly scalable and can take advantage of massively parallel computer systems. For example, the method has been shown to scale nearly linearly on up to 131 000 processor cores for calculations on large water clusters. There have been many applications of the FMO method to large molecular clusters, to biomolecules (e.g., proteins), and to materials that are used as heterogeneous catalysts. The effective fragment potential (EFP) method is a model potential approach that is fully derived from first principles and has no empirically fitted parameters. Consequently, an EFP can be generated for any molecule by a simple preparatory GAMESS calculation. The EFP method provides accurate descriptions of all types of intermolecular interactions, including Coulombic interactions, polarization/induction, exchange repulsion, dispersion, and charge transfer. The EFP method has been applied successfully to the study of liquid water, π-stacking in substituted benzenes and in DNA base pairs, solvent effects on positive and negative ions, electronic spectra and dynamics, non-adiabatic phenomena in electronic excited states, and nonlinear excited state properties. The effective fragment molecular orbital (EFMO) method is a merger of the FMO and EFP methods, in which interfragment interactions are described by the EFP potential, rather than the less accurate electrostatic potential. The use of EFP in this manner facilitates the use of a smaller value for the distance cut-off (Rcut). Rcut determines the distance at which EFP interactions replace fully quantum

  4. Accurate determination of antenna directivity

    DEFF Research Database (Denmark)

    Dich, Mikael

    1997-01-01

    The derivation of a formula for accurate estimation of the total radiated power from a transmitting antenna for which the radiated power density is known in a finite number of points on the far-field sphere is presented. The main application of the formula is determination of directivity from power......-pattern measurements. The derivation is based on the theory of spherical wave expansion of electromagnetic fields, which also establishes a simple criterion for the required number of samples of the power density. An array antenna consisting of Hertzian dipoles is used to test the accuracy and rate of convergence...

  5. Phylogenetic Distribution of Fungal Sterols

    Science.gov (United States)

    Weete, John D.; Abril, Maritza; Blackwell, Meredith

    2010-01-01

    Background Ergosterol has been considered the “fungal sterol” for almost 125 years; however, additional sterol data superimposed on a recent molecular phylogeny of kingdom Fungi reveals a different and more complex situation. Methodology/Principal Findings The interpretation of sterol distribution data in a modern phylogenetic context indicates that there is a clear trend from cholesterol and other Δ5 sterols in the earliest diverging fungal species to ergosterol in later diverging fungi. There are, however, deviations from this pattern in certain clades. Sterols of the diverse zoosporic and zygosporic forms exhibit structural diversity with cholesterol and 24-ethyl -Δ5 sterols in zoosporic taxa, and 24-methyl sterols in zygosporic fungi. For example, each of the three monophyletic lineages of zygosporic fungi has distinctive major sterols, ergosterol in Mucorales, 22-dihydroergosterol in Dimargaritales, Harpellales, and Kickxellales (DHK clade), and 24-methyl cholesterol in Entomophthorales. Other departures from ergosterol as the dominant sterol include: 24-ethyl cholesterol in Glomeromycota, 24-ethyl cholest-7-enol and 24-ethyl-cholesta-7,24(28)-dienol in rust fungi, brassicasterol in Taphrinales and hypogeous pezizalean species, and cholesterol in Pneumocystis. Conclusions/Significance Five dominant end products of sterol biosynthesis (cholesterol, ergosterol, 24-methyl cholesterol, 24-ethyl cholesterol, brassicasterol), and intermediates in the formation of 24-ethyl cholesterol, are major sterols in 175 species of Fungi. Although most fungi in the most speciose clades have ergosterol as a major sterol, sterols are more varied than currently understood, and their distribution supports certain clades of Fungi in current fungal phylogenies. In addition to the intellectual importance of understanding evolution of sterol synthesis in fungi, there is practical importance because certain antifungal drugs (e.g., azoles) target reactions in the synthesis of

  6. Phylogenetic distribution of fungal sterols.

    Directory of Open Access Journals (Sweden)

    John D Weete

    Full Text Available BACKGROUND: Ergosterol has been considered the "fungal sterol" for almost 125 years; however, additional sterol data superimposed on a recent molecular phylogeny of kingdom Fungi reveals a different and more complex situation. METHODOLOGY/PRINCIPAL FINDINGS: The interpretation of sterol distribution data in a modern phylogenetic context indicates that there is a clear trend from cholesterol and other Delta(5 sterols in the earliest diverging fungal species to ergosterol in later diverging fungi. There are, however, deviations from this pattern in certain clades. Sterols of the diverse zoosporic and zygosporic forms exhibit structural diversity with cholesterol and 24-ethyl -Delta(5 sterols in zoosporic taxa, and 24-methyl sterols in zygosporic fungi. For example, each of the three monophyletic lineages of zygosporic fungi has distinctive major sterols, ergosterol in Mucorales, 22-dihydroergosterol in Dimargaritales, Harpellales, and Kickxellales (DHK clade, and 24-methyl cholesterol in Entomophthorales. Other departures from ergosterol as the dominant sterol include: 24-ethyl cholesterol in Glomeromycota, 24-ethyl cholest-7-enol and 24-ethyl-cholesta-7,24(28-dienol in rust fungi, brassicasterol in Taphrinales and hypogeous pezizalean species, and cholesterol in Pneumocystis. CONCLUSIONS/SIGNIFICANCE: Five dominant end products of sterol biosynthesis (cholesterol, ergosterol, 24-methyl cholesterol, 24-ethyl cholesterol, brassicasterol, and intermediates in the formation of 24-ethyl cholesterol, are major sterols in 175 species of Fungi. Although most fungi in the most speciose clades have ergosterol as a major sterol, sterols are more varied than currently understood, and their distribution supports certain clades of Fungi in current fungal phylogenies. In addition to the intellectual importance of understanding evolution of sterol synthesis in fungi, there is practical importance because certain antifungal drugs (e.g., azoles target reactions in

  7. Phylogenetic approaches to natural product structure prediction.

    Science.gov (United States)

    Ziemert, Nadine; Jensen, Paul R

    2012-01-01

    Phylogenetics is the study of the evolutionary relatedness among groups of organisms. Molecular phylogenetics uses sequence data to infer these relationships for both organisms and the genes they maintain. With the large amount of publicly available sequence data, phylogenetic inference has become increasingly important in all fields of biology. In the case of natural product research, phylogenetic relationships are proving to be highly informative in terms of delineating the architecture and function of the genes involved in secondary metabolite biosynthesis. Polyketide synthases and nonribosomal peptide synthetases provide model examples in which individual domain phylogenies display different predictive capacities, resolving features ranging from substrate specificity to structural motifs associated with the final metabolic product. This chapter provides examples in which phylogeny has proven effective in terms of predicting functional or structural aspects of secondary metabolism. The basics of how to build a reliable phylogenetic tree are explained along with information about programs and tools that can be used for this purpose. Furthermore, it introduces the Natural Product Domain Seeker, a recently developed Web tool that employs phylogenetic logic to classify ketosynthase and condensation domains based on established enzyme architecture and biochemical function.

  8. A practical guide to phylogenetics for nonexperts.

    Science.gov (United States)

    O'Halloran, Damien

    2014-02-05

    Many researchers, across incredibly diverse foci, are applying phylogenetics to their research question(s). However, many researchers are new to this topic and so it presents inherent problems. Here we compile a practical introduction to phylogenetics for nonexperts. We outline in a step-by-step manner, a pipeline for generating reliable phylogenies from gene sequence datasets. We begin with a user-guide for similarity search tools via online interfaces as well as local executables. Next, we explore programs for generating multiple sequence alignments followed by protocols for using software to determine best-fit models of evolution. We then outline protocols for reconstructing phylogenetic relationships via maximum likelihood and Bayesian criteria and finally describe tools for visualizing phylogenetic trees. While this is not by any means an exhaustive description of phylogenetic approaches, it does provide the reader with practical starting information on key software applications commonly utilized by phylogeneticists. The vision for this article would be that it could serve as a practical training tool for researchers embarking on phylogenetic studies and also serve as an educational resource that could be incorporated into a classroom or teaching-lab.

  9. Maximizing the phylogenetic diversity of seed banks.

    Science.gov (United States)

    Griffiths, Kate E; Balding, Sharon T; Dickie, John B; Lewis, Gwilym P; Pearce, Tim R; Grenyer, Richard

    2015-04-01

    Ex situ conservation efforts such as those of zoos, botanical gardens, and seed banks will form a vital complement to in situ conservation actions over the coming decades. It is therefore necessary to pay the same attention to the biological diversity represented in ex situ conservation facilities as is often paid to protected-area networks. Building the phylogenetic diversity of ex situ collections will strengthen our capacity to respond to biodiversity loss. Since 2000, the Millennium Seed Bank Partnership has banked seed from 14% of the world's plant species. We assessed the taxonomic, geographic, and phylogenetic diversity of the Millennium Seed Bank collection of legumes (Leguminosae). We compared the collection with all known legume genera, their known geographic range (at country and regional levels), and a genus-level phylogeny of the legume family constructed for this study. Over half the phylogenetic diversity of legumes at the genus level was represented in the Millennium Seed Bank. However, pragmatic prioritization of species of economic importance and endangerment has led to the banking of a less-than-optimal phylogenetic diversity and prioritization of range-restricted species risks an underdispersed collection. The current state of the phylogenetic diversity of legumes in the Millennium Seed Bank could be substantially improved through the strategic banking of relatively few additional taxa. Our method draws on tools that are widely applied to in situ conservation planning, and it can be used to evaluate and improve the phylogenetic diversity of ex situ collections. © 2014 Society for Conservation Biology.

  10. How does cognition evolve? Phylogenetic comparative psychology

    Science.gov (United States)

    Matthews, Luke J.; Hare, Brian A.; Nunn, Charles L.; Anderson, Rindy C.; Aureli, Filippo; Brannon, Elizabeth M.; Call, Josep; Drea, Christine M.; Emery, Nathan J.; Haun, Daniel B. M.; Herrmann, Esther; Jacobs, Lucia F.; Platt, Michael L.; Rosati, Alexandra G.; Sandel, Aaron A.; Schroepfer, Kara K.; Seed, Amanda M.; Tan, Jingzhi; van Schaik, Carel P.; Wobber, Victoria

    2014-01-01

    Now more than ever animal studies have the potential to test hypotheses regarding how cognition evolves. Comparative psychologists have developed new techniques to probe the cognitive mechanisms underlying animal behavior, and they have become increasingly skillful at adapting methodologies to test multiple species. Meanwhile, evolutionary biologists have generated quantitative approaches to investigate the phylogenetic distribution and function of phenotypic traits, including cognition. In particular, phylogenetic methods can quantitatively (1) test whether specific cognitive abilities are correlated with life history (e.g., lifespan), morphology (e.g., brain size), or socio-ecological variables (e.g., social system), (2) measure how strongly phylogenetic relatedness predicts the distribution of cognitive skills across species, and (3) estimate the ancestral state of a given cognitive trait using measures of cognitive performance from extant species. Phylogenetic methods can also be used to guide the selection of species comparisons that offer the strongest tests of a priori predictions of cognitive evolutionary hypotheses (i.e., phylogenetic targeting). Here, we explain how an integration of comparative psychology and evolutionary biology will answer a host of questions regarding the phylogenetic distribution and history of cognitive traits, as well as the evolutionary processes that drove their evolution. PMID:21927850

  11. How does cognition evolve? Phylogenetic comparative psychology.

    Science.gov (United States)

    MacLean, Evan L; Matthews, Luke J; Hare, Brian A; Nunn, Charles L; Anderson, Rindy C; Aureli, Filippo; Brannon, Elizabeth M; Call, Josep; Drea, Christine M; Emery, Nathan J; Haun, Daniel B M; Herrmann, Esther; Jacobs, Lucia F; Platt, Michael L; Rosati, Alexandra G; Sandel, Aaron A; Schroepfer, Kara K; Seed, Amanda M; Tan, Jingzhi; van Schaik, Carel P; Wobber, Victoria

    2012-03-01

    Now more than ever animal studies have the potential to test hypotheses regarding how cognition evolves. Comparative psychologists have developed new techniques to probe the cognitive mechanisms underlying animal behavior, and they have become increasingly skillful at adapting methodologies to test multiple species. Meanwhile, evolutionary biologists have generated quantitative approaches to investigate the phylogenetic distribution and function of phenotypic traits, including cognition. In particular, phylogenetic methods can quantitatively (1) test whether specific cognitive abilities are correlated with life history (e.g., lifespan), morphology (e.g., brain size), or socio-ecological variables (e.g., social system), (2) measure how strongly phylogenetic relatedness predicts the distribution of cognitive skills across species, and (3) estimate the ancestral state of a given cognitive trait using measures of cognitive performance from extant species. Phylogenetic methods can also be used to guide the selection of species comparisons that offer the strongest tests of a priori predictions of cognitive evolutionary hypotheses (i.e., phylogenetic targeting). Here, we explain how an integration of comparative psychology and evolutionary biology will answer a host of questions regarding the phylogenetic distribution and history of cognitive traits, as well as the evolutionary processes that drove their evolution.

  12. Nodal distances for rooted phylogenetic trees.

    Science.gov (United States)

    Cardona, Gabriel; Llabrés, Mercè; Rosselló, Francesc; Valiente, Gabriel

    2010-08-01

    Dissimilarity measures for (possibly weighted) phylogenetic trees based on the comparison of their vectors of path lengths between pairs of taxa, have been present in the systematics literature since the early seventies. For rooted phylogenetic trees, however, these vectors can only separate non-weighted binary trees, and therefore these dissimilarity measures are metrics only on this class of rooted phylogenetic trees. In this paper we overcome this problem, by splitting in a suitable way each path length between two taxa into two lengths. We prove that the resulting splitted path lengths matrices single out arbitrary rooted phylogenetic trees with nested taxa and arcs weighted in the set of positive real numbers. This allows the definition of metrics on this general class of rooted phylogenetic trees by comparing these matrices through metrics in spaces M(n)(R) of real-valued n x n matrices. We conclude this paper by establishing some basic facts about the metrics for non-weighted phylogenetic trees defined in this way using L(p) metrics on M(n)(R), with p [epsilon] R(>0).

  13. Trends and concepts in fern classification

    Science.gov (United States)

    Christenhusz, Maarten J. M.; Chase, Mark W.

    2014-01-01

    Background and Aims Throughout the history of fern classification, familial and generic concepts have been highly labile. Many classifications and evolutionary schemes have been proposed during the last two centuries, reflecting different interpretations of the available evidence. Knowledge of fern structure and life histories has increased through time, providing more evidence on which to base ideas of possible relationships, and classification has changed accordingly. This paper reviews previous classifications of ferns and presents ideas on how to achieve a more stable consensus. Scope An historical overview is provided from the first to the most recent fern classifications, from which conclusions are drawn on past changes and future trends. The problematic concept of family in ferns is discussed, with a particular focus on how this has changed over time. The history of molecular studies and the most recent findings are also presented. Key Results Fern classification generally shows a trend from highly artificial, based on an interpretation of a few extrinsic characters, via natural classifications derived from a multitude of intrinsic characters, towards more evolutionary circumscriptions of groups that do not in general align well with the distribution of these previously used characters. It also shows a progression from a few broad family concepts to systems that recognized many more narrowly and highly controversially circumscribed families; currently, the number of families recognized is stabilizing somewhere between these extremes. Placement of many genera was uncertain until the arrival of molecular phylogenetics, which has rapidly been improving our understanding of fern relationships. As a collective category, the so-called ‘fern allies’ (e.g. Lycopodiales, Psilotaceae, Equisetaceae) were unsurprisingly found to be polyphyletic, and the term should be abandoned. Lycopodiaceae, Selaginellaceae and Isoëtaceae form a clade (the lycopods) that is

  14. Genome-based Taxonomic Classification of Bacteroidetes

    Directory of Open Access Journals (Sweden)

    Richard L. Hahnke

    2016-12-01

    Full Text Available The bacterial phylum Bacteroidetes, characterized by a distinct gliding motility, occurs in a broad variety of ecosystems, habitats, life styles and physiologies. Accordingly, taxonomic classification of the phylum, based on a limited number of features, proved difficult and controversial in the past, for example, when decisions were based on unresolved phylogenetic trees of the 16S rRNA gene sequence. Here we use a large collection of type-strain genomes from Bacteroidetes and closely related phyla for assessing their taxonomy based on the principles of phylogenetic classification and trees inferred from genome-scale data. No significant conflict between 16S rRNA gene and whole-genome phylogenetic analysis is found, whereas many but not all of the involved taxa are supported as monophyletic groups, particularly in the genome-scale trees. Phenotypic and phylogenomic features support the separation of Balneolaceae as new phylum Balneolaeota from Rhodothermaeota and of Saprospiraceae as new class Saprospiria from Chitinophagia. Epilithonimonas is nested within the older genus Chryseobacterium and without significant phenotypic differences; thus merging the two genera is proposed. Similarly, Vitellibacter is proposed to be included in Aequorivita. Flexibacter is confirmed as being heterogeneous and dissected, yielding six distinct genera. Hallella seregens is a later heterotypic synonym of Prevotella dentalis. Compared to values directly calculated from genome sequences, the G+C content mentioned in many species descriptions is too imprecise; moreover, corrected G+C content values have a significantly better fit to the phylogeny. Corresponding emendations of species descriptions are provided where necessary. Whereas most observed conflict with the current classification of Bacteroidetes is already visible in 16S rRNA gene trees, as expected whole-genome phylogenies are much better resolved.

  15. Genome-Based Taxonomic Classification of Bacteroidetes.

    Science.gov (United States)

    Hahnke, Richard L; Meier-Kolthoff, Jan P; García-López, Marina; Mukherjee, Supratim; Huntemann, Marcel; Ivanova, Natalia N; Woyke, Tanja; Kyrpides, Nikos C; Klenk, Hans-Peter; Göker, Markus

    2016-01-01

    The bacterial phylum Bacteroidetes, characterized by a distinct gliding motility, occurs in a broad variety of ecosystems, habitats, life styles, and physiologies. Accordingly, taxonomic classification of the phylum, based on a limited number of features, proved difficult and controversial in the past, for example, when decisions were based on unresolved phylogenetic trees of the 16S rRNA gene sequence. Here we use a large collection of type-strain genomes from Bacteroidetes and closely related phyla for assessing their taxonomy based on the principles of phylogenetic classification and trees inferred from genome-scale data. No significant conflict between 16S rRNA gene and whole-genome phylogenetic analysis is found, whereas many but not all of the involved taxa are supported as monophyletic groups, particularly in the genome-scale trees. Phenotypic and phylogenomic features support the separation of Balneolaceae as new phylum Balneolaeota from Rhodothermaeota and of Saprospiraceae as new class Saprospiria from Chitinophagia. Epilithonimonas is nested within the older genus Chryseobacterium and without significant phenotypic differences; thus merging the two genera is proposed. Similarly, Vitellibacter is proposed to be included in Aequorivita. Flexibacter is confirmed as being heterogeneous and dissected, yielding six distinct genera. Hallella seregens is a later heterotypic synonym of Prevotella dentalis. Compared to values directly calculated from genome sequences, the G+C content mentioned in many species descriptions is too imprecise; moreover, corrected G+C content values have a significantly better fit to the phylogeny. Corresponding emendations of species descriptions are provided where necessary. Whereas most observed conflict with the current classification of Bacteroidetes is already visible in 16S rRNA gene trees, as expected whole-genome phylogenies are much better resolved.

  16. Aphasia Classification Using Neural Networks

    DEFF Research Database (Denmark)

    Axer, H.; Jantzen, Jan; Berks, G.

    2000-01-01

    of the Aachen Aphasia Test (AAT). First a coarse classification was achieved by using an assessment of spontaneous speech of the patient. This classifier produced correct results in 87% of the test cases. For a second test, data analysis tools were used to select four features out of the 30 available test...... features to yield a more accurate diagnosis. This second classifier produced correct results in 92% of the test cases. This test requires four AAT scores as input for the multilayer perceptron. In practice, the second test requires hours of work on behalf of the clinician, whereas the first test can...

  17. Cluster Based Text Classification Model

    DEFF Research Database (Denmark)

    2011-01-01

    We propose a cluster based classification model for suspicious email detection and other text classification tasks. The text classification tasks comprise many training examples that require a complex classification model. Using clusters for classification makes the model simpler and increases th...... datasets. Our model also outperforms A Decision Cluster Classification (ADCC) and the Decision Cluster Forest Classification (DCFC) models on the Reuters-21578 dataset....

  18. A statistical approach to root system classification.

    Science.gov (United States)

    Bodner, Gernot; Leitner, Daniel; Nakhforoosh, Alireza; Sobotik, Monika; Moder, Karl; Kaul, Hans-Peter

    2013-01-01

    Plant root systems have a key role in ecology and agronomy. In spite of fast increase in root studies, still there is no classification that allows distinguishing among distinctive characteristics within the diversity of rooting strategies. Our hypothesis is that a multivariate approach for "plant functional type" identification in ecology can be applied to the classification of root systems. The classification method presented is based on a data-defined statistical procedure without a priori decision on the classifiers. The study demonstrates that principal component based rooting types provide efficient and meaningful multi-trait classifiers. The classification method is exemplified with simulated root architectures and morphological field data. Simulated root architectures showed that morphological attributes with spatial distribution parameters capture most distinctive features within root system diversity. While developmental type (tap vs. shoot-borne systems) is a strong, but coarse classifier, topological traits provide the most detailed differentiation among distinctive groups. Adequacy of commonly available morphologic traits for classification is supported by field data. Rooting types emerging from measured data, mainly distinguished by diameter/weight and density dominated types. Similarity of root systems within distinctive groups was the joint result of phylogenetic relation and environmental as well as human selection pressure. We concluded that the data-define classification is appropriate for integration of knowledge obtained with different root measurement methods and at various scales. Currently root morphology is the most promising basis for classification due to widely used common measurement protocols. To capture details of root diversity efforts in architectural measurement techniques are essential.

  19. Automated Decision Tree Classification of Corneal Shape

    Science.gov (United States)

    Twa, Michael D.; Parthasarathy, Srinivasan; Roberts, Cynthia; Mahmoud, Ashraf M.; Raasch, Thomas W.; Bullimore, Mark A.

    2011-01-01

    Purpose The volume and complexity of data produced during videokeratography examinations present a challenge of interpretation. As a consequence, results are often analyzed qualitatively by subjective pattern recognition or reduced to comparisons of summary indices. We describe the application of decision tree induction, an automated machine learning classification method, to discriminate between normal and keratoconic corneal shapes in an objective and quantitative way. We then compared this method with other known classification methods. Methods The corneal surface was modeled with a seventh-order Zernike polynomial for 132 normal eyes of 92 subjects and 112 eyes of 71 subjects diagnosed with keratoconus. A decision tree classifier was induced using the C4.5 algorithm, and its classification performance was compared with the modified Rabinowitz–McDonnell index, Schwiegerling’s Z3 index (Z3), Keratoconus Prediction Index (KPI), KISA%, and Cone Location and Magnitude Index using recommended classification thresholds for each method. We also evaluated the area under the receiver operator characteristic (ROC) curve for each classification method. Results Our decision tree classifier performed equal to or better than the other classifiers tested: accuracy was 92% and the area under the ROC curve was 0.97. Our decision tree classifier reduced the information needed to distinguish between normal and keratoconus eyes using four of 36 Zernike polynomial coefficients. The four surface features selected as classification attributes by the decision tree method were inferior elevation, greater sagittal depth, oblique toricity, and trefoil. Conclusions Automated decision tree classification of corneal shape through Zernike polynomials is an accurate quantitative method of classification that is interpretable and can be generated from any instrument platform capable of raw elevation data output. This method of pattern classification is extendable to other classification

  20. Accurate ab initio spin densities

    CERN Document Server

    Boguslawski, Katharina; Legeza, Örs; Reiher, Markus

    2012-01-01

    We present an approach for the calculation of spin density distributions for molecules that require very large active spaces for a qualitatively correct description of their electronic structure. Our approach is based on the density-matrix renormalization group (DMRG) algorithm to calculate the spin density matrix elements as basic quantity for the spatially resolved spin density distribution. The spin density matrix elements are directly determined from the second-quantized elementary operators optimized by the DMRG algorithm. As an analytic convergence criterion for the spin density distribution, we employ our recently developed sampling-reconstruction scheme [J. Chem. Phys. 2011, 134, 224101] to build an accurate complete-active-space configuration-interaction (CASCI) wave function from the optimized matrix product states. The spin density matrix elements can then also be determined as an expectation value employing the reconstructed wave function expansion. Furthermore, the explicit reconstruction of a CA...

  1. The Accurate Particle Tracer Code

    CERN Document Server

    Wang, Yulei; Qin, Hong; Yu, Zhi

    2016-01-01

    The Accurate Particle Tracer (APT) code is designed for large-scale particle simulations on dynamical systems. Based on a large variety of advanced geometric algorithms, APT possesses long-term numerical accuracy and stability, which are critical for solving multi-scale and non-linear problems. Under the well-designed integrated and modularized framework, APT serves as a universal platform for researchers from different fields, such as plasma physics, accelerator physics, space science, fusion energy research, computational mathematics, software engineering, and high-performance computation. The APT code consists of seven main modules, including the I/O module, the initialization module, the particle pusher module, the parallelization module, the field configuration module, the external force-field module, and the extendible module. The I/O module, supported by Lua and Hdf5 projects, provides a user-friendly interface for both numerical simulation and data analysis. A series of new geometric numerical methods...

  2. Accurate Modeling of Advanced Reflectarrays

    DEFF Research Database (Denmark)

    Zhou, Min

    Analysis and optimization methods for the design of advanced printed re ectarrays have been investigated, and the study is focused on developing an accurate and efficient simulation tool. For the analysis, a good compromise between accuracy and efficiency can be obtained using the spectral domain...... to the POT. The GDOT can optimize for the size as well as the orientation and position of arbitrarily shaped array elements. Both co- and cross-polar radiation can be optimized for multiple frequencies, dual polarization, and several feed illuminations. Several contoured beam reflectarrays have been designed...... using the GDOT to demonstrate its capabilities. To verify the accuracy of the GDOT, two offset contoured beam reflectarrays that radiate a high-gain beam on a European coverage have been designed and manufactured, and subsequently measured at the DTU-ESA Spherical Near-Field Antenna Test Facility...

  3. Accurate thickness measurement of graphene

    Science.gov (United States)

    Shearer, Cameron J.; Slattery, Ashley D.; Stapleton, Andrew J.; Shapter, Joseph G.; Gibson, Christopher T.

    2016-03-01

    Graphene has emerged as a material with a vast variety of applications. The electronic, optical and mechanical properties of graphene are strongly influenced by the number of layers present in a sample. As a result, the dimensional characterization of graphene films is crucial, especially with the continued development of new synthesis methods and applications. A number of techniques exist to determine the thickness of graphene films including optical contrast, Raman scattering and scanning probe microscopy techniques. Atomic force microscopy (AFM), in particular, is used extensively since it provides three-dimensional images that enable the measurement of the lateral dimensions of graphene films as well as the thickness, and by extension the number of layers present. However, in the literature AFM has proven to be inaccurate with a wide range of measured values for single layer graphene thickness reported (between 0.4 and 1.7 nm). This discrepancy has been attributed to tip-surface interactions, image feedback settings and surface chemistry. In this work, we use standard and carbon nanotube modified AFM probes and a relatively new AFM imaging mode known as PeakForce tapping mode to establish a protocol that will allow users to accurately determine the thickness of graphene films. In particular, the error in measuring the first layer is reduced from 0.1-1.3 nm to 0.1-0.3 nm. Furthermore, in the process we establish that the graphene-substrate adsorbate layer and imaging force, in particular the pressure the tip exerts on the surface, are crucial components in the accurate measurement of graphene using AFM. These findings can be applied to other 2D materials.

  4. Accurate thickness measurement of graphene.

    Science.gov (United States)

    Shearer, Cameron J; Slattery, Ashley D; Stapleton, Andrew J; Shapter, Joseph G; Gibson, Christopher T

    2016-03-29

    Graphene has emerged as a material with a vast variety of applications. The electronic, optical and mechanical properties of graphene are strongly influenced by the number of layers present in a sample. As a result, the dimensional characterization of graphene films is crucial, especially with the continued development of new synthesis methods and applications. A number of techniques exist to determine the thickness of graphene films including optical contrast, Raman scattering and scanning probe microscopy techniques. Atomic force microscopy (AFM), in particular, is used extensively since it provides three-dimensional images that enable the measurement of the lateral dimensions of graphene films as well as the thickness, and by extension the number of layers present. However, in the literature AFM has proven to be inaccurate with a wide range of measured values for single layer graphene thickness reported (between 0.4 and 1.7 nm). This discrepancy has been attributed to tip-surface interactions, image feedback settings and surface chemistry. In this work, we use standard and carbon nanotube modified AFM probes and a relatively new AFM imaging mode known as PeakForce tapping mode to establish a protocol that will allow users to accurately determine the thickness of graphene films. In particular, the error in measuring the first layer is reduced from 0.1-1.3 nm to 0.1-0.3 nm. Furthermore, in the process we establish that the graphene-substrate adsorbate layer and imaging force, in particular the pressure the tip exerts on the surface, are crucial components in the accurate measurement of graphene using AFM. These findings can be applied to other 2D materials.

  5. Influence of pansharpening techniques in obtaining accurate vegetation thematic maps

    Science.gov (United States)

    Ibarrola-Ulzurrun, Edurne; Gonzalo-Martin, Consuelo; Marcello-Ruiz, Javier

    2016-10-01

    In last decades, there have been a decline in natural resources, becoming important to develop reliable methodologies for their management. The appearance of very high resolution sensors has offered a practical and cost-effective means for a good environmental management. In this context, improvements are needed for obtaining higher quality of the information available in order to get reliable classified images. Thus, pansharpening enhances the spatial resolution of the multispectral band by incorporating information from the panchromatic image. The main goal in the study is to implement pixel and object-based classification techniques applied to the fused imagery using different pansharpening algorithms and the evaluation of thematic maps generated that serve to obtain accurate information for the conservation of natural resources. A vulnerable heterogenic ecosystem from Canary Islands (Spain) was chosen, Teide National Park, and Worldview-2 high resolution imagery was employed. The classes considered of interest were set by the National Park conservation managers. 7 pansharpening techniques (GS, FIHS, HCS, MTF based, Wavelet `à trous' and Weighted Wavelet `à trous' through Fractal Dimension Maps) were chosen in order to improve the data quality with the goal to analyze the vegetation classes. Next, different classification algorithms were applied at pixel-based and object-based approach, moreover, an accuracy assessment of the different thematic maps obtained were performed. The highest classification accuracy was obtained applying Support Vector Machine classifier at object-based approach in the Weighted Wavelet `à trous' through Fractal Dimension Maps fused image. Finally, highlight the difficulty of the classification in Teide ecosystem due to the heterogeneity and the small size of the species. Thus, it is important to obtain accurate thematic maps for further studies in the management and conservation of natural resources.

  6. Phylogenetic and functional assessment of orthologs inference projects and methods.

    Directory of Open Access Journals (Sweden)

    Adrian M Altenhoff

    2009-01-01

    Full Text Available Accurate genome-wide identification of orthologs is a central problem in comparative genomics, a fact reflected by the numerous orthology identification projects developed in recent years. However, only a few reports have compared their accuracy, and indeed, several recent efforts have not yet been systematically evaluated. Furthermore, orthology is typically only assessed in terms of function conservation, despite the phylogeny-based original definition of Fitch. We collected and mapped the results of nine leading orthology projects and methods (COG, KOG, Inparanoid, OrthoMCL, Ensembl Compara, Homologene, RoundUp, EggNOG, and OMA and two standard methods (bidirectional best-hit and reciprocal smallest distance. We systematically compared their predictions with respect to both phylogeny and function, using six different tests. This required the mapping of millions of sequences, the handling of hundreds of millions of predicted pairs of orthologs, and the computation of tens of thousands of trees. In phylogenetic analysis or in functional analysis where high specificity is required, we find that OMA and Homologene perform best. At lower functional specificity but higher coverage level, OrthoMCL outperforms Ensembl Compara, and to a lesser extent Inparanoid. Lastly, the large coverage of the recent EggNOG can be of interest to build broad functional grouping, but the method is not specific enough for phylogenetic or detailed function analyses. In terms of general methodology, we observe that the more sophisticated tree reconstruction/reconciliation approach of Ensembl Compara was at times outperformed by pairwise comparison approaches, even in phylogenetic tests. Furthermore, we show that standard bidirectional best-hit often outperforms projects with more complex algorithms. First, the present study provides guidance for the broad community of orthology data users as to which database best suits their needs. Second, it introduces new methodology

  7. Barcoding and Phylogenetic Inferences in Nine Mugilid Species (Pisces, Mugiliformes

    Directory of Open Access Journals (Sweden)

    Neonila Polyakova

    2013-10-01

    Full Text Available Accurate identification of fish and fish products, from eggs to adults, is important in many areas. Grey mullets of the family Mugilidae are distributed worldwide and inhabit marine, estuarine, and freshwater environments in all tropical and temperate regions. Various Mugilid species are commercially important species in fishery and aquaculture of many countries. For the present study we have chosen two Mugilid genes with different phylogenetic signals: relatively variable mitochondrial cytochrome oxidase subunit I (COI and conservative nuclear rhodopsin (RHO. We examined their diversity within and among 9 Mugilid species belonging to 4 genera, many of which have been examined from multiple specimens, with the goal of determining whether DNA barcoding can achieve unambiguous species recognition of Mugilid species. The data obtained showed that information based on COI sequences was diagnostic not only for species-level identification but also for recognition of intraspecific units, e.g., allopatric populations of circumtropical Mugil cephalus, or even native and acclimatized specimens of Chelon haematocheila. All RHO sequences appeared strictly species specific. Based on the data obtained, we conclude that COI, as well as RHO sequencing can be used to unambiguously identify fish species. Topologies of phylogeny based on RHO and COI sequences coincided with each other, while together they had a good phylogenetic signal.

  8. Inter-rater reliability of the EPUAP pressure ulcer classification system using photographs.

    NARCIS (Netherlands)

    Defloor, T.; Schoonhoven, L.

    2004-01-01

    BACKGROUND: Many classification systems for grading pressure ulcers are discussed in the literature. Correct identification and classification of a pressure ulcer is important for accurate reporting of the magnitude of the problem, and for timely prevention. The reliability of pressure ulcer classif

  9. Typology, classification and systematization of innovative projects and initiatives in the company

    Directory of Open Access Journals (Sweden)

    Baklanova Julia O.

    2012-04-01

    Full Text Available The author presents a comparison of definitions of typology, classification and systematization, and treats them as an example of innovative projects and initiatives of the company. The basis of typology and classification laid methodical Benko K., Mc Farlan. In order to obtain a more accurate result it is necessary to integrate the task typology, classification and systematization.

  10. Typology, classification and systematization of innovative projects and initiatives in the company

    OpenAIRE

    Baklanova Julia O.

    2012-01-01

    The author presents a comparison of definitions of typology, classification and systematization, and treats them as an example of innovative projects and initiatives of the company. The basis of typology and classification laid methodical Benko K., Mc Farlan. In order to obtain a more accurate result it is necessary to integrate the task typology, classification and systematization.

  11. Classification of cultivated plants.

    NARCIS (Netherlands)

    Brandenburg, W.A.

    1986-01-01

    Agricultural practice demands principles for classification, starting from the basal entity in cultivated plants: the cultivar. In establishing biosystematic relationships between wild, weedy and cultivated plants, the species concept needs re-examination. Combining of botanic classification, based

  12. Aircraft Operations Classification System

    Science.gov (United States)

    Harlow, Charles; Zhu, Weihong

    2001-01-01

    Accurate data is important in the aviation planning process. In this project we consider systems for measuring aircraft activity at airports. This would include determining the type of aircraft such as jet, helicopter, single engine, and multiengine propeller. Some of the issues involved in deploying technologies for monitoring aircraft operations are cost, reliability, and accuracy. In addition, the system must be field portable and acceptable at airports. A comparison of technologies was conducted and it was decided that an aircraft monitoring system should be based upon acoustic technology. A multimedia relational database was established for the study. The information contained in the database consists of airport information, runway information, acoustic records, photographic records, a description of the event (takeoff, landing), aircraft type, and environmental information. We extracted features from the time signal and the frequency content of the signal. A multi-layer feed-forward neural network was chosen as the classifier. Training and testing results were obtained. We were able to obtain classification results of over 90 percent for training and testing for takeoff events.

  13. Increased taxon sampling greatly reduces phylogenetic error.

    Science.gov (United States)

    Zwickl, Derrick J; Hillis, David M

    2002-08-01

    Several authors have argued recently that extensive taxon sampling has a positive and important effect on the accuracy of phylogenetic estimates. However, other authors have argued that there is little benefit of extensive taxon sampling, and so phylogenetic problems can or should be reduced to a few exemplar taxa as a means of reducing the computational complexity of the phylogenetic analysis. In this paper we examined five aspects of study design that may have led to these different perspectives. First, we considered the measurement of phylogenetic error across a wide range of taxon sample sizes, and conclude that the expected error based on randomly selecting trees (which varies by taxon sample size) must be considered in evaluating error in studies of the effects of taxon sampling. Second, we addressed the scope of the phylogenetic problems defined by different samples of taxa, and argue that phylogenetic scope needs to be considered in evaluating the importance of taxon-sampling strategies. Third, we examined the claim that fast and simple tree searches are as effective as more thorough searches at finding near-optimal trees that minimize error. We show that a more complete search of tree space reduces phylogenetic error, especially as the taxon sample size increases. Fourth, we examined the effects of simple versus complex simulation models on taxonomic sampling studies. Although benefits of taxon sampling are apparent for all models, data generated under more complex models of evolution produce higher overall levels of error and show greater positive effects of increased taxon sampling. Fifth, we asked if different phylogenetic optimality criteria show different effects of taxon sampling. Although we found strong differences in effectiveness of different optimality criteria as a function of taxon sample size, increased taxon sampling improved the results from all the common optimality criteria. Nonetheless, the method that showed the lowest overall

  14. Primate molecular phylogenetics in a genomic era.

    Science.gov (United States)

    Ting, Nelson; Sterner, Kirstin N

    2013-02-01

    A primary objective of molecular phylogenetics is to use molecular data to elucidate the evolutionary history of living organisms. Dr. Morris Goodman founded the journal Molecular Phylogenetics and Evolution as a forum where scientists could further our knowledge about the tree of life, and he recognized that the inference of species trees is a first and fundamental step to addressing many important evolutionary questions. In particular, Dr. Goodman was interested in obtaining a complete picture of the primate species tree in order to provide an evolutionary context for the study of human adaptations. A number of recent studies use multi-locus datasets to infer well-resolved and well-supported primate phylogenetic trees using consensus approaches (e.g., supermatrices). It is therefore tempting to assume that we have a complete picture of the primate tree, especially above the species level. However, recent theoretical and empirical work in the field of molecular phylogenetics demonstrates that consensus methods might provide a false sense of support at certain nodes. In this brief review we discuss the current state of primate molecular phylogenetics and highlight the importance of exploring the use of coalescent-based analyses that have the potential to better utilize information contained in multi-locus data.

  15. Worldwide phylogenetic relationship of avian poxviruses

    Science.gov (United States)

    Gyuranecz, Miklós; Foster, Jeffrey T.; Dán, Ádám; Ip, Hon S.; Egstad, Kristina F.; Parker, Patricia G.; Higashiguchi, Jenni M.; Skinner, Michael A.; Höfle, Ursula; Kreizinger, Zsuzsa; Dorrestein, Gerry M.; Solt, Szabolcs; Sós, Endre; Kim, Young Jun; Uhart, Marcela; Pereda, Ariel; González-Hein, Gisela; Hidalgo, Hector; Blanco, Juan-Manuel; Erdélyi, Károly

    2013-01-01

    Poxvirus infections have been found in 230 species of wild and domestic birds worldwide in both terrestrial and marine environments. This ubiquity raises the question of how infection has been transmitted and globally dispersed. We present a comprehensive global phylogeny of 111 novel poxvirus isolates in addition to all available sequences from GenBank. Phylogenetic analysis of Avipoxvirus genus has traditionally relied on one gene region (4b core protein). In this study we have expanded the analyses to include a second locus (DNA polymerase gene), allowing for a more robust phylogenetic framework, finer genetic resolution within specific groups and the detection of potential recombination. Our phylogenetic results reveal several major features of avipoxvirus evolution and ecology and propose an updated avipoxvirus taxonomy, including three novel subclades. The characterization of poxviruses from 57 species of birds in this study extends the current knowledge of their host range and provides the first evidence of the phylogenetic effect of genetic recombination of avipoxviruses. The repeated occurrence of avian family or order-specific grouping within certain clades (e.g. starling poxvirus, falcon poxvirus, raptor poxvirus, etc.) indicates a marked role of host adaptation, while the sharing of poxvirus species within prey-predator systems emphasizes the capacity for cross-species infection and limited host adaptation. Our study provides a broad and comprehensive phylogenetic analysis of the Avipoxvirus genus, an ecologically and environmentally important viral group, to formulate a genome sequencing strategy that will clarify avipoxvirus taxonomy.

  16. Fourier transform inequalities for phylogenetic trees.

    Science.gov (United States)

    Matsen, Frederick A

    2009-01-01

    Phylogenetic invariants are not the only constraints on site-pattern frequency vectors for phylogenetic trees. A mutation matrix, by its definition, is the exponential of a matrix with non-negative off-diagonal entries; this positivity requirement implies non-trivial constraints on the site-pattern frequency vectors. We call these additional constraints "edge-parameter inequalities". In this paper, we first motivate the edge-parameter inequalities by considering a pathological site-pattern frequency vector corresponding to a quartet tree with a negative internal edge. This site-pattern frequency vector nevertheless satisfies all of the constraints described up to now in the literature. We next describe two complete sets of edge-parameter inequalities for the group-based models; these constraints are square-free monomial inequalities in the Fourier transformed coordinates. These inequalities, along with the phylogenetic invariants, form a complete description of the set of site-pattern frequency vectors corresponding to bona fide trees. Said in mathematical language, this paper explicitly presents two finite lists of inequalities in Fourier coordinates of the form "monomial < or = 1", each list characterizing the phylogenetically relevant semialgebraic subsets of the phylogenetic varieties.

  17. Cirrhosis Classification Based on Texture Classification of Random Features

    Directory of Open Access Journals (Sweden)

    Hui Liu

    2014-01-01

    Full Text Available Accurate staging of hepatic cirrhosis is important in investigating the cause and slowing down the effects of cirrhosis. Computer-aided diagnosis (CAD can provide doctors with an alternative second opinion and assist them to make a specific treatment with accurate cirrhosis stage. MRI has many advantages, including high resolution for soft tissue, no radiation, and multiparameters imaging modalities. So in this paper, multisequences MRIs, including T1-weighted, T2-weighted, arterial, portal venous, and equilibrium phase, are applied. However, CAD does not meet the clinical needs of cirrhosis and few researchers are concerned with it at present. Cirrhosis is characterized by the presence of widespread fibrosis and regenerative nodules in the hepatic, leading to different texture patterns of different stages. So, extracting texture feature is the primary task. Compared with typical gray level cooccurrence matrix (GLCM features, texture classification from random features provides an effective way, and we adopt it and propose CCTCRF for triple classification (normal, early, and middle and advanced stage. CCTCRF does not need strong assumptions except the sparse character of image, contains sufficient texture information, includes concise and effective process, and makes case decision with high accuracy. Experimental results also illustrate the satisfying performance and they are also compared with typical NN with GLCM.

  18. Cirrhosis classification based on texture classification of random features.

    Science.gov (United States)

    Liu, Hui; Shao, Ying; Guo, Dongmei; Zheng, Yuanjie; Zhao, Zuowei; Qiu, Tianshuang

    2014-01-01

    Accurate staging of hepatic cirrhosis is important in investigating the cause and slowing down the effects of cirrhosis. Computer-aided diagnosis (CAD) can provide doctors with an alternative second opinion and assist them to make a specific treatment with accurate cirrhosis stage. MRI has many advantages, including high resolution for soft tissue, no radiation, and multiparameters imaging modalities. So in this paper, multisequences MRIs, including T1-weighted, T2-weighted, arterial, portal venous, and equilibrium phase, are applied. However, CAD does not meet the clinical needs of cirrhosis and few researchers are concerned with it at present. Cirrhosis is characterized by the presence of widespread fibrosis and regenerative nodules in the hepatic, leading to different texture patterns of different stages. So, extracting texture feature is the primary task. Compared with typical gray level cooccurrence matrix (GLCM) features, texture classification from random features provides an effective way, and we adopt it and propose CCTCRF for triple classification (normal, early, and middle and advanced stage). CCTCRF does not need strong assumptions except the sparse character of image, contains sufficient texture information, includes concise and effective process, and makes case decision with high accuracy. Experimental results also illustrate the satisfying performance and they are also compared with typical NN with GLCM.

  19. Teaching Molecular Phylogenetics through Investigating a Real-World Phylogenetic Problem

    Science.gov (United States)

    Zhang, Xiaorong

    2012-01-01

    A phylogenetics exercise is incorporated into the "Introduction to biocomputing" course, a junior-level course at Savannah State University. This exercise is designed to help students learn important concepts and practical skills in molecular phylogenetics through solving a real-world problem. In this application, students are required to identify…

  20. The Evolutionary Ecology of Plant Disease: A Phylogenetic Perspective.

    Science.gov (United States)

    Gilbert, Gregory S; Parker, Ingrid M

    2016-08-04

    An explicit phylogenetic perspective provides useful tools for phytopathology and plant disease ecology because the traits of both plants and microbes are shaped by their evolutionary histories. We present brief primers on phylogenetic signal and the analytical tools of phylogenetic ecology. We review the literature and find abundant evidence of phylogenetic signal in pathogens and plants for most traits involved in disease interactions. Plant nonhost resistance mechanisms and pathogen housekeeping functions are conserved at deeper phylogenetic levels, whereas molecular traits associated with rapid coevolutionary dynamics are more labile at branch tips. Horizontal gene transfer disrupts the phylogenetic signal for some microbial traits. Emergent traits, such as host range and disease severity, show clear phylogenetic signals. Therefore pathogen spread and disease impact are influenced by the phylogenetic structure of host assemblages. Phylogenetically rare species escape disease pressure. Phylogenetic tools could be used to develop predictive tools for phytosanitary risk analysis and reduce disease pressure in multispecies cropping systems.

  1. Random forest classification of etiologies for an orphan disease.

    Science.gov (United States)

    Speiser, Jaime Lynn; Durkalski, Valerie L; Lee, William M

    2015-02-28

    Classification of objects into pre-defined groups based on known information is a fundamental problem in the field of statistics. Although approaches for solving this problem exist, finding an accurate classification method can be challenging in an orphan disease setting, where data are minimal and often not normally distributed. The purpose of this paper is to illustrate the application of the random forest (RF) classification procedure in a real clinical setting and discuss typical questions that arise in the general classification framework as well as offer interpretations of RF results. This paper includes methods for assessing predictive performance, importance of predictor variables, and observation-specific information.

  2. A More Accurate Fourier Transform

    CERN Document Server

    Courtney, Elya

    2015-01-01

    Fourier transform methods are used to analyze functions and data sets to provide frequencies, amplitudes, and phases of underlying oscillatory components. Fast Fourier transform (FFT) methods offer speed advantages over evaluation of explicit integrals (EI) that define Fourier transforms. This paper compares frequency, amplitude, and phase accuracy of the two methods for well resolved peaks over a wide array of data sets including cosine series with and without random noise and a variety of physical data sets, including atmospheric $\\mathrm{CO_2}$ concentrations, tides, temperatures, sound waveforms, and atomic spectra. The FFT uses MIT's FFTW3 library. The EI method uses the rectangle method to compute the areas under the curve via complex math. Results support the hypothesis that EI methods are more accurate than FFT methods. Errors range from 5 to 10 times higher when determining peak frequency by FFT, 1.4 to 60 times higher for peak amplitude, and 6 to 10 times higher for phase under a peak. The ability t...

  3. SUMAC: Constructing Phylogenetic Supermatrices and Assessing Partially Decisive Taxon Coverage

    OpenAIRE

    William A. Freyman

    2015-01-01

    The amount of phylogenetically informative sequence data in GenBank is growing at an exponential rate, and large phylogenetic trees are increasingly used in research. Tools are needed to construct phylogenetic sequence matrices from GenBank data and evaluate the effect of missing data. Supermatrix Constructor (SUMAC) is a tool to data-mine GenBank, construct phylogenetic supermatrices, and assess the phylogenetic decisiveness of a matrix given the pattern of missing sequence data. SUMAC calcu...

  4. Phylogenetic analysis of the Trypanosoma genus based on the heat-shock protein 70 gene.

    Science.gov (United States)

    Fraga, Jorge; Fernández-Calienes, Aymé; Montalvo, Ana Margarita; Maes, Ilse; Deborggraeve, Stijn; Büscher, Philippe; Dujardin, Jean-Claude; Van der Auwera, Gert

    2016-09-01

    Trypanosome evolution was so far essentially studied on the basis of phylogenetic analyses of small subunit ribosomal RNA (SSU-rRNA) and glycosomal glyceraldehyde-3-phosphate dehydrogenase (gGAPDH) genes. We used for the first time the 70kDa heat-shock protein gene (hsp70) to investigate the phylogenetic relationships among 11 Trypanosoma species on the basis of 1380 nucleotides from 76 sequences corresponding to 65 strains. We also constructed a phylogeny based on combined datasets of SSU-rDNA, gGAPDH and hsp70 sequences. The obtained clusters can be correlated with the sections and subgenus classifications of mammal-infecting trypanosomes except for Trypanosoma theileri and Trypanosoma rangeli. Our analysis supports the classification of Trypanosoma species into clades rather than in sections and subgenera, some of which being polyphyletic. Nine clades were recognized: Trypanosoma carassi, Trypanosoma congolense, Trypanosoma cruzi, Trypanosoma grayi, Trypanosoma lewisi, T. rangeli, T. theileri, Trypanosoma vivax and Trypanozoon. These results are consistent with existing knowledge of the genus' phylogeny. Within the T. cruzi clade, three groups of T. cruzi discrete typing units could be clearly distinguished, corresponding to TcI, TcIII, and TcII+V+VI, while support for TcIV was lacking. Phylogenetic analyses based on hsp70 demonstrated that this molecular marker can be applied for discriminating most of the Trypanosoma species and clades.

  5. PhyTB: Phylogenetic tree visualisation and sample positioning for M. tuberculosis

    KAUST Repository

    Benavente, Ernest D

    2015-05-13

    Background Phylogenetic-based classification of M. tuberculosis and other bacterial genomes is a core analysis for studying evolutionary hypotheses, disease outbreaks and transmission events. Whole genome sequencing is providing new insights into the genomic variation underlying intra- and inter-strain diversity, thereby assisting with the classification and molecular barcoding of the bacteria. One roadblock to strain investigation is the lack of user-interactive solutions to interrogate and visualise variation within a phylogenetic tree setting. Results We have developed a web-based tool called PhyTB (http://pathogenseq.lshtm.ac.uk/phytblive/index.php webcite) to assist phylogenetic tree visualisation and identification of M. tuberculosis clade-informative polymorphism. Variant Call Format files can be uploaded to determine a sample position within the tree. A map view summarises the geographical distribution of alleles and strain-types. The utility of the PhyTB is demonstrated on sequence data from 1,601 M. tuberculosis isolates. Conclusion PhyTB contextualises M. tuberculosis genomic variation within epidemiological, geographical and phylogenic settings. Further tool utility is possible by incorporating large variants and phenotypic data (e.g. drug-resistance profiles), and an assessment of genotype-phenotype associations. Source code is available to develop similar websites for other organisms (http://sourceforge.net/projects/phylotrack webcite).

  6. ScripTree: scripting phylogenetic graphics.

    Science.gov (United States)

    Chevenet, François; Croce, Olivier; Hebrard, Maxime; Christen, Richard; Berry, Vincent

    2010-04-15

    There is a large amount of tools for interactive display of phylogenetic trees. However, there is a shortage of tools for the automation of tree rendering. Scripting phylogenetic graphics would enable the saving of graphical analyses involving numerous and complex tree handling operations and would allow the automation of repetitive tasks. ScripTree is a tool intended to fill this gap. It is an interpreter to be used in batch mode. Phylogenetic graphics instructions, related to tree rendering as well as tree annotation, are stored in a text file and processed in a sequential way. ScripTree can be used online or downloaded at www.scriptree.org, under the GPL license. ScripTree, written in Tcl/Tk, is a cross-platform application available for Windows and Unix-like systems including OS X. It can be used either as a stand-alone package or included in a bioinformatic pipeline and linked to a HTTP server.

  7. Phylogenetics, evolution, and medical importance of polyomaviruses.

    Science.gov (United States)

    Krumbholz, Andi; Bininda-Emonds, Olaf R P; Wutzler, Peter; Zell, Roland

    2009-09-01

    The increasing frequency of tissue transplantation, recent progress in the development and application of immunomodulators, and the depressingly high number of AIDS patients worldwide have placed human polyomaviruses, a group of pathogens that can become reactivated under the status of immunosuppression, suddenly in the spotlight. Since the first description of a polyomavirus a half-century ago in 1953, a multiplicity of human and animal polyomaviruses have been discovered. After reviewing the history of research into this group, with a special focus is made on the clinical importance of human polyomaviruses, we conclude by elucidating the phylogenetic relationships and thus evolutionary history of these viruses. Our phylogenetic analyses are based on all available putative polyomavirus species as well as including all subtypes, subgroups, and (sub)lineages of the human BK and JC polyomaviruses. Finally, we reveal that the hypothesis of a strict codivergence of polyomaviruses with their respective hosts does not represent a realistic assumption in light of phylogenetic findings presented here.

  8. Phylogenetic structure in tropical hummingbird communities

    DEFF Research Database (Denmark)

    Graham, Catherine H; Parra, Juan L; Rahbek, Carsten;

    2009-01-01

    composition of 189 hummingbird communities in Ecuador. We assessed how species and phylogenetic composition changed along environmental gradients and across biogeographic barriers. We show that humid, low-elevation communities are phylogenetically overdispersed (coexistence of distant relatives), a pattern...... an expensive means of locomotion at high elevations. We found that communities in the lowlands on opposite sides of the Andes tend to be phylogenetically similar despite their large differences in species composition, a pattern implicating the Andes as an important dispersal barrier. In contrast, along...... the steep environmental gradient between the lowlands and the Andes we found evidence that species turnover is comprised of relatively distantly related species. The integration of local and regional patterns of diversity across environmental gradients and biogeographic barriers provides insight...

  9. Phylogenetic invariants for group-based models

    CERN Document Server

    Donten-Bury, Maria

    2010-01-01

    In this paper we investigate properties of algebraic varieties representing group-based phylogenetic models. We give the (first) example of a nonnormal general group-based model for an abelian group. Following Kaie Kubjas we also determine some invariants of group-based models showing that the associated varieties do not have to be deformation equivalent. We propose a method of generating many phylogenetic invariants and in particular we show that our approach gives the whole ideal of the claw tree for 3-Kimura model under the assumption of the conjecture of Sturmfels and Sullivant. This, combined with the results of Sturmfels and Sullivant, would enable to determine all phylogenetic invariants for any tree for 3-Kimura model and possibly for other group-based models.

  10. Morphological and molecular convergences in mammalian phylogenetics.

    Science.gov (United States)

    Zou, Zhengting; Zhang, Jianzhi

    2016-09-02

    Phylogenetic trees reconstructed from molecular sequences are often considered more reliable than those reconstructed from morphological characters, in part because convergent evolution, which confounds phylogenetic reconstruction, is believed to be rarer for molecular sequences than for morphologies. However, neither the validity of this belief nor its underlying cause is known. Here comparing thousands of characters of each type that have been used for inferring the phylogeny of mammals, we find that on average morphological characters indeed experience much more convergences than amino acid sites, but this disparity is explained by fewer states per character rather than an intrinsically higher susceptibility to convergence for morphologies than sequences. We show by computer simulation and actual data analysis that a simple method for identifying and removing convergence-prone characters improves phylogenetic accuracy, potentially enabling, when necessary, the inclusion of morphologies and hence fossils for reliable tree inference.

  11. Visualizing Phylogenetic Treespace Using Cartographic Projections

    Science.gov (United States)

    Sundberg, Kenneth; Clement, Mark; Snell, Quinn

    Phylogenetic analysis is becoming an increasingly important tool for biological research. Applications include epidemiological studies, drug development, and evolutionary analysis. Phylogenetic search is a known NP-Hard problem. The size of the data sets which can be analyzed is limited by the exponential growth in the number of trees that must be considered as the problem size increases. A better understanding of the problem space could lead to better methods, which in turn could lead to the feasible analysis of more data sets. We present a definition of phylogenetic tree space and a visualization of this space that shows significant exploitable structure. This structure can be used to develop search methods capable of handling much larger datasets.

  12. Molecular phylogenetics of the hummingbird genus Coeligena.

    Science.gov (United States)

    Parra, Juan Luis; Remsen, J V; Alvarez-Rebolledo, Mauricio; McGuire, Jimmy A

    2009-11-01

    Advances in the understanding of biological radiations along tropical mountains depend on the knowledge of phylogenetic relationships among species. Here we present a species-level molecular phylogeny based on a multilocus dataset for the Andean hummingbird genus Coeligena. We compare this phylogeny to previous hypotheses of evolutionary relationships and use it as a framework to understand patterns in the evolution of sexual dichromatism and in the biogeography of speciation within the Andes. Previous phylogenetic hypotheses based mostly on similarities in coloration conflicted with our molecular phylogeny, emphasizing the unreliability of color characters for phylogenetic inference. Two major clades, one monochromatic and the other dichromatic, were found in Coeligena. Closely related species were either allopatric or parapatric on opposite mountain slopes. No sister lineages replaced each other along an elevational gradient. Our results indicate the importance of geographic isolation for speciation in this group and the potential interaction between isolation and sexual selection to promote diversification.

  13. Phylogenetic inference under varying proportions of indel-induced alignment gaps

    Directory of Open Access Journals (Sweden)

    Gadagkar Sudhindra R

    2009-08-01

    Full Text Available Abstract Background The effect of alignment gaps on phylogenetic accuracy has been the subject of numerous studies. In this study, we investigated the relationship between the total number of gapped sites and phylogenetic accuracy, when the gaps were introduced (by means of computer simulation to reflect indel (insertion/deletion events during the evolution of DNA sequences. The resulting (true alignments were subjected to commonly used gap treatment and phylogenetic inference methods. Results (1 In general, there was a strong – almost deterministic – relationship between the amount of gap in the data and the level of phylogenetic accuracy when the alignments were very "gappy", (2 gaps resulting from deletions (as opposed to insertions contributed more to the inaccuracy of phylogenetic inference, (3 the probabilistic methods (Bayesian, PhyML & "MLε, " a method implemented in DNAML in PHYLIP performed better at most levels of gap percentage when compared to parsimony (MP and distance (NJ methods, with Bayesian analysis being clearly the best, (4 methods that treat gapped sites as missing data yielded less accurate trees when compared to those that attribute phylogenetic signal to the gapped sites (by coding them as binary character data – presence/absence, or as in the MLε method, and (5 in general, the accuracy of phylogenetic inference depended upon the amount of available data when the gaps resulted from mainly deletion events, and the amount of missing data when insertion events were equally likely to have caused the alignment gaps. Conclusion When gaps in an alignment are a consequence of indel events in the evolution of the sequences, the accuracy of phylogenetic analysis is likely to improve if: (1 alignment gaps are categorized as arising from insertion events or deletion events and then treated separately in the analysis, (2 the evolutionary signal provided by indels is harnessed in the phylogenetic analysis, and (3 methods that

  14. Gallstone Classification in Western Countries.

    Science.gov (United States)

    Cariati, Andrea

    2015-12-01

    In order to compare gallstone disease data from India and Asian countries with Western countries, it is fundamental to follow a common gallstone classification. Gallstone disease has afflicted humans since the time of Egyptian kings, and gallstones have been found during autopsies on mummies. Gallstone prevalence in adult population ranges from 10 to 15 %. Gallstones in Western countries are distinguished into the following classes: cholesterol gallstones that contain more than 50 % of cholesterol (nearly 75 % of gallstones) and pigment gallstones that contain less than 30 % of cholesterol by weight, which can be subdivided into black pigment gallstones and brown pigment gallstones. It has been shown that ultrastructural analysis with scanning electron microscopy is useful in the classification and study of pigment gallstones. Moreover, x-ray diffractometry analysis and infrared spectroscopy of gallstones are of fundamental importance for an accurate stone analysis. An accurate study of gallstones is useful to understand gallstone pathogenesis. In fact, bacteria are not important in cholesterol gallstone nucleation and growth, but they are important in brown pigment gallstone formation. On the contrary, calcium bilirubinate is fundamental in black pigment gallstone formation and probably also plays an important role in cholesterol gallstone nucleation and growth.

  15. Probabilistic graphical model representation in phylogenetics.

    Science.gov (United States)

    Höhna, Sebastian; Heath, Tracy A; Boussau, Bastien; Landis, Michael J; Ronquist, Fredrik; Huelsenbeck, John P

    2014-09-01

    Recent years have seen a rapid expansion of the model space explored in statistical phylogenetics, emphasizing the need for new approaches to statistical model representation and software development. Clear communication and representation of the chosen model is crucial for: (i) reproducibility of an analysis, (ii) model development, and (iii) software design. Moreover, a unified, clear and understandable framework for model representation lowers the barrier for beginners and nonspecialists to grasp complex phylogenetic models, including their assumptions and parameter/variable dependencies. Graphical modeling is a unifying framework that has gained in popularity in the statistical literature in recent years. The core idea is to break complex models into conditionally independent distributions. The strength lies in the comprehensibility, flexibility, and adaptability of this formalism, and the large body of computational work based on it. Graphical models are well-suited to teach statistical models, to facilitate communication among phylogeneticists and in the development of generic software for simulation and statistical inference. Here, we provide an introduction to graphical models for phylogeneticists and extend the standard graphical model representation to the realm of phylogenetics. We introduce a new graphical model component, tree plates, to capture the changing structure of the subgraph corresponding to a phylogenetic tree. We describe a range of phylogenetic models using the graphical model framework and introduce modules to simplify the representation of standard components in large and complex models. Phylogenetic model graphs can be readily used in simulation, maximum likelihood inference, and Bayesian inference using, for example, Metropolis-Hastings or Gibbs sampling of the posterior distribution.

  16. Classification of refrigerants; Classification des fluides frigorigenes

    Energy Technology Data Exchange (ETDEWEB)

    NONE

    2001-07-01

    This document was made from the US standard ANSI/ASHRAE 34 published in 2001 and entitled 'designation and safety classification of refrigerants'. This classification allows to clearly organize in an international way the overall refrigerants used in the world thanks to a codification of the refrigerants in correspondence with their chemical composition. This note explains this codification: prefix, suffixes (hydrocarbons and derived fluids, azeotropic and non-azeotropic mixtures, various organic compounds, non-organic compounds), safety classification (toxicity, flammability, case of mixtures). (J.S.)

  17. Marine turtle mitogenome phylogenetics and evolution

    DEFF Research Database (Denmark)

    Duchene, Sebastián; Frey, Amy; Alfaro-Núñez, Luis Alonso

    2012-01-01

    . Analyses of partial mitochondrial sequences and some nuclear markers have revealed phylogenetic inconsistencies within Cheloniidae, especially regarding the placement of the flatback. Population genetic studies based on D-Loop sequences have shown considerable structuring in species with broad geographic...... to assess sea-turtle evolution with a large molecular dataset. We found variation in the length of the ATP8 gene and a highly variable site in ND4 near a proton translocation channel in the resulting protein. Complete mitogenomes show strong support and resolution for phylogenetic relationships among all...

  18. Phylogenetic Analysis of Mitochondrial Outer Membrane β-Barrel Channels

    Science.gov (United States)

    Wojtkowska, Małgorzata; Jąkalski, Marcin; Pieńkowska, Joanna R.; Stobienia, Olgierd; Karachitos, Andonis; Przytycka, Teresa M.; Weiner, January; Kmita, Hanna; Makałowski, Wojciech

    2012-01-01

    Transport of molecules across mitochondrial outer membrane is pivotal for a proper function of mitochondria. The transport pathways across the membrane are formed by ion channels that participate in metabolite exchange between mitochondria and cytoplasm (voltage-dependent anion-selective channel, VDAC) as well as in import of proteins encoded by nuclear genes (Tom40 and Sam50/Tob55). VDAC, Tom40, and Sam50/Tob55 are present in all eukaryotic organisms, encoded in the nuclear genome, and have β-barrel topology. We have compiled data sets of these protein sequences and studied their phylogenetic relationships with a special focus on the position of Amoebozoa. Additionally, we identified these protein-coding genes in Acanthamoeba castellanii and Dictyostelium discoideum to complement our data set and verify the phylogenetic position of these model organisms. Our analysis show that mitochondrial β-barrel channels from Archaeplastida (plants) and Opisthokonta (animals and fungi) experienced many duplication events that resulted in multiple paralogous isoforms and form well-defined monophyletic clades that match the current model of eukaryotic evolution. However, in representatives of Amoebozoa, Chromalveolata, and Excavata (former Protista), they do not form clearly distinguishable clades, although they locate basally to the plant and algae branches. In most cases, they do not posses paralogs and their sequences appear to have evolved quickly or degenerated. Consequently, the obtained phylogenies of mitochondrial outer membrane β-channels do not entirely reflect the recent eukaryotic classification system involving the six supergroups: Chromalveolata, Excavata, Archaeplastida, Rhizaria, Amoebozoa, and Opisthokonta. PMID:22155732

  19. Classification, disease, and diagnosis.

    Science.gov (United States)

    Jutel, Annemarie

    2011-01-01

    Classification shapes medicine and guides its practice. Understanding classification must be part of the quest to better understand the social context and implications of diagnosis. Classifications are part of the human work that provides a foundation for the recognition and study of illness: deciding how the vast expanse of nature can be partitioned into meaningful chunks, stabilizing and structuring what is otherwise disordered. This article explores the aims of classification, their embodiment in medical diagnosis, and the historical traditions of medical classification. It provides a brief overview of the aims and principles of classification and their relevance to contemporary medicine. It also demonstrates how classifications operate as social framing devices that enable and disable communication, assert and refute authority, and are important items for sociological study.

  20. Phylogenetic modeling of heterogeneous gene-expression microarray data from cancerous specimens.

    Science.gov (United States)

    Abu-Asab, Mones S; Chaouchi, Mohamed; Amri, Hakima

    2008-09-01

    The qualitative dimension of gene expression data and its heterogeneous nature in cancerous specimens can be accounted for by phylogenetic modeling that incorporates the directionality of altered gene expressions, complex patterns of expressions among a group of specimens, and data-based rather than specimen-based gene linkage. Our phylogenetic modeling approach is a double algorithmic technique that includes polarity assessment that brings out the qualitative value of the data, followed by maximum parsimony analysis that is most suitable for the data heterogeneity of cancer gene expression. We demonstrate that polarity assessment of expression values into derived and ancestral states, via outgroup comparison, reduces experimental noise; reveals dichotomously expressed asynchronous genes; and allows data pooling as well as comparability of intra- and interplatforms. Parsimony phylogenetic analysis of the polarized values produces a multidimensional classification of specimens into clades that reveal shared derived gene expressions (the synapomorphies); provides better assessment of ontogenic pathways and phyletic relatedness of specimens; efficiently utilizes dichotomously expressed genes; produces highly predictive class recognition; illustrates gene linkage and multiple developmental pathways; provides higher concordance between gene lists; and projects the direction of change among specimens. Further implication of this phylogenetic approach is that it may transform microarray into diagnostic, prognostic, and predictive tool.

  1. Automatic classification of blank substrate defects

    Science.gov (United States)

    Boettiger, Tom; Buck, Peter; Paninjath, Sankaranarayanan; Pereira, Mark; Ronald, Rob; Rost, Dan; Samir, Bhamidipati

    2014-10-01

    Mask preparation stages are crucial in mask manufacturing, since this mask is to later act as a template for considerable number of dies on wafer. Defects on the initial blank substrate, and subsequent cleaned and coated substrates, can have a profound impact on the usability of the finished mask. This emphasizes the need for early and accurate identification of blank substrate defects and the risk they pose to the patterned reticle. While Automatic Defect Classification (ADC) is a well-developed technology for inspection and analysis of defects on patterned wafers and masks in the semiconductors industry, ADC for mask blanks is still in the early stages of adoption and development. Calibre ADC is a powerful analysis tool for fast, accurate, consistent and automatic classification of defects on mask blanks. Accurate, automated classification of mask blanks leads to better usability of blanks by enabling defect avoidance technologies during mask writing. Detailed information on blank defects can help to select appropriate job-decks to be written on the mask by defect avoidance tools [1][4][5]. Smart algorithms separate critical defects from the potentially large number of non-critical defects or false defects detected at various stages during mask blank preparation. Mechanisms used by Calibre ADC to identify and characterize defects include defect location and size, signal polarity (dark, bright) in both transmitted and reflected review images, distinguishing defect signals from background noise in defect images. The Calibre ADC engine then uses a decision tree to translate this information into a defect classification code. Using this automated process improves classification accuracy, repeatability and speed, while avoiding the subjectivity of human judgment compared to the alternative of manual defect classification by trained personnel [2]. This paper focuses on the results from the evaluation of Automatic Defect Classification (ADC) product at MP Mask

  2. Comparative genomic analysis and phylogenetic position of Theileria equi

    Directory of Open Access Journals (Sweden)

    Kappmeyer Lowell S

    2012-11-01

    novel model describing the role of the EMA family in persistence. T. equi has lost the putative genes for host cell transformation, or the genes were acquired by T. parva and T. annulata after divergence from T. equi. Our analysis identified 50 genes that will be useful for definitive phylogenetic classification of T. equi and closely related organisms.

  3. Undergraduate Students’ Difficulties in Reading and Constructing Phylogenetic Tree

    Science.gov (United States)

    Sa'adah, S.; Tapilouw, F. S.; Hidayat, T.

    2017-02-01

    Representation is a very important communication tool to communicate scientific concepts. Biologists produce phylogenetic representation to express their understanding of evolutionary relationships. The phylogenetic tree is visual representation depict a hypothesis about the evolutionary relationship and widely used in the biological sciences. Phylogenetic tree currently growing for many disciplines in biology. Consequently, learning about phylogenetic tree become an important part of biological education and an interesting area for biology education research. However, research showed many students often struggle with interpreting the information that phylogenetic trees depict. The purpose of this study was to investigate undergraduate students’ difficulties in reading and constructing a phylogenetic tree. The method of this study is a descriptive method. In this study, we used questionnaires, interviews, multiple choice and open-ended questions, reflective journals and observations. The findings showed students experiencing difficulties, especially in constructing a phylogenetic tree. The students’ responds indicated that main reasons for difficulties in constructing a phylogenetic tree are difficult to placing taxa in a phylogenetic tree based on the data provided so that the phylogenetic tree constructed does not describe the actual evolutionary relationship (incorrect relatedness). Students also have difficulties in determining the sister group, character synapomorphy, autapomorphy from data provided (character table) and comparing among phylogenetic tree. According to them building the phylogenetic tree is more difficult than reading the phylogenetic tree. Finding this studies provide information to undergraduate instructor and students to overcome learning difficulties of reading and constructing phylogenetic tree.

  4. Quality-Oriented Classification of Aircraft Material Based on SVM

    Directory of Open Access Journals (Sweden)

    Hongxia Cai

    2014-01-01

    Full Text Available The existing material classification is proposed to improve the inventory management. However, different materials have the different quality-related attributes, especially in the aircraft industry. In order to reduce the cost without sacrificing the quality, we propose a quality-oriented material classification system considering the material quality character, Quality cost, and Quality influence. Analytic Hierarchy Process helps to make feature selection and classification decision. We use the improved Kraljic Portfolio Matrix to establish the three-dimensional classification model. The aircraft materials can be divided into eight types, including general type, key type, risk type, and leveraged type. Aiming to improve the classification accuracy of various materials, the algorithm of Support Vector Machine is introduced. Finally, we compare the SVM and BP neural network in the application. The results prove that the SVM algorithm is more efficient and accurate and the quality-oriented material classification is valuable.

  5. Causes, consequences and solutions of phylogenetic incongruence.

    Science.gov (United States)

    Som, Anup

    2015-05-01

    Phylogenetic analysis is used to recover the evolutionary history of species, genes or proteins. Understanding phylogenetic relationships between organisms is a prerequisite of almost any evolutionary study, as contemporary species all share a common history through their ancestry. Moreover, it is important because of its wide applications that include understanding genome organization, epidemiological investigations, predicting protein functions, and deciding the genes to be analyzed in comparative studies. Despite immense progress in recent years, phylogenetic reconstruction involves many challenges that create uncertainty with respect to the true evolutionary relationships of the species or genes analyzed. One of the most notable difficulties is the widespread occurrence of incongruence among methods and also among individual genes or different genomic regions. Presence of widespread incongruence inhibits successful revealing of evolutionary relationships and applications of phylogenetic analysis. In this article, I concisely review the effect of various factors that cause incongruence in molecular phylogenies, the advances in the field that resolved some factors, and explore unresolved factors that cause incongruence along with possible ways for tackling them. © The Author 2014. Published by Oxford University Press. For Permissions, please email: journals.permissions@oup.com.

  6. Phylogenetics in plant biotechnology: principles, obstacles and ...

    African Journals Online (AJOL)

    GREGO

    2007-03-19

    Mar 19, 2007 ... Their faster evolution may lead to more infor- mative characters. However .... the coefficient of variation was determined principally by the amount of ... phylogenetic context the priority switches from more samples to more ..... phylogeny algorithms under equal and unequal evolutionary rates. Mol. Biol. Evol.

  7. Constructing Student Problems in Phylogenetic Tree Construction.

    Science.gov (United States)

    Brewer, Steven D.

    Evolution is often equated with natural selection and is taught from a primarily functional perspective while comparative and historical approaches, which are critical for developing an appreciation of the power of evolutionary theory, are often neglected. This report describes a study of expert problem-solving in phylogenetic tree construction.…

  8. The phylogenetics of succession can guide restoration

    DEFF Research Database (Denmark)

    Shooner, Stephanie; Chisholm, Chelsea Lee; Davies, T. Jonathan

    2015-01-01

    Phylogenetic tools have increasingly been used in community ecology to describe the evolutionary relationships among co-occurring species. In studies of succession, such tools may allow us to identify the evolutionary lineages most suited for particular stages of succession and habitat rehabilita...

  9. The First Darwinian Phylogenetic Tree of Plants.

    Science.gov (United States)

    Hoßfeld, Uwe; Watts, Elizabeth; Levit, Georgy S

    2017-02-01

    In 1866, the German zoologist Ernst Haeckel (1834-1919) published the first Darwinian trees of life in the history of biology in his book General Morphology of Organisms. We take a specific look at the first phylogenetic trees for the plant kingdom that Haeckel created as part of this two-volume work. Copyright © 2016 Elsevier Ltd. All rights reserved.

  10. Quantifying MCMC exploration of phylogenetic tree space.

    Science.gov (United States)

    Whidden, Chris; Matsen, Frederick A

    2015-05-01

    In order to gain an understanding of the effectiveness of phylogenetic Markov chain Monte Carlo (MCMC), it is important to understand how quickly the empirical distribution of the MCMC converges to the posterior distribution. In this article, we investigate this problem on phylogenetic tree topologies with a metric that is especially well suited to the task: the subtree prune-and-regraft (SPR) metric. This metric directly corresponds to the minimum number of MCMC rearrangements required to move between trees in common phylogenetic MCMC implementations. We develop a novel graph-based approach to analyze tree posteriors and find that the SPR metric is much more informative than simpler metrics that are unrelated to MCMC moves. In doing so, we show conclusively that topological peaks do occur in Bayesian phylogenetic posteriors from real data sets as sampled with standard MCMC approaches, investigate the efficiency of Metropolis-coupled MCMC (MCMCMC) in traversing the valleys between peaks, and show that conditional clade distribution (CCD) can have systematic problems when there are multiple peaks.

  11. Multilocus phylogenetic analysis of the genus Aeromonas.

    Science.gov (United States)

    Martinez-Murcia, Antonio J; Monera, Arturo; Saavedra, M Jose; Oncina, Remedios; Lopez-Alvarez, Monserrate; Lara, Erica; Figueras, M Jose

    2011-05-01

    A broad multilocus phylogenetic analysis (MLPA) of the representative diversity of a genus offers the opportunity to incorporate concatenated inter-species phylogenies into bacterial systematics. Recent analyses based on single housekeeping genes have provided coherent phylogenies of Aeromonas. However, to date, a multi-gene phylogenetic analysis has never been tackled. In the present study, the intra- and inter-species phylogenetic relationships of 115 strains representing all Aeromonas species described to date were investigated by MLPA. The study included the independent analysis of seven single gene fragments (gyrB, rpoD, recA, dnaJ, gyrA, dnaX, and atpD), and the tree resulting from the concatenated 4705 bp sequence. The phylogenies obtained were consistent with each other, and clustering agreed with the Aeromonas taxonomy recognized to date. The highest clustering robustness was found for the concatenated tree (i.e. all Aeromonas species split into 100% bootstrap clusters). Both possible chronometric distortions and poor resolution encountered when using single-gene analysis were buffered in the concatenated MLPA tree. However, reliable phylogenetic species delineation required an MLPA including several "bona fide" strains representing all described species.

  12. Characterization of Escherichia coli Phylogenetic Groups ...

    African Journals Online (AJOL)

    high surface hydrophobicity, toxin (hemolysin and CNF) ... Triplex polymerase chain reaction was used to classify the phylogenetic groups; hemolysin ... was detected by combination disk method; AmpC was detected by AmpC disk test, ... Quick Response Code: ... norfloxacin (10 μg), amikacin (30 μg), gentamicin (10 μg),.

  13. The complete mitochondrial genome sequence of Acentrogobius sp. (Gobiiformes: Gobiidae) and phylogenetic studies of Gobiidae.

    Science.gov (United States)

    Yang, Qiu-Hua; Lin, Qi; He, Li-Bin; Huang, Rui-Fang; Lin, Ke-Bing; Ge, Hui; Wu, Jian-Shao; Zhou, Chen

    2016-07-01

    At present, few morphological descriptions are available for Acentrogobius species and there exist some confused issues on the species classification and phylogeny. In this study, we first determined and described the complete mitochondrial genome of Acentrogobius sp. The complete mitogenome sequence is 17 083 bp in length, containing 13 protein-coding genes, two rRNA genes, 22 tRNA genes, a putative control region (CR), and a light-strand replication origin (OL). The overall base composition is 28.9% A, 26.2% T, 28.5% C, and 16.4% G, with a slight AT bias (55.1%). To furthermore validate the new determined sequences, phylogenetic trees involving all the Gobiidae species available in GenBank database were constructed. These results are expected to provide useful molecular data for species identification and further phylogenetic studies of Gobiiformes.

  14. Phylogenetics of the Phlebotomine Sand Fly Group Verrucarum (Diptera: Psychodidae: Lutzomyia)

    Science.gov (United States)

    Cohnstaedt, Lee W.; Beati, Lorenza; Caceres, Abraham G.; Ferro, Cristina; Munstermann, Leonard E.

    2011-01-01

    Within the sand fly genus Lutzomyia, the Verrucarum species group contains several of the principal vectors of American cutaneous leishmaniasis and human bartonellosis in the Andean region of South America. The group encompasses 40 species for which the taxonomic status, phylogenetic relationships, and role of each species in disease transmission remain unresolved. Mitochondrial cytochrome c oxidase I (COI) phylogenetic analysis of a 667-bp fragment supported the morphological classification of the Verrucarum group into series. Genetic sequences from seven species were grouped in well-supported monophyletic lineages. Four species, however, clustered in two paraphyletic lineages that indicate conspecificity—the Lutzomyia longiflocosa–Lutzomyia sauroida pair and the Lutzomyia quasitownsendi–Lutzomyia torvida pair. COI sequences were also evaluated as a taxonomic tool based on interspecific genetic variability within the Verrucarum group and the intraspecific variability of one of its members, Lutzomyia verrucarum, across its known distribution. PMID:21633028

  15. Phylogenetic position of Oryzolejeunea (Lejeuneaceae,Marchantiophyta): Evidence from molecular markers and morphology

    Institute of Scientific and Technical Information of China (English)

    Wen YE; Yu-Mei WEI; Alfons SCH(A)FER-VERWIMP; Rui-Liang ZHU

    2013-01-01

    The systematic position of the small neotropical genus Oryzolejeunea (three spp.) has long been controversial.Phylogenetic analyses of molecular data for the present study using DNA markers (trnL,psbA,and a nuclear ribosomal internal transcribed spacer [nrITS] region) shows that the genus is nested in Lejeunea.The results not only reveal the phylogenetic position of Oryzolejeunea for the first time,but also challenge the taxonomic value of the proximal hyaline papilla as a key feature in Lejeunea.The present study shows the urgent need for a reassessment of the perimeters of the genus Lejeunea and its infrageneric classification.Three new combinations,namely Lejeunea saccatiloba,Lejeunea grolleana,and Lejeunea venezuelana,are proposed.

  16. Identification and Classification of Rhizobia by Matrix-Assisted Laser Desorption/Ionization Time-Of-Flight Mass Spectrometry.

    Science.gov (United States)

    Jia, Rui Zong; Zhang, Rong Juan; Wei, Qing; Chen, Wen Feng; Cho, Il Kyu; Chen, Wen Xin; Li, Qing X

    Mass spectrometry (MS) has been widely used for specific, sensitive and rapid analysis of proteins and has shown a high potential for bacterial identification and characterization. Type strains of four species of rhizobia and Escherichia coli DH5α were employed as reference bacteria to optimize various parameters for identification and classification of species of rhizobia by matrix-assisted laser desorption/ionization time-of-flight MS (MALDI TOF MS). The parameters optimized included culture medium states (liquid or solid), bacterial growth phases, colony storage temperature and duration, and protein data processing to enhance the bacterial identification resolution, accuracy and reliability. The medium state had little effects on the mass spectra of protein profiles. A suitable sampling time was between the exponential phase and the stationary phase. Consistent protein mass spectral profiles were observed for E. coli colonies pre-grown for 14 days and rhizobia for 21 days at 4°C or 21°C. A dendrogram of 75 rhizobial strains of 4 genera was constructed based on MALDI TOF mass spectra and the topological patterns agreed well with those in the 16S rDNA phylogenetic tree. The potential of developing a mass spectral database for all rhizobia species was assessed with blind samples. The entire process from sample preparation to accurate identification and classification of species required approximately one hour.

  17. Security classification of information

    Energy Technology Data Exchange (ETDEWEB)

    Quist, A.S.

    1993-04-01

    This document is the second of a planned four-volume work that comprehensively discusses the security classification of information. The main focus of Volume 2 is on the principles for classification of information. Included herein are descriptions of the two major types of information that governments classify for national security reasons (subjective and objective information), guidance to use when determining whether information under consideration for classification is controlled by the government (a necessary requirement for classification to be effective), information disclosure risks and benefits (the benefits and costs of classification), standards to use when balancing information disclosure risks and benefits, guidance for assigning classification levels (Top Secret, Secret, or Confidential) to classified information, guidance for determining how long information should be classified (classification duration), classification of associations of information, classification of compilations of information, and principles for declassifying and downgrading information. Rules or principles of certain areas of our legal system (e.g., trade secret law) are sometimes mentioned to .provide added support to some of those classification principles.

  18. Security classification of information

    Energy Technology Data Exchange (ETDEWEB)

    Quist, A.S.

    1989-09-01

    Certain governmental information must be classified for national security reasons. However, the national security benefits from classifying information are usually accompanied by significant costs -- those due to a citizenry not fully informed on governmental activities, the extra costs of operating classified programs and procuring classified materials (e.g., weapons), the losses to our nation when advances made in classified programs cannot be utilized in unclassified programs. The goal of a classification system should be to clearly identify that information which must be protected for national security reasons and to ensure that information not needing such protection is not classified. This document was prepared to help attain that goal. This document is the first of a planned four-volume work that comprehensively discusses the security classification of information. Volume 1 broadly describes the need for classification, the basis for classification, and the history of classification in the United States from colonial times until World War 2. Classification of information since World War 2, under Executive Orders and the Atomic Energy Acts of 1946 and 1954, is discussed in more detail, with particular emphasis on the classification of atomic energy information. Adverse impacts of classification are also described. Subsequent volumes will discuss classification principles, classification management, and the control of certain unclassified scientific and technical information. 340 refs., 6 tabs.

  19. pplacer: linear time maximum-likelihood and Bayesian phylogenetic placement of sequences onto a fixed reference tree

    Directory of Open Access Journals (Sweden)

    Kodner Robin B

    2010-10-01

    Full Text Available Abstract Background Likelihood-based phylogenetic inference is generally considered to be the most reliable classification method for unknown sequences. However, traditional likelihood-based phylogenetic methods cannot be applied to large volumes of short reads from next-generation sequencing due to computational complexity issues and lack of phylogenetic signal. "Phylogenetic placement," where a reference tree is fixed and the unknown query sequences are placed onto the tree via a reference alignment, is a way to bring the inferential power offered by likelihood-based approaches to large data sets. Results This paper introduces pplacer, a software package for phylogenetic placement and subsequent visualization. The algorithm can place twenty thousand short reads on a reference tree of one thousand taxa per hour per processor, has essentially linear time and memory complexity in the number of reference taxa, and is easy to run in parallel. Pplacer features calculation of the posterior probability of a placement on an edge, which is a statistically rigorous way of quantifying uncertainty on an edge-by-edge basis. It also can inform the user of the positional uncertainty for query sequences by calculating expected distance between placement locations, which is crucial in the estimation of uncertainty with a well-sampled reference tree. The software provides visualizations using branch thickness and color to represent number of placements and their uncertainty. A simulation study using reads generated from 631 COG alignments shows a high level of accuracy for phylogenetic placement over a wide range of alignment diversity, and the power of edge uncertainty estimates to measure placement confidence. Conclusions Pplacer enables efficient phylogenetic placement and subsequent visualization, making likelihood-based phylogenetics methodology practical for large collections of reads; it is freely available as source code, binaries, and a web service.

  20. Phyloclimatic modeling: combining phylogenetics and bioclimatic modeling.

    Science.gov (United States)

    Yesson, C; Culham, A

    2006-10-01

    We investigate the impact of past climates on plant diversification by tracking the "footprint" of climate change on a phylogenetic tree. Diversity within the cosmopolitan carnivorous plant genus Drosera (Droseraceae) is focused within Mediterranean climate regions. We explore whether this diversity is temporally linked to Mediterranean-type climatic shifts of the mid-Miocene and whether climate preferences are conservative over phylogenetic timescales. Phyloclimatic modeling combines environmental niche (bioclimatic) modeling with phylogenetics in order to study evolutionary patterns in relation to climate change. We present the largest and most complete such example to date using Drosera. The bioclimatic models of extant species demonstrate clear phylogenetic patterns; this is particularly evident for the tuberous sundews from southwestern Australia (subgenus Ergaleium). We employ a method for establishing confidence intervals of node ages on a phylogeny using replicates from a Bayesian phylogenetic analysis. This chronogram shows that many clades, including subgenus Ergaleium and section Bryastrum, diversified during the establishment of the Mediterranean-type climate. Ancestral reconstructions of bioclimatic models demonstrate a pattern of preference for this climate type within these groups. Ancestral bioclimatic models are projected into palaeo-climate reconstructions for the time periods indicated by the chronogram. We present two such examples that each generate plausible estimates of ancestral lineage distribution, which are similar to their current distributions. This is the first study to attempt bioclimatic projections on evolutionary time scales. The sundews appear to have diversified in response to local climate development. Some groups are specialized for Mediterranean climates, others show wide-ranging generalism. This demonstrates that Phyloclimatic modeling could be repeated for other plant groups and is fundamental to the understanding of

  1. Binets: Fundamental Building Blocks for Phylogenetic Networks.

    Science.gov (United States)

    van Iersel, Leo; Moulton, Vincent; de Swart, Eveline; Wu, Taoyang

    2017-05-01

    Phylogenetic networks are a generalization of evolutionary trees that are used by biologists to represent the evolution of organisms which have undergone reticulate evolution. Essentially, a phylogenetic network is a directed acyclic graph having a unique root in which the leaves are labelled by a given set of species. Recently, some approaches have been developed to construct phylogenetic networks from collections of networks on 2- and 3-leaved networks, which are known as binets and trinets, respectively. Here we study in more depth properties of collections of binets, one of the simplest possible types of networks into which a phylogenetic network can be decomposed. More specifically, we show that if a collection of level-1 binets is compatible with some binary network, then it is also compatible with a binary level-1 network. Our proofs are based on useful structural results concerning lowest stable ancestors in networks. In addition, we show that, although the binets do not determine the topology of the network, they do determine the number of reticulations in the network, which is one of its most important parameters. We also consider algorithmic questions concerning binets. We show that deciding whether an arbitrary set of binets is compatible with some network is at least as hard as the well-known graph isomorphism problem. However, if we restrict to level-1 binets, it is possible to decide in polynomial time whether there exists a binary network that displays all the binets. We also show that to find a network that displays a maximum number of the binets is NP-hard, but that there exists a simple polynomial-time 1/3-approximation algorithm for this problem. It is hoped that these results will eventually assist in the development of new methods for constructing phylogenetic networks from collections of smaller networks.

  2. Threat diversity will erode mammalian phylogenetic diversity in the near future.

    Directory of Open Access Journals (Sweden)

    Clémentine M A Jono

    Full Text Available To reduce the accelerating rate of phylogenetic diversity loss, many studies have searched for mechanisms that could explain why certain species are at risk, whereas others are not. In particular, it has been demonstrated that species might be affected by both extrinsic threat factors as well as intrinsic biological traits that could render a species more sensitive to extinction; here, we focus on extrinsic factors. Recently, the International Union for Conservation of Nature developed a new classification of threat types, including climate change, urbanization, pollution, agriculture and aquaculture, and harvesting/hunting. We have used this new classification to analyze two main factors that could explain the expected future loss of mammalian phylogenetic diversity: 1. differences in the type of threats that affect mammals and 2. differences in the number of major threats that accumulate for a single species. Our results showed that Cetartiodactyla, Diprotodontia, Monotremata, Perissodactyla, Primates, and Proboscidea could lose a high proportion of their current phylogenetic diversity in the coming decades. In contrast, Chiroptera, Didelphimorphia, and Rodentia could lose less phylogenetic diversity than expected if extinctions were random. Some mammalian clades, including Marsupiala, Chiroptera, and a subclade of Primates, are affected by particular threat types, most likely due solely to their geographic locations and associations with particular habitats. However, regardless of the geography, habitat, and taxon considered, it is not the threat type, but the threat diversity that determines the extinction risk for species and clades. Thus, some mammals might be randomly located in areas subjected to a large diversity of threats; they might also accumulate detrimental traits that render them sensitive to different threats, which is a characteristic that could be associated with large body size. Any action reducing threat diversity is

  3. 38 CFR 4.46 - Accurate measurement.

    Science.gov (United States)

    2010-07-01

    ... 38 Pensions, Bonuses, and Veterans' Relief 1 2010-07-01 2010-07-01 false Accurate measurement. 4... RATING DISABILITIES Disability Ratings The Musculoskeletal System § 4.46 Accurate measurement. Accurate measurement of the length of stumps, excursion of joints, dimensions and location of scars with respect...

  4. Ontologies vs. Classification Systems

    DEFF Research Database (Denmark)

    Madsen, Bodil Nistrup; Erdman Thomsen, Hanne

    2009-01-01

    What is an ontology compared to a classification system? Is a taxonomy a kind of classification system or a kind of ontology? These are questions that we meet when working with people from industry and public authorities, who need methods and tools for concept clarification, for developing meta d...... classification systems and meta data taxonomies, should be based on ontologies.......What is an ontology compared to a classification system? Is a taxonomy a kind of classification system or a kind of ontology? These are questions that we meet when working with people from industry and public authorities, who need methods and tools for concept clarification, for developing meta...... data sets or for obtaining advanced search facilities. In this paper we will present an attempt at answering these questions. We will give a presentation of various types of ontologies and briefly introduce terminological ontologies. Furthermore we will argue that classification systems, e.g. product...

  5. Classification of Spreadsheet Errors

    OpenAIRE

    Rajalingham, Kamalasen; Chadwick, David R.; Knight, Brian

    2008-01-01

    This paper describes a framework for a systematic classification of spreadsheet errors. This classification or taxonomy of errors is aimed at facilitating analysis and comprehension of the different types of spreadsheet errors. The taxonomy is an outcome of an investigation of the widespread problem of spreadsheet errors and an analysis of specific types of these errors. This paper contains a description of the various elements and categories of the classification and is supported by appropri...

  6. Automatic lexical classification: bridging research and practice.

    Science.gov (United States)

    Korhonen, Anna

    2010-08-13

    Natural language processing (NLP)--the automatic analysis, understanding and generation of human language by computers--is vitally dependent on accurate knowledge about words. Because words change their behaviour between text types, domains and sub-languages, a fully accurate static lexical resource (e.g. a dictionary, word classification) is unattainable. Researchers are now developing techniques that could be used to automatically acquire or update lexical resources from textual data. If successful, the automatic approach could considerably enhance the accuracy and portability of language technologies, such as machine translation, text mining and summarization. This paper reviews the recent and on-going research in automatic lexical acquisition. Focusing on lexical classification, it discusses the many challenges that still need to be met before the approach can benefit NLP on a large scale.

  7. Study for Updated Gout Classification Criteria

    DEFF Research Database (Denmark)

    Taylor, William J; Fransen, Jaap; Jansen, Tim L

    2015-01-01

    OBJECTIVE: To determine which clinical, laboratory, and imaging features most accurately distinguished gout from non-gout. METHODS: We performed a cross-sectional study of consecutive rheumatology clinic patients with ≥1 swollen joint or subcutaneous tophus. Gout was defined by synovial fluid or ...... (MTP1) joint ever involved (multivariate OR 2.30), location of currently tender joints in other foot/ankle (multivariate OR 2.28) or MTP1 joint (multivariate OR 2.82), serum urate level >6 mg/dl (0.36 mmoles/liter; multivariate OR 3.35), ultrasound double contour sign (multivariate OR 7...... been identified for further evaluation for new gout classification criteria. Ultrasound findings and degree of uricemia add discriminating value, and will significantly contribute to more accurate classification criteria....

  8. Text Classification Using Sentential Frequent Itemsets

    Institute of Scientific and Technical Information of China (English)

    Shi-Zhu Liu; He-Ping Hu

    2007-01-01

    Text classification techniques mostly rely on single term analysis of the document data set, while more concepts,especially the specific ones, are usually conveyed by set of terms. To achieve more accurate text classifier, more informative feature including frequent co-occurring words in the same sentence and their weights are particularly important in such scenarios. In this paper, we propose a novel approach using sentential frequent itemset, a concept comes from association rule mining, for text classification, which views a sentence rather than a document as a transaction, and uses a variable precision rough set based method to evaluate each sentential frequent itemset's contribution to the classification. Experiments over the Reuters and newsgroup corpus are carried out, which validate the practicability of the proposed system.

  9. AGN Zoo and Classifications of Active Galaxies

    Science.gov (United States)

    Mickaelian, Areg M.

    2015-07-01

    We review the variety of Active Galactic Nuclei (AGN) classes (so-called "AGN zoo") and classification schemes of galaxies by activity types based on their optical emission-line spectrum, as well as other parameters and other than optical wavelength ranges. A historical overview of discoveries of various types of active galaxies is given, including Seyfert galaxies, radio galaxies, QSOs, BL Lacertae objects, Starbursts, LINERs, etc. Various kinds of AGN diagnostics are discussed. All known AGN types and subtypes are presented and described to have a homogeneous classification scheme based on the optical emission-line spectra and in many cases, also other parameters. Problems connected with accurate classifications and open questions related to AGN and their classes are discussed and summarized.

  10. A new measure to study phylogenetic relations in the brown algal order Ectocarpales: The ``codon impact parameter"

    Indian Academy of Sciences (India)

    Smarajit Das; Jayprokas Chakrabarti; Zhumur Ghosh; Satyabrata Sahoo; Bibekanand Mallick

    2005-12-01

    We analyse forty-seven chloroplast genes of the large subunit of RuBisCO, from the algal order Ectocarpales, sourced from GenBank. Codon-usage weighted by the nucleotide base-bias defines our score called the codon-impact-parameter. This score is used to obtain phylogenetic relations amongst the 47 Ectocarpales. We compare our classification with the ones done earlier.

  11. Information gathering for CLP classification

    OpenAIRE

    Ida Marcello; Felice Giordano; Francesca Marina Costamagna

    2011-01-01

    Regulation 1272/2008 includes provisions for two types of classification: harmonised classification and self-classification. The harmonised classification of substances is decided at Community level and a list of harmonised classifications is included in the Annex VI of the classification, labelling and packaging Regulation (CLP). If a chemical substance is not included in the harmonised classification list it must be self-classified, based on available information, according to the requireme...

  12. BEASTling: A software tool for linguistic phylogenetics using BEAST 2

    Science.gov (United States)

    Forkel, Robert; Kaiping, Gereon A.; Atkinson, Quentin D.

    2017-01-01

    We present a new open source software tool called BEASTling, designed to simplify the preparation of Bayesian phylogenetic analyses of linguistic data using the BEAST 2 platform. BEASTling transforms comparatively short and human-readable configuration files into the XML files used by BEAST to specify analyses. By taking advantage of Creative Commons-licensed data from the Glottolog language catalog, BEASTling allows the user to conveniently filter datasets using names for recognised language families, to impose monophyly constraints so that inferred language trees are backward compatible with Glottolog classifications, or to assign geographic location data to languages for phylogeographic analyses. Support for the emerging cross-linguistic linked data format (CLDF) permits easy incorporation of data published in cross-linguistic linked databases into analyses. BEASTling is intended to make the power of Bayesian analysis more accessible to historical linguists without strong programming backgrounds, in the hopes of encouraging communication and collaboration between those developing computational models of language evolution (who are typically not linguists) and relevant domain experts. PMID:28796784

  13. Metrics for phylogenetic networks II: nodal and triplets metrics.

    Science.gov (United States)

    Cardona, Gabriel; Llabrés, Mercè; Rosselló, Francesc; Valiente, Gabriel

    2009-01-01

    The assessment of phylogenetic network reconstruction methods requires the ability to compare phylogenetic networks. This is the second in a series of papers devoted to the analysis and comparison of metrics for tree-child time consistent phylogenetic networks on the same set of taxa. In this paper, we generalize to phylogenetic networks two metrics that have already been introduced in the literature for phylogenetic trees: the nodal distance and the triplets distance. We prove that they are metrics on any class of tree-child time consistent phylogenetic networks on the same set of taxa, as well as some basic properties for them. To prove these results, we introduce a reduction/expansion procedure that can be used not only to establish properties of tree-child time consistent phylogenetic networks by induction, but also to generate all tree-child time consistent phylogenetic networks with a given number of leaves.

  14. Phylogenetic paleobiogeography of Late Ordovician Laurentian brachiopods

    Directory of Open Access Journals (Sweden)

    Jennifer E. Bauer

    2014-12-01

    Full Text Available Phylogenetic biogeographic analysis of four brachiopod genera was used to uncover large-scale geologic drivers of Late Ordovician biogeographic differentiation in Laurentia. Previously generated phylogenetic hypotheses were converted into area cladograms, ancestral geographic ranges were optimized and speciation events characterized as via dispersal or vicariance, when possible. Area relationships were reconstructed using Lieberman-modified Brooks Parsimony Analysis. The resulting area cladograms indicate tectonic and oceanographic changes were the primary geologic drivers of biogeographic patterns within the focal taxa. The Taconic tectophase contributed to the separation of the Appalachian and Central basins as well as the two midcontinent basins, whereas sea level rise following the Boda Event promoted interbasinal dispersal. Three migration pathways into the Cincinnati Basin were recognized, which supports the multiple pathway hypothesis for the Richmondian Invasion.

  15. Morphological Phylogenetics in the Genomic Age.

    Science.gov (United States)

    Lee, Michael S Y; Palci, Alessandro

    2015-10-05

    Evolutionary trees underpin virtually all of biology, and the wealth of new genomic data has enabled us to reconstruct them with increasing detail and confidence. While phenotypic (typically morphological) traits are becoming less important in reconstructing evolutionary trees, they still serve vital and unique roles in phylogenetics, even for living taxa for which vast amounts of genetic information are available. Morphology remains a powerful independent source of evidence for testing molecular clades, and - through fossil phenotypes - the primary means for time-scaling phylogenies. Morphological phylogenetics is therefore vital for transforming undated molecular topologies into dated evolutionary trees. However, if morphology is to be employed to its full potential, biologists need to start scrutinising phenotypes in a more objective fashion, models of phenotypic evolution need to be improved, and approaches for analysing phenotypic traits and fossils together with genomic data need to be refined.

  16. Alignment-free phylogenetics and population genetics.

    Science.gov (United States)

    Haubold, Bernhard

    2014-05-01

    Phylogenetics and population genetics are central disciplines in evolutionary biology. Both are based on comparative data, today usually DNA sequences. These have become so plentiful that alignment-free sequence comparison is of growing importance in the race between scientists and sequencing machines. In phylogenetics, efficient distance computation is the major contribution of alignment-free methods. A distance measure should reflect the number of substitutions per site, which underlies classical alignment-based phylogeny reconstruction. Alignment-free distance measures are either based on word counts or on match lengths, and I apply examples of both approaches to simulated and real data to assess their accuracy and efficiency. While phylogeny reconstruction is based on the number of substitutions, in population genetics, the distribution of mutations along a sequence is also considered. This distribution can be explored by match lengths, thus opening the prospect of alignment-free population genomics.

  17. Molecular phylogenetics of mastodon and Tyrannosaurus rex.

    Science.gov (United States)

    Organ, Chris L; Schweitzer, Mary H; Zheng, Wenxia; Freimark, Lisa M; Cantley, Lewis C; Asara, John M

    2008-04-25

    We report a molecular phylogeny for a nonavian dinosaur, extending our knowledge of trait evolution within nonavian dinosaurs into the macromolecular level of biological organization. Fragments of collagen alpha1(I) and alpha2(I) proteins extracted from fossil bones of Tyrannosaurus rex and Mammut americanum (mastodon) were analyzed with a variety of phylogenetic methods. Despite missing sequence data, the mastodon groups with elephant and the T. rex groups with birds, consistent with predictions based on genetic and morphological data for mastodon and on morphological data for T. rex. Our findings suggest that molecular data from long-extinct organisms may have the potential for resolving relationships at critical areas in the vertebrate evolutionary tree that have, so far, been phylogenetically intractable.

  18. A Phylogenetic Index for Cichlid Microsatellite Primers

    Directory of Open Access Journals (Sweden)

    Robert D. Kunkle

    2010-01-01

    Full Text Available Microsatellites abound in most organisms and have proven useful for a range of genetic and genomic studies. Once primers have been created, they can be applied to populations or taxa that have diverged from the source taxon. We use PCR amplification, in a 96-well format, to determine the presence and absence of 46 microsatellite loci in 13 cichlid species. At least one primer set amplified a product in each species tested, and some products were present in nearly all species. These results are compared to the known phylogenetic relationships among cichlids. While we do not address intraspecies variation, our results present a phylogenetic index for the success of microsatellite PCR primer product amplification, thus providing information regarding a collection of primers that are applicable to wide range of species. Through the use of such a uniform primer panel, the potential impact for cross species would be increased.

  19. Independent Comparison of Popular DPI Tools for Traffic Classification

    DEFF Research Database (Denmark)

    Bujlow, Tomasz; Carela-Español, Valentín; Barlet-Ros, Pere

    2015-01-01

    Deep Packet Inspection (DPI) is the state-of-the-art technology for traffic classification. According to the conventional wisdom, DPI is the most accurate classification technique. Consequently, most popular products, either commercial or open-source, rely on some sort of DPI for traffic classifi......Deep Packet Inspection (DPI) is the state-of-the-art technology for traffic classification. According to the conventional wisdom, DPI is the most accurate classification technique. Consequently, most popular products, either commercial or open-source, rely on some sort of DPI for traffic......, application and web service). We carefully built a labeled dataset with more than 750K flows, which contains traffic from popular applications. We used the Volunteer-Based System (VBS), developed at Aalborg University, to guarantee the correct labeling of the dataset. We released this dataset, including full...

  20. Phylogenetics and Computational Biology of Multigene Families

    Science.gov (United States)

    Liò, Pietro; Brilli, Matteo; Fani, Renato

    This chapter introduces the study of the major evolutionary forces operating in large gene families. The reconstruction of duplication history and phylogenetic analysis provide an interpretative framework of the evolution of multigene families. We present here two case studies, the first coming from Eukaryotes (chemokine receptors) and the second from Prokaryotes (TIM barrel proteins), showing how functional and structural constraints have shaped gene duplication events.

  1. Phylogenetic estimation with partial likelihood tensors

    CERN Document Server

    Sumner, J G

    2008-01-01

    We present an alternative method for calculating likelihoods in molecular phylogenetics. Our method is based on partial likelihood tensors, which are generalizations of partial likelihood vectors, as used in Felsenstein's approach. Exploiting a lexicographic sorting and partial likelihood tensors, it is possible to obtain significant computational savings. We show this on a range of simulated data by enumerating all numerical calculations that are required by our method and the standard approach.

  2. A phylogenetic analysis of Aquifex pyrophilus

    Science.gov (United States)

    Burggraf, S.; Olsen, G. J.; Stetter, K. O.; Woese, C. R.

    1992-01-01

    The 16S rRNA of the bacterion Aquifex pyrophilus, a microaerophilic, oxygen-reducing hyperthermophile, has been sequenced directly from the the PCR amplified gene. Phylogenetic analyses show the Aq. pyrophilus lineage to be probably the deepest (earliest) in the (eu)bacterial tree. The addition of this deep branching to the bacterial tree further supports the argument that the Bacteria are of thermophilic ancestry.

  3. MINER: software for phylogenetic motif identification

    OpenAIRE

    La, David; Livesay, Dennis R.

    2005-01-01

    MINER is web-based software for phylogenetic motif (PM) identification. PMs are sequence regions (fragments) that conserve the overall familial phylogeny. PMs have been shown to correspond to a wide variety of catalytic regions, substrate-binding sites and protein interfaces, making them ideal functional site predictions. The MINER output provides an intuitive interface for interactive PM sequence analysis and structural visualization. The web implementation of MINER is freely available at . ...

  4. Phylogenetic conservatism of environmental niches in mammals.

    Science.gov (United States)

    Cooper, Natalie; Freckleton, Rob P; Jetz, Walter

    2011-08-01

    Phylogenetic niche conservatism is the pattern where close relatives occupy similar niches, whereas distant relatives are more dissimilar. We suggest that niche conservatism will vary across clades in relation to their characteristics. Specifically, we investigate how conservatism of environmental niches varies among mammals according to their latitude, range size, body size and specialization. We use the Brownian rate parameter, σ(2), to measure the rate of evolution in key variables related to the ecological niche and define the more conserved group as the one with the slower rate of evolution. We find that tropical, small-ranged and specialized mammals have more conserved thermal niches than temperate, large-ranged or generalized mammals. Partitioning niche conservatism into its spatial and phylogenetic components, we find that spatial effects on niche variables are generally greater than phylogenetic effects. This suggests that recent evolution and dispersal have more influence on species' niches than more distant evolutionary events. These results have implications for our understanding of the role of niche conservatism in species richness patterns and for gauging the potential for species to adapt to global change.

  5. Incongruencies in Vaccinia Virus Phylogenetic Trees

    Directory of Open Access Journals (Sweden)

    Chad Smithson

    2014-10-01

    Full Text Available Over the years, as more complete poxvirus genomes have been sequenced, phylogenetic studies of these viruses have become more prevalent. In general, the results show similar relationships between the poxvirus species; however, some inconsistencies are notable. Previous analyses of the viral genomes contained within the vaccinia virus (VACV-Dryvax vaccine revealed that their phylogenetic relationships were sometimes clouded by low bootstrapping confidence. To analyze the VACV-Dryvax genomes in detail, a new tool-set was developed and integrated into the Base-By-Base bioinformatics software package. Analyses showed that fewer unique positions were present in each VACV-Dryvax genome than expected. A series of patterns, each containing several single nucleotide polymorphisms (SNPs were identified that were counter to the results of the phylogenetic analysis. The VACV genomes were found to contain short DNA sequence blocks that matched more distantly related clades. Additionally, similar non-conforming SNP patterns were observed in (1 the variola virus clade; (2 some cowpox clades; and (3 VACV-CVA, the direct ancestor of VACV-MVA. Thus, traces of past recombination events are common in the various orthopoxvirus clades, including those associated with smallpox and cowpox viruses.

  6. A Consistent Phylogenetic Backbone for the Fungi

    Science.gov (United States)

    Ebersberger, Ingo; de Matos Simoes, Ricardo; Kupczok, Anne; Gube, Matthias; Kothe, Erika; Voigt, Kerstin; von Haeseler, Arndt

    2012-01-01

    The kingdom of fungi provides model organisms for biotechnology, cell biology, genetics, and life sciences in general. Only when their phylogenetic relationships are stably resolved, can individual results from fungal research be integrated into a holistic picture of biology. However, and despite recent progress, many deep relationships within the fungi remain unclear. Here, we present the first phylogenomic study of an entire eukaryotic kingdom that uses a consistency criterion to strengthen phylogenetic conclusions. We reason that branches (splits) recovered with independent data and different tree reconstruction methods are likely to reflect true evolutionary relationships. Two complementary phylogenomic data sets based on 99 fungal genomes and 109 fungal expressed sequence tag (EST) sets analyzed with four different tree reconstruction methods shed light from different angles on the fungal tree of life. Eleven additional data sets address specifically the phylogenetic position of Blastocladiomycota, Ustilaginomycotina, and Dothideomycetes, respectively. The combined evidence from the resulting trees supports the deep-level stability of the fungal groups toward a comprehensive natural system of the fungi. In addition, our analysis reveals methodologically interesting aspects. Enrichment for EST encoded data—a common practice in phylogenomic analyses—introduces a strong bias toward slowly evolving and functionally correlated genes. Consequently, the generalization of phylogenomic data sets as collections of randomly selected genes cannot be taken for granted. A thorough characterization of the data to assess possible influences on the tree reconstruction should therefore become a standard in phylogenomic analyses. PMID:22114356

  7. Phylogenetic analysis of cubilin (CUBN) gene.

    Science.gov (United States)

    Shaik, Abjal Pasha; Alsaeed, Abbas H; Kiranmayee, S; Bammidi, Vk; Sultana, Asma

    2013-01-01

    Cubilin, (CUBN; also known as intrinsic factor-cobalamin receptor [Homo sapiens Entrez Pubmed ref NM_001081.3; NG_008967.1; GI: 119606627]), located in the epithelium of intestine and kidney acts as a receptor for intrinsic factor - vitamin B12 complexes. Mutations in CUBN may play a role in autosomal recessive megaloblastic anemia. The current study investigated the possible role of CUBN in evolution using phylogenetic testing. A total of 588 BLAST hits were found for the cubilin query sequence and these hits showed putative conserved domain, CUB superfamily (as on 27(th) Nov 2012). A first-pass phylogenetic tree was constructed to identify the taxa which most often contained the CUBN sequences. Following this, we narrowed down the search by manually deleting sequences which were not CUBN. A repeat phylogenetic analysis of 25 taxa was performed using PhyML, RAxML and TreeDyn softwares to confirm that CUBN is a conserved protein emphasizing its importance as an extracellular domain and being present in proteins mostly known to be involved in development in many chordate taxa but not found in prokaryotes, plants and yeast.. No horizontal gene transfers have been found between different taxa.

  8. Uncertain-tree: discriminating among competing approaches to the phylogenetic analysis of phenotype data

    Science.gov (United States)

    Tanner, Alastair R.; Fleming, James F.; Tarver, James E.; Pisani, Davide

    2017-01-01

    Morphological data provide the only means of classifying the majority of life's history, but the choice between competing phylogenetic methods for the analysis of morphology is unclear. Traditionally, parsimony methods have been favoured but recent studies have shown that these approaches are less accurate than the Bayesian implementation of the Mk model. Here we expand on these findings in several ways: we assess the impact of tree shape and maximum-likelihood estimation using the Mk model, as well as analysing data composed of both binary and multistate characters. We find that all methods struggle to correctly resolve deep clades within asymmetric trees, and when analysing small character matrices. The Bayesian Mk model is the most accurate method for estimating topology, but with lower resolution than other methods. Equal weights parsimony is more accurate than implied weights parsimony, and maximum-likelihood estimation using the Mk model is the least accurate method. We conclude that the Bayesian implementation of the Mk model should be the default method for phylogenetic estimation from phenotype datasets, and we explore the implications of our simulations in reanalysing several empirical morphological character matrices. A consequence of our finding is that high levels of resolution or the ability to classify species or groups with much confidence should not be expected when using small datasets. It is now necessary to depart from the traditional parsimony paradigms of constructing character matrices, towards datasets constructed explicitly for Bayesian methods. PMID:28077778

  9. Library Classification 2020

    Science.gov (United States)

    Harris, Christopher

    2013-01-01

    In this article the author explores how a new library classification system might be designed using some aspects of the Dewey Decimal Classification (DDC) and ideas from other systems to create something that works for school libraries in the year 2020. By examining what works well with the Dewey Decimal System, what features should be carried…

  10. Multiple sparse representations classification

    NARCIS (Netherlands)

    E. Plenge (Esben); S.K. Klein (Stefan); W.J. Niessen (Wiro); E. Meijering (Erik)

    2015-01-01

    textabstractSparse representations classification (SRC) is a powerful technique for pixelwise classification of images and it is increasingly being used for a wide variety of image analysis tasks. The method uses sparse representation and learned redundant dictionaries to classify image pixels. In t

  11. Library Classification 2020

    Science.gov (United States)

    Harris, Christopher

    2013-01-01

    In this article the author explores how a new library classification system might be designed using some aspects of the Dewey Decimal Classification (DDC) and ideas from other systems to create something that works for school libraries in the year 2020. By examining what works well with the Dewey Decimal System, what features should be carried…

  12. Classifier in Age classification

    Directory of Open Access Journals (Sweden)

    B. Santhi

    2012-12-01

    Full Text Available Face is the important feature of the human beings. We can derive various properties of a human by analyzing the face. The objective of the study is to design a classifier for age using facial images. Age classification is essential in many applications like crime detection, employment and face detection. The proposed algorithm contains four phases: preprocessing, feature extraction, feature selection and classification. The classification employs two class labels namely child and Old. This study addresses the limitations in the existing classifiers, as it uses the Grey Level Co-occurrence Matrix (GLCM for feature extraction and Support Vector Machine (SVM for classification. This improves the accuracy of the classification as it outperforms the existing methods.

  13. Enhancing Accuracy of Plant Leaf Classification Techniques

    Directory of Open Access Journals (Sweden)

    C. S. Sumathi

    2014-03-01

    Full Text Available Plants have become an important source of energy, and are a fundamental piece in the puzzle to solve the problem of global warming. Living beings also depend on plants for their food, hence it is of great importance to know about the plants growing around us and to preserve them. Automatic plant leaf classification is widely researched. This paper investigates the efficiency of learning algorithms of MLP for plant leaf classification. Incremental back propagation, Levenberg–Marquardt and batch propagation learning algorithms are investigated. Plant leaf images are examined using three different Multi-Layer Perceptron (MLP modelling techniques. Back propagation done in batch manner increases the accuracy of plant leaf classification. Results reveal that batch training is faster and more accurate than MLP with incremental training and Levenberg– Marquardt based learning for plant leaf classification. Various levels of semi-batch training used on 9 species of 15 sample each, a total of 135 instances show a roughly linear increase in classification accuracy.

  14. Comparative evolutionary diversity and phylogenetic structure across multiple forest dynamics plots: a mega-phylogeny approach

    Science.gov (United States)

    Erickson, David L.; Jones, Frank A.; Swenson, Nathan G.; Pei, Nancai; Bourg, Norman A.; Chen, Wenna; Davies, Stuart J.; Ge, Xue-jun; Hao, Zhanqing; Howe, Robert W.; Huang, Chun-Lin; Larson, Andrew J.; Lum, Shawn K. Y.; Lutz, James A.; Ma, Keping; Meegaskumbura, Madhava; Mi, Xiangcheng; Parker, John D.; Fang-Sun, I.; Wright, S. Joseph; Wolf, Amy T.; Ye, W.; Xing, Dingliang; Zimmerman, Jess K.; Kress, W. John

    2014-01-01

    Forest dynamics plots, which now span longitudes, latitudes, and habitat types across the globe, offer unparalleled insights into the ecological and evolutionary processes that determine how species are assembled into communities. Understanding phylogenetic relationships among species in a community has become an important component of assessing assembly processes. However, the application of evolutionary information to questions in community ecology has been limited in large part by the lack of accurate estimates of phylogenetic relationships among individual species found within communities, and is particularly limiting in comparisons between communities. Therefore, streamlining and maximizing the information content of these community phylogenies is a priority. To test the viability and advantage of a multi-community phylogeny, we constructed a multi-plot mega-phylogeny of 1347 species of trees across 15 forest dynamics plots in the ForestGEO network using DNA barcode sequence data (rbcL, matK, and psbA-trnH) and compared community phylogenies for each individual plot with respect to support for topology and branch lengths, which affect evolutionary inference of community processes. The levels of taxonomic differentiation across the phylogeny were examined by quantifying the frequency of resolved nodes throughout. In addition, three phylogenetic distance (PD) metrics that are commonly used to infer assembly processes were estimated for each plot [PD, Mean Phylogenetic Distance (MPD), and Mean Nearest Taxon Distance (MNTD)]. Lastly, we examine the partitioning of phylogenetic diversity among community plots through quantification of inter-community MPD and MNTD. Overall, evolutionary relationships were highly resolved across the DNA barcode-based mega-phylogeny, and phylogenetic resolution for each community plot was improved when estimated within the context of the mega-phylogeny. Likewise, when compared with phylogenies for individual plots, estimates of

  15. Multisensor Target Detection And Classification

    Science.gov (United States)

    Ruck, Dennis W.; Rogers, Steven K.; Mills, James P.; Kabrisky, Matthew

    1988-08-01

    In this paper a new approach to the detection and classification of tactical targets using a multifunction laser radar sensor is developed. Targets of interest are tanks, jeeps, trucks, and other vehicles. Doppler images are segmented by developing a new technique which compensates for spurious doppler returns. Relative range images are segmented using an approach based on range gradients. The resultant shapes in the segmented images are then classified using Zernike moment invariants as shape descriptors. Two classification decision rules are implemented: a classical statistical nearest-neighbor approach and a multilayer perceptron architecture. The doppler segmentation algorithm was applied to a set of 180 real sensor images. An accurate segmentation was obtained for 89 percent of the images. The new doppler segmentation proved to be a robust method, and the moment invariants were effective in discriminating the tactical targets. Tanks were classified correctly 86 percent of the time. The most important result of this research is the demonstration of the use of a new information processing architecture for image processing applications.

  16. COII ”long fragment” reliability in characterisation and classification of forensically important flies

    Directory of Open Access Journals (Sweden)

    Sanaa M. Aly

    2017-01-01

    Full Text Available Introduction : Molecular identification of collected flies is important in forensic entomological analysis guided with accurate evaluation of the chosen genetic marker. The selected mitochondrial DNA segments can be used to properly identify species. The aim of the present study was to determine the reliability of the 635-bp-long cytochrome oxidase II gene (COII in identification of forensically important flies. Material and methods: Forty-two specimens belonging to 11 species ( Calliphoridae: Chrysomya albiceps , C. rufifacies , C. megacephala , Lucilia sericata , L. cuprina ; Sarcophagidae: Sarcophaga carnaria , S. dux , S. albiceps , Wohlfahrtia nuba ; Muscidae: Musca domestica , M. autumnalis were analysed. The selected marker was amplified using PCR followed by sequencing. Nucleotide sequence divergences were calculated using the K2P (Kimura two-parameter distance model, and a NJ (neighbour-joining phylogenetic tree was constructed. Results : All examined specimens were assigned to the correct species, formed distinct monophyletic clades and ordered in accordance with their taxonomic classification. Intraspecific variation ranged from 0 to 1% and interspecific variation occurred between 2 and 20%. Conclusions : The 635-bp-long COII marker is suitable for clear differentiation and identification of forensically relevant flies.

  17. Photometric Supernova Classification with Machine Learning

    Science.gov (United States)

    Lochner, Michelle; McEwen, Jason D.; Peiris, Hiranya V.; Lahav, Ofer; Winter, Max K.

    2016-08-01

    Automated photometric supernova classification has become an active area of research in recent years in light of current and upcoming imaging surveys such as the Dark Energy Survey (DES) and the Large Synoptic Survey Telescope, given that spectroscopic confirmation of type for all supernovae discovered will be impossible. Here, we develop a multi-faceted classification pipeline, combining existing and new approaches. Our pipeline consists of two stages: extracting descriptive features from the light curves and classification using a machine learning algorithm. Our feature extraction methods vary from model-dependent techniques, namely SALT2 fits, to more independent techniques that fit parametric models to curves, to a completely model-independent wavelet approach. We cover a range of representative machine learning algorithms, including naive Bayes, k-nearest neighbors, support vector machines, artificial neural networks, and boosted decision trees (BDTs). We test the pipeline on simulated multi-band DES light curves from the Supernova Photometric Classification Challenge. Using the commonly used area under the curve (AUC) of the Receiver Operating Characteristic as a metric, we find that the SALT2 fits and the wavelet approach, with the BDTs algorithm, each achieve an AUC of 0.98, where 1 represents perfect classification. We find that a representative training set is essential for good classification, whatever the feature set or algorithm, with implications for spectroscopic follow-up. Importantly, we find that by using either the SALT2 or the wavelet feature sets with a BDT algorithm, accurate classification is possible purely from light curve data, without the need for any redshift information.

  18. Towards a new classification system for legumes: Progress report from the 6th International Legume Conference

    NARCIS (Netherlands)

    Pontes Coelho Borges, L.M.; Bruneau, A.; Cardoso, D.; Crisp, M.; Delgado-Salinas, A.; Doyle, J.J.; Egan, A.; Herendeen, P.S.; Hughes, C.; Kenicer, G.; Klitgaard, B.; Koenen, E.; Lavin, M.; Lewis, G.; Luckow, M.; Mackinder, B.; Malecot, V.; Miller, J.T.; Pennington, R.T.; Queiroz, de L.P.; Schrire, B.; Simon, M.F.; Steele, K.; Torke, B.; Wieringa, J.J.; Wojciechowski, M.F.; Boatwright, S.; Estrella, de la M.; Mansano, V.D.; Prado, D.E.; Stirton, C.; Wink, M.

    2013-01-01

    Legume systematists have been making great progress in understanding evolutionary relationships within the Leguminosae (Fabaceae), the third largest family of flowering plants. As the phylogenetic picture has become clearer, so too has the need for a revised classification of the family. The

  19. The New Higher Level Classification of Eukaryotes with Emphasis on the Taxonomy of Protists

    Science.gov (United States)

    SINA M. ADL; ALASTAIR G. B. SIMPSON; MARK A. FARMER; ROBERT A. ANDERSEN; O. ROGER ANDERSON; JOHN R. BARTA; SAMUEL S. BOWSER; GUY BRUGEROLLE; ROBERT A. FENSOME; SUZANNE FREDERICQ; TIMOTHY Y. JAMES; SERGEI KARPOV; PAUL KUGRENS; JOHN KRUG; CHRISTOPHER E. LANE; LOUISE A. LEWIS; JEAN LODGE; DENIS H. LYNN; DAVID G. MANN; RICHARD M. MCCOURT; LEONEL MENDOZA; ØJVIND MOESTRUP; SHARON E. MOZLEY-STANDRIDGE; THOMAS A. NERAD; CAROL A. SHEARER; ALEXEY V. SMIRNOV; FREDERICK W. SPIEGEL; MAX F.J.R. TAYLOR

    2005-01-01

    This revision of the classification of unicellular eukaryotes updates that of Levine et al. (1980) for the protozoa and expands it to include other protists. Whereas the previous revision was primarily to incorporate the results of ultrastructural studies, this revision incorporates results from both ultrastructural research since 1980 and molecular phylogenetic...

  20. Improved Surgical Site Infection (SSI) rate through accurately assessed surgical wounds.

    Science.gov (United States)

    John, Honeymol; Nimeri, Abdelrahman; Ellahham, Samer

    2015-01-01

    Sheikh Khalifa Medical City's (SKMC) Surgery Institute was identified as a high outlier in Surgical Site Infections (SSI) based on the American College of Surgeons National Surgical Quality Improvement Program (ACS NSQIP) - Semi-Annual Report (SAR) in January 2012. The aim of this project was to improve SSI rates through accurate wound classification. We identified SSI rate reduction as a performance improvement and safety priority at SKMC, a tertiary referral center. We used the American College of Surgeons National Surgical Quality Improvement Program (ACS NSQIP) best practice guidelines as a guide. ACS NSQIP is a clinical registry that provides risk-adjusted clinical outcome reports every six months. The rates of SSI are reported in an observed/expected ratio. The expected ratio is calculated based on the risk factors of the patients which include wound classification. We established a multidisciplinary SSI taskforce. The members of the SSI taskforce included the ACS NSQIP team members, quality, surgeons, nurses, infection control, IT, pharmacy, microbiology, and it was chaired by a colorectal surgeon. The taskforce focused on five areas: pre-op showering and hair removal, skin antisepsis, prophylactic antibiotics, peri-operative maintenance of glycaemia, and normothermia. We planned audits to evaluate our wound classification and our SSI rates based on the SAR. Our expected SSI rates in general surgery and the whole department were 2.52% and 1.70% respectively, while our observed SSI rates were 4.68% and 3.57% respectively, giving us a high outlier status with an odd's ratio of 1.72 and 2.03. Wound classifications were identified as an area of concern. For example, wound classifications were preoperatively selected based on the default wound classification of the booked procedure in the Electronic Medical Record (EMR) which led to under classifying wounds in many occasions. A total of 998 cases were reviewed, our rate of incorrect wound classification

  1. A multigene phylogenetic synthesis for the class Lecanoromycetes (Ascomycota): 1307 fungi representing 1139 infrageneric taxa, 317 genera and 66 families.

    Science.gov (United States)

    Miadlikowska, Jolanta; Kauff, Frank; Högnabba, Filip; Oliver, Jeffrey C; Molnár, Katalin; Fraker, Emily; Gaya, Ester; Hafellner, Josef; Hofstetter, Valérie; Gueidan, Cécile; Otálora, Mónica A G; Hodkinson, Brendan; Kukwa, Martin; Lücking, Robert; Björk, Curtis; Sipman, Harrie J M; Burgaz, Ana Rosa; Thell, Arne; Passo, Alfredo; Myllys, Leena; Goward, Trevor; Fernández-Brime, Samantha; Hestmark, Geir; Lendemer, James; Lumbsch, H Thorsten; Schmull, Michaela; Schoch, Conrad L; Sérusiaux, Emmanuël; Maddison, David R; Arnold, A Elizabeth; Lutzoni, François; Stenroos, Soili

    2014-10-01

    The Lecanoromycetes is the largest class of lichenized Fungi, and one of the most species-rich classes in the kingdom. Here we provide a multigene phylogenetic synthesis (using three ribosomal RNA-coding and two protein-coding genes) of the Lecanoromycetes based on 642 newly generated and 3329 publicly available sequences representing 1139 taxa, 317 genera, 66 families, 17 orders and five subclasses (four currently recognized: Acarosporomycetidae, Lecanoromycetidae, Ostropomycetidae, Umbilicariomycetidae; and one provisionarily recognized, 'Candelariomycetidae'). Maximum likelihood phylogenetic analyses on four multigene datasets assembled using a cumulative supermatrix approach with a progressively higher number of species and missing data (5-gene, 5+4-gene, 5+4+3-gene and 5+4+3+2-gene datasets) show that the current classification includes non-monophyletic taxa at various ranks, which need to be recircumscribed and require revisionary treatments based on denser taxon sampling and more loci. Two newly circumscribed orders (Arctomiales and Hymeneliales in the Ostropomycetidae) and three families (Ramboldiaceae and Psilolechiaceae in the Lecanorales, and Strangosporaceae in the Lecanoromycetes inc. sed.) are introduced. The potential resurrection of the families Eigleraceae and Lopadiaceae is considered here to alleviate phylogenetic and classification disparities. An overview of the photobionts associated with the main fungal lineages in the Lecanoromycetes based on available published records is provided. A revised schematic classification at the family level in the phylogenetic context of widely accepted and newly revealed relationships across Lecanoromycetes is included. The cumulative addition of taxa with an increasing amount of missing data (i.e., a cumulative supermatrix approach, starting with taxa for which sequences were available for all five targeted genes and ending with the addition of taxa for which only two genes have been sequenced) revealed

  2. Kappa Coefficients for Circular Classifications

    NARCIS (Netherlands)

    Warrens, Matthijs J.; Pratiwi, Bunga C.

    2016-01-01

    Circular classifications are classification scales with categories that exhibit a certain periodicity. Since linear scales have endpoints, the standard weighted kappas used for linear scales are not appropriate for analyzing agreement between two circular classifications. A family of kappa coefficie

  3. Classification of pmoA amplicon pyrosequences using BLAST and the lowest common ancestor method in MEGAN

    Directory of Open Access Journals (Sweden)

    Marc Gregory Dumont

    2014-02-01

    Full Text Available The classification of high-throughput sequencing data of protein-encoding genes is not as well established as for 16S rRNA. The objective of this work was to develop a simple and accurate method of classifying large datasets of pmoA sequences, a common marker for methanotrophic bacteria. A taxonomic system for pmoA was developed based on a phylogenetic analysis of available sequences. The taxonomy incorporates the known diversity of pmoA present in public databases, including both sequences from cultivated and uncultivated organisms. Representative sequences from closely related genes, such as those encoding the bacterial ammonia monooxygenase, were also included in the pmoA taxonomy. In total, 53 low-level taxa (genus-level are included. Using previously published datasets of high-throughput pmoA amplicon sequence data, we tested two approaches for classifying pmoA: a naïve Bayesian classifier and BLAST. Classification of pmoA sequences based on BLAST analyses was performed using the lowest common ancestor (LCA algorithm in MEGAN, a software program commonly used for the analysis of metagenomic data. Both the naïve Bayesian and BLAST methods were able to classify pmoA sequences and provided similar classifications; however, the naïve Bayesian classifier was prone to misclassifying contaminant sequences present in the datasets. Another advantage of the BLAST/LCA method was that it provided a user-interpretable output and enabled novelty detection at various levels, from highly divergent pmoA sequences to genus-level novelty.  

  4. Phylogenetic relationships of Proboscoida Broch, 1910 (Cnidaria, Hydrozoa): Are traditional morphological diagnostic characters relevant for the delimitation of lineages at the species, genus, and family levels?

    Science.gov (United States)

    Cunha, Amanda F; Collins, Allen G; Marques, Antonio C

    2017-01-01

    Overlapping variation of morphological characters can lead to misinterpretation in taxonomic diagnoses and the delimitation of different lineages. This is the case for hydrozoans that have traditionally been united in the family Campanulariidae, a group known for its wide morphological variation and complicated taxonomic history. In a recently proposed phylogenetic classification of leptothecate hydrozoans, this family was restricted to a more narrow sense while a larger clade containing most species traditionally classified in Campanulariidae, along with members of Bonneviellidae, was established as the suborder Proboscoida. We used molecular data to infer the phylogenetic relationships among campanulariids and assess the traditional classification of the family, as well as the new classification scheme for the group. The congruity and relevance of diagnostic characters were also evaluated. While mostly consistent with the new phylogenetic classification of Proboscoida, our increased taxon sampling resulted in some conflicts at the family level, specially regarding the monophyly of Clytiidae and Obeliidae. Considering the traditional classification, only Obeliidae is close to its original scope (as subfamily Obeliinae). At the genus level, Campanularia and Clytia are not monophyletic. Species with Obelia-like medusae do not form a monophyletic group, nor do species with fixed gonophores, indicating that these characters do not readily diagnose different genera. Finally, the species Orthopyxis integra, Clytia gracilis, and Obelia dichotoma are not monophyletic, suggesting that most of their current diagnostic characters are not informative for their delimitation. Several diagnostic characters in this group need to be reassessed, with emphasis on their variation, in order to have a consistent taxonomic and phylogenetic framework for the classification of campanulariid hydrozoans. Copyright © 2016 Elsevier Inc. All rights reserved.

  5. New classification criteria for gout: a framework for progress

    Science.gov (United States)

    Dalbeth, Nicola; Fransen, Jaap; Jansen, Tim L.; Neogi, Tuhina; Schumacher, H. Ralph; Taylor, William J.

    2013-01-01

    The definitive classification or diagnosis of gout normally relies upon the identification of MSU crystals in SF or from tophi. Where microscopic examination of SF is not available or is impractical, the best approach may differ depending upon the context. For many types of research, clinical classification criteria are necessary. The increasing prevalence of gout, advances in therapeutics and the development of international research collaborations to understand the impact, mechanisms and optimal treatment of this condition emphasize the need for accurate and uniform classification criteria for gout. Five clinical classification criteria for gout currently exist. However, none of the currently available criteria has been adequately validated. An international project is currently under way to develop new validated gout classification criteria. These criteria will be an essential step forward to advance the research agenda in the modern era of gout management. PMID:23611919

  6. A New Classification Method to Overcome Over-Branching

    Institute of Scientific and Technical Information of China (English)

    ZHOU Aoying(周傲英); QIAN Weining(钱卫宁); QIAN Hailei(钱海蕾); JIN Wen(金文)

    2002-01-01

    Classification is an important technique in data mining. The decision trees built by most of the existing classification algorithms commonly feature over-branching, which will lead to poor efficiency in the subsequent classification period. In this paper, we present a new value-oriented classification method, which aims at building accurately proper-sized decision trees while reducing over-branching as much as possible, based on the concepts of frequentpattern-node and exceptive-child-node. The experiments show that while using relevant analysis as pre-processing, our classification method, without loss of accuracy, can eliminate the over-branching greatly in decision trees more effectively and efficiently than other algorithms do.

  7. Intraregional classification of wine via ICP-MS elemental fingerprinting.

    Science.gov (United States)

    Coetzee, P P; van Jaarsveld, F P; Vanhaecke, F

    2014-12-01

    The feasibility of elemental fingerprinting in the classification of wines according to their provenance vineyard soil was investigated in the relatively small geographical area of a single wine district. Results for the Stellenbosch wine district (Western Cape Wine Region, South Africa), comprising an area of less than 1,000 km(2), suggest that classification of wines from different estates (120 wines from 23 estates) is indeed possible using accurate elemental data and multivariate statistical analysis based on a combination of principal component analysis, cluster analysis, and discriminant analysis. This is the first study to demonstrate the successful classification of wines at estate level in a single wine district in South Africa. The elements B, Ba, Cs, Cu, Mg, Rb, Sr, Tl and Zn were identified as suitable indicators. White and red wines were grouped in separate data sets to allow successful classification of wines. Correlation between wine classification and soil type distributions in the area was observed.

  8. Behavior Based Social Dimensions Extraction for Multi-Label Classification.

    Science.gov (United States)

    Li, Le; Xu, Junyi; Xiao, Weidong; Ge, Bin

    2016-01-01

    Classification based on social dimensions is commonly used to handle the multi-label classification task in heterogeneous networks. However, traditional methods, which mostly rely on the community detection algorithms to extract the latent social dimensions, produce unsatisfactory performance when community detection algorithms fail. In this paper, we propose a novel behavior based social dimensions extraction method to improve the classification performance in multi-label heterogeneous networks. In our method, nodes' behavior features, instead of community memberships, are used to extract social dimensions. By introducing Latent Dirichlet Allocation (LDA) to model the network generation process, nodes' connection behaviors with different communities can be extracted accurately, which are applied as latent social dimensions for classification. Experiments on various public datasets reveal that the proposed method can obtain satisfactory classification results in comparison to other state-of-the-art methods on smaller social dimensions.

  9. Cancer classification based on gene expression using neural networks.

    Science.gov (United States)

    Hu, H P; Niu, Z J; Bai, Y P; Tan, X H

    2015-12-21

    Based on gene expression, we have classified 53 colon cancer patients with UICC II into two groups: relapse and no relapse. Samples were taken from each patient, and gene information was extracted. Of the 53 samples examined, 500 genes were considered proper through analyses by S-Kohonen, BP, and SVM neural networks. Classification accuracy obtained by S-Kohonen neural network reaches 91%, which was more accurate than classification by BP and SVM neural networks. The results show that S-Kohonen neural network is more plausible for classification and has a certain feasibility and validity as compared with BP and SVM neural networks.

  10. A Note on Encodings of Phylogenetic Networks of Bounded Level

    CERN Document Server

    Gambette, Philippe

    2009-01-01

    Driven by the need for better models that allow one to shed light into the question how life's diversity has evolved, phylogenetic networks have now joined phylogenetic trees in the center of phylogenetics research. Like phylogenetic trees, such networks canonically induce collections of phylogenetic trees, clusters, and triplets, respectively. Thus it is not surprising that many network approaches aim to reconstruct a phylogenetic network from such collections. Related to the well-studied perfect phylogeny problem, the following question is of fundamental importance in this context: When does one of the above collections encode (i.e. uniquely describe) the network that induces it? In this note, we present a complete answer to this question for the special case of a level-1 (phylogenetic) network by characterizing those level-1 networks for which an encoding in terms of one (or equivalently all) of the above collections exists. Given that this type of network forms the first layer of the rich hierarchy of lev...

  11. Rapid and accurate pyrosequencing of angiosperm plastid genomes

    Directory of Open Access Journals (Sweden)

    Farmerie William G

    2006-08-01

    Full Text Available Abstract Background Plastid genome sequence information is vital to several disciplines in plant biology, including phylogenetics and molecular biology. The past five years have witnessed a dramatic increase in the number of completely sequenced plastid genomes, fuelled largely by advances in conventional Sanger sequencing technology. Here we report a further significant reduction in time and cost for plastid genome sequencing through the successful use of a newly available pyrosequencing platform, the Genome Sequencer 20 (GS 20 System (454 Life Sciences Corporation, to rapidly and accurately sequence the whole plastid genomes of the basal eudicot angiosperms Nandina domestica (Berberidaceae and Platanus occidentalis (Platanaceae. Results More than 99.75% of each plastid genome was simultaneously obtained during two GS 20 sequence runs, to an average depth of coverage of 24.6× in Nandina and 17.3× in Platanus. The Nandina and Platanus plastid genomes shared essentially identical gene complements and possessed the typical angiosperm plastid structure and gene arrangement. To assess the accuracy of the GS 20 sequence, over 45 kilobases of sequence were generated for each genome using conventional sequencing. Overall error rates of 0.043% and 0.031% were observed in GS 20 sequence for Nandina and Platanus, respectively. More than 97% of all observed errors were associated with homopolymer runs, with ~60% of all errors associated with homopolymer runs of 5 or more nucleotides and ~50% of all errors associated with regions of extensive homopolymer runs. No substitution errors were present in either genome. Error rates were generally higher in the single-copy and noncoding regions of both plastid genomes relative to the inverted repeat and coding regions. Conclusion Highly accurate and essentially complete sequence information was obtained for the Nandina and Platanus plastid genomes using the GS 20 System. More importantly, the high accuracy

  12. Identifying early events of gene expression in breast cancer with systems biology phylogenetics.

    Science.gov (United States)

    Abu-Asab, M S; Abu-Asab, N; Loffredo, C A; Clarke, R; Amri, H

    2013-01-01

    Advanced omics technologies such as deep sequencing and spectral karyotyping are revealing more of cancer heterogeneity at the genetic, genomic, gene expression, epigenetic, proteomic, and metabolomic levels. With this increasing body of emerging data, the task of data analysis becomes critical for mining and modeling to better understand the relevant underlying biological processes. However, the multiple levels of heterogeneity evident within and among populations, healthy and diseased, complicate the mining and interpretation of biological data, especially when dealing with hundreds to tens of thousands of variables. Heterogeneity occurs in many diseases, such as cancers, autism, macular degeneration, and others. In cancer, heterogeneity has hampered the search for validated biomarkers for early detection, and it has complicated the task of finding clonal (driver) and nonclonal (nonexpanded or passenger) aberrations. We show that subtyping of cancer (classification of specimens) should be an a priori step to the identification of early events of cancers. Studying early events in oncogenesis can be done on histologically normal tissues from diseased individuals (HNTDI), since they most likely have been exposed to the same mutagenic insults that caused the cancer in their neighboring tissues. Polarity assessment of HNTDI data variables by using healthy specimens as outgroup(s), followed by the application of parsimony phylogenetic analysis, produces a hierarchical classification of specimens that reveals the early events of the disease ontogeny within its subtypes as shared derived changes (abnormal changes) or synapomorphies in phylogenetic terminology. Copyright © 2013 S. Karger AG, Basel.

  13. Genome-wide identification and phylogenetic analysis of the ERF gene family in cucumbers

    Directory of Open Access Journals (Sweden)

    Lifang Hu

    2011-01-01

    Full Text Available Members of the ERF transcription-factor family participate in a number of biological processes, viz., responses to hormones, adaptation to biotic and abiotic stress, metabolism regulation, beneficial symbiotic interactions, cell differentiation and developmental processes. So far, no tissue-expression profile of any cucumber ERF protein has been reported in detail. Recent completion of the cucumber full-genome sequence has come to facilitate, not only genome-wide analysis of ERF family members in cucumbers themselves, but also a comparative analysis with those in Arabidopsis and rice. In this study, 103 hypothetical ERF family genes in the cucumber genome were identified, phylogenetic analysis indicating their classification into 10 groups, designated I to X. Motif analysis further indicated that most of the conserved motifs outside the AP2/ERF domain, are selectively distributed among the specific clades in the phylogenetic tree. From chromosomal localization and genome distribution analysis, it appears that tandem-duplication may have contributed to CsERF gene expansion. Intron/exon structure analysis indicated that a few CsERFs still conserved the former intron-position patterns existent in the common ancestor of monocots and eudicots. Expression analysis revealed the widespread distribution of the cucumber ERF gene family within plant tissues, thereby implying the probability of their performing various roles therein. Furthermore, members of some groups presented mutually similar expression patterns that might be related to their phylogenetic groups.

  14. Diversity of Clonostachys species assessed by molecular phylogenetics and MALDI-TOF mass spectrometry.

    Science.gov (United States)

    Abreu, Lucas M; Moreira, Gláucia M; Ferreira, Douglas; Rodrigues-Filho, Edson; Pfenning, Ludwig H

    2014-12-01

    We assessed the species diversity among 45 strains of Clonostachys from different substrates and localities in Brazil using molecular phylogenetics, and compared the results with the phenotypic classification of strains obtained from matrix-assisted laser desorption ionization time-of-flight mass spectrometry (MALDI-TOF MS). Phylogenetic analyses were based on beta tubulin (Tub), ITS-LSU rDNA, and a combined Tub-ITS DNA dataset. MALDI-TOF MS analyses were performed using intact conidia and conidiophores of strains cultivated on oatmeal agar and 4% malt extract agar. Six known species were identified: Clonostachys byssicola, Clonostachys candelabrum, Clonostachys pseudochroleuca, Clonostachys rhizophaga, Clonostachys rogersoniana, and Clonostachys rosea. Two clades and two singleton lineages did not correspond to known species represented in the reference DNA dataset and were identified as Clonostachys sp. 1-4. Multivariate cluster analyses of MALDI-TOF MS data classified the strains into eight clusters and three singletons, corresponding to the ten identified species plus one additional cluster containing two strains of C. rogersoniana that split from the other co-specific strains. The consistent results of MALDI-TOF MS supported the identification of strains assigned to C. byssicola and C. pseudochroleuca, which did not form well supported clades in all phylogenetic analyses, but formed distinct clusters in the MALDI-TOF dendrograms.

  15. Molecular phylogenetics and historical biogeography amid shifting continents in the cockles and giant clams (Bivalvia: Cardiidae).

    Science.gov (United States)

    Herrera, Nathanael D; Ter Poorten, Jan Johan; Bieler, Rüdiger; Mikkelsen, Paula M; Strong, Ellen E; Jablonski, David; Steppan, Scott J

    2015-12-01

    Reconstructing historical biogeography of the marine realm is complicated by indistinct barriers and, over deeper time scales, a dynamic landscape shaped by plate tectonics. Here we present the most extensive examination of model-based historical biogeography among marine invertebrates to date. We conducted the largest phylogenetic and molecular clock analyses to date for the bivalve family Cardiidae (cockles and giant clams) with three unlinked loci for 110 species representing 37 of the 50 genera. Ancestral ranges were reconstructed using the dispersal-extinction-cladogenesis (DEC) method with a time-stratified paleogeographic model wherein dispersal rates varied with shifting tectonics. Results were compared to previous classifications and the extensive paleontological record. Six of the eight prior subfamily groupings were found to be para- or polyphyletic. Cardiidae originated and subsequently diversified in the tropical Indo-Pacific starting in the Late Triassic. Eastern Atlantic species were mainly derived from the tropical Indo-Mediterranean region via the Tethys Sea. In contrast, the western Atlantic fauna was derived from Indo-Pacific clades. Our phylogenetic results demonstrated greater concordance with geography than did previous phylogenies based on morphology. Time-stratifying the DEC reconstruction improved the fit and was highly consistent with paleo-ocean currents and paleogeography. Lastly, combining molecular phylogenetics with a rich and well-documented fossil record allowed us to test the accuracy and precision of biogeographic range reconstructions.

  16. Molecular phylogenetics unveils the ancient evolutionary origins of the enigmatic fairy armadillos.

    Science.gov (United States)

    Delsuc, Frédéric; Superina, Mariella; Tilak, Marie-Ka; Douzery, Emmanuel J P; Hassanin, Alexandre

    2012-02-01

    Fairy armadillos or pichiciegos (Xenarthra, Dasypodidae) are among the most elusive mammals. Due to their subterranean and nocturnal lifestyle, their basic biology and evolutionary history remain virtually unknown. Two distinct species with allopatric distributions are recognized: Chlamyphorus truncatus is restricted to central Argentina, while Calyptophractus retusus occurs in the Gran Chaco of Argentina, Paraguay, and Bolivia. To test their monophyly and resolve their phylogenetic affinities within armadillos, we obtained sequence data from modern and museum specimens for two mitochondrial genes (12S RNA [MT-RNR1] and NADH dehydrogenase 1 [MT-ND1]) and two nuclear exons (breast cancer 1 early onset exon 11 [BRCA1] and von Willebrand factor exon 28 [VWF]). Phylogenetic analyses provided a reference phylogeny and timescale for living xenarthran genera. Our results reveal monophyletic pichiciegos as members of a major armadillo subfamily (Chlamyphorinae). Their strictly fossorial lifestyle probably evolved as a response to the Oligocene aridification that occurred in South America after their divergence from Tolypeutinae around 32 million years ago (Mya). The ancient divergence date (∼17Mya) for separation between the two species supports their taxonomic classification into distinct genera. The synchronicity with Middle Miocene marine incursions along the Paraná river basin suggests a vicariant origin for pichiciegos by the disruption of their ancestral range. Their phylogenetic distinctiveness and rarity in the wild argue in favor of high conservation priority.

  17. Phylogenetic relationships of 18 passerines based on Adenylate Kinase Intron 5 sequences

    Institute of Scientific and Technical Information of China (English)

    GUO Hui-yan; YU Hui-xin; BAI Su-ying; MA Yu-kun

    2008-01-01

    The 18 species of bird studied originally are known to belong to muscicapids, robins and sylviids of passerines, but some disputations are always present in their classification systems. In this experiment, phylogenetic relationships of 18 species of passerines were studied using Adenylate Kinase Intron 5 (AK5) sequences and DNA techniques. Through sequences analysis in comparison with each other, phylogenetic tree figures of 18 species of passerines were constructed using Neighbor-Joining (NJ) and Maximum-Parsimony (MP) methods . The results showed that sylviids should be listed as an independent family, while robins and flycatchers should be listed into Muscicapidae. Since the phylogenetic relationships between long-tailed tits and old world warblers are closer than that between long-tailed tits and parids, the long-tailed tits should be independent of paridae and be categorized into aegithalidae. Muscicapidae and Paridae are known to be two monophylitic families, but Sylviidae is not a monophyletic group. AK5 sequences had better efficacy in resolving close relationships of interspecies among intrageneric groups.

  18. Complete mitochondrial genomes elucidate phylogenetic relationships of the deep-sea octocoral families Coralliidae and Paragorgiidae

    Science.gov (United States)

    Figueroa, Diego F.; Baco, Amy R.

    2014-01-01

    In the past decade, molecular phylogenetic analyses of octocorals have shown that the current morphological taxonomic classification of these organisms needs to be revised. The latest phylogenetic analyses show that most octocorals can be divided into three main clades. One of these clades contains the families Coralliidae and Paragorgiidae. These families share several taxonomically important characters and it has been suggested that they may not be monophyletic; with the possibility of the Coralliidae being a derived branch of the Paragorgiidae. Uncertainty exists not only in the relationship of these two families, but also in the classification of the two genera that make up the Coralliidae, Corallium and Paracorallium. Molecular analyses suggest that the genus Corallium is paraphyletic, and it can be divided into two main clades, with the Paracorallium as members of one of these clades. In this study we sequenced the whole mitochondrial genome of five species of Paragorgia and of five species of Corallium to use in a phylogenetic analysis to achieve two main objectives; the first to elucidate the phylogenetic relationship between the Paragorgiidae and Coralliidae and the second to determine whether the genera Corallium and Paracorallium are monophyletic. Our results show that other members of the Coralliidae share the two novel mitochondrial gene arrangements found in a previous study in Corallium konojoi and Paracorallium japonicum; and that the Corallium konojoi arrangement is also found in the Paragorgiidae. Our phylogenetic reconstruction based on all the protein coding genes and ribosomal RNAs of the mitochondrial genome suggest that the Coralliidae are not a derived branch of the Paragorgiidae, but rather a monophyletic sister branch to the Paragorgiidae. While our manuscript was in review a study was published using morphological data and several fragments from mitochondrial genes to redefine the taxonomy of the Coralliidae. Paracorallium was subsumed

  19. Multilocus phylogeny of the New-World mud turtles (Kinosternidae) supports the traditional classification of the group.

    Science.gov (United States)

    Spinks, Phillip Q; Thomson, Robert C; Gidiş, Müge; Bradley Shaffer, H

    2014-07-01

    A goal of modern taxonomy is to develop classifications that reflect current phylogenetic relationships and are as stable as possible given the inherent uncertainties in much of the tree of life. Here, we provide an in-depth phylogenetic analysis, based on 14 nuclear loci comprising 10,305 base pairs of aligned sequence data from all but two species of the turtle family Kinosternidae, to determine whether recent proposed changes to the group's classification are justified and necessary. We conclude that those proposed changes were based on (1) mtDNA gene tree anomalies, (2) preliminary analyses that do not fully capture the breadth of geographic variation necessary to motivate taxonomic changes, and (3) changes in rank that are not motivated by non-monophyletic groups. Our recommendation, for this and other similar cases, is that taxonomic changes be made only when phylogenetic results that are statistically well-supported and corroborated by multiple independent lines of genetic evidence indicate that non-monophyletic groups are currently recognized and need to be corrected. We hope that other members of the phylogenetics community will join us in proposing taxonomic changes only when the strongest phylogenetic data demand such changes, and in so doing that we can move toward stable, phylogenetically informed classifications of lasting value.

  20. [Classification of human sleep stages based on EEG processing using hidden Markov models].

    Science.gov (United States)

    Doroshenkov, L G; Konyshev, V A; Selishchev, S V

    2007-01-01

    The goal of this work was to describe an automated system for classification of human sleep stages. Classification of sleep stages is an important problem of diagnosis and treatment of human sleep disorders. The developed classification method is based on calculation of characteristics of the main sleep rhythms. It uses hidden Markov models. The method is highly accurate and provides reliable identification of the main stages of sleep. The results of automatic classification are in good agreement with the results of sleep stage identification performed by an expert somnologist using Rechtschaffen and Kales rules. This substantiates the applicability of the developed classification system to clinical diagnosis.

  1. A phylum-level phylogenetic classification of zygomycete fungi based on genome-scale data

    Science.gov (United States)

    Zygomycete fungi were classified as a single phylum, Zygomycota, based on sexual reproduction by zygospores, frequent asexual reproduction by sporangia, absence of multicellular sporocarps, and production of coenocytic hyphae, all with some exceptions. Molecular phylogenies based on one or a few gen...

  2. Phylogenomics of zygomycete fungi: impacts on a phylogenetic classification of Kingdom Fungi

    Science.gov (United States)

    The zygomycetous fungi (”zygomycetes”) mark the major transition from zoosporic life histories of the common ancestor of Fungi and the earliest diverging chytrid lineages (Chytridiomycota and Blastocladiomycota). Their ecological and economic importance range from the earliest documented symbionts o...

  3. Automated protein subfamily identification and classification.

    Directory of Open Access Journals (Sweden)

    Duncan P Brown

    2007-08-01

    Full Text Available Function prediction by homology is widely used to provide preliminary functional annotations for genes for which experimental evidence of function is unavailable or limited. This approach has been shown to be prone to systematic error, including percolation of annotation errors through sequence databases. Phylogenomic analysis avoids these errors in function prediction but has been difficult to automate for high-throughput application. To address this limitation, we present a computationally efficient pipeline for phylogenomic classification of proteins. This pipeline uses the SCI-PHY (Subfamily Classification in Phylogenomics algorithm for automatic subfamily identification, followed by subfamily hidden Markov model (HMM construction. A simple and computationally efficient scoring scheme using family and subfamily HMMs enables classification of novel sequences to protein families and subfamilies. Sequences representing entirely novel subfamilies are differentiated from those that can be classified to subfamilies in the input training set using logistic regression. Subfamily HMM parameters are estimated using an information-sharing protocol, enabling subfamilies containing even a single sequence to benefit from conservation patterns defining the family as a whole or in related subfamilies. SCI-PHY subfamilies correspond closely to functional subtypes defined by experts and to conserved clades found by phylogenetic analysis. Extensive comparisons of subfamily and family HMM performances show that subfamily HMMs dramatically improve the separation between homologous and non-homologous proteins in sequence database searches. Subfamily HMMs also provide extremely high specificity of classification and can be used to predict entirely novel subtypes. The SCI-PHY Web server at http://phylogenomics.berkeley.edu/SCI-PHY/ allows users to upload a multiple sequence alignment for subfamily identification and subfamily HMM construction. Biologists wishing to

  4. Phylogenetic analysis of cubilin (CUBN) gene

    OpenAIRE

    Shaik, Abjal Pasha; Alsaeed, Abbas H; Kiranmayee, S; Bammidi, VK; Sultana, Asma

    2013-01-01

    Cubilin, (CUBN; also known as intrinsic factor-cobalamin receptor [Homo sapiens Entrez Pubmed ref NM_001081.3; NG_008967.1; GI: 119606627]), located in the epithelium of intestine and kidney acts as a receptor for intrinsic factor – vitamin B12 complexes. Mutations in CUBN may play a role in autosomal recessive megaloblastic anemia. The current study investigated the possible role of CUBN in evolution using phylogenetic testing. A total of 588 BLAST hits were found for the cubilin query seque...

  5. MINER: software for phylogenetic motif identification.

    Science.gov (United States)

    La, David; Livesay, Dennis R

    2005-07-01

    MINER is web-based software for phylogenetic motif (PM) identification. PMs are sequence regions (fragments) that conserve the overall familial phylogeny. PMs have been shown to correspond to a wide variety of catalytic regions, substrate-binding sites and protein interfaces, making them ideal functional site predictions. The MINER output provides an intuitive interface for interactive PM sequence analysis and structural visualization. The web implementation of MINER is freely available at http://www.pmap.csupomona.edu/MINER/. Source code is available to the academic community on request.

  6. Evidence of Statistical Inconsistency of Phylogenetic Methods in the Presence of Multiple Sequence Alignment Uncertainty.

    Science.gov (United States)

    Md Mukarram Hossain, A S; Blackburne, Benjamin P; Shah, Abhijeet; Whelan, Simon

    2015-07-01

    Evolutionary studies usually use a two-step process to investigate sequence data. Step one estimates a multiple sequence alignment (MSA) and step two applies phylogenetic methods to ask evolutionary questions of that MSA. Modern phylogenetic methods infer evolutionary parameters using maximum likelihood or Bayesian inference, mediated by a probabilistic substitution model that describes sequence change over a tree. The statistical properties of these methods mean that more data directly translates to an increased confidence in downstream results, providing the substitution model is adequate and the MSA is correct. Many studies have investigated the robustness of phylogenetic methods in the presence of substitution model misspecification, but few have examined the statistical properties of those methods when the MSA is unknown. This simulation study examines the statistical properties of the complete two-step process when inferring sequence divergence and the phylogenetic tree topology. Both nucleotide and amino acid analyses are negatively affected by the alignment step, both through inaccurate guide tree estimates and through overfitting to that guide tree. For many alignment tools these effects become more pronounced when additional sequences are added to the analysis. Nucleotide sequences are particularly susceptible, with MSA errors leading to statistical support for long-branch attraction artifacts, which are usually associated with gross substitution model misspecification. Amino acid MSAs are more robust, but do tend to arbitrarily resolve multifurcations in favor of the guide tree. No inference strategies produce consistently accurate estimates of divergence between sequences, although amino acid MSAs are again more accurate than their nucleotide counterparts. We conclude with some practical suggestions about how to limit the effect of MSA uncertainty on evolutionary inference.

  7. Two new rapid SNP-typing methods for classifying Mycobacterium tuberculosis complex into the main phylogenetic lineages.

    Directory of Open Access Journals (Sweden)

    David Stucki

    Full Text Available There is increasing evidence that strain variation in Mycobacterium tuberculosis complex (MTBC might influence the outcome of tuberculosis infection and disease. To assess genotype-phenotype associations, phylogenetically robust molecular markers and appropriate genotyping tools are required. Most current genotyping methods for MTBC are based on mobile or repetitive DNA elements. Because these elements are prone to convergent evolution, the corresponding genotyping techniques are suboptimal for phylogenetic studies and strain classification. By contrast, single nucleotide polymorphisms (SNP are ideal markers for classifying MTBC into phylogenetic lineages, as they exhibit very low degrees of homoplasy. In this study, we developed two complementary SNP-based genotyping methods to classify strains into the six main human-associated lineages of MTBC, the "Beijing" sublineage, and the clade comprising Mycobacterium bovis and Mycobacterium caprae. Phylogenetically informative SNPs were obtained from 22 MTBC whole-genome sequences. The first assay, referred to as MOL-PCR, is a ligation-dependent PCR with signal detection by fluorescent microspheres and a Luminex flow cytometer, which simultaneously interrogates eight SNPs. The second assay is based on six individual TaqMan real-time PCR assays for singleplex SNP-typing. We compared MOL-PCR and TaqMan results in two panels of clinical MTBC isolates. Both methods agreed fully when assigning 36 well-characterized strains into the main phylogenetic lineages. The sensitivity in allele-calling was 98.6% and 98.8% for MOL-PCR and TaqMan, respectively. Typing of an additional panel of 78 unknown clinical isolates revealed 99.2% and 100% sensitivity in allele-calling, respectively, and 100% agreement in lineage assignment between both methods. While MOL-PCR and TaqMan are both highly sensitive and specific, MOL-PCR is ideal for classification of isolates with no previous information, whereas TaqMan is faster

  8. 78 FR 54970 - Cotton Futures Classification: Optional Classification Procedure

    Science.gov (United States)

    2013-09-09

    ... process in March 2012 (77 FR 5379). When verified by a futures classification, Smith-Doxey data serves as... Classification: Optional Classification Procedure AGENCY: Agricultural Marketing Service, USDA. ACTION: Proposed... for the addition of an optional cotton futures classification procedure--identified and known...

  9. Pitch Based Sound Classification

    DEFF Research Database (Denmark)

    Nielsen, Andreas Brinch; Hansen, Lars Kai; Kjems, U

    2006-01-01

    A sound classification model is presented that can classify signals into music, noise and speech. The model extracts the pitch of the signal using the harmonic product spectrum. Based on the pitch estimate and a pitch error measure, features are created and used in a probabilistic model with soft......-max output function. Both linear and quadratic inputs are used. The model is trained on 2 hours of sound and tested on publicly available data. A test classification error below 0.05 with 1 s classification windows is achieved. Further more it is shown that linear input performs as well as a quadratic......, and that even though classification gets marginally better, not much is achieved by increasing the window size beyond 1 s....

  10. Learning Apache Mahout classification

    CERN Document Server

    Gupta, Ashish

    2015-01-01

    If you are a data scientist who has some experience with the Hadoop ecosystem and machine learning methods and want to try out classification on large datasets using Mahout, this book is ideal for you. Knowledge of Java is essential.

  11. [Classification of cardiomyopathy].

    Science.gov (United States)

    Asakura, Masanori; Kitakaze, Masafumi

    2014-01-01

    Cardiomyopathy is a group of cardiovascular diseases with poor prognosis. Some patients with dilated cardiomyopathy need heart transplantations due to severe heart failure. Some patients with hypertrophic cardiomyopathy die unexpectedly due to malignant ventricular arrhythmias. Various phenotypes of cardiomyopathies are due to the heterogeneous group of diseases. The classification of cardiomyopathies is important and indispensable in the clinical situation. However, their classification has not been established, because the causes of cardiomyopathies have not been fully elucidated. We usually use definition and classification offered by WHO/ISFC task force in 1995. Recently, several new definitions and classifications of the cardiomyopathies have been published by American Heart Association, European Society of Cardiology and Japanese Circulation Society.

  12. Carbohydrate terminology and classification

    National Research Council Canada - National Science Library

    Cummings, J H; Stephen, A M

    2007-01-01

    ...) and polysaccharides (DP> or =10). Within this classification, a number of terms are used such as mono- and disaccharides, polyols, oligosaccharides, starch, modified starch, non-starch polysaccharides, total carbohydrate, sugars, etc...

  13. Expected Classification Accuracy

    Directory of Open Access Journals (Sweden)

    Lawrence M. Rudner

    2005-08-01

    Full Text Available Every time we make a classification based on a test score, we should expect some number..of misclassifications. Some examinees whose true ability is within a score range will have..observed scores outside of that range. A procedure for providing a classification table of..true and expected scores is developed for polytomously scored items under item response..theory and applied to state assessment data. A simplified procedure for estimating the..table entries is also presented.

  14. Completion of the classification

    CERN Document Server

    Strade, Helmut

    2012-01-01

    This is the last of three volumes about ""Simple Lie Algebras over Fields of Positive Characteristic""by Helmut Strade, presenting the state of the art of the structure and classification of Lie algebras over fields of positive characteristic. In this monograph the proof of the Classification Theorem presented in the first volumeis concluded.Itcollects all the important results on the topic whichcan be found only in scatteredscientific literaturso far.

  15. Twitter content classification

    OpenAIRE

    2010-01-01

    This paper delivers a new Twitter content classification framework based sixteen existing Twitter studies and a grounded theory analysis of a personal Twitter history. It expands the existing understanding of Twitter as a multifunction tool for personal, profession, commercial and phatic communications with a split level classification scheme that offers broad categorization and specific sub categories for deeper insight into the real world application of the service.

  16. Epitope discovery with phylogenetic hidden Markov models.

    LENUS (Irish Health Repository)

    Lacerda, Miguel

    2010-05-01

    Existing methods for the prediction of immunologically active T-cell epitopes are based on the amino acid sequence or structure of pathogen proteins. Additional information regarding the locations of epitopes may be acquired by considering the evolution of viruses in hosts with different immune backgrounds. In particular, immune-dependent evolutionary patterns at sites within or near T-cell epitopes can be used to enhance epitope identification. We have developed a mutation-selection model of T-cell epitope evolution that allows the human leukocyte antigen (HLA) genotype of the host to influence the evolutionary process. This is one of the first examples of the incorporation of environmental parameters into a phylogenetic model and has many other potential applications where the selection pressures exerted on an organism can be related directly to environmental factors. We combine this novel evolutionary model with a hidden Markov model to identify contiguous amino acid positions that appear to evolve under immune pressure in the presence of specific host immune alleles and that therefore represent potential epitopes. This phylogenetic hidden Markov model provides a rigorous probabilistic framework that can be combined with sequence or structural information to improve epitope prediction. As a demonstration, we apply the model to a data set of HIV-1 protein-coding sequences and host HLA genotypes.

  17. A phylogenetic blueprint for a modern whale.

    Science.gov (United States)

    Gatesy, John; Geisler, Jonathan H; Chang, Joseph; Buell, Carl; Berta, Annalisa; Meredith, Robert W; Springer, Mark S; McGowen, Michael R

    2013-02-01

    The emergence of Cetacea in the Paleogene represents one of the most profound macroevolutionary transitions within Mammalia. The move from a terrestrial habitat to a committed aquatic lifestyle engendered wholesale changes in anatomy, physiology, and behavior. The results of this remarkable transformation are extant whales that include the largest, biggest brained, fastest swimming, loudest, deepest diving mammals, some of which can detect prey with a sophisticated echolocation system (Odontoceti - toothed whales), and others that batch feed using racks of baleen (Mysticeti - baleen whales). A broad-scale reconstruction of the evolutionary remodeling that culminated in extant cetaceans has not yet been based on integration of genomic and paleontological information. Here, we first place Cetacea relative to extant mammalian diversity, and assess the distribution of support among molecular datasets for relationships within Artiodactyla (even-toed ungulates, including Cetacea). We then merge trees derived from three large concatenations of molecular and fossil data to yield a composite hypothesis that encompasses many critical events in the evolutionary history of Cetacea. By combining diverse evidence, we infer a phylogenetic blueprint that outlines the stepwise evolutionary development of modern whales. This hypothesis represents a starting point for more detailed, comprehensive phylogenetic reconstructions in the future, and also highlights the synergistic interaction between modern (genomic) and traditional (morphological+paleontological) approaches that ultimately must be exploited to provide a rich understanding of evolutionary history across the entire tree of Life.

  18. A Distance Measure for Genome Phylogenetic Analysis

    Science.gov (United States)

    Cao, Minh Duc; Allison, Lloyd; Dix, Trevor

    Phylogenetic analyses of species based on single genes or parts of the genomes are often inconsistent because of factors such as variable rates of evolution and horizontal gene transfer. The availability of more and more sequenced genomes allows phylogeny construction from complete genomes that is less sensitive to such inconsistency. For such long sequences, construction methods like maximum parsimony and maximum likelihood are often not possible due to their intensive computational requirement. Another class of tree construction methods, namely distance-based methods, require a measure of distances between any two genomes. Some measures such as evolutionary edit distance of gene order and gene content are computational expensive or do not perform well when the gene content of the organisms are similar. This study presents an information theoretic measure of genetic distances between genomes based on the biological compression algorithm expert model. We demonstrate that our distance measure can be applied to reconstruct the consensus phylogenetic tree of a number of Plasmodium parasites from their genomes, the statistical bias of which would mislead conventional analysis methods. Our approach is also used to successfully construct a plausible evolutionary tree for the γ-Proteobacteria group whose genomes are known to contain many horizontally transferred genes.

  19. A molecular phylogenetic evaluation of the spizellomycetales.

    Science.gov (United States)

    Wakefield, William S; Powell, Martha J; Letcher, Peter M; Barr, Donald J S; Churchill, Perry F; Longcore, Joyce E; Chen, Shu-Fen

    2010-01-01

    Order Spizellomycetales was delineated based on a unique suite of zoospore ultrastructural characters and currently includes five genera and 14 validly published species, all of which have a propensity for soil habitats. We generated DNA sequences from small (SSU), large (LSU) and 5.8S ribosomal subunit genes to assess the monophyly of all genera and species in this order. The 53 cultures analyzed included isolates on which all described species were based, plus other spizellomycetalean cultures. Phylogenetic placement of these chytrids was explored with maximum parsimony and maximum likelihood analyses, both of which yielded comparable topologies. Kochiomyces, Powellomyces and Triparticalcar were monophyletic, while Gaertneriomyces and Spizellomyces were polyphyletic. Isolates, distinct from described species, clustered among each of the five genera, indicating that species diversity in genera is greater than currently recognized. One isolate formed a clade that included no described species, representing a new genus. Zoospore ultrastructural features and architecture seem to be good indicators of phylogenetic relationships, but finer scrutiny of characters such as kinetosome-associated structures (KAS) is needed to understand more clearly the diversity within this order as it is revised.

  20. Phylogenetic diversity of Mesorhizobium in chickpea

    Indian Academy of Sciences (India)

    Dong Hyun Kim; Mayank Kaashyap; Abhishek Rathore; Roma R Das; Swathi Parupalli; Hari D Upadhyaya; S Gopalakrishnan; Pooran M Gaur; Sarvjeet Singh; Jagmeet Kaur; Mohammad Yasin; Rajeev K Varshney

    2014-06-01

    Crop domestication, in general, has reduced genetic diversity in cultivated gene pool of chickpea (Cicer arietinum) as compared with wild species (C. reticulatum, C. bijugum). To explore impact of domestication on symbiosis, 10 accessions of chickpeas, including 4 accessions of C. arietinum, and 3 accessions of each of C. reticulatum and C. bijugum species, were selected and DNAs were extracted from their nodules. To distinguish chickpea symbiont, preliminary sequences analysis was attempted with 9 genes (16S rRNA, atpD, dnaJ, glnA, gyrB, nifH, nifK, nodD and recA) of which 3 genes (gyrB, nifK and nodD) were selected based on sufficient sequence diversity for further phylogenetic analysis. Phylogenetic analysis and sequence diversity for 3 genes demonstrated that sequences from C. reticulatum were more diverse. Nodule occupancy by dominant symbiont also indicated that C. reticulatum (60%) could have more various symbionts than cultivated chickpea (80%). The study demonstrated that wild chickpeas (C. reticulatum) could be used for selecting more diverse symbionts in the field conditions and it implies that chickpea domestication affected symbiosis negatively in addition to reducing genetic diversity.

  1. Comparative assessment of performance and genome dependence among phylogenetic profiling methods

    Directory of Open Access Journals (Sweden)

    Wu Jie

    2006-09-01

    Full Text Available Abstract Background The rapidly increasing speed with which genome sequence data can be generated will be accompanied by an exponential increase in the number of sequenced eukaryotes. With the increasing number of sequenced eukaryotic genomes comes a need for bioinformatic techniques to aid in functional annotation. Ideally, genome context based techniques such as proximity, fusion, and phylogenetic profiling, which have been so successful in prokaryotes, could be utilized in eukaryotes. Here we explore the application of phylogenetic profiling, a method that exploits the evolutionary co-occurrence of genes in the assignment of functional linkages, to eukaryotic genomes. Results In order to evaluate the performance of phylogenetic profiling in eukaryotes, we assessed the relative performance of commonly used profile construction techniques and genome compositions in predicting functional linkages in both prokaryotic and eukaryotic organisms. When predicting linkages in E. coli with a prokaryotic profile, the use of continuous values constructed from transformed BLAST bit-scores performed better than profiles composed of discretized E-values; the use of discretized E-values resulted in more accurate linkages when using S. cerevisiae as the query organism. Extending this analysis by incorporating several eukaryotic genomes in profiles containing a majority of prokaryotes resulted in similar overall accuracy, but with a surprising reduction in pathway diversity among the most significant linkages. Furthermore, the application of phylogenetic profiling using profiles composed of only eukaryotes resulted in the loss of the strong correlation between common KEGG pathway membership and profile similarity score. Profile construction methods, orthology definitions, ontology and domain complexity were explored as possible sources of the poor performance of eukaryotic profiles, but with no improvement in results. Conclusion Given the current set of

  2. The origin and diversification of eukaryotes: problems with molecular phylogenetics and molecular clock estimation.

    Science.gov (United States)

    Roger, Andrew J; Hug, Laura A

    2006-06-29

    Determining the relationships among and divergence times for the major eukaryotic lineages remains one of the most important and controversial outstanding problems in evolutionary biology. The sequencing and phylogenetic analyses of ribosomal RNA (rRNA) genes led to the first nearly comprehensive phylogenies of eukaryotes in the late 1980s, and supported a view where cellular complexity was acquired during the divergence of extant unicellular eukaryote lineages. More recently, however, refinements in analytical methods coupled with the availability of many additional genes for phylogenetic analysis showed that much of the deep structure of early rRNA trees was artefactual. Recent phylogenetic analyses of a multiple genes and the discovery of important molecular and ultrastructural phylogenetic characters have resolved eukaryotic diversity into six major hypothetical groups. Yet relationships among these groups remain poorly understood because of saturation of sequence changes on the billion-year time-scale, possible rapid radiations of major lineages, phylogenetic artefacts and endosymbiotic or lateral gene transfer among eukaryotes. Estimating the divergence dates between the major eukaryote lineages using molecular analyses is even more difficult than phylogenetic estimation. Error in such analyses comes from a myriad of sources including: (i) calibration fossil dates, (ii) the assumed phylogenetic tree, (iii) the nucleotide or amino acid substitution model, (iv) substitution number (branch length) estimates, (v) the model of how rates of evolution change over the tree, (vi) error inherent in the time estimates for a given model and (vii) how multiple gene data are treated. By reanalysing datasets from recently published molecular clock studies, we show that when errors from these various sources are properly accounted for, the confidence intervals on inferred dates can be very large. Furthermore, estimated dates of divergence vary hugely depending on the methods

  3. Laboratory Building for Accurate Determination of Plutonium

    Institute of Scientific and Technical Information of China (English)

    2008-01-01

    <正>The accurate determination of plutonium is one of the most important assay techniques of nuclear fuel, also the key of the chemical measurement transfer and the base of the nuclear material balance. An

  4. Phylogenetics of early branching eudicots: Comparing phylogenetic signal across plastid introns, spacers, and genes

    Institute of Scientific and Technical Information of China (English)

    Anna-Magdalena BARNISKE; Thomas BORSCH; Kai M(U)LLER; Michael KRUG; Andreas WORBERG; Christoph NEINHUIS; Dietmar QUANDT

    2012-01-01

    Recent phylogenetic analyses revealed a grade with Ranunculales,Sabiales,Proteales,Trochodendrales,and Buxales as first branching eudicots,with the respective positions of Proteales and Sabiales still lacking statistical confidence.As previous analyses of conserved plastid genes remain inconclusive,we aimed to use and evaluate a representative set of plastid introns (group Ⅰ:trnL; group Ⅱ:petD,rpll6,trnK) and intergenic spacers (trnL-F,petB-petD,atpB-rbcL,rps3-rpll6) in comparison to the rapidly evolving matK and slowly evolving atpB and rbcL genes.Overall patterns of microstructural mutations converged across genomic regions,underscoring the existence of a general mutational pattern throughout the plastid genome.Phylogenetic signal differed strongly between functionally and structurally different genomic regions and was highest in matK,followed by spacers,then group Ⅱ and group Ⅰ introns.The more conserved atpB and rbcL coding regions showed distinctly lower phylogenetic information content.Parsimony,maximum likelihood,and Bayesian phylogenetic analyses based on the combined dataset of non-coding and rapidly evolving regions (>14 000 aligned characters) converged to a backbone topology ofeudicots with Ranunculales branching first,a Proteales-Sabiales clade second,followed by Trochodendrales and Buxales.Gunnerales generally appeared as sister to all remaining core eudicots with maximum support.Our results show that a small number of intron and spacer sequences allow similar insights into phylogenetic relationships of eudicots compared to datasets of many combined genes.The non-coding proportion of the plastid genome thus can be considered an important information source for plastid phylogenomics.

  5. Phylogenetic positions of several amitochondriate protozoa-Evidence from phylogenetic analysis of DNA topoisomerase II

    Institute of Scientific and Technical Information of China (English)

    HE De; DONG Jiuhong; WEN Jianfan; XIN Dedong; LU Siqi

    2005-01-01

    Several groups of parasitic protozoa, as represented by Giardia, Trichomonas, Entamoeba and Microsporida, were once widely considered to be the most primitive extant eukaryotic group―Archezoa. The main evidence for this is their 'lacking mitochondria' and possessing some other primitive features between prokaryotes and eukaryotes, and being basal to all eukaryotes with mitochondria in phylogenies inferred from many molecules. Some authors even proposed that these organisms diverged before the endosymbiotic origin of mitochondria within eukaryotes. This view was once considered to be very significant to the study of origin and evolution of eukaryotic cells (eukaryotes). However, in recent years this has been challenged by accumulating evidence from new studies. Here the sequences of DNA topoisomerase II in G. lamblia, T. vaginalis and E. histolytica were identified first by PCR and sequencing, then combining with the sequence data of the microsporidia Encephalitozoon cunicul and other eukaryotic groups of different evolutionary positions from GenBank, phylogenetic trees were constructed by various methods to investigate the evolutionary positions of these amitochondriate protozoa. Our results showed that since the characteristics of DNA topoisomerase II make it avoid the defect of 'long-branch attraction' appearing in the previous phylogenetic analyses, our trees can not only reflect effectively the relationship of different major eukaryotic groups, which is widely accepted, but also reveal phylogenetic positions for these amitochondriate protozoa, which is different from the previous phylogenetic trees. They are not the earliest-branching eukaryotes, but diverged after some mitochondriate organisms such as kinetoplastids and mycetozoan; they are not a united group but occupy different phylogenetic positions. Combining with the recent cytological findings of mitochondria-like organelles in them, we think that though some of them (e.g. diplomonads, as represented

  6. Understanding the Code: keeping accurate records.

    Science.gov (United States)

    Griffith, Richard

    2015-10-01

    In his continuing series looking at the legal and professional implications of the Nursing and Midwifery Council's revised Code of Conduct, Richard Griffith discusses the elements of accurate record keeping under Standard 10 of the Code. This article considers the importance of accurate record keeping for the safety of patients and protection of district nurses. The legal implications of records are explained along with how district nurses should write records to ensure these legal requirements are met.

  7. Phylogenetic constraints in key functional traits behind species' climate niches

    DEFF Research Database (Denmark)

    Kellermann, Vanessa; Loeschcke, Volker; Hoffmann, Ary A;

    2012-01-01

    adapted to similar environments or alternatively phylogenetic inertia. For desiccation resistance, weak phylogenetic inertia was detected; ancestral trait reconstruction, however, revealed a deep divergence that could be traced back to the genus level. Despite drosophilids’ high evolutionary potential......) for 92–95 Drosophila species and assessed their importance for geographic distributions, while controlling for acclimation, phylogeny, and spatial autocorrelation. Employing an array of phylogenetic analyses, we documented moderate-to-strong phylogenetic signal in both desiccation and cold resistance....... Desiccation and cold resistance were clearly linked to species distributions because significant associations between traits and climatic variables persisted even after controlling for phylogeny. We used different methods to untangle whether phylogenetic signal reflected phylogenetically related species...

  8. A Novel Vehicle Classification Using Embedded Strain Gauge Sensors

    Directory of Open Access Journals (Sweden)

    Qi Wang

    2008-11-01

    Full Text Available Abstract: This paper presents a new vehicle classification and develops a traffic monitoring detector to provide reliable vehicle classification to aid traffic management systems. The basic principle of this approach is based on measuring the dynamic strain caused by vehicles across pavement to obtain the corresponding vehicle parameters – wheelbase and number of axles – to then accurately classify the vehicle. A system prototype with five embedded strain sensors was developed to validate the accuracy and effectiveness of the classification method. According to the special arrangement of the sensors and the different time a vehicle arrived at the sensors one can estimate the vehicle’s speed accurately, corresponding to the estimated vehicle wheelbase and number of axles. Because of measurement errors and vehicle characteristics, there is a lot of overlap between vehicle wheelbase patterns. Therefore, directly setting up a fixed threshold for vehicle classification often leads to low-accuracy results. Using the machine learning pattern recognition method to deal with this problem is believed as one of the most effective tools. In this study, support vector machines (SVMs were used to integrate the classification features extracted from the strain sensors to automatically classify vehicles into five types, ranging from small vehicles to combination trucks, along the lines of the Federal Highway Administration vehicle classification guide. Test bench and field experiments will be introduced in this paper. Two support vector machines classification algorithms (one-against-all, one-against-one are used to classify single sensor data and multiple sensor combination data. Comparison of the two classification method results shows that the classification accuracy is very close using single data or multiple data. Our results indicate that using multiclass SVM-based fusion multiple sensor data significantly improves

  9. Automatic selection of reference taxa for protein-protein interaction prediction with phylogenetic profiling

    DEFF Research Database (Denmark)

    Simonsen, Martin; Maetschke, S.R.; Ragan, M.A.

    2012-01-01

    Motivation: Phylogenetic profiling methods can achieve good accuracy in predicting protein–protein interactions, especially in prokaryotes. Recent studies have shown that the choice of reference taxa (RT) is critical for accurate prediction, but with more than 2500 fully sequenced taxa publicly......: We present three novel methods for automating the selection of RT, using machine learning based on known protein–protein interaction networks. One of these methods in particular, Tree-Based Search, yields greatly improved prediction accuracies. We further show that different methods for constituting...

  10. Association of Virulence Genotype with Phylogenetic Background in Comparison to Different Seropathotypes of Shiga Toxin-Producing Escherichia coli Isolates

    Science.gov (United States)

    Girardeau, Jean Pierre; Dalmasso, Alessandra; Bertin, Yolande; Ducrot, Christian; Bord, Séverine; Livrelli, Valérie; Vernozy-Rozand, Christine; Martin, Christine

    2005-01-01

    The distribution of virulent factors (VFs) in 287 Shiga toxin-producing Escherichia coli (STEC) strains that were classified according to Karmali et al. into five seropathotypes (M. A. Karmali, M. Mascarenhas, S. Shen, K. Ziebell, S. Johnson, R. Reid-Smith, J. Isaac-Renton, C. Clark, K. Rahn, and J. B. Kaper, J. Clin. Microbiol. 41:4930-4940, 2003) was investigated. The associations of VFs with phylogenetic background were assessed among the strains in comparison with the different seropathotypes. The phylogenetic analysis showed that STEC strains segregated mainly in phylogenetic group B1 (70%) and revealed the substantial prevalence (19%) of STEC belonging to phylogenetic group A (designated STEC-A). The presence of virulent clonal groups in seropathotypes that are associated with disease and their absence from seropathotypes that are not associated with disease support the concept of seropathotype classification. Although certain VFs (eae, stx2-EDL933, stx2-vha, and stx2-vhb) were concentrated in seropathotypes associated with disease, others (astA, HPI, stx1c, and stx2-NV206) were concentrated in seropathotypes that are not associated with disease. Taken together with the observation that the STEC-A group was exclusively composed of strains lacking eae recovered from seropathotypes that are not associated with disease, the “atypical” virulence pattern suggests that STEC-A strains comprise a distinct category of STEC strains. A practical benefit of our phylogenetic analysis of STEC strains is that phylogenetic group A status appears to be highly predictive of “nonvirulent” seropathotypes. PMID:16333104

  11. A common tendency for phylogenetic overdispersion in mammalian assemblages

    OpenAIRE

    Cooper, Natalie; RODRIGUEZ, JESUS; Purvis, Andy

    2008-01-01

    PUBLISHED Competition has long been proposed as an important force in structuring mammalian communities. Although early work recognised that competition has a phylogenetic dimension, only with recent increases in the availability of phylogenies have true phylogenetic investigations of mammalian community structure become possible. We test whether the phylogenetic structure of 142 assemblages from three mammalian clades (New World monkeys, North American ground squirrels and Australasian po...

  12. Best Practices for Data Sharing in Phylogenetic Research

    Science.gov (United States)

    Cranston, Karen; Harmon, Luke J.; O'Leary, Maureen A.; Lisle, Curtis

    2014-01-01

    As phylogenetic data becomes increasingly available, along with associated data on species’ genomes, traits, and geographic distributions, the need to ensure data availability and reuse become more and more acute. In this paper, we provide ten “simple rules” that we view as best practices for data sharing in phylogenetic research. These rules will help lead towards a future phylogenetics where data can easily be archived, shared, reused, and repurposed across a wide variety of projects. PMID:24987572

  13. Interobserver agreement of the old and the newly proposed ILAE epilepsy classification in children

    NARCIS (Netherlands)

    van Campen, Jolien S.; Jansen, Floor E.; Brouwer, Oebele F.; Nicolai, Joost; Braun, Kees P. J.

    Purpose Accurate classification of epileptic seizures, epilepsies, and epilepsy syndromes is mandatory in both clinical practice and epilepsy research. In 2010, the International League Against Epilepsy (ILAE) proposed a new classification scheme. The aim of this study is to determine whether

  14. Inter-rater reliability of the EPUAP pressure ulcer classification system using photographs.

    NARCIS (Netherlands)

    Defloor, T.; Schoonhoven, L.

    2004-01-01

    BACKGROUND: Many classification systems for grading pressure ulcers are discussed in the literature. Correct identification and classification of a pressure ulcer is important for accurate reporting of the magnitude of the problem, and for timely prevention. The reliability of pressure ulcer

  15. Segmentation Assisted Food Classification for Dietary Assessment.

    Science.gov (United States)

    Zhu, Fengqing; Bosch, Marc; Schap, Tusarebecca; Khanna, Nitin; Ebert, David S; Boushey, Carol J; Delp, Edward J

    2011-01-24

    Accurate methods and tools to assess food and nutrient intake are essential for the association between diet and health. Preliminary studies have indicated that the use of a mobile device with a built-in camera to obtain images of the food consumed may provide a less burdensome and more accurate method for dietary assessment. We are developing methods to identify food items using a single image acquired from the mobile device. Our goal is to automatically determine the regions in an image where a particular food is located (segmentation) and correctly identify the food type based on its features (classification or food labeling). Images of foods are segmented using Normalized Cuts based on intensity and color. Color and texture features are extracted from each segmented food region. Classification decisions for each segmented region are made using support vector machine methods. The segmentation of each food region is refined based on feedback from the output of classifier to provide more accurate estimation of the quantity of food consumed.

  16. Applications of phylogenetics to solve practical problems in insect conservation.

    Science.gov (United States)

    Buckley, Thomas R

    2016-12-01

    Phylogenetic approaches have much promise for the setting of conservation priorities and resource allocation. There has been significant development of analytical methods for the measurement of phylogenetic diversity within and among ecological communities as a way of setting conservation priorities. Application of these tools to insects has been low as has been the uptake by conservation managers. A critical reason for the lack of uptake includes the scarcity of detailed phylogenetic and species distribution data from much of insect diversity. Environmental DNA technologies offer a means for the high throughout collection of phylogenetic data across landscapes for conservation planning.

  17. Accelerating metagenomic read classification on CUDA-enabled GPUs.

    Science.gov (United States)

    Kobus, Robin; Hundt, Christian; Müller, André; Schmidt, Bertil

    2017-01-03

    Metagenomic sequencing studies are becoming increasingly popular with prominent examples including the sequencing of human microbiomes and diverse environments. A fundamental computational problem in this context is read classification; i.e. the assignment of each read to a taxonomic label. Due to the large number of reads produced by modern high-throughput sequencing technologies and the rapidly increasing number of available reference genomes software tools for fast and accurate metagenomic read classification are urgently needed. We present cuCLARK, a read-level classifier for CUDA-enabled GPUs, based on the fast and accurate classification of metagenomic sequences using reduced k-mers (CLARK) method. Using the processing power of a single Titan X GPU, cuCLARK can reach classification speeds of up to 50 million reads per minute. Corresponding speedups for species- (genus-)level classification range between 3.2 and 6.6 (3.7 and 6.4) compared to multi-threaded CLARK executed on a 16-core Xeon CPU workstation. cuCLARK can perform metagenomic read classification at superior speeds on CUDA-enabled GPUs. It is free software licensed under GPL and can be downloaded at https://github.com/funatiq/cuclark free of charge.

  18. GPCRTree: online hierarchical classification of GPCR function

    Directory of Open Access Journals (Sweden)

    Timmis Jon

    2008-08-01

    Full Text Available Abstract Background G protein-coupled receptors (GPCRs play important physiological roles transducing extracellular signals into intracellular responses. Approximately 50% of all marketed drugs target a GPCR. There remains considerable interest in effectively predicting the function of a GPCR from its primary sequence. Findings Using techniques drawn from data mining and proteochemometrics, an alignment-free approach to GPCR classification has been devised. It uses a simple representation of a protein's physical properties. GPCRTree, a publicly-available internet server, implements an algorithm that classifies GPCRs at the class, sub-family and sub-subfamily level. Conclusion A selective top-down classifier was developed which assigns sequences within a GPCR hierarchy. Compared to other publicly available GPCR prediction servers, GPCRTree is considerably more accurate at every level of classification. The server has been available online since March 2008 at URL: http://igrid-ext.cryst.bbk.ac.uk/gpcrtree/.

  19. Search techniques in intelligent classification systems

    CERN Document Server

    Savchenko, Andrey V

    2016-01-01

    A unified methodology for categorizing various complex objects is presented in this book. Through probability theory, novel asymptotically minimax criteria suitable for practical applications in imaging and data analysis are examined including the special cases such as the Jensen-Shannon divergence and the probabilistic neural network. An optimal approximate nearest neighbor search algorithm, which allows faster classification of databases is featured. Rough set theory, sequential analysis and granular computing are used to improve performance of the hierarchical classifiers. Practical examples in face identification (including deep neural networks), isolated commands recognition in voice control system and classification of visemes captured by the Kinect depth camera are included. This approach creates fast and accurate search procedures by using exact probability densities of applied dissimilarity measures. This book can be used as a guide for independent study and as supplementary material for a technicall...

  20. Prediction and classification of respiratory motion

    CERN Document Server

    Lee, Suk Jin

    2014-01-01

    This book describes recent radiotherapy technologies including tools for measuring target position during radiotherapy and tracking-based delivery systems. This book presents a customized prediction of respiratory motion with clustering from multiple patient interactions. The proposed method contributes to the improvement of patient treatments by considering breathing pattern for the accurate dose calculation in radiotherapy systems. Real-time tumor-tracking, where the prediction of irregularities becomes relevant, has yet to be clinically established. The statistical quantitative modeling for irregular breathing classification, in which commercial respiration traces are retrospectively categorized into several classes based on breathing pattern are discussed as well. The proposed statistical classification may provide clinical advantages to adjust the dose rate before and during the external beam radiotherapy for minimizing the safety margin. In the first chapter following the Introduction  to this book, we...

  1. The Shapley Value of Phylogenetic Trees

    CERN Document Server

    Haake, Claus-Jochen; Su, Francis Edward

    2007-01-01

    Every weighted tree corresponds naturally to a cooperative game that we call a "tree game"; it assigns to each subset of leaves the sum of the weights of the minimal subtree spanned by those leaves. In the context of phylogenetic trees, the leaves are species and this assignment captures the diversity present in the coalition of species considered. We consider the Shapley value of tree games and suggest a biological interpretation. We determine the linear transformation M that shows the dependence of the Shapley value on the edge weights of the tree, and we also compute a null space basis of M. Both depend on the "split counts" of the tree. Finally, we characterize the Shapley value on tree games by four axioms, a counterpart to Shapley's original theorem on the larger class of cooperative games.

  2. Tanglegrams: a Reduction Tool for Mathematical Phylogenetics.

    Science.gov (United States)

    Matsen, Frederick; Billey, Sara; Kas, Arnold; Konvalinka, Matjaz

    2016-10-03

    Many discrete mathematics problems in phylogenetics are defined in terms of the relative labeling of pairsof leaf-labeled trees. These relative labelings are naturally formalized as tanglegrams, which have previously been an object of study in coevolutionary analysis. Although there has been considerable work on planar drawings of tanglegrams, they have not been fully explored as combinatorial objects until recently. In this paper, we describe how many discrete mathematical questions on trees "factor" through a problem on tanglegrams, and how understanding that factoring can simplify analysis. Depending on the problem, it may be useful to consider a unordered version of tanglegrams, and/or their unrooted counterparts. For all of these definitions, we show how the isomorphism types of tanglegrams can be understood in terms of double cosets of the symmetric group, and we investigate their automorphisms. Understanding tanglegrams better will isolate the distinct problems on leaf-labeled pairs of trees and reveal natural symmetries of spaces associated with such problems.

  3. The rapidly changing landscape of insect phylogenetics.

    Science.gov (United States)

    Maddison, David R

    2016-12-01

    Insect phylogenetics is being profoundly changed by many innovations. Although rapid developments in genomics have center stage, key progress has been made in phenomics, field and museum science, digital databases and pipelines, analytical tools, and the culture of science. The importance of these methodological and cultural changes to the pace of inference of the hexapod Tree of Life is discussed. The innovations have the potential, when synthesized and mobilized in ways as yet unforeseen, to shine light on the million or more clades in insects, and infer their composition with confidence. There are many challenges to overcome before insects can enter the 'phylocognisant age', but because of the promise of genomics, phenomics, and informatics, that is now an imaginable future.

  4. Zika Virus: Emergence, Phylogenetics, Challenges, and Opportunities.

    Science.gov (United States)

    Rajah, Maaran M; Pardy, Ryan D; Condotta, Stephanie A; Richer, Martin J; Sagan, Selena M

    2016-11-11

    Zika virus (ZIKV) is an emerging arthropod-borne pathogen that has recently gained notoriety due to its rapid and ongoing geographic expansion and its novel association with neurological complications. Reports of ZIKV-associated Guillain-Barré syndrome as well as fetal microcephaly place emphasis on the need to develop preventative measures and therapeutics to combat ZIKV infection. Thus, it is imperative that models to study ZIKV replication and pathogenesis and the immune response are developed in conjunction with integrated vector control strategies to mount an efficient response to the pandemic. This paper summarizes the current state of knowledge on ZIKV, including the clinical features, phylogenetic analyses, pathogenesis, and the immune response to infection. Potential challenges in developing diagnostic tools, treatment, and prevention strategies are also discussed.

  5. Taxonomic review and phylogenetic analysis of Enchodontoidei.

    Science.gov (United States)

    Silva, Hilda M A; Gallo, Valéria

    2011-06-01

    Enchodontoidei are extinct marine teleost fishes with a long temporal range and a wide geographic distribution. As there has been no comprehensive phylogenetic study of this taxon, we performed a parsimony analysis using a data matrix with 87 characters, 31 terminal taxa for ingroup, and three taxa for outgroup. The analysis produced 93 equally parsimonious trees (L = 437 steps; CI = 0. 24; RI = 0. 49). The topology of the majority rule consensus tree was: (Sardinioides + Hemisaurida + (Nardorex + (Atolvorator + (Protostomias + Yabrudichthys ) + (Apateopholis + (Serrilepis + (Halec + Phylactocephalus ) + (Cimolichthys + (Prionolepis + ( (Eurypholis + Saurorhamphus ) + (Enchodus + (Paleolycus + Parenchodus ))))))) + ( (Ichthyotringa + Apateodus ) + (Rharbichthys + (Trachinocephalus + ( (Apuliadercetis + Brazilodercetis ) + (Benthesikyme + (Cyranichthys + Robertichthys ) + (Dercetis + Ophidercetis )) + (Caudadercetis + (Pelargorhynchus + (Nardodercetis + (Rhynchodercetis + (Dercetoides + Hastichthys )))))). The group Enchodontoidei is not monophyletic. Dercetidae form a clade supported by the presence of very reduced neural spines and possess a new composition. Enchodontidae are monophyletic by the presence of middorsal scutes, and Rharbichthys was excluded. Halecidae possess a new composition, with the exclusion of Hemisaurida. This taxon and Nardorex are Aulopiformes incertae sedis.

  6. Mitochondrial genome organization and vertebrate phylogenetics

    Directory of Open Access Journals (Sweden)

    Pereira Sérgio Luiz

    2000-01-01

    Full Text Available With the advent of DNA sequencing techniques the organization of the vertebrate mitochondrial genome shows variation between higher taxonomic levels. The most conserved gene order is found in placental mammals, turtles, fishes, some lizards and Xenopus. Birds, other species of lizards, crocodilians, marsupial mammals, snakes, tuatara, lamprey, and some other amphibians and one species of fish have gene orders that are less conserved. The most probable mechanism for new gene rearrangements seems to be tandem duplication and multiple deletion events, always associated with tRNA sequences. Some new rearrangements seem to be typical of monophyletic groups and the use of data from these groups may be useful for answering phylogenetic questions involving vertebrate higher taxonomic levels. Other features such as the secondary structure of tRNA, and the start and stop codons of protein-coding genes may also be useful in comparisons of vertebrate mitochondrial genomes.

  7. Inferring Phylogenetic Networks from Gene Order Data

    Directory of Open Access Journals (Sweden)

    Alexey Anatolievich Morozov

    2013-01-01

    Full Text Available Existing algorithms allow us to infer phylogenetic networks from sequences (DNA, protein or binary, sets of trees, and distance matrices, but there are no methods to build them using the gene order data as an input. Here we describe several methods to build split networks from the gene order data, perform simulation studies, and use our methods for analyzing and interpreting different real gene order datasets. All proposed methods are based on intermediate data, which can be generated from genome structures under study and used as an input for network construction algorithms. Three intermediates are used: set of jackknife trees, distance matrix, and binary encoding. According to simulations and case studies, the best intermediates are jackknife trees and distance matrix (when used with Neighbor-Net algorithm. Binary encoding can also be useful, but only when the methods mentioned above cannot be used.

  8. Estimating diversification rates from phylogenetic information.

    Science.gov (United States)

    Ricklefs, Robert E

    2007-11-01

    Patterns of species richness reflect the balance between speciation and extinction over the evolutionary history of life. These processes are influenced by the size and geographical complexity of regions, conditions of the environment, and attributes of individuals and species. Diversity within clades also depends on age and thus the time available for accumulating species. Estimating rates of diversification is key to understanding how these factors have shaped patterns of species richness. Several approaches to calculating both relative and absolute rates of speciation and extinction within clades are based on phylogenetic reconstructions of evolutionary relationships. As the size and quality of phylogenies increases, these approaches will find broader application. However, phylogeny reconstruction fosters a perceptual bias of continual increase in species richness, and the analysis of primarily large clades produces a data selection bias. Recognizing these biases will encourage the development of more realistic models of diversification and the regulation of species richness.

  9. S1 gene-based phylogeny of infectious bronchitis virus: An attempt to harmonize virus classification.

    Science.gov (United States)

    Valastro, Viviana; Holmes, Edward C; Britton, Paul; Fusaro, Alice; Jackwood, Mark W; Cattoli, Giovanni; Monne, Isabella

    2016-04-01

    Infectious bronchitis virus (IBV) is the causative agent of a highly contagious disease that results in severe economic losses to the global poultry industry. The virus exists in a wide variety of genetically distinct viral types, and both phylogenetic analysis and measures of pairwise similarity among nucleotide or amino acid sequences have been used to classify IBV strains. However, there is currently no consensus on the method by which IBV sequences should be compared, and heterogeneous genetic group designations that are inconsistent with phylogenetic history have been adopted, leading to the confusing coexistence of multiple genotyping schemes. Herein, we propose a simple and repeatable phylogeny-based classification system combined with an unambiguous and rationale lineage nomenclature for the assignment of IBV strains. By using complete nucleotide sequences of the S1 gene we determined the phylogenetic structure of IBV, which in turn allowed us to define 6 genotypes that together comprise 32 distinct viral lineages and a number of inter-lineage recombinants. Because of extensive rate variation among IBVs, we suggest that the inference of phylogenetic relationships alone represents a more appropriate criterion for sequence classification than pairwise sequence comparisons. The adoption of an internationally accepted viral nomenclature is crucial for future studies of IBV epidemiology and evolution, and the classification scheme presented here can be updated and revised novel S1 sequences should become available. Copyright © 2016 Elsevier B.V. All rights reserved.

  10. Phylogenetic insights into Andean plant diversification

    Directory of Open Access Journals (Sweden)

    Federico eLuebert

    2014-06-01

    Full Text Available Andean orogeny is considered as one of the most important events for the developmentof current plant diversity in South America. We compare available phylogenetic studies anddivergence time estimates for plant lineages that may have diversified in response to Andeanorogeny. The influence of the Andes on plant diversification is separated into four major groups:The Andes as source of new high-elevation habitats, as a vicariant barrier, as a North-Southcorridor and as generator of new environmental conditions outside the Andes. Biogeographicalrelationships between the Andes and other regions are also considered. Divergence timeestimates indicate that high-elevation lineages originated and diversified during or after the majorphases of Andean uplift (Mid-Miocene to Pliocene, although there are some exceptions. Asexpected, Andean mid-elevation lineages tend to be older than high-elevation groups. Mostclades with disjunct distribution on both sides of the Andes diverged during Andean uplift.Inner-Andean clades also tend to have divergence time during or after Andean uplift. This isinterpreted as evidence of vicariance. Dispersal along the Andes has been shown to occur ineither direction, mostly dated after the Andean uplift. Divergence time estimates of plant groupsoutside the Andes encompass a wider range of ages, indicating that the Andes may not benecessarily the cause of these diversifications. The Andes are biogeographically related to allneighbouring areas, especially Central America, with floristic interchanges in both directionssince Early Miocene times. Direct biogeographical relationships between the Andes and otherdisjunct regions have also been shown in phylogenetic studies, especially with the easternBrazilian highlands and North America. The history of the Andean flora is complex and plantdiversification has been driven by a variety of processes, including environmental change,adaptation, and biotic interactions

  11. Platelet-rich plasma: the PAW classification system.

    Science.gov (United States)

    DeLong, Jeffrey M; Russell, Ryan P; Mazzocca, Augustus D

    2012-07-01

    Platelet-rich plasma (PRP) has been the subject of hundreds of publications in recent years. Reports of its effects in tissue, both positive and negative, have generated great interest in the orthopaedic community. Protocols for PRP preparation vary widely between authors and are often not well documented in the literature, making results difficult to compare or replicate. A classification system is needed to more accurately compare protocols and results and effectively group studies together for meta-analysis. Although some classification systems have been proposed, no single system takes into account the multitude of variables that determine the efficacy of PRP. In this article we propose a simple method for organizing and comparing results in the literature. The PAW classification system is based on 3 components: (1) the absolute number of Platelets, (2) the manner in which platelet Activation occurs, and (3) the presence or absence of White cells. By analyzing these 3 variables, we are able to accurately compare publications.

  12. Application of fuzzy classification in modern primary dental care

    Directory of Open Access Journals (Sweden)

    Yauheni Veryha

    2005-03-01

    Full Text Available This paper describes a framework for implementing fuzzy classifications in primary dental care services. Dental practices aim to provide the highest quality services for their patients. To achieve this, it is important that dentists are able to obtain patients' opinions about their experiences in the dental practice and are able to accurately evaluate this. We propose the use of fuzzy classification to combine various assessment criteria into one general measure to assess patients' satisfaction with primary dental care services. The proposed framework can be used in conventional dental practice information systems and easily integrated with those already used. The benefits of using the proposed fuzzy classification approach include more flexible and accurate analysis of patients' feedback, combining verbal and numeric data. To confirm our theory, a prototype was developed based on the Microsoft TM SQL Server database management system for two criteria used in dental practices, namely making an appointment with a dentist and waiting time for dental care services.

  13. Convolutional Neural Networks for patient-specific ECG classification.

    Science.gov (United States)

    Kiranyaz, Serkan; Ince, Turker; Hamila, Ridha; Gabbouj, Moncef

    2015-01-01

    We propose a fast and accurate patient-specific electrocardiogram (ECG) classification and monitoring system using an adaptive implementation of 1D Convolutional Neural Networks (CNNs) that can fuse feature extraction and classification into a unified learner. In this way, a dedicated CNN will be trained for each patient by using relatively small common and patient-specific training data and thus it can also be used to classify long ECG records such as Holter registers in a fast and accurate manner. Alternatively, such a solution can conveniently be used for real-time ECG monitoring and early alert system on a light-weight wearable device. The experimental results demonstrate that the proposed system achieves a superior classification performance for the detection of ventricular ectopic beats (VEB) and supraventricular ectopic beats (SVEB).

  14. An analysis of network traffic classification for botnet detection

    DEFF Research Database (Denmark)

    Stevanovic, Matija; Pedersen, Jens Myrup

    2015-01-01

    Botnets represent one of the most serious threats to the Internet security today. This paper explores how can network traffic classification be used for accurate and efficient identification of botnet network activity at local and enterprise networks. The paper examines the effectiveness of detec......Botnets represent one of the most serious threats to the Internet security today. This paper explores how can network traffic classification be used for accurate and efficient identification of botnet network activity at local and enterprise networks. The paper examines the effectiveness...... of detecting botnet network traffic using three methods that target protocols widely considered as the main carriers of botnet Command and Control (C&C) and attack traffic, i.e. TCP, UDP and DNS. We propose three traffic classification methods based on capable Random Forests classifier. The proposed methods...

  15. Dilated Chi-Square : a novel interestingness measure to build accurate and compact decion list

    OpenAIRE

    Lan, Yu; Janssens, Davy; Chen, Guoqing; Wets, Geert

    2005-01-01

    Associative classification has aroused significant attention in recent years. This paper proposed a novel interestingness measure, named dilated chi-square, to statistically reveal the interdependence between the antecedents and the consequent of classificaton rules. Using dilated chi-square, instead of confidence, as the primary ranking criterion for rules under the framework of popular CBA algorithm, the adapted algorithm presented in this paper can empirically generate more accurate and mu...

  16. Supernova Photometric Classification Challenge

    CERN Document Server

    Kessler, Richard; Jha, Saurabh; Kuhlmann, Stephen

    2010-01-01

    We have publicly released a blinded mix of simulated SNe, with types (Ia, Ib, Ic, II) selected in proportion to their expected rate. The simulation is realized in the griz filters of the Dark Energy Survey (DES) with realistic observing conditions (sky noise, point spread function and atmospheric transparency) based on years of recorded conditions at the DES site. Simulations of non-Ia type SNe are based on spectroscopically confirmed light curves that include unpublished non-Ia samples donated from the Carnegie Supernova Project (CSP), the Supernova Legacy Survey (SNLS), and the Sloan Digital Sky Survey-II (SDSS-II). We challenge scientists to run their classification algorithms and report a type for each SN. A spectroscopically confirmed subset is provided for training. The goals of this challenge are to (1) learn the relative strengths and weaknesses of the different classification algorithms, (2) use the results to improve classification algorithms, and (3) understand what spectroscopically confirmed sub-...

  17. Classification in Medical Imaging

    DEFF Research Database (Denmark)

    Chen, Chen

    Classification is extensively used in the context of medical image analysis for the purpose of diagnosis or prognosis. In order to classify image content correctly, one needs to extract efficient features with discriminative properties and build classifiers based on these features. In addition......, a good metric is required to measure distance or similarity between feature points so that the classification becomes feasible. Furthermore, in order to build a successful classifier, one needs to deeply understand how classifiers work. This thesis focuses on these three aspects of classification...... to segment breast tissue and pectoral muscle area from the background in mammogram. The second focus is the choices of metric and its influence to the feasibility of a classifier, especially on k-nearest neighbors (k-NN) algorithm, with medical applications on breast cancer prediction and calcification...

  18. Classification of hand eczema

    DEFF Research Database (Denmark)

    Agner, T; Aalto-Korte, K; Andersen, K E;

    2015-01-01

    BACKGROUND: Classification of hand eczema (HE) is mandatory in epidemiological and clinical studies, and also important in clinical work. OBJECTIVES: The aim was to test a recently proposed classification system of HE in clinical practice in a prospective multicentre study. METHODS: Patients were...... HE, protein contact dermatitis/contact urticaria, hyperkeratotic endogenous eczema and vesicular endogenous eczema, respectively. An additional diagnosis was given if symptoms indicated that factors additional to the main diagnosis were of importance for the disease. RESULTS: Four hundred and twenty......%) could not be classified. 38% had one additional diagnosis and 26% had two or more additional diagnoses. Eczema on feet was found in 30% of the patients, statistically significantly more frequently associated with hyperkeratotic and vesicular endogenous eczema. CONCLUSION: We find that the classification...

  19. Acoustic classification of dwellings

    DEFF Research Database (Denmark)

    Berardi, Umberto; Rasmussen, Birgit

    2014-01-01

    Schemes for the classification of dwellings according to different building performances have been proposed in the last years worldwide. The general idea behind these schemes relates to the positive impact a higher label, and thus a better performance, should have. In particular, focusing on soun...... exchanging experiences about constructions fulfilling different classes, reducing trade barriers, and finally increasing the sound insulation of dwellings.......Schemes for the classification of dwellings according to different building performances have been proposed in the last years worldwide. The general idea behind these schemes relates to the positive impact a higher label, and thus a better performance, should have. In particular, focusing on sound...... insulation performance, national schemes for sound classification of dwellings have been developed in several European countries. These schemes define acoustic classes according to different levels of sound insulation. Due to the lack of coordination among countries, a significant diversity in terms...

  20. Classification problem in CBIR

    Directory of Open Access Journals (Sweden)

    Tatiana Jaworska

    2013-04-01

    Full Text Available At present a great deal of research is being done in different aspects of Content-Based Im-age Retrieval (CBIR. Image classification is one of the most important tasks in image re-trieval that must be dealt with. The primary issue we have addressed is: how can the fuzzy set theory be used to handle crisp image data. We propose fuzzy rule-based classification of image objects. To achieve this goal we have built fuzzy rule-based classifiers for crisp data. In this paper we present the results of fuzzy rule-based classification in our CBIR. Further-more, these results are used to construct a search engine taking into account data mining.

  1. Cellular image classification

    CERN Document Server

    Xu, Xiang; Lin, Feng

    2017-01-01

    This book introduces new techniques for cellular image feature extraction, pattern recognition and classification. The authors use the antinuclear antibodies (ANAs) in patient serum as the subjects and the Indirect Immunofluorescence (IIF) technique as the imaging protocol to illustrate the applications of the described methods. Throughout the book, the authors provide evaluations for the proposed methods on two publicly available human epithelial (HEp-2) cell datasets: ICPR2012 dataset from the ICPR'12 HEp-2 cell classification contest and ICIP2013 training dataset from the ICIP'13 Competition on cells classification by fluorescent image analysis. First, the reading of imaging results is significantly influenced by one’s qualification and reading systems, causing high intra- and inter-laboratory variance. The authors present a low-order LP21 fiber mode for optical single cell manipulation and imaging staining patterns of HEp-2 cells. A focused four-lobed mode distribution is stable and effective in optical...

  2. The paradox of atheoretical classification

    DEFF Research Database (Denmark)

    Hjørland, Birger

    2016-01-01

    A distinction can be made between “artificial classifications” and “natural classifications,” where artificial classifications may adequately serve some limited purposes, but natural classifications are overall most fruitful by allowing inference and thus many different purposes. There is strong...... support for the view that a natural classification should be based on a theory (and, of course, that the most fruitful theory provides the most fruitful classification). Nevertheless, atheoretical (or “descriptive”) classifications are often produced. Paradoxically, atheoretical classifications may...... be very successful. The best example of a successful “atheoretical” classification is probably the prestigious Diagnostic and Statistical Manual of Mental Disorders (DSM) since its third edition from 1980. Based on such successes one may ask: Should the claim that classifications ideally are natural...

  3. Information gathering for CLP classification.

    Science.gov (United States)

    Marcello, Ida; Giordano, Felice; Costamagna, Francesca Marina

    2011-01-01

    Regulation 1272/2008 includes provisions for two types of classification: harmonised classification and self-classification. The harmonised classification of substances is decided at Community level and a list of harmonised classifications is included in the Annex VI of the classification, labelling and packaging Regulation (CLP). If a chemical substance is not included in the harmonised classification list it must be self-classified, based on available information, according to the requirements of Annex I of the CLP Regulation. CLP appoints that the harmonised classification will be performed for carcinogenic, mutagenic or toxic to reproduction substances (CMR substances) and for respiratory sensitisers category 1 and for other hazard classes on a case-by-case basis. The first step of classification is the gathering of available and relevant information. This paper presents the procedure for gathering information and to obtain data. The data quality is also discussed.

  4. Information gathering for CLP classification

    Directory of Open Access Journals (Sweden)

    Ida Marcello

    2011-01-01

    Full Text Available Regulation 1272/2008 includes provisions for two types of classification: harmonised classification and self-classification. The harmonised classification of substances is decided at Community level and a list of harmonised classifications is included in the Annex VI of the classification, labelling and packaging Regulation (CLP. If a chemical substance is not included in the harmonised classification list it must be self-classified, based on available information, according to the requirements of Annex I of the CLP Regulation. CLP appoints that the harmonised classification will be performed for carcinogenic, mutagenic or toxic to reproduction substances (CMR substances and for respiratory sensitisers category 1 and for other hazard classes on a case-by-case basis. The first step of classification is the gathering of available and relevant information. This paper presents the procedure for gathering information and to obtain data. The data quality is also discussed.

  5. The complete chloroplast genome sequences of five Epimedium species: lights into phylogenetic and taxonomic analyses

    Directory of Open Access Journals (Sweden)

    Yanjun eZhang

    2016-03-01

    Full Text Available Epimedium L. is a phylogenetically and economically important genus in the family Berberidaceae. We here sequenced the complete chloroplast (cp genomes of four Epimedium species using Illumina sequencing technology via a combination of de novo and reference-guided assembly, which was also the first comprehensive cp genome analysis on Epimedium combining the cp genome sequence of E. koreanum previously reported. The five Epimedium cp genomes exhibited typical quadripartite and circular structure that was rather conserved in genomic structure and the synteny of gene order. However, these cp genomes presented obvious variations at the boundaries of the four regions because of the expansion and contraction of the inverted repeat (IR region and the single-copy (SC boundary regions. The trnQ-UUG duplication occurred in the five Epimedium cp genomes, which was not found in the other basal eudicotyledons. The rapidly evolving cp genome regions were detected among the five cp genomes, as well as the difference of simple sequence repeats (SSR and repeat sequence were identified. Phylogenetic relationships among the five Epimedium species based on their cp genomes showed accordance with the updated system of the genus on the whole, but reminded that the evolutionary relationships and the divisions of the genus need further investigation applying more evidences. The availability of these cp genomes provided valuable genetic information for accurately identifying species, taxonomy and phylogenetic resolution and evolution of Epimedium, and assist in exploration and utilization of Epimedium plants.

  6. Ascospore morphology is a poor predictor of the phylogenetic relationships of Neurospora and Gelasinospora.

    Science.gov (United States)

    Dettman, J R; Harbinski, F M; Taylor, J W

    2001-10-01

    The genera Neurospora and Gelasinospora are conventionally distinguished by differences in ascospore ornamentation, with elevated longitudinal ridges (ribs) separated by depressed grooves (veins) in Neurospora and spherical or oval indentations (pits) in Gelasinospora. The phylogenetic relationships of representatives of 12 Neurospora and 4 Gelasinospora species were assessed with the DNA sequences of four nuclear genes. Within the genus Neurospora, the 5 outbreeding conidiating species form a monophyletic group with N. discreta as the most divergent, and 4 of the homothallic species form a monophyletic group. In combined analysis, each of the conventionally defined Gelasinospora species was more closely related to a Neurospora species than to another Gelasinospora species. Evidently, the Neurospora and Gelasinospora species included in this study do not represent two clearly resolved monophyletic sister genera, but instead represent a polyphyletic group of taxa with close phylogenetic relationships and significant morphological similarities. Ascospore morphology, the character that the distinction between the genera Neurospora and Gelasinospora is based upon,was not an accurate predictor of phylogenetic relationships.

  7. Detecting taxonomic and phylogenetic signals in equid cheek teeth: towards new palaeontological and archaeological proxies

    Science.gov (United States)

    Cucchi, T.; Mohaseb, A.; Peigné, S.; Debue, K.; Orlando, L.; Mashkour, M.

    2017-04-01

    The Plio-Pleistocene evolution of Equus and the subsequent domestication of horses and donkeys remains poorly understood, due to the lack of phenotypic markers capable of tracing this evolutionary process in the palaeontological/archaeological record. Using images from 345 specimens, encompassing 15 extant taxa of equids, we quantified the occlusal enamel folding pattern in four mandibular cheek teeth with a single geometric morphometric protocol. We initially investigated the protocol accuracy by assigning each tooth to its correct anatomical position and taxonomic group. We then contrasted the phylogenetic signal present in each tooth shape with an exome-wide phylogeny from 10 extant equine species. We estimated the strength of the phylogenetic signal using a Brownian motion model of evolution with multivariate K statistic, and mapped the dental shape along the molecular phylogeny using an approach based on squared-change parsimony. We found clear evidence for the relevance of dental phenotypes to accurately discriminate all modern members of the genus Equus and capture their phylogenetic relationships. These results are valuable for both palaeontologists and zooarchaeologists exploring the spatial and temporal dynamics of the evolutionary history of the horse family, up to the latest domestication trajectories of horses and donkeys.

  8. morePhyML: improving the phylogenetic tree space exploration with PhyML 3.

    Science.gov (United States)

    Criscuolo, Alexis

    2011-12-01

    PhyML is a widely used Maximum Likelihood (ML) phylogenetic tree inference software based on a standard hill-climbing method. Starting from an initial tree, the version 3 of PhyML explores the tree space by using "Nearest Neighbor Interchange" (NNI) or "Subtree Pruning and Regrafting" (SPR) tree swapping techniques in order to find the ML phylogenetic tree. NNI-based local searches are fast but can often get trapped in local optima, whereas it is expected that the larger (but slower to cover) SPR-based neighborhoods will lead to trees with higher likelihood. Here, I verify that PhyML infers more likely trees with SPRs than with NNIs in almost all cases. However, I also show that the SPR-based local search of PhyML often does not succeed at locating the ML tree. To improve the tree space exploration, I deliver a script, named morePhyML, which allows escaping from local optima by performing character reweighting. This ML tree search strategy, named ratchet, often leads to higher likelihood estimates. Based on the analysis of a large number of amino acid and nucleotide data, I show that morePhyML allows inferring more accurate phylogenetic trees than several other recently developed ML tree inference softwares in many cases.

  9. Performance of criteria for selecting evolutionary models in phylogenetics: a comprehensive study based on simulated datasets

    Directory of Open Access Journals (Sweden)

    Luo Arong

    2010-08-01

    Full Text Available Abstract Background Explicit evolutionary models are required in maximum-likelihood and Bayesian inference, the two methods that are overwhelmingly used in phylogenetic studies of DNA sequence data. Appropriate selection of nucleotide substitution models is important because the use of incorrect models can mislead phylogenetic inference. To better understand the performance of different model-selection criteria, we used 33,600 simulated data sets to analyse the accuracy, precision, dissimilarity, and biases of the hierarchical likelihood-ratio test, Akaike information criterion, Bayesian information criterion, and decision theory. Results We demonstrate that the Bayesian information criterion and decision theory are the most appropriate model-selection criteria because of their high accuracy and precision. Our results also indicate that in some situations different models are selected by different criteria for the same dataset. Such dissimilarity was the highest between the hierarchical likelihood-ratio test and Akaike information criterion, and lowest between the Bayesian information criterion and decision theory. The hierarchical likelihood-ratio test performed poorly when the true model included a proportion of invariable sites, while the Bayesian information criterion and decision theory generally exhibited similar performance to each other. Conclusions Our results indicate that the Bayesian information criterion and decision theory should be preferred for model selection. Together with model-adequacy tests, accurate model selection will serve to improve the reliability of phylogenetic inference and related analyses.

  10. Bosniak Classification system

    DEFF Research Database (Denmark)

    Graumann, Ole; Osther, Susanne Sloth; Karstoft, Jens;

    2014-01-01

    . Purpose: To investigate the inter- and intra-observer agreement among experienced uroradiologists when categorizing complex renal cysts according to the Bosniak classification. Material and Methods: The original categories of 100 cystic renal masses were chosen as “Gold Standard” (GS), established...... to the calculated weighted κ all readers performed “very good” for both inter-observer and intra-observer variation. Most variation was seen in cysts catagorized as Bosniak II, IIF, and III. These results show that radiologists who evaluate complex renal cysts routinely may apply the Bosniak classification...

  11. Acoustic classification of dwellings

    DEFF Research Database (Denmark)

    Berardi, Umberto; Rasmussen, Birgit

    2014-01-01

    insulation performance, national schemes for sound classification of dwellings have been developed in several European countries. These schemes define acoustic classes according to different levels of sound insulation. Due to the lack of coordination among countries, a significant diversity in terms...... of descriptors, number of classes, and class intervals occurred between national schemes. However, a proposal “acoustic classification scheme for dwellings” has been developed recently in the European COST Action TU0901 with 32 member countries. This proposal has been accepted as an ISO work item. This paper...

  12. Classification of iconic images

    OpenAIRE

    Zrianina, Mariia; Kopf, Stephan

    2016-01-01

    Iconic images represent an abstract topic and use a presentation that is intuitively understood within a certain cultural context. For example, the abstract topic “global warming” may be represented by a polar bear standing alone on an ice floe. Such images are widely used in media and their automatic classification can help to identify high-level semantic concepts. This paper presents a system for the classification of iconic images. It uses a variation of the Bag of Visual Words approach wi...

  13. Classification problem in CBIR

    OpenAIRE

    Tatiana Jaworska

    2013-01-01

    At present a great deal of research is being done in different aspects of Content-Based Im-age Retrieval (CBIR). Image classification is one of the most important tasks in image re-trieval that must be dealt with. The primary issue we have addressed is: how can the fuzzy set theory be used to handle crisp image data. We propose fuzzy rule-based classification of image objects. To achieve this goal we have built fuzzy rule-based classifiers for crisp data. In this paper we present the results ...

  14. Latent classification models

    DEFF Research Database (Denmark)

    Langseth, Helge; Nielsen, Thomas Dyhre

    2005-01-01

    One of the simplest, and yet most consistently well-performing setof classifiers is the \\NB models. These models rely on twoassumptions: $(i)$ All the attributes used to describe an instanceare conditionally independent given the class of that instance,and $(ii)$ all attributes follow a specific...... parametric family ofdistributions.  In this paper we propose a new set of models forclassification in continuous domains, termed latent classificationmodels. The latent classification model can roughly be seen ascombining the \\NB model with a mixture of factor analyzers,thereby relaxing the assumptions...... classification model, and wedemonstrate empirically that the accuracy of the proposed model issignificantly higher than the accuracy of other probabilisticclassifiers....

  15. Minimum Error Entropy Classification

    CERN Document Server

    Marques de Sá, Joaquim P; Santos, Jorge M F; Alexandre, Luís A

    2013-01-01

    This book explains the minimum error entropy (MEE) concept applied to data classification machines. Theoretical results on the inner workings of the MEE concept, in its application to solving a variety of classification problems, are presented in the wider realm of risk functionals. Researchers and practitioners also find in the book a detailed presentation of practical data classifiers using MEE. These include multi‐layer perceptrons, recurrent neural networks, complexvalued neural networks, modular neural networks, and decision trees. A clustering algorithm using a MEE‐like concept is also presented. Examples, tests, evaluation experiments and comparison with similar machines using classic approaches, complement the descriptions.

  16. Constructing criticality by classification

    DEFF Research Database (Denmark)

    Machacek, Erika

    2017-01-01

    This paper explores the role of expertise, the nature of criticality, and their relationship to securitisation as mineral raw materials are classified. It works with the construction of risk along the liberal logic of security to explore how "key materials" are turned into "critical materials......, legitimizing a criticality discourse.Specifically, the paper introduces a typology delineating the inferences made by the experts from their produced recommendations in the classification of rare earth element criticality. The paper argues that the classification is a specific process of constructing risk...

  17. Undergraduate Students’ Initial Ability in Understanding Phylogenetic Tree

    Science.gov (United States)

    Sa'adah, S.; Hidayat, T.; Sudargo, Fransisca

    2017-04-01

    The Phylogenetic tree is a visual representation depicts a hypothesis about the evolutionary relationship among taxa. Evolutionary experts use this representation to evaluate the evidence for evolution. The phylogenetic tree is currently growing for many disciplines in biology. Consequently, learning about the phylogenetic tree has become an important part of biological education and an interesting area of biology education research. Skill to understanding and reasoning of the phylogenetic tree, (called tree thinking) is an important skill for biology students. However, research showed many students have difficulty in interpreting, constructing, and comparing among the phylogenetic tree, as well as experiencing a misconception in the understanding of the phylogenetic tree. Students are often not taught how to reason about evolutionary relationship depicted in the diagram. Students are also not provided with information about the underlying theory and process of phylogenetic. This study aims to investigate the initial ability of undergraduate students in understanding and reasoning of the phylogenetic tree. The research method is the descriptive method. Students are given multiple choice questions and an essay that representative by tree thinking elements. Each correct answer made percentages. Each student is also given questionnaires. The results showed that the undergraduate students’ initial ability in understanding and reasoning phylogenetic tree is low. Many students are not able to answer questions about the phylogenetic tree. Only 19 % undergraduate student who answered correctly on indicator evaluate the evolutionary relationship among taxa, 25% undergraduate student who answered correctly on indicator applying concepts of the clade, 17% undergraduate student who answered correctly on indicator determines the character evolution, and only a few undergraduate student who can construct the phylogenetic tree.

  18. Accurate tracking control in LOM application

    Institute of Scientific and Technical Information of China (English)

    2003-01-01

    The fabrication of accurate prototype from CAD model directly in short time depends on the accurate tracking control and reference trajectory planning in (Laminated Object Manufacture) LOM application. An improvement on contour accuracy is acquired by the introduction of a tracking controller and a trajectory generation policy. A model of the X-Y positioning system of LOM machine is developed as the design basis of tracking controller. The ZPETC (Zero Phase Error Tracking Controller) is used to eliminate single axis following error, thus reduce the contour error. The simulation is developed on a Maltab model based on a retrofitted LOM machine and the satisfied result is acquired.

  19. [Hard and soft classification method of multi-spectral remote sensing image based on adaptive thresholds].

    Science.gov (United States)

    Hu, Tan-Gao; Xu, Jun-Feng; Zhang, Deng-Rong; Wang, Jie; Zhang, Yu-Zhou

    2013-04-01

    Hard and soft classification techniques are the conventional methods of image classification for satellite data, but they have their own advantages and drawbacks. In order to obtain accurate classification results, we took advantages of both traditional hard classification methods (HCM) and soft classification models (SCM), and developed a new method called the hard and soft classification model (HSCM) based on adaptive threshold calculation. The authors tested the new method in land cover mapping applications. According to the results of confusion matrix, the overall accuracy of HCM, SCM, and HSCM is 71.06%, 67.86%, and 71.10%, respectively. And the kappa coefficient is 60.03%, 56.12%, and 60.07%, respectively. Therefore, the HSCM is better than HCM and SCM. Experimental results proved that the new method can obviously improve the land cover and land use classification accuracy.

  20. Ensemble polarimetric SAR image classification based on contextual sparse representation

    Science.gov (United States)

    Zhang, Lamei; Wang, Xiao; Zou, Bin; Qiao, Zhijun

    2016-05-01

    Polarimetric SAR image interpretation has become one of the most interesting topics, in which the construction of the reasonable and effective technique of image classification is of key importance. Sparse representation represents the data using the most succinct sparse atoms of the over-complete dictionary and the advantages of sparse representation also have been confirmed in the field of PolSAR classification. However, it is not perfect, like the ordinary classifier, at different aspects. So ensemble learning is introduced to improve the issue, which makes a plurality of different learners training and obtained the integrated results by combining the individual learner to get more accurate and ideal learning results. Therefore, this paper presents a polarimetric SAR image classification method based on the ensemble learning of sparse representation to achieve the optimal classification.

  1. A Syntactic Classification based Web Page Ranking Algorithm

    CERN Document Server

    Mukhopadhyay, Debajyoti; Kim, Young-Chon

    2011-01-01

    The existing search engines sometimes give unsatisfactory search result for lack of any categorization of search result. If there is some means to know the preference of user about the search result and rank pages according to that preference, the result will be more useful and accurate to the user. In the present paper a web page ranking algorithm is being proposed based on syntactic classification of web pages. Syntactic Classification does not bother about the meaning of the content of a web page. The proposed approach mainly consists of three steps: select some properties of web pages based on user's demand, measure them, and give different weightage to each property during ranking for different types of pages. The existence of syntactic classification is supported by running fuzzy c-means algorithm and neural network classification on a set of web pages. The change in ranking for difference in type of pages but for same query string is also being demonstrated.

  2. Automatic classification of time-variable X-ray sources

    CERN Document Server

    Lo, Kitty K; Murphy, Tara; Gaensler, B M

    2014-01-01

    To maximize the discovery potential of future synoptic surveys, especially in the field of transient science, it will be necessary to use automatic classification to identify some of the astronomical sources. The data mining technique of supervised classification is suitable for this problem. Here, we present a supervised learning method to automatically classify variable X-ray sources in the second \\textit{XMM-Newton} serendipitous source catalog (2XMMi-DR2). Random Forest is our classifier of choice since it is one of the most accurate learning algorithms available. Our training set consists of 873 variable sources and their features are derived from time series, spectra, and other multi-wavelength contextual information. The 10-fold cross validation accuracy of the training data is ${\\sim}$97% on a seven-class data set. We applied the trained classification model to 411 unknown variable 2XMM sources to produce a probabilistically classified catalog. Using the classification margin and the Random Forest der...

  3. Assessing Measures of Order Flow Toxicity via Perfect Trade Classification

    DEFF Research Database (Denmark)

    Andersen, Torben G.; Bondarenko, Oleg

    . The VPIN metric involves decomposing volume into active buys and sells. We use the best-bid-offer (BBO) files from the CME Group to construct (near) perfect trade classification measures for the E-mini S&P 500 futures contract. We investigate the accuracy of the ELO Bulk Volume Classification (BVC) scheme...... and find it inferior to a standard tick rule based on individual transactions. Moreover, when VPIN is constructed from accurate classification, it behaves in a diametrically opposite way to BVC-VPIN. We also find the latter to have forecast power for short-term volatility solely because it generates...... systematic classification errors that are correlated with trading volume and return volatility. When controlling for trading intensity and volatility, the BVC-VPIN measure has no incremental predictive power for future volatility. We conclude that VPIN is not suitable for measuring order flow imbalances....

  4. Land Cover Classification Using ALOS Imagery For Penang, Malaysia

    Science.gov (United States)

    Sim, C. K.; Abdullah, K.; MatJafri, M. Z.; Lim, H. S.

    2014-02-01

    This paper presents the potential of integrating optical and radar remote sensing data to improve automatic land cover mapping. The analysis involved standard image processing, and consists of spectral signature extraction and application of a statistical decision rule to identify land cover categories. A maximum likelihood classifier is utilized to determine different land cover categories. Ground reference data from sites throughout the study area are collected for training and validation. The land cover information was extracted from the digital data using PCI Geomatica 10.3.2 software package. The variations in classification accuracy due to a number of radar imaging processing techniques are studied. The relationship between the processing window and the land classification is also investigated. The classification accuracies from the optical and radar feature combinations are studied. Our research finds that fusion of radar and optical significantly improved classification accuracies. This study indicates that the land cover/use can be mapped accurately by using this approach.

  5. Classification using diffraction patterns for single-particle analysis

    Energy Technology Data Exchange (ETDEWEB)

    Hu, Hongli; Zhang, Kaiming [Department of Biophysics, the Health Science Centre, Peking University, Beijing 100191 (China); Meng, Xing, E-mail: xmeng101@gmail.com [Wadsworth Centre, New York State Department of Health, Albany, New York 12201 (United States)

    2016-05-15

    An alternative method has been assessed; diffraction patterns derived from the single particle data set were used to perform the first round of classification in creating the initial averages for proteins data with symmetrical morphology. The test protein set was a collection of Caenorhabditis elegans small heat shock protein 17 obtained by Cryo EM, which has a tetrahedral (12-fold) symmetry. It is demonstrated that the initial classification on diffraction patterns is workable as well as the real-space classification that is based on the phase contrast. The test results show that the information from diffraction patterns has the enough details to make the initial model faithful. The potential advantage using the alternative method is twofold, the ability to handle the sets with poor signal/noise or/and that break the symmetry properties. - Highlights: • New classification method. • Create the accurate initial model. • Better in handling noisy data.

  6. Flying insect detection and classification with inexpensive sensors.

    Science.gov (United States)

    Chen, Yanping; Why, Adena; Batista, Gustavo; Mafra-Neto, Agenor; Keogh, Eamonn

    2014-10-15

    An inexpensive, noninvasive system that could accurately classify flying insects would have important implications for entomological research, and allow for the development of many useful applications in vector and pest control for both medical and agricultural entomology. Given this, the last sixty years have seen many research efforts devoted to this task. To date, however, none of this research has had a lasting impact. In this work, we show that pseudo-acoustic optical sensors can produce superior data; that additional features, both intrinsic and extrinsic to the insect's flight behavior, can be exploited to improve insect classification; that a Bayesian classification approach allows to efficiently learn classification models that are very robust to over-fitting, and a general classification framework allows to easily incorporate arbitrary number of features. We demonstrate the findings with large-scale experiments that dwarf all previous works combined, as measured by the number of insects and the number of species considered.

  7. Spectral classification using convolutional neural networks

    CERN Document Server

    Hála, Pavel

    2014-01-01

    There is a great need for accurate and autonomous spectral classification methods in astrophysics. This thesis is about training a convolutional neural network (ConvNet) to recognize an object class (quasar, star or galaxy) from one-dimension spectra only. Author developed several scripts and C programs for datasets preparation, preprocessing and postprocessing of the data. EBLearn library (developed by Pierre Sermanet and Yann LeCun) was used to create ConvNets. Application on dataset of more than 60000 spectra yielded success rate of nearly 95%. This thesis conclusively proved great potential of convolutional neural networks and deep learning methods in astrophysics.

  8. [Classification of diabetes: an increasing heterogeneity].

    Science.gov (United States)

    Corcillo, Antonella; Corcillo Vionnet, Antonella; Jornayvaz, François R

    2015-06-03

    Diabetes mellitus is usually subdivided into type 1 and type 2. Despite precise criteria, distinction between these two types of diabetes can be difficult because of cases with superposition of the two classes. Adults aged 20 to 40 are particularly at risk of presenting an intermediary type of diabetes and thus are subject to misclassification. The distinction between these subtypes is relevant because of the therapeutic decision and the outcome which relies on insulin supply and therefore the evolution to insulin dependence. Thus, it seems important to review a new and more accurate classification of diabetes to offer a more appropriated care to patients.

  9. Identification of astigmatid mites using the second internal transcribed spacer (ITS2) region and its application for phylogenetic study.

    Science.gov (United States)

    Noge, Koji; Mori, Naoki; Tanaka, Chihiro; Nishida, Ritsuo; Tsuda, Mitsuya; Kuwahara, Yasumasa

    2005-01-01

    The second internal transcribed spacer (ITS2) of nuclear ribosomal DNA from 73 specimens of Astigmata was analyzed by PCR amplification and DNA sequencing. The length of the ITS2 region varied from 282 to 592 bp. The interspecific variation based on consensus sequences was more than 4.1%, while the intraspecific or intra-individual variation was from 0 to 5.7%. The variation between geographically separated populations (0-3.2%) was almost the same as the variation within strains. The sequences of the ITS2 region of Astigmata were concluded to be species-specific. The phylogenetic tree inferred from the ITS2 region supported Zachvatkin's morphological classification in the subfamily Rhizoglyphinae. The species-specific ITS2 sequence is useful for the species identification of astigmatid mites and for studying low-level phylogenetic relationships.

  10. Motif-Based Text Mining of Microbial Metagenome Redundancy Profiling Data for Disease Classification.

    Science.gov (United States)

    Wang, Yin; Li, Rudong; Zhou, Yuhua; Ling, Zongxin; Guo, Xiaokui; Xie, Lu; Liu, Lei

    2016-01-01

    Text data of 16S rRNA are informative for classifications of microbiota-associated diseases. However, the raw text data need to be systematically processed so that features for classification can be defined/extracted; moreover, the high-dimension feature spaces generated by the text data also pose an additional difficulty. Here we present a Phylogenetic Tree-Based Motif Finding algorithm (PMF) to analyze 16S rRNA text data. By integrating phylogenetic rules and other statistical indexes for classification, we can effectively reduce the dimension of the large feature spaces generated by the text datasets. Using the retrieved motifs in combination with common classification methods, we can discriminate different samples of both pneumonia and dental caries better than other existing methods. We extend the phylogenetic approaches to perform supervised learning on microbiota text data to discriminate the pathological states for pneumonia and dental caries. The results have shown that PMF may enhance the efficiency and reliability in analyzing high-dimension text data.

  11. Motif-Based Text Mining of Microbial Metagenome Redundancy Profiling Data for Disease Classification

    Directory of Open Access Journals (Sweden)

    Yin Wang

    2016-01-01

    Full Text Available Background. Text data of 16S rRNA are informative for classifications of microbiota-associated diseases. However, the raw text data need to be systematically processed so that features for classification can be defined/extracted; moreover, the high-dimension feature spaces generated by the text data also pose an additional difficulty. Results. Here we present a Phylogenetic Tree-Based Motif Finding algorithm (PMF to analyze 16S rRNA text data. By integrating phylogenetic rules and other statistical indexes for classification, we can effectively reduce the dimension of the large feature spaces generated by the text datasets. Using the retrieved motifs in combination with common classification methods, we can discriminate different samples of both pneumonia and dental caries better than other existing methods. Conclusions. We extend the phylogenetic approaches to perform supervised learning on microbiota text data to discriminate the pathological states for pneumonia and dental caries. The results have shown that PMF may enhance the efficiency and reliability in analyzing high-dimension text data.

  12. Validation and Classification of Web Services using Equalization Validation Classification

    Directory of Open Access Journals (Sweden)

    ALAMELU MUTHUKRISHNAN

    2012-12-01

    Full Text Available In the business process world, web services present a managed and middleware to connect huge number of services. Web service transaction is a mechanism to compose services with their desired quality parameters. If enormous transactions occur, the provider could not acquire the accurate data at the correct time. So it is necessary to reduce the overburden of web service t ransactions. In order to reduce the excess of transactions form customers to providers, this paper propose a new method called Equalization Validation Classification. This method introduces a new weight - reducing algorithm called Efficient Trim Down algorit hm to reduce the overburden of the incoming client requests. When this proposed algorithm is compared with Decision tree algorithms of (J48, Random Tree, Random Forest, AD Tree it produces a better accuracy and Validation than the existing algorithms. The proposed trimming method was analyzed with the Decision tree algorithms and the results implementation shows that the ETD algorithm provides better performance in terms of improved accuracy with Effective Validation. Therefore, the proposed method provide s a good gateway to reduce the overburden of the client requests in web services. Moreover analyzing the requests arrived from a vast number of clients and preventing the illegitimate requests save the service provider time

  13. Phylogenetic diversity (PD and biodiversity conservation: some bioinformatics challenges

    Directory of Open Access Journals (Sweden)

    Daniel P. Faith

    2006-01-01

    Full Text Available Biodiversity conservation addresses information challenges through estimations encapsulated in measures of diversity. A quantitative measure of phylogenetic diversity, “PD”, has been defined as the minimum total length of all the phylogenetic branches required to span a given set of taxa on the phylogenetic tree (Faith 1992a. While a recent paper incorrectly characterizes PD as not including information about deeper phylogenetic branches, PD applications over the past decade document the proper incorporation of shared deep branches when assessing the total PD of a set of taxa. Current PD applications to macroinvertebrate taxa in streams of New South Wales, Australia illustrate the practical importance of this definition. Phylogenetic lineages, often corresponding to new, “cryptic”, taxa, are restricted to a small number of stream localities. A recent case of human impact causing loss of taxa in one locality implies a higher PD value for another locality, because it now uniquely represents a deeper branch. This molecular-based phylogenetic pattern supports the use of DNA barcoding programs for biodiversity conservation planning. Here, PD assessments side-step the contentious use of barcoding-based “species” designations. Bio-informatics challenges include combining different phylogenetic evidence, optimization problems for conservation planning, and effective integration of phylogenetic information with environmental and socio-economic data.

  14. Student Interpretations of Phylogenetic Trees in an Introductory Biology Course

    Science.gov (United States)

    Dees, Jonathan; Momsen, Jennifer L.; Niemi, Jarad; Montplaisir, Lisa

    2014-01-01

    Phylogenetic trees are widely used visual representations in the biological sciences and the most important visual representations in evolutionary biology. Therefore, phylogenetic trees have also become an important component of biology education. We sought to characterize reasoning used by introductory biology students in interpreting taxa…

  15. A perl package and an alignment tool for phylogenetic networks.

    Science.gov (United States)

    Cardona, Gabriel; Rosselló, Francesc; Valiente, Gabriel

    2008-03-27

    Phylogenetic networks are a generalization of phylogenetic trees that allow for the representation of evolutionary events acting at the population level, like recombination between genes, hybridization between lineages, and lateral gene transfer. While most phylogenetics tools implement a wide range of algorithms on phylogenetic trees, there exist only a few applications to work with phylogenetic networks, none of which are open-source libraries, and they do not allow for the comparative analysis of phylogenetic networks by computing distances between them or aligning them. In order to improve this situation, we have developed a Perl package that relies on the BioPerl bundle and implements many algorithms on phylogenetic networks. We have also developed a Java applet that makes use of the aforementioned Perl package and allows the user to make simple experiments with phylogenetic networks without having to develop a program or Perl script by him or herself. The Perl package is available as part of the BioPerl bundle, and can also be downloaded. A web-based application is also available (see availability and requirements). The Perl package includes full documentation of all its features.

  16. A perl package and an alignment tool for phylogenetic networks

    Directory of Open Access Journals (Sweden)

    Valiente Gabriel

    2008-03-01

    Full Text Available Abstract Background Phylogenetic networks are a generalization of phylogenetic trees that allow for the representation of evolutionary events acting at the population level, like recombination between genes, hybridization between lineages, and lateral gene transfer. While most phylogenetics tools implement a wide range of algorithms on phylogenetic trees, there exist only a few applications to work with phylogenetic networks, none of which are open-source libraries, and they do not allow for the comparative analysis of phylogenetic networks by computing distances between them or aligning them. Results In order to improve this situation, we have developed a Perl package that relies on the BioPerl bundle and implements many algorithms on phylogenetic networks. We have also developed a Java applet that makes use of the aforementioned Perl package and allows the user to make simple experiments with phylogenetic networks without having to develop a program or Perl script by him or herself. Conclusion The Perl package is available as part of the BioPerl bundle, and can also be downloaded. A web-based application is also available (see availability and requirements. The Perl package includes full documentation of all its features.

  17. Utilization of complete chloroplast genomes for phylogenetic studies

    NARCIS (Netherlands)

    Ramlee, Shairul Izan Binti

    2016-01-01

    Chloroplast DNA sequence polymorphisms are a primary source of data in many plant phylogenetic studies. The chloroplast genome is relatively conserved in its evolution making it an ideal molecule to retain phylogenetic signals. The chloroplast genome is also largely, but not completely, free from ot

  18. PhyDesign: an online application for profiling phylogenetic informativeness

    Directory of Open Access Journals (Sweden)

    Townsend Jeffrey P

    2011-05-01

    Full Text Available Abstract Background The rapid increase in number of sequenced genomes for species across of the tree of life is revealing a diverse suite of orthologous genes that could potentially be employed to inform molecular phylogenetic studies that encompass broader taxonomic sampling. Optimal usage of this diversity of loci requires user-friendly tools to facilitate widespread cost-effective locus prioritization for phylogenetic sampling. The Townsend (2007 phylogenetic informativeness provides a unique empirical metric for guiding marker selection. However, no software or automated methodology to evaluate sequence alignments and estimate the phylogenetic informativeness metric has been available. Results Here, we present PhyDesign, a platform-independent online application that implements the Townsend (2007 phylogenetic informativeness analysis, providing a quantitative prediction of the utility of loci to solve specific phylogenetic questions. An easy-to-use interface facilitates uploading of alignments and ultrametric trees to calculate and depict profiles of informativeness over specified time ranges, and provides rankings of locus prioritization for epochs of interest. Conclusions By providing these profiles, PhyDesign facilitates locus prioritization increasing the efficiency of sequencing for phylogenetic purposes compared to traditional studies with more laborious and low capacity screening methods, as well as increasing the accuracy of phylogenetic studies. Together with a manual and sample files, the application is freely accessible at http://phydesign.townsend.yale.edu.

  19. Phylogenetic Analysis of Viridans Group Streptococci Causing Endocarditis ▿

    Science.gov (United States)

    Simmon, Keith E.; Hall, Lori; Woods, Christopher W.; Marco, Francesc; Miro, Jose M.; Cabell, Christopher; Hoen, Bruno; Marin, Mercedes; Utili, Riccardo; Giannitsioti, Efthymia; Doco-Lecompte, Thanh; Bradley, Suzanne; Mirrett, Stanley; Tambic, Arjana; Ryan, Suzanne; Gordon, David; Jones, Phillip; Korman, Tony; Wray, Dannah; Reller, L. Barth; Tripodi, Marie-Francoise; Plesiat, Patrick; Morris, Arthur J.; Lang, Selwyn; Murdoch, David R.; Petti, Cathy A.

    2008-01-01

    Identification of viridans group streptococci (VGS) to the species level is difficult because VGS exchange genetic material. We performed multilocus DNA target sequencing to assess phylogenetic concordance of VGS for a well-defined clinical syndrome. The hierarchy of sequence data was often discordant, underscoring the importance of establishing biological relevance for finer phylogenetic distinctions. PMID:18650347

  20. Phylogenetic analysis of viridans group streptococci causing endocarditis.

    Science.gov (United States)

    Simmon, Keith E; Hall, Lori; Woods, Christopher W; Marco, Francesc; Miro, Jose M; Cabell, Christopher; Hoen, Bruno; Marin, Mercedes; Utili, Riccardo; Giannitsioti, Efthymia; Doco-Lecompte, Thanh; Bradley, Suzanne; Mirrett, Stanley; Tambic, Arjana; Ryan, Suzanne; Gordon, David; Jones, Phillip; Korman, Tony; Wray, Dannah; Reller, L Barth; Tripodi, Marie-Francoise; Plesiat, Patrick; Morris, Arthur J; Lang, Selwyn; Murdoch, David R; Petti, Cathy A

    2008-09-01

    Identification of viridans group streptococci (VGS) to the species level is difficult because VGS exchange genetic material. We performed multilocus DNA target sequencing to assess phylogenetic concordance of VGS for a well-defined clinical syndrome. The hierarchy of sequence data was often discordant, underscoring the importance of establishing biological relevance for finer phylogenetic distinctions.

  1. Shark Teeth Classification

    Science.gov (United States)

    Brown, Tom; Creel, Sally; Lee, Velda

    2009-01-01

    On a recent autumn afternoon at Harmony Leland Elementary in Mableton, Georgia, students in a fifth-grade science class investigated the essential process of classification--the act of putting things into groups according to some common characteristics or attributes. While they may have honed these skills earlier in the week by grouping their own…

  2. Sandwich classification theorem

    Directory of Open Access Journals (Sweden)

    Alexey Stepanov

    2015-09-01

    Full Text Available The present note arises from the author's talk at the conference ``Ischia Group Theory 2014''. For subgroups FleN of a group G denote by Lat(F,N the set of all subgroups of N , containing F . Let D be a subgroup of G . In this note we study the lattice LL=Lat(D,G and the lattice LL ′ of subgroups of G , normalized by D . We say that LL satisfies sandwich classification theorem if LL splits into a disjoint union of sandwiches Lat(F,N G (F over all subgroups F such that the normal closure of D in F coincides with F . Here N G (F denotes the normalizer of F in G . A similar notion of sandwich classification is introduced for the lattice LL ′ . If D is perfect, i.,e. coincides with its commutator subgroup, then it turns out that sandwich classification theorem for LL and LL ′ are equivalent. We also show how to find basic subroup F of sandwiches for LL ′ and review sandwich classification theorems in algebraic groups over rings.

  3. Dynamic Latent Classification Model

    DEFF Research Database (Denmark)

    Zhong, Shengtong; Martínez, Ana M.; Nielsen, Thomas Dyhre

    as possible. Motivated by this problem setting, we propose a generative model for dynamic classification in continuous domains. At each time point the model can be seen as combining a naive Bayes model with a mixture of factor analyzers (FA). The latent variables of the FA are used to capture the dynamics...... in the process as well as modeling dependences between attributes....

  4. An automated cirrus classification

    Science.gov (United States)

    Gryspeerdt, Edward; Quaas, Johannes; Sourdeval, Odran; Goren, Tom

    2017-04-01

    Cirrus clouds play an important role in determining the radiation budget of the earth, but our understanding of the lifecycle and controls on cirrus clouds remains incomplete. Cirrus clouds can have very different properties and development depending on their environment, particularly during their formation. However, the relevant factors often cannot be distinguished using commonly retrieved satellite data products (such as cloud optical depth). In particular, the initial cloud phase has been identified as an important factor in cloud development, but although back-trajectory based methods can provide information on the initial cloud phase, they are computationally expensive and depend on the cloud parametrisations used in re-analysis products. In this work, a classification system (Identification and Classification of Cirrus, IC-CIR) is introduced. Using re-analysis and satellite data, cirrus clouds are separated in four main types: frontal, convective, orographic and in-situ. The properties of these classes show that this classification is able to provide useful information on the properties and initial phase of cirrus clouds, information that could not be provided by instantaneous satellite retrieved cloud properties alone. This classification is designed to be easily implemented in global climate models, helping to improve future comparisons between observations and models and reducing the uncertainty in cirrus clouds properties, leading to improved cloud parametrisations.

  5. Classifications in popular music

    NARCIS (Netherlands)

    van Venrooij, A.; Schmutz, V.; Wright, J.D.

    2015-01-01

    The categorical system of popular music, such as genre categories, is a highly differentiated and dynamic classification system. In this article we present work that studies different aspects of these categorical systems in popular music. Following the work of Paul DiMaggio, we focus on four questio

  6. Nearest convex hull classification

    NARCIS (Netherlands)

    G.I. Nalbantov (Georgi); P.J.F. Groenen (Patrick); J.C. Bioch (Cor)

    2006-01-01

    textabstractConsider the classification task of assigning a test object to one of two or more possible groups, or classes. An intuitive way to proceed is to assign the object to that class, to which the distance is minimal. As a distance measure to a class, we propose here to use the distance to the

  7. Principles for ecological classification

    Science.gov (United States)

    Dennis H. Grossman; Patrick Bourgeron; Wolf-Dieter N. Busch; David T. Cleland; William Platts; G. Ray; C. Robins; Gary Roloff

    1999-01-01

    The principal purpose of any classification is to relate common properties among different entities to facilitate understanding of evolutionary and adaptive processes. In the context of this volume, it is to facilitate ecosystem stewardship, i.e., to help support ecosystem conservation and management objectives.

  8. Improving Student Question Classification

    Science.gov (United States)

    Heiner, Cecily; Zachary, Joseph L.

    2009-01-01

    Students in introductory programming classes often articulate their questions and information needs incompletely. Consequently, the automatic classification of student questions to provide automated tutorial responses is a challenging problem. This paper analyzes 411 questions from an introductory Java programming course by reducing the natural…

  9. Classification of waste packages

    Energy Technology Data Exchange (ETDEWEB)

    Mueller, H.P.; Sauer, M.; Rojahn, T. [Versuchsatomkraftwerk GmbH, Kahl am Main (Germany)

    2001-07-01

    A barrel gamma scanning unit has been in use at the VAK for the classification of radioactive waste materials since 1998. The unit provides the facility operator with the data required for classification of waste barrels. Once these data have been entered into the AVK data processing system, the radiological status of raw waste as well as pre-treated and processed waste can be tracked from the point of origin to the point at which the waste is delivered to a final storage. Since the barrel gamma scanning unit was commissioned in 1998, approximately 900 barrels have been measured and the relevant data required for classification collected and analyzed. Based on the positive results of experience in the use of the mobile barrel gamma scanning unit, the VAK now offers the classification of barrels as a service to external users. Depending upon waste quantity accumulation, this measurement unit offers facility operators a reliable and time-saving and cost-effective means of identifying and documenting the radioactivity inventory of barrels scheduled for final storage. (orig.)

  10. Event Classification using Concepts

    NARCIS (Netherlands)

    Boer, M.H.T. de; Schutte, K.; Kraaij, W.

    2013-01-01

    The semantic gap is one of the challenges in the GOOSE project. In this paper a Semantic Event Classification (SEC) system is proposed as an initial step in tackling the semantic gap challenge in the GOOSE project. This system uses semantic text analysis, multiple feature detectors using the BoW

  11. Munitions Classification Library

    Science.gov (United States)

    2016-04-04

    the MM and TEMTADS 2x2 systems , with dynamic data handling for these systems on the horizon. Using UX-Analyze, a data processor can apply physics ... classification libraries. DRAFT 4 2.0 TECHNOLOGY Three different sensor systems were used during the initial phase of the data collection: a modified...

  12. Recurrent neural collective classification.

    Science.gov (United States)

    Monner, Derek D; Reggia, James A

    2013-12-01

    With the recent surge in availability of data sets containing not only individual attributes but also relationships, classification techniques that take advantage of predictive relationship information have gained in popularity. The most popular existing collective classification techniques have a number of limitations-some of them generate arbitrary and potentially lossy summaries of the relationship data, whereas others ignore directionality and strength of relationships. Popular existing techniques make use of only direct neighbor relationships when classifying a given entity, ignoring potentially useful information contained in expanded neighborhoods of radius greater than one. We present a new technique that we call recurrent neural collective classification (RNCC), which avoids arbitrary summarization, uses information about relationship directionality and strength, and through recursive encoding, learns to leverage larger relational neighborhoods around each entity. Experiments with synthetic data sets show that RNCC can make effective use of relationship data for both direct and expanded neighborhoods. Further experiments demonstrate that our technique outperforms previously published results of several collective classification methods on a number of real-world data sets.

  13. Event Classification using Concepts

    NARCIS (Netherlands)

    Boer, M.H.T. de; Schutte, K.; Kraaij, W.

    2013-01-01

    The semantic gap is one of the challenges in the GOOSE project. In this paper a Semantic Event Classification (SEC) system is proposed as an initial step in tackling the semantic gap challenge in the GOOSE project. This system uses semantic text analysis, multiple feature detectors using the BoW mod

  14. Phylogenetic relationships of Mesoamerican spider monkeys (Ateles geoffroyi): Molecular evidence suggests the need for a revised taxonomy.

    Science.gov (United States)

    Morales-Jimenez, Alba Lucia; Cortés-Ortiz, Liliana; Di Fiore, Anthony

    2015-01-01

    Mesoamerican spider monkeys (Ateles geoffroyi sensu lato) are widely distributed from Mexico to northern Colombia. This group of primates includes many allopatric forms with morphologically distinct pelage color and patterning, but its taxonomy and phylogenetic history are poorly understood. We explored the genetic relationships among the different forms of Mesoamerican spider monkeys using mtDNA sequence data, and we offer a new hypothesis for the evolutionary history of the group. We collected up to ∼800 bp of DNA sequence data from hypervariable region 1 (HV1) of the control region, or D-loop, of the mitochondrion for multiple putative subspecies of Ateles geoffroyi sensu lato. Both maximum likelihood and Bayesian reconstructions, using Ateles paniscus as an outgroup, showed that (1) A. fusciceps and A. geoffroyi form two different monophyletic groups and (2) currently recognized subspecies of A. geoffroyi are not monophyletic. Within A. geoffroyi, our phylogenetic analysis revealed little concordance between any of the classifications proposed for this taxon and their phylogenetic relationships, therefore a new classification is needed for this group. Several possible clades with recent divergence times (1.7-0.8 Ma) were identified within Ateles geoffroyi sensu lato. Some previously recognized taxa were not separated by our data (e.g., A. g. vellerosus and A. g. yucatanensis), while one distinct clade had never been described as a different evolutionary unit based on pelage or geography (Ateles geoffroyi ssp. indet. from El Salvador). Based on well-supported phylogenetic relationships, our results challenge previous taxonomic arrangements for Mesoamerican spider monkeys. We suggest a revised arrangement based on our data and call for a thorough taxonomic revision of this group. Copyright © 2014. Published by Elsevier Inc.

  15. Accurate Switched-Voltage voltage averaging circuit

    OpenAIRE

    金光, 一幸; 松本, 寛樹

    2006-01-01

    Abstract ###This paper proposes an accurate Switched-Voltage (SV) voltage averaging circuit. It is presented ###to compensated for NMOS missmatch error at MOS differential type voltage averaging circuit. ###The proposed circuit consists of a voltage averaging and a SV sample/hold (S/H) circuit. It can ###operate using nonoverlapping three phase clocks. Performance of this circuit is verified by PSpice ###simulations.

  16. Accurate overlaying for mobile augmented reality

    NARCIS (Netherlands)

    Pasman, W; van der Schaaf, A; Lagendijk, RL; Jansen, F.W.

    1999-01-01

    Mobile augmented reality requires accurate alignment of virtual information with objects visible in the real world. We describe a system for mobile communications to be developed to meet these strict alignment criteria using a combination of computer vision. inertial tracking and low-latency renderi

  17. Accurate overlaying for mobile augmented reality

    NARCIS (Netherlands)

    Pasman, W; van der Schaaf, A; Lagendijk, RL; Jansen, F.W.

    1999-01-01

    Mobile augmented reality requires accurate alignment of virtual information with objects visible in the real world. We describe a system for mobile communications to be developed to meet these strict alignment criteria using a combination of computer vision. inertial tracking and low-latency

  18. Site-specific time heterogeneity of the substitution process and its impact on phylogenetic inference

    Directory of Open Access Journals (Sweden)

    Philippe Hervé

    2011-01-01

    Full Text Available Abstract Background Model violations constitute the major limitation in inferring accurate phylogenies. Characterizing properties of the data that are not being correctly handled by current models is therefore of prime importance. One of the properties of protein evolution is the variation of the relative rate of substitutions across sites and over time, the latter is the phenomenon called heterotachy. Its effect on phylogenetic inference has recently obtained considerable attention, which led to the development of new models of sequence evolution. However, thus far focus has been on the quantitative heterogeneity of the evolutionary process, thereby overlooking more qualitative variations. Results We studied the importance of variation of the site-specific amino-acid substitution process over time and its possible impact on phylogenetic inference. We used the CAT model to define an infinite mixture of substitution processes characterized by equilibrium frequencies over the twenty amino acids, a useful proxy for qualitatively estimating the evolutionary process. Using two large datasets, we show that qualitative changes in site-specific substitution properties over time occurred significantly. To test whether this unaccounted qualitative variation can lead to an erroneous phylogenetic tree, we analyzed a concatenation of mitochondrial proteins in which Cnidaria and Porifera were erroneously grouped. The progressive removal of the sites with the most heterogeneous CAT profiles across clades led to the recovery of the monophyly of Eumetazoa (Cnidaria+Bilateria, suggesting that this heterogeneity can negatively influence phylogenetic inference. Conclusion The time-heterogeneity of the amino-acid replacement process is therefore an important evolutionary aspect that should be incorporated in future models of sequence change.

  19. Visualising very large phylogenetic trees in three dimensional hyperbolic space

    Directory of Open Access Journals (Sweden)

    Liberles David A

    2004-04-01

    Full Text Available Abstract Background Common existing phylogenetic tree visualisation tools are not able to display readable trees with more than a few thousand nodes. These existing methodologies are based in two dimensional space. Results We introduce the idea of visualising phylogenetic trees in three dimensional hyperbolic space with the Walrus graph visualisation tool and have developed a conversion tool that enables the conversion of standard phylogenetic tree formats to Walrus' format. With Walrus, it becomes possible to visualise and navigate phylogenetic trees with more than 100,000 nodes. Conclusion Walrus enables desktop visualisation of very large phylogenetic trees in 3 dimensional hyperbolic space. This application is potentially useful for visualisation of the tree of life and for functional genomics derivatives, like The Adaptive Evolution Database (TAED.

  20. Open Reading Frame Phylogenetic Analysis on the Cloud

    Directory of Open Access Journals (Sweden)

    Che-Lun Hung

    2013-01-01

    Full Text Available Phylogenetic analysis has become essential in researching the evolutionary relationships between viruses. These relationships are depicted on phylogenetic trees, in which viruses are grouped based on sequence similarity. Viral evolutionary relationships are identified from open reading frames rather than from complete sequences. Recently, cloud computing has become popular for developing internet-based bioinformatics tools. Biocloud is an efficient, scalable, and robust bioinformatics computing service. In this paper, we propose a cloud-based open reading frame phylogenetic analysis service. The proposed service integrates the Hadoop framework, virtualization technology, and phylogenetic analysis methods to provide a high-availability, large-scale bioservice. In a case study, we analyze the phylogenetic relationships among Norovirus. Evolutionary relationships are elucidated by aligning different open reading frame sequences. The proposed platform correctly identifies the evolutionary relationships between members of Norovirus.