WorldWideScience

Sample records for protein sequence polymorphisms

  1. The Number, Organization, and Size of Polymorphic Membrane Protein Coding Sequences as well as the Most Conserved Pmp Protein Differ within and across Chlamydia Species.

    Science.gov (United States)

    Van Lent, Sarah; Creasy, Heather Huot; Myers, Garry S A; Vanrompay, Daisy

    2016-01-01

    Variation is a central trait of the polymorphic membrane protein (Pmp) family. The number of pmp coding sequences differs between Chlamydia species, but it is unknown whether the number of pmp coding sequences is constant within a Chlamydia species. The level of conservation of the Pmp proteins has previously only been determined for Chlamydia trachomatis. As different Pmp proteins might be indispensible for the pathogenesis of different Chlamydia species, this study investigated the conservation of Pmp proteins both within and across C. trachomatis,C. pneumoniae,C. abortus, and C. psittaci. The pmp coding sequences were annotated in 16 C. trachomatis, 6 C. pneumoniae, 2 C. abortus, and 16 C. psittaci genomes. The number and organization of polymorphic membrane coding sequences differed within and across the analyzed Chlamydia species. The length of coding sequences of pmpA,pmpB, and pmpH was conserved among all analyzed genomes, while the length of pmpE/F and pmpG, and remarkably also of the subtype pmpD, differed among the analyzed genomes. PmpD, PmpA, PmpH, and PmpA were the most conserved Pmp in C. trachomatis,C. pneumoniae,C. abortus, and C. psittaci, respectively. PmpB was the most conserved Pmp across the 4 analyzed Chlamydia species. © 2016 S. Karger AG, Basel.

  2. Templated sequence insertion polymorphisms in the human genome

    Science.gov (United States)

    Onozawa, Masahiro; Aplan, Peter

    2016-11-01

    Templated Sequence Insertion Polymorphism (TSIP) is a recently described form of polymorphism recognized in the human genome, in which a sequence that is templated from a distant genomic region is inserted into the genome, seemingly at random. TSIPs can be grouped into two classes based on nucleotide sequence features at the insertion junctions; Class 1 TSIPs show features of insertions that are mediated via the LINE-1 ORF2 protein, including 1) target-site duplication (TSD), 2) polyadenylation 10-30 nucleotides downstream of a “cryptic” polyadenylation signal, and 3) preference for insertion at a 5’-TTTT/A-3’ sequence. In contrast, class 2 TSIPs show features consistent with repair of a DNA double-strand break via insertion of a DNA “patch” that is derived from a distant genomic region. Survey of a large number of normal human volunteers demonstrates that most individuals have 25-30 TSIPs, and that these TSIPs track with specific geographic regions. Similar to other forms of human polymorphism, we suspect that these TSIPs may be important for the generation of human diversity and genetic diseases.

  3. Yersinia pestis Caf1 Protein: Effect of Sequence Polymorphism on Intrinsic Disorder Propensity, Serological Cross-Reactivity and Cross-Protectivity of Isoforms.

    Directory of Open Access Journals (Sweden)

    Pavel Kh Kopylov

    Full Text Available Yersinia pestis Caf1 is a multifunctional protein responsible for antiphagocytic activity and is a key protective antigen. It is generally conserved between globally distributed Y. pestis strains, but Y. pestis subsp. microtus biovar caucasica strains circulating within populations of common voles in Georgia and Armenia were reported to carry a single substitution of alanine to serine. We investigated polymorphism of the Caf1 sequences among other Y. pestis subsp. microtus strains, which have a limited virulence in guinea pigs and in humans. Sequencing of caf1 genes from 119 Y. pestis strains belonging to different biovars within subsp. microtus showed that the Caf1 proteins exist in three isoforms, the global type Caf1NT1 (Ala48 Phe117, type Caf1NT2 (Ser48 Phe117 found in Transcaucasian-highland and Pre-Araks natural plague foci #4-7, and a novel Caf1NT3 type (Ala48 Val117 endemic in Dagestan-highland natural plague focus #39. Both minor types are the progenies of the global isoform. In this report, Caf1 polymorphism was analyzed by comparing predicted intrinsic disorder propensities and potential protein-protein interactivities of the three Caf1 isoforms. The analysis revealed that these properties of Caf1 protein are minimally affected by its polymorphism. All protein isoforms could be equally detected by an immunochromatography test for plague at the lowest protein concentration tested (1.0 ng/mL, which is the detection limit. When compared to the classic Caf1NT1 isoform, the endemic Caf1NT2 or Caf1NT3 had lower immunoreactivity in ELISA and lower indices of self- and cross-protection. Despite a visible reduction in cross-protection between all Caf1 isoforms, our data suggest that polymorphism in the caf1 gene may not allow the carriers of Caf1NT2 or Caf1NT3 variants escaping from the Caf1NT1-mediated immunity to plague in the case of a low-dose flea-borne infection.

  4. Polymorphism Sequence - JSNP | LSDB Archive [Life Science Database Archive metadata

    Lifescience Database Archive (English)

    Full Text Available List Contact us JSNP Polymorphism Sequence Data detail Data name Polymorphism Sequence DOI 10.18908/lsdba.nb...dc00114-001 Description of data contents Information on polymorphisms (SNPs and insertions/deletions) and th...se Name database name JSNP_SNP: single nucleotide polymorphism JSNP_InsDel_IND: insertion/deletion JSNP_InsD...ved allele observed 3' Flanking Sequence 3' flanking sequence Offset in Flanking Sequence position of the polymorphism...uence Accession No. accession No. of the sequence for polymorphism screening Offset in Record position of the polymorphism

  5. Comparative analysis of the prion protein gene sequences in African lion.

    Science.gov (United States)

    Wu, Chang-De; Pang, Wan-Yong; Zhao, De-Ming

    2006-10-01

    The prion protein gene of African lion (Panthera Leo) was first cloned and polymorphisms screened. The results suggest that the prion protein gene of eight African lions is highly homogenous. The amino acid sequences of the prion protein (PrP) of all samples tested were identical. Four single nucleotide polymorphisms (C42T, C81A, C420T, T600C) in the prion protein gene (Prnp) of African lion were found, but no amino acid substitutions. Sequence analysis showed that the higher homology is observed to felis catus AF003087 (96.7%) and to sheep number M31313.1 (96.2%) Genbank accessed. With respect to all the mammalian prion protein sequences compared, the African lion prion protein sequence has three amino acid substitutions. The homology might in turn affect the potential intermolecular interactions critical for cross species transmission of prion disease.

  6. The first report of prion-related protein gene (PRNT) polymorphisms in goat.

    Science.gov (United States)

    Kim, Yong-Chan; Jeong, Byung-Hoon

    2017-06-01

    Prion protein is encoded by the prion protein gene (PRNP). Polymorphisms of several members of the prion gene family have shown association with prion diseases in several species. Recent studies on a novel member of the prion gene family in rams have shown that prion-related protein gene (PRNT) has a linkage with codon 26 of prion-like protein (PRND). In a previous study, codon 26 polymorphism of PRND has shown connection with PRNP haplotype which is strongly associated with scrapie vulnerability. In addition, the genotype of a single nucleotide polymorphism (SNP) at codon 26 of PRND is related to fertilisation capacity. These findings necessitate studies on the SNP of PRNT gene which is connected with PRND. In goat, several polymorphism studies have been performed for PRNP, PRND, and shadow of prion protein gene (SPRN). However, polymorphism on PRNT has not been reported. Hence, the objective of this study was to determine the genotype and allelic distribution of SNPs of PRNT in 238 Korean native goats and compare PRNT DNA sequences between Korean native goats and several ruminant species. A total of five SNPs, including PRNT c.-114G > T, PRNT c.-58A > G in the upstream of PRNT gene, PRNT c.71C > T (p.Ala24Val) and PRNT c.102G > A in the open reading frame (ORF) and c.321C > T in the downstream of PRNT gene, were found in this study. All five SNPs of caprine PRNT gene in Korean native goat are in complete linkage disequilibrium (LD) with a D' value of 1.0. Interestingly, comparative sequence analysis of the PRNT gene revealed five mismatches between DNA sequences of Korean native goats and those of goats deposited in the GenBank. Korean native black goats also showed 5 mismatches in PRNT ORF with cattle. To the best of our knowledge, this is the first genetic research of the PRNT gene in goat.

  7. Variability of the protein sequences of lcrV between epidemic and atypical rhamnose-positive strains of Yersinia pestis.

    Science.gov (United States)

    Anisimov, Andrey P; Panfertsev, Evgeniy A; Svetoch, Tat'yana E; Dentovskaya, Svetlana V

    2007-01-01

    Sequencing of lcrV genes and comparison of the deduced amino acid sequences from ten Y. pestis strains belonging mostly to the group of atypical rhamnose-positive isolates (non-pestis subspecies or pestoides group) showed that the LcrV proteins analyzed could be classified into five sequence types. This classification was based on major amino acid polymorphisms among LcrV proteins in the four "hot points" of the protein sequences. Some additional minor polymorphisms were found throughout these sequence types. The "hot points" corresponded to amino acids 18 (Lys --> Asn), 72 (Lys --> Arg), 273 (Cys --> Ser), and 324-326 (Ser-Gly-Lys --> Arg) in the LcrV sequence of the reference Y. pestis strain CO92. One possible explanation for polymorphism in amino acid sequences of LcrV among different strains is that strain-specific variation resulted from adaptation of the plague pathogen to different rodent and lagomorph hosts.

  8. Methylation sensitive-sequence related amplified polymorphism (MS ...

    African Journals Online (AJOL)

    DR NJ TONUKARI

    2011-04-25

    Apr 25, 2011 ... Sequence-related amplified polymorphism (SRAP) is a simple but an efficient gene amplification marker system for both .... Each polymorphic band reflecting different methylation status at the ... After boiling for 5 min in the water, the .... CpG dinucleotides in the open reading frame of a testicular germ cell-.

  9. Evidence of Alternative Cystatin C Signal Sequence Cleavage Which Is Influenced by the A25T Polymorphism.

    Directory of Open Access Journals (Sweden)

    Annie Nguyen

    Full Text Available Cystatin C (Cys C is a small, potent, cysteine protease inhibitor. An Ala25Thr (A25T polymorphism in Cys C has been associated with both macular degeneration and late-onset Alzheimer's disease. Previously, studies have suggested that this polymorphism may compromise the secretion of Cys C. Interestingly, we found that untagged A25T, A25T tagged C-terminally with FLAG, or A25T FLAG followed by green fluorescent protein (GFP, were all secreted as efficiently from immortalized human cells as their wild-type (WT counterparts (e.g., 112%, 100%, and 88% of WT levels from HEK-293T cells, respectively. Supporting these observations, WT and A25T Cys C variants also showed similar intracellular steady state levels. Furthermore, A25T Cys C did not activate the unfolded protein response and followed the same canonical endoplasmic reticulum (ER-Golgi trafficking pathway as WT Cys C. WT Cys C has been shown to undergo signal sequence cleavage between residues Gly26 and Ser27. While the A25T polymorphism did not affect Cys C secretion, we hypothesized that it may alter where the Cys C signal sequence is preferentially cleaved. Under normal conditions, WT and A25T Cys C have the same signal sequence cleavage site after Gly26 (referred to as 'site 2' cleavage. However, in particular circumstances when the residues around site 2 are modified (such as by the presence of an N-terminal FLAG tag immediately after Gly26, or by a Gly26Lys (G26K mutation, A25T has a significantly higher likelihood than WT Cys C of alternative signal sequence cleavage after Ala20 ('site 1' or even earlier in the Cys C sequence. Overall, our results indicate that the A25T polymorphism does not cause a significant reduction in Cys C secretion, but instead predisposes the protein to be cleaved at an alternative signal sequence cleavage site if site 2 is hindered. Additional N-terminal amino acids resulting from alternative signal sequence cleavage may, in turn, affect the protease

  10. Conserved hypothetical protein Rv1977 in Mycobacterium tuberculosis strains contains sequence polymorphisms and might be involved in ongoing immune evasion.

    Science.gov (United States)

    Jiang, Yi; Liu, Haican; Wang, Xuezhi; Li, Guilian; Qiu, Yan; Dou, Xiangfeng; Wan, Kanglin

    2015-01-01

    Host immune pressure and associated parasite immune evasion are key features of host-pathogen co-evolution. A previous study showed that human T cell epitopes of Mycobacterium tuberculosis are evolutionarily hyperconserved and thus it was deduced that M. tuberculosis lacks antigenic variation and immune evasion. Here, we selected 151 clinical Mycobacterium tuberculosis isolates from China, amplified gene encoding Rv1977 and compared the sequences. The results showed that Rv1977, a conserved hypothetical protein, is not conserved in M. tuberculosis strains and there are polymorphisms existed in the protein. Some mutations, especially one frameshift mutation, occurred in the antigen Rv1977, which is uncommon in M.tb strains and may lead to the protein function altering. Mutations and deletion in the gene all affect one of three T cell epitopes and the changed T cell epitope contained more than one variable position, which may suggest ongoing immune evasion.

  11. Bioinformatic Analysis of Deleterious Non-Synonymous Single Nucleotide Polymorphisms (nsSNPs in the Coding Regions of Human Prion Protein Gene (PRNP

    Directory of Open Access Journals (Sweden)

    Kourosh Bamdad

    2016-12-01

    Full Text Available Background & Objective: Single nucleotide polymorphisms are the cause of genetic variation to living organisms. Single nucleotide polymorphisms alter residues in the protein sequence. In this investigation, the relationship between prion protein gene polymorphisms and its relevance to pathogenicity was studied. Material & Method: Amino acid sequence of the main isoform from the human prion protein gene (PRNP was extracted from UniProt database and evaluated by FoldAmyloid and AmylPred servers. All non-synonymous single nucleotide polymorphisms (nsSNPs from SNP database (dbSNP were further analyzed by bioinformatics servers including SIFT, PolyPhen-2, I-Mutant-3.0, PANTHER, SNPs & GO, PHD-SNP, Meta-SNP, and MutPred to determine the most damaging nsSNPs. Results: The results of the first structure analyses by FoldAmyloid and AmylPerd servers implied that regions including 5-15, 174-178, 180-184, 211-217, and 240-252 were the most sensitive parts of the protein sequence to amyloidosis. Screening all nsSNPs of the main protein isoform using bioinformatic servers revealed that substitution of Aspartic acid with Valine at position 178 (ID code: rs11538766 was the most deleterious nsSNP in the protein structure. Conclusion:  Substitution of the Aspartic acid with Valine at position 178 (D178V was the most pathogenic mutation in the human prion protein gene. Analyses from the MutPred server also showed that beta-sheets’ increment in the secondary structure was the main reason behind the molecular mechanism of the prion protein aggregation.

  12. Full-length sequencing and identification of novel polymorphisms in ...

    Indian Academy of Sciences (India)

    The aim of this work was to sequence the entirecoding region of ACACA gene in Valle del Belice sheep breed to identify polymorphic sites. A total of 51 coding exons of ACACA gene were sequenced in 32 individuals of Valle del Belice sheep breed. Sequencing analysis and alignment of obtained sequences showed the ...

  13. Sequencing genes in silico using single nucleotide polymorphisms

    Directory of Open Access Journals (Sweden)

    Zhang Xinyi

    2012-01-01

    Full Text Available Abstract Background The advent of high throughput sequencing technology has enabled the 1000 Genomes Project Pilot 3 to generate complete sequence data for more than 906 genes and 8,140 exons representing 697 subjects. The 1000 Genomes database provides a critical opportunity for further interpreting disease associations with single nucleotide polymorphisms (SNPs discovered from genetic association studies. Currently, direct sequencing of candidate genes or regions on a large number of subjects remains both cost- and time-prohibitive. Results To accelerate the translation from discovery to functional studies, we propose an in silico gene sequencing method (ISS, which predicts phased sequences of intragenic regions, using SNPs. The key underlying idea of our method is to infer diploid sequences (a pair of phased sequences/alleles at every functional locus utilizing the deep sequencing data from the 1000 Genomes Project and SNP data from the HapMap Project, and to build prediction models using flanking SNPs. Using this method, we have developed a database of prediction models for 611 known genes. Sequence prediction accuracy for these genes is 96.26% on average (ranges 79%-100%. This database of prediction models can be enhanced and scaled up to include new genes as the 1000 Genomes Project sequences additional genes on additional individuals. Applying our predictive model for the KCNJ11 gene to the Wellcome Trust Case Control Consortium (WTCCC Type 2 diabetes cohort, we demonstrate how the prediction of phased sequences inferred from GWAS SNP genotype data can be used to facilitate interpretation and identify a probable functional mechanism such as protein changes. Conclusions Prior to the general availability of routine sequencing of all subjects, the ISS method proposed here provides a time- and cost-effective approach to broadening the characterization of disease associated SNPs and regions, and facilitating the prioritization of candidate

  14. Shotgun protein sequencing.

    Energy Technology Data Exchange (ETDEWEB)

    Faulon, Jean-Loup Michel; Heffelfinger, Grant S.

    2009-06-01

    A novel experimental and computational technique based on multiple enzymatic digestion of a protein or protein mixture that reconstructs protein sequences from sequences of overlapping peptides is described in this SAND report. This approach, analogous to shotgun sequencing of DNA, is to be used to sequence alternative spliced proteins, to identify post-translational modifications, and to sequence genetically engineered proteins.

  15. The effects of non-synonymous single nucleotide polymorphisms (nsSNPs) on protein-protein interactions.

    Science.gov (United States)

    Yates, Christopher M; Sternberg, Michael J E

    2013-11-01

    Non-synonymous single nucleotide polymorphisms (nsSNPs) are single base changes leading to a change to the amino acid sequence of the encoded protein. Many of these variants are associated with disease, so nsSNPs have been well studied, with studies looking at the effects of nsSNPs on individual proteins, for example, on stability and enzyme active sites. In recent years, the impact of nsSNPs upon protein-protein interactions has also been investigated, giving a greater insight into the mechanisms by which nsSNPs can lead to disease. In this review, we summarize these studies, looking at the various mechanisms by which nsSNPs can affect protein-protein interactions. We focus on structural changes that can impair interaction, changes to disorder, gain of interaction, and post-translational modifications before looking at some examples of nsSNPs at human-pathogen protein-protein interfaces and the analysis of nsSNPs from a network perspective. © 2013.

  16. Single nucleotide polymorphism analysis of ubiquitin extension protein genes (ubq) of gossypium arboreum and gossypium herbaceum in comparison with arabidopsis thaliana

    International Nuclear Information System (INIS)

    Shaheen, T.; Zafar, Y.; Rahman, M.

    2014-01-01

    Single nucleotide polymorphism analysis is an expedient way to study polymorphisms at genomic level. In the present study we have explored Ubiquitin extension protein gene of G. arboreum (A2) and G. herbaceum (A1) of cotton which is a multiple copy gene. We have found SNPs at 16 positions in 200 bp region within A genome of cotton indicating frequency of SNPs 1/13 bp. Both sequences from cotton have shown maximum similarity with UBQ5 and UBQ6 of Arabidopsis thaliana. Sequence obtained from G. arboreum has shown SNPs at 28 positions in comparison with each UBQ5 and UBQ6 of Arabidopsis thaliana while sequence obtained from G. herbaceum has shown SNPs at 31 positions in comparison with each UBQ5 and UBQ6 of Arabidopsis thaliana. In conclusion although during pace of evolution ubiquitin extension protein genes of both A genome species have got some mutations from nature but still most of their sequence is similar. Single nucleotide polymorphism study can prove a vital tool to identify gene type in case of Multicopy genes. (author)

  17. Bovine proteins containing poly-glutamine repeats are often polymorphic and enriched for components of transcriptional regulatory complexes

    LENUS (Irish Health Repository)

    Whan, Vicki

    2010-11-23

    Abstract Background About forty human diseases are caused by repeat instability mutations. A distinct subset of these diseases is the result of extreme expansions of polymorphic trinucleotide repeats; typically CAG repeats encoding poly-glutamine (poly-Q) tracts in proteins. Polymorphic repeat length variation is also apparent in human poly-Q encoding genes from normal individuals. As these coding sequence repeats are subject to selection in mammals, it has been suggested that normal variations in some of these typically highly conserved genes are implicated in morphological differences between species and phenotypic variations within species. At present, poly-Q encoding genes in non-human mammalian species are poorly documented, as are their functions and propensities for polymorphic variation. Results The current investigation identified 178 bovine poly-Q encoding genes (Q ≥ 5) and within this group, 26 genes with orthologs in both human and mouse that did not contain poly-Q repeats. The bovine poly-Q encoding genes typically had ubiquitous expression patterns although there was bias towards expression in epithelia, brain and testes. They were also characterised by unusually large sizes. Analysis of gene ontology terms revealed that the encoded proteins were strongly enriched for functions associated with transcriptional regulation and many contributed to physical interaction networks in the nucleus where they presumably act cooperatively in transcriptional regulatory complexes. In addition, the coding sequence CAG repeats in some bovine genes impacted mRNA splicing thereby generating unusual transcriptional diversity, which in at least one instance was tissue-specific. The poly-Q encoding genes were prioritised using multiple criteria for their likelihood of being polymorphic and then the highest ranking group was experimentally tested for polymorphic variation within a cattle diversity panel. Extensive and meiotically stable variation was identified

  18. PSSRdb: a relational database of polymorphic simple sequence repeats extracted from prokaryotic genomes.

    Science.gov (United States)

    Kumar, Pankaj; Chaitanya, Pasumarthy S; Nagarajaram, Hampapathalu A

    2011-01-01

    PSSRdb (Polymorphic Simple Sequence Repeats database) (http://www.cdfd.org.in/PSSRdb/) is a relational database of polymorphic simple sequence repeats (PSSRs) extracted from 85 different species of prokaryotes. Simple sequence repeats (SSRs) are the tandem repeats of nucleotide motifs of the sizes 1-6 bp and are highly polymorphic. SSR mutations in and around coding regions affect transcription and translation of genes. Such changes underpin phase variations and antigenic variations seen in some bacteria. Although SSR-mediated phase variation and antigenic variations have been well-studied in some bacteria there seems a lot of other species of prokaryotes yet to be investigated for SSR mediated adaptive and other evolutionary advantages. As a part of our on-going studies on SSR polymorphism in prokaryotes we compared the genome sequences of various strains and isolates available for 85 different species of prokaryotes and extracted a number of SSRs showing length variations and created a relational database called PSSRdb. This database gives useful information such as location of PSSRs in genomes, length variation across genomes, the regions harboring PSSRs, etc. The information provided in this database is very useful for further research and analysis of SSRs in prokaryotes.

  19. Simple sequence repeats in Neurospora crassa: distribution, polymorphism and evolutionary inference

    Directory of Open Access Journals (Sweden)

    Park Jongsun

    2008-01-01

    Full Text Available Abstract Background Simple sequence repeats (SSRs have been successfully used for various genetic and evolutionary studies in eukaryotic systems. The eukaryotic model organism Neurospora crassa is an excellent system to study evolution and biological function of SSRs. Results We identified and characterized 2749 SSRs of 963 SSR types in the genome of N. crassa. The distribution of tri-nucleotide (nt SSRs, the most common SSRs in N. crassa, was significantly biased in exons. We further characterized the distribution of 19 abundant SSR types (AST, which account for 71% of total SSRs in the N. crassa genome, using a Poisson log-linear model. We also characterized the size variation of SSRs among natural accessions using Polymorphic Index Content (PIC and ANOVA analyses and found that there are genome-wide, chromosome-dependent and local-specific variations. Using polymorphic SSRs, we have built linkage maps from three line-cross populations. Conclusion Taking our computational, statistical and experimental data together, we conclude that 1 the distributions of the SSRs in the sequenced N. crassa genome differ systematically between chromosomes as well as between SSR types, 2 the size variation of tri-nt SSRs in exons might be an important mechanism in generating functional variation of proteins in N. crassa, 3 there are different levels of evolutionary forces in variation of amino acid repeats, and 4 SSRs are stable molecular markers for genetic studies in N. crassa.

  20. Genetic polymorphism of blood proteins in a population of shetland ponies

    NARCIS (Netherlands)

    Buis, R.C.

    1976-01-01

    Genetic variation of proteins (protein polymorphism) is widespread among many animal species. The biological significance of protein polymorphism has been the subject of many studies. This variation has a supporting function for population genetic studies as a source of genetic markers. In

  1. Do prion protein gene polymorphisms induce apoptosis in non ...

    Indian Academy of Sciences (India)

    2016-08-26

    Aug 26, 2016 ... Genetic variations such as single nucleotide polymorphisms (SNPs) in prion protein coding gene, Prnp, greatly affect susceptibility to prion diseases in mammals. Here, the coding region of Prnp was screened for polymorphisms in redeared turtle, Trachemys scripta. Four polymorphisms, L203V, N205I, ...

  2. Analysis of von Willebrand factor A domain-related protein (WARP polymorphism in temperate and tropical Plasmodium vivax field isolates

    Directory of Open Access Journals (Sweden)

    Zakeri Sedigheh

    2009-06-01

    Full Text Available Abstract Background The identification of key molecules is crucial for designing transmission-blocking vaccines (TBVs, among those ookinete micronemal proteins are candidate as a general class of malaria transmission-blocking targets. Here, the sequence analysis of an extra-cellular malaria protein expressed in ookinetes, named von Willebrand factor A domain-related protein (WARP, is reported in 91 Plasmodium vivax isolates circulating in different regions of Iran. Methods Clinical isolates were collected from north temperate and southern tropical regions in Iran. Primers have been designed based on P. vivax sequence (ctg_6991 which amplified a fragment of about 1044 bp with no size variation. Direct sequencing of PCR products was used to determine polymorphism and further bioinformatics analysis in P. vivax sexual stage antigen, pvwarp. Results Amplified pvwarp gene showed 886 bp in size, with no intron. BLAST analysis showed a similarity of 98–100% to P. vivax Sal-I strain; however, Iranian isolates had 2 bp mismatches in 247 and 531 positions that were non-synonymous substitution [T (ACT to A (GCT and R (AGA to S (AGT] in comparison with the Sal-I sequence. Conclusion This study presents the first large-scale survey on pvwarp polymorphism in the world, which provides baseline data for developing WARP-based TBV against both temperate and tropical P. vivax isolates.

  3. Salivary protein polymorphisms and risk of dental caries: a systematic review.

    Science.gov (United States)

    Lips, Andrea; Antunes, Leonardo Santos; Antunes, Lívia Azeredo; Pintor, Andrea Vaz Braga; Santos, Diana Amado Baptista Dos; Bachinski, Rober; Küchler, Erika Calvano; Alves, Gutemberg Gomes

    2017-06-05

    Dental caries is an oral pathology associated with both lifestyle and genetic factors. The caries process can be influenced by salivary composition, which includes ions and proteins. Studies have described associations between salivary protein polymorphisms and dental caries experience, while others have shown no association with salivary proteins genetic variability. The aim of this study is to assess the influence of salivary protein polymorphisms on the risk of dental caries by means of a systematic review of the current literature. An electronic search was performed in PubMed, Scopus, and Virtual Health Library. The following search terms were used: "dental caries susceptibility," "dental caries," "polymorphism, genetics," "saliva," "proteins," and "peptides." Related MeSH headings and free terms were included. The inclusion criteria comprised clinical investigations of subjects with and without caries. After application of these eligibility criteria, the selected articles were qualified by assessing their methodological quality. Initially, 338 articles were identified from the electronic databases after exclusion of duplicates. Exclusion criteria eliminated 322 articles, and 16 remained for evaluation. Eleven articles found a consistent association between salivary protein polymorphisms and risk of dental caries, for proteins related to antimicrobial activity (beta defensin 1 and lysozyme-like protein), pH control (carbonic anhydrase VI), and bacterial colonization/adhesion (lactotransferrin, mucin, and proline-rich protein Db). This systematic review demonstrated an association between genetic polymorphisms and risk of dental caries for most of the salivary proteins.

  4. Identification of a polymorphic collagen-like protein in the crustacean bacteria Pasteuria ramosa.

    Science.gov (United States)

    Mouton, Laurence; Traunecker, Emmanuel; McElroy, Kerensa; Du Pasquier, Louis; Ebert, Dieter

    2009-12-01

    Pasteuria ramosa is a spore-forming bacterium that infects Daphnia species. Previous results demonstrated a high specificity of host clone/parasite genotype interactions. Surface proteins of bacteria often play an important role in attachment to host cells prior to infection. We analyzed surface proteins of P. ramosa spores by two-dimensional gel electrophoresis. For the first time, we prove that two isolates selected for their differences in infectivity reveal few but clear-cut differences in protein patterns. Using internal sequencing and LC/MS/MS, we identified a collagen-like protein named Pcl1a (Pasteuria collagen-like protein 1a). This protein, reconstructed with the help of Pasteuria genome sequences, contains three domains: a 75-amino-acid amino-terminal domain with a potential transmembrane helix domain, a central collagen-like region (CLR) containing Gly-Xaa-Yaa (GXY) repeats, and a 7-amino-acid carboxy-terminal domain. The CLR region is polymorphic among the two isolates with amino-acid substitutions and a variable number of GXY triplets. Collagen-like proteins are rare in prokaryotes, although they have been described in several pathogenic bacteria, including Bacillus cereus, Bacillus anthracis and Bacillus thuringiensis, closely related to Pasteuria species, in which they could be involved in the adherence of bacteria to host cells.

  5. Bioinformatic Analysis of Chlamydia trachomatis Polymorphic Membrane Proteins PmpE, PmpF, PmpG and PmpH as Potential Vaccine Antigens.

    Directory of Open Access Journals (Sweden)

    Alexandra Nunes

    Full Text Available Chlamydia trachomatis is the most important infectious cause of infertility in women with important implications in public health and for which a vaccine is urgently needed. Recent immunoproteomic vaccine studies found that four polymorphic membrane proteins (PmpE, PmpF, PmpG and PmpH are immunodominant, recognized by various MHC class II haplotypes and protective in mouse models. In the present study, we aimed to evaluate genetic and protein features of Pmps (focusing on the N-terminal 600 amino acids where MHC class II epitopes were mapped in order to understand antigen variation that may emerge following vaccine induced immune selection. We used several bioinformatics platforms to study: i Pmps' phylogeny and genetic polymorphism; ii the location and distribution of protein features (GGA(I, L/FxxN motifs and cysteine residues that may impact pathogen-host interactions and protein conformation; and iii the existence of phase variation mechanisms that may impact Pmps' expression. We used a well-characterized collection of 53 fully-sequenced strains that represent the C. trachomatis serovars associated with the three disease groups: ocular (N=8, epithelial-genital (N=25 and lymphogranuloma venereum (LGV (N=20. We observed that PmpF and PmpE are highly polymorphic between LGV and epithelial-genital strains, and also within populations of the latter. We also found heterogeneous representation among strains for GGA(I, L/FxxN motifs and cysteine residues, suggesting possible alterations in adhesion properties, tissue specificity and immunogenicity. PmpG and, to a lesser extent, PmpH revealed low polymorphism and high conservation of protein features among the genital strains (including the LGV group. Uniquely among the four Pmps, pmpG has regulatory sequences suggestive of phase variation. In aggregate, the results suggest that PmpG may be the lead vaccine candidate because of sequence conservation but may need to be paired with another protective

  6. The V279F polymorphism might change protein character and ...

    African Journals Online (AJOL)

    Polymorphisms at the protein level are also estimated to correlate with increased risk factors for heart attacks. One such polymorphism is the V279F polymorphism in Lp-PLA2 which results in a change in enzyme performance capability. This in turn implies a reduced risk of acute myocardial infarct (AMI) in Korean and ...

  7. Sequence polymorphisms in Pvs48/45 and Pvs47 gametocyte and gamete surface proteins in Plasmodium vivax isolated in Korea.

    Science.gov (United States)

    Woo, Mi Kyung; Kim, Kyeong Ah; Kim, JuYeon; Oh, Jun Seo; Han, Eun Taek; An, Seong Soo A; Lim, Chae Seung

    2013-05-01

    Nucleotide sequence analyses of the Pvs48/45 and Pvs47 genes were conducted in 46 malaria patients from the Republic of Korea (ROK) (n = 40) and returning travellers from India (n = 3) and Indonesia (n = 3). The domain structures, which were based on cysteine residue position and secondary protein structure, were similar between Plasmodium vivax (Pvs48/45 and Pvs47) and Plasmodium falciparum (Pfs48/45 and Pfs47). In comparison to the Sal-1 reference strain (Pvs48/45, PVX_083235 and Pvs47, PVX_083240), Korean isolates revealed seven polymorphisms (E35K, H211N, K250N, D335Y, A376T, I380T and K418R) in Pvs48/45. These isolates could be divided into five haplotypes with the two major types having frequencies of 47.5% and 20%, respectivelfy. In Pvs47, 10 polymorphisms (F22L, F24L, K27E, D31N, V230I, M233I, E240D, I262T, I273M and A373V) were found and they could be divided into four haplotypes with one major type having a frequency of 75%. The Pvs48/45 isolates from India showed a unique amino acid substitution site (K26R). Compared to the Sal-1 and ROK isolates, the Pvs47 isolates from travellers returning from India and Indonesia had amino acid substitutions (S57T and I262K). The current data may contribute to the development of the malaria transmission-blocking vaccine in future clinical trials.

  8. Molecular mechanisms for protein-encoded inheritance

    Science.gov (United States)

    Wiltzius, Jed J. W.; Landau, Meytal; Nelson, Rebecca; Sawaya, Michael R.; Apostol, Marcin I.; Goldschmidt, Lukasz; Soriaga, Angela B.; Cascio, Duilio; Rajashankar, Kanagalaghatta; Eisenberg, David

    2013-01-01

    Strains are phenotypic variants, encoded by nucleic acid sequences in chromosomal inheritance and by protein “conformations” in prion inheritance and transmission. But how is a protein “conformation” stable enough to endure transmission between cells or organisms? Here new polymorphic crystal structures of segments of prion and other amyloid proteins offer structural mechanisms for prion strains. In packing polymorphism, prion strains are encoded by alternative packings (polymorphs) of β-sheets formed by the same segment of a protein; in a second mechanism, segmental polymorphism, prion strains are encoded by distinct β-sheets built from different segments of a protein. Both forms of polymorphism can produce enduring “conformations,” capable of encoding strains. These molecular mechanisms for transfer of information into prion strains share features with the familiar mechanism for transfer of information by nucleic acid inheritance, including sequence specificity and recognition by non-covalent bonds. PMID:19684598

  9. Partial Sequence Analysis of Merozoite Surface Proteine-3α Gene in Plasmodium vivax Isolates from Malarious Areas of Iran

    Directory of Open Access Journals (Sweden)

    H Mirhendi

    2008-12-01

    Full Text Available Background: Approximately 85-90% of malaria infections in Iran are attributed to Plasmodium vivax, while little is known about the genetic of the parasite and its strain types in this region. This study was designed and performed for describing genetic characteristics of Plasmodium vivax population of Iran based on the merozoite surface protein-3α gene sequence. Methods: Through a descriptive study we analyzed partial P. vivax merozoite surface protein-3α gene sequences from 17 clinical P. vivax isolates collected from malarious areas of Iran. Genomic DNA was extracted by Q1Aamp® DNA blood mini kit, amplified through nested PCR for a partial nucleotide sequence of PvMSP-3 gene in P. vivax. PCR-amplified products were sequenced with an ABI Prism Perkin-Elmer 310 sequencer machine and the data were analyzed with clustal W software. Results: Analysis of PvMSP-3 gene sequences demonstrated extensive polymorphisms, but the sequence identity between isolates of same types was relatively high. We identified specific insertions and deletions for the types A, B and C variants of P. vivax in our isolates. In phylogenetic comparison of geographically separated isolates, there was not a significant geo­graphical branching of the parasite populations. Conclusion: The highly polymorphic nature of isolates suggests that more investigations of the PvMSP-3 gene are needed to explore its vaccine potential.

  10. Protein sequence comparison and protein evolution

    Energy Technology Data Exchange (ETDEWEB)

    Pearson, W.R. [Univ. of Virginia, Charlottesville, VA (United States). Dept. of Biochemistry

    1995-12-31

    This tutorial was one of eight tutorials selected to be presented at the Third International Conference on Intelligent Systems for Molecular Biology which was held in the United Kingdom from July 16 to 19, 1995. This tutorial examines how the information conserved during the evolution of a protein molecule can be used to infer reliably homology, and thus a shared proteinfold and possibly a shared active site or function. The authors start by reviewing a geological/evolutionary time scale. Next they look at the evolution of several protein families. During the tutorial, these families will be used to demonstrate that homologous protein ancestry can be inferred with confidence. They also examine different modes of protein evolution and consider some hypotheses that have been presented to explain the very earliest events in protein evolution. The next part of the tutorial will examine the technical aspects of protein sequence comparison. Both optimal and heuristic algorithms and their associated parameters that are used to characterize protein sequence similarities are discussed. Perhaps more importantly, they survey the statistics of local similarity scores, and how these statistics can both be used to improve the selectivity of a search and to evaluate the significance of a match. They them examine distantly related members of three protein families, the serine proteases, the glutathione transferases, and the G-protein-coupled receptors (GCRs). Finally, the discuss how sequence similarity can be used to examine internal repeated or mosaic structures in proteins.

  11. Genetic polymorphism and natural selection of Duffy binding protein of Plasmodium vivax Myanmar isolates

    Science.gov (United States)

    2012-01-01

    Background Plasmodium vivax Duffy binding protein (PvDBP) plays an essential role in erythrocyte invasion and a potential asexual blood stage vaccine candidate antigen against P. vivax. The polymorphic nature of PvDBP, particularly amino terminal cysteine-rich region (PvDBPII), represents a major impediment to the successful design of a protective vaccine against vivax malaria. In this study, the genetic polymorphism and natural selection at PvDBPII among Myanmar P. vivax isolates were analysed. Methods Fifty-four P. vivax infected blood samples collected from patients in Myanmar were used. The region flanking PvDBPII was amplified by PCR, cloned into Escherichia coli, and sequenced. The polymorphic characters and natural selection of the region were analysed using the DnaSP and MEGA4 programs. Results Thirty-two point mutations (28 non-synonymous and four synonymous mutations) were identified in PvDBPII among the Myanmar P. vivax isolates. Sequence analyses revealed that 12 different PvDBPII haplotypes were identified in Myanmar P. vivax isolates and that the region has evolved under positive natural selection. High selective pressure preferentially acted on regions identified as B- and T-cell epitopes of PvDBPII. Recombination may also be played a role in the resulting genetic diversity of PvDBPII. Conclusions PvDBPII of Myanmar P. vivax isolates displays a high level of genetic polymorphism and is under selective pressure. Myanmar P. vivax isolates share distinct types of PvDBPII alleles that are different from those of other geographical areas. These results will be useful for understanding the nature of the P. vivax population in Myanmar and for development of PvDBPII-based vaccine. PMID:22380592

  12. Sequence polymorphisms in Pvs48/45 and Pvs47 gametocyte and gamete surface proteins in Plasmodium vivax isolated in Korea

    Directory of Open Access Journals (Sweden)

    Mi Kyung Woo

    2013-05-01

    Full Text Available Nucleotide sequence analyses of the Pvs48/45 and Pvs47 genes were conducted in 46 malaria patients from the Republic of Korea (ROK (n = 40 and returning travellers from India (n = 3 and Indonesia (n = 3. The domain structures, which were based on cysteine residue position and secondary protein structure, were similar between Plasmodium vivax (Pvs48/45 and Pvs47 and Plasmodium falciparum (Pfs48/45 and Pfs47. In comparison to the Sal-1 reference strain (Pvs48/45, PVX_083235 and Pvs47, PVX_083240, Korean isolates revealed seven polymorphisms (E35K, H211N, K250N, D335Y, A376T, I380T and K418R in Pvs48/45. These isolates could be divided into five haplotypes with the two major types having frequencies of 47.5% and 20%, respectivelfy. In Pvs47, 10 polymorphisms (F22L, F24L, K27E, D31N, V230I, M233I, E240D, I262T, I273M and A373V were found and they could be divided into four haplotypes with one major type having a frequency of 75%. The Pvs48/45 isolates from India showed a unique amino acid substitution site (K26R. Compared to the Sal-1 and ROK isolates, the Pvs47 isolates from travellers returning from India and Indonesia had amino acid substitutions (S57T and I262K. The current data may contribute to the development of the malaria transmission-blocking vaccine in future clinical trials.

  13. The Large Subunit rDNA Sequence of Plasmodiophora brassicae Does not Contain Intra-species Polymorphism.

    Science.gov (United States)

    Schwelm, Arne; Berney, Cédric; Dixelius, Christina; Bass, David; Neuhauser, Sigrid

    2016-12-01

    Clubroot disease caused by Plasmodiophora brassicae is one of the most important diseases of cultivated brassicas. P. brassicae occurs in pathotypes which differ in the aggressiveness towards their Brassica host plants. To date no DNA based method to distinguish these pathotypes has been described. In 2011 polymorphism within the 28S rDNA of P. brassicae was reported which potentially could allow to distinguish pathotypes without the need of time-consuming bioassays. However, isolates of P. brassicae from around the world analysed in this study do not show polymorphism in their LSU rDNA sequences. The previously described polymorphism most likely derived from soil inhabiting Cercozoa more specifically Neoheteromita-like glissomonads. Here we correct the LSU rDNA sequence of P. brassicae. By using FISH we demonstrate that our newly generated sequence belongs to the causal agent of clubroot disease. Copyright © 2016 The Authors. Published by Elsevier GmbH.. All rights reserved.

  14. Systematic examination of polymorphism in amyloid fibrils by molecular-dynamics simulation.

    Science.gov (United States)

    Berryman, Joshua T; Radford, Sheena E; Harris, Sarah A

    2011-05-04

    Amyloid fibrils often exhibit polymorphism. Polymorphs are formed when proteins or peptides with identical sequences self-assemble into fibrils containing substantially different arrangements of the β-strands. We used atomistic molecular-dynamics simulation to examine the thermodynamic stability of a amyloid fibrils in different polymorphic forms by performing a systematic investigation of sequence and symmetry space for a series of peptides with a range of physicochemical properties. We show that the stability of fibrils depends on both sequence and the symmetry because these factors determine the availability of favorable interactions between the peptide strands within a sheet and in intersheet packing. By performing a detailed analysis of these interactions as a function of symmetry, we obtained a series of simple design rules that can be used to determine which polymorphs of a given sequence are most likely to form thermodynamically stable fibrils. These rules can potentially be employed to design peptide sequences that aggregate into a preferred polymorphic form for nanotechnological purposes. Copyright © 2011 Biophysical Society. Published by Elsevier Inc. All rights reserved.

  15. HIV protein sequence hotspots for crosstalk with host hub proteins.

    Directory of Open Access Journals (Sweden)

    Mahdi Sarmady

    Full Text Available HIV proteins target host hub proteins for transient binding interactions. The presence of viral proteins in the infected cell results in out-competition of host proteins in their interaction with hub proteins, drastically affecting cell physiology. Functional genomics and interactome datasets can be used to quantify the sequence hotspots on the HIV proteome mediating interactions with host hub proteins. In this study, we used the HIV and human interactome databases to identify HIV targeted host hub proteins and their host binding partners (H2. We developed a high throughput computational procedure utilizing motif discovery algorithms on sets of protein sequences, including sequences of HIV and H2 proteins. We identified as HIV sequence hotspots those linear motifs that are highly conserved on HIV sequences and at the same time have a statistically enriched presence on the sequences of H2 proteins. The HIV protein motifs discovered in this study are expressed by subsets of H2 host proteins potentially outcompeted by HIV proteins. A large subset of these motifs is involved in cleavage, nuclear localization, phosphorylation, and transcription factor binding events. Many such motifs are clustered on an HIV sequence in the form of hotspots. The sequential positions of these hotspots are consistent with the curated literature on phenotype altering residue mutations, as well as with existing binding site data. The hotspot map produced in this study is the first global portrayal of HIV motifs involved in altering the host protein network at highly connected hub nodes.

  16. Molecular basis of the apolipoprotein H (beta 2-glycoprotein I) protein polymorphism

    DEFF Research Database (Denmark)

    Sanghera, Dharambir K; Kristensen, Torsten; Hamman, Richard F

    1997-01-01

    Apolipoprotein H (apoH, protein; APOH, gene) is considered to be an essential cofactor for the binding of certain antiphospholipid autoantibodies to anionic phospholipids. APOH exhibits a genetically determined structural polymorphism due to the presence of three common alleles (APOH*1, APOH*2...... was observed sporadically in blacks (0.008), it was present at a polymorphic frequency in Hispanics (0.027) and non-Hispanic whites (0.059). The identification of the molecular basis of the APOH protein polymorphism will help to elucidate the structural – functional relationship of apoH in the production...

  17. [Sequence polymorphism of mtDNA HVR Iand HVR II of Oroqen ethnic group in Inner Mongolia].

    Science.gov (United States)

    Yan, Chun-Xia; Chen, Feng; Dang, Yong-Hui; Li, Tao; Zheng, Hai-Bo; Chen, Teng; Li, Sheng-Bin

    2008-04-01

    Venous blood samples from 50 unrelated Oroqen individuals living in Inner Mongolia were collected and their mtDNA HVR I and HVR II sequences were detected by using ABI PRISM377 sequencers. The number of polymorphic loci, haplotype, haplotype frequence, average nucleotide variability and other polymorphic parameters were calculated. Based on Oroqen mtDNA sequence data obtained in our experiments and published data, genetic distance between Oroqen ethnic group and other populations were computered by Nei's measure. Phylogenetic tree was constructed by Neighbor Joining method. Comparing with Anderson sequence, 52 polymorphic loci in HVR I and 24 loci in HVR II were found in Oroqen mtDNA sequence, 38 and 27 haplotypes were defined herewith. Haplotype diversity and average nucleotide variability were 0.964+/-0.018 and 7.379 in HVR I, 0.929+/-0.019 and 2.408 in HVR II respectively. Fst and dA genetic distance between 12 populations were calculated based on HVR I sequence, and their relative coefficients were 0.993(P HVR I and HVR II in Oroqen ethnic group has some specificities compared with that of other populations. These data provide a useful tool in forensic identification, population genetic study and other research fields.

  18. Male accessory gland secretory protein polymorphism in natural ...

    Indian Academy of Sciences (India)

    [Ravi Ram K. and Ramesh S. R. 2007 Male accessory gland secretory protein polymorphism in natural ..... quence of species-specific genetic responses to variations in .... Eberhard W. G. 1996 Female control: sexual selection by cryptic.

  19. IMPACT OF MANNOSE-BINDING PROTEIN GENE POLYMORPHISMS IN OMANI SICKLE CELL DISEASE PATIENTS

    Directory of Open Access Journals (Sweden)

    Mathew Zachariah

    2016-02-01

    Full Text Available Objectives: Our objective was to study mannose binding protein (MBP polymorphisms in exonic and promoter region and correlate associated infections and vasoocculsive (VOC episodes, since MBP plays an important role in innate immunity by activating the complement system. Methods: We studied the genetic polymorphisms in the Exon 1 (alleles A/O and promoter region (alleles Y/X; H/L, P/Q of the MBL2 gene, in sickle cell disease (SCD patients as increased incidence of infections is seen in these patients. A PCR-based, targeted genomic DNA sequencing of MBL2 was used to study 68 SCD Omani patients and 44 controls (voluntary blood donors. Results: The observed frequencies of MBL2 promoter polymorphism (-221, Y/X were 44.4% and 20.5% for the heterozygous genotype Y/X and 3.2% and 2.2% for the homozygous (X/X respectively between SCD patients and controls. MBL2 Exon1 gene mutations were 29.4% and 50% for the heterozygous genotype A/O and 5.9% and 6.8% respectively for the homozygous (O/O genotype between SCD patients and controls. The distribution of variant MBL2 polymorphisms did not show any correlation in SCD patients with or without vasoocculsive crisis (VOC attacks (p=0.162; OR-0.486; CI=0.177 -1.33, however, it was correlated with infections (p=0.0162; OR-3.55; CI 1.25-10.04. Conclusions: Although the frequency of the genotypes and haplotypes of MBL2 in SCD patients did not differ from controls, overall in the SCD patient cohort the increased representation of variant alleles was significantly correlated with infections (p<0.05. However, these variant MBL2 polymorphisms did not seem to play a significant role in the VOC episodes in this SCD cohort. Keywords: Mannose-binding lectin, polymorphism, promoter, Sickle cell disease, MBL2, MBP

  20. Characterisation of a large family of polymorphic collagen-like proteins in the endospore-forming bacterium Pasteuria ramosa.

    Science.gov (United States)

    McElroy, Kerensa; Mouton, Laurence; Du Pasquier, Louis; Qi, Weihong; Ebert, Dieter

    2011-09-01

    Collagen-like proteins containing glycine-X-Y repeats have been identified in several pathogenic bacteria potentially involved in virulence. Recently, a collagen-like surface protein, Pcl1a, was identified in Pasteuria ramosa, a spore-forming parasite of Daphnia. Here we characterise 37 novel putative P. ramosa collagen-like protein genes (PCLs). PCR amplification and sequencing across 10 P. ramosa strains showed they were polymorphic, distinguishing genotypes matching known differences in Daphnia/P. ramosa interaction specificity. Thirty PCLs could be divided into four groups based on sequence similarity, conserved N- and C-terminal regions and G-X-Y repeat structure. Group 1, Group 2 and Group 3 PCLs formed triplets within the genome, with one member from each group represented in each triplet. Maximum-likelihood trees suggested that these groups arose through multiple instances of triplet duplication. For Group 1, 2, 3 and 4 PCLs, X was typically proline and Y typically threonine, consistent with other bacterial collagen-like proteins. The amino acid composition of Pcl2 closely resembled Pcl1a, with X typically being glutamic acid or aspartic acid and Y typically being lysine or glutamine. Pcl2 also showed sequence similarity to Pcl1a and contained a predicted signal peptide, cleavage site and transmembrane domain, suggesting that it is a surface protein. Copyright © 2011 Institut Pasteur. Published by Elsevier Masson SAS. All rights reserved.

  1. Generating markers based on biotic stress of protein system in and tandem repeats sequence for Aquilaria sp

    International Nuclear Information System (INIS)

    Azhar Mohamad; Muhammad Hanif Azhari N; Siti Norhayati Ismail

    2014-01-01

    Aquilaria sp. belongs to the Thymelaeaceae family and is well distributed in Asia region. The species has multipurpose use from root to shoot and is an economically important crop, which generates wide interest in understanding genetic diversity of the species. Knowledge on DNA-based markers has become a prerequisite for more effective application of molecular marker techniques in breeding and mapping programs. In this work, both targeted genes and tandem repeat sequences were used for DNA fingerprinting in Aquilaria sp. A total of 100 ISSR (inter simple sequence repeat) primers and 50 combination pairs of specific primers derived from conserved region of a specific protein known as system in were optimized. 38 ISSR primers were found affirmative for polymorphism evaluation study and were generated from both specific and degenerate ISSR primers. And one utmost combination of system in primers showed significant results in distinguishing the Aquilaria sp. In conclusion, polymorphism derived from ISSR profiling and targeted stress genes of protein system in proved as a powerful approach for identification and molecular classification of Aquilaria sp. which will be useful for diversification in identifying any mutant lines derived from nature. (author)

  2. Pregnancy-associated plasma protein A gene polymorphism in pregnant women with preeclampsia and intrauterine growth restriction

    Directory of Open Access Journals (Sweden)

    Sultan Ozkan

    2015-10-01

    Full Text Available Preeclampsia (PE and intrauterine growth restriction (IUGR are still among the most commonly researched titles in perinatology. To shed light on their etiology, new prevention and treatment strategies are the major targets of studies. In this study, we aimed to investigate the relation between gene polymorphism of one of the products of trophoblasts, pregnancy-associated plasma protein A (PAPP-A and PE/IUGR.A total of 147 women (IUGR, n = 61; PE, n = 47; IUGR + PE, n = 37; eclampsia, n = 2 were compared with 103 controls with respect to the sequencing of exon 14 of the PAPP-A gene to detect (rs7020782 polymorphism. Genotypes “AA” and “CC” were given in the event of A or C allele homozygosity and “AC” in A and C allele heterozygosity. Our findings revealed that the rate of AA, CC homozygotes, and AC heterozygotes did not differ between groups. Moreover, there was no difference in the distribution of PAPP-A genotypes among the patients with IUGR, PE, IUGR + PE, or eclampsia. Finally, birth weight, rate of the presence of proteinuria, and total protein excretion on 24-hour urine were similar in the subgroups of AA, AC, and CC genotypes in the study group. Our study demonstrated no association between PAPP-A gene rs7020782 polymorphism and PE/IUGR.

  3. Pregnancy-associated plasma protein A gene polymorphism in pregnant women with preeclampsia and intrauterine growth restriction.

    Science.gov (United States)

    Ozkan, Sultan; Sanhal, Cem Yasar; Yeniel, Ozgur; Arslan Ates, Esra; Ergenoglu, Mete; Bınbır, Birol; Onay, Huseyin; Ozkınay, Ferda; Sagol, Sermet

    2015-10-01

    Preeclampsia (PE) and intrauterine growth restriction (IUGR) are still among the most commonly researched titles in perinatology. To shed light on their etiology, new prevention and treatment strategies are the major targets of studies. In this study, we aimed to investigate the relation between gene polymorphism of one of the products of trophoblasts, pregnancy-associated plasma protein A (PAPP-A) and PE/IUGR.A total of 147 women (IUGR, n = 61; PE, n = 47; IUGR + PE, n = 37; eclampsia, n = 2) were compared with 103 controls with respect to the sequencing of exon 14 of the PAPP-A gene to detect (rs7020782) polymorphism. Genotypes "AA" and "CC" were given in the event of A or C allele homozygosity and "AC" in A and C allele heterozygosity. Our findings revealed that the rate of AA, CC homozygotes, and AC heterozygotes did not differ between groups. Moreover, there was no difference in the distribution of PAPP-A genotypes among the patients with IUGR, PE, IUGR + PE, or eclampsia. Finally, birth weight, rate of the presence of proteinuria, and total protein excretion on 24-hour urine were similar in the subgroups of AA, AC, and CC genotypes in the study group. Our study demonstrated no association between PAPP-A gene rs7020782 polymorphism and PE/IUGR. Copyright © 2015. Published by Elsevier Taiwan.

  4. Distribution and linkage disequilibrium analysis of polymorphisms of ...

    Indian Academy of Sciences (India)

    2Five-Star Animal Health Pharmaceutical Factory of Jilian Province, 5333 Xi'an Road, Changchun, Jilin 130062,. People's .... showed that GH peptide sequence of mature protein in. Rongjiang pig ..... 1, white. (PIC value <0.25, low polymorphism; 0.25 protein and energy meta-.

  5. Polymorphisms of the prion protein gene Arabi sheep breed in Iran ...

    African Journals Online (AJOL)

    Ovine scrapie is a neurodegenerative disease caused by polymorphisms of the prion protein gene (Prnp); especially the amino acid residue alterations at codons 136, 154, and 174, in sheep have been found to be associated with susceptibility to scrapie disease. We studied Prnp polymorphisms in local sheep of ...

  6. Induced protein polymorphisms and nutritional quality of gamma irradiation mutants of sorghum

    Energy Technology Data Exchange (ETDEWEB)

    Mehlo, Luke, E-mail: LMehlo@csir.co.za [CSIR Biosciences, Meiring Naude Road, P.O. Box 395, Pretoria 0001 (South Africa); Mbambo, Zodwa [CSIR Biosciences, Meiring Naude Road, P.O. Box 395, Pretoria 0001 (South Africa); Microbiology Discipline, School of Life Sciences, University of KwaZulu-Natal (Westville Campus), Private Bag X54001, Durban 4000 (South Africa); Bado, Souleymane [Plant Breeding and Genetics Laboratory – Joint FAO/IAEA Agriculture and Biotechnology Laboratory, International Atomic Energy Agency Laboratories, A-2444 Seibersdorf (Austria); Lin, Johnson [Microbiology Discipline, School of Life Sciences, University of KwaZulu-Natal (Westville Campus), Private Bag X54001, Durban 4000 (South Africa); Moagi, Sydwell M.; Buthelezi, Sindisiwe; Stoychev, Stoyan; Chikwamba, Rachel [CSIR Biosciences, Meiring Naude Road, P.O. Box 395, Pretoria 0001 (South Africa)

    2013-09-15

    Highlights: • We analyse kafirin protein polymorphisms induced by gamma irradiation in sorghum. • One mutant with suppressed kafirins in the endosperm accumulated them in the germ. • Kafirin polymorphisms were associated with high levels of free amino acids. • Nutritional value of sorghum can be improved significantly by induced mutations. - Abstract: Physical and biochemical analysis of protein polymorphisms in seed storage proteins of a mutant population of sorghum revealed a mutant with redirected accumulation of kafirin proteins in the germ. The change in storage proteins was accompanied by an unusually high level accumulation of free lysine and other essential amino acids in the endosperm. This mutant further displayed a significant suppression in the synthesis and accumulation of the 27 kDa γ-, 24 kDa α-A1 and the 22 kDa α-A2 kafirins in the endosperm. The suppression of kafirins was counteracted by an upsurge in the synthesis and accumulation of albumins, globulins and other proteins. The data collectively suggest that sorghum has huge genetic potential for nutritional biofortification and that induced mutations can be used as an effective tool in achieving premium nutrition in staple cereals.

  7. Induced protein polymorphisms and nutritional quality of gamma irradiation mutants of sorghum

    International Nuclear Information System (INIS)

    Mehlo, Luke; Mbambo, Zodwa; Bado, Souleymane; Lin, Johnson; Moagi, Sydwell M.; Buthelezi, Sindisiwe; Stoychev, Stoyan; Chikwamba, Rachel

    2013-01-01

    Highlights: • We analyse kafirin protein polymorphisms induced by gamma irradiation in sorghum. • One mutant with suppressed kafirins in the endosperm accumulated them in the germ. • Kafirin polymorphisms were associated with high levels of free amino acids. • Nutritional value of sorghum can be improved significantly by induced mutations. - Abstract: Physical and biochemical analysis of protein polymorphisms in seed storage proteins of a mutant population of sorghum revealed a mutant with redirected accumulation of kafirin proteins in the germ. The change in storage proteins was accompanied by an unusually high level accumulation of free lysine and other essential amino acids in the endosperm. This mutant further displayed a significant suppression in the synthesis and accumulation of the 27 kDa γ-, 24 kDa α-A1 and the 22 kDa α-A2 kafirins in the endosperm. The suppression of kafirins was counteracted by an upsurge in the synthesis and accumulation of albumins, globulins and other proteins. The data collectively suggest that sorghum has huge genetic potential for nutritional biofortification and that induced mutations can be used as an effective tool in achieving premium nutrition in staple cereals

  8. [Sequence polymorphisms of the mitochondrial DNA HVR I and HVR II regions in the Deng populations from Tibet in China].

    Science.gov (United States)

    Kang, Longli; Zhang, Xiaofeng; Liu, Kai; Zhao, Jianmin

    2009-12-01

    To analyze the sequence polymorphisms of the mitochondrial DNA hypervariable regions I (HVR I) and HVR II in the Deng population in Linzhi area of Tibet. mtDNAs obtained from 119 unrelated individuals were amplified and directly sequenced. One hundred and ten variable sites were identified, including nucleotide transitions, transversions, and insertions. In the HVR I region (nt16024-nt16365), 68 polymorphic sites and 119 haplotypes were observed, the genetic diversity was 0.9916. In the HVR II (nt73-nt340) region, 42 polymorphic sites and 113 haplotypes were observed, and the genetic diversity was 0.9907. The random match probability of the HVR I and HVR II regions were 0.0084 and 0.0093, respectively. When combining the HVR I and HVR II regions, 119 different haplotypes were found. The combined match probability of two unrelated persons having the same sequence was 0.0084. There are some unique polymorphic loci in the Deng population. There are different genetic structures between Chinese and other Asian populations in the mitochondrial DNA D-loop region. Sequence polymorphism of mitochondrial DNA HVR I and HVR II can be used as a genetic marker for forensic individual identification and genetic analysis.

  9. Novel algorithms for protein sequence analysis

    NARCIS (Netherlands)

    Ye, Kai

    2008-01-01

    Each protein is characterized by its unique sequential order of amino acids, the so-called protein sequence. Biology”s paradigm is that this order of amino acids determines the protein”s architecture and function. In this thesis, we introduce novel algorithms to analyze protein sequences. Chapter 1

  10. Polymorphisms in promoter sequences of MDM2, p53, and p16INK4a genes in normal Japanese individuals

    Directory of Open Access Journals (Sweden)

    Yasuhito Ohsaka

    2010-01-01

    Full Text Available Research has been conducted to identify sequence polymorphisms of gene promoter regions in patients and control subjects, including normal individuals, and to determine the influence of these polymorphisms on transcriptional regulation in cells that express wild-type or mutant p53. In this study we isolated genomic DNA from whole blood of healthy Japanese individuals and sequenced the promoter regions of the MDM2, p53, and p16INK4a genes. We identified polymorphisms comprising 3 nucleotide substitutions at exon 1 and intron 1 regions of the MDM2 gene and 1 nucleotide insertion at a poly(C nucleotide position in the p53 gene. The Japanese individuals also exhibited p16INK4a polymorphisms at several positions, including position -191. Reporter gene analysis by using luciferase revealed that the polymorphisms of MDM2, p53, and p16INK4a differentially altered luciferase activities in several cell lines, including the Colo320DM, U251, and T98G cell lines expressing mutant p53. Our results indicate that the promoter sequences of these genes differ among normal Japanese individuals and that polymorphisms can alter gene transcription activity.

  11. Comprehensive search for intra- and inter-specific sequence polymorphisms among coding envelope genes of retroviral origin found in the human genome: genes and pseudogenes

    Directory of Open Access Journals (Sweden)

    Vasilescu Alexandre

    2005-09-01

    Full Text Available Abstract Background The human genome carries a high load of proviral-like sequences, called Human Endogenous Retroviruses (HERVs, which are the genomic traces of ancient infections by active retroviruses. These elements are in most cases defective, but open reading frames can still be found for the retroviral envelope gene, with sixteen such genes identified so far. Several of them are conserved during primate evolution, having possibly been co-opted by their host for a physiological role. Results To characterize further their status, we presently sequenced 12 of these genes from a panel of 91 Caucasian individuals. Genomic analyses reveal strong sequence conservation (only two non synonymous Single Nucleotide Polymorphisms [SNPs] for the two HERV-W and HERV-FRD envelope genes, i.e. for the two genes specifically expressed in the placenta and possibly involved in syncytiotrophoblast formation. We further show – using an ex vivo fusion assay for each allelic form – that none of these SNPs impairs the fusogenic function. The other envelope proteins disclose variable polymorphisms, with the occurrence of a stop codon and/or frameshift for most – but not all – of them. Moreover, the sequence conservation analysis of the orthologous genes that can be found in primates shows that three env genes have been maintained in a fully coding state throughout evolution including envW and envFRD. Conclusion Altogether, the present study strongly suggests that some but not all envelope encoding sequences are bona fide genes. It also provides new tools to elucidate the possible role of endogenous envelope proteins as susceptibility factors in a number of pathologies where HERVs have been suspected to be involved.

  12. Nanomechanical properties of distinct fibrillar polymorphs of the protein α-synuclein

    Science.gov (United States)

    Makky, Ali; Bousset, Luc; Polesel-Maris, Jérôme; Melki, Ronald

    2016-11-01

    Alpha-synuclein (α-Syn) is a small presynaptic protein of 140 amino acids. Its pathologic intracellular aggregation within the central nervous system yields protein fibrillar inclusions named Lewy bodies that are the hallmarks of Parkinson’s disease (PD). In solution, pure α-Syn adopts an intrinsically disordered structure and assembles into fibrils that exhibit considerable morphological heterogeneity depending on their assembly conditions. We recently established tightly controlled experimental conditions allowing the assembly of α-Syn into highly homogeneous and pure polymorphs. The latter exhibited differences in their shape, their structure but also in their functional properties. We have conducted an AFM study at high resolution and performed a statistical analysis of fibrillar α-Syn shape and thermal fluctuations to calculate the persistence length to further assess the nanomechanical properties of α-Syn polymorphs. Herein, we demonstrated quantitatively that distinct polymorphs made of the same protein (wild-type α-Syn) show significant differences in their morphology (height, width and periodicity) and physical properties (persistence length, bending rigidity and axial Young’s modulus).

  13. Analysis of polymorphisms in milk proteins from cloned and sexually reproduced goats.

    Science.gov (United States)

    Xing, H; Shao, B; Gu, Y Y; Yuan, Y G; Zhang, T; Zang, J; Cheng, Y

    2015-12-08

    This study evaluates the relationship between the genotype and milk protein components in goats. Milk samples were collected from cloned goats and normal white goats during different postpartum (or abortion) phases. Two cloned goats, originated from the same somatic line of goat mammary gland epithelial cells, and three sexually reproduced normal white goats with no genetic relationships were used as the control. The goats were phylogenetically analyzed by polymerase chain reaction-restriction fragment length polymorphism. The milk protein components were identified by sodium dodecyl sulfate polyacrylamide gel electrophoresis. The results indicated that despite the genetic fingerprints being identical, the milk protein composition differed between the two cloned goats. The casein content of cloned goat C-50 was significantly higher than that of cloned goat C-4. Conversely, although the genetic fingerprints of the normal white goats N-1, N-2, and N-3 were not identical, the milk protein profiles did not differ significantly in their milk samples (obtained on postpartum day 15, 20, 25, 30, and 150). These results indicated an association between milk protein phenotypes and genetic polymorphisms, epigenetic regulation, and/or non-chromosomal factors. This study extends the knowledge of goat milk protein polymorphisms, and provides new strategies for the breeding of high milk-yielding goats.

  14. Sequence based polymorphic (SBP marker technology for targeted genomic regions: its application in generating a molecular map of the Arabidopsis thaliana genome

    Directory of Open Access Journals (Sweden)

    Sahu Binod B

    2012-01-01

    Full Text Available Abstract Background Molecular markers facilitate both genotype identification, essential for modern animal and plant breeding, and the isolation of genes based on their map positions. Advancements in sequencing technology have made possible the identification of single nucleotide polymorphisms (SNPs for any genomic regions. Here a sequence based polymorphic (SBP marker technology for generating molecular markers for targeted genomic regions in Arabidopsis is described. Results A ~3X genome coverage sequence of the Arabidopsis thaliana ecotype, Niederzenz (Nd-0 was obtained by applying Illumina's sequencing by synthesis (Solexa technology. Comparison of the Nd-0 genome sequence with the assembled Columbia-0 (Col-0 genome sequence identified putative single nucleotide polymorphisms (SNPs throughout the entire genome. Multiple 75 base pair Nd-0 sequence reads containing SNPs and originating from individual genomic DNA molecules were the basis for developing co-dominant SBP markers. SNPs containing Col-0 sequences, supported by transcript sequences or sequences from multiple BAC clones, were compared to the respective Nd-0 sequences to identify possible restriction endonuclease enzyme site variations. Small amplicons, PCR amplified from both ecotypes, were digested with suitable restriction enzymes and resolved on a gel to reveal the sequence based polymorphisms. By applying this technology, 21 SBP markers for the marker poor regions of the Arabidopsis map representing polymorphisms between Col-0 and Nd-0 ecotypes were generated. Conclusions The SBP marker technology described here allowed the development of molecular markers for targeted genomic regions of Arabidopsis. It should facilitate isolation of co-dominant molecular markers for targeted genomic regions of any animal or plant species, whose genomic sequences have been assembled. This technology will particularly facilitate the development of high density molecular marker maps, essential for

  15. Polymorphisms in the prion protein gene and in the doppel gene increase susceptibility for Creutzfeldt-Jakob disease

    NARCIS (Netherlands)

    E.A. Croes (Esther); B.Z. Alizadeh (Behrooz); A.M. Bertoli Avella (Aida); T.A.M. Rademaker (Tessa); J. Vergeer-Drop (Jeannette); B. Dermaut (Bart); J.J. Houwing-Duistermaat (Jeanine); D.P.W.M. Wientjens (Dorothee); A. Hofman (Albert); C. van Broeckhoven (Christine); C.M. van Duijn (Cornelia)

    2004-01-01

    textabstractThe prion protein gene (PRNP) plays a central role in the origin of Creutzfeldt-Jakob disease (CJD), but there is growing interest in other polymorphisms that may be involved in CJD. Polymorphisms upstream of PRNP that may modulate the prion protein production as well as polymorphisms in

  16. GenProBiS: web server for mapping of sequence variants to protein binding sites.

    Science.gov (United States)

    Konc, Janez; Skrlj, Blaz; Erzen, Nika; Kunej, Tanja; Janezic, Dusanka

    2017-07-03

    Discovery of potentially deleterious sequence variants is important and has wide implications for research and generation of new hypotheses in human and veterinary medicine, and drug discovery. The GenProBiS web server maps sequence variants to protein structures from the Protein Data Bank (PDB), and further to protein-protein, protein-nucleic acid, protein-compound, and protein-metal ion binding sites. The concept of a protein-compound binding site is understood in the broadest sense, which includes glycosylation and other post-translational modification sites. Binding sites were defined by local structural comparisons of whole protein structures using the Protein Binding Sites (ProBiS) algorithm and transposition of ligands from the similar binding sites found to the query protein using the ProBiS-ligands approach with new improvements introduced in GenProBiS. Binding site surfaces were generated as three-dimensional grids encompassing the space occupied by predicted ligands. The server allows intuitive visual exploration of comprehensively mapped variants, such as human somatic mis-sense mutations related to cancer and non-synonymous single nucleotide polymorphisms from 21 species, within the predicted binding sites regions for about 80 000 PDB protein structures using fast WebGL graphics. The GenProBiS web server is open and free to all users at http://genprobis.insilab.org. © The Author(s) 2017. Published by Oxford University Press on behalf of Nucleic Acids Research.

  17. Detection of Ribosomal DNA Sequence Polymorphisms in the Protist Plasmodiophora brassicae for the Identification of Geographical Isolates

    Directory of Open Access Journals (Sweden)

    Rawnak Laila

    2017-01-01

    Full Text Available Clubroot is a soil-borne disease caused by the protist Plasmodiophora brassicae (P. brassicae. It is one of the most economically important diseases of Brassica rapa and other cruciferous crops as it can cause remarkable yield reductions. Understanding P. brassicae genetics, and developing efficient molecular markers, is essential for effective detection of harmful races of this pathogen. Samples from 11 Korean field populations of P. brassicae (geographic isolates, collected from nine different locations in South Korea, were used in this study. Genomic DNA was extracted from the clubroot-infected samples to sequence the ribosomal DNA. Primers and probes for P. brassicae were designed using a ribosomal DNA gene sequence from a Japanese strain available in GenBank (accession number AB526843; isolate NGY. The nuclear ribosomal DNA (rDNA sequence of P. brassicae, comprising 6932 base pairs (bp, was cloned and sequenced and found to include the small subunits (SSUs and a large subunit (LSU, internal transcribed spacers (ITS1 and ITS2, and a 5.8s. Sequence variation was observed in both the SSU and LSU. Four markers showed useful differences in high-resolution melting analysis to identify nucleotide polymorphisms including single- nucleotide polymorphisms (SNPs, oligonucleotide polymorphisms, and insertions/deletions (InDels. A combination of three markers was able to distinguish the geographical isolates into two groups.

  18. Processing of Chlamydia abortus polymorphic membrane protein 18D during the chlamydial developmental cycle.

    Directory of Open Access Journals (Sweden)

    Nick M Wheelhouse

    Full Text Available Chlamydia possess a unique family of autotransporter proteins known as the Polymorphic membrane proteins (Pmps. While the total number of pmp genes varies between Chlamydia species, all encode a single pmpD gene. In both Chlamydia trachomatis (C. trachomatis and C. pneumoniae, the PmpD protein is proteolytically cleaved on the cell surface. The current study was carried out to determine the cleavage patterns of the PmpD protein in the animal pathogen C. abortus (termed Pmp18D.Using antibodies directed against different regions of Pmp18D, proteomic techniques revealed that the mature protein was cleaved on the cell surface, resulting in a100 kDa N-terminal product and a 60 kDa carboxy-terminal protein. The N-terminal protein was further processed into 84, 76 and 73 kDa products. Clustering analysis resolved PmpD proteins into three distinct clades with C. abortus Pmp18D, being most similar to those originating from C. psittaci, C. felis and C. caviae.This study indicates that C. abortus Pmp18D is proteolytically processed at the cell surface similar to the proteins of C. trachomatis and C. pneumoniae. However, patterns of cleavage are species-specific, with low sequence conservation of PmpD across the genus. The absence of conserved domains indicates that the function of the PmpD molecule in chlamydia remains to be elucidated.

  19. Processing of Chlamydia abortus polymorphic membrane protein 18D during the chlamydial developmental cycle.

    Science.gov (United States)

    Wheelhouse, Nick M; Sait, Michelle; Aitchison, Kevin; Livingstone, Morag; Wright, Frank; McLean, Kevin; Inglis, Neil F; Smith, David G E; Longbottom, David

    2012-01-01

    Chlamydia possess a unique family of autotransporter proteins known as the Polymorphic membrane proteins (Pmps). While the total number of pmp genes varies between Chlamydia species, all encode a single pmpD gene. In both Chlamydia trachomatis (C. trachomatis) and C. pneumoniae, the PmpD protein is proteolytically cleaved on the cell surface. The current study was carried out to determine the cleavage patterns of the PmpD protein in the animal pathogen C. abortus (termed Pmp18D). Using antibodies directed against different regions of Pmp18D, proteomic techniques revealed that the mature protein was cleaved on the cell surface, resulting in a100 kDa N-terminal product and a 60 kDa carboxy-terminal protein. The N-terminal protein was further processed into 84, 76 and 73 kDa products. Clustering analysis resolved PmpD proteins into three distinct clades with C. abortus Pmp18D, being most similar to those originating from C. psittaci, C. felis and C. caviae. This study indicates that C. abortus Pmp18D is proteolytically processed at the cell surface similar to the proteins of C. trachomatis and C. pneumoniae. However, patterns of cleavage are species-specific, with low sequence conservation of PmpD across the genus. The absence of conserved domains indicates that the function of the PmpD molecule in chlamydia remains to be elucidated.

  20. Development of cleaved amplified polymorphic sequence markers and a CAPS-based genetic linkage map in watermelon (Citrullus lanatus [Thunb.] Matsum. and Nakai) constructed using whole-genome re-sequencing data.

    Science.gov (United States)

    Liu, Shi; Gao, Peng; Zhu, Qianglong; Luan, Feishi; Davis, Angela R; Wang, Xiaolu

    2016-03-01

    Cleaved amplified polymorphic sequence (CAPS) markers are useful tools for detecting single nucleotide polymorphisms (SNPs). This study detected and converted SNP sites into CAPS markers based on high-throughput re-sequencing data in watermelon, for linkage map construction and quantitative trait locus (QTL) analysis. Two inbred lines, Cream of Saskatchewan (COS) and LSW-177 had been re-sequenced and analyzed by Perl self-compiled script for CAPS marker development. 88.7% and 78.5% of the assembled sequences of the two parental materials could map to the reference watermelon genome, respectively. Comparative assembled genome data analysis provided 225,693 and 19,268 SNPs and indels between the two materials. 532 pairs of CAPS markers were designed with 16 restriction enzymes, among which 271 pairs of primers gave distinct bands of the expected length and polymorphic bands, via PCR and enzyme digestion, with a polymorphic rate of 50.94%. Using the new CAPS markers, an initial CAPS-based genetic linkage map was constructed with the F2 population, spanning 1836.51 cM with 11 linkage groups and 301 markers. 12 QTLs were detected related to fruit flesh color, length, width, shape index, and brix content. These newly CAPS markers will be a valuable resource for breeding programs and genetic studies of watermelon.

  1. Identification of new polymorphic regions and differentiation of cultivated olives (Olea europaea L.) through plastome sequence comparison

    Science.gov (United States)

    2010-01-01

    Background The cultivated olive (Olea europaea L.) is the most agriculturally important species of the Oleaceae family. Although many studies have been performed on plastid polymorphisms to evaluate taxonomy, phylogeny and phylogeography of Olea subspecies, only few polymorphic regions discriminating among the agronomically and economically important olive cultivars have been identified. The objective of this study was to sequence the entire plastome of olive and analyze many potential polymorphic regions to develop new inter-cultivar genetic markers. Results The complete plastid genome of the olive cultivar Frantoio was determined by direct sequence analysis using universal and novel PCR primers designed to amplify all overlapping regions. The chloroplast genome of the olive has an organisation and gene order that is conserved among numerous Angiosperm species and do not contain any of the inversions, gene duplications, insertions, inverted repeat expansions and gene/intron losses that have been found in the chloroplast genomes of the genera Jasminum and Menodora, from the same family as Olea. The annotated sequence was used to evaluate the content of coding genes, the extent, and distribution of repeated and long dispersed sequences and the nucleotide composition pattern. These analyses provided essential information for structural, functional and comparative genomic studies in olive plastids. Furthermore, the alignment of the olive plastome sequence to those of other varieties and species identified 30 new organellar polymorphisms within the cultivated olive. Conclusions In addition to identifying mutations that may play a functional role in modifying the metabolism and adaptation of olive cultivars, the new chloroplast markers represent a valuable tool to assess the level of olive intercultivar plastome variation for use in population genetic analysis, phylogenesis, cultivar characterisation and DNA food tracking. PMID:20868482

  2. Use of designed sequences in protein structure recognition.

    Science.gov (United States)

    Kumar, Gayatri; Mudgal, Richa; Srinivasan, Narayanaswamy; Sandhya, Sankaran

    2018-05-09

    Knowledge of the protein structure is a pre-requisite for improved understanding of molecular function. The gap in the sequence-structure space has increased in the post-genomic era. Grouping related protein sequences into families can aid in narrowing the gap. In the Pfam database, structure description is provided for part or full-length proteins of 7726 families. For the remaining 52% of the families, information on 3-D structure is not yet available. We use the computationally designed sequences that are intermediately related to two protein domain families, which are already known to share the same fold. These strategically designed sequences enable detection of distant relationships and here, we have employed them for the purpose of structure recognition of protein families of yet unknown structure. We first measured the success rate of our approach using a dataset of protein families of known fold and achieved a success rate of 88%. Next, for 1392 families of yet unknown structure, we made structural assignments for part/full length of the proteins. Fold association for 423 domains of unknown function (DUFs) are provided as a step towards functional annotation. The results indicate that knowledge-based filling of gaps in protein sequence space is a lucrative approach for structure recognition. Such sequences assist in traversal through protein sequence space and effectively function as 'linkers', where natural linkers between distant proteins are unavailable. This article was reviewed by Oliviero Carugo, Christine Orengo and Srikrishna Subramanian.

  3. Genetic polymorphism of horse serum protein 3 (SP3).

    Science.gov (United States)

    Juneja, R K; Sandberg, K; Kuryl, J; Gahne, B

    1989-01-01

    Two-dimensional agarose gel (pH 8.6)-horizontal polyacrylamide gel (pH 9.0) electrophoresis of horse serum samples, followed by general protein staining, revealed genetic polymorphism of an unidentified protein tentatively designated serum protein 3 (SP3). The SP3 fractions appeared distinctly when a 14% concentration of acrylamide was used in the separation gels. The 2-D mobilities of SP3 fractions were quite similar to that of albumin. Family data were consistent with the hypothesis that the observed SP3 phenotypes were controlled by four co-dominant, autosomal alleles (D, F, I, S). Evidence was provided that the F allele can be further divided into two alleles (F1 and F2); the mobilities of F1 and F2 variants were very similar. Each of the SP3 alleles gave rise to one fraction and each of the heterozygous types showed two fractions. More than 600 horses representing five different breeds (Swedish Trotter, North-Swedish Trotter, Thoroughbred, Arab and Polish Tarpan) were typed for SP3, and allele frequency estimates were calculated. SP3 was highly polymorphic in all breeds studied.

  4. Overlapping genomic sequences: a treasure trove of single-nucleotide polymorphisms.

    Science.gov (United States)

    Taillon-Miller, P; Gu, Z; Li, Q; Hillier, L; Kwok, P Y

    1998-07-01

    An efficient strategy to develop a dense set of single-nucleotide polymorphism (SNP) markers is to take advantage of the human genome sequencing effort currently under way. Our approach is based on the fact that bacterial artificial chromosomes (BACs) and P1-based artificial chromosomes (PACs) used in long-range sequencing projects come from diploid libraries. If the overlapping clones sequenced are from different lineages, one is comparing the sequences from 2 homologous chromosomes in the overlapping region. We have analyzed in detail every SNP identified while sequencing three sets of overlapping clones found on chromosome 5p15.2, 7q21-7q22, and 13q12-13q13. In the 200.6 kb of DNA sequence analyzed in these overlaps, 153 SNPs were identified. Computer analysis for repetitive elements and suitability for STS development yielded 44 STSs containing 68 SNPs for further study. All 68 SNPs were confirmed to be present in at least one of the three (Caucasian, African-American, Hispanic) populations studied. Furthermore, 42 of the SNPs tested (62%) were informative in at least one population, 32 (47%) were informative in two or more populations, and 23 (34%) were informative in all three populations. These results clearly indicate that developing SNP markers from overlapping genomic sequence is highly efficient and cost effective, requiring only the two simple steps of developing STSs around the known SNPs and characterizing them in the appropriate populations.

  5. Improving accuracy of protein-protein interaction prediction by considering the converse problem for sequence representation

    Directory of Open Access Journals (Sweden)

    Wang Yong

    2011-10-01

    Full Text Available Abstract Background With the development of genome-sequencing technologies, protein sequences are readily obtained by translating the measured mRNAs. Therefore predicting protein-protein interactions from the sequences is of great demand. The reason lies in the fact that identifying protein-protein interactions is becoming a bottleneck for eventually understanding the functions of proteins, especially for those organisms barely characterized. Although a few methods have been proposed, the converse problem, if the features used extract sufficient and unbiased information from protein sequences, is almost untouched. Results In this study, we interrogate this problem theoretically by an optimization scheme. Motivated by the theoretical investigation, we find novel encoding methods for both protein sequences and protein pairs. Our new methods exploit sufficiently the information of protein sequences and reduce artificial bias and computational cost. Thus, it significantly outperforms the available methods regarding sensitivity, specificity, precision, and recall with cross-validation evaluation and reaches ~80% and ~90% accuracy in Escherichia coli and Saccharomyces cerevisiae respectively. Our findings here hold important implication for other sequence-based prediction tasks because representation of biological sequence is always the first step in computational biology. Conclusions By considering the converse problem, we propose new representation methods for both protein sequences and protein pairs. The results show that our method significantly improves the accuracy of protein-protein interaction predictions.

  6. Mapping of Micro-Tom BAC-End Sequences to the Reference Tomato Genome Reveals Possible Genome Rearrangements and Polymorphisms

    Science.gov (United States)

    Asamizu, Erika; Shirasawa, Kenta; Hirakawa, Hideki; Sato, Shusei; Tabata, Satoshi; Yano, Kentaro; Ariizumi, Tohru; Shibata, Daisuke; Ezura, Hiroshi

    2012-01-01

    A total of 93,682 BAC-end sequences (BESs) were generated from a dwarf model tomato, cv. Micro-Tom. After removing repetitive sequences, the BESs were similarity searched against the reference tomato genome of a standard cultivar, “Heinz 1706.” By referring to the “Heinz 1706” physical map and by eliminating redundant or nonsignificant hits, 28,804 “unique pair ends” and 8,263 “unique ends” were selected to construct hypothetical BAC contigs. The total physical length of the BAC contigs was 495, 833, 423 bp, covering 65.3% of the entire genome. The average coverage of euchromatin and heterochromatin was 58.9% and 67.3%, respectively. From this analysis, two possible genome rearrangements were identified: one in chromosome 2 (inversion) and the other in chromosome 3 (inversion and translocation). Polymorphisms (SNPs and Indels) between the two cultivars were identified from the BLAST alignments. As a result, 171,792 polymorphisms were mapped on 12 chromosomes. Among these, 30,930 polymorphisms were found in euchromatin (1 per 3,565 bp) and 140,862 were found in heterochromatin (1 per 2,737 bp). The average polymorphism density in the genome was 1 polymorphism per 2,886 bp. To facilitate the use of these data in Micro-Tom research, the BAC contig and polymorphism information are available in the TOMATOMICS database. PMID:23227037

  7. Uncoupling protein 2 G(-866A polymorphism: a new gene polymorphism associated with C-reactive protein in type 2 diabetic patients C-reactive protein in type 2 diabetic patients

    Directory of Open Access Journals (Sweden)

    Cocozza Sergio

    2010-10-01

    Full Text Available Abstract Background This study evaluated the relationship between the G(-866A polymorphism of the uncoupling protein 2 (UCP2 gene and high-sensitivity C reactive protein (hs-CRP plasma levels in diabetic patients. Methods We studied 383 unrelated people with type 2 diabetes aged 40-70 years. Anthropometry, fasting lipids, glucose, HbA1c, and hs-CRP were measured. Participants were genotyped for the G (-866A polymorphism of the uncoupling protein 2 gene. Results Hs-CRP (mg/L increased progressively across the three genotype groups AA, AG, or GG, being respectively 3.0 ± 3.2, 3.6 ± 5.0, and 4.8 ± 5.3 (p for trend = 0.03. Since hs-CRP values were not significantly different between AA and AG genotype, these two groups were pooled for further analyses. Compared to participants with the AA/AG genotypes, homozygotes for the G allele (GG genotype had significantly higher hs-CRP levels (4.8 ± 5.3 vs 3.5 ± 4.7 mg/L, p = 0.01 and a larger proportion (53.9% vs 46.1%, p = 0.013 of elevated hs-CRP (> 2 mg/L. This was not explained by major confounders such as age, gender, BMI, waist circumference, HbA1c, smoking, or medications use which were comparable in the two genotype groups. Conclusions The study shows for the first time, in type 2 diabetic patients, a significant association of hs-CRP levels with the G(-866A polymorphism of UCP2 beyond the effect of major confounders.

  8. Proximal Region of the Gene Encoding Cytadherence-Related Protein Permits Molecular Typing of Mycoplasma genitalium Clinical Strains by PCR-Restriction Fragment Length Polymorphism

    Science.gov (United States)

    Musatovova, Oxana; Herrera, Caleb; Baseman, Joel B.

    2006-01-01

    Restriction fragment length polymorphism (RFLP) analysis of the PCR-amplified proximal region of the gene encoding cytadherence accessory protein P110 (MG192) revealed DNA sequence divergences among 54 Mycoplasma genitalium clinical strains isolated from the genitourinary tracts of women attending a sexually transmitted disease-related health clinic, plus one from the respiratory tract and one from synovial fluid. Seven of 56 (12.5%) strains exhibited RFLPs following digestion of the proximal region with restriction endonuclease MboI or RsaI, or both. No sequence variability was detected in the distal portion of the gene. PMID:16455921

  9. THE POLYMORPHISM OF THE SUS4 SUCROSE SYNTHASE DOMAIN SEQUENCES IN RUSSIAN, BELORUSSIAN AND KAZAKH POTATO CULTIVARS

    Directory of Open Access Journals (Sweden)

    M. A. Slugina

    2016-01-01

    Full Text Available The potato is one of the main strategic crops in the Russian Federation, Belarus and Kazakhstan. Currently, we have achieved significant advances in the understanding of metabolic mechanism of carbohydrate and interconversion «sucrose – starch» in potato tubers. Sucrose synthase (Sus is a key enzyme in the breakdown of sucrose. Sucrose synthase (Sus is catalyzing a reversible reaction of conversion sucrose and UDP into fructose and UDP-glucose. The identification and subsequent characterization of the genes encoding plant sucrose synthase is the first step towards understanding their physiological roles and metabolic mechanism involved in carbohydrate accumulation in potato tubers. In the present work the nucleotide and amino acid polymorphism of the Sus4 gene fragments containing sequences of the sucrose synthase domain were analyzed. Sus4 gene fragments (intron III – exon VI in 9 potato cultivars of Russian, Kazakh and Belarusian breeding were analyzed. The polymorphism of the Sus4 sucrose synthase domain sequences was first examined. The length of analyzed fragment varied from 977 b.p. (cultivars Favorit, Karasaiskii, Miras to 1013 b.p. (cultivars Zorochka, Manifest, Elisaveta, Bashkirskii. It was demonstrated that the examined sequences contained point mutations, as well as insertions and deletions. The common polymorphism level was 5.82%. It was shown that the examined sequences contained 58 SNPs and 4 indels. The most variable were introns IV (12.4% and V (9.18%. The most variable was exon IV. 7 allelic variants were detected. 6 different amino acid sequences specific to different varieties were also identified.

  10. Unraveling the sequence-dependent polymorphic behavior of d(CpG) steps in B-DNA.

    Science.gov (United States)

    Dans, Pablo Daniel; Faustino, Ignacio; Battistini, Federica; Zakrzewska, Krystyna; Lavery, Richard; Orozco, Modesto

    2014-10-01

    We have made a detailed study of one of the most surprising sources of polymorphism in B-DNA: the high twist/low twist (HT/LT) conformational change in the d(CpG) base pair step. Using extensive computations, complemented with database analysis, we were able to characterize the twist polymorphism in the d(CpG) step in all the possible tetranucleotide environment. We found that twist polymorphism is coupled with BI/BII transitions, and, quite surprisingly, with slide polymorphism in the neighboring step. Unexpectedly, the penetration of cations into the minor groove of the d(CpG) step seems to be the key element in promoting twist transitions. The tetranucleotide environment also plays an important role in the sequence-dependent d(CpG) polymorphism. In this connection, we have detected a previously unexplored intramolecular C-H···O hydrogen bond interaction that stabilizes the low twist state when 3'-purines flank the d(CpG) step. This work explains a coupled mechanism involving several apparently uncorrelated conformational transitions that has only been partially inferred by earlier experimental or theoretical studies. Our results provide a complete description of twist polymorphism in d(CpG) steps and a detailed picture of the molecular choreography associated with this conformational change. © The Author(s) 2014. Published by Oxford University Press on behalf of Nucleic Acids Research.

  11. ‘‘Blind'' mapping of genic DNA sequence polymorphisms in Lolium perenne L. by high resolution melting curve analysis

    DEFF Research Database (Denmark)

    Studer, Bruno; Jensen, Louise Bach; Fiil, Alice

    2009-01-01

    High resolution melting curve analysis (HRM) measures dissociation of double stranded DNA of a PCR product amplified in the presence of a saturating fluorescence dye. Recently, HRM proved successful to genotype DNA sequence polymorphisms such as SSRs and SNPs based on the shape of the melting...... curves. In this study, HRM was used for simultaneous screening and genotyping of genic DNA sequence polymorphisms identified in the Lolium perenne F2 mapping population VrnA. Melting profiles of PCR products amplified from previously published gene loci and from a novel gene putatively involved...

  12. Nonlinear deterministic structures and the randomness of protein sequences

    CERN Document Server

    Huang Yan Zhao

    2003-01-01

    To clarify the randomness of protein sequences, we make a detailed analysis of a set of typical protein sequences representing each structural classes by using nonlinear prediction method. No deterministic structures are found in these protein sequences and this implies that they behave as random sequences. We also give an explanation to the controversial results obtained in previous investigations.

  13. Single-nucleotide polymorphism discovery by high-throughput sequencing in sorghum

    Directory of Open Access Journals (Sweden)

    White Frank F

    2011-07-01

    Full Text Available Abstract Background Eight diverse sorghum (Sorghum bicolor L. Moench accessions were subjected to short-read genome sequencing to characterize the distribution of single-nucleotide polymorphisms (SNPs. Two strategies were used for DNA library preparation. Missing SNP genotype data were imputed by local haplotype comparison. The effect of library type and genomic diversity on SNP discovery and imputation are evaluated. Results Alignment of eight genome equivalents (6 Gb to the public reference genome revealed 283,000 SNPs at ≥82% confirmation probability. Sequencing from libraries constructed to limit sequencing to start at defined restriction sites led to genotyping 10-fold more SNPs in all 8 accessions, and correctly imputing 11% more missing data, than from semirandom libraries. The SNP yield advantage of the reduced-representation method was less than expected, since up to one fifth of reads started at noncanonical restriction sites and up to one third of restriction sites predicted in silico to yield unique alignments were not sampled at near-saturation. For imputation accuracy, the availability of a genomically similar accession in the germplasm panel was more important than panel size or sequencing coverage. Conclusions A sequence quantity of 3 million 50-base reads per accession using a BsrFI library would conservatively provide satisfactory genotyping of 96,000 sorghum SNPs. For most reliable SNP-genotype imputation in shallowly sequenced genomes, germplasm panels should consist of pairs or groups of genomically similar entries. These results may help in designing strategies for economical genotyping-by-sequencing of large numbers of plant accessions.

  14. Nucleation phenomena in protein folding: the modulating role of protein sequence

    International Nuclear Information System (INIS)

    Travasso, Rui D M; FaIsca, Patricia F N; Gama, Margarida M Telo da

    2007-01-01

    For the vast majority of naturally occurring, small, single-domain proteins, folding is often described as a two-state process that lacks detectable intermediates. This observation has often been rationalized on the basis of a nucleation mechanism for protein folding whose basic premise is the idea that, after completion of a specific set of contacts forming the so-called folding nucleus, the native state is achieved promptly. Here we propose a methodology to identify folding nuclei in small lattice polymers and apply it to the study of protein molecules with a chain length of N = 48. To investigate the extent to which protein topology is a robust determinant of the nucleation mechanism, we compare the nucleation scenario of a native-centric model with that of a sequence-specific model sharing the same native fold. To evaluate the impact of the sequence's finer details in the nucleation mechanism, we consider the folding of two non-homologous sequences. We conclude that, in a sequence-specific model, the folding nucleus is, to some extent, formed by the most stable contacts in the protein and that the less stable linkages in the folding nucleus are solely determined by the fold's topology. We have also found that, independently of the protein sequence, the folding nucleus performs the same 'topological' function. This unifying feature of the nucleation mechanism results from the residues forming the folding nucleus being distributed along the protein chain in a similar and well-defined manner that is determined by the fold's topological features

  15. WildSpan: mining structured motifs from protein sequences

    Directory of Open Access Journals (Sweden)

    Chen Chien-Yu

    2011-03-01

    Full Text Available Abstract Background Automatic extraction of motifs from biological sequences is an important research problem in study of molecular biology. For proteins, it is desired to discover sequence motifs containing a large number of wildcard symbols, as the residues associated with functional sites are usually largely separated in sequences. Discovering such patterns is time-consuming because abundant combinations exist when long gaps (a gap consists of one or more successive wildcards are considered. Mining algorithms often employ constraints to narrow down the search space in order to increase efficiency. However, improper constraint models might degrade the sensitivity and specificity of the motifs discovered by computational methods. We previously proposed a new constraint model to handle large wildcard regions for discovering functional motifs of proteins. The patterns that satisfy the proposed constraint model are called W-patterns. A W-pattern is a structured motif that groups motif symbols into pattern blocks interleaved with large irregular gaps. Considering large gaps reflects the fact that functional residues are not always from a single region of protein sequences, and restricting motif symbols into clusters corresponds to the observation that short motifs are frequently present within protein families. To efficiently discover W-patterns for large-scale sequence annotation and function prediction, this paper first formally introduces the problem to solve and proposes an algorithm named WildSpan (sequential pattern mining across large wildcard regions that incorporates several pruning strategies to largely reduce the mining cost. Results WildSpan is shown to efficiently find W-patterns containing conserved residues that are far separated in sequences. We conducted experiments with two mining strategies, protein-based and family-based mining, to evaluate the usefulness of W-patterns and performance of WildSpan. The protein-based mining mode

  16. LDsplit: screening for cis-regulatory motifs stimulating meiotic recombination hotspots by analysis of DNA sequence polymorphisms.

    Science.gov (United States)

    Yang, Peng; Wu, Min; Guo, Jing; Kwoh, Chee Keong; Przytycka, Teresa M; Zheng, Jie

    2014-02-17

    As a fundamental genomic element, meiotic recombination hotspot plays important roles in life sciences. Thus uncovering its regulatory mechanisms has broad impact on biomedical research. Despite the recent identification of the zinc finger protein PRDM9 and its 13-mer binding motif as major regulators for meiotic recombination hotspots, other regulators remain to be discovered. Existing methods for finding DNA sequence motifs of recombination hotspots often rely on the enrichment of co-localizations between hotspots and short DNA patterns, which ignore the cross-individual variation of recombination rates and sequence polymorphisms in the population. Our objective in this paper is to capture signals encoded in genetic variations for the discovery of recombination-associated DNA motifs. Recently, an algorithm called "LDsplit" has been designed to detect the association between single nucleotide polymorphisms (SNPs) and proximal meiotic recombination hotspots. The association is measured by the difference of population recombination rates at a hotspot between two alleles of a candidate SNP. Here we present an open source software tool of LDsplit, with integrative data visualization for recombination hotspots and their proximal SNPs. Applying LDsplit on SNPs inside an established 7-mer motif bound by PRDM9 we observed that SNP alleles preserving the original motif tend to have higher recombination rates than the opposite alleles that disrupt the motif. Running on SNP windows around hotspots each containing an occurrence of the 7-mer motif, LDsplit is able to guide the established motif finding algorithm of MEME to recover the 7-mer motif. In contrast, without LDsplit the 7-mer motif could not be identified. LDsplit is a software tool for the discovery of cis-regulatory DNA sequence motifs stimulating meiotic recombination hotspots by screening and narrowing down to hotspot associated SNPs. It is the first computational method that utilizes the genetic variation of

  17. HomPPI: a class of sequence homology based protein-protein interface prediction methods

    Directory of Open Access Journals (Sweden)

    Dobbs Drena

    2011-06-01

    Full Text Available Abstract Background Although homology-based methods are among the most widely used methods for predicting the structure and function of proteins, the question as to whether interface sequence conservation can be effectively exploited in predicting protein-protein interfaces has been a subject of debate. Results We studied more than 300,000 pair-wise alignments of protein sequences from structurally characterized protein complexes, including both obligate and transient complexes. We identified sequence similarity criteria required for accurate homology-based inference of interface residues in a query protein sequence. Based on these analyses, we developed HomPPI, a class of sequence homology-based methods for predicting protein-protein interface residues. We present two variants of HomPPI: (i NPS-HomPPI (Non partner-specific HomPPI, which can be used to predict interface residues of a query protein in the absence of knowledge of the interaction partner; and (ii PS-HomPPI (Partner-specific HomPPI, which can be used to predict the interface residues of a query protein with a specific target protein. Our experiments on a benchmark dataset of obligate homodimeric complexes show that NPS-HomPPI can reliably predict protein-protein interface residues in a given protein, with an average correlation coefficient (CC of 0.76, sensitivity of 0.83, and specificity of 0.78, when sequence homologs of the query protein can be reliably identified. NPS-HomPPI also reliably predicts the interface residues of intrinsically disordered proteins. Our experiments suggest that NPS-HomPPI is competitive with several state-of-the-art interface prediction servers including those that exploit the structure of the query proteins. The partner-specific classifier, PS-HomPPI can, on a large dataset of transient complexes, predict the interface residues of a query protein with a specific target, with a CC of 0.65, sensitivity of 0.69, and specificity of 0.70, when homologs of

  18. A DEL phenotype attributed to RHD Exon 9 sequence deletion: slipped-strand mispairing and blood group polymorphisms.

    Science.gov (United States)

    Lopez, Genghis H; Turner, Robyn M; McGowan, Eunike C; Schoeman, Elizna M; Scott, Stacy A; O'Brien, Helen; Millard, Glenda M; Roulis, Eileen V; Allen, Amanda J; Liew, Yew-Wah; Flower, Robert L; Hyland, Catherine A

    2018-03-01

    The RhD blood group antigen is extremely polymorphic and the DEL phenotype represents one such class of polymorphisms. The DEL phenotype prevalent in East Asian populations arises from a synonymous substitution defined as RHD*1227A. However, initially, based on genomic and cDNA studies, the genetic basis for a DEL phenotype in Taiwan was attributed to a deletion of RHD Exon 9 that was never verified at the genomic level by any other independent group. Here we investigate the genetic basis for a Caucasian donor with a DEL partial D phenotype and compare the genomic findings to those initial molecular studies. The 3'-region of the RHD gene was amplified by long-range polymerase chain reaction (PCR) for massively parallel sequencing. Primers were designed to encompass a deletion, flanking Exon 9, by standard PCR for Sanger sequencing. Targeted sequencing of exons and flanking introns was also performed. Genomic DNA exhibited a 1012-bp deletion spanning from Intron 8, across Exon 9 into Intron 9. The deletion breakpoints occurred between two 25-bp repeat motifs flanking Exon 9 such that one repeat sequence remained. Deletion mutations bordered by repeat sequences are a hallmark of slipped-strand mispairing (SSM) event. We propose this genetic mechanism generated the germline deletion in the Caucasian donor. Extensive studies show that the RHD*1227A is the most prevalent DEL allele in East Asian populations and may have confounded the initial molecular studies. Review of the literature revealed that the SSM model explains some of the extreme polymorphisms observed in the clinically significant RhD blood group antigen. © 2017 AABB.

  19. Protein 3D structure computed from evolutionary sequence variation.

    Directory of Open Access Journals (Sweden)

    Debora S Marks

    Full Text Available The evolutionary trajectory of a protein through sequence space is constrained by its function. Collections of sequence homologs record the outcomes of millions of evolutionary experiments in which the protein evolves according to these constraints. Deciphering the evolutionary record held in these sequences and exploiting it for predictive and engineering purposes presents a formidable challenge. The potential benefit of solving this challenge is amplified by the advent of inexpensive high-throughput genomic sequencing.In this paper we ask whether we can infer evolutionary constraints from a set of sequence homologs of a protein. The challenge is to distinguish true co-evolution couplings from the noisy set of observed correlations. We address this challenge using a maximum entropy model of the protein sequence, constrained by the statistics of the multiple sequence alignment, to infer residue pair couplings. Surprisingly, we find that the strength of these inferred couplings is an excellent predictor of residue-residue proximity in folded structures. Indeed, the top-scoring residue couplings are sufficiently accurate and well-distributed to define the 3D protein fold with remarkable accuracy.We quantify this observation by computing, from sequence alone, all-atom 3D structures of fifteen test proteins from different fold classes, ranging in size from 50 to 260 residues, including a G-protein coupled receptor. These blinded inferences are de novo, i.e., they do not use homology modeling or sequence-similar fragments from known structures. The co-evolution signals provide sufficient information to determine accurate 3D protein structure to 2.7-4.8 Å C(α-RMSD error relative to the observed structure, over at least two-thirds of the protein (method called EVfold, details at http://EVfold.org. This discovery provides insight into essential interactions constraining protein evolution and will facilitate a comprehensive survey of the universe of

  20. Polymorphism in the Mr 32,000 Rh protein purified from Rh(D)-positive and -negative erythrocytes

    International Nuclear Information System (INIS)

    Saboori, A.M.; Smith, B.L.; Agre, P.

    1988-01-01

    A M r 32,000 integral membrane protein has previously been identified on erythrocytes bearing the Rh(D) antigen and is thought to contain the antigenic variations responsible for the different Rh phenotypes. To study it on a biochemical level, a simple large-scale method was developed to purify the M r 32,000 Rh protein from multiple units of Rh(D)-positive and -negative blood. Erythrocyte membrane vesicles were solubilized in NaDodSO 4 , and a tracer of immunoprecipitated 125 I surface-labeled Rh protein was added. The Rh protein was purified to homogeneity by hydroxylapatite chromatography followed by preparative NaDodSO 4 /PAGE. Approximately 25 nmol of pure Rh protein was recovered from each unit of Rh(D)-positive and -negative blood. Rh protein purified from both Rh phenotypes appeared similar by one-dimensional NaDodSO 4 /PAGE, and the N-terminal amino acid sequences for the first 20 residues were identical. Rh proteins purified from Rh(D)-positive and -negative blood were compared by two-dimensional iodopeptide mapping after 125 I-labeling and α-chymotrypsin digestion. The peptide maps were very similar. These data indicate that a similar core Rh protein exists in both Rh(D)-positive and -negative erythrocytes, and the Rh proteins from erythrocytes with different Rh phenotypes contain distinct structural polymorphisms

  1. Protein Function Prediction Based on Sequence and Structure Information

    KAUST Repository

    Smaili, Fatima Z.

    2016-05-25

    The number of available protein sequences in public databases is increasing exponentially. However, a significant fraction of these sequences lack functional annotation which is essential to our understanding of how biological systems and processes operate. In this master thesis project, we worked on inferring protein functions based on the primary protein sequence. In the approach we follow, 3D models are first constructed using I-TASSER. Functions are then deduced by structurally matching these predicted models, using global and local similarities, through three independent enzyme commission (EC) and gene ontology (GO) function libraries. The method was tested on 250 “hard” proteins, which lack homologous templates in both structure and function libraries. The results show that this method outperforms the conventional prediction methods based on sequence similarity or threading. Additionally, our method could be improved even further by incorporating protein-protein interaction information. Overall, the method we use provides an efficient approach for automated functional annotation of non-homologous proteins, starting from their sequence.

  2. Polymorphisms of the prion protein gene Arabi sheep breed in Iran

    African Journals Online (AJOL)

    Novin

    2011-11-09

    Nov 9, 2011 ... Key words: Prion protein gene (Prnp), polymorphisms, susceptibility, scrapie. INTRODUCTION. Scrapie is an invariably fatal .... consists of a DNA denaturation step (5 min at 95°C), followed by 30 amplification cycles ...

  3. FOLATE CYCLE GENE POLYMORPHISM AND ENDOGENOUS PEPTIDES IN CHILDREN WITH COW’S MILK PROTEIN ALLERGY

    Directory of Open Access Journals (Sweden)

    T. A. Shumatova

    2016-01-01

    Full Text Available Folate cycle gene polymorphisms and the levels of endogenous antimicrobial peptides and proteins in the blood and coprofiltrates were studied in 45 children aged 3 to 12 months with cow’s milk protein allergy. The polymorphic variants of the MTHFR, MTRR, and MTR genes were shown to be considered as a risk factor for the development of allergy. There was a significant increase in the levels of zonulin, β-defensin 2, transthyretin, and eosinophil cationic protein in the coprofiltrates and in those of eotaxin, fatty acidbinding proteins, and membrane permeability-increasing protein in the serum (p<0.05. The finding can improve the diagnosis of the disease for a predictive purpose for the evaluation of the efficiency of performed therapy.

  4. Codon 129 polymorphism of prion protein gene in is not a risk factor for Alzheimer's disease

    Directory of Open Access Journals (Sweden)

    Jerusa Smid

    2013-07-01

    Full Text Available Interaction of prion protein and amyloid-b oligomers has been demonstrated recently. Homozygosity at prion protein gene (PRNP codon 129 is associated with higher risk for Creutzfeldt-Jakob disease. This polymorphism has been addressed as a possible risk factor in Alzheimer disease (AD. Objective To describe the association between codon 129 polymorphisms and AD. Methods We investigated the association of codon 129 polymorphism of PRNP in 99 AD patients and 111 controls, and the association between this polymorphism and cognitive performance. Other polymorphisms of PRNP and additive effect of apolipoprotein E gene (ApoE were evaluated. Results Codon 129 genotype distribution in AD 45.5% methionine (MM, 42.2% methionine valine (MV, 12.1% valine (VV; and 39.6% MM, 50.5% MV, 9.9% VV among controls (p>0.05. There were no differences of cognitive performance concerning codon 129. Stratification according to ApoE genotype did not reveal difference between groups. Conclusion Codon 129 polymorphism is not a risk factor for AD in Brazilian patients.

  5. [open quotes]Cryptic[close quotes] repeating triplets of purines and pyrimidines (cRRY(i)) are frequent and polymorphic: Analysis of coding cRRY(i) in the proopiomelanocortin (POMC) and TATA-binding protein (TBP) genes

    Energy Technology Data Exchange (ETDEWEB)

    Gostout, B.; Qiang Liu; Sommer, S.S. (Mayo Clinic/Foundation, Rochester, MN (United States))

    1993-06-01

    Triplets of the form of purine, purine, pyrimidine (RRY(i)) are enhanced in frequency in the genomes of primates, rodents, and bacteria. Some RRY(i) are [open quotes]cryptic[close quotes] repeats (cRRY(i)) in which no one tandem run of a trinucleotide predominates. A search of human GenBank sequence revealed that the sequences of cRRY(i) are highly nonrandom. Three randomly chosen human cRRY(i) were sequenced in search of polymorphic alleles. Multiple polymorphic alleles were found in cRRY(i) in the coding regions of the genes for proopiomelanocortin (POMC) and TATA-binding protein (TBP). The highly polymorphic TBP cRRY(i) was characterized in detail. Direct sequencing of 157 unrelated human alleles demonstrated the presence of 20 different alleles which resulted in 29--40 consecutive glutamines in the amino-terminal region of TBP. These alleles are differently distributed among the races. PCR was used to screen 1,846 additional alleles in order to characterize more fully the range of variation in the population. Three additional alleles were discovered, but there was no example of a substantial sequence amplification as is seen in the repeat sequences associated with X-linked spinal and bulbar muscular atrophy, myotonic dystrophy, or the fragile-X syndrome. The structure of the TBP cRRY(i) is conserved in the five monkey species examined. In the chimpanzee, examination of four individuals revealed that the cRRY(i) was highly polymorphic, but the pattern of polymorphism differed from that in humans. The TBP cRRY(i) displays both similarities with and differences from the previously described RRY(i) in the coding sequence of the androgen receptor. The data suggest how simple tandem repeats could evolve from cryptic repeats. 18 refs., 3 figs., 6 tabs.

  6. Association of the macrophage activating factor (MAF) precursor activity with polymorphism in vitamin D-binding protein.

    Science.gov (United States)

    Nagasawa, Hideko; Sasaki, Hideyuki; Uto, Yoshihiro; Kubo, Shinichi; Hori, Hitoshi

    2004-01-01

    Serum vitamin D-binding protein (Gc protein or DBP) is a highly expressed polymorphic protein, which is a precursor of the inflammation-primed macrophage activating factor, GcMAF, by a cascade of carbohydrate processing reactions. In order to elucidate the relationship between Gc polymorphism and GcMAF precursor activity, we estimated the phagocytic ability of three homotypes of Gc protein, Gc1F-1F, Gc1S-1S and Gc2-2, through processing of their carbohydrate moiety. We performed Gc typing of human serum samples by isoelectric focusing (IEF). Gc protein from human serum was purified by affinity chromatography with 25-hydroxyvitamin D3-sepharose. A phagocytosis assay of Gc proteins, modified using beta-glycosidase and sialidase, was carried out. The Gc1F-1F phenotype was revealed to possess Galbeta1-4GalNAc linkage by the analysis of GcMAF precursor activity using beta1-4 linkage-specific galactosidase from jack bean. The GcMAF precursor activity of the Gc1F-1F phenotype was highest among three Gc homotypes. The Gc polymorphism and carbohydrate diversity of Gc protein are significant for its pleiotropic effects.

  7. Evaluation of genetic diversity in Chinese kale (Brassica oleracea L. var. alboglabra Bailey) by using rapid amplified polymorphic DNA and sequence-related amplified polymorphism markers.

    Science.gov (United States)

    Zhang, J; Zhang, L G

    2014-02-14

    Chinese kale is an original Chinese vegetable of the Cruciferae family. To select suitable parents for hybrid breeding, we thoroughly analyzed the genetic diversity of Chinese kale. Random amplified polymorphic DNA (RAPD) and sequence-related amplified polymorphism (SRAP) molecular markers were used to evaluate the genetic diversity across 21 Chinese kale accessions from AVRDC and Guangzhou in China. A total of 104 bands were detected by 11 RAPD primers, of which 66 (63.5%) were polymorphic, and 229 polymorphic bands (68.4%) were observed in 335 bands amplified by 17 SRAP primer combinations. The dendrogram showed the grouping of the 21 accessions into 4 main clusters based on RAPD data, and into 6 clusters based on SRAP and combined data (RAPD + SRAP). The clustering of accessions based on SRAP data was consistent with petal colors. The Mantel test indicated a poor fit for the RAPD and SRAP data (r = 0.16). These results have an important implication for Chinese kale germplasm characterization and improvement.

  8. Detection of Egg Production of Tegal Duck by Blood Protein Polymorphism

    Directory of Open Access Journals (Sweden)

    Ismoyowati Ismoyowati

    2008-05-01

    Full Text Available The aim of this research was to study the effect of transfferine, albumine, and haemoglobine loci to egg production characteristic of Tegal duck.  100 lying of Tegal ducks keeping by batteray-pen were used in this study.  Individual egg production was recorded until period of 120 days. Blood protein polymorphism analysed by electrophoresis method, and blood sample taken from each ducks.. Egg production and transfferine albumine, and haemoglobine phenotipe on electrophoresis gel were observed in this study.  Genotipe and gene frequencies and genetic variant were applied in data analysis. The result showed that (1 in the transferine locus were identified 3 aleles forming 4 genotipes (TfAA,TfAB, TfBB, and TfBC, (2 in albumine were identified 3 aleles forming 5 genotipes (AlbAA, AlbAB, AlbAC, AlbBB and AlbBC and (3 haemoglobine locus were identified 6 aleles forming 4 genotipes ((HbAA, HbAB, HbAC, HbBB, HbBC dan HbCC.  This study demostrated that B gene frequenci in transfferine, albumine and haemoglonine loci was highest than A and C gene frequency.  Tegal Duck with AA genotipe on all loci had higher egg production than BB and CC homozigote.  This research revealed that the most efective of selection method by haemoglobine protein polymorphism. (Animal Production 10(2: 122-128 (2008   Key Words: Tegal duck, egg production, selection, blood protein polymorphism

  9. GuiTope: an application for mapping random-sequence peptides to protein sequences.

    Science.gov (United States)

    Halperin, Rebecca F; Stafford, Phillip; Emery, Jack S; Navalkar, Krupa Arun; Johnston, Stephen Albert

    2012-01-03

    Random-sequence peptide libraries are a commonly used tool to identify novel ligands for binding antibodies, other proteins, and small molecules. It is often of interest to compare the selected peptide sequences to the natural protein binding partners to infer the exact binding site or the importance of particular residues. The ability to search a set of sequences for similarity to a set of peptides may sometimes enable the prediction of an antibody epitope or a novel binding partner. We have developed a software application designed specifically for this task. GuiTope provides a graphical user interface for aligning peptide sequences to protein sequences. All alignment parameters are accessible to the user including the ability to specify the amino acid frequency in the peptide library; these frequencies often differ significantly from those assumed by popular alignment programs. It also includes a novel feature to align di-peptide inversions, which we have found improves the accuracy of antibody epitope prediction from peptide microarray data and shows utility in analyzing phage display datasets. Finally, GuiTope can randomly select peptides from a given library to estimate a null distribution of scores and calculate statistical significance. GuiTope provides a convenient method for comparing selected peptide sequences to protein sequences, including flexible alignment parameters, novel alignment features, ability to search a database, and statistical significance of results. The software is available as an executable (for PC) at http://www.immunosignature.com/software and ongoing updates and source code will be available at sourceforge.net.

  10. GuiTope: an application for mapping random-sequence peptides to protein sequences

    Directory of Open Access Journals (Sweden)

    Halperin Rebecca F

    2012-01-01

    Full Text Available Abstract Background Random-sequence peptide libraries are a commonly used tool to identify novel ligands for binding antibodies, other proteins, and small molecules. It is often of interest to compare the selected peptide sequences to the natural protein binding partners to infer the exact binding site or the importance of particular residues. The ability to search a set of sequences for similarity to a set of peptides may sometimes enable the prediction of an antibody epitope or a novel binding partner. We have developed a software application designed specifically for this task. Results GuiTope provides a graphical user interface for aligning peptide sequences to protein sequences. All alignment parameters are accessible to the user including the ability to specify the amino acid frequency in the peptide library; these frequencies often differ significantly from those assumed by popular alignment programs. It also includes a novel feature to align di-peptide inversions, which we have found improves the accuracy of antibody epitope prediction from peptide microarray data and shows utility in analyzing phage display datasets. Finally, GuiTope can randomly select peptides from a given library to estimate a null distribution of scores and calculate statistical significance. Conclusions GuiTope provides a convenient method for comparing selected peptide sequences to protein sequences, including flexible alignment parameters, novel alignment features, ability to search a database, and statistical significance of results. The software is available as an executable (for PC at http://www.immunosignature.com/software and ongoing updates and source code will be available at sourceforge.net.

  11. Overlapping Genomic Sequences: A Treasure Trove of Single-Nucleotide Polymorphisms

    Science.gov (United States)

    Taillon-Miller, Patricia; Gu, Zhijie; Li, Qun; Hillier, LaDeana; Kwok, Pui-Yan

    1998-01-01

    An efficient strategy to develop a dense set of single-nucleotide polymorphism (SNP) markers is to take advantage of the human genome sequencing effort currently under way. Our approach is based on the fact that bacterial artificial chromosomes (BACs) and P1-based artificial chromosomes (PACs) used in long-range sequencing projects come from diploid libraries. If the overlapping clones sequenced are from different lineages, one is comparing the sequences from 2 homologous chromosomes in the overlapping region. We have analyzed in detail every SNP identified while sequencing three sets of overlapping clones found on chromosome 5p15.2, 7q21–7q22, and 13q12–13q13. In the 200.6 kb of DNA sequence analyzed in these overlaps, 153 SNPs were identified. Computer analysis for repetitive elements and suitability for STS development yielded 44 STSs containing 68 SNPs for further study. All 68 SNPs were confirmed to be present in at least one of the three (Caucasian, African-American, Hispanic) populations studied. Furthermore, 42 of the SNPs tested (62%) were informative in at least one population, 32 (47%) were informative in two or more populations, and 23 (34%) were informative in all three populations. These results clearly indicate that developing SNP markers from overlapping genomic sequence is highly efficient and cost effective, requiring only the two simple steps of developing STSs around the known SNPs and characterizing them in the appropriate populations. [The sequence data described in this paper have been submitted to the GenBank data library under accession nos. AC003015 (for GS113423), AC002380 (GS330J10), AC000066 (RG293F11), AC003086 (RG104F04), AC002525 (257C22A), and U73331 (96A18A).] PMID:9685323

  12. Dynamics of domain coverage of the protein sequence universe

    Science.gov (United States)

    2012-01-01

    Background The currently known protein sequence space consists of millions of sequences in public databases and is rapidly expanding. Assigning sequences to families leads to a better understanding of protein function and the nature of the protein universe. However, a large portion of the current protein space remains unassigned and is referred to as its “dark matter”. Results Here we suggest that true size of “dark matter” is much larger than stated by current definitions. We propose an approach to reducing the size of “dark matter” by identifying and subtracting regions in protein sequences that are not likely to contain any domain. Conclusions Recent improvements in computational domain modeling result in a decrease, albeit slowly, in the relative size of “dark matter”; however, its absolute size increases substantially with the growth of sequence data. PMID:23157439

  13. Dynamics of domain coverage of the protein sequence universe

    Directory of Open Access Journals (Sweden)

    Rekapalli Bhanu

    2012-11-01

    Full Text Available Abstract Background The currently known protein sequence space consists of millions of sequences in public databases and is rapidly expanding. Assigning sequences to families leads to a better understanding of protein function and the nature of the protein universe. However, a large portion of the current protein space remains unassigned and is referred to as its “dark matter”. Results Here we suggest that true size of “dark matter” is much larger than stated by current definitions. We propose an approach to reducing the size of “dark matter” by identifying and subtracting regions in protein sequences that are not likely to contain any domain. Conclusions Recent improvements in computational domain modeling result in a decrease, albeit slowly, in the relative size of “dark matter”; however, its absolute size increases substantially with the growth of sequence data.

  14. Advances in molecular modeling of human cytochrome P450 polymorphism.

    Science.gov (United States)

    Martiny, Virginie Y; Miteva, Maria A

    2013-11-01

    Cytochrome P450 (CYP) is a supergene family of metabolizing enzymes involved in the phase I metabolism of drugs and endogenous compounds. CYP oxidation often leads to inactive drug metabolites or to highly toxic or carcinogenic metabolites involved in adverse drug reactions (ADR). During the last decade, the impact of CYP polymorphism in various drug responses and ADR has been demonstrated. Of the drugs involved in ADR, 56% are metabolized by polymorphic phase I metabolizing enzymes, 86% among them being CYP. Here, we review the major CYP polymorphic forms, their impact for drug response and current advances in molecular modeling of CYP polymorphism. We focus on recent studies exploring CYP polymorphism performed by the use of sequence-based and/or protein-structure-based computational approaches. The importance of understanding the molecular mechanisms related to CYP polymorphism and drug response at the atomic level is outlined. © 2013.

  15. MIPS: a database for genomes and protein sequences.

    Science.gov (United States)

    Mewes, H W; Frishman, D; Güldener, U; Mannhaupt, G; Mayer, K; Mokrejs, M; Morgenstern, B; Münsterkötter, M; Rudd, S; Weil, B

    2002-01-01

    The Munich Information Center for Protein Sequences (MIPS-GSF, Neuherberg, Germany) continues to provide genome-related information in a systematic way. MIPS supports both national and European sequencing and functional analysis projects, develops and maintains automatically generated and manually annotated genome-specific databases, develops systematic classification schemes for the functional annotation of protein sequences, and provides tools for the comprehensive analysis of protein sequences. This report updates the information on the yeast genome (CYGD), the Neurospora crassa genome (MNCDB), the databases for the comprehensive set of genomes (PEDANT genomes), the database of annotated human EST clusters (HIB), the database of complete cDNAs from the DHGP (German Human Genome Project), as well as the project specific databases for the GABI (Genome Analysis in Plants) and HNB (Helmholtz-Netzwerk Bioinformatik) networks. The Arabidospsis thaliana database (MATDB), the database of mitochondrial proteins (MITOP) and our contribution to the PIR International Protein Sequence Database have been described elsewhere [Schoof et al. (2002) Nucleic Acids Res., 30, 91-93; Scharfe et al. (2000) Nucleic Acids Res., 28, 155-158; Barker et al. (2001) Nucleic Acids Res., 29, 29-32]. All databases described, the protein analysis tools provided and the detailed descriptions of our projects can be accessed through the MIPS World Wide Web server (http://mips.gsf.de).

  16. Prediction of Protein Hotspots from Whole Protein Sequences by a Random Projection Ensemble System

    Directory of Open Access Journals (Sweden)

    Jinjian Jiang

    2017-07-01

    Full Text Available Hotspot residues are important in the determination of protein-protein interactions, and they always perform specific functions in biological processes. The determination of hotspot residues is by the commonly-used method of alanine scanning mutagenesis experiments, which is always costly and time consuming. To address this issue, computational methods have been developed. Most of them are structure based, i.e., using the information of solved protein structures. However, the number of solved protein structures is extremely less than that of sequences. Moreover, almost all of the predictors identified hotspots from the interfaces of protein complexes, seldom from the whole protein sequences. Therefore, determining hotspots from whole protein sequences by sequence information alone is urgent. To address the issue of hotspot predictions from the whole sequences of proteins, we proposed an ensemble system with random projections using statistical physicochemical properties of amino acids. First, an encoding scheme involving sequence profiles of residues and physicochemical properties from the AAindex1 dataset is developed. Then, the random projection technique was adopted to project the encoding instances into a reduced space. Then, several better random projections were obtained by training an IBk classifier based on the training dataset, which were thus applied to the test dataset. The ensemble of random projection classifiers is therefore obtained. Experimental results showed that although the performance of our method is not good enough for real applications of hotspots, it is very promising in the determination of hotspot residues from whole sequences.

  17. A novel approach to sequence validating protein expression clones with automated decision making

    Directory of Open Access Journals (Sweden)

    Mohr Stephanie E

    2007-06-01

    Full Text Available Abstract Background Whereas the molecular assembly of protein expression clones is readily automated and routinely accomplished in high throughput, sequence verification of these clones is still largely performed manually, an arduous and time consuming process. The ultimate goal of validation is to determine if a given plasmid clone matches its reference sequence sufficiently to be "acceptable" for use in protein expression experiments. Given the accelerating increase in availability of tens of thousands of unverified clones, there is a strong demand for rapid, efficient and accurate software that automates clone validation. Results We have developed an Automated Clone Evaluation (ACE system – the first comprehensive, multi-platform, web-based plasmid sequence verification software package. ACE automates the clone verification process by defining each clone sequence as a list of multidimensional discrepancy objects, each describing a difference between the clone and its expected sequence including the resulting polypeptide consequences. To evaluate clones automatically, this list can be compared against user acceptance criteria that specify the allowable number of discrepancies of each type. This strategy allows users to re-evaluate the same set of clones against different acceptance criteria as needed for use in other experiments. ACE manages the entire sequence validation process including contig management, identifying and annotating discrepancies, determining if discrepancies correspond to polymorphisms and clone finishing. Designed to manage thousands of clones simultaneously, ACE maintains a relational database to store information about clones at various completion stages, project processing parameters and acceptance criteria. In a direct comparison, the automated analysis by ACE took less time and was more accurate than a manual analysis of a 93 gene clone set. Conclusion ACE was designed to facilitate high throughput clone sequence

  18. Sequence-Related Amplified Polymorphism (SRAP Markers: A Potential Resource for Studies in Plant Molecular Biology

    Directory of Open Access Journals (Sweden)

    Daniel W. H. Robarts

    2014-07-01

    Full Text Available In the past few decades, many investigations in the field of plant biology have employed selectively neutral, multilocus, dominant markers such as inter-simple sequence repeat (ISSR, random-amplified polymorphic DNA (RAPD, and amplified fragment length polymorphism (AFLP to address hypotheses at lower taxonomic levels. More recently, sequence-related amplified polymorphism (SRAP markers have been developed, which are used to amplify coding regions of DNA with primers targeting open reading frames. These markers have proven to be robust and highly variable, on par with AFLP, and are attained through a significantly less technically demanding process. SRAP markers have been used primarily for agronomic and horticultural purposes, developing quantitative trait loci in advanced hybrids and assessing genetic diversity of large germplasm collections. Here, we suggest that SRAP markers should be employed for research addressing hypotheses in plant systematics, biogeography, conservation, ecology, and beyond. We provide an overview of the SRAP literature to date, review descriptive statistics of SRAP markers in a subset of 171 publications, and present relevant case studies to demonstrate the applicability of SRAP markers to the diverse field of plant biology. Results of these selected works indicate that SRAP markers have the potential to enhance the current suite of molecular tools in a diversity of fields by providing an easy-to-use. highly variable marker with inherent biological significance.

  19. Genomic polymorphism, recombination, and linkage disequilibrium in human major histocompatibility complex-encoded antigen-processing genes.

    Science.gov (United States)

    van Endert, P M; Lopez, M T; Patel, S D; Monaco, J J; McDevitt, H O

    1992-01-01

    Recently, two subunits of a large cytosolic protease and two putative peptide transporter proteins were found to be encoded by genes within the class II region of the major histocompatibility complex (MHC). These genes have been suggested to be involved in the processing of antigenic proteins for presentation by MHC class I molecules. Because of the high degree of polymorphism in MHC genes, and previous evidence for both functional and polypeptide sequence polymorphism in the proteins encoded by the antigen-processing genes, we tested DNA from 27 consanguineous human cell lines for genomic polymorphism by restriction fragment length polymorphism (RFLP) analysis. These studies demonstrate a strong linkage disequilibrium between TAP1 and LMP2 RFLPs. Moreover, RFLPs, as well as a polymorphic stop codon in the telomeric TAP2 gene, appear to be in linkage disequilibrium with HLA-DR alleles and RFLPs in the HLA-DO gene. A high rate of recombination, however, seems to occur in the center of the complex, between the TAP1 and TAP2 genes. Images PMID:1360671

  20. Gene Unprediction with Spurio: A tool to identify spurious protein sequences.

    Science.gov (United States)

    Höps, Wolfram; Jeffryes, Matt; Bateman, Alex

    2018-01-01

    We now have access to the sequences of tens of millions of proteins. These protein sequences are essential for modern molecular biology and computational biology. The vast majority of protein sequences are derived from gene prediction tools and have no experimental supporting evidence for their translation.  Despite the increasing accuracy of gene prediction tools there likely exists a large number of spurious protein predictions in the sequence databases.  We have developed the Spurio tool to help identify spurious protein predictions in prokaryotes.  Spurio searches the query protein sequence against a prokaryotic nucleotide database using tblastn and identifies homologous sequences. The tblastn matches are used to score the query sequence's likelihood of being a spurious protein prediction using a Gaussian process model. The most informative feature is the appearance of stop codons within the presumed translation of homologous DNA sequences. Benchmarking shows that the Spurio tool is able to distinguish spurious from true proteins. However, transposon proteins are prone to be predicted as spurious because of the frequency of degraded homologs found in the DNA sequence databases. Our initial experiments suggest that less than 1% of the proteins in the UniProtKB sequence database are likely to be spurious and that Spurio is able to identify over 60 times more spurious proteins than the AntiFam resource. The Spurio software and source code is available under an MIT license at the following URL: https://bitbucket.org/bateman-group/spurio.

  1. Structure and Sequence Search on Aptamer-Protein Docking

    Science.gov (United States)

    Xiao, Jiajie; Bonin, Keith; Guthold, Martin; Salsbury, Freddie

    2015-03-01

    Interactions between proteins and deoxyribonucleic acid (DNA) play a significant role in the living systems, especially through gene regulation. However, short nucleic acids sequences (aptamers) with specific binding affinity to specific proteins exhibit clinical potential as therapeutics. Our capillary and gel electrophoresis selection experiments show that specific sequences of aptamers can be selected that bind specific proteins. Computationally, given the experimentally-determined structure and sequence of a thrombin-binding aptamer, we can successfully dock the aptamer onto thrombin in agreement with experimental structures of the complex. In order to further study the conformational flexibility of this thrombin-binding aptamer and to potentially develop a predictive computational model of aptamer-binding, we use GPU-enabled molecular dynamics simulations to both examine the conformational flexibility of the aptamer in the absence of binding to thrombin, and to determine our ability to fold an aptamer. This study should help further de-novo predictions of aptamer sequences by enabling the study of structural and sequence-dependent effects on aptamer-protein docking specificity.

  2. AlignMe—a membrane protein sequence alignment web server

    Science.gov (United States)

    Stamm, Marcus; Staritzbichler, René; Khafizov, Kamil; Forrest, Lucy R.

    2014-01-01

    We present a web server for pair-wise alignment of membrane protein sequences, using the program AlignMe. The server makes available two operational modes of AlignMe: (i) sequence to sequence alignment, taking two sequences in fasta format as input, combining information about each sequence from multiple sources and producing a pair-wise alignment (PW mode); and (ii) alignment of two multiple sequence alignments to create family-averaged hydropathy profile alignments (HP mode). For the PW sequence alignment mode, four different optimized parameter sets are provided, each suited to pairs of sequences with a specific similarity level. These settings utilize different types of inputs: (position-specific) substitution matrices, secondary structure predictions and transmembrane propensities from transmembrane predictions or hydrophobicity scales. In the second (HP) mode, each input multiple sequence alignment is converted into a hydrophobicity profile averaged over the provided set of sequence homologs; the two profiles are then aligned. The HP mode enables qualitative comparison of transmembrane topologies (and therefore potentially of 3D folds) of two membrane proteins, which can be useful if the proteins have low sequence similarity. In summary, the AlignMe web server provides user-friendly access to a set of tools for analysis and comparison of membrane protein sequences. Access is available at http://www.bioinfo.mpg.de/AlignMe PMID:24753425

  3. The relationship of protein conservation and sequence length

    Directory of Open Access Journals (Sweden)

    Panchenko Anna R

    2002-11-01

    Full Text Available Abstract Background In general, the length of a protein sequence is determined by its function and the wide variance in the lengths of an organism's proteins reflects the diversity of specific functional roles for these proteins. However, additional evolutionary forces that affect the length of a protein may be revealed by studying the length distributions of proteins evolving under weaker functional constraints. Results We performed sequence comparisons to distinguish highly conserved and poorly conserved proteins from the bacterium Escherichia coli, the archaeon Archaeoglobus fulgidus, and the eukaryotes Saccharomyces cerevisiae, Drosophila melanogaster, and Homo sapiens. For all organisms studied, the conserved and nonconserved proteins have strikingly different length distributions. The conserved proteins are, on average, longer than the poorly conserved ones, and the length distributions for the poorly conserved proteins have a relatively narrow peak, in contrast to the conserved proteins whose lengths spread over a wider range of values. For the two prokaryotes studied, the poorly conserved proteins approximate the minimal length distribution expected for a diverse range of structural folds. Conclusions There is a relationship between protein conservation and sequence length. For all the organisms studied, there seems to be a significant evolutionary trend favoring shorter proteins in the absence of other, more specific functional constraints.

  4. Modeling genetic imprinting effects of DNA sequences with multilocus polymorphism data

    Directory of Open Access Journals (Sweden)

    Staud Roland

    2009-08-01

    Full Text Available Abstract Single nucleotide polymorphisms (SNPs represent the most widespread type of DNA sequence variation in the human genome and they have recently emerged as valuable genetic markers for revealing the genetic architecture of complex traits in terms of nucleotide combination and sequence. Here, we extend an algorithmic model for the haplotype analysis of SNPs to estimate the effects of genetic imprinting expressed at the DNA sequence level. The model provides a general procedure for identifying the number and types of optimal DNA sequence variants that are expressed differently due to their parental origin. The model is used to analyze a genetic data set collected from a pain genetics project. We find that DNA haplotype GAC from three SNPs, OPRKG36T (with two alleles G and T, OPRKA843G (with alleles A and G, and OPRKC846T (with alleles C and T, at the kappa-opioid receptor, triggers a significant effect on pain sensitivity, but with expression significantly depending on the parent from which it is inherited (p = 0.008. With a tremendous advance in SNP identification and automated screening, the model founded on haplotype discovery and statistical inference may provide a useful tool for genetic analysis of any quantitative trait with complex inheritance.

  5. MIPS: a database for protein sequences and complete genomes.

    Science.gov (United States)

    Mewes, H W; Hani, J; Pfeiffer, F; Frishman, D

    1998-01-01

    The MIPS group [Munich Information Center for Protein Sequences of the German National Center for Environment and Health (GSF)] at the Max-Planck-Institute for Biochemistry, Martinsried near Munich, Germany, is involved in a number of data collection activities, including a comprehensive database of the yeast genome, a database reflecting the progress in sequencing the Arabidopsis thaliana genome, the systematic analysis of other small genomes and the collection of protein sequence data within the framework of the PIR-International Protein Sequence Database (described elsewhere in this volume). Through its WWW server (http://www.mips.biochem.mpg.de ) MIPS provides access to a variety of generic databases, including a database of protein families as well as automatically generated data by the systematic application of sequence analysis algorithms. The yeast genome sequence and its related information was also compiled on CD-ROM to provide dynamic interactive access to the 16 chromosomes of the first eukaryotic genome unraveled. PMID:9399795

  6. Can Natural Proteins Designed with ‘Inverted’ Peptide Sequences Adopt Native-Like Protein Folds?

    Science.gov (United States)

    Sridhar, Settu; Guruprasad, Kunchur

    2014-01-01

    We have carried out a systematic computational analysis on a representative dataset of proteins of known three-dimensional structure, in order to evaluate whether it would possible to ‘swap’ certain short peptide sequences in naturally occurring proteins with their corresponding ‘inverted’ peptides and generate ‘artificial’ proteins that are predicted to retain native-like protein fold. The analysis of 3,967 representative proteins from the Protein Data Bank revealed 102,677 unique identical inverted peptide sequence pairs that vary in sequence length between 5–12 and 18 amino acid residues. Our analysis illustrates with examples that such ‘artificial’ proteins may be generated by identifying peptides with ‘similar structural environment’ and by using comparative protein modeling and validation studies. Our analysis suggests that natural proteins may be tolerant to accommodating such peptides. PMID:25210740

  7. Canine olfactory receptor gene polymorphism and its relation to odor detection performance by sniffer dogs.

    Science.gov (United States)

    Lesniak, Anna; Walczak, Marta; Jezierski, Tadeusz; Sacharczuk, Mariusz; Gawkowski, Maciej; Jaszczak, Kazimierz

    2008-01-01

    The outstanding sensitivity of the canine olfactory system has been acknowledged by using sniffer dogs in military and civilian service for detection of a variety of odors. It is hypothesized that the canine olfactory ability is determined by polymorphisms in olfactory receptor (OR) genes. We investigated 5 OR genes for polymorphic sites which might affect the olfactory ability of service dogs in different fields of specific substance detection. All investigated OR DNA sequences proved to have allelic variants, the majority of which lead to protein sequence alteration. Homozygous individuals at 2 gene loci significantly differed in their detection skills from other genotypes. This suggests a role of specific alleles in odor detection and a linkage between single-nucleotide polymorphism and odor recognition efficiency.

  8. The SWISS-PROT protein sequence data bank: current status.

    OpenAIRE

    Bairoch, A; Boeckmann, B

    1994-01-01

    SWISS-PROT is an annotated protein sequence database established in 1986 and maintained collaboratively, since 1988, by the Department of Medical Biochemistry of the University of Geneva and the EMBL Data Library. The SWISS-PROT protein sequence data bank consist of sequence entries. Sequence entries are composed of different lines types, each with their own format. For standardization purposes the format of SWISS-PROT follows as closely as possible that of the EMBL Nucleotide Sequence Databa...

  9. Identification and Evaluation of Single-Nucleotide Polymorphisms in Allotetraploid Peanut (Arachis hypogaea L.) Based on Amplicon Sequencing Combined with High Resolution Melting (HRM) Analysis.

    Science.gov (United States)

    Hong, Yanbin; Pandey, Manish K; Liu, Ying; Chen, Xiaoping; Liu, Hong; Varshney, Rajeev K; Liang, Xuanqiang; Huang, Shangzhi

    2015-01-01

    The cultivated peanut (Arachis hypogaea L.) is an allotetraploid (AABB) species derived from the A-genome (Arachis duranensis) and B-genome (Arachis ipaensis) progenitors. Presence of two versions of a DNA sequence based on the two progenitor genomes poses a serious technical and analytical problem during single nucleotide polymorphism (SNP) marker identification and analysis. In this context, we have analyzed 200 amplicons derived from expressed sequence tags (ESTs) and genome survey sequences (GSS) to identify SNPs in a panel of genotypes consisting of 12 cultivated peanut varieties and two diploid progenitors representing the ancestral genomes. A total of 18 EST-SNPs and 44 genomic-SNPs were identified in 12 peanut varieties by aligning the sequence of A. hypogaea with diploid progenitors. The average frequency of sequence polymorphism was higher for genomic-SNPs than the EST-SNPs with one genomic-SNP every 1011 bp as compared to one EST-SNP every 2557 bp. In order to estimate the potential and further applicability of these identified SNPs, 96 peanut varieties were genotyped using high resolution melting (HRM) method. Polymorphism information content (PIC) values for EST-SNPs ranged between 0.021 and 0.413 with a mean of 0.172 in the set of peanut varieties, while genomic-SNPs ranged between 0.080 and 0.478 with a mean of 0.249. Total 33 SNPs were used for polymorphism detection among the parents and 10 selected lines from mapping population Y13Zh (Zhenzhuhei × Yueyou13). Of the total 33 SNPs, nine SNPs showed polymorphism in the mapping population Y13Zh, and seven SNPs were successfully mapped into five linkage groups. Our results showed that SNPs can be identified in allotetraploid peanut with high accuracy through amplicon sequencing and HRM assay. The identified SNPs were very informative and can be used for different genetic and breeding applications in peanut.

  10. Partial sequence determination of metabolically labeled radioactive proteins and peptides

    International Nuclear Information System (INIS)

    Anderson, C.W.

    1982-01-01

    The author has used the sequence analysis of radioactive proteins and peptides to approach several problems during the past few years. They, in collaboration with others, have mapped precisely several adenovirus proteins with respect to the nucleotide sequence of the adenovirus genome; identified hitherto missed proteins encoded by bacteriophage MS2 and by simian virus 40; analyzed the aminoterminal maturation of several virus proteins; determined the cleavage sites for processing of the poliovirus polyprotein; and analyzed the mechanism of frameshifting by excess normal tRNAs during cell-free protein synthesis. This chapter is designed to aid those without prior experience at protein sequence determinations. It is based primarily on the experience gained in the studies cited above, which made use of the Beckman 890 series automated protein sequencers

  11. Correlation between protein sequence similarity and x-ray diffraction quality in the protein data bank.

    Science.gov (United States)

    Lu, Hui-Meng; Yin, Da-Chuan; Ye, Ya-Jing; Luo, Hui-Min; Geng, Li-Qiang; Li, Hai-Sheng; Guo, Wei-Hong; Shang, Peng

    2009-01-01

    As the most widely utilized technique to determine the 3-dimensional structure of protein molecules, X-ray crystallography can provide structure of the highest resolution among the developed techniques. The resolution obtained via X-ray crystallography is known to be influenced by many factors, such as the crystal quality, diffraction techniques, and X-ray sources, etc. In this paper, the authors found that the protein sequence could also be one of the factors. We extracted information of the resolution and the sequence of proteins from the Protein Data Bank (PDB), classified the proteins into different clusters according to the sequence similarity, and statistically analyzed the relationship between the sequence similarity and the best resolution obtained. The results showed that there was a pronounced correlation between the sequence similarity and the obtained resolution. These results indicate that protein structure itself is one variable that may affect resolution when X-ray crystallography is used.

  12. Taxonomic colouring of phylogenetic trees of protein sequences

    Directory of Open Access Journals (Sweden)

    Andrade-Navarro Miguel A

    2006-02-01

    Full Text Available Abstract Background Phylogenetic analyses of protein families are used to define the evolutionary relationships between homologous proteins. The interpretation of protein-sequence phylogenetic trees requires the examination of the taxonomic properties of the species associated to those sequences. However, there is no online tool to facilitate this interpretation, for example, by automatically attaching taxonomic information to the nodes of a tree, or by interactively colouring the branches of a tree according to any combination of taxonomic divisions. This is especially problematic if the tree contains on the order of hundreds of sequences, which, given the accelerated increase in the size of the protein sequence databases, is a situation that is becoming common. Results We have developed PhyloView, a web based tool for colouring phylogenetic trees upon arbitrary taxonomic properties of the species represented in a protein sequence phylogenetic tree. Provided that the tree contains SwissProt, SpTrembl, or GenBank protein identifiers, the tool retrieves the taxonomic information from the corresponding database. A colour picker displays a summary of the findings and allows the user to associate colours to the leaves of the tree according to any number of taxonomic partitions. Then, the colours are propagated to the branches of the tree. Conclusion PhyloView can be used at http://www.ogic.ca/projects/phyloview/. A tutorial, the software with documentation, and GPL licensed source code, can be accessed at the same web address.

  13. Semi-Supervised Learning for Classification of Protein Sequence Data

    Directory of Open Access Journals (Sweden)

    Brian R. King

    2008-01-01

    Full Text Available Protein sequence data continue to become available at an exponential rate. Annotation of functional and structural attributes of these data lags far behind, with only a small fraction of the data understood and labeled by experimental methods. Classification methods that are based on semi-supervised learning can increase the overall accuracy of classifying partly labeled data in many domains, but very few methods exist that have shown their effect on protein sequence classification. We show how proven methods from text classification can be applied to protein sequence data, as we consider both existing and novel extensions to the basic methods, and demonstrate restrictions and differences that must be considered. We demonstrate comparative results against the transductive support vector machine, and show superior results on the most difficult classification problems. Our results show that large repositories of unlabeled protein sequence data can indeed be used to improve predictive performance, particularly in situations where there are fewer labeled protein sequences available, and/or the data are highly unbalanced in nature.

  14. Development, characterization and cross species amplification of polymorphic microsatellite markers from expressed sequence tags of turmeric (Curcuma longa L.).

    Science.gov (United States)

    Siju, S; Dhanya, K; Syamkumar, S; Sasikumar, B; Sheeja, T E; Bhat, A I; Parthasarathy, V A

    2010-02-01

    Expressed sequence tags (ESTs) from turmeric (Curcuma longa L.) were used for the screening of type and frequency of Class I (hypervariable) simple sequence repeats (SSRs). A total of 231 microsatellite repeats were detected from 12,593 EST sequences of turmeric after redundancy elimination. The average density of Class I SSRs accounts to one SSR per 17.96 kb of EST. Mononucleotides were the most abundant class of microsatellite repeat in turmeric ESTs followed by trinucleotides. A robust set of 17 polymorphic EST-SSRs were developed and used for evaluating 20 turmeric accessions. The number of alleles detected ranged from 3 to 8 per loci. The developed markers were also evaluated in 13 related species of C. longa confirming high rate (100%) of cross species transferability. The polymorphic microsatellite markers generated from this study could be used for genetic diversity analysis and resolving the taxonomic confusion prevailing in the genus.

  15. Sparse networks of directly coupled, polymorphic, and functional side chains in allosteric proteins.

    Science.gov (United States)

    Soltan Ghoraie, Laleh; Burkowski, Forbes; Zhu, Mu

    2015-03-01

    Recent studies have highlighted the role of coupled side-chain fluctuations alone in the allosteric behavior of proteins. Moreover, examination of X-ray crystallography data has recently revealed new information about the prevalence of alternate side-chain conformations (conformational polymorphism), and attempts have been made to uncover the hidden alternate conformations from X-ray data. Hence, new computational approaches are required that consider the polymorphic nature of the side chains, and incorporate the effects of this phenomenon in the study of information transmission and functional interactions of residues in a molecule. These studies can provide a more accurate understanding of the allosteric behavior. In this article, we first present a novel approach to generate an ensemble of conformations and an efficient computational method to extract direct couplings of side chains in allosteric proteins, and provide sparse network representations of the couplings. We take the side-chain conformational polymorphism into account, and show that by studying the intrinsic dynamics of an inactive structure, we are able to construct a network of functionally crucial residues. Second, we show that the proposed method is capable of providing a magnified view of the coupled and conformationally polymorphic residues. This model reveals couplings between the alternate conformations of a coupled residue pair. To the best of our knowledge, this is the first computational method for extracting networks of side chains' alternate conformations. Such networks help in providing a detailed image of side-chain dynamics in functionally important and conformationally polymorphic sites, such as binding and/or allosteric sites. © 2014 Wiley Periodicals, Inc.

  16. Surfactant protein D multimerization and gene polymorphism in COPD and asthma

    DEFF Research Database (Denmark)

    Fakih, Dalia; Akiki, Zeina; Junker, Kirsten

    2018-01-01

    BACKGROUND AND OBJECTIVE: A structural single nucleotide polymorphism rs721917 in the surfactant protein D (SP-D) gene, known as Met11Thr, was reported to influence the circulating levels and degree of multimerization of SP-D and was associated with both COPD and atopy in asthma. Moreover, diseas...

  17. Association of Polymorphism in Gene of Pregnancy-Associated Plasma Protein A (PAPPA and Preeclampsia

    Directory of Open Access Journals (Sweden)

    Nasrin Moghaddam

    2018-02-01

    Full Text Available Background: Preeclampsia is a common disorder of pregnancy. Current study was conducted to determine the association of polymorphism in gene of pregnancy-associated plasma protein A (PAPPA and preeclampsia. Methods: In this prospective cohort study, 134 pregnant women were consecutively enrolled and the blood sampling was performed for genetic analysis in a single lab. Then the subjects were followed-up for preeclampsia and it was seen that 34 women developed preeclampsia and the polymorphism of PAPPA gene was compared between those with and without preeclampsia. Results: The results demonstrated that despite twice higher proportion of CC condition of PAPPA in those with preeclampsia in comparison with those with normal pregnancy, there was no significant difference between two groups (P > 0.05. Conclusions: Totally, according to the obtained results, it may be concluded that polymorphism of pregnancy-associated plasma protein A is not related to occurrence of preeclampsia in pregnant women.

  18. Sequence motifs in MADS transcription factors responsible for specificity and diversification of protein-protein interaction.

    Directory of Open Access Journals (Sweden)

    Aalt D J van Dijk

    Full Text Available Protein sequences encompass tertiary structures and contain information about specific molecular interactions, which in turn determine biological functions of proteins. Knowledge about how protein sequences define interaction specificity is largely missing, in particular for paralogous protein families with high sequence similarity, such as the plant MADS domain transcription factor family. In comparison to the situation in mammalian species, this important family of transcription regulators has expanded enormously in plant species and contains over 100 members in the model plant species Arabidopsis thaliana. Here, we provide insight into the mechanisms that determine protein-protein interaction specificity for the Arabidopsis MADS domain transcription factor family, using an integrated computational and experimental approach. Plant MADS proteins have highly similar amino acid sequences, but their dimerization patterns vary substantially. Our computational analysis uncovered small sequence regions that explain observed differences in dimerization patterns with reasonable accuracy. Furthermore, we show the usefulness of the method for prediction of MADS domain transcription factor interaction networks in other plant species. Introduction of mutations in the predicted interaction motifs demonstrated that single amino acid mutations can have a large effect and lead to loss or gain of specific interactions. In addition, various performed bioinformatics analyses shed light on the way evolution has shaped MADS domain transcription factor interaction specificity. Identified protein-protein interaction motifs appeared to be strongly conserved among orthologs, indicating their evolutionary importance. We also provide evidence that mutations in these motifs can be a source for sub- or neo-functionalization. The analyses presented here take us a step forward in understanding protein-protein interactions and the interplay between protein sequences and

  19. Identification and characterization of transcript polymorphisms in soybean lines varying in oil composition and content.

    Science.gov (United States)

    Goettel, Wolfgang; Xia, Eric; Upchurch, Robert; Wang, Ming-Li; Chen, Pengyin; An, Yong-Qiang Charles

    2014-04-23

    Variation in seed oil composition and content among soybean varieties is largely attributed to differences in transcript sequences and/or transcript accumulation of oil production related genes in seeds. Discovery and analysis of sequence and expression variations in these genes will accelerate soybean oil quality improvement. In an effort to identify these variations, we sequenced the transcriptomes of soybean seeds from nine lines varying in oil composition and/or total oil content. Our results showed that 69,338 distinct transcripts from 32,885 annotated genes were expressed in seeds. A total of 8,037 transcript expression polymorphisms and 50,485 transcript sequence polymorphisms (48,792 SNPs and 1,693 small Indels) were identified among the lines. Effects of the transcript polymorphisms on their encoded protein sequences and functions were predicted. The studies also provided independent evidence that the lack of FAD2-1A gene activity and a non-synonymous SNP in the coding sequence of FAB2C caused elevated oleic acid and stearic acid levels in soybean lines M23 and FAM94-41, respectively. As a proof-of-concept, we developed an integrated RNA-seq and bioinformatics approach to identify and functionally annotate transcript polymorphisms, and demonstrated its high effectiveness for discovery of genetic and transcript variations that result in altered oil quality traits. The collection of transcript polymorphisms coupled with their predicted functional effects will be a valuable asset for further discovery of genes, gene variants, and functional markers to improve soybean oil quality.

  20. How Many Protein Sequences Fold to a Given Structure? A Coevolutionary Analysis.

    Science.gov (United States)

    Tian, Pengfei; Best, Robert B

    2017-10-17

    Quantifying the relationship between protein sequence and structure is key to understanding the protein universe. A fundamental measure of this relationship is the total number of amino acid sequences that can fold to a target protein structure, known as the "sequence capacity," which has been suggested as a proxy for how designable a given protein fold is. Although sequence capacity has been extensively studied using lattice models and theory, numerical estimates for real protein structures are currently lacking. In this work, we have quantitatively estimated the sequence capacity of 10 proteins with a variety of different structures using a statistical model based on residue-residue co-evolution to capture the variation of sequences from the same protein family. Remarkably, we find that even for the smallest protein folds, such as the WW domain, the number of foldable sequences is extremely large, exceeding the Avogadro constant. In agreement with earlier theoretical work, the calculated sequence capacity is positively correlated with the size of the protein, or better, the density of contacts. This allows the absolute sequence capacity of a given protein to be approximately predicted from its structure. On the other hand, the relative sequence capacity, i.e., normalized by the total number of possible sequences, is an extremely tiny number and is strongly anti-correlated with the protein length. Thus, although there may be more foldable sequences for larger proteins, it will be much harder to find them. Lastly, we have correlated the evolutionary age of proteins in the CATH database with their sequence capacity as predicted by our model. The results suggest a trade-off between the opposing requirements of high designability and the likelihood of a novel fold emerging by chance. Published by Elsevier Inc.

  1. Soluble polymorphic bank vole prion proteins induced by co-expression of quiescin sulfhydryl oxidase in E. coli and their aggregation behaviors.

    Science.gov (United States)

    Abskharon, Romany; Dang, Johnny; Elfarash, Ameer; Wang, Zerui; Shen, Pingping; Zou, Lewis S; Hassan, Sedky; Wang, Fei; Fujioka, Hisashi; Steyaert, Jan; Mulaj, Mentor; Surewicz, Witold K; Castilla, Joaquín; Wohlkonig, Alexandre; Zou, Wen-Quan

    2017-10-04

    The infectious prion protein (PrP Sc or prion) is derived from its cellular form (PrP C ) through a conformational transition in animal and human prion diseases. Studies have shown that the interspecies conversion of PrP C to PrP Sc is largely swayed by species barriers, which is mainly deciphered by the sequence and conformation of the proteins among species. However, the bank vole PrP C (BVPrP) is highly susceptible to PrP Sc from different species. Transgenic mice expressing BVPrP with the polymorphic isoleucine (109I) but methionine (109M) at residue 109 spontaneously develop prion disease. To explore the mechanism underlying the unique susceptibility and convertibility, we generated soluble BVPrP by co-expression of BVPrP with Quiescin sulfhydryl oxidase (QSOX) in Escherichia coli. Interestingly, rBVPrP-109M and rBVPrP-109I exhibited distinct seeded aggregation pathways and aggregate morphologies upon seeding of mouse recombinant PrP fibrils, as monitored by thioflavin T fluorescence and electron microscopy. Moreover, they displayed different aggregation behaviors induced by seeding of hamster and mouse prion strains under real-time quaking-induced conversion. Our results suggest that QSOX facilitates the formation of soluble prion protein and provide further evidence that the polymorphism at residue 109 of QSOX-induced BVPrP may be a determinant in mediating its distinct convertibility and susceptibility.

  2. Predicting the tolerated sequences for proteins and protein interfaces using RosettaBackrub flexible backbone design.

    Directory of Open Access Journals (Sweden)

    Colin A Smith

    Full Text Available Predicting the set of sequences that are tolerated by a protein or protein interface, while maintaining a desired function, is useful for characterizing protein interaction specificity and for computationally designing sequence libraries to engineer proteins with new functions. Here we provide a general method, a detailed set of protocols, and several benchmarks and analyses for estimating tolerated sequences using flexible backbone protein design implemented in the Rosetta molecular modeling software suite. The input to the method is at least one experimentally determined three-dimensional protein structure or high-quality model. The starting structure(s are expanded or refined into a conformational ensemble using Monte Carlo simulations consisting of backrub backbone and side chain moves in Rosetta. The method then uses a combination of simulated annealing and genetic algorithm optimization methods to enrich for low-energy sequences for the individual members of the ensemble. To emphasize certain functional requirements (e.g. forming a binding interface, interactions between and within parts of the structure (e.g. domains can be reweighted in the scoring function. Results from each backbone structure are merged together to create a single estimate for the tolerated sequence space. We provide an extensive description of the protocol and its parameters, all source code, example analysis scripts and three tests applying this method to finding sequences predicted to stabilize proteins or protein interfaces. The generality of this method makes many other applications possible, for example stabilizing interactions with small molecules, DNA, or RNA. Through the use of within-domain reweighting and/or multistate design, it may also be possible to use this method to find sequences that stabilize particular protein conformations or binding interactions over others.

  3. Sequence-related amplified polymorphism (SRAP) markers: A potential resource for studies in plant molecular biology1

    Science.gov (United States)

    Robarts, Daniel W. H.; Wolfe, Andrea D.

    2014-01-01

    In the past few decades, many investigations in the field of plant biology have employed selectively neutral, multilocus, dominant markers such as inter-simple sequence repeat (ISSR), random-amplified polymorphic DNA (RAPD), and amplified fragment length polymorphism (AFLP) to address hypotheses at lower taxonomic levels. More recently, sequence-related amplified polymorphism (SRAP) markers have been developed, which are used to amplify coding regions of DNA with primers targeting open reading frames. These markers have proven to be robust and highly variable, on par with AFLP, and are attained through a significantly less technically demanding process. SRAP markers have been used primarily for agronomic and horticultural purposes, developing quantitative trait loci in advanced hybrids and assessing genetic diversity of large germplasm collections. Here, we suggest that SRAP markers should be employed for research addressing hypotheses in plant systematics, biogeography, conservation, ecology, and beyond. We provide an overview of the SRAP literature to date, review descriptive statistics of SRAP markers in a subset of 171 publications, and present relevant case studies to demonstrate the applicability of SRAP markers to the diverse field of plant biology. Results of these selected works indicate that SRAP markers have the potential to enhance the current suite of molecular tools in a diversity of fields by providing an easy-to-use, highly variable marker with inherent biological significance. PMID:25202637

  4. Sequence-related amplified polymorphism (SRAP) markers: A potential resource for studies in plant molecular biology(1.).

    Science.gov (United States)

    Robarts, Daniel W H; Wolfe, Andrea D

    2014-07-01

    In the past few decades, many investigations in the field of plant biology have employed selectively neutral, multilocus, dominant markers such as inter-simple sequence repeat (ISSR), random-amplified polymorphic DNA (RAPD), and amplified fragment length polymorphism (AFLP) to address hypotheses at lower taxonomic levels. More recently, sequence-related amplified polymorphism (SRAP) markers have been developed, which are used to amplify coding regions of DNA with primers targeting open reading frames. These markers have proven to be robust and highly variable, on par with AFLP, and are attained through a significantly less technically demanding process. SRAP markers have been used primarily for agronomic and horticultural purposes, developing quantitative trait loci in advanced hybrids and assessing genetic diversity of large germplasm collections. Here, we suggest that SRAP markers should be employed for research addressing hypotheses in plant systematics, biogeography, conservation, ecology, and beyond. We provide an overview of the SRAP literature to date, review descriptive statistics of SRAP markers in a subset of 171 publications, and present relevant case studies to demonstrate the applicability of SRAP markers to the diverse field of plant biology. Results of these selected works indicate that SRAP markers have the potential to enhance the current suite of molecular tools in a diversity of fields by providing an easy-to-use, highly variable marker with inherent biological significance.

  5. CMsearch: simultaneous exploration of protein sequence space and structure space improves not only protein homology detection but also protein structure prediction

    KAUST Repository

    Cui, Xuefeng

    2016-06-15

    Motivation: Protein homology detection, a fundamental problem in computational biology, is an indispensable step toward predicting protein structures and understanding protein functions. Despite the advances in recent decades on sequence alignment, threading and alignment-free methods, protein homology detection remains a challenging open problem. Recently, network methods that try to find transitive paths in the protein structure space demonstrate the importance of incorporating network information of the structure space. Yet, current methods merge the sequence space and the structure space into a single space, and thus introduce inconsistency in combining different sources of information. Method: We present a novel network-based protein homology detection method, CMsearch, based on cross-modal learning. Instead of exploring a single network built from the mixture of sequence and structure space information, CMsearch builds two separate networks to represent the sequence space and the structure space. It then learns sequence–structure correlation by simultaneously taking sequence information, structure information, sequence space information and structure space information into consideration. Results: We tested CMsearch on two challenging tasks, protein homology detection and protein structure prediction, by querying all 8332 PDB40 proteins. Our results demonstrate that CMsearch is insensitive to the similarity metrics used to define the sequence and the structure spaces. By using HMM–HMM alignment as the sequence similarity metric, CMsearch clearly outperforms state-of-the-art homology detection methods and the CASP-winning template-based protein structure prediction methods.

  6. EST2Prot: Mapping EST sequences to proteins

    Directory of Open Access Journals (Sweden)

    Lin David M

    2006-03-01

    Full Text Available Abstract Background EST libraries are used in various biological studies, from microarray experiments to proteomic and genetic screens. These libraries usually contain many uncharacterized ESTs that are typically ignored since they cannot be mapped to known genes. Consequently, new discoveries are possibly overlooked. Results We describe a system (EST2Prot that uses multiple elements to map EST sequences to their corresponding protein products. EST2Prot uses UniGene clusters, substring analysis, information about protein coding regions in existing DNA sequences and protein database searches to detect protein products related to a query EST sequence. Gene Ontology terms, Swiss-Prot keywords, and protein similarity data are used to map the ESTs to functional descriptors. Conclusion EST2Prot extends and significantly enriches the popular UniGene mapping by utilizing multiple relations between known biological entities. It produces a mapping between ESTs and proteins in real-time through a simple web-interface. The system is part of the Biozon database and is accessible at http://biozon.org/tools/est/.

  7. Adaptive compressive learning for prediction of protein-protein interactions from primary sequence.

    Science.gov (United States)

    Zhang, Ya-Nan; Pan, Xiao-Yong; Huang, Yan; Shen, Hong-Bin

    2011-08-21

    Protein-protein interactions (PPIs) play an important role in biological processes. Although much effort has been devoted to the identification of novel PPIs by integrating experimental biological knowledge, there are still many difficulties because of lacking enough protein structural and functional information. It is highly desired to develop methods based only on amino acid sequences for predicting PPIs. However, sequence-based predictors are often struggling with the high-dimensionality causing over-fitting and high computational complexity problems, as well as the redundancy of sequential feature vectors. In this paper, a novel computational approach based on compressed sensing theory is proposed to predict yeast Saccharomyces cerevisiae PPIs from primary sequence and has achieved promising results. The key advantage of the proposed compressed sensing algorithm is that it can compress the original high-dimensional protein sequential feature vector into a much lower but more condensed space taking the sparsity property of the original signal into account. What makes compressed sensing much more attractive in protein sequence analysis is its compressed signal can be reconstructed from far fewer measurements than what is usually considered necessary in traditional Nyquist sampling theory. Experimental results demonstrate that proposed compressed sensing method is powerful for analyzing noisy biological data and reducing redundancy in feature vectors. The proposed method represents a new strategy of dealing with high-dimensional protein discrete model and has great potentiality to be extended to deal with many other complicated biological systems. Copyright © 2011 Elsevier Ltd. All rights reserved.

  8. Design of Protein Multi-specificity Using an Independent Sequence Search Reduces the Barrier to Low Energy Sequences.

    Directory of Open Access Journals (Sweden)

    Alexander M Sevy

    2015-07-01

    Full Text Available Computational protein design has found great success in engineering proteins for thermodynamic stability, binding specificity, or enzymatic activity in a 'single state' design (SSD paradigm. Multi-specificity design (MSD, on the other hand, involves considering the stability of multiple protein states simultaneously. We have developed a novel MSD algorithm, which we refer to as REstrained CONvergence in multi-specificity design (RECON. The algorithm allows each state to adopt its own sequence throughout the design process rather than enforcing a single sequence on all states. Convergence to a single sequence is encouraged through an incrementally increasing convergence restraint for corresponding positions. Compared to MSD algorithms that enforce (constrain an identical sequence on all states the energy landscape is simplified, which accelerates the search drastically. As a result, RECON can readily be used in simulations with a flexible protein backbone. We have benchmarked RECON on two design tasks. First, we designed antibodies derived from a common germline gene against their diverse targets to assess recovery of the germline, polyspecific sequence. Second, we design "promiscuous", polyspecific proteins against all binding partners and measure recovery of the native sequence. We show that RECON is able to efficiently recover native-like, biologically relevant sequences in this diverse set of protein complexes.

  9. Nonlinear analysis of sequence symmetry of beta-trefoil family proteins

    Energy Technology Data Exchange (ETDEWEB)

    Li Mingfeng [Biomolecular Physics and Modeling Group, Department of Physics, Huazhong University of Science and Technology, Wuhan 430074, Hubei (China); Huang Yanzhao [Biomolecular Physics and Modeling Group, Department of Physics, Huazhong University of Science and Technology, Wuhan 430074, Hubei (China); Xu Ruizhen [Biomolecular Physics and Modeling Group, Department of Physics, Huazhong University of Science and Technology, Wuhan 430074, Hubei (China); Xiao Yi [Biomolecular Physics and Modeling Group, Department of Physics, Huazhong University of Science and Technology, Wuhan 430074, Hubei (China)]. E-mail: yxiao@mail.hust.edu.cn

    2005-07-01

    The tertiary structures of proteins of beta-trefoil family have three-fold quasi-symmetry while their amino acid sequences appear almost at random. In the present paper we show that these amino acid sequences have hidden symmetries in fact and furthermore the degrees of these hidden symmetries are the same as those of their tertiary structures. We shall present a modified recurrence plot to reveal hidden symmetries in protein sequences. Our results can explain the contradiction in sequence-structure relations of proteins of beta-trefoil family.

  10. Correlated mutations in protein sequences: Phylogenetic and structural effects

    Energy Technology Data Exchange (ETDEWEB)

    Lapedes, A.S. [Los Alamos National Lab., NM (United States). Theoretical Div.]|[Santa Fe Inst., NM (United States); Giraud, B.G. [C.E.N. Saclay, Gif/Yvette (France). Service Physique Theorique; Liu, L.C. [Los Alamos National Lab., NM (United States). Theoretical Div.; Stormo, G.D. [Univ. of Colorado, Boulder, CO (United States). Dept. of Molecular, Cellular and Developmental Biology

    1998-12-01

    Covariation analysis of sets of aligned sequences for RNA molecules is relatively successful in elucidating RNA secondary structure, as well as some aspects of tertiary structure. Covariation analysis of sets of aligned sequences for protein molecules is successful in certain instances in elucidating certain structural and functional links, but in general, pairs of sites displaying highly covarying mutations in protein sequences do not necessarily correspond to sites that are spatially close in the protein structure. In this paper the authors identify two reasons why naive use of covariation analysis for protein sequences fails to reliably indicate sequence positions that are spatially proximate. The first reason involves the bias introduced in calculation of covariation measures due to the fact that biological sequences are generally related by a non-trivial phylogenetic tree. The authors present a null-model approach to solve this problem. The second reason involves linked chains of covariation which can result in pairs of sites displaying significant covariation even though they are not spatially proximate. They present a maximum entropy solution to this classic problem of causation versus correlation. The methodologies are validated in simulation.

  11. Genome sequence of herpes simplex virus 1 strain KOS.

    Science.gov (United States)

    Macdonald, Stuart J; Mostafa, Heba H; Morrison, Lynda A; Davido, David J

    2012-06-01

    Herpes simplex virus type 1 (HSV-1) strain KOS has been extensively used in many studies to examine HSV-1 replication, gene expression, and pathogenesis. Notably, strain KOS is known to be less pathogenic than the first sequenced genome of HSV-1, strain 17. To understand the genotypic differences between KOS and other phenotypically distinct strains of HSV-1, we sequenced the viral genome of strain KOS. When comparing strain KOS to strain 17, there are at least 1,024 small nucleotide polymorphisms (SNPs) and 172 insertions/deletions (indels). The polymorphisms observed in the KOS genome will likely provide insights into the genes, their protein products, and the cis elements that regulate the biology of this HSV-1 strain.

  12. Definition of novel GP6 polymorphisms and major difference in haplotype frequencies between populations by a combination of in-depth exon resequencing and genotyping with tag single nucleotide polymorphisms.

    Science.gov (United States)

    Watkins, N A; O'Connor, M N; Rankin, A; Jennings, N; Wilson, E; Harmer, I J; Davies, L; Smethurst, P A; Dudbridge, F; Farndale, R W; Ouwehand, W H

    2006-06-01

    Common genetic variants of cell surface receptors contribute to differences in functional responses and disease susceptibility. We have previously shown that single nucleotide polymorphisms (SNPs) in platelet glycoprotein VI (GP6) determine the extent of response to agonist. In addition, SNPs in the GP6 gene have been proposed as risk factors for coronary artery disease. To completely characterize genetic variation in the GP6 gene we generated a high-resolution SNP map by sequencing the promoter, exons and consensus splice sequences in 94 non-related Caucasoids. In addition, we sequenced DNA encoding the ligand-binding domains of GP6 from non-human primates to determine the level of evolutionary conservation. Eighteen SNPs were identified, six of which encoded amino acid substitutions in the mature form of the protein. The single non-synonymous SNP identified in the exons encoding the ligand-binding domains, encoding for a 103Leu > Val substitution, resulted in reduced ligand binding. Two common protein isoforms were confirmed in Caucasoid with frequencies of 0.82 and 0.15. Variation at the GP6 locus was characterized further by determining SNP frequency in over 2000 individuals from different ethnic backgrounds. The SNPs were polymorphic in all populations studied although significant differences in allele frequencies were observed. Twelve additional GP6 protein isoforms were identified from the genotyping results and, despite extensive variation in GP6, the sequence of the ligand-binding domains is conserved. Sequences from non-human primates confirmed this observation. These data provide valuable information for the optimal selection of genetic variants for use in future association studies.

  13. DNA Polymorphism of Insulin-like Growth Factor-binding Protein-3 Gene and Its Association with Cashmere Traits in Cashmere Goats

    Directory of Open Access Journals (Sweden)

    Haiying Liu

    2012-11-01

    Full Text Available Insulin-like growth factor binding protein-3 (IGFBP-3 gene is important for regulation of growth and development in mammals. The present investigation was carried out to study DNA polymorphism by PCR-RFLP of IGFBP-3 gene and its effect on fibre traits of Chinese Inner Mongolian cashmere goats. The fibre traits data investigated were cashmere fibre diameter, combed cashmere weight, cashmere fibre length and guard hair length. Four hundred and forty-four animals were used to detect polymorphisms in the hircine IGFBP-3 gene. A 316-bp fragment of the IGFBP-3 gene in exon 2 was amplified and digested with HaeIII restriction enzyme. Three patterns of restriction fragments were observed in the populations. The frequency of AA, AB and BB genotypes was 0.58, 0.33 and 0.09 respectively. The allelic frequency of the A and B allele was 0.75 and 0.25 respectively. Nucleotide sequencing revealed a C>G transition in the exon 2 region of the IGFBP-3 gene resulting in R158G change which caused the polymorphism. Least squares analysis revealed a significant effect of genotypes on cashmere weight (p0.05. The animals of AB and BB genotypes showed higher cashmere weight, cashmere fibre length and hair length than the animals possessing AA genotype. These results suggested that polymorphisms in the hircine IGFBP-3 gene might be a potential molecular marker for cashmere weight in cashmere goats.

  14. Sequence characterization of heat shock protein gene of Cyclospora cayetanensis isolates from Nepal, Mexico, and Peru.

    Science.gov (United States)

    Sulaiman, Irshad M; Torres, Patricia; Simpson, Steven; Kerdahi, Khalil; Ortega, Ynes

    2013-04-01

    We have described the development of a 2-step nested PCR protocol based on the characterization of the 70-kDa heat shock protein (HSP70) gene for rapid detection of the human-pathogenic Cyclospora cayetanensis parasite. We tested and validated these newly designed primer sets by PCR amplification followed by nucleotide sequencing of PCR-amplified HSP70 fragments belonging to 16 human C. cayetanensis isolates from 3 different endemic regions that include Nepal, Mexico, and Peru. No genetic polymorphism was observed among the isolates at the characterized regions of the HSP70 locus. This newly developed HSP70 gene-based nested PCR protocol provides another useful genetic marker for the rapid detection of C. cayetanensis in the future.

  15. Development of highly polymorphic simple sequence repeat markers using genome-wide microsatellite variant analysis in Foxtail millet [Setaria italica (L.) P. Beauv].

    Science.gov (United States)

    Zhang, Shuo; Tang, Chanjuan; Zhao, Qiang; Li, Jing; Yang, Lifang; Qie, Lufeng; Fan, Xingke; Li, Lin; Zhang, Ning; Zhao, Meicheng; Liu, Xiaotong; Chai, Yang; Zhang, Xue; Wang, Hailong; Li, Yingtao; Li, Wen; Zhi, Hui; Jia, Guanqing; Diao, Xianmin

    2014-01-28

    Foxtail millet (Setaria italica (L.) Beauv.) is an important gramineous grain-food and forage crop. It is grown worldwide for human and livestock consumption. Its small genome and diploid nature have led to foxtail millet fast becoming a novel model for investigating plant architecture, drought tolerance and C4 photosynthesis of grain and bioenergy crops. Therefore, cost-effective, reliable and highly polymorphic molecular markers covering the entire genome are required for diversity, mapping and functional genomics studies in this model species. A total of 5,020 highly repetitive microsatellite motifs were isolated from the released genome of the genotype 'Yugu1' by sequence scanning. Based on sequence comparison between S. italica and S. viridis, a set of 788 SSR primer pairs were designed. Of these primers, 733 produced reproducible amplicons and were polymorphic among 28 Setaria genotypes selected from diverse geographical locations. The number of alleles detected by these SSR markers ranged from 2 to 16, with an average polymorphism information content of 0.67. The result obtained by neighbor-joining cluster analysis of 28 Setaria genotypes, based on Nei's genetic distance of the SSR data, showed that these SSR markers are highly polymorphic and effective. A large set of highly polymorphic SSR markers were successfully and efficiently developed based on genomic sequence comparison between different genotypes of the genus Setaria. The large number of new SSR markers and their placement on the physical map represent a valuable resource for studying diversity, constructing genetic maps, functional gene mapping, QTL exploration and molecular breeding in foxtail millet and its closely related species.

  16. C-reactive protein gene polymorphisms and myocardial infarction risk: a meta-analysis and meta-regression.

    Science.gov (United States)

    Zhu, Yanbin; Liu, Tongku; He, Haitao; Sun, Yuqing; Zhuo, Fengling

    2013-12-01

    C-reactive protein (CRP), the classic acute-phase protein, plays an important role in the etiology of myocardial infarction (MI). Emerging evidence has shown that the common polymorphisms in the CRP gene may influence an individual's susceptibility to MI; but individually published studies showed inconclusive results. This meta-analysis aimed to derive a more precise estimation of the associations between CRP gene polymorphisms and MI risk. A literature search of PubMed, Embase, Web of Science, and China BioMedicine (CBM) databases was conducted on articles published before June 1st, 2013. Crude odds ratio (OR) with 95% confidence interval (CI) were calculated. Nine case-control studies were included with a total of 2992 MI patients and 4711 healthy controls. The meta-analysis results indicated that CRP rs3093059 (T>C) polymorphism was associated with decreased risk of MI, especially among Asian populations. However, similar associations were not observed in CRP rs1800947 (G>C) and rs2794521 (G>A) polymorphisms (all p>0.05) among both Asian and Caucasian populations. Univariate and multivariate meta-regression analyses showed that ethnicity may be a major source of heterogeneity. No publication bias was detected in this meta-analysis. In conclusion, the current meta-analysis indicates that CRP rs3093059 (T>C) polymorphism may be associated with decreased risk of MI, especially among Asian populations.

  17. Erratum Associations of POU1F1 gene polymorphisms and protein ...

    Indian Academy of Sciences (India)

    Associations of POU1F1 gene polymorphisms and protein structure changes with growth traits and blood metabolites in two Iranian sheep breeds. Mostafa Sadeghi, Ali Jalil-Sarghale and Mohammed Moradi-Shahrbabak. J. Genet. 93, 831–835. The erratum published in the March 2015 issue to this article did not point out ...

  18. Matrix Gla Protein polymorphism, but not concentrations, is associated with radiographic hand osteoarthritis

    Science.gov (United States)

    Objective. Factors associated with mineralization and osteophyte formation in osteoarthritis (OA) are incompletely understood. Genetic polymorphisms of matrix Gla protein (MGP), a mineralization inhibitor, have been associated clinically with conditions of abnormal calcification. We therefore evalua...

  19. Single nucleotide polymorphism in transcriptional regulatory regions and expression of environmentally responsive genes

    International Nuclear Information System (INIS)

    Wang, Xuting; Tomso, Daniel J.; Liu Xuemei; Bell, Douglas A.

    2005-01-01

    Single nucleotide polymorphisms (SNPs) in the human genome are DNA sequence variations that can alter an individual's response to environmental exposure. SNPs in gene coding regions can lead to changes in the biological properties of the encoded protein. In contrast, SNPs in non-coding gene regulatory regions may affect gene expression levels in an allele-specific manner, and these functional polymorphisms represent an important but relatively unexplored class of genetic variation. The main challenge in analyzing these SNPs is a lack of robust computational and experimental methods. Here, we first outline mechanisms by which genetic variation can impact gene regulation, and review recent findings in this area; then, we describe a methodology for bioinformatic discovery and functional analysis of regulatory SNPs in cis-regulatory regions using the assembled human genome sequence and databases on sequence polymorphism and gene expression. Our method integrates SNP and gene databases and uses a set of computer programs that allow us to: (1) select SNPs, from among the >9 million human SNPs in the NCBI dbSNP database, that are similar to cis-regulatory element (RE) consensus sequences; (2) map the selected dbSNP entries to the human genome assembly in order to identify polymorphic REs near gene start sites; (3) prioritize the candidate polymorphic RE containing genes by searching the existing genotype and gene expression data sets. The applicability of this system has been demonstrated through studies on p53 responsive elements and is being extended to additional pathways and environmentally responsive genes

  20. Finding the right coverage : The impact of coverage and sequence quality on single nucleotide polymorphism genotyping error rates

    NARCIS (Netherlands)

    Fountain, Emily D.; Pauli, Jonathan N.; Reid, Brendan N.; Palsboll, Per J.; Peery, M. Zachariah

    Restriction-enzyme-based sequencing methods enable the genotyping of thousands of single nucleotide polymorphism (SNP) loci in nonmodel organisms. However, in contrast to traditional genetic markers, genotyping error rates in SNPs derived from restriction-enzyme-based methods remain largely unknown.

  1. Investigation of heart proteome of different consomic mouse strains. Testing the effect of polymorphisms on the proteome-wide trans-variation of proteins

    Directory of Open Access Journals (Sweden)

    Stefanie Forler

    2015-06-01

    Full Text Available We investigated to which extent polymorphisms of an individual affect the proteomic network. Consomic mouse strains (CS were used to study the trans-effect of the cis-variant (polymorphic proteins of the strain PWD/Ph on the proteins of the host strain C57BL/6J. The cardiac proteome of ten CSs was analyzed by 2-DE and MS. Cis-variant PWD proteins altered a high number of C57BL/6J proteins, but the number of trans-variant proteins differed considerably between different CSs. Cardiac hypertrophy was induced in CSs. We found that high variability of the proteome, as induced by polymorphisms in CS14, acts protective against the complex disease.

  2. Deep sequencing methods for protein engineering and design.

    Science.gov (United States)

    Wrenbeck, Emily E; Faber, Matthew S; Whitehead, Timothy A

    2017-08-01

    The advent of next-generation sequencing (NGS) has revolutionized protein science, and the development of complementary methods enabling NGS-driven protein engineering have followed. In general, these experiments address the functional consequences of thousands of protein variants in a massively parallel manner using genotype-phenotype linked high-throughput functional screens followed by DNA counting via deep sequencing. We highlight the use of information rich datasets to engineer protein molecular recognition. Examples include the creation of multiple dual-affinity Fabs targeting structurally dissimilar epitopes and engineering of a broad germline-targeted anti-HIV-1 immunogen. Additionally, we highlight the generation of enzyme fitness landscapes for conducting fundamental studies of protein behavior and evolution. We conclude with discussion of technological advances. Copyright © 2016 Elsevier Ltd. All rights reserved.

  3. Utilization of a cloned alphoid repeating sequence of human DNA in the study of polymorphism of chromosomal heterochromatin regions

    International Nuclear Information System (INIS)

    Kruminya, A.R.; Kroshkina, V.G.; Yurov, Yu.B.; Aleksandrov, I.A.; Mitkevich, S.P.; Gindilis, V.M.

    1988-01-01

    The chromosomal distribution of the cloned PHS05 fragment of human alphoid DNA was studied by in situ hybridization in 38 individuals. It was shown that this DNA fraction is primarily localized in the pericentric regions of practically all chromosomes of the set. Significant interchromosomal differences and a weakly expressed interindividual polymorphism were discovered in the copying ability of this class of repeating DNA sequences; associations were not found between the results of hybridization and the pattern of Q-polymorphism

  4. G-protein-coupled estrogen receptor-30 gene polymorphisms are associated with uterine leiomyoma risk.

    Science.gov (United States)

    Kasap, Burcu; Öztürk Turhan, Nilgün; Edgünlü, Tuba; Duran, Müzeyyen; Akbaba, Eren; Öner, Gökalp

    2016-01-06

    The G-protein-coupled estrogen receptor (GPR30, GPER-1) is a member of the G-protein-coupled receptor 1 family and is expressed significantly in uterine leiomyomas. To understand the relationship between GPR30 single nucleotide polymorphisms and the risk of leiomyoma, we measured the follicle-stimulating hormone (FSH) and estradiol (E2) levels of 78 perimenopausal healthy women and 111 perimenopausal women with leiomyomas. The participants' leiomyoma number and volume were recorded. DNA was extracted from whole blood with a GeneJET Genomic DNA Purification Kit. An amplification-refractory mutation system polymerase chain reaction approach was used for genotyping of the GPR30 gene (rs3808350, rs3808351, and rs11544331). The differences in genotype and allele frequencies between the leiomyoma and control groups were calculated using the chi-square (χ2) and Fischer's exact test. The median FSH level was higher in controls (63 vs. 10 IU/L, p=0.000), whereas the median E2 level was higher in the leiomyoma group (84 vs. 9.1 pg/mL, p=0.000). The G allele of rs3808351 and the GG genotype of both the rs3808350 and rs3808351 polymorphisms and the GGC haplotype increased the risk of developing leiomyoma. There was no significant difference in genotype frequencies or leiomyoma volume. However, the GG genotype of the GPR30 rs3808351 polymorphism and G allele of the GPR30 rs3808351 polymorphism were associated with the risk of having a single leiomyoma. Our results suggest that the presence of the GG genotype of the GPR30 rs3808351 polymorphism and the G allele of the GPR30 rs3808351 polymorphism affect the characteristics and development of leiomyomas in the Turkish population.

  5. GNBP domain of Anopheles darlingi: are polymorphic inversions and gene variation related to adaptive evolution?

    Science.gov (United States)

    Bridi, L C; Rafael, M S

    2016-02-01

    Anopheles darlingi is the main malaria vector in humans in South America. In the Amazon basin, it lives along the banks of rivers and lakes, which responds to the annual hydrological cycle (dry season and rainy season). In these breeding sites, the larvae of this mosquito feed on decomposing organic and microorganisms, which can be pathogenic and trigger the activation of innate immune system pathways, such as proteins Gram-negative binding protein (GNBP). Such environmental changes affect the occurrence of polymorphic inversions especially at the heterozygote frequency, which confer adaptative advantage compared to homozygous inversions. We mapped the GNBP probe to the An. darlingi 2Rd inversion by fluorescent in situ hybridization (FISH), which was a good indicator of the GNBP immune response related to the chromosomal polymorphic inversions and adaptative evolution. To better understand the evolutionary relations and time of divergence of the GNBP of An. darlingi, we compared it with nine other mosquito GNBPs. The results of the phylogenetic analysis of the GNBP sequence between the species of mosquitoes demonstrated three clades. Clade I and II included the GNBPB5 sequence, and clade III the sequence of GNBPB1. Most of these sequences of GNBP analyzed were homologous with that of subfamily B, including that of An. gambiae (87 %), therefore suggesting that GNBP of An. darling belongs to subfamily B. This work helps us understand the role of inversion polymorphism in evolution of An. darlingi.

  6. DNA sequence polymorphisms within the bovine guanine nucleotide-binding protein Gs subunit alpha (Gsα-encoding (GNAS genomic imprinting domain are associated with performance traits

    Directory of Open Access Journals (Sweden)

    Mullen Michael P

    2011-01-01

    Full Text Available Abstract Background Genes which are epigenetically regulated via genomic imprinting can be potential targets for artificial selection during animal breeding. Indeed, imprinted loci have been shown to underlie some important quantitative traits in domestic mammals, most notably muscle mass and fat deposition. In this candidate gene study, we have identified novel associations between six validated single nucleotide polymorphisms (SNPs spanning a 97.6 kb region within the bovine guanine nucleotide-binding protein Gs subunit alpha gene (GNAS domain on bovine chromosome 13 and genetic merit for a range of performance traits in 848 progeny-tested Holstein-Friesian sires. The mammalian GNAS domain consists of a number of reciprocally-imprinted, alternatively-spliced genes which can play a major role in growth, development and disease in mice and humans. Based on the current annotation of the bovine GNAS domain, four of the SNPs analysed (rs43101491, rs43101493, rs43101485 and rs43101486 were located upstream of the GNAS gene, while one SNP (rs41694646 was located in the second intron of the GNAS gene. The final SNP (rs41694656 was located in the first exon of transcripts encoding the putative bovine neuroendocrine-specific protein NESP55, resulting in an aspartic acid-to-asparagine amino acid substitution at amino acid position 192. Results SNP genotype-phenotype association analyses indicate that the single intronic GNAS SNP (rs41694646 is associated (P ≤ 0.05 with a range of performance traits including milk yield, milk protein yield, the content of fat and protein in milk, culled cow carcass weight and progeny carcass conformation, measures of animal body size, direct calving difficulty (i.e. difficulty in calving due to the size of the calf and gestation length. Association (P ≤ 0.01 with direct calving difficulty (i.e. due to calf size and maternal calving difficulty (i.e. due to the maternal pelvic width size was also observed at the rs

  7. Prevalence of single nucleotide polymorphism among 27 diverse alfalfa genotypes as assessed by transcriptome sequencing

    Directory of Open Access Journals (Sweden)

    Li Xuehui

    2012-10-01

    Full Text Available Abstract Background Alfalfa, a perennial, outcrossing species, is a widely planted forage legume producing highly nutritious biomass. Currently, improvement of cultivated alfalfa mainly relies on recurrent phenotypic selection. Marker assisted breeding strategies can enhance alfalfa improvement efforts, particularly if many genome-wide markers are available. Transcriptome sequencing enables efficient high-throughput discovery of single nucleotide polymorphism (SNP markers for a complex polyploid species. Result The transcriptomes of 27 alfalfa genotypes, including elite breeding genotypes, parents of mapping populations, and unimproved wild genotypes, were sequenced using an Illumina Genome Analyzer IIx. De novo assembly of quality-filtered 72-bp reads generated 25,183 contigs with a total length of 26.8 Mbp and an average length of 1,065 bp, with an average read depth of 55.9-fold for each genotype. Overall, 21,954 (87.2% of the 25,183 contigs represented 14,878 unique protein accessions. Gene ontology (GO analysis suggested that a broad diversity of genes was represented in the resulting sequences. The realignment of individual reads to the contigs enabled the detection of 872,384 SNPs and 31,760 InDels. High resolution melting (HRM analysis was used to validate 91% of 192 putative SNPs identified by sequencing. Both allelic variants at about 95% of SNP sites identified among five wild, unimproved genotypes are still present in cultivated alfalfa, and all four US breeding programs also contain a high proportion of these SNPs. Thus, little evidence exists among this dataset for loss of significant DNA sequence diversity from either domestication or breeding of alfalfa. Structure analysis indicated that individuals from the subspecies falcata, the diploid subspecies caerulea, and the tetraploid subspecies sativa (cultivated tetraploid alfalfa were clearly separated. Conclusion We used transcriptome sequencing to discover large numbers of SNPs

  8. Biophysical and structural considerations for protein sequence evolution

    Directory of Open Access Journals (Sweden)

    Grahnen Johan A

    2011-12-01

    Full Text Available Abstract Background Protein sequence evolution is constrained by the biophysics of folding and function, causing interdependence between interacting sites in the sequence. However, current site-independent models of sequence evolutions do not take this into account. Recent attempts to integrate the influence of structure and biophysics into phylogenetic models via statistical/informational approaches have not resulted in expected improvements in model performance. This suggests that further innovations are needed for progress in this field. Results Here we develop a coarse-grained physics-based model of protein folding and binding function, and compare it to a popular informational model. We find that both models violate the assumption of the native sequence being close to a thermodynamic optimum, causing directional selection away from the native state. Sampling and simulation show that the physics-based model is more specific for fold-defining interactions that vary less among residue type. The informational model diffuses further in sequence space with fewer barriers and tends to provide less support for an invariant sites model, although amino acid substitutions are generally conservative. Both approaches produce sequences with natural features like dN/dS Conclusions Simple coarse-grained models of protein folding can describe some natural features of evolving proteins but are currently not accurate enough to use in evolutionary inference. This is partly due to improper packing of the hydrophobic core. We suggest possible improvements on the representation of structure, folding energy, and binding function, as regards both native and non-native conformations, and describe a large number of possible applications for such a model.

  9. Systematic analysis of protein identity between Zika virus and other arthropod-borne viruses.

    Science.gov (United States)

    Chang, Hsiao-Han; Huber, Roland G; Bond, Peter J; Grad, Yonatan H; Camerini, David; Maurer-Stroh, Sebastian; Lipsitch, Marc

    2017-07-01

    To analyse the proportions of protein identity between Zika virus and dengue, Japanese encephalitis, yellow fever, West Nile and chikungunya viruses as well as polymorphism between different Zika virus strains. We used published protein sequences for the Zika virus and obtained protein sequences for the other viruses from the National Center for Biotechnology Information (NCBI) protein database or the NCBI virus variation resource. We used BLASTP to find regions of identity between viruses. We quantified the identity between the Zika virus and each of the other viruses, as well as within-Zika virus polymorphism for all amino acid k -mers across the proteome, with k ranging from 6 to 100. We assessed accessibility of protein fragments by calculating the solvent accessible surface area for the envelope and nonstructural-1 (NS1) proteins. In total, we identified 294 Zika virus protein fragments with both low proportion of identity with other viruses and low levels of polymorphisms among Zika virus strains. The list includes protein fragments from all Zika virus proteins, except NS3. NS4A has the highest number (190 k -mers) of protein fragments on the list. We provide a candidate list of protein fragments that could be used when developing a sensitive and specific serological test to detect previous Zika virus infections.

  10. Single nucleotide polymorphisms (SNPs in coding regions of canine dopamine- and serotonin-related genes

    Directory of Open Access Journals (Sweden)

    Lingaas Frode

    2008-01-01

    Full Text Available Abstract Background Polymorphism in genes of regulating enzymes, transporters and receptors of the neurotransmitters of the central nervous system have been associated with altered behaviour, and single nucleotide polymorphisms (SNPs represent the most frequent type of genetic variation. The serotonin and dopamine signalling systems have a central influence on different behavioural phenotypes, both of invertebrates and vertebrates, and this study was undertaken in order to explore genetic variation that may be associated with variation in behaviour. Results Single nucleotide polymorphisms in canine genes related to behaviour were identified by individually sequencing eight dogs (Canis familiaris of different breeds. Eighteen genes from the dopamine and the serotonin systems were screened, revealing 34 SNPs distributed in 14 of the 18 selected genes. A total of 24,895 bp coding sequence was sequenced yielding an average frequency of one SNP per 732 bp (1/732. A total of 11 non-synonymous SNPs (nsSNPs, which may be involved in alteration of protein function, were detected. Of these 11 nsSNPs, six resulted in a substitution of amino acid residue with concomitant change in structural parameters. Conclusion We have identified a number of coding SNPs in behaviour-related genes, several of which change the amino acids of the proteins. Some of the canine SNPs exist in codons that are evolutionary conserved between five compared species, and predictions indicate that they may have a functional effect on the protein. The reported coding SNP frequency of the studied genes falls within the range of SNP frequencies reported earlier in the dog and other mammalian species. Novel SNPs are presented and the results show a significant genetic variation in expressed sequences in this group of genes. The results can contribute to an improved understanding of the genetics of behaviour.

  11. MIPS: a database for protein sequences, homology data and yeast genome information.

    Science.gov (United States)

    Mewes, H W; Albermann, K; Heumann, K; Liebl, S; Pfeiffer, F

    1997-01-01

    The MIPS group (Martinsried Institute for Protein Sequences) at the Max-Planck-Institute for Biochemistry, Martinsried near Munich, Germany, collects, processes and distributes protein sequence data within the framework of the tripartite association of the PIR-International Protein Sequence Database (,). MIPS contributes nearly 50% of the data input to the PIR-International Protein Sequence Database. The database is distributed on CD-ROM together with PATCHX, an exhaustive supplement of unique, unverified protein sequences from external sources compiled by MIPS. Through its WWW server (http://www.mips.biochem.mpg.de/ ) MIPS permits internet access to sequence databases, homology data and to yeast genome information. (i) Sequence similarity results from the FASTA program () are stored in the FASTA database for all proteins from PIR-International and PATCHX. The database is dynamically maintained and permits instant access to FASTA results. (ii) Starting with FASTA database queries, proteins have been classified into families and superfamilies (PROT-FAM). (iii) The HPT (hashed position tree) data structure () developed at MIPS is a new approach for rapid sequence and pattern searching. (iv) MIPS provides access to the sequence and annotation of the complete yeast genome (), the functional classification of yeast genes (FunCat) and its graphical display, the 'Genome Browser' (). A CD-ROM based on the JAVA programming language providing dynamic interactive access to the yeast genome and the related protein sequences has been compiled and is available on request. PMID:9016498

  12. Formatt: Correcting protein multiple structural alignments by incorporating sequence alignment

    Directory of Open Access Journals (Sweden)

    Daniels Noah M

    2012-10-01

    Full Text Available Abstract Background The quality of multiple protein structure alignments are usually computed and assessed based on geometric functions of the coordinates of the backbone atoms from the protein chains. These purely geometric methods do not utilize directly protein sequence similarity, and in fact, determining the proper way to incorporate sequence similarity measures into the construction and assessment of protein multiple structure alignments has proved surprisingly difficult. Results We present Formatt, a multiple structure alignment based on the Matt purely geometric multiple structure alignment program, that also takes into account sequence similarity when constructing alignments. We show that Formatt outperforms Matt and other popular structure alignment programs on the popular HOMSTRAD benchmark. For the SABMark twilight zone benchmark set that captures more remote homology, Formatt and Matt outperform other programs; depending on choice of embedded sequence aligner, Formatt produces either better sequence and structural alignments with a smaller core size than Matt, or similarly sized alignments with better sequence similarity, for a small cost in average RMSD. Conclusions Considering sequence information as well as purely geometric information seems to improve quality of multiple structure alignments, though defining what constitutes the best alignment when sequence and structural measures would suggest different alignments remains a difficult open question.

  13. Polymorphism in ovine ANXA9 gene and physic-chemical properties and the fraction of protein in milk.

    Science.gov (United States)

    Pecka-Kiełb, Ewa; Czerniawska-Piątkowska, Ewa; Kowalewska-Łuczak, Inga; Vasil, Milan

    2018-04-16

    Annexin A9 (ANXA9) is a specific fatty acid transport protein. ANXA9 gene is expressed in various tissues, including secretory tissue and mammary glands. The association between three SNPs of the ANXA9 gene and sheep's milk compositions was assessed. Genotype analysis was performed with the use of PCR-RFLP method. The studied ANXA9 polymorphisms had the following MAF (Major Allele Frequency): SNP1: allele G 0,66; SNP2: allele G 0,54; SNP3: allele C 0,57. The study found the most desired profile of protein fractions, namely an increased kappa-casein fractions and a decreased level of whey protein in sheep's milk for SNP1 and SNP3 polymorphisms. Sheep with the SNP1 GA genotype had the highest (P <0.05) content of fat and dry matter in milk. AXNA9 gene polymorphism did not influence the levels of protein, lactose or urea in sheep's milk. The information contained in this study may be useful for determining the impact of the ANXA9 gene on sheep's milk. The ANXA9 SNP1 and SNP3 polymorphisms results could be included in the breeding programs to select the sheep with the genotypes ensuring the highest kappa-casein levels in milk. However, it is worth conducting further research on ANXA9 and milk composition in larger herds of animals and various breeds of sheep. This article is protected by copyright. All rights reserved.

  14. Single-molecule protein sequencing through fingerprinting: computational assessment

    Science.gov (United States)

    Yao, Yao; Docter, Margreet; van Ginkel, Jetty; de Ridder, Dick; Joo, Chirlmin

    2015-10-01

    Proteins are vital in all biological systems as they constitute the main structural and functional components of cells. Recent advances in mass spectrometry have brought the promise of complete proteomics by helping draft the human proteome. Yet, this commonly used protein sequencing technique has fundamental limitations in sensitivity. Here we propose a method for single-molecule (SM) protein sequencing. A major challenge lies in the fact that proteins are composed of 20 different amino acids, which demands 20 molecular reporters. We computationally demonstrate that it suffices to measure only two types of amino acids to identify proteins and suggest an experimental scheme using SM fluorescence. When achieved, this highly sensitive approach will result in a paradigm shift in proteomics, with major impact in the biological and medical sciences.

  15. Single-molecule protein sequencing through fingerprinting: computational assessment

    International Nuclear Information System (INIS)

    Yao, Yao; Docter, Margreet; Van Ginkel, Jetty; Joo, Chirlmin; De Ridder, Dick

    2015-01-01

    Proteins are vital in all biological systems as they constitute the main structural and functional components of cells. Recent advances in mass spectrometry have brought the promise of complete proteomics by helping draft the human proteome. Yet, this commonly used protein sequencing technique has fundamental limitations in sensitivity. Here we propose a method for single-molecule (SM) protein sequencing. A major challenge lies in the fact that proteins are composed of 20 different amino acids, which demands 20 molecular reporters. We computationally demonstrate that it suffices to measure only two types of amino acids to identify proteins and suggest an experimental scheme using SM fluorescence. When achieved, this highly sensitive approach will result in a paradigm shift in proteomics, with major impact in the biological and medical sciences. (paper)

  16. Identification of prion protein gene polymorphisms in goats from Italian scrapie outbreaks

    NARCIS (Netherlands)

    Acutis, P.L.; Bossers, A.; Priem, J.; Riina, M.V.; Peletto, S.; Mazza, M.; Casalone, C.; Forloni, G.; Ru, G.; Caramelli, M.

    2006-01-01

    Susceptibility to scrapie in sheep is influenced by polymorphisms of the prion protein (PrP) gene, whereas no strong association between genetics and scrapie has yet been determined in goats due to the limited number of studies on these animals. In this case¿control study on 177 goats from six

  17. Genome-wide association study using high-density single nucleotide polymorphism arrays and whole-genome sequences for clinical mastitis traits in dairy cattle.

    Science.gov (United States)

    Sahana, G; Guldbrandtsen, B; Thomsen, B; Holm, L-E; Panitz, F; Brøndum, R F; Bendixen, C; Lund, M S

    2014-11-01

    Mastitis is a mammary disease that frequently affects dairy cattle. Despite considerable research on the development of effective prevention and treatment strategies, mastitis continues to be a significant issue in bovine veterinary medicine. To identify major genes that affect mastitis in dairy cattle, 6 chromosomal regions on Bos taurus autosome (BTA) 6, 13, 16, 19, and 20 were selected from a genome scan for 9 mastitis phenotypes using imputed high-density single nucleotide polymorphism arrays. Association analyses using sequence-level variants for the 6 targeted regions were carried out to map causal variants using whole-genome sequence data from 3 breeds. The quantitative trait loci (QTL) discovery population comprised 4,992 progeny-tested Holstein bulls, and QTL were confirmed in 4,442 Nordic Red and 1,126 Jersey cattle. The targeted regions were imputed to the sequence level. The highest association signal for clinical mastitis was observed on BTA 6 at 88.97 Mb in Holstein cattle and was confirmed in Nordic Red cattle. The peak association region on BTA 6 contained 2 genes: vitamin D-binding protein precursor (GC) and neuropeptide FF receptor 2 (NPFFR2), which, based on known biological functions, are good candidates for affecting mastitis. However, strong linkage disequilibrium in this region prevented conclusive determination of the causal gene. A different QTL on BTA 6 located at 88.32 Mb in Holstein cattle affected mastitis. In addition, QTL on BTA 13 and 19 were confirmed to segregate in Nordic Red cattle and QTL on BTA 16 and 20 were confirmed in Jersey cattle. Although several candidate genes were identified in these targeted regions, it was not possible to identify a gene or polymorphism as the causal factor for any of these regions. Copyright © 2014 American Dairy Science Association. Published by Elsevier Inc. All rights reserved.

  18. Prediction of Protein Structural Classes for Low-Similarity Sequences Based on Consensus Sequence and Segmented PSSM

    Directory of Open Access Journals (Sweden)

    Yunyun Liang

    2015-01-01

    Full Text Available Prediction of protein structural classes for low-similarity sequences is useful for understanding fold patterns, regulation, functions, and interactions of proteins. It is well known that feature extraction is significant to prediction of protein structural class and it mainly uses protein primary sequence, predicted secondary structure sequence, and position-specific scoring matrix (PSSM. Currently, prediction solely based on the PSSM has played a key role in improving the prediction accuracy. In this paper, we propose a novel method called CSP-SegPseP-SegACP by fusing consensus sequence (CS, segmented PsePSSM, and segmented autocovariance transformation (ACT based on PSSM. Three widely used low-similarity datasets (1189, 25PDB, and 640 are adopted in this paper. Then a 700-dimensional (700D feature vector is constructed and the dimension is decreased to 224D by using principal component analysis (PCA. To verify the performance of our method, rigorous jackknife cross-validation tests are performed on 1189, 25PDB, and 640 datasets. Comparison of our results with the existing PSSM-based methods demonstrates that our method achieves the favorable and competitive performance. This will offer an important complementary to other PSSM-based methods for prediction of protein structural classes for low-similarity sequences.

  19. Sequence polymorphism can produce serious artefacts in real-time PCR assays: hard lessons from Pacific oysters

    Directory of Open Access Journals (Sweden)

    Camara Mark D

    2008-05-01

    Full Text Available Abstract Background Since it was first described in the mid-1990s, quantitative real time PCR (Q-PCR has been widely used in many fields of biomedical research and molecular diagnostics. This method is routinely used to validate whole transcriptome analyses such as DNA microarrays, suppressive subtractive hybridization (SSH or differential display techniques such as cDNA-AFLP (Amplification Fragment Length Polymorphism. Despite efforts to optimize the methodology, misleading results are still possible, even when standard optimization approaches are followed. Results As part of a larger project aimed at elucidating transcriptome-level responses of Pacific oysters (Crassostrea gigas to various environmental stressors, we used microarrays and cDNA-AFLP to identify Expressed Sequence Tag (EST fragments that are differentially expressed in response to bacterial challenge in two heat shock tolerant and two heat shock sensitive full-sib oyster families. We then designed primers for these differentially expressed ESTs in order to validate the results using Q-PCR. For two of these ESTs we tested fourteen primer pairs each and using standard optimization methods (i.e. melt-curve analysis to ensure amplification of a single product, determined that of the fourteen primer pairs tested, six and nine pairs respectively amplified a single product and were thus acceptable for further testing. However, when we used these primers, we obtained different statistical outcomes among primer pairs, raising unexpected but serious questions about their reliability. We hypothesize that as a consequence of high levels of sequence polymorphism in Pacific oysters, Q-PCR amplification is sub-optimal in some individuals because sequence variants in priming sites results in poor primer binding and amplification in some individuals. This issue is similar to the high frequency of null alleles observed for microsatellite markers in Pacific oysters. Conclusion This study highlights

  20. Scoring protein relationships in functional interaction networks predicted from sequence data.

    Directory of Open Access Journals (Sweden)

    Gaston K Mazandu

    Full Text Available UNLABELLED: The abundance of diverse biological data from various sources constitutes a rich source of knowledge, which has the power to advance our understanding of organisms. This requires computational methods in order to integrate and exploit these data effectively and elucidate local and genome wide functional connections between protein pairs, thus enabling functional inferences for uncharacterized proteins. These biological data are primarily in the form of sequences, which determine functions, although functional properties of a protein can often be predicted from just the domains it contains. Thus, protein sequences and domains can be used to predict protein pair-wise functional relationships, and thus contribute to the function prediction process of uncharacterized proteins in order to ensure that knowledge is gained from sequencing efforts. In this work, we introduce information-theoretic based approaches to score protein-protein functional interaction pairs predicted from protein sequence similarity and conserved protein signature matches. The proposed schemes are effective for data-driven scoring of connections between protein pairs. We applied these schemes to the Mycobacterium tuberculosis proteome to produce a homology-based functional network of the organism with a high confidence and coverage. We use the network for predicting functions of uncharacterised proteins. AVAILABILITY: Protein pair-wise functional relationship scores for Mycobacterium tuberculosis strain CDC1551 sequence data and python scripts to compute these scores are available at http://web.cbio.uct.ac.za/~gmazandu/scoringschemes.

  1. Inverse statistical physics of protein sequences: a key issues review.

    Science.gov (United States)

    Cocco, Simona; Feinauer, Christoph; Figliuzzi, Matteo; Monasson, Rémi; Weigt, Martin

    2018-03-01

    In the course of evolution, proteins undergo important changes in their amino acid sequences, while their three-dimensional folded structure and their biological function remain remarkably conserved. Thanks to modern sequencing techniques, sequence data accumulate at unprecedented pace. This provides large sets of so-called homologous, i.e. evolutionarily related protein sequences, to which methods of inverse statistical physics can be applied. Using sequence data as the basis for the inference of Boltzmann distributions from samples of microscopic configurations or observables, it is possible to extract information about evolutionary constraints and thus protein function and structure. Here we give an overview over some biologically important questions, and how statistical-mechanics inspired modeling approaches can help to answer them. Finally, we discuss some open questions, which we expect to be addressed over the next years.

  2. Ultra-fast evaluation of protein energies directly from sequence.

    Directory of Open Access Journals (Sweden)

    Gevorg Grigoryan

    2006-06-01

    Full Text Available The structure, function, stability, and many other properties of a protein in a fixed environment are fully specified by its sequence, but in a manner that is difficult to discern. We present a general approach for rapidly mapping sequences directly to their energies on a pre-specified rigid backbone, an important sub-problem in computational protein design and in some methods for protein structure prediction. The cluster expansion (CE method that we employ can, in principle, be extended to model any computable or measurable protein property directly as a function of sequence. Here we show how CE can be applied to the problem of computational protein design, and use it to derive excellent approximations of physical potentials. The approach provides several attractive advantages. First, following a one-time derivation of a CE expansion, the amount of time necessary to evaluate the energy of a sequence adopting a specified backbone conformation is reduced by a factor of 10(7 compared to standard full-atom methods for the same task. Second, the agreement between two full-atom methods that we tested and their CE sequence-based expressions is very high (root mean square deviation 1.1-4.7 kcal/mol, R2 = 0.7-1.0. Third, the functional form of the CE energy expression is such that individual terms of the expansion have clear physical interpretations. We derived expressions for the energies of three classic protein design targets-a coiled coil, a zinc finger, and a WW domain-as functions of sequence, and examined the most significant terms. Single-residue and residue-pair interactions are sufficient to accurately capture the energetics of the dimeric coiled coil, whereas higher-order contributions are important for the two more globular folds. For the task of designing novel zinc-finger sequences, a CE-derived energy function provides significantly better solutions than a standard design protocol, in comparable computation time. Given these advantages

  3. Conformational effects of a common codon 751 polymorphism on the C-terminal domain of the xeroderma pigmentosum D protein

    Directory of Open Access Journals (Sweden)

    Monaco Regina

    2009-01-01

    Full Text Available Aim: The xeroderma pigmentosum D (XPD protein is a DNA helicase involved in the repair of DNA damage, including nucleotide excision repair (NER and transcription-coupled repair (TCR. The C-terminal domain of XPD has been implicated in interactions with other components of the TFIIH complex, and it is also the site of a common genetic polymorphism in XPD at amino acid residue 751 (Lys->Gln. Some evidence suggests that this polymorphism may alter DNA repair capacity and increase cancer risk. The aim of this study was to investigate whether these effects could be attributable to conformational changes in XPD induced by the polymorphism. Materials and Methods: Molecular dynamics techniques were used to predict the structure of the wild-type and polymorphic forms of the C-terminal domain of XPD and differences in structure produced by the polymorphic substitution were determined. Results: The results indicate that, although the general configuration of both proteins is similar, the substitution produces a significant conformational change immediately N-terminal to the site of the polymorphism. Conclusion: These results provide support for the hypothesis that this polymorphism in XPD could affect DNA repair capability, and hence cancer risk, by altering the structure of the C-terminal domain.

  4. ProteinSplit: splitting of multi-domain proteins using prediction of ordered and disordered regions in protein sequences for virtual structural genomics

    International Nuclear Information System (INIS)

    Wyrwicz, Lucjan S; Koczyk, Grzegorz; Rychlewski, Leszek; Plewczynski, Dariusz

    2007-01-01

    The annotation of protein folds within newly sequenced genomes is the main target for semi-automated protein structure prediction (virtual structural genomics). A large number of automated methods have been developed recently with very good results in the case of single-domain proteins. Unfortunately, most of these automated methods often fail to properly predict the distant homology between a given multi-domain protein query and structural templates. Therefore a multi-domain protein should be split into domains in order to overcome this limitation. ProteinSplit is designed to identify protein domain boundaries using a novel algorithm that predicts disordered regions in protein sequences. The software utilizes various sequence characteristics to assess the local propensity of a protein to be disordered or ordered in terms of local structure stability. These disordered parts of a protein are likely to create interdomain spacers. Because of its speed and portability, the method was successfully applied to several genome-wide fold annotation experiments. The user can run an automated analysis of sets of proteins or perform semi-automated multiple user projects (saving the results on the server). Additionally the sequences of predicted domains can be sent to the Bioinfo.PL Protein Structure Prediction Meta-Server for further protein three-dimensional structure and function prediction. The program is freely accessible as a web service at http://lucjan.bioinfo.pl/proteinsplit together with detailed benchmark results on the critical assessment of a fully automated structure prediction (CAFASP) set of sequences. The source code of the local version of protein domain boundary prediction is available upon request from the authors

  5. Rapid identification of sequences for orphan enzymes to power accurate protein annotation.

    Directory of Open Access Journals (Sweden)

    Kevin R Ramkissoon

    Full Text Available The power of genome sequencing depends on the ability to understand what those genes and their proteins products actually do. The automated methods used to assign functions to putative proteins in newly sequenced organisms are limited by the size of our library of proteins with both known function and sequence. Unfortunately this library grows slowly, lagging well behind the rapid increase in novel protein sequences produced by modern genome sequencing methods. One potential source for rapidly expanding this functional library is the "back catalog" of enzymology--"orphan enzymes," those enzymes that have been characterized and yet lack any associated sequence. There are hundreds of orphan enzymes in the Enzyme Commission (EC database alone. In this study, we demonstrate how this orphan enzyme "back catalog" is a fertile source for rapidly advancing the state of protein annotation. Starting from three orphan enzyme samples, we applied mass-spectrometry based analysis and computational methods (including sequence similarity networks, sequence and structural alignments, and operon context analysis to rapidly identify the specific sequence for each orphan while avoiding the most time- and labor-intensive aspects of typical sequence identifications. We then used these three new sequences to more accurately predict the catalytic function of 385 previously uncharacterized or misannotated proteins. We expect that this kind of rapid sequence identification could be efficiently applied on a larger scale to make enzymology's "back catalog" another powerful tool to drive accurate genome annotation.

  6. Rapid Identification of Sequences for Orphan Enzymes to Power Accurate Protein Annotation

    Science.gov (United States)

    Ojha, Sunil; Watson, Douglas S.; Bomar, Martha G.; Galande, Amit K.; Shearer, Alexander G.

    2013-01-01

    The power of genome sequencing depends on the ability to understand what those genes and their proteins products actually do. The automated methods used to assign functions to putative proteins in newly sequenced organisms are limited by the size of our library of proteins with both known function and sequence. Unfortunately this library grows slowly, lagging well behind the rapid increase in novel protein sequences produced by modern genome sequencing methods. One potential source for rapidly expanding this functional library is the “back catalog” of enzymology – “orphan enzymes,” those enzymes that have been characterized and yet lack any associated sequence. There are hundreds of orphan enzymes in the Enzyme Commission (EC) database alone. In this study, we demonstrate how this orphan enzyme “back catalog” is a fertile source for rapidly advancing the state of protein annotation. Starting from three orphan enzyme samples, we applied mass-spectrometry based analysis and computational methods (including sequence similarity networks, sequence and structural alignments, and operon context analysis) to rapidly identify the specific sequence for each orphan while avoiding the most time- and labor-intensive aspects of typical sequence identifications. We then used these three new sequences to more accurately predict the catalytic function of 385 previously uncharacterized or misannotated proteins. We expect that this kind of rapid sequence identification could be efficiently applied on a larger scale to make enzymology’s “back catalog” another powerful tool to drive accurate genome annotation. PMID:24386392

  7. Complete cDNA sequence coding for human docking protein

    Energy Technology Data Exchange (ETDEWEB)

    Hortsch, M; Labeit, S; Meyer, D I

    1988-01-11

    Docking protein (DP, or SRP receptor) is a rough endoplasmic reticulum (ER)-associated protein essential for the targeting and translocation of nascent polypeptides across this membrane. It specifically interacts with a cytoplasmic ribonucleoprotein complex, the signal recognition particle (SRP). The nucleotide sequence of cDNA encoding the entire human DP and its deduced amino acid sequence are given.

  8. Designing sequence to control protein function in an EF-hand protein.

    Science.gov (United States)

    Bunick, Christopher G; Nelson, Melanie R; Mangahas, Sheryll; Hunter, Michael J; Sheehan, Jonathan H; Mizoue, Laura S; Bunick, Gerard J; Chazin, Walter J

    2004-05-19

    The extent of conformational change that calcium binding induces in EF-hand proteins is a key biochemical property specifying Ca(2+) sensor versus signal modulator function. To understand how differences in amino acid sequence lead to differences in the response to Ca(2+) binding, comparative analyses of sequence and structures, combined with model building, were used to develop hypotheses about which amino acid residues control Ca(2+)-induced conformational changes. These results were used to generate a first design of calbindomodulin (CBM-1), a calbindin D(9k) re-engineered with 15 mutations to respond to Ca(2+) binding with a conformational change similar to that of calmodulin. The gene for CBM-1 was synthesized, and the protein was expressed and purified. Remarkably, this protein did not exhibit any non-native-like molten globule properties despite the large number of mutations and the nonconservative nature of some of them. Ca(2+)-induced changes in CD intensity and in the binding of the hydrophobic probe, ANS, implied that CBM-1 does undergo Ca(2+) sensorlike conformational changes. The X-ray crystal structure of Ca(2+)-CBM-1 determined at 1.44 A resolution reveals the anticipated increase in hydrophobic surface area relative to the wild-type protein. A nascent calmodulin-like hydrophobic docking surface was also found, though it is occluded by the inter-EF-hand loop. The results from this first calbindomodulin design are discussed in terms of progress toward understanding the relationships between amino acid sequence, protein structure, and protein function for EF-hand CaBPs, as well as the additional mutations for the next CBM design.

  9. Polymorphism of Paramecium pentaurelia (Ciliophora, Oligohymenophorea) strains revealed by rDNA and mtDNA sequences.

    Science.gov (United States)

    Przyboś, Ewa; Tarcz, Sebastian; Greczek-Stachura, Magdalena; Surmacz, Marta

    2011-05-01

    Paramecium pentaurelia is one of 15 known sibling species of the Paramecium aurelia complex. It is recognized as a species showing no intra-specific differentiation on the basis of molecular fingerprint analyses, whereas the majority of other species are polymorphic. This study aimed at assessing genetic polymorphism within P. pentaurelia including new strains recently found in Poland (originating from two water bodies, different years, seasons, and clones of one strain) as well as strains collected from distant habitats (USA, Europe, Asia), and strains representing other species of the complex. We compared two DNA fragments: partial sequences (349 bp) of the LSU rDNA and partial sequences (618 bp) of cytochrome B gene. A correlation between the geographical origin of the strains and the genetic characteristics of their genotypes was not observed. Different genotypes were found in Kraków in two types of water bodies (Opatkowice-natural pond; Jordan's Park-artificial pond). Haplotype diversity within a single water body was not recorded. Likewise, seasonal haplotype differences between the strains within the artificial water body, as well as differences between clones originating from one strain, were not detected. The clustering of some strains belonging to different species was observed in the phylogenies. Copyright © 2010 Elsevier GmbH. All rights reserved.

  10. Polymorphism at codon 36 of the p53 gene.

    Science.gov (United States)

    Felix, C A; Brown, D L; Mitsudomi, T; Ikagaki, N; Wong, A; Wasserman, R; Womer, R B; Biegel, J A

    1994-01-01

    A polymorphism at codon 36 in exon 4 of the p53 gene was identified by single strand conformation polymorphism (SSCP) analysis and direct sequencing of genomic DNA PCR products. The polymorphic allele, present in the heterozygous state in genomic DNAs of four of 100 individuals (4%), changes the codon 36 CCG to CCA, eliminates a FinI restriction site and creates a BccI site. Including this polymorphism there are four known polymorphisms in the p53 coding sequence.

  11. Prediction of protein-protein interaction sites in sequences and 3D structures by random forests.

    Directory of Open Access Journals (Sweden)

    Mile Sikić

    2009-01-01

    Full Text Available Identifying interaction sites in proteins provides important clues to the function of a protein and is becoming increasingly relevant in topics such as systems biology and drug discovery. Although there are numerous papers on the prediction of interaction sites using information derived from structure, there are only a few case reports on the prediction of interaction residues based solely on protein sequence. Here, a sliding window approach is combined with the Random Forests method to predict protein interaction sites using (i a combination of sequence- and structure-derived parameters and (ii sequence information alone. For sequence-based prediction we achieved a precision of 84% with a 26% recall and an F-measure of 40%. When combined with structural information, the prediction performance increases to a precision of 76% and a recall of 38% with an F-measure of 51%. We also present an attempt to rationalize the sliding window size and demonstrate that a nine-residue window is the most suitable for predictor construction. Finally, we demonstrate the applicability of our prediction methods by modeling the Ras-Raf complex using predicted interaction sites as target binding interfaces. Our results suggest that it is possible to predict protein interaction sites with quite a high accuracy using only sequence information.

  12. Complete cDNA sequence and amino acid analysis of a bovine ribonuclease K6 gene.

    Science.gov (United States)

    Pietrowski, D; Förster, M

    2000-01-01

    The complete cDNA sequence of a ribonuclease k6 gene of Bos Taurus has been determined. It codes for a protein with 154 amino acids and contains the invariant cysteine, histidine and lysine residues as well as the characteristic motifs specific to ribonuclease active sites. The deduced protein sequence is 27 residues longer than other known ribonucleases k6 and shows amino acids exchanges which could reflect a strain specificity or polymorphism within the bovine genome. Based on sequence similarity we have termed the identified gene bovine ribonuclease k6 b (brk6b).

  13. Repeat Sequence Proteins as Matrices for Nanocomposites

    Energy Technology Data Exchange (ETDEWEB)

    Drummy, L.; Koerner, H; Phillips, D; McAuliffe, J; Kumar, M; Farmer, B; Vaia, R; Naik, R

    2009-01-01

    Recombinant protein-inorganic nanocomposites comprised of exfoliated Na+ montmorillonite (MMT) in a recombinant protein matrix based on silk-like and elastin-like amino acid motifs (silk elastin-like protein (SELP)) were formed via a solution blending process. Charged residues along the protein backbone are shown to dominate long-range interactions, whereas the SELP repeat sequence leads to local protein/MMT compatibility. Up to a 50% increase in room temperature modulus and a comparable decrease in high temperature coefficient of thermal expansion occur for cast films containing 2-10 wt.% MMT.

  14. Analysis of long-range correlation in sequences data of proteins

    OpenAIRE

    ADRIANA ISVORAN; LAURA UNIPAN; DANA CRACIUN; VASILE MORARIU

    2007-01-01

    The results presented here suggest the existence of correlations in the sequence data of proteins. 32 proteins, both globular and fibrous, both monomeric and polymeric, were analyzed. The primary structures of these proteins were treated as time series. Three spatial series of data for each sequence of a protein were generated from numerical correspondences between each amino acid and a physical property associated with it, i.e., its electric charge, its polar character and its dipole moment....

  15. Species-specific markers for the differential diagnosis of Trypanosoma cruzi and Trypanosoma rangeli and polymorphisms detection in Trypanosoma rangeli.

    Science.gov (United States)

    Ferreira, Keila Adriana Magalhães; Fajardo, Emanuella Francisco; Baptista, Rodrigo P; Macedo, Andrea Mara; Lages-Silva, Eliane; Ramírez, Luis Eduardo; Pedrosa, André Luiz

    2014-06-01

    Trypanosoma cruzi and Trypanosoma rangeli are kinetoplastid parasites which are able to infect humans in Central and South America. Misdiagnosis between these trypanosomes can be avoided by targeting barcoding sequences or genes of each organism. This work aims to analyze the feasibility of using species-specific markers for identification of intraspecific polymorphisms and as target for diagnostic methods by PCR. Accordingly, primers which are able to specifically detect T. cruzi or T. rangeli genomic DNA were characterized. The use of intergenic regions, generally divergent in the trypanosomatids, and the serine carboxypeptidase gene were successful. Using T. rangeli genomic sequences for the identification of group-specific polymorphisms and a polymorphic AT(n) dinucleotide repeat permitted the classification of the strains into two groups, which are entirely coincident with T. rangeli main lineages, KP1 (+) and KP1 (-), previously determined by kinetoplast DNA (kDNA) characterization. The sequences analyzed totalize 622 bp (382 bp represent a hypothetical protein sequence, and 240 bp represent an anonymous sequence), and of these, 581 (93.3%) are conserved sites and 41 bp (6.7%) are polymorphic, with 9 transitions (21.9%), 2 transversions (4.9%), and 30 (73.2%) insertion/deletion events. Taken together, the species-specific markers analyzed may be useful for the development of new strategies for the accurate diagnosis of infections. Furthermore, the identification of T. rangeli polymorphisms has a direct impact in the understanding of the population structure of this parasite.

  16. Genome-wide data-mining of candidate human splice translational efficiency polymorphisms (STEPs and an online database.

    Directory of Open Access Journals (Sweden)

    Christopher A Raistrick

    2010-10-01

    Full Text Available Variation in pre-mRNA splicing is common and in some cases caused by genetic variants in intronic splicing motifs. Recent studies into the insulin gene (INS discovered a polymorphism in a 5' non-coding intron that influences the likelihood of intron retention in the final mRNA, extending the 5' untranslated region and maintaining protein quality. Retention was also associated with increased insulin levels, suggesting that such variants--splice translational efficiency polymorphisms (STEPs--may relate to disease phenotypes through differential protein expression. We set out to explore the prevalence of STEPs in the human genome and validate this new category of protein quantitative trait loci (pQTL using publicly available data.Gene transcript and variant data were collected and mined for candidate STEPs in motif regions. Sequences from transcripts containing potential STEPs were analysed for evidence of splice site recognition and an effect in expressed sequence tags (ESTs. 16 publicly released genome-wide association data sets of common diseases were searched for association to candidate polymorphisms with HapMap frequency data. Our study found 3324 candidate STEPs lying in motif sequences of 5' non-coding introns and further mining revealed 170 with transcript evidence of intron retention. 21 potential STEPs had EST evidence of intron retention or exon extension, as well as population frequency data for comparison.Results suggest that the insulin STEP was not a unique example and that many STEPs may occur genome-wide with potentially causal effects in complex disease. An online database of STEPs is freely accessible at http://dbstep.genes.org.uk/.

  17. Sequence analysis reveals how G protein-coupled receptors transduce the signal to the G protein.

    NARCIS (Netherlands)

    Oliveira, L.; Paiva, P.B.; Paiva, A.C.; Vriend, G.

    2003-01-01

    Sequence entropy-variability plots based on alignments of very large numbers of sequences-can indicate the location in proteins of the main active site and modulator sites. In the previous article in this issue, we applied this observation to a series of well-studied proteins and concluded that it

  18. Association of impulsivity and polymorphic microRNA-641 target sites in the SNAP-25 gene.

    Directory of Open Access Journals (Sweden)

    Nóra Németh

    Full Text Available Impulsivity is a personality trait of high impact and is connected with several types of maladaptive behavior and psychiatric diseases, such as attention deficit hyperactivity disorder, alcohol and drug abuse, as well as pathological gambling and mood disorders. Polymorphic variants of the SNAP-25 gene emerged as putative genetic components of impulsivity, as SNAP-25 protein plays an important role in the central nervous system, and its SNPs are associated with several psychiatric disorders. In this study we aimed to investigate if polymorphisms in the regulatory regions of the SNAP-25 gene are in association with normal variability of impulsivity. Genotypes and haplotypes of two polymorphisms in the promoter (rs6077690 and rs6039769 and two SNPs in the 3' UTR (rs3746544 and rs1051312 of the SNAP-25 gene were determined in a healthy Hungarian population (N = 901 using PCR-RFLP or real-time PCR in combination with sequence specific probes. Significant association was found between the T-T 3' UTR haplotype and impulsivity, whereas no association could be detected with genotypes or haplotypes of the promoter loci. According to sequence alignment, the polymorphisms in the 3' UTR of the gene alter the binding site of microRNA-641, which was analyzed by luciferase reporter system. It was observed that haplotypes altering one or two nucleotides in the binding site of the seed region of microRNA-641 significantly increased the amount of generated protein in vitro. These findings support the role of polymorphic SNAP-25 variants both at psychogenetic and molecular biological levels.

  19. PTPN22 -1123G>C polymorphism and anti-cyclic citrullinated protein antibodies in rheumatoid arthritis.

    Science.gov (United States)

    Muñoz-Valle, José Francisco; Padilla-Gutiérrez, Jorge Ramón; Hernández-Bello, Jorge; Ruiz-Noa, Yeniley; Valle, Yeminia; Palafox-Sánchez, Claudia Azucena; Parra-Rojas, Isela; Gutiérrez-Ureña, Sergio Ramón; Rangel-Villalobos, Hector

    2017-08-10

    The protein tyrosine phosphatase non-receptor type 22 (PTPN22) gene encodes an important negative regulator of T-cell activation, lymphoid-specific phosphatase -Lyp- and has been associated with different autoimmune disorders. The PTPN22 -1123G>C polymorphism appears to affect the transcriptional control of this gene, but to date, the biological significance of this polymorphisms on rheumatoid arthritis (RA) risk remains unknown. We evaluate the association of PTPN22 -1123G>C polymorphism with anti-cyclic citrullinated protein antibodies (anti-CCP) and risk for RA in population from Western Mexico. A transversal analytic study, which enrolled 300 RA patients classified according to ACR-EULAR criteria and 300 control subjects (CS) was conducted. The -1123 G>C polymorphism was genotyped by PCR-RFLP. The anti-CCP antibodies levels were quantified by ELISA kit. We found a higher prevalence of homozygous PTPN22 -1123CC genotype in CS than in RA patients (OR 0.41; 95% confidence interval 0.24-0.71; P=.001), suggesting a potential protective effect against RA. Concerning anti-CCP levels, the CC genotype carriers showed the lowest median levels in RA (P<.05). The PTPN22 -1123CC genotype is a protector factor to RA in a Mexican-mestizo population and is associated with low anti-CCP antibodies. Copyright © 2017 Elsevier España, S.L.U. All rights reserved.

  20. Peptide Pattern Recognition for high-throughput protein sequence analysis and clustering

    DEFF Research Database (Denmark)

    Busk, Peter Kamp

    2017-01-01

    Large collections of protein sequences with divergent sequences are tedious to analyze for understanding their phylogenetic or structure-function relation. Peptide Pattern Recognition is an algorithm that was developed to facilitate this task but the previous version does only allow a limited...... number of sequences as input. I implemented Peptide Pattern Recognition as a multithread software designed to handle large numbers of sequences and perform analysis in a reasonable time frame. Benchmarking showed that the new implementation of Peptide Pattern Recognition is twenty times faster than...... the previous implementation on a small protein collection with 673 MAP kinase sequences. In addition, the new implementation could analyze a large protein collection with 48,570 Glycosyl Transferase family 20 sequences without reaching its upper limit on a desktop computer. Peptide Pattern Recognition...

  1. M protein typing of Thai group A streptococcal isolates by PCR-Restriction fragment length polymorphism analysis

    Directory of Open Access Journals (Sweden)

    Good Michael F

    2005-10-01

    Full Text Available Abstract Background Group A streptococcal (GAS infections can lead to the development of severe post-infectious sequelae, such as rheumatic fever (RF and rheumatic heart disease (RHD. RF and RHD are a major health concern in developing countries, and in indigenous populations of developed nations. The majority of GAS isolates are M protein-nontypeable (MNT by standard serotyping. However, GAS typing is a necessary tool in the epidemiologically analysis of GAS and provides useful information for vaccine development. Although DNA sequencing is the most conclusive method for M protein typing, this is not a feasible approach especially in developing countries. To overcome this problem, we have developed a polymerase chain reaction-restriction fragment length polymorphism (PCR-RFLP-based assay for molecular typing the M protein gene (emm of GAS. Results Using one pair of primers, 13 known GAS M types showed one to four bands of PCR products and after digestion with Alu I, they gave different RFLP patterns. Of 106 GAS isolates examined from the normal Thai population and from patients with GAS-associated complications including RHD, 95 isolates gave RFLP patterns that corresponded to the 13 known M types. Only 11 isolates gave RFLP patterns that differed from the 13 known M types. These were then analyzed by DNA sequencing and six additional M types were identified. In addition, we found that M93 GAS was the most common M type in the population studied, and is consistent with a previous study of Thai GAS isolates. Conclusion PCR-RFLP analysis has the potential for the rapid screening of different GAS M types and is therefore considerably advantageous as an alternative M typing approach in developing countries in which GAS is endemic.

  2. [Interspecific polymorphism of the glucosyltransferase domain of the sucrose synthase gene in the genus Malus and related species of Rosaceae].

    Science.gov (United States)

    Boris, K V; Kochieva, E Z; Kudryavtsev, A M

    2014-12-01

    The sequences that encode the main functional glucosyltransferase domain of sucrose synthase genes have been identified for the first time in 14 species of the genus Malus and related species of the family Rosaceae, and their polymorphism was investigated. Single nucleotide substitutions leading to amino acid substitutions in the protein sequence, including the conservative transmembrane motif sequence common to all sucrose synthase genes of higher plants, were detected in the studied sequences.

  3. Discovery and mapping of a new expressed sequence tag-single nucleotide polymorphism and simple sequence repeat panel for large-scale genetic studies and breeding of Theobroma cacao L.

    Science.gov (United States)

    Allegre, Mathilde; Argout, Xavier; Boccara, Michel; Fouet, Olivier; Roguet, Yolande; Bérard, Aurélie; Thévenin, Jean Marc; Chauveau, Aurélie; Rivallan, Ronan; Clement, Didier; Courtois, Brigitte; Gramacho, Karina; Boland-Augé, Anne; Tahi, Mathias; Umaharan, Pathmanathan; Brunel, Dominique; Lanaud, Claire

    2012-01-01

    Theobroma cacao is an economically important tree of several tropical countries. Its genetic improvement is essential to provide protection against major diseases and improve chocolate quality. We discovered and mapped new expressed sequence tag-single nucleotide polymorphism (EST-SNP) and simple sequence repeat (SSR) markers and constructed a high-density genetic map. By screening 149 650 ESTs, 5246 SNPs were detected in silico, of which 1536 corresponded to genes with a putative function, while 851 had a clear polymorphic pattern across a collection of genetic resources. In addition, 409 new SSR markers were detected on the Criollo genome. Lastly, 681 new EST-SNPs and 163 new SSRs were added to the pre-existing 418 co-dominant markers to construct a large consensus genetic map. This high-density map and the set of new genetic markers identified in this study are a milestone in cocoa genomics and for marker-assisted breeding. The data are available at http://tropgenedb.cirad.fr. PMID:22210604

  4. Genomic DNA Enrichment Using Sequence Capture Microarrays: a Novel Approach to Discover Sequence Nucleotide Polymorphisms (SNP) in Brassica napus L

    Science.gov (United States)

    Clarke, Wayne E.; Parkin, Isobel A.; Gajardo, Humberto A.; Gerhardt, Daniel J.; Higgins, Erin; Sidebottom, Christine; Sharpe, Andrew G.; Snowdon, Rod J.; Federico, Maria L.; Iniguez-Luy, Federico L.

    2013-01-01

    Targeted genomic selection methodologies, or sequence capture, allow for DNA enrichment and large-scale resequencing and characterization of natural genetic variation in species with complex genomes, such as rapeseed canola (Brassica napus L., AACC, 2n=38). The main goal of this project was to combine sequence capture with next generation sequencing (NGS) to discover single nucleotide polymorphisms (SNPs) in specific areas of the B. napus genome historically associated (via quantitative trait loci –QTL– analysis) to traits of agronomical and nutritional importance. A 2.1 million feature sequence capture platform was designed to interrogate DNA sequence variation across 47 specific genomic regions, representing 51.2 Mb of the Brassica A and C genomes, in ten diverse rapeseed genotypes. All ten genotypes were sequenced using the 454 Life Sciences chemistry and to assess the effect of increased sequence depth, two genotypes were also sequenced using Illumina HiSeq chemistry. As a result, 589,367 potentially useful SNPs were identified. Analysis of sequence coverage indicated a four-fold increased representation of target regions, with 57% of the filtered SNPs falling within these regions. Sixty percent of discovered SNPs corresponded to transitions while 40% were transversions. Interestingly, fifty eight percent of the SNPs were found in genic regions while 42% were found in intergenic regions. Further, a high percentage of genic SNPs was found in exons (65% and 64% for the A and C genomes, respectively). Two different genotyping assays were used to validate the discovered SNPs. Validation rates ranged from 61.5% to 84% of tested SNPs, underpinning the effectiveness of this SNP discovery approach. Most importantly, the discovered SNPs were associated with agronomically important regions of the B. napus genome generating a novel data resource for research and breeding this crop species. PMID:24312619

  5. Elman RNN based classification of proteins sequences on account of their mutual information.

    Science.gov (United States)

    Mishra, Pooja; Nath Pandey, Paras

    2012-10-21

    In the present work we have employed the method of estimating residue correlation within the protein sequences, by using the mutual information (MI) of adjacent residues, based on structural and solvent accessibility properties of amino acids. The long range correlation between nonadjacent residues is improved by constructing a mutual information vector (MIV) for a single protein sequence, like this each protein sequence is associated with its corresponding MIVs. These MIVs are given to Elman RNN to obtain the classification of protein sequences. The modeling power of MIV was shown to be significantly better, giving a new approach towards alignment free classification of protein sequences. We also conclude that sequence structural and solvent accessible property based MIVs are better predictor. Copyright © 2012 Elsevier Ltd. All rights reserved.

  6. A resource of genome-wide single-nucleotide polymorphisms generated by RAD tag sequencing in the critically endangered European eel

    DEFF Research Database (Denmark)

    Pujolar, J.M.; Jacobsen, M.W.; Frydenberg, J.

    2013-01-01

    Reduced representation genome sequencing such as restriction-site-associated DNA (RAD) sequencing is finding increased use to identify and genotype large numbers of single-nucleotide polymorphisms (SNPs) in model and nonmodel species. We generated a unique resource of novel SNP markers for the Eu...... 425 loci and 376 918 associated SNPs provides a valuable tool for future population genetics and genomics studies and allows for targeting specific genes and particularly interesting regions of the eel genome...

  7. Predicting membrane protein types by fusing composite protein sequence features into pseudo amino acid composition.

    Science.gov (United States)

    Hayat, Maqsood; Khan, Asifullah

    2011-02-21

    Membrane proteins are vital type of proteins that serve as channels, receptors, and energy transducers in a cell. Prediction of membrane protein types is an important research area in bioinformatics. Knowledge of membrane protein types provides some valuable information for predicting novel example of the membrane protein types. However, classification of membrane protein types can be both time consuming and susceptible to errors due to the inherent similarity of membrane protein types. In this paper, neural networks based membrane protein type prediction system is proposed. Composite protein sequence representation (CPSR) is used to extract the features of a protein sequence, which includes seven feature sets; amino acid composition, sequence length, 2 gram exchange group frequency, hydrophobic group, electronic group, sum of hydrophobicity, and R-group. Principal component analysis is then employed to reduce the dimensionality of the feature vector. The probabilistic neural network (PNN), generalized regression neural network, and support vector machine (SVM) are used as classifiers. A high success rate of 86.01% is obtained using SVM for the jackknife test. In case of independent dataset test, PNN yields the highest accuracy of 95.73%. These classifiers exhibit improved performance using other performance measures such as sensitivity, specificity, Mathew's correlation coefficient, and F-measure. The experimental results show that the prediction performance of the proposed scheme for classifying membrane protein types is the best reported, so far. This performance improvement may largely be credited to the learning capabilities of neural networks and the composite feature extraction strategy, which exploits seven different properties of protein sequences. The proposed Mem-Predictor can be accessed at http://111.68.99.218/Mem-Predictor. Copyright © 2010 Elsevier Ltd. All rights reserved.

  8. Surfactant protein B polymorphisms are associated with severe respiratory syncytial virus infection, but not with asthma

    Directory of Open Access Journals (Sweden)

    Heinzmann Andrea

    2007-05-01

    Full Text Available Abstract Background Surfactant proteins (SP are important for the innate host defence and essential for a physiological lung function. Several linkage and association studies have investigated the genes coding for different surfactant proteins in the context of pulmonary diseases such as chronic obstructive pulmonary disease or respiratory distress syndrome of preterm infants. In this study we tested whether SP-B was in association with two further pulmonary diseases in children, i. e. severe infections caused by respiratory syncytial virus and bronchial asthma. Methods We chose to study five polymorphisms in SP-B: rs2077079 in the promoter region; rs1130866 leading to the amino acid exchange T131I; rs2040349 in intron 8; rs3024801 leading to L176F and rs3024809 resulting in R272H. Statistical analyses made use of the Armitage's trend test for single polymorphisms and FAMHAP and FASTEHPLUS for haplotype analyses. Results The polymorphisms rs3024801 and rs3024809 were not present in our study populations. The three other polymorphisms were common and in tight linkage disequilibrium with each other. They did not show association with bronchial asthma or severe RSV infection in the analyses of single polymorphisms. However, haplotypes analyses revealed association of SP-B with severe RSV infection (p = 0.034. Conclusion Thus our results indicate a possible involvement of SP-B in the genetic predisposition to severe RSV infections in the German population. In order to determine which of the three polymorphisms constituting the haplotypes is responsible for the association, further case control studies on large populations are necessary. Furthermore, functional analysis need to be conducted.

  9. Sequence-Based Introgression Mapping Identifies Candidate White Mold Tolerance Genes in Common Bean

    Directory of Open Access Journals (Sweden)

    Sujan Mamidi

    2016-07-01

    Full Text Available White mold, caused by the necrotrophic fungus (Lib. de Bary, is a major disease of common bean ( L.. WM7.1 and WM8.3 are two quantitative trait loci (QTL with major effects on tolerance to the pathogen. Advanced backcross populations segregating individually for either of the two QTL, and a recombinant inbred (RI population segregating for both QTL were used to fine map and confirm the genetic location of the QTL. The QTL intervals were physically mapped using the reference common bean genome sequence, and the physical intervals for each QTL were further confirmed by sequence-based introgression mapping. Using whole-genome sequence data from susceptible and tolerant DNA pools, introgressed regions were identified as those with significantly higher numbers of single-nucleotide polymorphisms (SNPs relative to the whole genome. By combining the QTL and SNP data, WM7.1 was located to a 660-kb region that contained 41 gene models on the proximal end of chromosome Pv07, while the WM8.3 introgression was narrowed to a 1.36-Mb region containing 70 gene models. The most polymorphic candidate gene in the WM7.1 region encodes a BEACH-domain protein associated with apoptosis. Within the WM8.3 interval, a receptor-like protein with the potential to recognize pathogen effectors was the most polymorphic gene. The use of gene and sequence-based mapping identified two candidate genes whose putative functions are consistent with the current model of pathogenicity.

  10. Genome-Wide Analysis of Simple Sequence Repeats and Efficient Development of Polymorphic SSR Markers Based on Whole Genome Re-Sequencing of Multiple Isolates of the Wheat Stripe Rust Fungus.

    Directory of Open Access Journals (Sweden)

    Huaiyong Luo

    Full Text Available The biotrophic parasitic fungus Puccinia striiformis f. sp. tritici (Pst causes stripe rust, a devastating disease of wheat, endangering global food security. Because the Pst population is highly dynamic, it is difficult to develop wheat cultivars with durable and highly effective resistance. Simple sequence repeats (SSRs are widely used as molecular markers in genetic studies to determine population structure in many organisms. However, only a small number of SSR markers have been developed for Pst. In this study, a total of 4,792 SSR loci were identified using the whole genome sequences of six isolates from different regions of the world, with a marker density of one SSR per 22.95 kb. The majority of the SSRs were di- and tri-nucleotide repeats. A database containing 1,113 SSR markers were established. Through in silico comparison, the previously reported SSR markers were found mainly in exons, whereas the SSR markers in the database were mostly in intergenic regions. Furthermore, 105 polymorphic SSR markers were confirmed in silico by their identical positions and nucleotide variations with INDELs identified among the six isolates. When 104 in silico polymorphic SSR markers were used to genotype 21 Pst isolates, 84 produced the target bands, and 82 of them were polymorphic and revealed the genetic relationships among the isolates. The results show that whole genome re-sequencing of multiple isolates provides an ideal resource for developing SSR markers, and the newly developed SSR markers are useful for genetic and population studies of the wheat stripe rust fungus.

  11. Genome-Wide Analysis of Simple Sequence Repeats and Efficient Development of Polymorphic SSR Markers Based on Whole Genome Re-Sequencing of Multiple Isolates of the Wheat Stripe Rust Fungus.

    Science.gov (United States)

    Luo, Huaiyong; Wang, Xiaojie; Zhan, Gangming; Wei, Guorong; Zhou, Xinli; Zhao, Jing; Huang, Lili; Kang, Zhensheng

    2015-01-01

    The biotrophic parasitic fungus Puccinia striiformis f. sp. tritici (Pst) causes stripe rust, a devastating disease of wheat, endangering global food security. Because the Pst population is highly dynamic, it is difficult to develop wheat cultivars with durable and highly effective resistance. Simple sequence repeats (SSRs) are widely used as molecular markers in genetic studies to determine population structure in many organisms. However, only a small number of SSR markers have been developed for Pst. In this study, a total of 4,792 SSR loci were identified using the whole genome sequences of six isolates from different regions of the world, with a marker density of one SSR per 22.95 kb. The majority of the SSRs were di- and tri-nucleotide repeats. A database containing 1,113 SSR markers were established. Through in silico comparison, the previously reported SSR markers were found mainly in exons, whereas the SSR markers in the database were mostly in intergenic regions. Furthermore, 105 polymorphic SSR markers were confirmed in silico by their identical positions and nucleotide variations with INDELs identified among the six isolates. When 104 in silico polymorphic SSR markers were used to genotype 21 Pst isolates, 84 produced the target bands, and 82 of them were polymorphic and revealed the genetic relationships among the isolates. The results show that whole genome re-sequencing of multiple isolates provides an ideal resource for developing SSR markers, and the newly developed SSR markers are useful for genetic and population studies of the wheat stripe rust fungus.

  12. Single nucleotide polymorphism analysis of Korean native chickens using next generation sequencing data.

    Science.gov (United States)

    Seo, Dong-Won; Oh, Jae-Don; Jin, Shil; Song, Ki-Duk; Park, Hee-Bok; Heo, Kang-Nyeong; Shin, Younhee; Jung, Myunghee; Park, Junhyung; Jo, Cheorun; Lee, Hak-Kyo; Lee, Jun-Heon

    2015-02-01

    There are five native chicken lines in Korea, which are mainly classified by plumage colors (black, white, red, yellow, gray). These five lines are very important genetic resources in the Korean poultry industry. Based on a next generation sequencing technology, whole genome sequence and reference assemblies were performed using Gallus_gallus_4.0 (NCBI) with whole genome sequences from these lines to identify common and novel single nucleotide polymorphisms (SNPs). We obtained 36,660,731,136 ± 1,257,159,120 bp of raw sequence and average 26.6-fold of 25-29 billion reference assembly sequences representing 97.288 % coverage. Also, 4,006,068 ± 97,534 SNPs were observed from 29 autosomes and the Z chromosome and, of these, 752,309 SNPs are the common SNPs across lines. Among the identified SNPs, the number of novel- and known-location assigned SNPs was 1,047,951 ± 14,956 and 2,948,648 ± 81,414, respectively. The number of unassigned known SNPs was 1,181 ± 150 and unassigned novel SNPs was 8,238 ± 1,019. Synonymous SNPs, non-synonymous SNPs, and SNPs having character changes were 26,266 ± 1,456, 11,467 ± 604, 8,180 ± 458, respectively. Overall, 443,048 ± 26,389 SNPs in each bird were identified by comparing with dbSNP in NCBI. The presently obtained genome sequence and SNP information in Korean native chickens have wide applications for further genome studies such as genetic diversity studies to detect causative mutations for economic and disease related traits.

  13. The HMMER Web Server for Protein Sequence Similarity Search.

    Science.gov (United States)

    Prakash, Ananth; Jeffryes, Matt; Bateman, Alex; Finn, Robert D

    2017-12-08

    Protein sequence similarity search is one of the most commonly used bioinformatics methods for identifying evolutionarily related proteins. In general, sequences that are evolutionarily related share some degree of similarity, and sequence-search algorithms use this principle to identify homologs. The requirement for a fast and sensitive sequence search method led to the development of the HMMER software, which in the latest version (v3.1) uses a combination of sophisticated acceleration heuristics and mathematical and computational optimizations to enable the use of profile hidden Markov models (HMMs) for sequence analysis. The HMMER Web server provides a common platform by linking the HMMER algorithms to databases, thereby enabling the search for homologs, as well as providing sequence and functional annotation by linking external databases. This unit describes three basic protocols and two alternate protocols that explain how to use the HMMER Web server using various input formats and user defined parameters. © 2017 by John Wiley & Sons, Inc. Copyright © 2017 John Wiley & Sons, Inc.

  14. Quantiprot - a Python package for quantitative analysis of protein sequences.

    Science.gov (United States)

    Konopka, Bogumił M; Marciniak, Marta; Dyrka, Witold

    2017-07-17

    The field of protein sequence analysis is dominated by tools rooted in substitution matrices and alignments. A complementary approach is provided by methods of quantitative characterization. A major advantage of the approach is that quantitative properties defines a multidimensional solution space, where sequences can be related to each other and differences can be meaningfully interpreted. Quantiprot is a software package in Python, which provides a simple and consistent interface to multiple methods for quantitative characterization of protein sequences. The package can be used to calculate dozens of characteristics directly from sequences or using physico-chemical properties of amino acids. Besides basic measures, Quantiprot performs quantitative analysis of recurrence and determinism in the sequence, calculates distribution of n-grams and computes the Zipf's law coefficient. We propose three main fields of application of the Quantiprot package. First, quantitative characteristics can be used in alignment-free similarity searches, and in clustering of large and/or divergent sequence sets. Second, a feature space defined by quantitative properties can be used in comparative studies of protein families and organisms. Third, the feature space can be used for evaluating generative models, where large number of sequences generated by the model can be compared to actually observed sequences.

  15. Analysis of the intronic single nucleotide polymorphism rs#466452 of the nephrin gene in patients with diabetic nephropathy

    Directory of Open Access Journals (Sweden)

    RODRIGO GONZÁLEZ

    2009-01-01

    Full Text Available We present the analysis of an intronic polymorphism of the nephrin gene and its relationship to the development of diabetic nephropathy in a study of diabetes type 1 and type 2 patients. The frequency of the single nucleotide polymorphism rs#466452 in the nephrin gene was determined in 231 patients and control subjects. The C/T status of the polymorphism was assessed using restriction enzyme digestions and the nephrin transcript from a kidney biopsy was examined. Association between the polymorphism and clinical parameters was evaluated using multivaríate correspondence analysis. A bioinformatics analysis of the single nucleotide polymorphism rs#466452 suggested the appearance of a splicing enhancer sequence in intron 24 of the nephrin gene and a modification of proteins that bind to this sequence. However, no change in the splicing of a nephrin transcript from a renal biopsy was found. No association was found between the polymorphism and diabetes or degree of renal damage in diabetes type 1 or 2 patients. The single nucleotide polymorphism rs#466452 of the nephrin gene seems to be neutral in relation to diabetes and the development of diabetic nephropathy, and does not affect the splicing of a nephrin transcript, in spite of a splicing enhancer site.

  16. Single nucleotide polymorphism in Egyptian cattle insulin-like growth factor binding protein-3 gene

    Directory of Open Access Journals (Sweden)

    Othman E. Othman

    2014-12-01

    It is concluded that the IGFBP-3/HaeIII polymorphism may be utilized as a good marker for genetic differentiation between cattle animals for different body functions such as growth, metabolism, reproduction, immunity and energy balance. The nucleotide sequences of Egyptian cattle IGFBP-3 A and C alleles were submitted to GenBank with the accession numbers KF899893 and KF899894, respectively.

  17. New in protein structure and function annotation: hotspots, single nucleotide polymorphisms and the 'Deep Web'.

    Science.gov (United States)

    Bromberg, Yana; Yachdav, Guy; Ofran, Yanay; Schneider, Reinhard; Rost, Burkhard

    2009-05-01

    The rapidly increasing quantity of protein sequence data continues to widen the gap between available sequences and annotations. Comparative modeling suggests some aspects of the 3D structures of approximately half of all known proteins; homology- and network-based inferences annotate some aspect of function for a similar fraction of the proteome. For most known protein sequences, however, there is detailed knowledge about neither their function nor their structure. Comprehensive efforts towards the expert curation of sequence annotations have failed to meet the demand of the rapidly increasing number of available sequences. Only the automated prediction of protein function in the absence of homology can close the gap between available sequences and annotations in the foreseeable future. This review focuses on two novel methods for automated annotation, and briefly presents an outlook on how modern web software may revolutionize the field of protein sequence annotation. First, predictions of protein binding sites and functional hotspots, and the evolution of these into the most successful type of prediction of protein function from sequence will be discussed. Second, a new tool, comprehensive in silico mutagenesis, which contributes important novel predictions of function and at the same time prepares for the onset of the next sequencing revolution, will be described. While these two new sub-fields of protein prediction represent the breakthroughs that have been achieved methodologically, it will then be argued that a different development might further change the way biomedical researchers benefit from annotations: modern web software can connect the worldwide web in any browser with the 'Deep Web' (ie, proprietary data resources). The availability of this direct connection, and the resulting access to a wealth of data, may impact drug discovery and development more than any existing method that contributes to protein annotation.

  18. Porcine lung surfactant protein B gene (SFTPB)

    DEFF Research Database (Denmark)

    Cirera Salicio, Susanna; Fredholm, Merete

    2008-01-01

    The porcine surfactant protein B (SFTPB) is a single copy gene on chromosome 3. Three different cDNAs for the SFTPB have been isolated and sequenced. Nucleotide sequence comparison revealed six nonsynonymous single nucleotide polymorphisms (SNPs), four synonymous SNPs and an in-frame deletion of 69...... bp in the region coding for the active protein. Northern analysis showed lung-specific expression of three different isoforms of the SFTPB transcript. The expression level for the SFTPB gene is low in 50 days-old fetus and it increases during lung development. Quantitative real-time polymerase chain...

  19. Osteocalcin protein sequences of Neanderthals and modern primates.

    Science.gov (United States)

    Nielsen-Marsh, Christina M; Richards, Michael P; Hauschka, Peter V; Thomas-Oates, Jane E; Trinkaus, Erik; Pettitt, Paul B; Karavanic, Ivor; Poinar, Hendrik; Collins, Matthew J

    2005-03-22

    We report here protein sequences of fossil hominids, from two Neanderthals dating to approximately 75,000 years old from Shanidar Cave in Iraq. These sequences, the oldest reported fossil primate protein sequences, are of bone osteocalcin, which was extracted and sequenced by using MALDI-TOF/TOF mass spectrometry. Through a combination of direct sequencing and peptide mass mapping, we determined that Neanderthals have an osteocalcin amino acid sequence that is identical to that of modern humans. We also report complete osteocalcin sequences for chimpanzee (Pan troglodytes) and gorilla (Gorilla gorilla gorilla) and a partial sequence for orangutan (Pongo pygmaeus), all of which are previously unreported. We found that the osteocalcin sequences of Neanderthals, modern human, chimpanzee, and orangutan are unusual among mammals in that the ninth amino acid is proline (Pro-9), whereas most species have hydroxyproline (Hyp-9). Posttranslational hydroxylation of Pro-9 in osteocalcin by prolyl-4-hydroxylase requires adequate concentrations of vitamin C (l-ascorbic acid), molecular O(2), Fe(2+), and 2-oxoglutarate, and also depends on enzyme recognition of the target proline substrate consensus sequence Leu-Gly-Ala-Pro-9-Ala-Pro-Tyr occurring in most mammals. In five species with Pro-9-Val-10, hydroxylation is blocked, whereas in gorilla there is a mixture of Pro-9 and Hyp-9. We suggest that the absence of hydroxylation of Pro-9 in Pan, Pongo, and Homo may reflect response to a selective pressure related to a decline in vitamin C in the diet during omnivorous dietary adaptation, either independently or through the common ancestor of these species.

  20. Sequence protein identification by randomized sequence database and transcriptome mass spectrometry (SPIDER-TMS): from manual to automatic application of a 'de novo sequencing' approach.

    Science.gov (United States)

    Pascale, Raffaella; Grossi, Gerarda; Cruciani, Gabriele; Mecca, Giansalvatore; Santoro, Donatello; Sarli Calace, Renzo; Falabella, Patrizia; Bianco, Giuliana

    Sequence protein identification by a randomized sequence database and transcriptome mass spectrometry software package has been developed at the University of Basilicata in Potenza (Italy) and designed to facilitate the determination of the amino acid sequence of a peptide as well as an unequivocal identification of proteins in a high-throughput manner with enormous advantages of time, economical resource and expertise. The software package is a valid tool for the automation of a de novo sequencing approach, overcoming the main limits and a versatile platform useful in the proteomic field for an unequivocal identification of proteins, starting from tandem mass spectrometry data. The strength of this software is that it is a user-friendly and non-statistical approach, so protein identification can be considered unambiguous.

  1. Characterization of profilin polymorphism in pollen with a focus on multifunctionality.

    Directory of Open Access Journals (Sweden)

    Jose C Jimenez-Lopez

    Full Text Available Profilin, a multigene family involved in actin dynamics, is a multiple partners-interacting protein, as regard of the presence of at least of three binding domains encompassing actin, phosphoinositide lipids, and poly-L-proline interacting patches. In addition, pollen profilins are important allergens in several species like Olea europaea L. (Ole e 2, Betula pendula (Bet v 2, Phleum pratense (Phl p 12, Zea mays (Zea m 12 and Corylus avellana (Cor a 2. In spite of the biological and clinical importance of these molecules, variability in pollen profilin sequences has been poorly pointed out up until now. In this work, a relatively high number of pollen profilin sequences have been cloned, with the aim of carrying out an extensive characterization of their polymorphism among 24 olive cultivars and the above mentioned plant species. Our results indicate a high level of variability in the sequences analyzed. Quantitative intra-specific/varietal polymorphism was higher in comparison to inter-specific/cultivars comparisons. Multi-optional posttranslational modifications, e.g. phosphorylation sites, physicochemical properties, and partners-interacting functional residues have been shown to be affected by profilin polymorphism. As a result of this variability, profilins yielded a clear taxonomic separation between the five plant species. Profilin family multifunctionality might be inferred by natural variation through profilin isovariants generated among olive germplasm, as a result of polymorphism. The high variability might result in both differential profilin properties and differences in the regulation of the interaction with natural partners, affecting the mechanisms underlying the transmission of signals throughout signaling pathways in response to different stress environments. Moreover, elucidating the effect of profilin polymorphism in adaptive responses like actin dynamics, and cellular behavior, represents an exciting research goal for the

  2. AFLP fragment isolation technique as a method to produce random sequences for single nucleotide polymorphism discovery in the green turtle, Chelonia mydas.

    Science.gov (United States)

    Roden, Suzanne E; Dutton, Peter H; Morin, Phillip A

    2009-01-01

    The green sea turtle, Chelonia mydas, was used as a case study for single nucleotide polymorphism (SNP) discovery in a species that has little genetic sequence information available. As green turtles have a complex population structure, additional nuclear markers other than microsatellites could add to our understanding of their complex life history. Amplified fragment length polymorphism technique was used to generate sets of random fragments of genomic DNA, which were then electrophoretically separated with precast gels, stained with SYBR green, excised, and directly sequenced. It was possible to perform this method without the use of polyacrylamide gels, radioactive or fluorescent labeled primers, or hybridization methods, reducing the time, expense, and safety hazards of SNP discovery. Within 13 loci, 2547 base pairs were screened, resulting in the discovery of 35 SNPs. Using this method, it was possible to yield a sufficient number of loci to screen for SNP markers without the availability of prior sequence information.

  3. Blood Protein Polymorphism of Small Cattle Bred in Armenia

    Directory of Open Access Journals (Sweden)

    Gayane Marmaryan

    2017-05-01

    Full Text Available ABSTRACT Biochemical and genetic markers have not yet been used in selection and breeding of agricultural animals in Armenia. The objective behind the experiments was to assess the small cattle bred in Armenia- the semi-coarse wool and semi-fine wool sheep, as well as goats of different genotypes characterized by different milk productivity, polymorphic blood proteins-transferrin, ceruloplasmin, aiming to use them in breeding. Another objective was to study the genetic distance between the researched breeds and crossbreeds aiming to reveal the genetic similarity. The findings of research studies come to show that goats with polymorphic transferrin locus and a big set of genotypes are characterized by higher milk productivity; therefore it is recommended that they be used in selection as a supplementary milk productivity marker. The genetic distance between the researched goats comprised 0,60 which comes to show that semi-fine wool sheep were also bred with the help of semi-coarse wool crossbreeds. The high coefficients of genetic distance of crossbreed goats compared to both the highly milk producing imported breeds and the locals are indicative of inherited high adaptability qualities of locals and high milk producing quality of pure breeds.

  4. On the relationship between residue structural environment and sequence conservation in proteins.

    Science.gov (United States)

    Liu, Jen-Wei; Lin, Jau-Ji; Cheng, Chih-Wen; Lin, Yu-Feng; Hwang, Jenn-Kang; Huang, Tsun-Tsao

    2017-09-01

    Residues that are crucial to protein function or structure are usually evolutionarily conserved. To identify the important residues in protein, sequence conservation is estimated, and current methods rely upon the unbiased collection of homologous sequences. Surprisingly, our previous studies have shown that the sequence conservation is closely correlated with the weighted contact number (WCN), a measure of packing density for residue's structural environment, calculated only based on the C α positions of a protein structure. Moreover, studies have shown that sequence conservation is correlated with environment-related structural properties calculated based on different protein substructures, such as a protein's all atoms, backbone atoms, side-chain atoms, or side-chain centroid. To know whether the C α atomic positions are adequate to show the relationship between residue environment and sequence conservation or not, here we compared C α atoms with other substructures in their contributions to the sequence conservation. Our results show that C α positions are substantially equivalent to the other substructures in calculations of various measures of residue environment. As a result, the overlapping contributions between C α atoms and the other substructures are high, yielding similar structure-conservation relationship. Take the WCN as an example, the average overlapping contribution to sequence conservation is 87% between C α and all-atom substructures. These results indicate that only C α atoms of a protein structure could reflect sequence conservation at the residue level. © 2017 Wiley Periodicals, Inc.

  5. SAAS: Short Amino Acid Sequence - A Promising Protein Secondary Structure Prediction Method of Single Sequence

    Directory of Open Access Journals (Sweden)

    Zhou Yuan Wu

    2013-07-01

    Full Text Available In statistical methods of predicting protein secondary structure, many researchers focus on single amino acid frequencies in α-helices, β-sheets, and so on, or the impact near amino acids on an amino acid forming a secondary structure. But the paper considers a short sequence of amino acids (3, 4, 5 or 6 amino acids as integer, and statistics short sequence's probability forming secondary structure. Also, many researchers select low homologous sequences as statistical database. But this paper select whole PDB database. In this paper we propose a strategy to predict protein secondary structure using simple statistical method. Numerical computation shows that, short amino acids sequence as integer to statistics, which can easy see trend of short sequence forming secondary structure, and it will work well to select large statistical database (whole PDB database without considering homologous, and Q3 accuracy is ca. 74% using this paper proposed simple statistical method, but accuracy of others statistical methods is less than 70%.

  6. Computational identification of MoRFs in protein sequences.

    Science.gov (United States)

    Malhis, Nawar; Gsponer, Jörg

    2015-06-01

    Intrinsically disordered regions of proteins play an essential role in the regulation of various biological processes. Key to their regulatory function is the binding of molecular recognition features (MoRFs) to globular protein domains in a process known as a disorder-to-order transition. Predicting the location of MoRFs in protein sequences with high accuracy remains an important computational challenge. In this study, we introduce MoRFCHiBi, a new computational approach for fast and accurate prediction of MoRFs in protein sequences. MoRFCHiBi combines the outcomes of two support vector machine (SVM) models that take advantage of two different kernels with high noise tolerance. The first, SVMS, is designed to extract maximal information from the general contrast in amino acid compositions between MoRFs, their surrounding regions (Flanks), and the remainders of the sequences. The second, SVMT, is used to identify similarities between regions in a query sequence and MoRFs of the training set. We evaluated the performance of our predictor by comparing its results with those of two currently available MoRF predictors, MoRFpred and ANCHOR. Using three test sets that have previously been collected and used to evaluate MoRFpred and ANCHOR, we demonstrate that MoRFCHiBi outperforms the other predictors with respect to different evaluation metrics. In addition, MoRFCHiBi is downloadable and fast, which makes it useful as a component in other computational prediction tools. http://www.chibi.ubc.ca/morf/. © The Author 2015. Published by Oxford University Press. All rights reserved. For Permissions, please email: journals.permissions@oup.com.

  7. cDNA, genomic sequence cloning and overexpression of ribosomal protein S25 gene (RPS25) from the Giant Panda.

    Science.gov (United States)

    Hao, Yan-Zhe; Hou, Wan-Ru; Hou, Yi-Ling; Du, Yu-Jie; Zhang, Tian; Peng, Zheng-Song

    2009-11-01

    RPS25 is a component of the 40S small ribosomal subunit encoded by RPS25 gene, which is specific to eukaryotes. Studies in reference to RPS25 gene from animals were handful. The Giant Panda (Ailuropoda melanoleuca), known as a "living fossil", are increasingly concerned by the world community. Studies on RPS25 of the Giant Panda could provide scientific data for inquiring into the hereditary traits of the gene and formulating the protective strategy for the Giant Panda. The cDNA of the RPS25 cloned from Giant Panda is 436 bp in size, containing an open reading frame of 378 bp encoding 125 amino acids. The length of the genomic sequence is 1,992 bp, which was found to possess four exons and three introns. Alignment analysis indicated that the nucleotide sequence of the coding sequence shows a high homology to those of Homo sapiens, Bos taurus, Mus musculus and Rattus norvegicus as determined by Blast analysis, 92.6, 94.4, 89.2 and 91.5%, respectively. Primary structure analysis revealed that the molecular weight of the putative RPS25 protein is 13.7421 kDa with a theoretical pI 10.12. Topology prediction showed there is one N-glycosylation site, one cAMP and cGMP-dependent protein kinase phosphorylation site, two Protein kinase C phosphorylation sites and one Tyrosine kinase phosphorylation site in the RPS25 protein of the Giant Panda. The RPS25 gene was overexpressed in E. coli BL21 and Western Blotting of the RPS25 protein was also done. The results indicated that the RPS25 gene can be really expressed in E. coli and the RPS25 protein fusioned with the N-terminally his-tagged form gave rise to the accumulation of an expected 17.4 kDa polypeptide. The cDNA and the genomic sequence of RPS25 were cloned successfully for the first time from the Giant Panda using RT-PCR technology and Touchdown-PCR, respectively, which were both sequenced and analyzed preliminarily; then the cDNA of the RPS25 gene was overexpressed in E. coli BL21 and immunoblotted, which is the first

  8. The impact of particle preparation methods and polymorphic stability of lipid excipients on protein distribution in microparticles

    DEFF Research Database (Denmark)

    Liu, Jingying; Christophersen, Philip C; Yang, Mingshi

    2017-01-01

    OBJECTIVE: The present study aimed at elucidating the influence of polymorphic stability of lipid excipients on the physicochemical characters of different solid lipid microparticles (SLM), with the focus on the alteration of protein distribution in SLM. METHODS: Labeled lysozyme was incorporated...... provides updated knowledge for rational development of lipid-based formulations for oral delivery of peptide or protein drugs.......OBJECTIVE: The present study aimed at elucidating the influence of polymorphic stability of lipid excipients on the physicochemical characters of different solid lipid microparticles (SLM), with the focus on the alteration of protein distribution in SLM. METHODS: Labeled lysozyme was incorporated...... into SLM prepared with different excipients, i.e. trimyristin (TG14), glyceryl distearate (GDS), and glyceryl monostearate (GMS), by water-oil-water (w/o/w) or solid-oil-water (s/o/w) method. The distribution of lysozyme in SLM and the release of the protein from SLM were evaluated by confocal laser...

  9. A polyvalent hybrid protein elicits antibodies against the diverse allelic types of block 2 in Plasmodium falciparum merozoite surface protein 1.

    Science.gov (United States)

    Tetteh, Kevin K A; Conway, David J

    2011-10-13

    Merozoite surface protein 1 (MSP1) of Plasmodium falciparum has been implicated as an important target of acquired immunity, and candidate components for a vaccine include polymorphic epitopes in the N-terminal polymorphic block 2 region. We designed a polyvalent hybrid recombinant protein incorporating sequences of the three major allelic types of block 2 together with a composite repeat sequence of one of the types and N-terminal flanking T cell epitopes, and compared this with a series of recombinant proteins containing modular sub-components and similarly expressed in Escherichia coli. Immunogenicity of the full polyvalent hybrid protein was tested in both mice and rabbits, and comparative immunogenicity studies of the sub-component modules were performed in mice. The full hybrid protein induced high titre antibodies against each of the major block 2 allelic types expressed as separate recombinant proteins and against a wide range of allelic types naturally expressed by a panel of diverse P. falciparum isolates, while the sub-component modules had partial antigenic coverage as expected. This encourages further development and evaluation of the full MSP1 block 2 polyvalent hybrid protein as a candidate blood-stage component of a malaria vaccine. Copyright © 2011 Elsevier Ltd. All rights reserved.

  10. Polymorphism of BMP4 gene in Indian goat breeds differing in prolificacy.

    Science.gov (United States)

    Sharma, Rekha; Ahlawat, Sonika; Maitra, A; Roy, Manoranjan; Mandakmale, S; Tantia, M S

    2013-12-10

    Bone morphogenetic proteins (BMPs) are members of the TGF-β (transforming growth factor-beta) superfamily, of which BMP4 is the most important due to its crucial role in follicular growth and differentiation, cumulus expansion and ovulation. Reproduction is a crucial trait in goat breeding and based on the important role of BMP4 gene in reproduction it was considered as a possible candidate gene for the prolificacy of goats. The objective of the present study was to detect polymorphism in intronic, exonic and 3' un-translated regions of BMP4 gene in Indian goats. Nine different goat breeds (Barbari, Beetal, Black Bengal, Malabari, Jakhrana (Twinning>40%), Osmanabadi, Sangamneri (Twinning 20-30%), Sirohi and Ganjam (Twinning<10%)) differing in prolificacy and geographic distribution were employed for polymorphism scanning. Cattle sequence (AC_000167.1) was used to design primers for the amplification of a targeted region followed by direct DNA sequencing to identify the genetic variations. Single nucleotide polymorphisms (SNPs) were not detected in exon 3, the intronic region and the 3' flanking region. A SNP (G1534A) was identified in exon 2. It was a non-synonymous mutation resulting in an arginine to lysine change in a corresponding protein sequence. G to A transition at the 1534 locus revealed two genotypes GG and GA in the nine investigated goat breeds. The GG genotype was predominant with a genotype frequency of 0.98. The GA genotype was present in the Black Bengal as well as Jakhrana breed with a genotype frequency of 0.02. A microsatellite was identified in the 3' flanking region, only 20 nucleotides downstream from the termination site of the coding region, as a short sequence with more than nineteen continuous and repeated CA dinucleotides. Since the gene is highly evolutionarily conserved, identification of a non-synonymous SNP (G1534A) in the coding region gains further importance. To our knowledge, this is the first report of a mutation in the coding

  11. Associations of ECP (eosinophil cationic protein-gene polymorphisms to allergy, asthma, smoke habits and lung function in two Estonian and Swedish sub cohorts of the ECRHS II study

    Directory of Open Access Journals (Sweden)

    Janson Christer

    2010-06-01

    Full Text Available Abstract Background The Eosinophil Cationic Protein (ECP is a potent multifunctional protein. Three common polymorphisms are present in the ECP gene, which determine the function and production of the protein. The aim was to study the relationship of these ECP gene polymorphisms to signs and symptoms of allergy and asthma in a community based cohort (The European Community Respiratory Health Survey (ECRHS. Methods Swedish and Estonian subjects (n = 757 were selected from the larger cohort of the ECRHS II study cohort. The prevalence of the gene polymorphisms ECP434(G>C (rs2073342, ECP562(G>C (rs2233860 and ECP c.-38(A>C (rs2233859 were analysed by DNA sequencing and/or real-time PCR and related to questionnaire-based information of allergy, asthma, smoking habits and to lung functions. Results Genotype prevalence showed both ethnic and gender differences. Close associations were found between the ECP434(G>C and ECP562(G>C genotypes and smoking habits, lung function and expression of allergic symptoms. Non-allergic asthma was associated with an increased prevalence of the ECP434GG genotype. The ECP c.-38(A>C genotypes were independently associated to the subject being atopic. Conclusion Our results show associations of symptoms of allergy and asthma to ECP-genotypes, but also to smoking habits. ECP may be involved in impairment of lung functions in disease. Gender, ethnicity and smoking habits are major confounders in the evaluations of genetic associations to allergy and asthma.

  12. Single nucleotide polymorphism discovery in rainbow trout by deep sequencing of a reduced representation library

    Directory of Open Access Journals (Sweden)

    Salem Mohamed

    2009-11-01

    Full Text Available Abstract Background To enhance capabilities for genomic analyses in rainbow trout, such as genomic selection, a large suite of polymorphic markers that are amenable to high-throughput genotyping protocols must be identified. Expressed Sequence Tags (ESTs have been used for single nucleotide polymorphism (SNP discovery in salmonids. In those strategies, the salmonid semi-tetraploid genomes often led to assemblies of paralogous sequences and therefore resulted in a high rate of false positive SNP identification. Sequencing genomic DNA using primers identified from ESTs proved to be an effective but time consuming methodology of SNP identification in rainbow trout, therefore not suitable for high throughput SNP discovery. In this study, we employed a high-throughput strategy that used pyrosequencing technology to generate data from a reduced representation library constructed with genomic DNA pooled from 96 unrelated rainbow trout that represent the National Center for Cool and Cold Water Aquaculture (NCCCWA broodstock population. Results The reduced representation library consisted of 440 bp fragments resulting from complete digestion with the restriction enzyme HaeIII; sequencing produced 2,000,000 reads providing an average 6 fold coverage of the estimated 150,000 unique genomic restriction fragments (300,000 fragment ends. Three independent data analyses identified 22,022 to 47,128 putative SNPs on 13,140 to 24,627 independent contigs. A set of 384 putative SNPs, randomly selected from the sets produced by the three analyses were genotyped on individual fish to determine the validation rate of putative SNPs among analyses, distinguish apparent SNPs that actually represent paralogous loci in the tetraploid genome, examine Mendelian segregation, and place the validated SNPs on the rainbow trout linkage map. Approximately 48% (183 of the putative SNPs were validated; 167 markers were successfully incorporated into the rainbow trout linkage map. In

  13. Single nucleotide polymorphism discovery in rainbow trout by deep sequencing of a reduced representation library.

    Science.gov (United States)

    Sánchez, Cecilia Castaño; Smith, Timothy P L; Wiedmann, Ralph T; Vallejo, Roger L; Salem, Mohamed; Yao, Jianbo; Rexroad, Caird E

    2009-11-25

    To enhance capabilities for genomic analyses in rainbow trout, such as genomic selection, a large suite of polymorphic markers that are amenable to high-throughput genotyping protocols must be identified. Expressed Sequence Tags (ESTs) have been used for single nucleotide polymorphism (SNP) discovery in salmonids. In those strategies, the salmonid semi-tetraploid genomes often led to assemblies of paralogous sequences and therefore resulted in a high rate of false positive SNP identification. Sequencing genomic DNA using primers identified from ESTs proved to be an effective but time consuming methodology of SNP identification in rainbow trout, therefore not suitable for high throughput SNP discovery. In this study, we employed a high-throughput strategy that used pyrosequencing technology to generate data from a reduced representation library constructed with genomic DNA pooled from 96 unrelated rainbow trout that represent the National Center for Cool and Cold Water Aquaculture (NCCCWA) broodstock population. The reduced representation library consisted of 440 bp fragments resulting from complete digestion with the restriction enzyme HaeIII; sequencing produced 2,000,000 reads providing an average 6 fold coverage of the estimated 150,000 unique genomic restriction fragments (300,000 fragment ends). Three independent data analyses identified 22,022 to 47,128 putative SNPs on 13,140 to 24,627 independent contigs. A set of 384 putative SNPs, randomly selected from the sets produced by the three analyses were genotyped on individual fish to determine the validation rate of putative SNPs among analyses, distinguish apparent SNPs that actually represent paralogous loci in the tetraploid genome, examine Mendelian segregation, and place the validated SNPs on the rainbow trout linkage map. Approximately 48% (183) of the putative SNPs were validated; 167 markers were successfully incorporated into the rainbow trout linkage map. In addition, 2% of the sequences from the

  14. MannDB – A microbial database of automated protein sequence analyses and evidence integration for protein characterization

    Directory of Open Access Journals (Sweden)

    Kuczmarski Thomas A

    2006-10-01

    Full Text Available Abstract Background MannDB was created to meet a need for rapid, comprehensive automated protein sequence analyses to support selection of proteins suitable as targets for driving the development of reagents for pathogen or protein toxin detection. Because a large number of open-source tools were needed, it was necessary to produce a software system to scale the computations for whole-proteome analysis. Thus, we built a fully automated system for executing software tools and for storage, integration, and display of automated protein sequence analysis and annotation data. Description MannDB is a relational database that organizes data resulting from fully automated, high-throughput protein-sequence analyses using open-source tools. Types of analyses provided include predictions of cleavage, chemical properties, classification, features, functional assignment, post-translational modifications, motifs, antigenicity, and secondary structure. Proteomes (lists of hypothetical and known proteins are downloaded and parsed from Genbank and then inserted into MannDB, and annotations from SwissProt are downloaded when identifiers are found in the Genbank entry or when identical sequences are identified. Currently 36 open-source tools are run against MannDB protein sequences either on local systems or by means of batch submission to external servers. In addition, BLAST against protein entries in MvirDB, our database of microbial virulence factors, is performed. A web client browser enables viewing of computational results and downloaded annotations, and a query tool enables structured and free-text search capabilities. When available, links to external databases, including MvirDB, are provided. MannDB contains whole-proteome analyses for at least one representative organism from each category of biological threat organism listed by APHIS, CDC, HHS, NIAID, USDA, USFDA, and WHO. Conclusion MannDB comprises a large number of genomes and comprehensive protein

  15. Nonlinear analysis of sequence repeats of multi-domain proteins

    Energy Technology Data Exchange (ETDEWEB)

    Huang Yanzhao [Biomolecular Physics and Modeling Group, Department of Physics, Huazhong University of Science and Technology, Wuhan 430074, Hubei (China); Li Mingfeng [Biomolecular Physics and Modeling Group, Department of Physics, Huazhong University of Science and Technology, Wuhan 430074, Hubei (China); Xiao Yi [Biomolecular Physics and Modeling Group, Department of Physics, Huazhong University of Science and Technology, Wuhan 430074, Hubei (China)]. E-mail: lmf_bill@sina.com

    2007-11-15

    Many multi-domain proteins have repetitive three-dimensional structures but nearly-random amino acid sequences. In the present paper, by using a modified recurrence plot proposed by us previously, we show that these amino acid sequences have hidden repetitions in fact. These results indicate that the repetitive domain structures are encoded by the repetitive sequences. This also gives a method to detect the repetitive domain structures directly from amino acid sequences.

  16. The SWISS-PROT protein sequence data bank

    OpenAIRE

    Bairoch, Amos; Boeckmann, Brigitte

    1992-01-01

    SWISS-PROT is an annotated protein sequence database established in 1986 and maintained collaboratively, since 1988, by the Department of Medical Biochemistry of the University of Geneva and the EMBL Data Library

  17. Using sequence similarity networks for visualization of relationships across diverse protein superfamilies.

    Directory of Open Access Journals (Sweden)

    Holly J Atkinson

    Full Text Available The dramatic increase in heterogeneous types of biological data--in particular, the abundance of new protein sequences--requires fast and user-friendly methods for organizing this information in a way that enables functional inference. The most widely used strategy to link sequence or structure to function, homology-based function prediction, relies on the fundamental assumption that sequence or structural similarity implies functional similarity. New tools that extend this approach are still urgently needed to associate sequence data with biological information in ways that accommodate the real complexity of the problem, while being accessible to experimental as well as computational biologists. To address this, we have examined the application of sequence similarity networks for visualizing functional trends across protein superfamilies from the context of sequence similarity. Using three large groups of homologous proteins of varying types of structural and functional diversity--GPCRs and kinases from humans, and the crotonase superfamily of enzymes--we show that overlaying networks with orthogonal information is a powerful approach for observing functional themes and revealing outliers. In comparison to other primary methods, networks provide both a good representation of group-wise sequence similarity relationships and a strong visual and quantitative correlation with phylogenetic trees, while enabling analysis and visualization of much larger sets of sequences than trees or multiple sequence alignments can easily accommodate. We also define important limitations and caveats in the application of these networks. As a broadly accessible and effective tool for the exploration of protein superfamilies, sequence similarity networks show great potential for generating testable hypotheses about protein structure-function relationships.

  18. Using sequence similarity networks for visualization of relationships across diverse protein superfamilies.

    Science.gov (United States)

    Atkinson, Holly J; Morris, John H; Ferrin, Thomas E; Babbitt, Patricia C

    2009-01-01

    The dramatic increase in heterogeneous types of biological data--in particular, the abundance of new protein sequences--requires fast and user-friendly methods for organizing this information in a way that enables functional inference. The most widely used strategy to link sequence or structure to function, homology-based function prediction, relies on the fundamental assumption that sequence or structural similarity implies functional similarity. New tools that extend this approach are still urgently needed to associate sequence data with biological information in ways that accommodate the real complexity of the problem, while being accessible to experimental as well as computational biologists. To address this, we have examined the application of sequence similarity networks for visualizing functional trends across protein superfamilies from the context of sequence similarity. Using three large groups of homologous proteins of varying types of structural and functional diversity--GPCRs and kinases from humans, and the crotonase superfamily of enzymes--we show that overlaying networks with orthogonal information is a powerful approach for observing functional themes and revealing outliers. In comparison to other primary methods, networks provide both a good representation of group-wise sequence similarity relationships and a strong visual and quantitative correlation with phylogenetic trees, while enabling analysis and visualization of much larger sets of sequences than trees or multiple sequence alignments can easily accommodate. We also define important limitations and caveats in the application of these networks. As a broadly accessible and effective tool for the exploration of protein superfamilies, sequence similarity networks show great potential for generating testable hypotheses about protein structure-function relationships.

  19. Polymorphic DNA sequences of the fungal honey bee pathogen Ascosphaera apis

    DEFF Research Database (Denmark)

    Jensen, Annette B; Welker, Dennis L; Kryger, Per

    2012-01-01

    The pathogenic fungus Ascosphaera apis is ubiquitous in honey bee populations. We used the draft genome assembly of this pathogen to search for polymorphic intergenic loci that could be used to differentiate haplotypes. Primers were developed for five such loci, and the species specificities were...... verified using DNA from nine closely related species. The sequence variation was compared among 12 A. apis isolates at each of these loci, and two additional loci, the internal transcribed spacer of the ribosomal RNA (ITS) and a variable part of the elongation factor 1α (Ef1α). The degree of variation...... was then compared among the different loci, and three were found to have the greatest detection power for identifying A. apis haplotypes. The described loci can help to resolve strain differences and population genetic structures, to elucidate host–pathogen interaction and to test evolutionary hypotheses...

  20. c.1810C>T Polymorphism of NTRK1 Gene is associated with reduced Survival in Neuroblastoma Patients

    International Nuclear Information System (INIS)

    Lipska, Beata S; Biernat, Wojciech; Limon, Janusz; Drożynska, Elżbieta; Scaruffi, Paola; Tonini, Gian Paolo; Iżycka-Świeszewska, Ewa; Ziętkiewicz, Szymon; Balcerska, Anna; Perek, Danuta; Chybicka, Alicja

    2009-01-01

    TrkA (encoded by NTRK1 gene), the high-affinity tyrosine kinase receptor for neurotrophins, is involved in neural crest cell differentiation. Its expression has been reported to be associated with a favourable prognosis in neuroblastoma. Therefore, the entire coding sequence of NTRK1 gene has been analysed in order to identify mutations and/or polymorphisms which may alter TrkA receptor expression. DNA was extracted from neuroblastomas of 55 Polish and 114 Italian patients and from peripheral blood leukocytes of 158 healthy controls. Denaturing High-Performance Liquid Chromatography (DHPLC) and Single-Strand Conformation Polymorphism (SSCP) analysis were used to screen for sequence variants. Genetic changes were confirmed by direct sequencing and correlated with biological and clinical data. Three previously reported and nine new single nucleotide polymorphisms were detected. c.1810C>T polymorphism present in 8.7% of cases was found to be an independent marker of disease recurrence (OR = 13.3; p = 0.009) associated with lower survival rates (HR = 4.45 p = 0.041). c.1810C>T polymorphism's unfavourable prognostic value was most significant in patients under 18 months of age with no MYCN amplification (HR = 26; p = 0.008). In-silico analysis of the c.1810C>T polymorphism suggests that the substitution of the corresponding amino acid residue within the conservative region of the tyrosine kinase domain might theoretically interfere with the functioning of the TrkA protein. NTRK1 c.1810C>T polymorphism appears to be a new independent prognostic factor of poor outcome in neuroblastoma, especially in children under 18 months of age with no MYCN amplification

  1. Genetic alterations of hepatocellular carcinoma by random amplified polymorphic DNA analysis and cloning sequencing of tumor differential DNA fragment

    Science.gov (United States)

    Xian, Zhi-Hong; Cong, Wen-Ming; Zhang, Shu-Hui; Wu, Meng-Chao

    2005-01-01

    AIM: To study the genetic alterations and their association with clinicopathological characteristics of hepatocellular carcinoma (HCC), and to find the tumor related DNA fragments. METHODS: DNA isolated from tumors and corresponding noncancerous liver tissues of 56 HCC patients was amplified by random amplified polymorphic DNA (RAPD) with 10 random 10-mer arbitrary primers. The RAPD bands showing obvious differences in tumor tissue DNA corresponding to that of normal tissue were separated, purified, cloned and sequenced. DNA sequences were analyzed and compared with GenBank data. RESULTS: A total of 56 cases of HCC were demonstrated to have genetic alterations, which were detected by at least one primer. The detestability of genetic alterations ranged from 20% to 70% in each case, and 17.9% to 50% in each primer. Serum HBV infection, tumor size, histological grade, tumor capsule, as well as tumor intrahepatic metastasis, might be correlated with genetic alterations on certain primers. A band with a higher intensity of 480 bp or so amplified fragments in tumor DNA relative to normal DNA could be seen in 27 of 56 tumor samples using primer 4. Sequence analysis of these fragments showed 91% homology with Homo sapiens double homeobox protein DUX10 gene. CONCLUSION: Genetic alterations are a frequent event in HCC, and tumor related DNA fragments have been found in this study, which may be associated with hepatocarcin-ogenesis. RAPD is an effective method for the identification and analysis of genetic alterations in HCC, and may provide new information for further evaluating the molecular mechanism of hepatocarcinogenesis. PMID:15996039

  2. Direct chloroplast sequencing: comparison of sequencing platforms and analysis tools for whole chloroplast barcoding.

    Directory of Open Access Journals (Sweden)

    Marta Brozynska

    Full Text Available Direct sequencing of total plant DNA using next generation sequencing technologies generates a whole chloroplast genome sequence that has the potential to provide a barcode for use in plant and food identification. Advances in DNA sequencing platforms may make this an attractive approach for routine plant identification. The HiSeq (Illumina and Ion Torrent (Life Technology sequencing platforms were used to sequence total DNA from rice to identify polymorphisms in the whole chloroplast genome sequence of a wild rice plant relative to cultivated rice (cv. Nipponbare. Consensus chloroplast sequences were produced by mapping sequence reads to the reference rice chloroplast genome or by de novo assembly and mapping of the resulting contigs to the reference sequence. A total of 122 polymorphisms (SNPs and indels between the wild and cultivated rice chloroplasts were predicted by these different sequencing and analysis methods. Of these, a total of 102 polymorphisms including 90 SNPs were predicted by both platforms. Indels were more variable with different sequencing methods, with almost all discrepancies found in homopolymers. The Ion Torrent platform gave no apparent false SNP but was less reliable for indels. The methods should be suitable for routine barcoding using appropriate combinations of sequencing platform and data analysis.

  3. PredPPCrys: accurate prediction of sequence cloning, protein production, purification and crystallization propensity from protein sequences using multi-step heterogeneous feature fusion and selection.

    Directory of Open Access Journals (Sweden)

    Huilin Wang

    Full Text Available X-ray crystallography is the primary approach to solve the three-dimensional structure of a protein. However, a major bottleneck of this method is the failure of multi-step experimental procedures to yield diffraction-quality crystals, including sequence cloning, protein material production, purification, crystallization and ultimately, structural determination. Accordingly, prediction of the propensity of a protein to successfully undergo these experimental procedures based on the protein sequence may help narrow down laborious experimental efforts and facilitate target selection. A number of bioinformatics methods based on protein sequence information have been developed for this purpose. However, our knowledge on the important determinants of propensity for a protein sequence to produce high diffraction-quality crystals remains largely incomplete. In practice, most of the existing methods display poorer performance when evaluated on larger and updated datasets. To address this problem, we constructed an up-to-date dataset as the benchmark, and subsequently developed a new approach termed 'PredPPCrys' using the support vector machine (SVM. Using a comprehensive set of multifaceted sequence-derived features in combination with a novel multi-step feature selection strategy, we identified and characterized the relative importance and contribution of each feature type to the prediction performance of five individual experimental steps required for successful crystallization. The resulting optimal candidate features were used as inputs to build the first-level SVM predictor (PredPPCrys I. Next, prediction outputs of PredPPCrys I were used as the input to build second-level SVM classifiers (PredPPCrys II, which led to significantly enhanced prediction performance. Benchmarking experiments indicated that our PredPPCrys method outperforms most existing procedures on both up-to-date and previous datasets. In addition, the predicted crystallization

  4. Microwave-assisted acid and base hydrolysis of intact proteins containing disulfide bonds for protein sequence analysis by mass spectrometry.

    Science.gov (United States)

    Reiz, Bela; Li, Liang

    2010-09-01

    Controlled hydrolysis of proteins to generate peptide ladders combined with mass spectrometric analysis of the resultant peptides can be used for protein sequencing. In this paper, two methods of improving the microwave-assisted protein hydrolysis process are described to enable rapid sequencing of proteins containing disulfide bonds and increase sequence coverage, respectively. It was demonstrated that proteins containing disulfide bonds could be sequenced by MS analysis by first performing hydrolysis for less than 2 min, followed by 1 h of reduction to release the peptides originally linked by disulfide bonds. It was shown that a strong base could be used as a catalyst for microwave-assisted protein hydrolysis, producing complementary sequence information to that generated by microwave-assisted acid hydrolysis. However, using either acid or base hydrolysis, amide bond breakages in small regions of the polypeptide chains of the model proteins (e.g., cytochrome c and lysozyme) were not detected. Dynamic light scattering measurement of the proteins solubilized in an acid or base indicated that protein-protein interaction or aggregation was not the cause of the failure to hydrolyze certain amide bonds. It was speculated that there were some unknown local structures that might play a role in preventing an acid or base from reacting with the peptide bonds therein. 2010 American Society for Mass Spectrometry. Published by Elsevier Inc. All rights reserved.

  5. Sequence-based prediction of protein-binding sites in DNA: comparative study of two SVM models.

    Science.gov (United States)

    Park, Byungkyu; Im, Jinyong; Tuvshinjargal, Narankhuu; Lee, Wook; Han, Kyungsook

    2014-11-01

    As many structures of protein-DNA complexes have been known in the past years, several computational methods have been developed to predict DNA-binding sites in proteins. However, its inverse problem (i.e., predicting protein-binding sites in DNA) has received much less attention. One of the reasons is that the differences between the interaction propensities of nucleotides are much smaller than those between amino acids. Another reason is that DNA exhibits less diverse sequence patterns than protein. Therefore, predicting protein-binding DNA nucleotides is much harder than predicting DNA-binding amino acids. We computed the interaction propensity (IP) of nucleotide triplets with amino acids using an extensive dataset of protein-DNA complexes, and developed two support vector machine (SVM) models that predict protein-binding nucleotides from sequence data alone. One SVM model predicts protein-binding nucleotides using DNA sequence data alone, and the other SVM model predicts protein-binding nucleotides using both DNA and protein sequences. In a 10-fold cross-validation with 1519 DNA sequences, the SVM model that uses DNA sequence data only predicted protein-binding nucleotides with an accuracy of 67.0%, an F-measure of 67.1%, and a Matthews correlation coefficient (MCC) of 0.340. With an independent dataset of 181 DNAs that were not used in training, it achieved an accuracy of 66.2%, an F-measure 66.3% and a MCC of 0.324. Another SVM model that uses both DNA and protein sequences achieved an accuracy of 69.6%, an F-measure of 69.6%, and a MCC of 0.383 in a 10-fold cross-validation with 1519 DNA sequences and 859 protein sequences. With an independent dataset of 181 DNAs and 143 proteins, it showed an accuracy of 67.3%, an F-measure of 66.5% and a MCC of 0.329. Both in cross-validation and independent testing, the second SVM model that used both DNA and protein sequence data showed better performance than the first model that used DNA sequence data. To the best of

  6. Analysis of long-range correlation in sequences data of proteins

    Directory of Open Access Journals (Sweden)

    ADRIANA ISVORAN

    2007-04-01

    Full Text Available The results presented here suggest the existence of correlations in the sequence data of proteins. 32 proteins, both globular and fibrous, both monomeric and polymeric, were analyzed. The primary structures of these proteins were treated as time series. Three spatial series of data for each sequence of a protein were generated from numerical correspondences between each amino acid and a physical property associated with it, i.e., its electric charge, its polar character and its dipole moment. For each series, the spectral coefficient, the scaling exponent and the Hurst coefficient were determined. The values obtained for these coefficients revealed non-randomness in the series of data.

  7. Automatic discovery of cross-family sequence features associated with protein function

    Directory of Open Access Journals (Sweden)

    Krings Andrea

    2006-01-01

    Full Text Available Abstract Background Methods for predicting protein function directly from amino acid sequences are useful tools in the study of uncharacterised protein families and in comparative genomics. Until now, this problem has been approached using machine learning techniques that attempt to predict membership, or otherwise, to predefined functional categories or subcellular locations. A potential drawback of this approach is that the human-designated functional classes may not accurately reflect the underlying biology, and consequently important sequence-to-function relationships may be missed. Results We show that a self-supervised data mining approach is able to find relationships between sequence features and functional annotations. No preconceived ideas about functional categories are required, and the training data is simply a set of protein sequences and their UniProt/Swiss-Prot annotations. The main technical aspect of the approach is the co-evolution of amino acid-based regular expressions and keyword-based logical expressions with genetic programming. Our experiments on a strictly non-redundant set of eukaryotic proteins reveal that the strongest and most easily detected sequence-to-function relationships are concerned with targeting to various cellular compartments, which is an area already well studied both experimentally and computationally. Of more interest are a number of broad functional roles which can also be correlated with sequence features. These include inhibition, biosynthesis, transcription and defence against bacteria. Despite substantial overlaps between these functions and their corresponding cellular compartments, we find clear differences in the sequence motifs used to predict some of these functions. For example, the presence of polyglutamine repeats appears to be linked more strongly to the "transcription" function than to the general "nuclear" function/location. Conclusion We have developed a novel and useful approach for

  8. Protein sequences from mastodon and Tyrannosaurus rex revealed by mass spectrometry.

    Science.gov (United States)

    Asara, John M; Schweitzer, Mary H; Freimark, Lisa M; Phillips, Matthew; Cantley, Lewis C

    2007-04-13

    Fossilized bones from extinct taxa harbor the potential for obtaining protein or DNA sequences that could reveal evolutionary links to extant species. We used mass spectrometry to obtain protein sequences from bones of a 160,000- to 600,000-year-old extinct mastodon (Mammut americanum) and a 68-million-year-old dinosaur (Tyrannosaurus rex). The presence of T. rex sequences indicates that their peptide bonds were remarkably stable. Mass spectrometry can thus be used to determine unique sequences from ancient organisms from peptide fragmentation patterns, a valuable tool to study the evolution and adaptation of ancient taxa from which genomic sequences are unlikely to be obtained.

  9. Structural insights and ab initio sequencing within the DING proteins family

    International Nuclear Information System (INIS)

    Elias, Mikael; Liebschner, Dorothee; Gotthard, Guillaume; Chabriere, Eric

    2011-01-01

    DING proteins constitute a recently discovered protein family that is ubiquitous in eukaryotes. The structural insights and the physiological involvements of these intriguing proteins are hereby deciphered. DING proteins constitute an intriguing family of phosphate-binding proteins that was identified in a wide range of organisms, from prokaryotes and archae to eukaryotes. Despite their seemingly ubiquitous occurrence in eukaryotes, their encoding genes are missing from sequenced genomes. Such a lack has considerably hampered functional studies. In humans, these proteins have been related to several diseases, like atherosclerosis, kidney stones, inflammation processes and HIV inhibition. The human phosphate binding protein is a human representative of the DING family that was serendipitously discovered from human plasma. An original approach was developed to determine ab initio the complete and exact sequence of this 38 kDa protein by utilizing mass spectrometry and X-ray data in tandem. Taking advantage of this first complete eukaryotic DING sequence, a immunohistochemistry study was undertaken to check the presence of DING proteins in various mice tissues, revealing that these proteins are widely expressed. Finally, the structure of a bacterial representative from Pseudomonas fluorescens was solved at sub-angstrom resolution, allowing the molecular mechanism of the phosphate binding in these high-affinity proteins to be elucidated

  10. Structural insights and ab initio sequencing within the DING proteins family

    Energy Technology Data Exchange (ETDEWEB)

    Elias, Mikael, E-mail: mikael.elias@weizmann.ac.il [Weizmann Institute of Science, Rehovot (Israel); Liebschner, Dorothee [CRM2, Nancy Université (France); Gotthard, Guillaume; Chabriere, Eric [AFMB, Université Aix-Marseille II (France)

    2011-01-01

    DING proteins constitute a recently discovered protein family that is ubiquitous in eukaryotes. The structural insights and the physiological involvements of these intriguing proteins are hereby deciphered. DING proteins constitute an intriguing family of phosphate-binding proteins that was identified in a wide range of organisms, from prokaryotes and archae to eukaryotes. Despite their seemingly ubiquitous occurrence in eukaryotes, their encoding genes are missing from sequenced genomes. Such a lack has considerably hampered functional studies. In humans, these proteins have been related to several diseases, like atherosclerosis, kidney stones, inflammation processes and HIV inhibition. The human phosphate binding protein is a human representative of the DING family that was serendipitously discovered from human plasma. An original approach was developed to determine ab initio the complete and exact sequence of this 38 kDa protein by utilizing mass spectrometry and X-ray data in tandem. Taking advantage of this first complete eukaryotic DING sequence, a immunohistochemistry study was undertaken to check the presence of DING proteins in various mice tissues, revealing that these proteins are widely expressed. Finally, the structure of a bacterial representative from Pseudomonas fluorescens was solved at sub-angstrom resolution, allowing the molecular mechanism of the phosphate binding in these high-affinity proteins to be elucidated.

  11. Solving Classification Problems for Large Sets of Protein Sequences with the Example of Hox and ParaHox Proteins

    Directory of Open Access Journals (Sweden)

    Stefanie D. Hueber

    2016-02-01

    Full Text Available Phylogenetic methods are key to providing models for how a given protein family evolved. However, these methods run into difficulties when sequence divergence is either too low or too high. Here, we provide a case study of Hox and ParaHox proteins so that additional insights can be gained using a new computational approach to help solve old classification problems. For two (Gsx and Cdx out of three ParaHox proteins the assignments differ between the currently most established view and four alternative scenarios. We use a non-phylogenetic, pairwise-sequence-similarity-based method to assess which of the previous predictions, if any, are best supported by the sequence-similarity relationships between Hox and ParaHox proteins. The overall sequence-similarities show Gsx to be most similar to Hox2–3, and Cdx to be most similar to Hox4–8. The results indicate that a purely pairwise-sequence-similarity-based approach can provide additional information not only when phylogenetic inference methods have insufficient information to provide reliable classifications (as was shown previously for central Hox proteins, but also when the sequence variation is so high that the resulting phylogenetic reconstructions are likely plagued by long-branch-attraction artifacts.

  12. Protein Science by DNA Sequencing: How Advances in Molecular Biology Are Accelerating Biochemistry.

    Science.gov (United States)

    Higgins, Sean A; Savage, David F

    2018-01-09

    A fundamental goal of protein biochemistry is to determine the sequence-function relationship, but the vastness of sequence space makes comprehensive evaluation of this landscape difficult. However, advances in DNA synthesis and sequencing now allow researchers to assess the functional impact of every single mutation in many proteins, but challenges remain in library construction and the development of general assays applicable to a diverse range of protein functions. This Perspective briefly outlines the technical innovations in DNA manipulation that allow massively parallel protein biochemistry and then summarizes the methods currently available for library construction and the functional assays of protein variants. Areas in need of future innovation are highlighted with a particular focus on assay development and the use of computational analysis with machine learning to effectively traverse the sequence-function landscape. Finally, applications in the fundamentals of protein biochemistry, disease prediction, and protein engineering are presented.

  13. An algorithm to find all palindromic sequences in proteins

    Indian Academy of Sciences (India)

    2013-01-20

    Jan 20, 2013 ... 1976; Karrer and Gall 1976; Vogt and Braun 1976) and (iii) in the formation of hairpin loops in the newly transcribed RNA. Palindromic sequences are observed in various classes of proteins like histones (Cheng et al. 1989), prion proteins (Sulkowski 1992; Kazim 1993),. DNA-binding proteins (Suzuki 1992; ...

  14. 3D representations of amino acids—applications to protein sequence comparison and classification

    Directory of Open Access Journals (Sweden)

    Jie Li

    2014-08-01

    Full Text Available The amino acid sequence of a protein is the key to understanding its structure and ultimately its function in the cell. This paper addresses the fundamental issue of encoding amino acids in ways that the representation of such a protein sequence facilitates the decoding of its information content. We show that a feature-based representation in a three-dimensional (3D space derived from amino acid substitution matrices provides an adequate representation that can be used for direct comparison of protein sequences based on geometry. We measure the performance of such a representation in the context of the protein structural fold prediction problem. We compare the results of classifying different sets of proteins belonging to distinct structural folds against classifications of the same proteins obtained from sequence alone or directly from structural information. We find that sequence alone performs poorly as a structure classifier. We show in contrast that the use of the three dimensional representation of the sequences significantly improves the classification accuracy. We conclude with a discussion of the current limitations of such a representation and with a description of potential improvements.

  15. Sequence- and interactome-based prediction of viral protein hotspots targeting host proteins: a case study for HIV Nef.

    Directory of Open Access Journals (Sweden)

    Mahdi Sarmady

    Full Text Available Virus proteins alter protein pathways of the host toward the synthesis of viral particles by breaking and making edges via binding to host proteins. In this study, we developed a computational approach to predict viral sequence hotspots for binding to host proteins based on sequences of viral and host proteins and literature-curated virus-host protein interactome data. We use a motif discovery algorithm repeatedly on collections of sequences of viral proteins and immediate binding partners of their host targets and choose only those motifs that are conserved on viral sequences and highly statistically enriched among binding partners of virus protein targeted host proteins. Our results match experimental data on binding sites of Nef to host proteins such as MAPK1, VAV1, LCK, HCK, HLA-A, CD4, FYN, and GNB2L1 with high statistical significance but is a poor predictor of Nef binding sites on highly flexible, hoop-like regions. Predicted hotspots recapture CD8 cell epitopes of HIV Nef highlighting their importance in modulating virus-host interactions. Host proteins potentially targeted or outcompeted by Nef appear crowding the T cell receptor, natural killer cell mediated cytotoxicity, and neurotrophin signaling pathways. Scanning of HIV Nef motifs on multiple alignments of hepatitis C protein NS5A produces results consistent with literature, indicating the potential value of the hotspot discovery in advancing our understanding of virus-host crosstalk.

  16. Sequence-based prediction of protein protein interaction using a deep-learning algorithm.

    Science.gov (United States)

    Sun, Tanlin; Zhou, Bo; Lai, Luhua; Pei, Jianfeng

    2017-05-25

    Protein-protein interactions (PPIs) are critical for many biological processes. It is therefore important to develop accurate high-throughput methods for identifying PPI to better understand protein function, disease occurrence, and therapy design. Though various computational methods for predicting PPI have been developed, their robustness for prediction with external datasets is unknown. Deep-learning algorithms have achieved successful results in diverse areas, but their effectiveness for PPI prediction has not been tested. We used a stacked autoencoder, a type of deep-learning algorithm, to study the sequence-based PPI prediction. The best model achieved an average accuracy of 97.19% with 10-fold cross-validation. The prediction accuracies for various external datasets ranged from 87.99% to 99.21%, which are superior to those achieved with previous methods. To our knowledge, this research is the first to apply a deep-learning algorithm to sequence-based PPI prediction, and the results demonstrate its potential in this field.

  17. Prediction of protein hydration sites from sequence by modular neural networks

    DEFF Research Database (Denmark)

    Ehrlich, L.; Reczko, M.; Bohr, Henrik

    1998-01-01

    The hydration properties of a protein are important determinants of its structure and function. Here, modular neural networks are employed to predict ordered hydration sites using protein sequence information. First, secondary structure and solvent accessibility are predicted from sequence with two...... separate neural networks. These predictions are used as input together with protein sequences for networks predicting hydration of residues, backbone atoms and sidechains. These networks are teined with protein crystal structures. The prediction of hydration is improved by adding information on secondary...... structure and solvent accessibility and, using actual values of these properties, redidue hydration can be predicted to 77% accuracy with a Metthews coefficient of 0.43. However, predicted property data with an accuracy of 60-70% result in less than half the improvement in predictive performance observed...

  18. Sequence polymorphism in an insect RNA virus field population: A snapshot from a single point in space and time reveals stochastic differences among and within individual hosts

    Energy Technology Data Exchange (ETDEWEB)

    Stenger, Drake C., E-mail: drake.stenger@ars.usda.gov [USDA, Agricultural Research Service, San Joaquin Valley Agricultural Sciences Center, 9611 South Riverbend Ave., Parlier, CA 93648-9757 (United States); Krugner, Rodrigo [USDA, Agricultural Research Service, San Joaquin Valley Agricultural Sciences Center, 9611 South Riverbend Ave., Parlier, CA 93648-9757 (United States); Nouri, Shahideh; Ferriol, Inmaculada; Falk, Bryce W. [Department of Plant Pathology, University of California, Davis, CA 95616 (United States); Sisterson, Mark S. [USDA, Agricultural Research Service, San Joaquin Valley Agricultural Sciences Center, 9611 South Riverbend Ave., Parlier, CA 93648-9757 (United States)

    2016-11-15

    Population structure of Homalodisca coagulata Virus-1 (HoCV-1) among and within field-collected insects sampled from a single point in space and time was examined. Polymorphism in complete consensus sequences among single-insect isolates was dominated by synonymous substitutions. The mutant spectrum of the C2 helicase region within each single-insect isolate was unique and dominated by nonsynonymous singletons. Bootstrapping was used to correct the within-isolate nonsynonymous:synonymous arithmetic ratio (N:S) for RT-PCR error, yielding an N:S value ~one log-unit greater than that of consensus sequences. Probability of all possible single-base substitutions for the C2 region predicted N:S values within 95% confidence limits of the corrected within-isolate N:S when the only constraint imposed was viral polymerase error bias for transitions over transversions. These results indicate that bottlenecks coupled with strong negative/purifying selection drive consensus sequences toward neutral sequence space, and that most polymorphism within single-insect isolates is composed of newly-minted mutations sampled prior to selection. -- Highlights: •Sampling protocol minimized differential selection/history among isolates. •Polymorphism among consensus sequences dominated by negative/purifying selection. •Within-isolate N:S ratio corrected for RT-PCR error by bootstrapping. •Within-isolate mutant spectrum dominated by new mutations yet to undergo selection.

  19. Protein Function Prediction Based on Sequence and Structure Information

    KAUST Repository

    Smaili, Fatima Z.

    2016-01-01

    operate. In this master thesis project, we worked on inferring protein functions based on the primary protein sequence. In the approach we follow, 3D models are first constructed using I-TASSER. Functions are then deduced by structurally matching

  20. Genetic analysis of 430 Chinese Cynodon dactylon accessions using sequence-related amplified polymorphism markers.

    Science.gov (United States)

    Huang, Chunqiong; Liu, Guodao; Bai, Changjun; Wang, Wenqiang

    2014-10-21

    Although Cynodon dactylon (C. dactylon) is widely distributed in China, information on its genetic diversity within the germplasm pool is limited. The objective of this study was to reveal the genetic variation and relationships of 430 C. dactylon accessions collected from 22 Chinese provinces using sequence-related amplified polymorphism (SRAP) markers. Fifteen primer pairs were used to amplify specific C. dactylon genomic sequences. A total of 481 SRAP fragments were generated, with fragment sizes ranging from 260-1800 base pairs (bp). Genetic similarity coefficients (GSC) among the 430 accessions averaged 0.72 and ranged from 0.53-0.96. Cluster analysis conducted by two methods, namely the unweighted pair-group method with arithmetic averages (UPGMA) and principle coordinate analysis (PCoA), separated the accessions into eight distinct groups. Our findings verify that Chinese C. dactylon germplasms have rich genetic diversity, which is an excellent basis for C. dactylon breeding for new cultivars.

  1. In Silico Characterization of Pectate Lyase Protein Sequences from Different Source Organisms

    Directory of Open Access Journals (Sweden)

    Amit Kumar Dubey

    2010-01-01

    Full Text Available A total of 121 protein sequences of pectate lyases were subjected to homology search, multiple sequence alignment, phylogenetic tree construction, and motif analysis. The phylogenetic tree constructed revealed different clusters based on different source organisms representing bacterial, fungal, plant, and nematode pectate lyases. The multiple accessions of bacterial, fungal, nematode, and plant pectate lyase protein sequences were placed closely revealing a sequence level similarity. The multiple sequence alignment of these pectate lyase protein sequences from different source organisms showed conserved regions at different stretches with maximum homology from amino acid residues 439–467, 715–816, and 829–910 which could be used for designing degenerate primers or probes specific for pectate lyases. The motif analysis revealed a conserved Pec_Lyase_C domain uniformly observed in all pectate lyases irrespective of variable sources suggesting its possible role in structural and enzymatic functions.

  2. Sequence alignment reveals possible MAPK docking motifs on HIV proteins.

    Directory of Open Access Journals (Sweden)

    Perry Evans

    Full Text Available Over the course of HIV infection, virus replication is facilitated by the phosphorylation of HIV proteins by human ERK1 and ERK2 mitogen-activated protein kinases (MAPKs. MAPKs are known to phosphorylate their substrates by first binding with them at a docking site. Docking site interactions could be viable drug targets because the sequences guiding them are more specific than phosphorylation consensus sites. In this study we use multiple bioinformatics tools to discover candidate MAPK docking site motifs on HIV proteins known to be phosphorylated by MAPKs, and we discuss the possibility of targeting docking sites with drugs. Using sequence alignments of HIV proteins of different subtypes, we show that MAPK docking patterns previously described for human proteins appear on the HIV matrix, Tat, and Vif proteins in a strain dependent manner, but are absent from HIV Rev and appear on all HIV Nef strains. We revise the regular expressions of previously annotated MAPK docking patterns in order to provide a subtype independent motif that annotates all HIV proteins. One revision is based on a documented human variant of one of the substrate docking motifs, and the other reduces the number of required basic amino acids in the standard docking motifs from two to one. The proposed patterns are shown to be consistent with in silico docking between ERK1 and the HIV matrix protein. The motif usage on HIV proteins is sufficiently different from human proteins in amino acid sequence similarity to allow for HIV specific targeting using small-molecule drugs.

  3. Feature Selection and the Class Imbalance Problem in Predicting Protein Function from Sequence

    NARCIS (Netherlands)

    Al-Shahib, A.; Breitling, R.; Gilbert, D.

    2005-01-01

    Abstract: When the standard approach to predict protein function by sequence homology fails, other alternative methods can be used that require only the amino acid sequence for predicting function. One such approach uses machine learning to predict protein function directly from amino acid sequence

  4. Functional analysis of bipartite begomovirus coat protein promoter sequences

    International Nuclear Information System (INIS)

    Lacatus, Gabriela; Sunter, Garry

    2008-01-01

    We demonstrate that the AL2 gene of Cabbage leaf curl virus (CaLCuV) activates the CP promoter in mesophyll and acts to derepress the promoter in vascular tissue, similar to that observed for Tomato golden mosaic virus (TGMV). Binding studies indicate that sequences mediating repression and activation of the TGMV and CaLCuV CP promoter specifically bind different nuclear factors common to Nicotiana benthamiana, spinach and tomato. However, chromatin immunoprecipitation demonstrates that TGMV AL2 can interact with both sequences independently. Binding of nuclear protein(s) from different crop species to viral sequences conserved in both bipartite and monopartite begomoviruses, including TGMV, CaLCuV, Pepper golden mosaic virus and Tomato yellow leaf curl virus suggests that bipartite begomoviruses bind common host factors to regulate the CP promoter. This is consistent with a model in which AL2 interacts with different components of the cellular transcription machinery that bind viral sequences important for repression and activation of begomovirus CP promoters

  5. Protein sequencing via nanopore based devices: a nanofluidics perspective

    Science.gov (United States)

    Chinappi, Mauro; Cecconi, Fabio

    2018-05-01

    Proteins perform a huge number of central functions in living organisms, thus all the new techniques allowing their precise, fast and accurate characterization at single-molecule level certainly represent a burst in proteomics with important biomedical impact. In this review, we describe the recent progresses in the developing of nanopore based devices for protein sequencing. We start with a critical analysis of the main technical requirements for nanopore protein sequencing, summarizing some ideas and methodologies that have recently appeared in the literature. In the last sections, we focus on the physical modelling of the transport phenomena occurring in nanopore based devices. The multiscale nature of the problem is discussed and, in this respect, some of the main possible computational approaches are illustrated.

  6. Plasminogen activator inhibitor-2 polymorphism associates with recurrent coronary event risk in patients with high HDL and C-reactive protein levels.

    Directory of Open Access Journals (Sweden)

    James P Corsetti

    Full Text Available The objective of this work was to investigate whether fibrinolysis plays a role in establishing recurrent coronary event risk in a previously identified group of postinfarction patients. This group of patients was defined as having concurrently high levels of high-density lipoprotein cholesterol (HDL-C and C-reactive protein (CRP and was previously demonstrated to be at high-risk for recurrent coronary events. Potential risk associations of a genetic polymorphism of plasminogen activator inhibitor-2 (PAI-2 were probed as well as potential modulatory effects on such risk of a polymorphism of low-density lipoprotein receptor related protein (LRP-1, a scavenger receptor known to be involved in fibrinolysis in the context of cellular internalization of plasminogen activator/plansminogen activator inhibitor complexes. To this end, Cox multivariable modeling was performed as a function of genetic polymorphisms of PAI-2 (SERPINB, rs6095 and LRP-1 (LRP1, rs1800156 as well as a set of clinical parameters, blood biomarkers, and genetic polymorphisms previously demonstrated to be significantly and independently associated with risk in the study population including cholesteryl ester transfer protein (CETP, rs708272, p22phox (CYBA, rs4673, and thrombospondin-4 (THBS4, rs1866389. Risk association was demonstrated for the reference allele of the PAI-2 polymorphism (hazard ratio 0.41 per allele, 95% CI 0.20-0.84, p=0.014 along with continued significant risk associations for the p22phox and thrombospondin-4 polymorphisms. Additionally, further analysis revealed interaction of the LRP-1 and PAI-2 polymorphisms in generating differential risk that was illustrated using Kaplan-Meier survival analysis. We conclude from the study that fibrinolysis likely plays a role in establishing recurrent coronary risk in postinfarction patients with concurrently high levels of HDL-C and CRP as manifested by differential effects on risk by polymorphisms of several genes linked

  7. Amino acid and structural variability of Yersinia pestis LcrV protein

    Energy Technology Data Exchange (ETDEWEB)

    Anisimov, A P; Dentovskaya, S V; Panfertsev, E A; Svetoch, T E; Kopylov, P K; Segelke, B W; Zemla, A; Telepnev, M V; Motin, V L

    2009-11-09

    The LcrV protein is a multifunctional virulence factor and protective antigen of the plague bacterium which is generally conserved between the epidemic strains of Yersinia pestis. They investigated the diversity in the LcrV sequences among non-epidemic Y. pestis strains which have a limited virulence in selected animal models and for humans. Sequencing of lcrV genes from ten Y. pestis strains belonging to different phylogenetic groups (subspecies) showed that the LcrV proteins possess four major variable hotspots at positions 18, 72, 273, and 324-326. These major variations, together with other minor substitutions in amino acid sequences, allowed them to classify the LcrV alleles into five sequence types (A-E). They observed that the strains of different Y. pestis subspecies can have the same typ of LcrV, and different types of LcrV can exist within the same natural plague focus. The LcrV polymorphisms were structurally analyzed by comparing the modeled structures of LcrV from all available strains. All changes except one occurred either in flexible regions or on the surface of the protein, but local chemical properties (i.e. those of a hydrophobic, hydrophilic, amphipathic, or charged nature) were conserved across all of the strains. Polymorphisms in flexible and surface regions are likely subject to less selective pressure, and have a limited impact on the structure. In contrast, the substitution of tryptophan at position 113 with either glutamic acid or glycine likely has a serious influence on the regional structure of the protein, and these mutations might have an effect on the function of LcrV. The polymorphisms at positions 18, 72 and 273 were accountable for differences in oligomerization of LcrV. The importance of the latter property in emergence of epidemic strains of Y. pestis during evolution of this pathogen will need to be further investigated.

  8. UET: a database of evolutionarily-predicted functional determinants of protein sequences that cluster as functional sites in protein structures.

    Science.gov (United States)

    Lua, Rhonald C; Wilson, Stephen J; Konecki, Daniel M; Wilkins, Angela D; Venner, Eric; Morgan, Daniel H; Lichtarge, Olivier

    2016-01-04

    The structure and function of proteins underlie most aspects of biology and their mutational perturbations often cause disease. To identify the molecular determinants of function as well as targets for drugs, it is central to characterize the important residues and how they cluster to form functional sites. The Evolutionary Trace (ET) achieves this by ranking the functional and structural importance of the protein sequence positions. ET uses evolutionary distances to estimate functional distances and correlates genotype variations with those in the fitness phenotype. Thus, ET ranks are worse for sequence positions that vary among evolutionarily closer homologs but better for positions that vary mostly among distant homologs. This approach identifies functional determinants, predicts function, guides the mutational redesign of functional and allosteric specificity, and interprets the action of coding sequence variations in proteins, people and populations. Now, the UET database offers pre-computed ET analyses for the protein structure databank, and on-the-fly analysis of any protein sequence. A web interface retrieves ET rankings of sequence positions and maps results to a structure to identify functionally important regions. This UET database integrates several ways of viewing the results on the protein sequence or structure and can be found at http://mammoth.bcm.tmc.edu/uet/. © The Author(s) 2015. Published by Oxford University Press on behalf of Nucleic Acids Research.

  9. The sheep growth hormone gene polymorphism and its effects on milk traits.

    Science.gov (United States)

    Dettori, Maria Luisa; Pazzola, Michele; Pira, Emanuela; Paschino, Pietro; Vacca, Giuseppe Massimo

    2015-05-01

    Growth hormone (GH) is encoded by the GH gene, which may be single copy or duplicate in sheep. The two copies of the sheep GH gene (GH1/GH2-N and GH2-Z) were entirely sequenced in one 106 ewes of Sarda breed, in order to highlight sequence polymorphisms and investigate possible association between genetic variants and milk traits. Milk traits included milk yield, fat, protein, casein and lactose percentage. We evidenced 75 nucleotide changes. Transcription factor binding site prediction revealed two sequences potentially recognised by the pituitary-specific transcription factor POU1FI at the GH1/GH2-N gene, which were lost at the promoter of GH2-Z, which might explain the different tissues of expression of GH1/GH2-N (pituitary) and GH2-Z (placenta). Significant differences in milk traits were observed among genotypes at polymorphic loci only for the GH2-Z gene. Sheep with homozygote genotype ss748770547 CC had higher fat percentage (P < 0.01) than TT. SNP ss748770547 was part of a potential transcription factor binding site for C/EBP alpha (CCAAT/Enhancer Binding Protein), which is involved in the regulation of adipogenesis and adipoblast differentiation. SNP ss748770547, located in the GH2-Z gene 5' flanking region, may be a causal mutation affecting milk fat content. These findings might contribute to the knowledge of the sheep GH locus and might be useful in selection processes in sheep.

  10. Genetic polymorphisms and protein structures in growth hormone, growth hormone receptor, ghrelin, insulin-like growth factor 1 and leptin in Mehraban sheep.

    Science.gov (United States)

    Bahrami, A; Behzadi, Sh; Miraei-Ashtiani, S R; Roh, S-G; Katoh, K

    2013-09-15

    The somatotropic axis, the control system for growth hormone (GH) secretion and its endogenous factors involved in the regulation of metabolism and energy partitioning, has promising potentials for producing economically valuable traits in farm animals. Here we investigated single nucleotide polymorphisms (SNPs) of the genes of factors involved in the somatotropic axis for growth hormone (GH1), growth hormone receptor (GHR), ghrelin (GHRL), insulin-like growth factor 1 (IGF-I) and leptin (LEP), using polymerase chain reaction-single-strand conformation polymorphism (PCR-SSCP) and DNA sequencing methods in 452 individual Mehraban sheep. A nonradioactive method to allow SSCP detection was used for genomic DNA and PCR amplification of six fragments: exons 4 and 5 of GH1; exon 10 of GH receptor (GHR); exon 1 of ghrelin (GHRL); exon 1 of insulin-like growth factor-I (IGF-I), and exon 3 of leptin (LEP). Polymorphisms were detected in five of the six PCR products. Two electrophoretic patterns were detected for GH1 exon 4. Five conformational patterns were detected for GH1 exon 5 and LEP exon 3, and three for IGF-I exon 1. Only GHR and GHRL were monomorphic. Changes in protein structures due to variable SNPs were also analyzed. The results suggest that Mehraban sheep, a major breed that is important for the animal industry in Middle East countries, has high genetic variability, opening interesting prospects for future selection programs and preservation strategies. Copyright © 2013 Elsevier B.V. All rights reserved.

  11. CMsearch: simultaneous exploration of protein sequence space and structure space improves not only protein homology detection but also protein structure prediction

    KAUST Repository

    Cui, Xuefeng; Lu, Zhiwu; Wang, Sheng; Jing-Yan Wang, Jim; Gao, Xin

    2016-01-01

    Motivation: Protein homology detection, a fundamental problem in computational biology, is an indispensable step toward predicting protein structures and understanding protein functions. Despite the advances in recent decades on sequence alignment

  12. Improving pairwise comparison of protein sequences with domain co-occurrence

    Science.gov (United States)

    Gascuel, Olivier

    2018-01-01

    Comparing and aligning protein sequences is an essential task in bioinformatics. More specifically, local alignment tools like BLAST are widely used for identifying conserved protein sub-sequences, which likely correspond to protein domains or functional motifs. However, to limit the number of false positives, these tools are used with stringent sequence-similarity thresholds and hence can miss several hits, especially for species that are phylogenetically distant from reference organisms. A solution to this problem is then to integrate additional contextual information to the procedure. Here, we propose to use domain co-occurrence to increase the sensitivity of pairwise sequence comparisons. Domain co-occurrence is a strong feature of proteins, since most protein domains tend to appear with a limited number of other domains on the same protein. We propose a method to take this information into account in a typical BLAST analysis and to construct new domain families on the basis of these results. We used Plasmodium falciparum as a case study to evaluate our method. The experimental findings showed an increase of 14% of the number of significant BLAST hits and an increase of 25% of the proteome area that can be covered with a domain. Our method identified 2240 new domains for which, in most cases, no model of the Pfam database could be linked. Moreover, our study of the quality of the new domains in terms of alignment and physicochemical properties show that they are close to that of standard Pfam domains. Source code of the proposed approach and supplementary data are available at: https://gite.lirmm.fr/menichelli/pairwise-comparison-with-cooccurrence PMID:29293498

  13. Discriminating Microbial Species Using Protein Sequence Properties and Machine Learning

    NARCIS (Netherlands)

    Shahib, Ali Al-; Gilbert, David; Breitling, Rainer

    2007-01-01

    Much work has been done to identify species-specific proteins in sequenced genomes and hence to determine their function. We assumed that such proteins have specific physico-chemical properties that will discriminate them from proteins in other species. In this paper, we examine the validity of this

  14. Protein secondary structure prediction for a single-sequence using hidden semi-Markov models

    Directory of Open Access Journals (Sweden)

    Borodovsky Mark

    2006-03-01

    Full Text Available Abstract Background The accuracy of protein secondary structure prediction has been improving steadily towards the 88% estimated theoretical limit. There are two types of prediction algorithms: Single-sequence prediction algorithms imply that information about other (homologous proteins is not available, while algorithms of the second type imply that information about homologous proteins is available, and use it intensively. The single-sequence algorithms could make an important contribution to studies of proteins with no detected homologs, however the accuracy of protein secondary structure prediction from a single-sequence is not as high as when the additional evolutionary information is present. Results In this paper, we further refine and extend the hidden semi-Markov model (HSMM initially considered in the BSPSS algorithm. We introduce an improved residue dependency model by considering the patterns of statistically significant amino acid correlation at structural segment borders. We also derive models that specialize on different sections of the dependency structure and incorporate them into HSMM. In addition, we implement an iterative training method to refine estimates of HSMM parameters. The three-state-per-residue accuracy and other accuracy measures of the new method, IPSSP, are shown to be comparable or better than ones for BSPSS as well as for PSIPRED, tested under the single-sequence condition. Conclusions We have shown that new dependency models and training methods bring further improvements to single-sequence protein secondary structure prediction. The results are obtained under cross-validation conditions using a dataset with no pair of sequences having significant sequence similarity. As new sequences are added to the database it is possible to augment the dependency structure and obtain even higher accuracy. Current and future advances should contribute to the improvement of function prediction for orphan proteins inscrutable

  15. A functional CD86 polymorphism associated with asthma and related allergic disorders

    DEFF Research Database (Denmark)

    Corydon, Thomas Juhl; Haagerup, Annette; Jensen, Thomas Gryesten

    2007-01-01

    BACKGROUND: Several studies have documented a substantial genetic component in the aetiology of allergic diseases and a number of atopy susceptibility loci have been suggested. One of these loci is 3q21, at which linkage to multiple atopy phenotypes has been reported. This region harbours the CD8......, and specifically the Ile179Val polymorphism, may be a novel aetiological factor in the development of asthma and related allergic disorders....... gene encoding the costimulatory B7.2 protein. The costimulatory system, consisting of receptor proteins, cytokines and associated factors, activates T cells and regulates the immune response upon allergen challenge. METHODS: We sequenced the CD86 gene in patients with atopy from 10 families that showed...... evidence of linkage to 3q21. Identified polymorphisms were analysed in a subsequent family-based association study of two independent Danish samples, respectively comprising 135 and 100 trios of children with atopy and their parents. Functional analysis of the costimulatory effect on cytokine production...

  16. Adhesive proteins of stalked and acorn barnacles display homology with low sequence similarities.

    Directory of Open Access Journals (Sweden)

    Jaimie-Leigh Jonker

    Full Text Available Barnacle adhesion underwater is an important phenomenon to understand for the prevention of biofouling and potential biotechnological innovations, yet so far, identifying what makes barnacle glue proteins 'sticky' has proved elusive. Examination of a broad range of species within the barnacles may be instructive to identify conserved adhesive domains. We add to extensive information from the acorn barnacles (order Sessilia by providing the first protein analysis of a stalked barnacle adhesive, Lepas anatifera (order Lepadiformes. It was possible to separate the L. anatifera adhesive into at least 10 protein bands using SDS-PAGE. Intense bands were present at approximately 30, 70, 90 and 110 kilodaltons (kDa. Mass spectrometry for protein identification was followed by de novo sequencing which detected 52 peptides of 7-16 amino acids in length. None of the peptides matched published or unpublished transcriptome sequences, but some amino acid sequence similarity was apparent between L. anatifera and closely-related Dosima fascicularis. Antibodies against two acorn barnacle proteins (ab-cp-52k and ab-cp-68k showed cross-reactivity in the adhesive glands of L. anatifera. We also analysed the similarity of adhesive proteins across several barnacle taxa, including Pollicipes pollicipes (a stalked barnacle in the order Scalpelliformes. Sequence alignment of published expressed sequence tags clearly indicated that P. pollicipes possesses homologues for the 19 kDa and 100 kDa proteins in acorn barnacles. Homology aside, sequence similarity in amino acid and gene sequences tended to decline as taxonomic distance increased, with minimum similarities of 18-26%, depending on the gene. The results indicate that some adhesive proteins (e.g. 100 kDa are more conserved within barnacles than others (20 kDa.

  17. Relationships between residue Voronoi volume and sequence conservation in proteins.

    Science.gov (United States)

    Liu, Jen-Wei; Cheng, Chih-Wen; Lin, Yu-Feng; Chen, Shao-Yu; Hwang, Jenn-Kang; Yen, Shih-Chung

    2018-02-01

    Functional and biophysical constraints can cause different levels of sequence conservation in proteins. Previously, structural properties, e.g., relative solvent accessibility (RSA) and packing density of the weighted contact number (WCN), have been found to be related to protein sequence conservation (CS). The Voronoi volume has recently been recognized as a new structural property of the local protein structural environment reflecting CS. However, for surface residues, it is sensitive to water molecules surrounding the protein structure. Herein, we present a simple structural determinant termed the relative space of Voronoi volume (RSV); it uses the Voronoi volume and the van der Waals volume of particular residues to quantify the local structural environment. RSV (range, 0-1) is defined as (Voronoi volume-van der Waals volume)/Voronoi volume of the target residue. The concept of RSV describes the extent of available space for every protein residue. RSV and Voronoi profiles with and without water molecules (RSVw, RSV, VOw, and VO) were compared for 554 non-homologous proteins. RSV (without water) showed better Pearson's correlations with CS than did RSVw, VO, or VOw values. The mean correlation coefficient between RSV and CS was 0.51, which is comparable to the correlation between RSA and CS (0.49) and that between WCN and CS (0.56). RSV is a robust structural descriptor with and without water molecules and can quantitatively reflect evolutionary information in a single protein structure. Therefore, it may represent a practical structural determinant to study protein sequence, structure, and function relationships. Copyright © 2017 Elsevier B.V. All rights reserved.

  18. Simple sequence repeat marker development from bacterial artificial chromosome end sequences and expressed sequence tags of flax (Linum usitatissimum L.).

    Science.gov (United States)

    Cloutier, Sylvie; Miranda, Evelyn; Ward, Kerry; Radovanovic, Natasa; Reimer, Elsa; Walichnowski, Andrzej; Datla, Raju; Rowland, Gordon; Duguid, Scott; Ragupathy, Raja

    2012-08-01

    Flax is an important oilseed crop in North America and is mostly grown as a fibre crop in Europe. As a self-pollinated diploid with a small estimated genome size of ~370 Mb, flax is well suited for fast progress in genomics. In the last few years, important genetic resources have been developed for this crop. Here, we describe the assessment and comparative analyses of 1,506 putative simple sequence repeats (SSRs) of which, 1,164 were derived from BAC-end sequences (BESs) and 342 from expressed sequence tags (ESTs). The SSRs were assessed on a panel of 16 flax accessions with 673 (58 %) and 145 (42 %) primer pairs being polymorphic in the BESs and ESTs, respectively. With 818 novel polymorphic SSR primer pairs reported in this study, the repertoire of available SSRs in flax has more than doubled from the combined total of 508 of all previous reports. Among nucleotide motifs, trinucleotides were the most abundant irrespective of the class, but dinucleotides were the most polymorphic. SSR length was also positively correlated with polymorphism. Two dinucleotide (AT/TA and AG/GA) and two trinucleotide (AAT/ATA/TAA and GAA/AGA/AAG) motifs and their iterations, different from those reported in many other crops, accounted for more than half of all the SSRs and were also more polymorphic (63.4 %) than the rest of the markers (42.7 %). This improved resource promises to be useful in genetic, quantitative trait loci (QTL) and association mapping as well as for anchoring the physical/genetic map with the whole genome shotgun reference sequence of flax.

  19. Understanding gene sequence variation in the context of transcription regulation in yeast.

    Directory of Open Access Journals (Sweden)

    Irit Gat-Viks

    2010-01-01

    Full Text Available DNA sequence polymorphism in a regulatory protein can have a widespread transcriptional effect. Here we present a computational approach for analyzing modules of genes with a common regulation that are affected by specific DNA polymorphisms. We identify such regulatory-linkage modules by integrating genotypic and expression data for individuals in a segregating population with complementary expression data of strains mutated in a variety of regulatory proteins. Our procedure searches simultaneously for groups of co-expressed genes, for their common underlying linkage interval, and for their shared regulatory proteins. We applied the method to a cross between laboratory and wild strains of S. cerevisiae, demonstrating its ability to correctly suggest modules and to outperform extant approaches. Our results suggest that middle sporulation genes are under the control of polymorphism in the sporulation-specific tertiary complex Sum1p/Rfm1p/Hst1p. In another example, our analysis reveals novel inter-relations between Swi3 and two mitochondrial inner membrane proteins underlying variation in a module of aerobic cellular respiration genes. Overall, our findings demonstrate that this approach provides a useful framework for the systematic mapping of quantitative trait loci and their role in gene expression variation.

  20. Frequency of polymorphisms and protein expression of cyclin-dependent kinase inhibitor 1A (CDKN1A in central nervous system tumors

    Directory of Open Access Journals (Sweden)

    Mev Dominguez Valentin

    Full Text Available CONTEXT AND OBJECTIVE: Genetic investigation of central nervous system (CNS tumors provides valuable information about the genes regulating proliferation, differentiation, angiogenesis, migration and apoptosis in the CNS. The aim of our study was to determine the prevalence of genetic polymorphisms (codon 31 and 3' untranslated region, 3'UTR and protein expression of the cyclin-dependent kinase inhibitor 1A (CDKN1A gene in patients with and without CNS tumors. DESIGN AND SETTING: Analytical cross-sectional study with a control group, at the Molecular Biology Laboratory, Pediatric Oncology Department, Hospital das Clínicas de Ribeirão Preto. METHODS: 41 patients with CNS tumors and a control group of 161 subjects without cancer and paires for sex, age and ethnicity were genotyped using polymerase chain reaction-restriction fragment length polymorphism (PCR-RFLP. Protein analysis was performed on 36 patients with CNS tumors, using the Western Blotting technique. RESULTS: The frequencies of the heterozygote (Ser/Arg and polymorphic homozygote (Arg/Arg genotypes of codon 31 in the control subjects were 28.0% and 1.2%, respectively. However, the 3'UTR site presented frequencies of 24.2% (C/T and 0.6% (T/T. These frequencies were not statistically different (P > 0.05 from those seen in the patients with CNS tumors (19.4% and 0.0%, codon 31; 15.8% and 2.6%, 3'UTR site. Regarding the protein expression in ependymomas, 66.67% did not express the protein CDKN1A. The results for medulloblastomas and astrocytomas were similar: neither of them expressed the protein (57.14% and 61.54%, respectively. CONCLUSION: No significant differences in protein expression patterns or polymorphisms of CDKN1A in relation to the three types of CNS tumors were observed among Brazilian subjects.

  1. Efficient Feature Selection and Classification of Protein Sequence Data in Bioinformatics

    Science.gov (United States)

    Faye, Ibrahima; Samir, Brahim Belhaouari; Md Said, Abas

    2014-01-01

    Bioinformatics has been an emerging area of research for the last three decades. The ultimate aims of bioinformatics were to store and manage the biological data, and develop and analyze computational tools to enhance their understanding. The size of data accumulated under various sequencing projects is increasing exponentially, which presents difficulties for the experimental methods. To reduce the gap between newly sequenced protein and proteins with known functions, many computational techniques involving classification and clustering algorithms were proposed in the past. The classification of protein sequences into existing superfamilies is helpful in predicting the structure and function of large amount of newly discovered proteins. The existing classification results are unsatisfactory due to a huge size of features obtained through various feature encoding methods. In this work, a statistical metric-based feature selection technique has been proposed in order to reduce the size of the extracted feature vector. The proposed method of protein classification shows significant improvement in terms of performance measure metrics: accuracy, sensitivity, specificity, recall, F-measure, and so forth. PMID:25045727

  2. Proteomic studies related to genetic determinants of variability in protein concentrations

    NARCIS (Netherlands)

    Horvatovich, Peter; Franke, Lude; Bischoff, Rainer

    2014-01-01

    Genetic variation has multiple effects on the proteome. It may influence the expression level of proteins, modify their sequences through single nucleotide polymorphisms, the occurrence of allelic variants, or alternative splicing (ASP) events. This perspective paper summarizes the major effects of

  3. Characterization of Indian and exotic quality protein maize (QPM ...

    African Journals Online (AJOL)

    Polymorphism analysis and genetic diversity of normal maize and quality protein maize (QPM) inbreds among locally well adapted germplasm is a prerequisite for hybrid maize breeding program. The diversity analyses of 48 maize accessions including Indian and exotic germplasm using 75 simple sequence repeat (SSR) ...

  4. Genome-Wide Single-Nucleotide Polymorphisms Discovery and High-Density Genetic Map Construction in Cauliflower Using Specific-Locus Amplified Fragment Sequencing

    Science.gov (United States)

    Zhao, Zhenqing; Gu, Honghui; Sheng, Xiaoguang; Yu, Huifang; Wang, Jiansheng; Huang, Long; Wang, Dan

    2016-01-01

    Molecular markers and genetic maps play an important role in plant genomics and breeding studies. Cauliflower is an important and distinctive vegetable; however, very few molecular resources have been reported for this species. In this study, a novel, specific-locus amplified fragment (SLAF) sequencing strategy was employed for large-scale single nucleotide polymorphism (SNP) discovery and high-density genetic map construction in a double-haploid, segregating population of cauliflower. A total of 12.47 Gb raw data containing 77.92 M pair-end reads were obtained after processing and 6815 polymorphic SLAFs between the two parents were detected. The average sequencing depths reached 52.66-fold for the female parent and 49.35-fold for the male parent. Subsequently, these polymorphic SLAFs were used to genotype the population and further filtered based on several criteria to construct a genetic linkage map of cauliflower. Finally, 1776 high-quality SLAF markers, including 2741 SNPs, constituted the linkage map with average data integrity of 95.68%. The final map spanned a total genetic length of 890.01 cM with an average marker interval of 0.50 cM, and covered 364.9 Mb of the reference genome. The markers and genetic map developed in this study could provide an important foundation not only for comparative genomics studies within Brassica oleracea species but also for quantitative trait loci identification and molecular breeding of cauliflower. PMID:27047515

  5. C-reactive Protein −717A>G and −286C>T>A Gene Polymorphism and Ischemic Stroke

    OpenAIRE

    Yan Liu; Pei-Liang Geng; Fu-Qin Yan; Tong Chen; Wei Wang; Xu-Dong Tang; Jing-Chen Zheng; Wei-Ping Wu; Zhen-Fu Wang

    2015-01-01

    Background: Inflammation plays a pivotal role in the formation and progression of ischemic stroke. Recently, more and more epidemiological studies have focused on the association between C-reactive protein (CRP) −717A > G and −286C > T > A genetic polymorphisms and ischemic stroke. However, the findings of these researches are not conclusive. Methods: We performed a meta-analysis to determine whether these two polymorphisms are associated with the risk of ischemic stroke. Eligible studies...

  6. The first report of RPSA polymorphisms, also called 37/67 kDa LRP/LR gene, in sporadic Creutzfeldt-Jakob disease (CJD

    Directory of Open Access Journals (Sweden)

    Jeong Byung-Hoon

    2011-08-01

    Full Text Available Abstract Background Although polymorphisms of PRNP, the gene encoding prion protein, are known as a determinant affecting prion disease susceptibility, other genes also influence prion incubation time. This finding offers the opportunity to identify other genetic or environmental factor (s modulating susceptibility to prion disease. Ribosomal protein SA (RPSA, also called 37 kDa laminin receptor precursor (LRP/67 kDa laminin receptor (LR, acts as a receptor for laminin, viruses and prion proteins. The binding/internalization of prion protein is dependent for LRP/LR. Methods To identify other susceptibility genes involved in prion disease, we performed genetic analysis of RPSA. For this case-control study, we included 180 sporadic Creutzfeldt-Jakob disease (CJD patients and 189 healthy Koreans. We investigated genotype and allele frequencies of polymorphism on RPSA by direct sequencing or restriction fragment length polymorphism (RFLP analysis. Results We observed four single nucleotide polymorphisms (SNPs, including -8T>C (rs1803893 in the 5'-untranslated region (UTR of exon 2, 134-32C>T (rs3772138 in the intron, 519G>A (rs2269350 in the intron and 793+58C>T (rs2723 in the intron on the RPSA. The 519G>A (at codon 173 is located in the direct PrP binding site. The genotypes and allele frequencies of the RPSA polymorphisms showed no significant differences between the controls and sporadic CJD patients. Conclusion These results suggest that these RPSA polymorphisms have no direct influence on the susceptibility to sporadic CJD. This was the first genetic association study of the polymorphisms of RPSA gene with sporadic CJD.

  7. Prediction by graph theoretic measures of structural effects in proteins arising from non-synonymous single nucleotide polymorphisms.

    Directory of Open Access Journals (Sweden)

    Tammy M K Cheng

    Full Text Available Recent analyses of human genome sequences have given rise to impressive advances in identifying non-synonymous single nucleotide polymorphisms (nsSNPs. By contrast, the annotation of nsSNPs and their links to diseases are progressing at a much slower pace. Many of the current approaches to analysing disease-associated nsSNPs use primarily sequence and evolutionary information, while structural information is relatively less exploited. In order to explore the potential of such information, we developed a structure-based approach, Bongo (Bonds ON Graph, to predict structural effects of nsSNPs. Bongo considers protein structures as residue-residue interaction networks and applies graph theoretical measures to identify the residues that are critical for maintaining structural stability by assessing the consequences on the interaction network of single point mutations. Our results show that Bongo is able to identify mutations that cause both local and global structural effects, with a remarkably low false positive rate. Application of the Bongo method to the prediction of 506 disease-associated nsSNPs resulted in a performance (positive predictive value, PPV, 78.5% similar to that of PolyPhen (PPV, 77.2% and PANTHER (PPV, 72.2%. As the Bongo method is solely structure-based, our results indicate that the structural changes resulting from nsSNPs are closely associated to their pathological consequences.

  8. Protein Z G79A polymorphism in Turkish pediatriccerebral infarct patients.

    Science.gov (United States)

    Öztürk, Ayşenur; Eğin, Yonca; Deda, Gülhis; Teber, Serap; Akar, Nejat

    2008-09-05

    Protein Z (PZ) plays an enhancer role in coagulation as an anticoagulant. This is the first study in which G79A polymorphism investigated in Turkish paediatric stroke patients. Ninety-one paediatric stroke patients with cerebral ischemia and 70 control subjects were analyzed for PZ G79A and also FVL, PT mutations. PZ 79 'A' allele in homozygous state was found in five patients (5,5%), while it was found only in one control subject (1,4%) and it was seemed as a risk factor for peadiatric ischemia [OR=3,94 (0,44-35,1)]. When patients and controls who had FVL and PT carriers were excluded, AA genotype carried a risk [OR=3,88 (0,41-36,5)]. Also plasma protein Z levels measured in 21 stroke patients and 52 controls. Plasma protein Z levels were not different between stroke patients (500,95 ngmL-1±158,35) and controls (447,34 ngmL-1±165,97). But the plasma levels of protein Z was decreased in patients with AA genotype. Our data showed that carrying 79 AA genotype could be a genetic risk factor for cerebral infarct in peadiatric patients.

  9. Linkage of congenital isolated adrenocorticotropic hormone deficiency to the corticotropin releasing hormone locus using simple sequence repeat polymorphisms

    Energy Technology Data Exchange (ETDEWEB)

    Kyllo, J.H.; Collins, M.M.; Vetter, K.L. [Univ. of Iowa College of Medicine, Iowa City, IA (United States)] [and others

    1996-03-29

    Genetic screening techniques using simple sequence repeat polymorphisms were applied to investigate the molecular nature of congenital isolated adrenocorticotropic hormone (ACTH) deficiency. We hypothesize that this rare cause of hypocortisolism shared by a brother and sister with two unaffected sibs and unaffected parents is inherited as an autosomal recessive single gene mutation. Genes involved in the hypothalamic-pituitary axis controlling cortisol sufficiency were investigated for a causal role in this disorder. Southern blotting showed no detectable mutations of the gene encoding pro-opiomelanocortin (POMC), the ACTH precursor. Other candidate genes subsequently considered were those encoding neuroendocrine convertase-1, and neuroendocrine convertase-2 (NEC-1, NEC-2), and corticotropin releasing hormone (CRH). Tests for linkage were performed using polymorphic di- and tetranucleotide simple sequence repeat markers flanking the reported map locations for POMC, NEC-1, NEC-2, and CRH. The chromosomal haplotypes determined by the markers flanking the loci for POMC, NEC-1, and NEC-2 were not compatible with linkage. However, 22 individual markers defining the chromosomal haplotypes flanking CRH were compatible with linkage of the disorder to the immediate area of this gene of chromosome 8. Based on these data, we hypothesize that the ACTH deficiency in this family is due to an abnormality of CRH gene structure or expression. These results illustrate the useful application of high density genetic maps constructed with simple sequence repeat markers for inclusion/exclusion studies of candidate genes in even very small nuclear families segregating for unusual phenotypes. 25 refs., 5 figs., 2 tabs.

  10. Sequence and organization of the rhoptry-associated-protein-1 (rap-1) locus for the sheep hemoprotozoan Babesia sp. BQ1 Lintan (B. motasi phylogenetic group).

    Science.gov (United States)

    Niu, Qingli; Bonsergent, Claire; Guan, Guiquan; Yin, Hong; Malandrin, Laurence

    2013-11-15

    Babesiosis is a frequent infection of animals worldwide by tick borne pathogen Babesia, and several species are responsible for ovine babesiosis. Recently, several Babesia motasi-like isolates were described in sheep in China. In this study, we sequenced the multigenic rap-1 gene locus of one of these isolates, Babesia sp. BQ1 Lintan. The RAP-1 proteins are involved in the process of red blood cells invasion and thus represent a potential target for vaccine development. A complex composition and organization of the rap-1 locus was discovered with: (1) the presence of 3 different types of rap-1 sequences (rap-1a, rap-1b and rap-1c); (2) the presence of multiple copies of rap-1a and rap-1b; (3) polymorphism among the rap-1a copies, with two classes (named rap-1a61 and rap-1a67) having a similarity of 95.7%, each class represented by two close variants; (4) polymorphism between rap-1a61-1 and rap-1a61-2 limited to three nucleotide positions; (5) a difference of eight nucleotides between rap-1a67-1 and rap-1a67-2 from position 1270 to the putative stop site of rap-1a67-1 which might produce two putative proteins of slightly different sizes; (6) the ratio of rap-1a copies corresponding to one rap-1a67, one rap-1a61-1 and one rap-1a61-2; (7) the presence of three different intergenic regions separating rap-1a, rap-1b and rap-1c; (8) interspacing of the rap-1a copies with rap-1b copies; and (9) the terminal position of rap-1c in the locus. A 31kb locus composed of 6 rap-1a sequences interspaced with 5 rap-1b sequences and with a terminal rap-1c copy was hypothesized. A strikingly similar sequence composition (rap-1a, rap-1b and rap-1c), as well as strong gene identities and similar locus organization with B. bigemina were found and highlight the conservation of synteny at this locus in this phylogenetic clade. Copyright © 2013 Elsevier B.V. All rights reserved.

  11. Investigating Correlation between Protein Sequence Similarity and Semantic Similarity Using Gene Ontology Annotations.

    Science.gov (United States)

    Ikram, Najmul; Qadir, Muhammad Abdul; Afzal, Muhammad Tanvir

    2018-01-01

    Sequence similarity is a commonly used measure to compare proteins. With the increasing use of ontologies, semantic (function) similarity is getting importance. The correlation between these measures has been applied in the evaluation of new semantic similarity methods, and in protein function prediction. In this research, we investigate the relationship between the two similarity methods. The results suggest absence of a strong correlation between sequence and semantic similarities. There is a large number of proteins with low sequence similarity and high semantic similarity. We observe that Pearson's correlation coefficient is not sufficient to explain the nature of this relationship. Interestingly, the term semantic similarity values above 0 and below 1 do not seem to play a role in improving the correlation. That is, the correlation coefficient depends only on the number of common GO terms in proteins under comparison, and the semantic similarity measurement method does not influence it. Semantic similarity and sequence similarity have a distinct behavior. These findings are of significant effect for future works on protein comparison, and will help understand the semantic similarity between proteins in a better way.

  12. Thermodynamic description of polymorphism in Q- and N-rich peptide aggregates revealed by atomistic simulation.

    Science.gov (United States)

    Berryman, Joshua T; Radford, Sheena E; Harris, Sarah A

    2009-07-08

    Amyloid fibrils are long, helically symmetric protein aggregates that can display substantial variation (polymorphism), including alterations in twist and structure at the beta-strand and protofilament levels, even when grown under the same experimental conditions. The structural and thermodynamic origins of this behavior are not yet understood. We performed molecular-dynamics simulations to determine the thermodynamic properties of different polymorphs of the peptide GNNQQNY, modeling fibrils containing different numbers of protofilaments based on the structure of amyloid-like cross-beta crystals of this peptide. We also modeled fibrils with new orientations of the side chains, as well as a de novo designed structure based on antiparallel beta-strands. The simulations show that these polymorphs are approximately isoenergetic under a range of conditions. Structural analysis reveals a dynamic reorganization of electrostatics and hydrogen bonding in the main and side chains of the Gln and Asn residues that characterize this peptide sequence. Q/N-rich stretches are found in several amyloidogenic proteins and peptides, including the yeast prions Sup35-N and Ure2p, as well as in the human poly-Q disease proteins, including the ataxins and huntingtin. Based on our results, we propose that these residues imbue a unique structural plasticity to the amyloid fibrils that they comprise, rationalizing the ability of proteins enriched in these amino acids to form prion strains with heritable and different phenotypic traits.

  13. Sequence heterogeneity accelerates protein search for targets on DNA

    International Nuclear Information System (INIS)

    Shvets, Alexey A.; Kolomeisky, Anatoly B.

    2015-01-01

    The process of protein search for specific binding sites on DNA is fundamentally important since it marks the beginning of all major biological processes. We present a theoretical investigation that probes the role of DNA sequence symmetry, heterogeneity, and chemical composition in the protein search dynamics. Using a discrete-state stochastic approach with a first-passage events analysis, which takes into account the most relevant physical-chemical processes, a full analytical description of the search dynamics is obtained. It is found that, contrary to existing views, the protein search is generally faster on DNA with more heterogeneous sequences. In addition, the search dynamics might be affected by the chemical composition near the target site. The physical origins of these phenomena are discussed. Our results suggest that biological processes might be effectively regulated by modifying chemical composition, symmetry, and heterogeneity of a genome

  14. Sequence heterogeneity accelerates protein search for targets on DNA

    Energy Technology Data Exchange (ETDEWEB)

    Shvets, Alexey A.; Kolomeisky, Anatoly B., E-mail: tolya@rice.edu [Department of Chemistry and Center for Theoretical Biological Physics, Rice University, Houston, Texas 77005 (United States)

    2015-12-28

    The process of protein search for specific binding sites on DNA is fundamentally important since it marks the beginning of all major biological processes. We present a theoretical investigation that probes the role of DNA sequence symmetry, heterogeneity, and chemical composition in the protein search dynamics. Using a discrete-state stochastic approach with a first-passage events analysis, which takes into account the most relevant physical-chemical processes, a full analytical description of the search dynamics is obtained. It is found that, contrary to existing views, the protein search is generally faster on DNA with more heterogeneous sequences. In addition, the search dynamics might be affected by the chemical composition near the target site. The physical origins of these phenomena are discussed. Our results suggest that biological processes might be effectively regulated by modifying chemical composition, symmetry, and heterogeneity of a genome.

  15. Residue-Specific Side-Chain Polymorphisms via Particle Belief Propagation.

    Science.gov (United States)

    Ghoraie, Laleh Soltan; Burkowski, Forbes; Li, Shuai Cheng; Zhu, Mu

    2014-01-01

    Protein side chains populate diverse conformational ensembles in crystals. Despite much evidence that there is widespread conformational polymorphism in protein side chains, most of the X-ray crystallography data are modeled by single conformations in the Protein Data Bank. The ability to extract or to predict these conformational polymorphisms is of crucial importance, as it facilitates deeper understanding of protein dynamics and functionality. In this paper, we describe a computational strategy capable of predicting side-chain polymorphisms. Our approach extends a particular class of algorithms for side-chain prediction by modeling the side-chain dihedral angles more appropriately as continuous rather than discrete variables. Employing a new inferential technique known as particle belief propagation, we predict residue-specific distributions that encode information about side-chain polymorphisms. Our predicted polymorphisms are in relatively close agreement with results from a state-of-the-art approach based on X-ray crystallography data, which characterizes the conformational polymorphisms of side chains using electron density information, and has successfully discovered previously unmodeled conformations.

  16. Protein sequences clustering of herpes virus by using Tribe Markov clustering (Tribe-MCL)

    Science.gov (United States)

    Bustamam, A.; Siswantining, T.; Febriyani, N. L.; Novitasari, I. D.; Cahyaningrum, R. D.

    2017-07-01

    The herpes virus can be found anywhere and one of the important characteristics is its ability to cause acute and chronic infection at certain times so as a result of the infection allows severe complications occurred. The herpes virus is composed of DNA containing protein and wrapped by glycoproteins. In this work, the Herpes viruses family is classified and analyzed by clustering their protein-sequence using Tribe Markov Clustering (Tribe-MCL) algorithm. Tribe-MCL is an efficient clustering method based on the theory of Markov chains, to classify protein families from protein sequences using pre-computed sequence similarity information. We implement the Tribe-MCL algorithm using an open source program of R. We select 24 protein sequences of Herpes virus obtained from NCBI database. The dataset consists of three types of glycoprotein B, F, and H. Each type has eight herpes virus that infected humans. Based on our simulation using different inflation factor r=1.5, 2, 3 we find a various number of the clusters results. The greater the inflation factor the greater the number of their clusters. Each protein will grouped together in the same type of protein.

  17. A machine learning approach for the identification of odorant binding proteins from sequence-derived properties

    Directory of Open Access Journals (Sweden)

    Suganthan PN

    2007-09-01

    Full Text Available Abstract Background Odorant binding proteins (OBPs are believed to shuttle odorants from the environment to the underlying odorant receptors, for which they could potentially serve as odorant presenters. Although several sequence based search methods have been exploited for protein family prediction, less effort has been devoted to the prediction of OBPs from sequence data and this area is more challenging due to poor sequence identity between these proteins. Results In this paper, we propose a new algorithm that uses Regularized Least Squares Classifier (RLSC in conjunction with multiple physicochemical properties of amino acids to predict odorant-binding proteins. The algorithm was applied to the dataset derived from Pfam and GenDiS database and we obtained overall prediction accuracy of 97.7% (94.5% and 98.4% for positive and negative classes respectively. Conclusion Our study suggests that RLSC is potentially useful for predicting the odorant binding proteins from sequence-derived properties irrespective of sequence similarity. Our method predicts 92.8% of 56 odorant binding proteins non-homologous to any protein in the swissprot database and 97.1% of the 414 independent dataset proteins, suggesting the usefulness of RLSC method for facilitating the prediction of odorant binding proteins from sequence information.

  18. Two sequence-ready contigs spanning the two copies of a 200-kb duplication on human 21q: partial sequence and polymorphisms.

    Science.gov (United States)

    Potier, M; Dutriaux, A; Orti, R; Groet, J; Gibelin, N; Karadima, G; Lutfalla, G; Lynn, A; Van Broeckhoven, C; Chakravarti, A; Petersen, M; Nizetic, D; Delabar, J; Rossier, J

    1998-08-01

    Physical mapping across a duplication can be a tour de force if the region is larger than the size of a bacterial clone. This was the case of the 170- to 275-kb duplication present on the long arm of chromosome 21 in normal human at 21q11.1 (proximal region) and at 21q22.1 (distal region), which we described previously. We have constructed sequence-ready contigs of the two copies of the duplication of which all the clones are genuine representatives of one copy or the other. This required the identification of four duplicon polymorphisms that are copy-specific and nonallelic variations in the sequence of the STSs. Thirteen STSs were mapped inside the duplicated region and 5 outside but close to the boundaries. Among these STSs 10 were end clones from YACs, PACs, or cosmids, and the average interval between two markers in the duplicated region was 16 kb. Eight PACs and cosmids showing minimal overlaps were selected in both copies of the duplication. Comparative sequence analysis along the duplication showed three single-basepair changes between the two copies over 659 bp sequenced (4 STSs), suggesting that the duplication is recent (less than 4 mya). Two CpG islands were located in the duplication, but no genes were identified after a 36-kb cosmid from the proximal copy of the duplication was sequenced. The homology of this chromosome 21 duplicated region with the pericentromeric regions of chromosomes 13, 2, and 18 suggests that the mechanism involved is probably similar to pericentromeric-directed mechanisms described in interchromosomal duplications. Copyright 1998 Academic Press.

  19. Inter Simple Sequence Repeat DNA (ISSR) Polymorphism Utility in Haploid Nicotiana Alata Irradiated Plants for Finding Markers Associated with Gamma Irradiation and Salinity

    International Nuclear Information System (INIS)

    El-Fiki, A.; Adly, M.; El-Metabteb, G.

    2017-01-01

    Nicotiana alata is an ornamental plant. It is a member of family Solanasea. Tobacco (Nicotiana spp.) is one of the most important commercial crops in the world. Wild Nicotiana species, as a store house of genes for several diseases and pests, in addition to genes for several important phytochemicals and quality traits which are not present in cultivated varieties. Inter simple sequence repeat DNA (ISSR) analysis was used to determine the degree of genetic variation in treated haploid Nicotiana alata plants. Total genomic DNAs from different treated haploid plant lets were amplified using five specific primers. All primers were polymorphic. A total of 209 bands were amplified of which 135 (59.47%) polymorphic across the radiation treatments. Whilst, the level of polymorphism among the salinity treatments were 181 (85.6 %). Whereas, the polymorphism among the combined effects between gamma radiation doses and salinity concentrations were 283 ( 73.95% ). Treatments relationships were estimated through cluster analysis (UPGMA) based on ISSR data

  20. Lactobacillus strain diversity based on partial hsp60 gene sequences and design of PCR-restriction fragment length polymorphism assays for species identification and differentiation.

    Science.gov (United States)

    Blaiotta, Giuseppe; Fusco, Vincenzina; Ercolini, Danilo; Aponte, Maria; Pepe, Olimpia; Villani, Francesco

    2008-01-01

    A phylogenetic tree showing diversities among 116 partial (499-bp) Lactobacillus hsp60 (groEL, encoding a 60-kDa heat shock protein) nucleotide sequences was obtained and compared to those previously described for 16S rRNA and tuf gene sequences. The topology of the tree produced in this study showed a Lactobacillus species distribution similar, but not identical, to those previously reported. However, according to the most recent systematic studies, a clear differentiation of 43 single-species clusters was detected/identified among the sequences analyzed. The slightly higher variability of the hsp60 nucleotide sequences than of the 16S rRNA sequences offers better opportunities to design or develop molecular assays allowing identification and differentiation of either distant or very closely related Lactobacillus species. Therefore, our results suggest that hsp60 can be considered an excellent molecular marker for inferring the taxonomy and phylogeny of members of the genus Lactobacillus and that the chosen primers can be used in a simple PCR procedure allowing the direct sequencing of the hsp60 fragments. Moreover, in this study we performed a computer-aided restriction endonuclease analysis of all 499-bp hsp60 partial sequences and we showed that the PCR-restriction fragment length polymorphism (RFLP) patterns obtainable by using both endonucleases AluI and TacI (in separate reactions) can allow identification and differentiation of all 43 Lactobacillus species considered, with the exception of the pair L. plantarum/L. pentosus. However, the latter species can be differentiated by further analysis with Sau3AI or MseI. The hsp60 PCR-RFLP approach was efficiently applied to identify and to differentiate a total of 110 wild Lactobacillus strains (including closely related species, such as L. casei and L. rhamnosus or L. plantarum and L. pentosus) isolated from cheese and dry-fermented sausages.

  1. Dog Y chromosomal DNA sequence: identification, sequencing and SNP discovery

    Directory of Open Access Journals (Sweden)

    Kirkness Ewen

    2006-10-01

    Full Text Available Abstract Background Population genetic studies of dogs have so far mainly been based on analysis of mitochondrial DNA, describing only the history of female dogs. To get a picture of the male history, as well as a second independent marker, there is a need for studies of biallelic Y-chromosome polymorphisms. However, there are no biallelic polymorphisms reported, and only 3200 bp of non-repetitive dog Y-chromosome sequence deposited in GenBank, necessitating the identification of dog Y chromosome sequence and the search for polymorphisms therein. The genome has been only partially sequenced for one male dog, disallowing mapping of the sequence into specific chromosomes. However, by comparing the male genome sequence to the complete female dog genome sequence, candidate Y-chromosome sequence may be identified by exclusion. Results The male dog genome sequence was analysed by Blast search against the human genome to identify sequences with a best match to the human Y chromosome and to the female dog genome to identify those absent in the female genome. Candidate sequences were then tested for male specificity by PCR of five male and five female dogs. 32 sequences from the male genome, with a total length of 24 kbp, were identified as male specific, based on a match to the human Y chromosome, absence in the female dog genome and male specific PCR results. 14437 bp were then sequenced for 10 male dogs originating from Europe, Southwest Asia, Siberia, East Asia, Africa and America. Nine haplotypes were found, which were defined by 14 substitutions. The genetic distance between the haplotypes indicates that they originate from at least five wolf haplotypes. There was no obvious trend in the geographic distribution of the haplotypes. Conclusion We have identified 24159 bp of dog Y-chromosome sequence to be used for population genetic studies. We sequenced 14437 bp in a worldwide collection of dogs, identifying 14 SNPs for future SNP analyses, and

  2. Polymorphism discovery and allele frequency estimation using high-throughput DNA sequencing of target-enriched pooled DNA samples

    Directory of Open Access Journals (Sweden)

    Mullen Michael P

    2012-01-01

    Full Text Available Abstract Background The central role of the somatotrophic axis in animal post-natal growth, development and fertility is well established. Therefore, the identification of genetic variants affecting quantitative traits within this axis is an attractive goal. However, large sample numbers are a pre-requisite for the identification of genetic variants underlying complex traits and although technologies are improving rapidly, high-throughput sequencing of large numbers of complete individual genomes remains prohibitively expensive. Therefore using a pooled DNA approach coupled with target enrichment and high-throughput sequencing, the aim of this study was to identify polymorphisms and estimate allele frequency differences across 83 candidate genes of the somatotrophic axis, in 150 Holstein-Friesian dairy bulls divided into two groups divergent for genetic merit for fertility. Results In total, 4,135 SNPs and 893 indels were identified during the resequencing of the 83 candidate genes. Nineteen percent (n = 952 of variants were located within 5' and 3' UTRs. Seventy-two percent (n = 3,612 were intronic and 9% (n = 464 were exonic, including 65 indels and 236 SNPs resulting in non-synonymous substitutions (NSS. Significant (P ® MassARRAY. No significant differences (P > 0.1 were observed between the two methods for any of the 43 SNPs across both pools (i.e., 86 tests in total. Conclusions The results of the current study support previous findings of the use of DNA sample pooling and high-throughput sequencing as a viable strategy for polymorphism discovery and allele frequency estimation. Using this approach we have characterised the genetic variation within genes of the somatotrophic axis and related pathways, central to mammalian post-natal growth and development and subsequent lactogenesis and fertility. We have identified a large number of variants segregating at significantly different frequencies between cattle groups divergent for calving

  3. Aligning protein sequence and analysing substitution pattern using ...

    Indian Academy of Sciences (India)

    Prakash

    Aligning protein sequences using a score matrix has became a routine but valuable method in modern biological ..... the amino acids according to their substitution behaviour ...... which may cause great change (e.g. prolonging the helix) in.

  4. Partner-Mediated Polymorphism of an Intrinsically Disordered Protein.

    Science.gov (United States)

    Bignon, Christophe; Troilo, Francesca; Gianni, Stefano; Longhi, Sonia

    2017-11-29

    Intrinsically disordered proteins (IDPs) recognize their partners through molecular recognition elements (MoREs). The MoRE of the C-terminal intrinsically disordered domain of the measles virus nucleoprotein (N TAIL ) is partly pre-configured as an α-helix in the free form and undergoes α-helical folding upon binding to the X domain (XD) of the viral phosphoprotein. Beyond XD, N TAIL also binds the major inducible heat shock protein 70 (hsp70). So far, no structural information is available for the N TAIL /hsp70 complex. Using mutational studies combined with a protein complementation assay based on green fluorescent protein reconstitution, we have investigated both N TAIL /XD and N TAIL /hsp70 interactions. Although the same N TAIL region binds the two partners, the binding mechanisms are different. Hsp70 binding is much more tolerant of MoRE substitutions than XD, and the majority of substitutions lead to an increased N TAIL /hsp70 interaction strength. Furthermore, while an increased and a decreased α-helicity of the MoRE lead to enhanced and reduced interaction strength with XD, respectively, the impact on hsp70 binding is negligible, suggesting that the MoRE does not adopt an α-helical conformation once bound to hsp70. Here, by showing that the α-helical conformation sampled by the free form of the MoRE does not systematically commit it to adopt an α-helical conformation in the bound form, we provide an example of partner-mediated polymorphism of an IDP and of the relative insensitiveness of the bound structure to the pre-recognition state. The present results therefore contribute to shed light on the molecular mechanisms by which IDPs recognize different partners. Copyright © 2017 Elsevier Ltd. All rights reserved.

  5. Chaos game representation of functional protein sequences, and simulation and multifractal analysis of induced measures

    International Nuclear Information System (INIS)

    Zu-Guo, Yu; Qian-Jun, Xiao; Long, Shi; Jun-Wu, Yu; Anh, Vo

    2010-01-01

    Investigating the biological function of proteins is a key aspect of protein studies. Bioinformatic methods become important for studying the biological function of proteins. In this paper, we first give the chaos game representation (CGR) of randomly-linked functional protein sequences, then propose the use of the recurrent iterated function systems (RIFS) in fractal theory to simulate the measure based on their chaos game representations. This method helps to extract some features of functional protein sequences, and furthermore the biological functions of these proteins. Then multifractal analysis of the measures based on the CGRs of randomly-linked functional protein sequences are performed. We find that the CGRs have clear fractal patterns. The numerical results show that the RIFS can simulate the measure based on the CGR very well. The relative standard error and the estimated probability matrix in the RIFS do not depend on the order to link the functional protein sequences. The estimated probability matrices in the RIFS with different biological functions are evidently different. Hence the estimated probability matrices in the RIFS can be used to characterise the difference among linked functional protein sequences with different biological functions. From the values of the D q curves, one sees that these functional protein sequences are not completely random. The D q of all linked functional proteins studied are multifractal-like and sufficiently smooth for the C q (analogous to specific heat) curves to be meaningful. Furthermore, the D q curves of the measure μ based on their CGRs for different orders to link the functional protein sequences are almost identical if q ≥ 0. Finally, the C q curves of all linked functional proteins resemble a classical phase transition at a critical point. (cross-disciplinary physics and related areas of science and technology)

  6. The SBASE protein domain library, release 8.0: a collection of annotated protein sequence segments.

    Science.gov (United States)

    Murvai, J; Vlahovicek, K; Barta, E; Pongor, S

    2001-01-01

    SBASE 8.0 is the eighth release of the SBASE library of protein domain sequences that contains 294 898 annotated structural, functional, ligand-binding and topogenic segments of proteins, cross-referenced to most major sequence databases and sequence pattern collections. The entries are clustered into over 2005 statistically validated domain groups (SBASE-A) and 595 non-validated groups (SBASE-B), provided with several WWW-based search and browsing facilities for online use. A domain-search facility was developed, based on non-parametric pattern recognition methods, including artificial neural networks. SBASE 8.0 is freely available by anonymous 'ftp' file transfer from ftp.icgeb.trieste.it. Automated searching of SBASE can be carried out with the WWW servers http://www.icgeb.trieste.it/sbase/ and http://sbase.abc. hu/sbase/.

  7. The nucleotide sequence of human transition protein 1 cDNA

    Energy Technology Data Exchange (ETDEWEB)

    Luerssen, H; Hoyer-Fender, S; Engel, W [Universitaet Goettingen (West Germany)

    1988-08-11

    The authors have screened a human testis cDNA library with an oligonucleotide of 81 mer prepared according to a part of the published nucleotide sequence of the rat transition protein TP 1. They have isolated a cDNA clone with the length of 441 bp containing the coding region of 162 bp for human transition protein 1. There is about 84% homology in the coding region of the sequence compared to rat. The human cDNA-clone encodes a polypeptide of 54 amino acids of which 7 are different to that of rat.

  8. Rapid evolution of the sequences and gene repertoires of secreted proteins in bacteria.

    Directory of Open Access Journals (Sweden)

    Teresa Nogueira

    Full Text Available Proteins secreted to the extracellular environment or to the periphery of the cell envelope, the secretome, play essential roles in foraging, antagonistic and mutualistic interactions. We hypothesize that arms races, genetic conflicts and varying selective pressures should lead to the rapid change of sequences and gene repertoires of the secretome. The analysis of 42 bacterial pan-genomes shows that secreted, and especially extracellular proteins, are predominantly encoded in the accessory genome, i.e. among genes not ubiquitous within the clade. Genes encoding outer membrane proteins might engage more frequently in intra-chromosomal gene conversion because they are more often in multi-genic families. The gene sequences encoding the secretome evolve faster than the rest of the genome and in particular at non-synonymous positions. Cell wall proteins in Firmicutes evolve particularly fast when compared with outer membrane proteins of Proteobacteria. Virulence factors are over-represented in the secretome, notably in outer membrane proteins, but cell localization explains more of the variance in substitution rates and gene repertoires than sequence homology to known virulence factors. Accordingly, the repertoires and sequences of the genes encoding the secretome change fast in the clades of obligatory and facultative pathogens and also in the clades of mutualists and free-living bacteria. Our study shows that cell localization shapes genome evolution. In agreement with our hypothesis, the repertoires and the sequences of genes encoding secreted proteins evolve fast. The particularly rapid change of extracellular proteins suggests that these public goods are key players in bacterial adaptation.

  9. Cloning and sequence analysis of cDNA coding for rat nucleolar protein C23

    International Nuclear Information System (INIS)

    Ghaffari, S.H.; Olson, M.O.J.

    1986-01-01

    Using synthetic oligonucleotides as primers and probes, the authors have isolated and sequenced cDNA clones encoding protein C23, a putative nucleolus organizer protein. Poly(A + ) RNA was isolated from rat Novikoff hepatoma cells and enriched in C23 mRNA by sucrose density gradient ultracentrifugation. Two deoxyoligonuleotides, a 48- and a 27-mer, were synthesized on the basis of amino acid sequence from the C-terminal half of protein C23 and cDNA sequence data from CHO cell protein. The 48-mer was used a primer for synthesis of cDNA which was then inserted into plasmid pUC9. Transformed bacterial colonies were screened by hybridization with 32 P labeled 27-mer. Two clones among 5000 gave a strong positive signal. Plasmid DNAs from these clones were purified and characterized by blotting and nucleotide sequence analysis. The length of C23 mRNA was estimated to be 3200 bases in a northern blot analysis. The sequence of a 267 b.p. insert shows high homology with the CHO cDNA with only 9 nucleotide differences and an identical amino acid sequence. These studies indicate that this region of the protein is highly conserved

  10. Integrated analysis of RNA-binding protein complexes using in vitro selection and high-throughput sequencing and sequence specificity landscapes (SEQRS).

    Science.gov (United States)

    Lou, Tzu-Fang; Weidmann, Chase A; Killingsworth, Jordan; Tanaka Hall, Traci M; Goldstrohm, Aaron C; Campbell, Zachary T

    2017-04-15

    RNA-binding proteins (RBPs) collaborate to control virtually every aspect of RNA function. Tremendous progress has been made in the area of global assessment of RBP specificity using next-generation sequencing approaches both in vivo and in vitro. Understanding how protein-protein interactions enable precise combinatorial regulation of RNA remains a significant problem. Addressing this challenge requires tools that can quantitatively determine the specificities of both individual proteins and multimeric complexes in an unbiased and comprehensive way. One approach utilizes in vitro selection, high-throughput sequencing, and sequence-specificity landscapes (SEQRS). We outline a SEQRS experiment focused on obtaining the specificity of a multi-protein complex between Drosophila RBPs Pumilio (Pum) and Nanos (Nos). We discuss the necessary controls in this type of experiment and examine how the resulting data can be complemented with structural and cell-based reporter assays. Additionally, SEQRS data can be integrated with functional genomics data to uncover biological function. Finally, we propose extensions of the technique that will enhance our understanding of multi-protein regulatory complexes assembled onto RNA. Copyright © 2016 Elsevier Inc. All rights reserved.

  11. Next-Generation Sequencing for Binary Protein-Protein Interactions

    Directory of Open Access Journals (Sweden)

    Bernhard eSuter

    2015-12-01

    Full Text Available The yeast two-hybrid (Y2H system exploits host cell genetics in order to display binary protein-protein interactions (PPIs via defined and selectable phenotypes. Numerous improvements have been made to this method, adapting the screening principle for diverse applications, including drug discovery and the scale-up for proteome wide interaction screens in human and other organisms. Here we discuss a systematic workflow and analysis scheme for screening data generated by Y2H and related assays that includes high-throughput selection procedures, readout of comprehensive results via next-generation sequencing (NGS, and the interpretation of interaction data via quantitative statistics. The novel assays and tools will serve the broader scientific community to harness the power of NGS technology to address PPI networks in health and disease. We discuss examples of how this next-generation platform can be applied to address specific questions in diverse fields of biology and medicine.

  12. Impacts of Nonsynonymous Single Nucleotide Polymorphisms of Adiponectin Receptor 1 Gene on Corresponding Protein Stability: A Computational Approach

    Directory of Open Access Journals (Sweden)

    Md. Abu Saleh

    2016-01-01

    Full Text Available Despite the reported association of adiponectin receptor 1 (ADIPOR1 gene mutations with vulnerability to several human metabolic diseases, there is lack of computational analysis on the functional and structural impacts of single nucleotide polymorphisms (SNPs of the human ADIPOR1 at protein level. Therefore, sequence- and structure-based computational tools were employed in this study to functionally and structurally characterize the coding nsSNPs of ADIPOR1 gene listed in the dbSNP database. Our in silico analysis by SIFT, nsSNPAnalyzer, PolyPhen-2, Fathmm, I-Mutant 2.0, SNPs&GO, PhD-SNP, PANTHER, and SNPeffect tools identified the nsSNPs with distorting functional impacts, namely, rs765425383 (A348G, rs752071352 (H341Y, rs759555652 (R324L, rs200326086 (L224F, and rs766267373 (L143P from 74 nsSNPs of ADIPOR1 gene. Finally the aforementioned five deleterious nsSNPs were introduced using Swiss-PDB Viewer package within the X-ray crystal structure of ADIPOR1 protein, and changes in free energy for these mutations were computed. Although increased free energy was observed for all the mutants, the nsSNP H341Y caused the highest energy increase amongst all. RMSD and TM scores predicted that mutants were structurally similar to wild type protein. Our analyses suggested that the aforementioned variants especially H341Y could directly or indirectly destabilize the amino acid interactions and hydrogen bonding networks of ADIPOR1.

  13. Exploring sequence characteristics related to high-level production of secreted proteins in Aspergillus niger.

    Directory of Open Access Journals (Sweden)

    Bastiaan A van den Berg

    Full Text Available Protein sequence features are explored in relation to the production of over-expressed extracellular proteins by fungi. Knowledge on features influencing protein production and secretion could be employed to improve enzyme production levels in industrial bioprocesses via protein engineering. A large set, over 600 homologous and nearly 2,000 heterologous fungal genes, were overexpressed in Aspergillus niger using a standardized expression cassette and scored for high versus no production. Subsequently, sequence-based machine learning techniques were applied for identifying relevant DNA and protein sequence features. The amino-acid composition of the protein sequence was found to be most predictive and interpretation revealed that, for both homologous and heterologous gene expression, the same features are important: tyrosine and asparagine composition was found to have a positive correlation with high-level production, whereas for unsuccessful production, contributions were found for methionine and lysine composition. The predictor is available online at http://bioinformatics.tudelft.nl/hipsec. Subsequent work aims at validating these findings by protein engineering as a method for increasing expression levels per gene copy.

  14. Sequence variability is correlated with weak immunogenicity in Streptococcus pyogenes M protein

    Science.gov (United States)

    Lannergård, Jonas; Kristensen, Bodil M; Gustafsson, Mattias C U; Persson, Jenny J; Norrby-Teglund, Anna; Stålhammar-Carlemalm, Margaretha; Lindahl, Gunnar

    2015-01-01

    The M protein of Streptococcus pyogenes, a major bacterial virulence factor, has an amino-terminal hypervariable region (HVR) that is a target for type-specific protective antibodies. Intriguingly, the HVR elicits a weak antibody response, indicating that it escapes host immunity by two mechanisms, sequence variability and weak immunogenicity. However, the properties influencing the immunogenicity of regions in an M protein remain poorly understood. Here, we studied the antibody response to different regions of the classical M1 and M5 proteins, in which not only the HVR but also the adjacent fibrinogen-binding B repeat region exhibits extensive sequence divergence. Analysis of antisera from S. pyogenes-infected patients, infected mice, and immunized mice showed that both the HVR and the B repeat region elicited weak antibody responses, while the conserved carboxy-terminal part was immunodominant. Thus, we identified a correlation between sequence variability and weak immunogenicity for M protein regions. A potential explanation for the weak immunogenicity was provided by the demonstration that protease digestion selectively eliminated the HVR-B part from whole M protein-expressing bacteria. These data support a coherent model, in which the entire variable HVR-B part evades antibody attack, not only by sequence variability but also by weak immunogenicity resulting from protease attack. PMID:26175306

  15. Comparative analysis of DNA methylation polymorphism in drought sensitive (HPKC2) and tolerant (HPK4) genotypes of horse Gram (Macrotyloma uniflorum).

    Science.gov (United States)

    Bhardwaj, Jyoti; Mahajan, Monika; Yadav, Sudesh Kumar

    2013-08-01

    DNA methylation is known as an epigenetic modification that affects gene expression in plants. Variation in CpG methylation behavior was studied in two natural horse gram (Macrotyloma uniflorum [Lam.] Verdc.) genotypes, HPKC2 (drought-sensitive) and HPK4 (drought-tolerant). The methylation pattern in both genotypes was studied through methylation-sensitive amplified polymorphism. The results revealed that methylation was higher in HPKC2 (10.1%) than in HPK4 (8.6%). Sequencing demonstrated sequence homology with the DRE binding factor (cbf1), the POZ/BTB protein, and the Ty1-copia retrotransposon among some of the polymorphic fragments showing alteration in methylation behavior. Differences in DNA methylation patterns could explain the differential drought tolerance and the epigenetic signature of these two horse gram genotypes.

  16. JACOP: A simple and robust method for the automated classification of protein sequences with modular architecture

    Directory of Open Access Journals (Sweden)

    Pagni Marco

    2005-08-01

    Full Text Available Abstract Background Whole-genome sequencing projects are rapidly producing an enormous number of new sequences. Consequently almost every family of proteins now contains hundreds of members. It has thus become necessary to develop tools, which classify protein sequences automatically and also quickly and reliably. The difficulty of this task is intimately linked to the mechanism by which protein sequences diverge, i.e. by simultaneous residue substitutions, insertions and/or deletions and whole domain reorganisations (duplications/swapping/fusion. Results Here we present a novel approach, which is based on random sampling of sub-sequences (probes out of a set of input sequences. The probes are compared to the input sequences, after a normalisation step; the results are used to partition the input sequences into homogeneous groups of proteins. In addition, this method provides information on diagnostic parts of the proteins. The performance of this method is challenged by two data sets. The first one contains the sequences of prokaryotic lyases that could be arranged as a multiple sequence alignment. The second one contains all proteins from Swiss-Prot Release 36 with at least one Src homology 2 (SH2 domain – a classical example for proteins with modular architecture. Conclusion The outcome of our method is robust, highly reproducible as shown using bootstrap and resampling validation procedures. The results are essentially coherent with the biology. This method depends solely on well-established publicly available software and algorithms.

  17. Molecular effects of autoimmune-risk promoter polymorphisms on expression, exon choice, and translational efficiency of interferon regulatory factor 5.

    Science.gov (United States)

    Clark, Daniel N; Lambert, Jared P; Till, Rodney E; Argueta, Lissenya B; Greenhalgh, Kathryn E; Henrie, Brandon; Bills, Trieste; Hawkley, Tyson F; Roznik, Marinya G; Sloan, Jason M; Mayhew, Vera; Woodland, Loc; Nelson, Eric P; Tsai, Meng-Hsuan; Poole, Brian D

    2014-05-01

    The rs2004640 single nucleotide polymorphism and the CGGGG copy-number variant (rs77571059) are promoter polymorphisms within interferon regulatory factor 5 (IRF5). They have been implicated as susceptibility factors for several autoimmune diseases. IRF5 uses alternative promoter splicing, where any of 4 first exons begin the mRNA. The CGGGG indel is in exon 1A's promoter; the rs2004640 allele creates a splicing recognition site, enabling usage of exon 1B. This study aimed at characterizing alterations in IRF5 mRNA due to these polymorphisms. Cells with risk polymorphisms exhibited ~2-fold higher levels of IRF5 mRNA and protein, but demonstrated no change in mRNA stability. Quantitative PCR demonstrated decreased usage of exons 1C and 1D in cell lines with the risk polymorphisms. RNA folding analysis revealed a hairpin in exon 1B; mutational analysis showed that the hairpin shape decreased translation 5-fold. Although translation of mRNA that uses exon 1B is low due to a hairpin, increased IRF5 mRNA levels in individuals with the rs2004640 risk allele lead to higher overall protein expression. In addition, several new splice variants of IRF5 were sequenced. IRF5's promoter polymorphisms alter first exon usage and increase transcription levels. High levels of IRF5 may bias the immune system toward autoimmunity.

  18. Mitochondrial and nuclear sequence polymorphisms reveal geographic structuring in Amazonian populations of Echinococcus vogeli (Cestoda: Taeniidae).

    Science.gov (United States)

    Santos, Guilherme B; Soares, Manoel do C P; de F Brito, Elisabete M; Rodrigues, André L; Siqueira, Nilton G; Gomes-Gouvêa, Michele S; Alves, Max M; Carneiro, Liliane A; Malheiros, Andreza P; Póvoa, Marinete M; Zaha, Arnaldo; Haag, Karen L

    2012-12-01

    To date, nothing is known about the genetic diversity of the Echinococcus neotropical species, Echinococcus vogeli and Echinococcus oligarthrus. Here we used mitochondrial and nuclear DNA sequence polymorphisms to uncover the genetic structure, transmission and history of E. vogeli in the Brazilian Amazon, based on a sample of 38 isolates obtained from human and wild animal hosts. We confirm that the parasite is partially synanthropic and show that its populations are diverse. Furthermore, significant geographical structuring is found, with western and eastern populations being genetically divergent. Copyright © 2012 Australian Society for Parasitology Inc. Published by Elsevier Ltd. All rights reserved.

  19. Sequence charge decoration dictates coil-globule transition in intrinsically disordered proteins.

    Science.gov (United States)

    Firman, Taylor; Ghosh, Kingshuk

    2018-03-28

    We present an analytical theory to compute conformations of heteropolymers-applicable to describe disordered proteins-as a function of temperature and charge sequence. The theory describes coil-globule transition for a given protein sequence when temperature is varied and has been benchmarked against the all-atom Monte Carlo simulation (using CAMPARI) of intrinsically disordered proteins (IDPs). In addition, the model quantitatively shows how subtle alterations of charge placement in the primary sequence-while maintaining the same charge composition-can lead to significant changes in conformation, even as drastic as a coil (swelled above a purely random coil) to globule (collapsed below a random coil) and vice versa. The theory provides insights on how to control (enhance or suppress) these changes by tuning the temperature (or solution condition) and charge decoration. As an application, we predict the distribution of conformations (at room temperature) of all naturally occurring IDPs in the DisProt database and notice significant size variation even among IDPs with a similar composition of positive and negative charges. Based on this, we provide a new diagram-of-states delineating the sequence-conformation relation for proteins in the DisProt database. Next, we study the effect of post-translational modification, e.g., phosphorylation, on IDP conformations. Modifications as little as two-site phosphorylation can significantly alter the size of an IDP with everything else being constant (temperature, salt concentration, etc.). However, not all possible modification sites have the same effect on protein conformations; there are certain "hot spots" that can cause maximal change in conformation. The location of these "hot spots" in the parent sequence can readily be identified by using a sequence charge decoration metric originally introduced by Sawle and Ghosh. The ability of our model to predict conformations (both expanded and collapsed states) of IDPs at a high

  20. Polymorphisms in the ghrelin gene and their associations with milk yield and quality in water buffaloes.

    Science.gov (United States)

    Gil, F M M; de Camargo, G M F; Pablos de Souza, F R; Cardoso, D F; Fonseca, P D S; Zetouni, L; Braz, C U; Aspilcueta-Borquis, R R; Tonhati, H

    2013-05-01

    Ghrelin is a gastrointestinal hormone that acts in releasing growth hormone and influences the body general metabolism. It has been proposed as a candidate gene for traits such as growth, carcass quality, and milk production of livestock because it influences feed intake. In this context, the aim of this study was to verify the existence of polymorphisms in the ghrelin gene and their associations with milk, fat and protein yield, and percentage in water buffaloes (Bubalus bubalis). A group of 240 animals was studied. Five primer pairs were used and 11 single nucleotide polymorphisms (SNP) were found in the ghrelin gene by sequencing. The animals were genotyped for 8 SNP by PCR-RFLP. The SNP g.960G>A and g.778C>T were associated with fat yield and the SNP g.905T>C was associated with fat yield and percentage and protein percentage. These SNP are located in intronic regions of DNA and may be in noncoding RNA sites or affect transcriptional efciency. The ghrelin gene in buffaloes influences milk fat and protein synthesis. The polymorphisms observed can be used as molecular markers to assist selection. Copyright © 2013 American Dairy Science Association. Published by Elsevier Inc. All rights reserved.

  1. Matrix-Gla Protein rs4236 [A/G] gene polymorphism and serum and GCF levels of MGP in patients with subgingival dental calculus.

    Science.gov (United States)

    Doğan, Gülnihal Emrem; Demir, Turgut; Aksoy, Hülya; Sağlam, Ebru; Laloğlu, Esra; Yildirim, Abdulkadir

    2016-10-01

    Matrix-Gla Protein (MGP) is one of the major Gla-containing protein associated with calcification process. It also has a high affinity for Ca 2+ and hydroxyapatite. In this study we aimed to evaluate the MGP rs4236 [A/G] gene polymorphism in association with subgingival dental calculus. Also a possible relationship between MGP gene polymorphism and serum and GCF levels of MGP were examined. MGP rs4236 [A/G] gene polymorphism was investigated in 110 patients with or without subgingival dental calculus, using polymerase chain reaction-restriction fragment length polymorphism (PCR-RFLP) techniques. Additionally, serum and GCF levels of MGP of the patients were compared according to subgingival dental calculus. Comparison of patients with and without subgingival dental calculus showed no statistically significant difference in MGP rs4236 [A/G] gene polymorphism (p=0.368). MGP concentrations in GCF of patients with subgingival dental calculus were statistically higher than those without subgingival dental calculus (p=0.032). However, a significant association was not observed between the genotypes of AA, AG and GG of the MGP rs4236 gene and the serum and GCF concentrations of MGP in subjects. In this study, it was found that MGP rs4236 [A/G] gene polymorphism was not to be associated with subgingival dental calculus. Also, that GCF MGP levels were detected higher in patients with subgingival dental calculus than those without subgingival dental calculus independently of polymorphism, may be the effect of adaptive mechanism to inhibit calculus formation. Copyright © 2016 Elsevier Ltd. All rights reserved.

  2. Human lymphocyte polymorphisms detected by quantitative two-dimensional electrophoresis

    International Nuclear Information System (INIS)

    Goldman, D.; Merril, C.R.

    1983-01-01

    A survey of 186 soluble lymphocyte proteins for genetic polymorphism was carried out utilizing two-dimensional electrophoresis of 14 C-labeled phytohemagglutinin (PHA)-stimulated human lymphocyte proteins. Nineteen of these proteins exhibited positional variation consistent with independent genetic polymorphism in a primary sample of 28 individuals. Each of these polymorphisms was characterized by quantitative gene-dosage dependence insofar as the heterozygous phenotype expressed approximately 50% of each allelic gene product as was seen in homozygotes. Patterns observed were also identical in monozygotic twins, replicate samples, and replicate gels. The three expected phenotypes (two homozygotes and a heterozygote) were observed in each of 10 of these polymorphisms while the remaining nine had one of the homozygous classes absent. The presence of the three phenotypes, the demonstration of gene-dosage dependence, and our own and previous pedigree analysis of certain of these polymorphisms supports the genetic basis of these variants. Based on this data, the frequency of polymorphic loci for man is: P . 19/186 . .102, and the average heterozygosity is .024. This estimate is approximately 1/3 to 1/2 the rate of polymorphism previously estimated for man in other studies using one-dimensional electrophoresis of isozyme loci. The newly described polymorphisms and others which should be detectable in larger protein surveys with two-dimensional electrophoresis hold promise as genetic markers of the human genome for use in gene mapping and pedigree analyses

  3. The polymorphic integumentary mucin B.1 from Xenopus laevis contains the short consensus repeat.

    Science.gov (United States)

    Probst, J C; Hauser, F; Joba, W; Hoffmann, W

    1992-03-25

    The frog integumentary mucin B.1 (FIM-B.1), discovered by molecular cloning, contains a cysteine-rich C-terminal domain which is homologous with von Willebrand factor. With the help of the polymerase chain reaction, we now characterize a contiguous region 5' to the von Willebrand factor domain containing the short consensus repeat typical of many proteins from the complement system. Multiple transcripts have been cloned, which originate from a single animal and differ by a variable number of tandem repeats (rep-33 sequences). These different transcripts probably originate solely from two genes and are generated presumably by alternative splicing of an huge array of functional cassettes. This model is supported by analysis of genomic FIM-B.1 sequences from Xenopus laevis. Here, rep-33 sequences are arranged in an interrupted array of individual units. Additionally, results of Southern analysis revealed genetic polymorphism between different animals which is predicted to be within the tandem repeats. A first investigation of the predicted mucins with the help of a specific antibody against a synthetic peptide determined the molecular mass of FIM-B.1 to greater than 200 kDa. Here again, genetic polymorphism between different animals is detected.

  4. MODexplorer: an integrated tool for exploring protein sequence, structure and function relationships.

    KAUST Repository

    Kosinski, Jan; Barbato, Alessandro; Tramontano, Anna

    2013-01-01

    SUMMARY: MODexplorer is an integrated tool aimed at exploring the sequence, structural and functional diversity in protein families useful in homology modeling and in analyzing protein families in general. It takes as input either the sequence or the structure of a protein and provides alignments with its homologs along with a variety of structural and functional annotations through an interactive interface. The annotations include sequence conservation, similarity scores, ligand-, DNA- and RNA-binding sites, secondary structure, disorder, crystallographic structure resolution and quality scores of models implied by the alignments to the homologs of known structure. MODexplorer can be used to analyze sequence and structural conservation among the structures of similar proteins, to find structures of homologs solved in different conformational state or with different ligands and to transfer functional annotations. Furthermore, if the structure of the query is not known, MODexplorer can be used to select the modeling templates taking all this information into account and to build a comparative model. AVAILABILITY AND IMPLEMENTATION: Freely available on the web at http://modorama.biocomputing.it/modexplorer. Website implemented in HTML and JavaScript with all major browsers supported. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

  5. MODexplorer: an integrated tool for exploring protein sequence, structure and function relationships.

    KAUST Repository

    Kosinski, Jan

    2013-02-08

    SUMMARY: MODexplorer is an integrated tool aimed at exploring the sequence, structural and functional diversity in protein families useful in homology modeling and in analyzing protein families in general. It takes as input either the sequence or the structure of a protein and provides alignments with its homologs along with a variety of structural and functional annotations through an interactive interface. The annotations include sequence conservation, similarity scores, ligand-, DNA- and RNA-binding sites, secondary structure, disorder, crystallographic structure resolution and quality scores of models implied by the alignments to the homologs of known structure. MODexplorer can be used to analyze sequence and structural conservation among the structures of similar proteins, to find structures of homologs solved in different conformational state or with different ligands and to transfer functional annotations. Furthermore, if the structure of the query is not known, MODexplorer can be used to select the modeling templates taking all this information into account and to build a comparative model. AVAILABILITY AND IMPLEMENTATION: Freely available on the web at http://modorama.biocomputing.it/modexplorer. Website implemented in HTML and JavaScript with all major browsers supported. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.

  6. Application of native signal sequences for recombinant proteins secretion in Pichia pastoris

    DEFF Research Database (Denmark)

    Borodina, Irina; Do, Duy Duc; Eriksen, Jens C.

    Background Methylotrophic yeast Pichia pastoris is widely used for recombinant protein production, largely due to its ability to secrete correctly folded heterologous proteins to the fermentation medium. Secretion is usually achieved by cloning the recombinant gene after a leader sequence, where...... alpha‐mating factor (MF) prepropeptide from Saccharomyces cerevisiae is most commonly used. Our aim was to test whether signal peptides from P. pastoris native secreted proteins could be used to direct secretion of recombinant proteins. Results Eleven native signal peptides from P. pastoris were tested...... by optimization of expression of three different proteins in P. pastoris. Conclusions Native signal peptides from P. pastoris can be used to direct secretion of recombinant proteins. A novel USER‐based P. pastoris system allows easy cloning of protein‐coding gene with the promoter and leader sequence of choice....

  7. Prediction of glutathionylation sites in proteins using minimal sequence information and their experimental validation.

    Science.gov (United States)

    Pal, Debojyoti; Sharma, Deepak; Kumar, Mukesh; Sandur, Santosh K

    2016-09-01

    S-glutathionylation of proteins plays an important role in various biological processes and is known to be protective modification during oxidative stress. Since, experimental detection of S-glutathionylation is labor intensive and time consuming, bioinformatics based approach is a viable alternative. Available methods require relatively longer sequence information, which may prevent prediction if sequence information is incomplete. Here, we present a model to predict glutathionylation sites from pentapeptide sequences. It is based upon differential association of amino acids with glutathionylated and non-glutathionylated cysteines from a database of experimentally verified sequences. This data was used to calculate position dependent F-scores, which measure how a particular amino acid at a particular position may affect the likelihood of glutathionylation event. Glutathionylation-score (G-score), indicating propensity of a sequence to undergo glutathionylation, was calculated using position-dependent F-scores for each amino-acid. Cut-off values were used for prediction. Our model returned an accuracy of 58% with Matthew's correlation-coefficient (MCC) value of 0.165. On an independent dataset, our model outperformed the currently available model, in spite of needing much less sequence information. Pentapeptide motifs having high abundance among glutathionylated proteins were identified. A list of potential glutathionylation hotspot sequences were obtained by assigning G-scores and subsequent Protein-BLAST analysis revealed a total of 254 putative glutathionable proteins, a number of which were already known to be glutathionylated. Our model predicted glutathionylation sites in 93.93% of experimentally verified glutathionylated proteins. Outcome of this study may assist in discovering novel glutathionylation sites and finding candidate proteins for glutathionylation.

  8. RSARF: Prediction of residue solvent accessibility from protein sequence using random forest method

    KAUST Repository

    Ganesan, Pugalenthi; Kandaswamy, Krishna Kumar Umar; Chou -, Kuochen; Vivekanandan, Saravanan; Kolatkar, Prasanna R.

    2012-01-01

    Prediction of protein structure from its amino acid sequence is still a challenging problem. The complete physicochemical understanding of protein folding is essential for the accurate structure prediction. Knowledge of residue solvent accessibility gives useful insights into protein structure prediction and function prediction. In this work, we propose a random forest method, RSARF, to predict residue accessible surface area from protein sequence information. The training and testing was performed using 120 proteins containing 22006 residues. For each residue, buried and exposed state was computed using five thresholds (0%, 5%, 10%, 25%, and 50%). The prediction accuracy for 0%, 5%, 10%, 25%, and 50% thresholds are 72.9%, 78.25%, 78.12%, 77.57% and 72.07% respectively. Further, comparison of RSARF with other methods using a benchmark dataset containing 20 proteins shows that our approach is useful for prediction of residue solvent accessibility from protein sequence without using structural information. The RSARF program, datasets and supplementary data are available at http://caps.ncbs.res.in/download/pugal/RSARF/. - See more at: http://www.eurekaselect.com/89216/article#sthash.pwVGFUjq.dpuf

  9. Alternative splicing affects the targeting sequence of peroxisome proteins in Arabidopsis.

    Science.gov (United States)

    An, Chuanjing; Gao, Yuefang; Li, Jinyu; Liu, Xiaomin; Gao, Fuli; Gao, Hongbo

    2017-07-01

    A systematic analysis of the Arabidopsis genome in combination with localization experiments indicates that alternative splicing affects the peroxisomal targeting sequence of at least 71 genes in Arabidopsis. Peroxisomes are ubiquitous eukaryotic cellular organelles that play a key role in diverse metabolic functions. All peroxisome proteins are encoded by nuclear genes and target to peroxisomes mainly through two types of targeting signals: peroxisomal targeting signal type 1 (PTS1) and PTS2. Alternative splicing (AS) is a process occurring in all eukaryotes by which a single pre-mRNA can generate multiple mRNA variants, often encoding proteins with functional differences. However, the effects of AS on the PTS1 or PTS2 and the targeting of the protein were rarely studied, especially in plants. Here, we systematically analyzed the genome of Arabidopsis, and found that the C-terminal targeting sequence PTS1 of 66 genes and the N-terminal targeting sequence PTS2 of 5 genes are affected by AS. Experimental determination of the targeting of selected protein isoforms further demonstrated that AS at both the 5' and 3' region of a gene can affect the inclusion of PTS2 and PTS1, respectively. This work underscores the importance of AS on the global regulation of peroxisome protein targeting.

  10. Always look on both sides: phylogenetic information conveyed by simple sequence repeat allele sequences.

    Directory of Open Access Journals (Sweden)

    Stéphanie Barthe

    Full Text Available Simple sequence repeat (SSR markers are widely used tools for inferences about genetic diversity, phylogeography and spatial genetic structure. Their applications assume that variation among alleles is essentially caused by an expansion or contraction of the number of repeats and that, accessorily, mutations in the target sequences follow the stepwise mutation model (SMM. Generally speaking, PCR amplicon sizes are used as direct indicators of the number of SSR repeats composing an allele with the data analysis either ignoring the extent of allele size differences or assuming that there is a direct correlation between differences in amplicon size and evolutionary distance. However, without precisely knowing the kind and distribution of polymorphism within an allele (SSR and the associated flanking region (FR sequences, it is hard to say what kind of evolutionary message is conveyed by such a synthetic descriptor of polymorphism as DNA amplicon size. In this study, we sequenced several SSR alleles in multiple populations of three divergent tree genera and disentangled the types of polymorphisms contained in each portion of the DNA amplicon containing an SSR. The patterns of diversity provided by amplicon size variation, SSR variation itself, insertions/deletions (indels, and single nucleotide polymorphisms (SNPs observed in the FRs were compared. Amplicon size variation largely reflected SSR repeat number. The amount of variation was as large in FRs as in the SSR itself. The former contributed significantly to the phylogenetic information and sometimes was the main source of differentiation among individuals and populations contained by FR and SSR regions of SSR markers. The presence of mutations occurring at different rates within a marker's sequence offers the opportunity to analyse evolutionary events occurring on various timescales, but at the same time calls for caution in the interpretation of SSR marker data when the distribution of within

  11. Protein sequence annotation in the genome era: the annotation concept of SWISS-PROT+TREMBL.

    Science.gov (United States)

    Apweiler, R; Gateau, A; Contrino, S; Martin, M J; Junker, V; O'Donovan, C; Lang, F; Mitaritonna, N; Kappus, S; Bairoch, A

    1997-01-01

    SWISS-PROT is a curated protein sequence database which strives to provide a high level of annotation, a minimal level of redundancy and high level of integration with other databases. Ongoing genome sequencing projects have dramatically increased the number of protein sequences to be incorporated into SWISS-PROT. Since we do not want to dilute the quality standards of SWISS-PROT by incorporating sequences without proper sequence analysis and annotation, we cannot speed up the incorporation of new incoming data indefinitely. However, as we also want to make the sequences available as fast as possible, we introduced TREMBL (TRanslation of EMBL nucleotide sequence database), a supplement to SWISS-PROT. TREMBL consists of computer-annotated entries in SWISS-PROT format derived from the translation of all coding sequences (CDS) in the EMBL nucleotide sequence database, except for CDS already included in SWISS-PROT. While TREMBL is already of immense value, its computer-generated annotation does not match the quality of SWISS-PROTs. The main difference is in the protein functional information attached to sequences. With this in mind, we are dedicating substantial effort to develop and apply computer methods to enhance the functional information attached to TREMBL entries.

  12. A sequence-based survey of the complex structural organization of tumor genomes

    Energy Technology Data Exchange (ETDEWEB)

    Collins, Colin; Raphael, Benjamin J.; Volik, Stanislav; Yu, Peng; Wu, Chunxiao; Huang, Guiqing; Linardopoulou, Elena V.; Trask, Barbara J.; Waldman, Frederic; Costello, Joseph; Pienta, Kenneth J.; Mills, Gordon B.; Bajsarowicz, Krystyna; Kobayashi, Yasuko; Sridharan, Shivaranjani; Paris, Pamela; Tao, Quanzhou; Aerni, Sarah J.; Brown, Raymond P.; Bashir, Ali; Gray, Joe W.; Cheng, Jan-Fang; de Jong, Pieter; Nefedov, Mikhail; Ried, Thomas; Padilla-Nash, Hesed M.; Collins, Colin C.

    2008-04-03

    The genomes of many epithelial tumors exhibit extensive chromosomal rearrangements. All classes of genome rearrangements can be identified using End Sequencing Profiling (ESP), which relies on paired-end sequencing of cloned tumor genomes. In this study, brain, breast, ovary and prostate tumors along with three breast cancer cell lines were surveyed with ESP yielding the largest available collection of sequence-ready tumor genome breakpoints and providing evidence that some rearrangements may be recurrent. Sequencing and fluorescence in situ hybridization (FISH) confirmed translocations and complex tumor genome structures that include coamplification and packaging of disparate genomic loci with associated molecular heterogeneity. Comparison of the tumor genomes suggests recurrent rearrangements. Some are likely to be novel structural polymorphisms, whereas others may be bona fide somatic rearrangements. A recurrent fusion transcript in breast tumors and a constitutional fusion transcript resulting from a segmental duplication were identified. Analysis of end sequences for single nucleotide polymorphisms (SNPs) revealed candidate somatic mutations and an elevated rate of novel SNPs in an ovarian tumor. These results suggest that the genomes of many epithelial tumors may be far more dynamic and complex than previously appreciated and that genomic fusions including fusion transcripts and proteins may be common, possibly yielding tumor-specific biomarkers and therapeutic targets.

  13. Hapsembler: An Assembler for Highly Polymorphic Genomes

    Science.gov (United States)

    Donmez, Nilgun; Brudno, Michael

    As whole genome sequencing has become a routine biological experiment, algorithms for assembly of whole genome shotgun data has become a topic of extensive research, with a plethora of off-the-shelf methods that can reconstruct the genomes of many organisms. Simultaneously, several recently sequenced genomes exhibit very high polymorphism rates. For these organisms genome assembly remains a challenge as most assemblers are unable to handle highly divergent haplotypes in a single individual. In this paper we describe Hapsembler, an assembler for highly polymorphic genomes, which makes use of paired reads. Our experiments show that Hapsembler produces accurate and contiguous assemblies of highly polymorphic genomes, while performing on par with the leading tools on haploid genomes. Hapsembler is available for download at http://compbio.cs.toronto.edu/hapsembler.

  14. Experimental Rugged Fitness Landscape in Protein Sequence Space

    OpenAIRE

    HAYASHI, Yuuki; 相田, 拓洋; TOYOTA, Hitoshi; 伏見, 譲; URABE, Itaru; YOMO, Tetsuya

    2006-01-01

    The fitness landscape in sequence space determines the process of biomolecular evolution. To plot the fitness landscape of protein function, we carried out in vitro molecular evolution beginning with a defective fd phage carrying a random polypeptide of 139 amino acids in place of the g3p minor coat protein D2 domain, which is essential for phage infection. After 20 cycles of random substitution at sites 12-130 of the initial random polypeptide and selection for infectivity, the selected phag...

  15. A method for partitioning the information contained in a protein sequence between its structure and function.

    Science.gov (United States)

    Possenti, Andrea; Vendruscolo, Michele; Camilloni, Carlo; Tiana, Guido

    2018-05-23

    Proteins employ the information stored in the genetic code and translated into their sequences to carry out well-defined functions in the cellular environment. The possibility to encode for such functions is controlled by the balance between the amount of information supplied by the sequence and that left after that the protein has folded into its structure. We study the amount of information necessary to specify the protein structure, providing an estimate that keeps into account the thermodynamic properties of protein folding. We thus show that the information remaining in the protein sequence after encoding for its structure (the 'information gap') is very close to what needed to encode for its function and interactions. Then, by predicting the information gap directly from the protein sequence, we show that it may be possible to use these insights from information theory to discriminate between ordered and disordered proteins, to identify unknown functions, and to optimize artificially-designed protein sequences. This article is protected by copyright. All rights reserved. © 2018 Wiley Periodicals, Inc.

  16. Sequence variability is correlated with weak immunogenicity in Streptococcus pyogenes M protein.

    Science.gov (United States)

    Lannergård, Jonas; Kristensen, Bodil M; Gustafsson, Mattias C U; Persson, Jenny J; Norrby-Teglund, Anna; Stålhammar-Carlemalm, Margaretha; Lindahl, Gunnar

    2015-10-01

    The M protein of Streptococcus pyogenes, a major bacterial virulence factor, has an amino-terminal hypervariable region (HVR) that is a target for type-specific protective antibodies. Intriguingly, the HVR elicits a weak antibody response, indicating that it escapes host immunity by two mechanisms, sequence variability and weak immunogenicity. However, the properties influencing the immunogenicity of regions in an M protein remain poorly understood. Here, we studied the antibody response to different regions of the classical M1 and M5 proteins, in which not only the HVR but also the adjacent fibrinogen-binding B repeat region exhibits extensive sequence divergence. Analysis of antisera from S. pyogenes-infected patients, infected mice, and immunized mice showed that both the HVR and the B repeat region elicited weak antibody responses, while the conserved carboxy-terminal part was immunodominant. Thus, we identified a correlation between sequence variability and weak immunogenicity for M protein regions. A potential explanation for the weak immunogenicity was provided by the demonstration that protease digestion selectively eliminated the HVR-B part from whole M protein-expressing bacteria. These data support a coherent model, in which the entire variable HVR-B part evades antibody attack, not only by sequence variability but also by weak immunogenicity resulting from protease attack. © 2015 The Authors. MicrobiologyOpen published by John Wiley & Sons Ltd.

  17. Seeing the trees through the forest : sequence-based homo- and heteromeric protein-protein interaction sites prediction using random forest

    NARCIS (Netherlands)

    Hou, Qingzhen; De Geest, Paul F.G.; Vranken, Wim F.; Heringa, Jaap; Feenstra, K. Anton

    2017-01-01

    Motivation: Genome sequencing is producing an ever-increasing amount of associated protein sequences. Few of these sequences have experimentally validated annotations, however, and computational predictions are becoming increasingly successful in producing such annotations. One key challenge remains

  18. Combining protein sequence, structure, and dynamics: A novel approach for functional evolution analysis of PAS domain superfamily.

    Science.gov (United States)

    Dong, Zheng; Zhou, Hongyu; Tao, Peng

    2018-02-01

    PAS domains are widespread in archaea, bacteria, and eukaryota, and play important roles in various functions. In this study, we aim to explore functional evolutionary relationship among proteins in the PAS domain superfamily in view of the sequence-structure-dynamics-function relationship. We collected protein sequences and crystal structure data from RCSB Protein Data Bank of the PAS domain superfamily belonging to three biological functions (nucleotide binding, photoreceptor activity, and transferase activity). Protein sequences were aligned and then used to select sequence-conserved residues and build phylogenetic tree. Three-dimensional structure alignment was also applied to obtain structure-conserved residues. The protein dynamics were analyzed using elastic network model (ENM) and validated by molecular dynamics (MD) simulation. The result showed that the proteins with same function could be grouped by sequence similarity, and proteins in different functional groups displayed statistically significant difference in their vibrational patterns. Interestingly, in all three functional groups, conserved amino acid residues identified by sequence and structure conservation analysis generally have a lower fluctuation than other residues. In addition, the fluctuation of conserved residues in each biological function group was strongly correlated with the corresponding biological function. This research suggested a direct connection in which the protein sequences were related to various functions through structural dynamics. This is a new attempt to delineate functional evolution of proteins using the integrated information of sequence, structure, and dynamics. © 2017 The Protein Society.

  19. Milk protein-gum tragacanth mixed gels: effect of heat-treatment sequence.

    Science.gov (United States)

    Hatami, Masoud; Nejatian, Mohammad; Mohammadifar, Mohammad Amin; Pourmand, Hanieh

    2014-01-30

    The aim of this study was to investigate the role of the heat-treatment sequence of biopolymer mixtures as a formulation parameter on the acid-induced gelation of tri-polymeric systems composed of sodium caseinate (Na-caseinate), whey protein concentrate (WPC), and gum tragacanth (GT). This was studied by applying four sequences of heat treatment: (A) co-heating all three biopolymers; (B) heating the milk-protein dispersion and the GT dispersion separately; (C) heating the dispersion containing Na-caseinate and GT together and heating whey protein alone; and (D) co-heating whey protein with GT and heating Na-caseinate alone. According to small-deformation rheological measurements, the strength of the mixed-gel network decreased in the order: C>B>D>A samples. SEM micrographs show that the network of sample C is much more homogenous, coarse and dense than sample A, while the networks of samples B and D are of intermediate density. The heat-treatment sequence of the biopolymer mixtures as a formulation parameter thus offers an opportunity to control the microstructure and rheological properties of mixed gels. Copyright © 2013 Elsevier Ltd. All rights reserved.

  20. Discovery and Evaluation of Polymorphisms in the and Promoter Regions for Risk of Korean Lung Cancer

    Directory of Open Access Journals (Sweden)

    Jae Sook Sung

    2012-09-01

    Full Text Available AKT is a signal transduction protein that plays a central role in the tumorigenesis. There are 3 mammalian isoforms of this serine/threonine protein kinase-AKT1, AKT2, and AKT3-showing a broad tissue distribution. We first discovered 2 novel polymorphisms (AKT2 -9826 C/G and AKT3 -811 A/G, and we confirmed 6 known polymorphisms (AKT2 -9473 C/T, AKT2 -9151 C/T, AKT2 -9025 C/T, AKT2 -8618G/A, AKT3 -675 A/-, and AKT3 -244 C/T of the AKT2 and AKT3 promoter region in 24 blood samples of Korean lung cancer patients using direct sequencing. To evaluate the role of AKT2 and AKT3 polymorphisms in the risk of Korean lung cancer, genotypes of the AKT2 and AKT3 polymorphisms (AKT2 -9826 C/G, AKT2 -9473 C/T, AKT2 -9151 C/T, AKT2 -9025 C/T, AKT2 -8618G/A, and AKT3 -675 A/- were determined in 360 lung cancer patients and 360 normal controls. Statistical analyses revealed that the genotypes and haplotypes in the AKT2 and AKT3 promoter regions were not significantly associated with the risk of lung cancer in the Korean population. These results suggest that polymorphisms of the AKT2 and AKT3 promoter regions do not contribute to the genetic susceptibility to lung cancer in the Korean population.

  1. A scalable double-barcode sequencing platform for characterization of dynamic protein-protein interactions.

    Science.gov (United States)

    Schlecht, Ulrich; Liu, Zhimin; Blundell, Jamie R; St Onge, Robert P; Levy, Sasha F

    2017-05-25

    Several large-scale efforts have systematically catalogued protein-protein interactions (PPIs) of a cell in a single environment. However, little is known about how the protein interactome changes across environmental perturbations. Current technologies, which assay one PPI at a time, are too low throughput to make it practical to study protein interactome dynamics. Here, we develop a highly parallel protein-protein interaction sequencing (PPiSeq) platform that uses a novel double barcoding system in conjunction with the dihydrofolate reductase protein-fragment complementation assay in Saccharomyces cerevisiae. PPiSeq detects PPIs at a rate that is on par with current assays and, in contrast with current methods, quantitatively scores PPIs with enough accuracy and sensitivity to detect changes across environments. Both PPI scoring and the bulk of strain construction can be performed with cell pools, making the assay scalable and easily reproduced across environments. PPiSeq is therefore a powerful new tool for large-scale investigations of dynamic PPIs.

  2. OPAL: prediction of MoRF regions in intrinsically disordered protein sequences.

    Science.gov (United States)

    Sharma, Ronesh; Raicar, Gaurav; Tsunoda, Tatsuhiko; Patil, Ashwini; Sharma, Alok

    2018-06-01

    Intrinsically disordered proteins lack stable 3-dimensional structure and play a crucial role in performing various biological functions. Key to their biological function are the molecular recognition features (MoRFs) located within long disordered regions. Computationally identifying these MoRFs from disordered protein sequences is a challenging task. In this study, we present a new MoRF predictor, OPAL, to identify MoRFs in disordered protein sequences. OPAL utilizes two independent sources of information computed using different component predictors. The scores are processed and combined using common averaging method. The first score is computed using a component MoRF predictor which utilizes composition and sequence similarity of MoRF and non-MoRF regions to detect MoRFs. The second score is calculated using half-sphere exposure (HSE), solvent accessible surface area (ASA) and backbone angle information of the disordered protein sequence, using information from the amino acid properties of flanks surrounding the MoRFs to distinguish MoRF and non-MoRF residues. OPAL is evaluated using test sets that were previously used to evaluate MoRF predictors, MoRFpred, MoRFchibi and MoRFchibi-web. The results demonstrate that OPAL outperforms all the available MoRF predictors and is the most accurate predictor available for MoRF prediction. It is available at http://www.alok-ai-lab.com/tools/opal/. ashwini@hgc.jp or alok.sharma@griffith.edu.au. Supplementary data are available at Bioinformatics online.

  3. EST-PAC a web package for EST annotation and protein sequence prediction

    Directory of Open Access Journals (Sweden)

    Strahm Yvan

    2006-10-01

    Full Text Available Abstract With the decreasing cost of DNA sequencing technology and the vast diversity of biological resources, researchers increasingly face the basic challenge of annotating a larger number of expressed sequences tags (EST from a variety of species. This typically consists of a series of repetitive tasks, which should be automated and easy to use. The results of these annotation tasks need to be stored and organized in a consistent way. All these operations should be self-installing, platform independent, easy to customize and amenable to using distributed bioinformatics resources available on the Internet. In order to address these issues, we present EST-PAC a web oriented multi-platform software package for expressed sequences tag (EST annotation. EST-PAC provides a solution for the administration of EST and protein sequence annotations accessible through a web interface. Three aspects of EST annotation are automated: 1 searching local or remote biological databases for sequence similarities using Blast services, 2 predicting protein coding sequence from EST data and, 3 annotating predicted protein sequences with functional domain predictions. In practice, EST-PAC integrates the BLASTALL suite, EST-Scan2 and HMMER in a relational database system accessible through a simple web interface. EST-PAC also takes advantage of the relational database to allow consistent storage, powerful queries of results and, management of the annotation process. The system allows users to customize annotation strategies and provides an open-source data-management environment for research and education in bioinformatics.

  4. Identification of sequence-related amplified polymorphism markers linked to the red leaf trait in ornamental kale (Brassica oleracea L. var. acephala).

    Science.gov (United States)

    Wang, Y S; Liu, Z Y; Li, Y F; Zhang, Y; Yang, X F; Feng, H

    2013-04-02

    Artistic diversiform leaf color is an important agronomic trait that affects the market value of ornamental kale. In the present study, genetic analysis showed that a single-dominant gene, Re (red leaf), determines the red leaf trait in ornamental kale. An F2 population consisting of 500 individuals from the cross of a red leaf double-haploid line 'D05' with a white leaf double-haploid line 'D10' was analyzed for the red leaf trait. By combining bulked segregant analysis and sequence-related amplified polymorphism technology, we identified 3 markers linked to the Re/re locus. A genetic map of the Re locus was constructed using these sequence-related amplified polymorphism markers. Two of the markers, Me8Em4 and Me8Em17, were located on one side of Re/re at distances of 2.2 and 6.4 cM, whereas the other marker, Me9Em11, was located on the other side of Re/re at a distance of 3.7 cM. These markers could be helpful for the subsequent cloning of the red trait gene and marker-assisted selection in ornamental kale breeding programs.

  5. DeepGO: predicting protein functions from sequence and interactions using a deep ontology-aware classifier

    KAUST Repository

    Kulmanov, Maxat

    2017-09-27

    Motivation A large number of protein sequences are becoming available through the application of novel high-throughput sequencing technologies. Experimental functional characterization of these proteins is time-consuming and expensive, and is often only done rigorously for few selected model organisms. Computational function prediction approaches have been suggested to fill this gap. The functions of proteins are classified using the Gene Ontology (GO), which contains over 40 000 classes. Additionally, proteins have multiple functions, making function prediction a large-scale, multi-class, multi-label problem. Results We have developed a novel method to predict protein function from sequence. We use deep learning to learn features from protein sequences as well as a cross-species protein–protein interaction network. Our approach specifically outputs information in the structure of the GO and utilizes the dependencies between GO classes as background information to construct a deep learning model. We evaluate our method using the standards established by the Computational Assessment of Function Annotation (CAFA) and demonstrate a significant improvement over baseline methods such as BLAST, in particular for predicting cellular locations.

  6. Gene identification and protein classification in microbial metagenomic sequence data via incremental clustering

    Directory of Open Access Journals (Sweden)

    Li Weizhong

    2008-04-01

    Full Text Available Abstract Background The identification and study of proteins from metagenomic datasets can shed light on the roles and interactions of the source organisms in their communities. However, metagenomic datasets are characterized by the presence of organisms with varying GC composition, codon usage biases etc., and consequently gene identification is challenging. The vast amount of sequence data also requires faster protein family classification tools. Results We present a computational improvement to a sequence clustering approach that we developed previously to identify and classify protein coding genes in large microbial metagenomic datasets. The clustering approach can be used to identify protein coding genes in prokaryotes, viruses, and intron-less eukaryotes. The computational improvement is based on an incremental clustering method that does not require the expensive all-against-all compute that was required by the original approach, while still preserving the remote homology detection capabilities. We present evaluations of the clustering approach in protein-coding gene identification and classification, and also present the results of updating the protein clusters from our previous work with recent genomic and metagenomic sequences. The clustering results are available via CAMERA, (http://camera.calit2.net. Conclusion The clustering paradigm is shown to be a very useful tool in the analysis of microbial metagenomic data. The incremental clustering method is shown to be much faster than the original approach in identifying genes, grouping sequences into existing protein families, and also identifying novel families that have multiple members in a metagenomic dataset. These clusters provide a basis for further studies of protein families.

  7. Amino acid sequences of ribosomal proteins S11 from Bacillus stearothermophilus and S19 from Halobacterium marismortui. Comparison of the ribosomal protein S11 family.

    Science.gov (United States)

    Kimura, M; Kimura, J; Hatakeyama, T

    1988-11-21

    The complete amino acid sequences of ribosomal proteins S11 from the Gram-positive eubacterium Bacillus stearothermophilus and of S19 from the archaebacterium Halobacterium marismortui have been determined. A search for homologous sequences of these proteins revealed that they belong to the ribosomal protein S11 family. Homologous proteins have previously been sequenced from Escherichia coli as well as from chloroplast, yeast and mammalian ribosomes. A pairwise comparison of the amino acid sequences showed that Bacillus protein S11 shares 68% identical residues with S11 from Escherichia coli and a slightly lower homology (52%) with the homologous chloroplast protein. The halophilic protein S19 is more related to the eukaryotic (45-49%) than to the eubacterial counterparts (35%).

  8. Efficient use of unlabeled data for protein sequence classification: a comparative study.

    Science.gov (United States)

    Kuksa, Pavel; Huang, Pai-Hsi; Pavlovic, Vladimir

    2009-04-29

    Recent studies in computational primary protein sequence analysis have leveraged the power of unlabeled data. For example, predictive models based on string kernels trained on sequences known to belong to particular folds or superfamilies, the so-called labeled data set, can attain significantly improved accuracy if this data is supplemented with protein sequences that lack any class tags-the unlabeled data. In this study, we present a principled and biologically motivated computational framework that more effectively exploits the unlabeled data by only using the sequence regions that are more likely to be biologically relevant for better prediction accuracy. As overly-represented sequences in large uncurated databases may bias the estimation of computational models that rely on unlabeled data, we also propose a method to remove this bias and improve performance of the resulting classifiers. Combined with state-of-the-art string kernels, our proposed computational framework achieves very accurate semi-supervised protein remote fold and homology detection on three large unlabeled databases. It outperforms current state-of-the-art methods and exhibits significant reduction in running time. The unlabeled sequences used under the semi-supervised setting resemble the unpolished gemstones; when used as-is, they may carry unnecessary features and hence compromise the classification accuracy but once cut and polished, they improve the accuracy of the classifiers considerably.

  9. TRDistiller: a rapid filter for enrichment of sequence datasets with proteins containing tandem repeats.

    Science.gov (United States)

    Richard, François D; Kajava, Andrey V

    2014-06-01

    The dramatic growth of sequencing data evokes an urgent need to improve bioinformatics tools for large-scale proteome analysis. Over the last two decades, the foremost efforts of computer scientists were devoted to proteins with aperiodic sequences having globular 3D structures. However, a large portion of proteins contain periodic sequences representing arrays of repeats that are directly adjacent to each other (so called tandem repeats or TRs). These proteins frequently fold into elongated fibrous structures carrying different fundamental functions. Algorithms specific to the analysis of these regions are urgently required since the conventional approaches developed for globular domains have had limited success when applied to the TR regions. The protein TRs are frequently not perfect, containing a number of mutations, and some of them cannot be easily identified. To detect such "hidden" repeats several algorithms have been developed. However, the most sensitive among them are time-consuming and, therefore, inappropriate for large scale proteome analysis. To speed up the TR detection we developed a rapid filter that is based on the comparison of composition and order of short strings in the adjacent sequence motifs. Tests show that our filter discards up to 22.5% of proteins which are known to be without TRs while keeping almost all (99.2%) TR-containing sequences. Thus, we are able to decrease the size of the initial sequence dataset enriching it with TR-containing proteins which allows a faster subsequent TR detection by other methods. The program is available upon request. Copyright © 2014 Elsevier Inc. All rights reserved.

  10. Hypertension-Related Gene Polymorphisms of G-Protein-Coupled Receptor Kinase 4 Are Associated with NT-proBNP Concentration in Normotensive Healthy Adults

    Directory of Open Access Journals (Sweden)

    Junichi Yatabe

    2012-01-01

    Full Text Available G protein-coupled receptor kinase 4 (GRK4 with activating polymorphisms desensitize the natriuric renal tubular D1 dopamine receptor, and these GRK4 polymorphisms are strongly associated with salt sensitivity and hypertension. Meanwhile, N-terminal pro-B-type natriuretic peptide (NT-proBNP may be useful in detecting slight volume expansion. However, relations between hypertension-related gene polymorphisms including GRK4 and cardiovascular indices such as NT-proBNP are not clear, especially in healthy subjects. Therefore, various hypertension-related polymorphisms and cardiovascular indices were analyzed in 97 normotensive, healthy Japanese adults. NT-proBNP levels were significantly higher in subjects with two or more GRK4 polymorphic alleles. Other hypertension-related gene polymorphisms, such as those of renin-angiotensin-aldosterone system genes, did not correlate with NT-proBNP. There was no significant association between any of the hypertension-related gene polymorphisms and central systolic blood pressure, cardioankle vascular index, augmentation index, plasma aldosterone concentration, or an oxidative stress marker, urinary 8-OHdG. Normotensive individuals with GRK4 polymorphisms show increased serum NT-proBNP concentration and may be at a greater risk of developing hypertension and cardiovascular disease.

  11. Do prion protein gene polymorphisms induce apoptosis in non

    Indian Academy of Sciences (India)

    To elucidate the relationship between the SNPs and apoptosis, TUNEL assays and active caspase-3 immunodetection techniques in brain sections of the polymorphic samples were performed. The results revealed that TUNEL-positive cells and active caspase-3-positive cells in the turtles with four polymorphisms were ...

  12. Sifting through genomes with iterative-sequence clustering produces a large, phylogenetically diverse protein-family resource.

    Science.gov (United States)

    Sharpton, Thomas J; Jospin, Guillaume; Wu, Dongying; Langille, Morgan G I; Pollard, Katherine S; Eisen, Jonathan A

    2012-10-13

    New computational resources are needed to manage the increasing volume of biological data from genome sequencing projects. One fundamental challenge is the ability to maintain a complete and current catalog of protein diversity. We developed a new approach for the identification of protein families that focuses on the rapid discovery of homologous protein sequences. We implemented fully automated and high-throughput procedures to de novo cluster proteins into families based upon global alignment similarity. Our approach employs an iterative clustering strategy in which homologs of known families are sifted out of the search for new families. The resulting reduction in computational complexity enables us to rapidly identify novel protein families found in new genomes and to perform efficient, automated updates that keep pace with genome sequencing. We refer to protein families identified through this approach as "Sifting Families," or SFams. Our analysis of ~10.5 million protein sequences from 2,928 genomes identified 436,360 SFams, many of which are not represented in other protein family databases. We validated the quality of SFam clustering through statistical as well as network topology-based analyses. We describe the rapid identification of SFams and demonstrate how they can be used to annotate genomes and metagenomes. The SFam database catalogs protein-family quality metrics, multiple sequence alignments, hidden Markov models, and phylogenetic trees. Our source code and database are publicly available and will be subject to frequent updates (http://edhar.genomecenter.ucdavis.edu/sifting_families/).

  13. Extreme sequence divergence but conserved ligand-binding specificity in Streptococcus pyogenes M protein.

    Directory of Open Access Journals (Sweden)

    2006-05-01

    Full Text Available Many pathogenic microorganisms evade host immunity through extensive sequence variability in a protein region targeted by protective antibodies. In spite of the sequence variability, a variable region commonly retains an important ligand-binding function, reflected in the presence of a highly conserved sequence motif. Here, we analyze the limits of sequence divergence in a ligand-binding region by characterizing the hypervariable region (HVR of Streptococcus pyogenes M protein. Our studies were focused on HVRs that bind the human complement regulator C4b-binding protein (C4BP, a ligand that confers phagocytosis resistance. A previous comparison of C4BP-binding HVRs identified residue identities that could be part of a binding motif, but the extended analysis reported here shows that no residue identities remain when additional C4BP-binding HVRs are included. Characterization of the HVR in the M22 protein indicated that two relatively conserved Leu residues are essential for C4BP binding, but these residues are probably core residues in a coiled-coil, implying that they do not directly contribute to binding. In contrast, substitution of either of two relatively conserved Glu residues, predicted to be solvent-exposed, had no effect on C4BP binding, although each of these changes had a major effect on the antigenic properties of the HVR. Together, these findings show that HVRs of M proteins have an extraordinary capacity for sequence divergence and antigenic variability while retaining a specific ligand-binding function.

  14. DeepGO: predicting protein functions from sequence and interactions using a deep ontology-aware classifier

    KAUST Repository

    Kulmanov, Maxat; Khan, Mohammed Asif; Hoehndorf, Robert

    2017-01-01

    A large number of protein sequences are becoming available through the application of novel high-throughput sequencing technologies. Experimental functional characterization of these proteins is time-consuming and expensive, and is often

  15. Development of cleaved amplified polymorphic sequence (CAPS) and high-resolution melting (HRM) markers from the chloroplast genome of Glycyrrhiza species.

    Science.gov (United States)

    Jo, Ick-Hyun; Sung, Jwakyung; Hong, Chi-Eun; Raveendar, Sebastin; Bang, Kyong-Hwan; Chung, Jong-Wook

    2018-05-01

    Licorice ( Glycyrrhiza glabra ) is an important medicinal crop often used as health foods or medicine worldwide. The molecular genetics of licorice is under scarce owing to lack of molecular markers. Here, we have developed cleaved amplified polymorphic sequence (CAPS) and high-resolution melting (HRM) markers based on single nucleotide polymorphisms (SNP) by comparing the chloroplast genomes of two Glycyrrhiza species ( G. glabra and G. lepidota ). The CAPS and HRM markers were tested for diversity analysis with 24 Glycyrrhiza accessions. The restriction profiles generated with CAPS markers classified the accessions (2-4 genotypes) and melting curves (2-3) were obtained from the HRM markers. The number of alleles and major allele frequency were 2-6 and 0.31-0.92, respectively. The genetic distance and polymorphism information content values were 0.16-0.76 and 0.15-0.72, respectively. The phylogenetic relationships among the 24 accessions were estimated using a dendrogram, which classified them into four clades. Except clade III, the remaining three clades included the same species, confirming interspecies genetic correlation. These 18 CAPS and HRM markers might be helpful for genetic diversity assessment and rapid identification of licorice species.

  16. RTA, a candidate G protein-coupled receptor: Cloning, sequencing, and tissue distribution

    International Nuclear Information System (INIS)

    Ross, P.C.; Figler, R.A.; Corjay, M.H.; Barber, C.M.; Adam, N.; Harcus, D.R.; Lynch, K.R.

    1990-01-01

    Genomic and cDNA clones, encoding a protein that is a member of the guanine nucleotide-binding regulatory protein (G protein)-coupled receptor superfamily, were isolated by screening rat genomic and thoracic aorta cDNA libraries with an oligonucleotide encoding a highly conserved region of the M 1 muscarinic acetylcholine receptor. Sequence analyses of these clones showed that they encode a 343-amino acid protein (named RTA). The RTA gene is single copy, as demonstrated by restriction mapping and Southern blotting of genomic clones and rat genomic DNA. RTA RNA sequences are relatively abundant throughout the gut, vas deferens, uterus, and aorta but are only barely detectable (on Northern blots) in liver, kidney, lung, and salivary gland. In the rat brain, RTA sequences are markedly abundant in the cerebellum. TRA is most closely related to the mas oncogene (34% identity), which has been suggested to be a forebrain angiotensin receptor. They conclude that RTA is not an angiotensin receptor; to date, they have been unable to identify its ligand

  17. Sequencing of 50 human exomes reveals adaptation to high altitude

    DEFF Research Database (Denmark)

    Yi, Xin; Liang, Yu; Huerta-Sanchez, Emilia

    2010-01-01

    Residents of the Tibetan Plateau show heritable adaptations to extreme altitude. We sequenced 50 exomes of ethnic Tibetans, encompassing coding sequences of 92% of human genes, with an average coverage of 18x per individual. Genes showing population-specific allele frequency changes, which repres...... in genetic adaptation to high altitude.......Residents of the Tibetan Plateau show heritable adaptations to extreme altitude. We sequenced 50 exomes of ethnic Tibetans, encompassing coding sequences of 92% of human genes, with an average coverage of 18x per individual. Genes showing population-specific allele frequency changes, which...... represent strong candidates for altitude adaptation, were identified. The strongest signal of natural selection came from endothelial Per-Arnt-Sim (PAS) domain protein 1 (EPAS1), a transcription factor involved in response to hypoxia. One single-nucleotide polymorphism (SNP) at EPAS1 shows a 78% frequency...

  18. Discovery of human inversion polymorphisms by comparative analysis of human and chimpanzee DNA sequence assemblies.

    Directory of Open Access Journals (Sweden)

    2005-10-01

    Full Text Available With a draft genome-sequence assembly for the chimpanzee available, it is now possible to perform genome-wide analyses to identify, at a submicroscopic level, structural rearrangements that have occurred between chimpanzees and humans. The goal of this study was to investigate chromosomal regions that are inverted between the chimpanzee and human genomes. Using the net alignments for the builds of the human and chimpanzee genome assemblies, we identified a total of 1,576 putative regions of inverted orientation, covering more than 154 mega-bases of DNA. The DNA segments are distributed throughout the genome and range from 23 base pairs to 62 mega-bases in length. For the 66 inversions more than 25 kilobases (kb in length, 75% were flanked on one or both sides by (often unrelated segmental duplications. Using PCR and fluorescence in situ hybridization we experimentally validated 23 of 27 (85% semi-randomly chosen regions; the largest novel inversion confirmed was 4.3 mega-bases at human Chromosome 7p14. Gorilla was used as an out-group to assign ancestral status to the variants. All experimentally validated inversion regions were then assayed against a panel of human samples and three of the 23 (13% regions were found to be polymorphic in the human genome. These polymorphic inversions include 730 kb (at 7p22, 13 kb (at 7q11, and 1 kb (at 16q24 fragments with a 5%, 30%, and 48% minor allele frequency, respectively. Our results suggest that inversions are an important source of variation in primate genome evolution. The finding of at least three novel inversion polymorphisms in humans indicates this type of structural variation may be a more common feature of our genome than previously realized.

  19. Complete chloroplast genome sequence of a major allogamous forage species, perennial ryegrass (Lolium perenne L.).

    Science.gov (United States)

    Diekmann, Kerstin; Hodkinson, Trevor R; Wolfe, Kenneth H; van den Bekerom, Rob; Dix, Philip J; Barth, Susanne

    2009-06-01

    Lolium perenne L. (perennial ryegrass) is globally one of the most important forage and grassland crops. We sequenced the chloroplast (cp) genome of Lolium perenne cultivar Cashel. The L. perenne cp genome is 135 282 bp with a typical quadripartite structure. It contains genes for 76 unique proteins, 30 tRNAs and four rRNAs. As in other grasses, the genes accD, ycf1 and ycf2 are absent. The genome is of average size within its subfamily Pooideae and of medium size within the Poaceae. Genome size differences are mainly due to length variations in non-coding regions. However, considerable length differences of 1-27 codons in comparison of L. perenne to other Poaceae and 1-68 codons among all Poaceae were also detected. Within the cp genome of this outcrossing cultivar, 10 insertion/deletion polymorphisms and 40 single nucleotide polymorphisms were detected. Two of the polymorphisms involve tiny inversions within hairpin structures. By comparing the genome sequence with RT-PCR products of transcripts for 33 genes, 31 mRNA editing sites were identified, five of them unique to Lolium. The cp genome sequence of L. perenne is available under Accession number AM777385 at the European Molecular Biology Laboratory, National Center for Biotechnology Information and DNA DataBank of Japan.

  20. Sequence-specific capture of protein-DNA complexes for mass spectrometric protein identification.

    Directory of Open Access Journals (Sweden)

    Cheng-Hsien Wu

    Full Text Available The regulation of gene transcription is fundamental to the existence of complex multicellular organisms such as humans. Although it is widely recognized that much of gene regulation is controlled by gene-specific protein-DNA interactions, there presently exists little in the way of tools to identify proteins that interact with the genome at locations of interest. We have developed a novel strategy to address this problem, which we refer to as GENECAPP, for Global ExoNuclease-based Enrichment of Chromatin-Associated Proteins for Proteomics. In this approach, formaldehyde cross-linking is employed to covalently link DNA to its associated proteins; subsequent fragmentation of the DNA, followed by exonuclease digestion, produces a single-stranded region of the DNA that enables sequence-specific hybridization capture of the protein-DNA complex on a solid support. Mass spectrometric (MS analysis of the captured proteins is then used for their identification and/or quantification. We show here the development and optimization of GENECAPP for an in vitro model system, comprised of the murine insulin-like growth factor-binding protein 1 (IGFBP1 promoter region and FoxO1, a member of the forkhead rhabdomyosarcoma (FoxO subfamily of transcription factors, which binds specifically to the IGFBP1 promoter. This novel strategy provides a powerful tool for studies of protein-DNA and protein-protein interactions.

  1. Automated sequence-specific protein NMR assignment using the memetic algorithm MATCH

    International Nuclear Information System (INIS)

    Volk, Jochen; Herrmann, Torsten; Wuethrich, Kurt

    2008-01-01

    MATCH (Memetic Algorithm and Combinatorial Optimization Heuristics) is a new memetic algorithm for automated sequence-specific polypeptide backbone NMR assignment of proteins. MATCH employs local optimization for tracing partial sequence-specific assignments within a global, population-based search environment, where the simultaneous application of local and global optimization heuristics guarantees high efficiency and robustness. MATCH thus makes combined use of the two predominant concepts in use for automated NMR assignment of proteins. Dynamic transition and inherent mutation are new techniques that enable automatic adaptation to variable quality of the experimental input data. The concept of dynamic transition is incorporated in all major building blocks of the algorithm, where it enables switching between local and global optimization heuristics at any time during the assignment process. Inherent mutation restricts the intrinsically required randomness of the evolutionary algorithm to those regions of the conformation space that are compatible with the experimental input data. Using intact and artificially deteriorated APSY-NMR input data of proteins, MATCH performed sequence-specific resonance assignment with high efficiency and robustness

  2. Allelic sequence variations in the hypervariable region of a T-cell receptor β chain: Correlation with restriction fragment length polymorphism in human families and populations

    International Nuclear Information System (INIS)

    Robinson, M.A.

    1989-01-01

    Direct sequence analysis of the human T-cell antigen receptor (TCR) V β1 variable gene identified a single base-pair allelic variation (C/G) located within the coding region. This change results in substitution of a histidine (CAC) for a glutamine (CAG) at position 48 of the TCR β chain, a position predicted to be in the TCR antigen binding site. The V β1 polymorphism was found by DNA sequence analysis of V β1 genes from seven unrelated individuals; V β1 genes were amplified by the polymerase chain reaction, the amplified fragments were cloned into M13 phage vectors, and sequences were determined. To determined the inheritance patterns of the V β1 substitution and to test correlation with V β1 restriction fragment length polymorphism detected with Pvu II and Taq I, allele-specific oligonucleotides were constructed and used to characterize amplified DNA samples. Seventy unrelated individuals and six families were tested for both restriction fragment length polymorphism and for the V β1 substitution. The correlation was also tested using amplified, size-selected, Pvu II- and Taq I-digested DNA samples from heterozygotes. Pvu II allele 1 (61/70) and Taq I allele 1 (66/70) were found to be correlated with the substitution giving rise to a histidine at position 48. Because there are exceptions to the correlation, the use of specific probes to characterize allelic forms of TCR variable genes will provide important tools for studies of basic TCR genetics and disease associations

  3. Sorting of a HaloTag protein that has only a signal peptide sequence into exocrine secretory granules without protein aggregation.

    Science.gov (United States)

    Fujita-Yoshigaki, Junko; Matsuki-Fukushima, Miwako; Yokoyama, Megumi; Katsumata-Kato, Osamu

    2013-11-15

    The mechanism involved in the sorting and accumulation of secretory cargo proteins, such as amylase, into secretory granules of exocrine cells remains to be solved. To clarify that sorting mechanism, we expressed a reporter protein HaloTag fused with partial sequences of salivary amylase protein in primary cultured parotid acinar cells. We found that a HaloTag protein fused with only the signal peptide sequence (Met(1)-Ala(25)) of amylase, termed SS25H, colocalized well with endogenous amylase, which was confirmed by immunofluorescence microscopy. Percoll-density gradient centrifugation of secretory granule fractions shows that the distributions of amylase and SS25H were similar. These results suggest that SS25H is transported to secretory granules and is not discriminated from endogenous amylase by the machinery that functions to remove proteins other than granule cargo from immature granules. Another reporter protein, DsRed2, that has the same signal peptide sequence also colocalized with amylase, suggesting that the sorting to secretory granules is not dependent on a characteristic of the HaloTag protein. Whereas Blue Native PAGE demonstrates that endogenous amylase forms a high-molecular-weight complex, SS25H does not participate in the complex and does not form self-aggregates. Nevertheless, SS25H was released from cells by the addition of a β-adrenergic agonist, isoproterenol, which also induces amylase secretion. These results indicate that addition of the signal peptide sequence, which is necessary for the translocation in the endoplasmic reticulum, is sufficient for the transportation and storage of cargo proteins in secretory granules of exocrine cells.

  4. The N-terminal sequence of ribosomal protein L10 from the archaebacterium Halobacterium marismortui and its relationship to eubacterial protein L6 and other ribosomal proteins.

    Science.gov (United States)

    Dijk, J; van den Broek, R; Nasiulas, G; Beck, A; Reinhardt, R; Wittmann-Liebold, B

    1987-08-01

    The amino-terminal sequence of ribosomal protein L10 from Halobacterium marismortui has been determined up to residue 54, using both a liquid- and a gas-phase sequenator. The two sequences are in good agreement. The protein is clearly homologous to protein HcuL10 from the related strain Halobacterium cutirubrum. Furthermore, a weaker but distinct homology to ribosomal protein L6 from Escherichia coli and Bacillus stearothermophilus can be detected. In addition to 7 identical amino acids in the first 36 residues in all four sequences a number of conservative replacements occurs, of mainly hydrophobic amino acids. In this common region the pattern of conserved amino acids suggests the presence of a beta-alpha fold as it occurs in ribosomal proteins L12 and L30. Furthermore, several potential cases of homology to other ribosomal components of the three ur-kingdoms have been found.

  5. Sequence of a cDNA encoding turtle high mobility group 1 protein.

    Science.gov (United States)

    Zheng, Jifang; Hu, Bi; Wu, Duansheng

    2005-07-01

    In order to understand sequence information about turtle HMG1 gene, a cDNA encoding HMG1 protein of the Chinese soft-shell turtle (Pelodiscus sinensis) was amplified by RT-PCR from kidney total RNA, and was cloned, sequenced and analyzed. The results revealed that the open reading frame (ORF) of turtle HMG1 cDNA is 606 bp long. The ORF codifies 202 amino acid residues, from which two DNA-binding domains and one polyacidic region are derived. The DNA-binding domains share higher amino acid identity with homologues sequences of chicken (96.5%) and mammalian (74%) than homologues sequence of rainbow trout (67%). The polyacidic region shows 84.6% amino acid homology with the equivalent region of chicken HMG1 cDNA. Turtle HMG1 protein contains 3 Cys residues located at completely conserved positions. Conservation in sequence and structure suggests that the functions of turtle HMG1 cDNA may be highly conserved during evolution. To our knowledge, this is the first report of HMG1 cDNA sequence in any reptilian.

  6. Characterisation of Toxoplasma gondii isolates using polymerase chain reaction (PCR) and restriction fragment length polymorphism (RFLP) of the non-coding Toxoplasma gondii (TGR)-gene sequences

    DEFF Research Database (Denmark)

    Høgdall, Estrid; Vuust, Jens; Lind, Peter

    2000-01-01

    of using TGR gene variants as markers to distinguish among T. gondii isolates from different animals and different geographical sources. Based on the band patterns obtained by restriction fragment length polymorphism (RFLP) analysis of the polymerase chain reaction (PCR) amplified TGR sequences, the T...

  7. Typing of canine parvovirus isolates using mini-sequencing based single nucleotide polymorphism analysis.

    Science.gov (United States)

    Naidu, Hariprasad; Subramanian, B Mohana; Chinchkar, Shankar Ramchandra; Sriraman, Rajan; Rana, Samir Kumar; Srinivasan, V A

    2012-05-01

    The antigenic types of canine parvovirus (CPV) are defined based on differences in the amino acids of the major capsid protein VP2. Type specificity is conferred by a limited number of amino acid changes and in particular by few nucleotide substitutions. PCR based methods are not particularly suitable for typing circulating variants which differ in a few specific nucleotide substitutions. Assays for determining SNPs can detect efficiently nucleotide substitutions and can thus be adapted to identify CPV types. In the present study, CPV typing was performed by single nucleotide extension using the mini-sequencing technique. A mini-sequencing signature was established for all the four CPV types (CPV2, 2a, 2b and 2c) and feline panleukopenia virus. The CPV typing using the mini-sequencing reaction was performed for 13 CPV field isolates and the two vaccine strains available in our repository. All the isolates had been typed earlier by full-length sequencing of the VP2 gene. The typing results obtained from mini-sequencing matched completely with that of sequencing. Typing could be achieved with less than 100 copies of standard plasmid DNA constructs or ≤10¹ FAID₅₀ of virus by mini-sequencing technique. The technique was also efficient for detecting multiple types in mixed infections. Copyright © 2012 Elsevier B.V. All rights reserved.

  8. C-reactive Protein -717A>G and -286C>T>A Gene Polymorphism and Ischemic Stroke.

    Science.gov (United States)

    Liu, Yan; Geng, Pei-Liang; Yan, Fu-Qin; Chen, Tong; Wang, Wei; Tang, Xu-Dong; Zheng, Jing-Chen; Wu, Wei-Ping; Wang, Zhen-Fu

    2015-06-20

    Inflammation plays a pivotal role in the formation and progression of ischemic stroke. Recently, more and more epidemiological studies have focused on the association between C-reactive protein (CRP) -717A > G and -286C > T > A genetic polymorphisms and ischemic stroke. However, the findings of these researches are not conclusive. We performed a meta-analysis to determine whether these two polymorphisms are associated with the risk of ischemic stroke. Eligible studies were identified from the database of PubMed, Medline, Embase, Web of Science, CNKI, Weipu, and Wanfang. We calculated odds ratios (ORs) with 95% confidence intervals (CIs) to assess the strength of the association. Four articles were included in our study, including 1926 cases and 2678 controls for -717A > G polymorphism, 652 cases and 1103 controls for -286C > T > A polymorphism. The results of meta-analysis showed that single nucleotide polymorphism (SNP) -717A > G was not significantly associated with the risk of ischemic stroke (GG vs. AA, OR = 1.12, 95% CI = 0.83-1.50, P = 0.207; GG + GA vs. AA, OR = 1.04, 95% CI = 0.93-1.17, P = 0.533; GG vs. GA + AA, OR = 1.10, 95% CI = 0.82-1.47, P = 0.220). Meta-analysis of SNP - 286C > T > A also demonstrated no statistical evidence of a significant association with the risk of ischemic stroke (AA vs. CC, OR = 0.86, 95% CI = 0.59-1.25, P = 0.348; AA vs. CC, OR = 0.92, 95% CI = 0.80-1.06, P = 0.609; AA vs. CC, OR = 0.89, 95% CI = 0.62-1.30, P = 0.374). This meta-analysis demonstrated little evidence to support a role of CRP gene -717A > G, -286C > T > A polymorphisms in ischemic stroke predisposition. However, to draw comprehensive and more reliable conclusions, further larger studies are needed to validate the association between CRP gene polymorphisms and ischemic stroke in various ethnic groups.

  9. Polymorphic sequence-characterized codominant loci in the chestnut blight fungus, Cryphonectria parasitica

    Science.gov (United States)

    J. E. Davis; Thomas L. Kubisiak; M. G. Milgroom

    2005-01-01

    Studies on the population biology of the chestnut blight fungus, Cryphonectria parasitica, have previously been carried out with dominant restriction fragment length polymorphism (RFLP) fingerprinting markers. In this study, we described the development of 11 condominant markers from randomly amplified polymorphic DNAs (RAPDs). RAPD fragments were...

  10. DNA sequence polymorphisms in a panel of eight candidate bovine imprinted genes and their association with performance traits in Irish Holstein-Friesian cattle

    Directory of Open Access Journals (Sweden)

    Mullen Michael P

    2010-10-01

    Full Text Available Abstract Background Studies in mice and humans have shown that imprinted genes, whereby expression from one of the two parentally inherited alleles is attenuated or completely silenced, have a major effect on mammalian growth, metabolism and physiology. More recently, investigations in livestock species indicate that genes subject to this type of epigenetic regulation contribute to, or are associated with, several performance traits, most notably muscle mass and fat deposition. In the present study, a candidate gene approach was adopted to assess 17 validated single nucleotide polymorphisms (SNPs and their association with a range of performance traits in 848 progeny-tested Irish Holstein-Friesian artificial insemination sires. These SNPs are located proximal to, or within, the bovine orthologs of eight genes (CALCR, GRB10, PEG3, PHLDA2, RASGRF1, TSPAN32, ZIM2 and ZNF215 that have been shown to be imprinted in cattle or in at least one other mammalian species (i.e. human/mouse/pig/sheep. Results Heterozygosities for all SNPs analysed ranged from 0.09 to 0.46 and significant deviations from Hardy-Weinberg proportions (P ≤ 0.01 were observed at four loci. Phenotypic associations (P ≤ 0.05 were observed between nine SNPs proximal to, or within, six of the eight analysed genes and a number of performance traits evaluated, including milk protein percentage, somatic cell count, culled cow and progeny carcass weight, angularity, body conditioning score, progeny carcass conformation, body depth, rump angle, rump width, animal stature, calving difficulty, gestation length and calf perinatal mortality. Notably, SNPs within the imprinted paternally expressed gene 3 (PEG3 gene cluster were associated (P ≤ 0.05 with calving, calf performance and fertility traits, while a single SNP in the zinc finger protein 215 gene (ZNF215 was associated with milk protein percentage (P ≤ 0.05, progeny carcass weight (P ≤ 0.05, culled cow carcass weight (P ≤ 0

  11. DNA sequence polymorphisms in a panel of eight candidate bovine imprinted genes and their association with performance traits in Irish Holstein-Friesian cattle

    Science.gov (United States)

    2010-01-01

    Background Studies in mice and humans have shown that imprinted genes, whereby expression from one of the two parentally inherited alleles is attenuated or completely silenced, have a major effect on mammalian growth, metabolism and physiology. More recently, investigations in livestock species indicate that genes subject to this type of epigenetic regulation contribute to, or are associated with, several performance traits, most notably muscle mass and fat deposition. In the present study, a candidate gene approach was adopted to assess 17 validated single nucleotide polymorphisms (SNPs) and their association with a range of performance traits in 848 progeny-tested Irish Holstein-Friesian artificial insemination sires. These SNPs are located proximal to, or within, the bovine orthologs of eight genes (CALCR, GRB10, PEG3, PHLDA2, RASGRF1, TSPAN32, ZIM2 and ZNF215) that have been shown to be imprinted in cattle or in at least one other mammalian species (i.e. human/mouse/pig/sheep). Results Heterozygosities for all SNPs analysed ranged from 0.09 to 0.46 and significant deviations from Hardy-Weinberg proportions (P ≤ 0.01) were observed at four loci. Phenotypic associations (P ≤ 0.05) were observed between nine SNPs proximal to, or within, six of the eight analysed genes and a number of performance traits evaluated, including milk protein percentage, somatic cell count, culled cow and progeny carcass weight, angularity, body conditioning score, progeny carcass conformation, body depth, rump angle, rump width, animal stature, calving difficulty, gestation length and calf perinatal mortality. Notably, SNPs within the imprinted paternally expressed gene 3 (PEG3) gene cluster were associated (P ≤ 0.05) with calving, calf performance and fertility traits, while a single SNP in the zinc finger protein 215 gene (ZNF215) was associated with milk protein percentage (P ≤ 0.05), progeny carcass weight (P ≤ 0.05), culled cow carcass weight (P ≤ 0.01), angularity (P

  12. Sifting through genomes with iterative-sequence clustering produces a large, phylogenetically diverse protein-family resource

    Directory of Open Access Journals (Sweden)

    Sharpton Thomas J

    2012-10-01

    Full Text Available Abstract Background New computational resources are needed to manage the increasing volume of biological data from genome sequencing projects. One fundamental challenge is the ability to maintain a complete and current catalog of protein diversity. We developed a new approach for the identification of protein families that focuses on the rapid discovery of homologous protein sequences. Results We implemented fully automated and high-throughput procedures to de novo cluster proteins into families based upon global alignment similarity. Our approach employs an iterative clustering strategy in which homologs of known families are sifted out of the search for new families. The resulting reduction in computational complexity enables us to rapidly identify novel protein families found in new genomes and to perform efficient, automated updates that keep pace with genome sequencing. We refer to protein families identified through this approach as “Sifting Families,” or SFams. Our analysis of ~10.5 million protein sequences from 2,928 genomes identified 436,360 SFams, many of which are not represented in other protein family databases. We validated the quality of SFam clustering through statistical as well as network topology–based analyses. Conclusions We describe the rapid identification of SFams and demonstrate how they can be used to annotate genomes and metagenomes. The SFam database catalogs protein-family quality metrics, multiple sequence alignments, hidden Markov models, and phylogenetic trees. Our source code and database are publicly available and will be subject to frequent updates (http://edhar.genomecenter.ucdavis.edu/sifting_families/.

  13. Genetic polymorphisms in the glutamate-rich protein of Plasmodium falciparum field isolates from a malaria-endemic area of Brazil

    DEFF Research Database (Denmark)

    Pratt-Riccio, Lilian Rose; Perce-da-Silva, Daiana de Souza; Lima-Junior, Josué da Costa

    2013-01-01

    The genetic diversity displayed by Plasmodium falciparum, the most deadly Plasmodium species, is a significant obstacle for effective malaria vaccine development. In this study, we identified genetic polymorphisms in P. falciparum glutamate-rich protein (GLURP), which is currently being tested in...

  14. Novel fluorescent sequence-related amplified polymorphism(FSRAP markers for the construction of a genetic linkage map of wheat(Triticum aestivum L.

    Directory of Open Access Journals (Sweden)

    Zhao Lingbo

    2017-01-01

    Full Text Available Novel fluorescent sequence-related amplified polymorphism (FSRAP markers were developed based on the SRAP molecular marker. Then, the FSRAP markers were used to construct the genetic map of a wheat (Triticum aestivumL. recombinant inbred line population derived from a Chuanmai 42×Chuannong 16 cross. Reproducibility and polymorphism tests indicated that the FSRAP markers have repeatability and better reflect the polymorphism of wheat varieties compared with SRAP markers. A total of 430 polymorphic loci between Chuanmai 42 and Chuannong 16 were detected with 189 FSRAP primer combinations. A total of 281 FSARP markers and 39 SSR markers re classified into 20 linkage groups. The maps spanned a total length of 2499.3cM with an average distance of 7.81cM between markers. A total of 201 markers were mapped on the B genome and covered a distance of 1013cM. On the A genome, 84 markers were mapped and covered a distance of 849.6cM. On the D genome, however, only 35 markers were mapped and covered a distance of 636.7cM. No FSRAP markers were distributed on the 7D chromosome. The results of the present study revealed that the novel FSRAP markers can be used to generate dense, uniform genetic maps of wheat.

  15. Polymorphisms in the Prion Protein Gene of cattle breeds from Brazil

    Directory of Open Access Journals (Sweden)

    Cristiane C. Sanches

    Full Text Available ABSTRACT: One of the alterations that occur in the PRNP gene in bovines is the insertion/deletion (indel of base sequences in specific regions, such as indels of 12-base pairs (bp in intron 1 and of 23- bp in the promoter region. The deletion allele of 23 bp is associated with susceptibility to bovine spongiform encephalopathy (BSE as well as the presence of the deletion allele of 12 bp. In the present study, the variability of nucleotides in the promoter region and intron 1 of the PRNP gene was genotyped for the Angus, Canchim, Nellore and Simmental bovine breeds to identify the genotype profiles of resistance and/or susceptibility to BSE in each animal. Genomic DNA was extracted for amplification of the target regions of the PRNP gene using polymerase chain reaction (PCR and specific primers. The PCR products were submitted to electrophoresis in agarose gel 3% and sequencing for genotyping. With the exception of the Angus breed, most breeds exhibited a higher frequency of deletion alleles for 12 bp and 23 bp in comparison to their respective insertion alleles for both regions. These results represent an important contribution to understanding the formation process of the Brazilian herd in relation to bovine PRNP gene polymorphisms.

  16. Pooled Enrichment Sequencing Identifies Diversity and Evolutionary Pressures at NLR Resistance Genes within a Wild Tomato Population

    Science.gov (United States)

    Stam, Remco; Scheikl, Daniela; Tellier, Aurélien

    2016-01-01

    Nod-like receptors (NLRs) are nucleotide-binding domain and leucine-rich repeats containing proteins that are important in plant resistance signaling. Many of the known pathogen resistance (R) genes in plants are NLRs and they can recognize pathogen molecules directly or indirectly. As such, divergence and copy number variants at these genes are found to be high between species. Within populations, positive and balancing selection are to be expected if plants coevolve with their pathogens. In order to understand the complexity of R-gene coevolution in wild nonmodel species, it is necessary to identify the full range of NLRs and infer their evolutionary history. Here we investigate and reveal polymorphism occurring at 220 NLR genes within one population of the partially selfing wild tomato species Solanum pennellii. We use a combination of enrichment sequencing and pooling ten individuals, to specifically sequence NLR genes in a resource and cost-effective manner. We focus on the effects which different mapping and single nucleotide polymorphism calling software and settings have on calling polymorphisms in customized pooled samples. Our results are accurately verified using Sanger sequencing of polymorphic gene fragments. Our results indicate that some NLRs, namely 13 out of 220, have maintained polymorphism within our S. pennellii population. These genes show a wide range of πN/πS ratios and differing site frequency spectra. We compare our observed rate of heterozygosity with expectations for this selfing and bottlenecked population. We conclude that our method enables us to pinpoint NLR genes which have experienced natural selection in their habitat. PMID:27189991

  17. Correlation of PTPN11 polymorphism at intron 3 with gastric cancer

    Directory of Open Access Journals (Sweden)

    Li ZHANG

    2011-05-01

    Full Text Available Objective To explore the correlation of protein-tyrosine-phosphatase nonreceptor-type 11(PTPN11 polymorphism at intron 3 with gastric cancer in Chinese population,and the feasibility and accuracy of employing mastrix-assisted laser desorption ionization time-of-flight mass spectrogram(MALDI-TOF-MS in genotyping of this SNP.Methods Two hundred and forty-seven patients with gastric cancer,212 cancer-free patients and 160 cord blood samples were enrolled in present study.Genotypes of PTPN11 G/A polymorphism at intron 3 were determined by MALDI-TOF-MS analysis,and direct sequencing of PCR products with 20 samples of the gene locus was done for checking the accuracy of MALDI-TOF-MS.Histological examination,Helicobacter pylori culture,rapid urease test,serum anti-H.pylori antibodies(ELISA and urease colloidal gold test were performed to evaluate H.pylori infection.Results Direct sequencing of 20 random selected samples were well consistent with the MALDI-TOF-MS results.The rates of H.pylori infection were 73.68% in gastric cancer patients and 75.47% in cancer-free patients,implying no significant difference between the two groups.The distributions of genotypes were in Hardy Weinberg equilibrium in both gastric cancer patients and controls.There were no significant differences in the genotype frequencies between the 2 groups(P>0.05.Compared with the GG genotype,GA+AA genotype could not influence the risk of gastric cancer.When stratified for status,PTPN11 polymorphism was not associated with age,gender and H.pylori infection states in both cancer patients and controls.Conclusion It seems that PTPN11 G/A polymorphism at intron 3 has no affection on the risk of gastric cancer in Chinese population.

  18. In silico polymorphism analysis for the development of simple sequence repeat and transposon markers and construction of linkage map in cultivated peanut

    Directory of Open Access Journals (Sweden)

    Shirasawa Kenta

    2012-06-01

    Full Text Available Abstract Background Peanut (Arachis hypogaea is an autogamous allotetraploid legume (2n = 4x = 40 that is widely cultivated as a food and oil crop. More than 6,000 DNA markers have been developed in Arachis spp., but high-density linkage maps useful for genetics, genomics, and breeding have not been constructed due to extremely low genetic diversity. Polymorphic marker loci are useful for the construction of such high-density linkage maps. The present study used in silico analysis to develop simple sequence repeat-based and transposon-based markers. Results The use of in silico analysis increased the efficiency of polymorphic marker development by more than 3-fold. In total, 926 (34.2% of 2,702 markers showed polymorphisms between parental lines of the mapping population. Linkage analysis of the 926 markers along with 253 polymorphic markers selected from 4,449 published markers generated 21 linkage groups covering 2,166.4 cM with 1,114 loci. Based on the map thus produced, 23 quantitative trait loci (QTLs for 15 agronomical traits were detected. Another linkage map with 326 loci was also constructed and revealed a relationship between the genotypes of the FAD2 genes and the ratio of oleic/linoleic acid in peanut seed. Conclusions In silico analysis of polymorphisms increased the efficiency of polymorphic marker development, and contributed to the construction of high-density linkage maps in cultivated peanut. The resultant maps were applicable to QTL analysis. Marker subsets and linkage maps developed in this study should be useful for genetics, genomics, and breeding in Arachis. The data are available at the Kazusa DNA Marker Database (http://marker.kazusa.or.jp.

  19. Synthetic biology for the directed evolution of protein biocatalysts: navigating sequence space intelligently

    Science.gov (United States)

    Currin, Andrew; Swainston, Neil; Day, Philip J.

    2015-01-01

    The amino acid sequence of a protein affects both its structure and its function. Thus, the ability to modify the sequence, and hence the structure and activity, of individual proteins in a systematic way, opens up many opportunities, both scientifically and (as we focus on here) for exploitation in biocatalysis. Modern methods of synthetic biology, whereby increasingly large sequences of DNA can be synthesised de novo, allow an unprecedented ability to engineer proteins with novel functions. However, the number of possible proteins is far too large to test individually, so we need means for navigating the ‘search space’ of possible protein sequences efficiently and reliably in order to find desirable activities and other properties. Enzymologists distinguish binding (K d) and catalytic (k cat) steps. In a similar way, judicious strategies have blended design (for binding, specificity and active site modelling) with the more empirical methods of classical directed evolution (DE) for improving k cat (where natural evolution rarely seeks the highest values), especially with regard to residues distant from the active site and where the functional linkages underpinning enzyme dynamics are both unknown and hard to predict. Epistasis (where the ‘best’ amino acid at one site depends on that or those at others) is a notable feature of directed evolution. The aim of this review is to highlight some of the approaches that are being developed to allow us to use directed evolution to improve enzyme properties, often dramatically. We note that directed evolution differs in a number of ways from natural evolution, including in particular the available mechanisms and the likely selection pressures. Thus, we stress the opportunities afforded by techniques that enable one to map sequence to (structure and) activity in silico, as an effective means of modelling and exploring protein landscapes. Because known landscapes may be assessed and reasoned about as a whole

  20. Genetic basis of olfactory cognition: extremely high level of DNA sequence polymorphism in promoter regions of the human olfactory receptor genes revealed using the 1000 Genomes Project dataset.

    Science.gov (United States)

    Ignatieva, Elena V; Levitsky, Victor G; Yudin, Nikolay S; Moshkin, Mikhail P; Kolchanov, Nikolay A

    2014-01-01

    The molecular mechanism of olfactory cognition is very complicated. Olfactory cognition is initiated by olfactory receptor proteins (odorant receptors), which are activated by olfactory stimuli (ligands). Olfactory receptors are the initial player in the signal transduction cascade producing a nerve impulse, which is transmitted to the brain. The sensitivity to a particular ligand depends on the expression level of multiple proteins involved in the process of olfactory cognition: olfactory receptor proteins, proteins that participate in signal transduction cascade, etc. The expression level of each gene is controlled by its regulatory regions, and especially, by the promoter [a region of DNA about 100-1000 base pairs long located upstream of the transcription start site (TSS)]. We analyzed single nucleotide polymorphisms using human whole-genome data from the 1000 Genomes Project and revealed an extremely high level of single nucleotide polymorphisms in promoter regions of olfactory receptor genes and HLA genes. We hypothesized that the high level of polymorphisms in olfactory receptor promoters was responsible for the diversity in regulatory mechanisms controlling the expression levels of olfactory receptor proteins. Such diversity of regulatory mechanisms may cause the great variability of olfactory cognition of numerous environmental olfactory stimuli perceived by human beings (air pollutants, human body odors, odors in culinary etc.). In turn, this variability may provide a wide range of emotional and behavioral reactions related to the vast variety of olfactory stimuli.

  1. Genomic analysis of an attenuated Chlamydia abortus live vaccine strain reveals defects in central metabolism and surface proteins.

    Science.gov (United States)

    Burall, L S; Rodolakis, A; Rekiki, A; Myers, G S A; Bavoil, P M

    2009-09-01

    Comparative genomic analysis of a wild-type strain of the ovine pathogen Chlamydia abortus and its nitrosoguanidine-induced, temperature-sensitive, virulence-attenuated live vaccine derivative identified 22 single nucleotide polymorphisms unique to the mutant, including nine nonsynonymous mutations, one leading to a truncation of pmpG, which encodes a polymorphic membrane protein, and two intergenic mutations potentially affecting promoter sequences. Other nonsynonymous mutations mapped to a pmpG pseudogene and to predicted coding sequences encoding a putative lipoprotein, a sigma-54-dependent response regulator, a PhoH-like protein, a putative export protein, two tRNA synthetases, and a putative serine hydroxymethyltransferase. One of the intergenic mutations putatively affects transcription of two divergent genes encoding pyruvate kinase and a putative SOS response nuclease, respectively. These observations suggest that the temperature-sensitive phenotype and associated virulence attenuation of the vaccine strain result from disrupted metabolic activity due to altered pyruvate kinase expression and/or alteration in the function of one or more membrane proteins, most notably PmpG and a putative lipoprotein.

  2. Search for methylation-sensitive amplification polymorphisms in mutant figs

    OpenAIRE

    Rodrigues, M. G F; Martins, A. B G [UNESP; Bertoni, B. W.; Figueira, A.; Giuliatti, S.

    2013-01-01

    Fig (Ficus carica) breeding programs that use conventional approaches to develop new cultivars are rare, owing to limited genetic variability and the difficulty in obtaining plants via gamete fusion. Cytosine methylation in plants leads to gene repression, thereby affecting transcription without changing the DNA sequence. Previous studies using random amplification of polymorphic DNA and amplified fragment length polymorphism markers revealed no polymorphisms among select fig mutants that ori...

  3. Characterization of European Yersinia enterocolitica 1A strains using restriction fragment length polymorphism and multilocus sequence analysis.

    Science.gov (United States)

    Murros, A; Säde, E; Johansson, P; Korkeala, H; Fredriksson-Ahomaa, M; Björkroth, J

    2016-10-01

    Yersinia enterocolitica is currently divided into two subspecies: subsp. enterocolitica including highly pathogenic strains of biotype 1B and subsp. palearctica including nonpathogenic strains of biotype 1A and moderately pathogenic strains of biotypes 2-5. In this work, we characterized 162 Y. enterocolitica strains of biotype 1A and 50 strains of biotypes 2-4 isolated from human, animal and food samples by restriction fragment length polymorphism using the HindIII restriction enzyme. Phylogenetic relatedness of 20 representative Y. enterocolitica strains including 15 biotype 1A strains was further studied by the multilocus sequence analysis of four housekeeping genes (glnA, gyrB, recA and HSP60). In all the analyses, biotype 1A strains formed a separate genomic group, which differed from Y. enterocolitica subsp. enterocolitica and from the strains of biotypes 2-4 of Y. enterocolitica subsp. palearctica. Based on these results, biotype 1A strains considered nonpathogenic should not be included in subspecies palearctica containing pathogenic strains of biotypes 2-5. Yersinia enterocolitica strains are currently divided into six biotypes and two subspecies. Strains of biotype 1A, which are phenotypically and genotypically very heterogeneous, are classified as subspecies palearctica. In this study, European Y. enterocolitica 1A strains isolated from both human and nonhuman sources were characterized using restriction fragment length polymorphism and multilocus sequence analysis. The European biotype 1A strains formed a separate group, which differed from strains belonging to subspecies enterocolitica and palearctica. This may indicate that the current division between the two subspecies is not sufficient considering the strain diversity within Y. enterocolitica. © 2016 The Society for Applied Microbiology.

  4. DNA-binding proteins from marine bacteria expand the known sequence diversity of TALE-like repeats.

    Science.gov (United States)

    de Lange, Orlando; Wolf, Christina; Thiel, Philipp; Krüger, Jens; Kleusch, Christian; Kohlbacher, Oliver; Lahaye, Thomas

    2015-11-16

    Transcription Activator-Like Effectors (TALEs) of Xanthomonas bacteria are programmable DNA binding proteins with unprecedented target specificity. Comparative studies into TALE repeat structure and function are hindered by the limited sequence variation among TALE repeats. More sequence-diverse TALE-like proteins are known from Ralstonia solanacearum (RipTALs) and Burkholderia rhizoxinica (Bats), but RipTAL and Bat repeats are conserved with those of TALEs around the DNA-binding residue. We study two novel marine-organism TALE-like proteins (MOrTL1 and MOrTL2), the first to date of non-terrestrial origin. We have assessed their DNA-binding properties and modelled repeat structures. We found that repeats from these proteins mediate sequence specific DNA binding conforming to the TALE code, despite low sequence similarity to TALE repeats, and with novel residues around the BSR. However, MOrTL1 repeats show greater sequence discriminating power than MOrTL2 repeats. Sequence alignments show that there are only three residues conserved between repeats of all TALE-like proteins including the two new additions. This conserved motif could prove useful as an identifier for future TALE-likes. Additionally, comparing MOrTL repeats with those of other TALE-likes suggests a common evolutionary origin for the TALEs, RipTALs and Bats. © The Author(s) 2015. Published by Oxford University Press on behalf of Nucleic Acids Research.

  5. CISAPS: Complex Informational Spectrum for the Analysis of Protein Sequences

    Directory of Open Access Journals (Sweden)

    Charalambos Chrysostomou

    2015-01-01

    Full Text Available Complex informational spectrum analysis for protein sequences (CISAPS and its web-based server are developed and presented. As recent studies show, only the use of the absolute spectrum in the analysis of protein sequences using the informational spectrum analysis is proven to be insufficient. Therefore, CISAPS is developed to consider and provide results in three forms including absolute, real, and imaginary spectrum. Biologically related features to the analysis of influenza A subtypes as presented as a case study in this study can also appear individually either in the real or imaginary spectrum. As the results presented, protein classes can present similarities or differences according to the features extracted from CISAPS web server. These associations are probable to be related with the protein feature that the specific amino acid index represents. In addition, various technical issues such as zero-padding and windowing that may affect the analysis are also addressed. CISAPS uses an expanded list of 611 unique amino acid indices where each one represents a different property to perform the analysis. This web-based server enables researchers with little knowledge of signal processing methods to apply and include complex informational spectrum analysis to their work.

  6. Genotyping via Sequence Related Amplified Polymorphism Markers in Fusarium culmorum

    Directory of Open Access Journals (Sweden)

    Işıl Melis Zümrüt

    2018-02-01

    Full Text Available Fusarium culmorum is predominating causal agent of head blight (HB and root rot (RR in cereals worldwide. Since F. culmorum has a great level of genetic diversity and the parasexual stage is assumed for this phytopathogen, characterization of isolates from different regions is significant step in food safety and controlling the HB. In this study, it was aimed to characterize totally 37 F. culmorum isolates from Turkey via sequence related amplified polymorphism (SRAP marker based genotyping. MAT-1/MAT-2 type assay was also used in order to reveal intraspecific variation in F. culmorum. MAT-1 and MAT-2 specific primer pairs for mating assays resulted in 210 and 260 bp bands, respectively. 11 of isolates were belonged to MAT-1 type whereas 19 samples were of MAT-2. Remaining 7 samples yielded both amplicons. Totally 9 SRAP primer sets yielded amplicons from all isolates. Genetic similarity values were ranged from 39 to 94.7%. Total band number was 127 and PCR product sizes were in the range of 0.1-2.5 kb. Amplicon numbers for individuals were ranged from 1 to 16. According to data obtained from current study, SRAP based genotyping is powerful tool for supporting the data obtained from investigations including phenotypic and agro-ecological characteristics. Findings showed that SRAP-based markers could be useful in F. culmorum characterization.

  7. Identification of a novel Plasmopara halstedii elicitor protein combining de novo peptide sequencing algorithms and RACE-PCR

    Directory of Open Access Journals (Sweden)

    Madlung Johannes

    2010-05-01

    Full Text Available Abstract Background Often high-quality MS/MS spectra of tryptic peptides do not match to any database entry because of only partially sequenced genomes and therefore, protein identification requires de novo peptide sequencing. To achieve protein identification of the economically important but still unsequenced plant pathogenic oomycete Plasmopara halstedii, we first evaluated the performance of three different de novo peptide sequencing algorithms applied to a protein digests of standard proteins using a quadrupole TOF (QStar Pulsar i. Results The performance order of the algorithms was PEAKS online > PepNovo > CompNovo. In summary, PEAKS online correctly predicted 45% of measured peptides for a protein test data set. All three de novo peptide sequencing algorithms were used to identify MS/MS spectra of tryptic peptides of an unknown 57 kDa protein of P. halstedii. We found ten de novo sequenced peptides that showed homology to a Phytophthora infestans protein, a closely related organism of P. halstedii. Employing a second complementary approach, verification of peptide prediction and protein identification was performed by creation of degenerate primers for RACE-PCR and led to an ORF of 1,589 bp for a hypothetical phosphoenolpyruvate carboxykinase. Conclusions Our study demonstrated that identification of proteins within minute amounts of sample material improved significantly by combining sensitive LC-MS methods with different de novo peptide sequencing algorithms. In addition, this is the first study that verified protein prediction from MS data by also employing a second complementary approach, in which RACE-PCR led to identification of a novel elicitor protein in P. halstedii.

  8. Genetic polymorphism in postoperative sepsis after open heart surgery in infants.

    Science.gov (United States)

    Fakhri, Dicky; Djauzi, Samsuridjal; Murni, Tri Wahyu; Rachmat, Jusuf; Harahap, Alida Roswita; Rahayuningsih, Sri Endah; Mansyur, Muchtaruddin; Santoso, Anwar

    2016-05-01

    Sepsis is one of the complications following open heart surgery. Toll-like receptor 2 and toll-interacting protein polymorphism influence the immune response after open heart surgery. This study aimed to assess the genetic distribution of toll-like receptor 2 N199N and toll-interacting protein rs5743867 polymorphism in the development of postoperative sepsis. A prospective cohort study was conducted in 108 children open heart surgery with a Basic Aristotle score ≥6. Patients with an accompanying congenital anomaly, human immunodeficiency virus infection, or history of previous open heart surgery were excluded. The patients' nutritional status and genetic polymorphism were assessed prior to surgery. The results of genetic polymorphism were obtained through genotyping. Patients' ages on the day of surgery and cardiopulmonary bypass times were recorded. The diagnosis of sepsis was established according to Surviving Sepsis Campaign criteria. Postoperative sepsis was observed in 21% of patients. There were 92.6% patients with toll-like receptor 2 N199N polymorphism and 52.8% with toll-interacting protein rs5743867 polymorphism. Toll-like receptor 2 N199N polymorphism tends to increase the risk of sepsis (odds ratio = 1.974; 95% confidence interval: 0.23-16.92; p = 0.504), while toll-interacting protein rs5743867 polymorphism tends to decrease the risk of sepsis (odds ratio = 0.496; 95% confidence interval: 0.19-1.27; p = 0.139) in infants open heart surgery. © The Author(s) 2016.

  9. Prediction of flexible/rigid regions from protein sequences using k-spaced amino acid pairs

    Directory of Open Access Journals (Sweden)

    Ruan Jishou

    2007-04-01

    Full Text Available Abstract Background Traditionally, it is believed that the native structure of a protein corresponds to a global minimum of its free energy. However, with the growing number of known tertiary (3D protein structures, researchers have discovered that some proteins can alter their structures in response to a change in their surroundings or with the help of other proteins or ligands. Such structural shifts play a crucial role with respect to the protein function. To this end, we propose a machine learning method for the prediction of the flexible/rigid regions of proteins (referred to as FlexRP; the method is based on a novel sequence representation and feature selection. Knowledge of the flexible/rigid regions may provide insights into the protein folding process and the 3D structure prediction. Results The flexible/rigid regions were defined based on a dataset, which includes protein sequences that have multiple experimental structures, and which was previously used to study the structural conservation of proteins. Sequences drawn from this dataset were represented based on feature sets that were proposed in prior research, such as PSI-BLAST profiles, composition vector and binary sequence encoding, and a newly proposed representation based on frequencies of k-spaced amino acid pairs. These representations were processed by feature selection to reduce the dimensionality. Several machine learning methods for the prediction of flexible/rigid regions and two recently proposed methods for the prediction of conformational changes and unstructured regions were compared with the proposed method. The FlexRP method, which applies Logistic Regression and collocation-based representation with 95 features, obtained 79.5% accuracy. The two runner-up methods, which apply the same sequence representation and Support Vector Machines (SVM and Naïve Bayes classifiers, obtained 79.2% and 78.4% accuracy, respectively. The remaining considered methods are

  10. Neanderthal and Denisova tooth protein variants in present-day humans.

    Directory of Open Access Journals (Sweden)

    Clément Zanolli

    Full Text Available Environment parameters, diet and genetic factors interact to shape tooth morphostructure. In the human lineage, archaic and modern hominins show differences in dental traits, including enamel thickness, but variability also exists among living populations. Several polymorphisms, in particular in the non-collagenous extracellular matrix proteins of the tooth hard tissues, like enamelin, are involved in dental structure variation and defects and may be associated with dental disorders or susceptibility to caries. To gain insights into the relationships between tooth protein polymorphisms and dental structural morphology and defects, we searched for non-synonymous polymorphisms in tooth proteins from Neanderthal and Denisova hominins. The objective was to identify archaic-specific missense variants that may explain the dental morphostructural variability between extinct and modern humans, and to explore their putative impact on present-day dental phenotypes. Thirteen non-collagenous extracellular matrix proteins specific to hard dental tissues have been selected, searched in the publicly available sequence databases of Neanderthal and Denisova individuals and compared with modern human genome data. A total of 16 non-synonymous polymorphisms were identified in 6 proteins (ameloblastin, amelotin, cementum protein 1, dentin matrix acidic phosphoprotein 1, enamelin and matrix Gla protein. Most of them are encoded by dentin and enamel genes located on chromosome 4, previously reported to show signs of archaic introgression within Africa. Among the variants shared with modern humans, two are ancestral (common with apes and one is the derived enamelin major variant, T648I (rs7671281, associated with a thinner enamel and specific to the Homo lineage. All the others are specific to Neanderthals and Denisova, and are found at a very low frequency in modern Africans or East and South Asians, suggesting that they may be related to particular dental traits or

  11. Rapid detection and purification of sequence specific DNA binding proteins using magnetic separation

    Directory of Open Access Journals (Sweden)

    TIJANA SAVIC

    2006-02-01

    Full Text Available In this paper, a method for the rapid identification and purification of sequence specific DNA binding proteins based on magnetic separation is presented. This method was applied to confirm the binding of the human recombinant USF1 protein to its putative binding site (E-box within the human SOX3 protomer. It has been shown that biotinylated DNA attached to streptavidin magnetic particles specifically binds the USF1 protein in the presence of competitor DNA. It has also been demonstrated that the protein could be successfully eluted from the beads, in high yield and with restored DNA binding activity. The advantage of these procedures is that they could be applied for the identification and purification of any high-affinity sequence-specific DNA binding protein with only minor modifications.

  12. BCL2-like 11 intron 2 deletion polymorphism is not associated with non-small cell lung cancer risk and prognosis.

    Science.gov (United States)

    Cho, Eun Na; Kim, Eun Young; Jung, Ji Ye; Kim, Arum; Oh, In Jae; Kim, Young Chul; Chang, Yoon Soo

    2015-10-01

    BCL2-Like 11(BIM), which encodes a BH3-only protein, is a major pro-apoptotic molecule that facilitates cell death. We hypothesized that a BIM intron 2 deletion polymorphism increases lung cancer risk and predicts poor prognosis in non-small lung cancer (NSCLC) patients. We prospectively recruited 450 lung cancer patients and 1:1 age, sex, and smoking status matched control subjects from February 2013 to April 2014 among patients treated at Severance, Gangnam Severance, and Chonnam Hwasoon Hospital. The presence of a 2903-bp genomic DNA deletion polymorphism of intron 2 of BIM was analyzed by PCR and validated by sequencing. Odds ratios were calculated by chi-square tests and survival analysis with Kaplan-Meier estimation. Sixty-nine out of 450 (15.3%) lung cancer patients carried the BIM deletion polymorphism, while 66 out of 450 (14.7%) control subjects carried the BIM deletion polymorphism, with an odds ratio of for lung cancer of 1.054 (95% CI; 0.731-1.519). We categorized 406 NSCLC patients according to the presence of the polymorphism and found that there were no statistically significant differences in age, sex, histologic type, or stage between subjects with and without the deletion polymorphism. The BIM deletion polymorphism did not influence overall survival (OS) or progression free survival (PFS) in our sample (OS; 37.6 vs 34.4 months (P=0.759), PFS; 49.6 vs 26.0 months (P=0.434)). These findings indicate that the BIM deletion polymorphism is common in Korean NSCLC patients but does not significantly affect the intrinsic biologic function of BH3-only protein. Furthermore, the BIM deletion polymorphism did not predict clinical outcomes in patients with NSCLC. Copyright © 2015 Elsevier Ireland Ltd. All rights reserved.

  13. Fast computational methods for predicting protein structure from primary amino acid sequence

    Science.gov (United States)

    Agarwal, Pratul Kumar [Knoxville, TN

    2011-07-19

    The present invention provides a method utilizing primary amino acid sequence of a protein, energy minimization, molecular dynamics and protein vibrational modes to predict three-dimensional structure of a protein. The present invention also determines possible intermediates in the protein folding pathway. The present invention has important applications to the design of novel drugs as well as protein engineering. The present invention predicts the three-dimensional structure of a protein independent of size of the protein, overcoming a significant limitation in the prior art.

  14. Analysis of correlations between sites in models of protein sequences

    International Nuclear Information System (INIS)

    Giraud, B.G.; Lapedes, A.; Liu, L.C.

    1998-01-01

    A criterion based on conditional probabilities, related to the concept of algorithmic distance, is used to detect correlated mutations at noncontiguous sites on sequences. We apply this criterion to the problem of analyzing correlations between sites in protein sequences; however, the analysis applies generally to networks of interacting sites with discrete states at each site. Elementary models, where explicit results can be derived easily, are introduced. The number of states per site considered ranges from 2, illustrating the relation to familiar classical spin systems, to 20 states, suitable for representing amino acids. Numerical simulations show that the criterion remains valid even when the genetic history of the data samples (e.g., protein sequences), as represented by a phylogenetic tree, introduces nonindependence between samples. Statistical fluctuations due to finite sampling are also investigated and do not invalidate the criterion. A subsidiary result is found: The more homogeneous a population, the more easily its average properties can drift from the properties of its ancestor. copyright 1998 The American Physical Society

  15. Sequence walkers: a graphical method to display how binding proteins interact with DNA or RNA sequences | Center for Cancer Research

    Science.gov (United States)

    A graphical method is presented for displaying how binding proteins and other macromolecules interact with individual bases of nucleotide sequences. Characters representing the sequence are either oriented normally and placed above a line indicating favorable contact, or upside-down and placed below the line indicating unfavorable contact. The positive or negative height of each letter shows the contribution of that base to the average sequence conservation of the binding site, as represented by a sequence logo.

  16. DeepGO: predicting protein functions from sequence and interactions using a deep ontology-aware classifier.

    Science.gov (United States)

    Kulmanov, Maxat; Khan, Mohammed Asif; Hoehndorf, Robert; Wren, Jonathan

    2018-02-15

    A large number of protein sequences are becoming available through the application of novel high-throughput sequencing technologies. Experimental functional characterization of these proteins is time-consuming and expensive, and is often only done rigorously for few selected model organisms. Computational function prediction approaches have been suggested to fill this gap. The functions of proteins are classified using the Gene Ontology (GO), which contains over 40 000 classes. Additionally, proteins have multiple functions, making function prediction a large-scale, multi-class, multi-label problem. We have developed a novel method to predict protein function from sequence. We use deep learning to learn features from protein sequences as well as a cross-species protein-protein interaction network. Our approach specifically outputs information in the structure of the GO and utilizes the dependencies between GO classes as background information to construct a deep learning model. We evaluate our method using the standards established by the Computational Assessment of Function Annotation (CAFA) and demonstrate a significant improvement over baseline methods such as BLAST, in particular for predicting cellular locations. Web server: http://deepgo.bio2vec.net, Source code: https://github.com/bio-ontology-research-group/deepgo. robert.hoehndorf@kaust.edu.sa. Supplementary data are available at Bioinformatics online. © The Author(s) 2017. Published by Oxford University Press.

  17. Determining and comparing protein function in Bacterial genome sequences

    DEFF Research Database (Denmark)

    Vesth, Tammi Camilla

    of this class have very little homology to other known genomes making functional annotation based on sequence similarity very difficult. Inspired in part by this analysis, an approach for comparative functional annotation was created based public sequenced genomes, CMGfunc. Functionally related groups......In November 2013, there was around 21.000 different prokaryotic genomes sequenced and publicly available, and the number is growing daily with another 20.000 or more genomes expected to be sequenced and deposited by the end of 2014. An important part of the analysis of this data is the functional...... annotation of genes – the descriptions assigned to genes that describe the likely function of the encoded proteins. This process is limited by several factors, including the definition of a function which can be more or less specific as well as how many genes can actually be assigned a function based...

  18. Sequencing and annotation of the chloroplast DNAs and identification of polymorphisms distinguishing normal male-fertile and male-sterile cytoplasms of onion.

    Science.gov (United States)

    von Kohn, Christopher; Kiełkowska, Agnieszka; Havey, Michael J

    2013-12-01

    Male-sterile (S) cytoplasm of onion is an alien cytoplasm introgressed into onion in antiquity and is widely used for hybrid seed production. Owing to the biennial generation time of onion, classical crossing takes at least 4 years to classify cytoplasms as S or normal (N) male-fertile. Molecular markers in the organellar DNAs that distinguish N and S cytoplasms are useful to reduce the time required to classify onion cytoplasms. In this research, we completed next-generation sequencing of the chloroplast DNAs of N- and S-cytoplasmic onions; we assembled and annotated the genomes in addition to identifying polymorphisms that distinguish these cytoplasms. The sizes (153 538 and 153 355 base pairs) and GC contents (36.8%) were very similar for the chloroplast DNAs of N and S cytoplasms, respectively, as expected given their close phylogenetic relationship. The size difference was primarily due to small indels in intergenic regions and a deletion in the accD gene of N-cytoplasmic onion. The structures of the onion chloroplast DNAs were similar to those of most land plants with large and small single copy regions separated by inverted repeats. Twenty-eight single nucleotide polymorphisms, two polymorphic restriction-enzyme sites, and one indel distributed across 20 chloroplast genes in the large and small single copy regions were selected and validated using diverse onion populations previously classified as N or S cytoplasmic using restriction fragment length polymorphisms. Although cytoplasmic male sterility is likely associated with the mitochondrial DNA, maternal transmission of the mitochondrial and chloroplast DNAs allows for polymorphisms in either genome to be useful for classifying onion cytoplasms to aid the development of hybrid onion cultivars.

  19. A synonymous polymorphic variation in ACADM exon 11 affects splicing efficiency and may affect fatty acid oxidation

    DEFF Research Database (Denmark)

    Bruun, Gitte Hoffmann; Doktor, Thomas Koed; Andresen, Brage Storstein

    2013-01-01

    beta-oxidation of medium-chain fatty acids. We examined the functional basis for this association and identified linkage between rs211718 and the intragenic synonymous polymorphic variant c.1161A>G in ACADM exon 11 (rs1061337). Employing minigene studies we show that the c.1161A allele is associated......, perhaps due to improved splicing. This study is a proof of principle that synonymous SNPs are not neutral. By changing the binding sites for splicing regulatory proteins they can have significant effects on pre-mRNA splicing and thus protein function. In addition, this study shows that for a sequence...

  20. FASTERp: A Feature Array Search Tool for Estimating Resemblance of Protein Sequences

    Energy Technology Data Exchange (ETDEWEB)

    Macklin, Derek; Egan, Rob; Wang, Zhong

    2014-03-14

    Metagenome sequencing efforts have provided a large pool of billions of genes for identifying enzymes with desirable biochemical traits. However, homology search with billions of genes in a rapidly growing database has become increasingly computationally impractical. Here we present our pilot efforts to develop a novel alignment-free algorithm for homology search. Specifically, we represent individual proteins as feature vectors that denote the presence or absence of short kmers in the protein sequence. Similarity between feature vectors is then computed using the Tanimoto score, a distance metric that can be rapidly computed on bit string representations of feature vectors. Preliminary results indicate good correlation with optimal alignment algorithms (Spearman r of 0.87, ~;;1,000,000 proteins from Pfam), as well as with heuristic algorithms such as BLAST (Spearman r of 0.86, ~;;1,000,000 proteins). Furthermore, a prototype of FASTERp implemented in Python runs approximately four times faster than BLAST on a small scale dataset (~;;1000 proteins). We are optimizing and scaling to improve FASTERp to enable rapid homology searches against billion-protein databases, thereby enabling more comprehensive gene annotation efforts.

  1. muBLASTP: database-indexed protein sequence search on multicore CPUs.

    Science.gov (United States)

    Zhang, Jing; Misra, Sanchit; Wang, Hao; Feng, Wu-Chun

    2016-11-04

    The Basic Local Alignment Search Tool (BLAST) is a fundamental program in the life sciences that searches databases for sequences that are most similar to a query sequence. Currently, the BLAST algorithm utilizes a query-indexed approach. Although many approaches suggest that sequence search with a database index can achieve much higher throughput (e.g., BLAT, SSAHA, and CAFE), they cannot deliver the same level of sensitivity as the query-indexed BLAST, i.e., NCBI BLAST, or they can only support nucleotide sequence search, e.g., MegaBLAST. Due to different challenges and characteristics between query indexing and database indexing, the existing techniques for query-indexed search cannot be used into database indexed search. muBLASTP, a novel database-indexed BLAST for protein sequence search, delivers identical hits returned to NCBI BLAST. On Intel Haswell multicore CPUs, for a single query, the single-threaded muBLASTP achieves up to a 4.41-fold speedup for alignment stages, and up to a 1.75-fold end-to-end speedup over single-threaded NCBI BLAST. For a batch of queries, the multithreaded muBLASTP achieves up to a 5.7-fold speedups for alignment stages, and up to a 4.56-fold end-to-end speedup over multithreaded NCBI BLAST. With a newly designed index structure for protein database and associated optimizations in BLASTP algorithm, we re-factored BLASTP algorithm for modern multicore processors that achieves much higher throughput with acceptable memory footprint for the database index.

  2. Analysis of alkaptonuria (AKU) mutations and polymorphisms reveals that the CCC sequence motif is a mutational hot spot in the homogentisate 1,2 dioxygenase gene (HGO).

    Science.gov (United States)

    Beltrán-Valero de Bernabé, D; Jimenez, F J; Aquaron, R; Rodríguez de Córdoba, S

    1999-01-01

    We recently showed that alkaptonuria (AKU) is caused by loss-of-function mutations in the homogentisate 1,2 dioxygenase gene (HGO). Herein we describe haplotype and mutational analyses of HGO in seven new AKU pedigrees. These analyses identified two novel single-nucleotide polymorphisms (INV4+31A-->G and INV11+18A-->G) and six novel AKU mutations (INV1-1G-->A, W60G, Y62C, A122D, P230T, and D291E), which further illustrates the remarkable allelic heterogeneity found in AKU. Reexamination of all 29 mutations and polymorphisms thus far described in HGO shows that these nucleotide changes are not randomly distributed; the CCC sequence motif and its inverted complement, GGG, are preferentially mutated. These analyses also demonstrated that the nucleotide substitutions in HGO do not involve CpG dinucleotides, which illustrates important differences between HGO and other genes for the occurrence of mutation at specific short-sequence motifs. Because the CCC sequence motifs comprise a significant proportion (34.5%) of all mutated bases that have been observed in HGO, we conclude that the CCC triplet is a mutational hot spot in HGO. PMID:10205262

  3. Isolation and N-terminal sequencing of a novel cadmium-binding protein from Boletus edulis

    Science.gov (United States)

    Collin-Hansen, C.; Andersen, R. A.; Steinnes, E.

    2003-05-01

    A Cd-binding protein was isolated from the popular edible mushroom Boletus edulis, which is a hyperaccumulator of both Cd and Hg. Wild-growing samples of B. edulis were collected from soils rich in Cd. Cd radiotracer was added to the crude protein preparation obtained from ethanol precipitation of heat-treated cytosol. Proteins were then further separated in two consecutive steps; gel filtration and anion exchange chromatography. In both steps the Cd radiotracer profile showed only one distinct peak, which corresponded well with the profiles of endogenous Cd obtained by atomic absorption spectrophotometry (AAS). Concentrations of the essential elements Cu and Zn were low in the protein fractions high in Cd. N-terminal sequencing performed on the Cd-binding protein fractions revealed a protein with a novel amino acid sequence, which contained aromatic amino acids as well as proline. Both the N-terminal sequencing and spectrofluorimetric analysis with EDTA and ABD-F (4-aminosulfonyl-7-fluoro-2, 1, 3-benzoxadiazole) failed to detect cysteine in the Cd-binding fractions. These findings conclude that the novel protein does not belong to the metallothionein family. The results suggest a role for the protein in Cd transport and storage, and they are of importance in view of toxicology and food chemistry, but also for environmental protection.

  4. Sequence charge decoration dictates coil-globule transition in intrinsically disordered proteins

    Science.gov (United States)

    Firman, Taylor; Ghosh, Kingshuk

    2018-03-01

    We present an analytical theory to compute conformations of heteropolymers—applicable to describe disordered proteins—as a function of temperature and charge sequence. The theory describes coil-globule transition for a given protein sequence when temperature is varied and has been benchmarked against the all-atom Monte Carlo simulation (using CAMPARI) of intrinsically disordered proteins (IDPs). In addition, the model quantitatively shows how subtle alterations of charge placement in the primary sequence—while maintaining the same charge composition—can lead to significant changes in conformation, even as drastic as a coil (swelled above a purely random coil) to globule (collapsed below a random coil) and vice versa. The theory provides insights on how to control (enhance or suppress) these changes by tuning the temperature (or solution condition) and charge decoration. As an application, we predict the distribution of conformations (at room temperature) of all naturally occurring IDPs in the DisProt database and notice significant size variation even among IDPs with a similar composition of positive and negative charges. Based on this, we provide a new diagram-of-states delineating the sequence-conformation relation for proteins in the DisProt database. Next, we study the effect of post-translational modification, e.g., phosphorylation, on IDP conformations. Modifications as little as two-site phosphorylation can significantly alter the size of an IDP with everything else being constant (temperature, salt concentration, etc.). However, not all possible modification sites have the same effect on protein conformations; there are certain "hot spots" that can cause maximal change in conformation. The location of these "hot spots" in the parent sequence can readily be identified by using a sequence charge decoration metric originally introduced by Sawle and Ghosh. The ability of our model to predict conformations (both expanded and collapsed states) of IDPs at

  5. Genomic DNA sequence and cytosine methylation changes of adult rice leaves after seeds space flight

    Science.gov (United States)

    Shi, Jinming

    In this study, cytosine methylation on CCGG site and genomic DNA sequence changes of adult leaves of rice after seeds space flight were detected by methylation-sensitive amplification polymorphism (MSAP) and Amplified fragment length polymorphism (AFLP) technique respectively. Rice seeds were planted in the trial field after 4 days space flight on the shenzhou-6 Spaceship of China. Adult leaves of space-treated rice including 8 plants chosen randomly and 2 plants with phenotypic mutation were used for AFLP and MSAP analysis. Polymorphism of both DNA sequence and cytosine methylation were detected. For MSAP analysis, the average polymorphic frequency of the on-ground controls, space-treated plants and mutants are 1.3%, 3.1% and 11% respectively. For AFLP analysis, the average polymorphic frequencies are 1.4%, 2.9%and 8%respectively. Total 27 and 22 polymorphic fragments were cloned sequenced from MSAP and AFLP analysis respectively. Nine of the 27 fragments from MSAP analysis show homology to coding sequence. For the 22 polymorphic fragments from AFLP analysis, no one shows homology to mRNA sequence and eight fragments show homology to repeat region or retrotransposon sequence. These results suggest that although both genomic DNA sequence and cytosine methylation status can be effected by space flight, the genomic region homology to the fragments from genome DNA and cytosine methylation analysis were different.

  6. Discovery of novel MHC-class I alleles and haplotypes in Filipino cynomolgus macaques (Macaca fascicularis) by pyrosequencing and Sanger sequencing: Mafa-class I polymorphism.

    Science.gov (United States)

    Shiina, Takashi; Yamada, Yukiho; Aarnink, Alice; Suzuki, Shingo; Masuya, Anri; Ito, Sayaka; Ido, Daisuke; Yamanaka, Hisashi; Iwatani, Chizuru; Tsuchiya, Hideaki; Ishigaki, Hirohito; Itoh, Yasushi; Ogasawara, Kazumasa; Kulski, Jerzy K; Blancher, Antoine

    2015-10-01

    Although the low polymorphism of the major histocompatibility complex (MHC) transplantation genes in the Filipino cynomolgus macaque (Macaca fascicularis) is expected to have important implications in the selection and breeding of animals for medical research, detailed polymorphism information is still lacking for many of the duplicated class I genes. To better elucidate the degree and types of MHC polymorphisms and haplotypes in the Filipino macaque population, we genotyped 127 unrelated animals by the Sanger sequencing method and high-resolution pyrosequencing and identified 112 different alleles, 28 at cynomolgus macaque MHC (Mafa)-A, 54 at Mafa-B, 12 at Mafa-I, 11 at Mafa-E, and seven at Mafa-F alleles, of which 56 were newly described. Of them, the newly discovered Mafa-A8*01:01 lineage allele had low nucleotide similarities (Filipino macaque population would identify these and other high-frequency Mafa-class I haplotypes that could be used as MHC control animals for the benefit of biomedical research.

  7. Methylation Sensitive Amplification Polymorphism Sequencing (MSAP-Seq)-A Method for High-Throughput Analysis of Differentially Methylated CCGG Sites in Plants with Large Genomes.

    Science.gov (United States)

    Chwialkowska, Karolina; Korotko, Urszula; Kosinska, Joanna; Szarejko, Iwona; Kwasniewski, Miroslaw

    2017-01-01

    Epigenetic mechanisms, including histone modifications and DNA methylation, mutually regulate chromatin structure, maintain genome integrity, and affect gene expression and transposon mobility. Variations in DNA methylation within plant populations, as well as methylation in response to internal and external factors, are of increasing interest, especially in the crop research field. Methylation Sensitive Amplification Polymorphism (MSAP) is one of the most commonly used methods for assessing DNA methylation changes in plants. This method involves gel-based visualization of PCR fragments from selectively amplified DNA that are cleaved using methylation-sensitive restriction enzymes. In this study, we developed and validated a new method based on the conventional MSAP approach called Methylation Sensitive Amplification Polymorphism Sequencing (MSAP-Seq). We improved the MSAP-based approach by replacing the conventional separation of amplicons on polyacrylamide gels with direct, high-throughput sequencing using Next Generation Sequencing (NGS) and automated data analysis. MSAP-Seq allows for global sequence-based identification of changes in DNA methylation. This technique was validated in Hordeum vulgare . However, MSAP-Seq can be straightforwardly implemented in different plant species, including crops with large, complex and highly repetitive genomes. The incorporation of high-throughput sequencing into MSAP-Seq enables parallel and direct analysis of DNA methylation in hundreds of thousands of sites across the genome. MSAP-Seq provides direct genomic localization of changes and enables quantitative evaluation. We have shown that the MSAP-Seq method specifically targets gene-containing regions and that a single analysis can cover three-quarters of all genes in large genomes. Moreover, MSAP-Seq's simplicity, cost effectiveness, and high-multiplexing capability make this method highly affordable. Therefore, MSAP-Seq can be used for DNA methylation analysis in crop

  8. Methylation Sensitive Amplification Polymorphism Sequencing (MSAP-Seq—A Method for High-Throughput Analysis of Differentially Methylated CCGG Sites in Plants with Large Genomes

    Directory of Open Access Journals (Sweden)

    Karolina Chwialkowska

    2017-11-01

    Full Text Available Epigenetic mechanisms, including histone modifications and DNA methylation, mutually regulate chromatin structure, maintain genome integrity, and affect gene expression and transposon mobility. Variations in DNA methylation within plant populations, as well as methylation in response to internal and external factors, are of increasing interest, especially in the crop research field. Methylation Sensitive Amplification Polymorphism (MSAP is one of the most commonly used methods for assessing DNA methylation changes in plants. This method involves gel-based visualization of PCR fragments from selectively amplified DNA that are cleaved using methylation-sensitive restriction enzymes. In this study, we developed and validated a new method based on the conventional MSAP approach called Methylation Sensitive Amplification Polymorphism Sequencing (MSAP-Seq. We improved the MSAP-based approach by replacing the conventional separation of amplicons on polyacrylamide gels with direct, high-throughput sequencing using Next Generation Sequencing (NGS and automated data analysis. MSAP-Seq allows for global sequence-based identification of changes in DNA methylation. This technique was validated in Hordeum vulgare. However, MSAP-Seq can be straightforwardly implemented in different plant species, including crops with large, complex and highly repetitive genomes. The incorporation of high-throughput sequencing into MSAP-Seq enables parallel and direct analysis of DNA methylation in hundreds of thousands of sites across the genome. MSAP-Seq provides direct genomic localization of changes and enables quantitative evaluation. We have shown that the MSAP-Seq method specifically targets gene-containing regions and that a single analysis can cover three-quarters of all genes in large genomes. Moreover, MSAP-Seq's simplicity, cost effectiveness, and high-multiplexing capability make this method highly affordable. Therefore, MSAP-Seq can be used for DNA methylation

  9. Characterization of single nucleotide polymorphism markers for eelgrass (Zostera marina)

    NARCIS (Netherlands)

    Ferber, Steven; Reusch, Thorsten B. H.; Stam, Wytze T.; Olsen, Jeanine L.

    We characterized 37 single nucleotide polymorphism (SNP) makers for eelgrass Zostera marina. SNP markers were developed using existing EST (expressed sequence tag)-libraries to locate polymorphic loci and develop primers from the functional expressed genes that are deposited in The ZOSTERA database

  10. Exploring Sequence Characteristics Related to High- Level Production of Secreted Proteins in Aspergillus niger

    NARCIS (Netherlands)

    Van den Berg, B.A.; Reinders, M.J.T.; Hulsman, M.; Wu, L.; Pel, H.J.; Roubos, J.A.; De Ridder, D.

    2012-01-01

    Protein sequence features are explored in relation to the production of over-expressed extracellular proteins by fungi. Knowledge on features influencing protein production and secretion could be employed to improve enzyme production levels in industrial bioprocesses via protein engineering. A large

  11. Low level of sequence diversity at merozoite surface protein-1 locus of Plasmodium ovale curtisi and P. ovale wallikeri from Thai isolates.

    Science.gov (United States)

    Putaporntip, Chaturong; Hughes, Austin L; Jongwutiwes, Somchai

    2013-01-01

    The merozoite surface protein-1 (MSP-1) is a candidate target for the development of blood stage vaccines against malaria. Polymorphism in MSP-1 can be useful as a genetic marker for strain differentiation in malarial parasites. Although sequence diversity in the MSP-1 locus has been extensively analyzed in field isolates of Plasmodium falciparum and P. vivax, the extent of variation in its homologues in P. ovale curtisi and P. ovale wallikeri, remains unknown. Analysis of the mitochondrial cytochrome b sequences of 10 P. ovale isolates from symptomatic malaria patients from diverse endemic areas of Thailand revealed co-existence of P. ovale curtisi (n = 5) and P. ovale wallikeri (n = 5). Direct sequencing of the PCR-amplified products encompassing the entire coding region of MSP-1 of P. ovale curtisi (PocMSP-1) and P. ovale wallikeri (PowMSP-1) has identified 3 imperfect repeated segments in the former and one in the latter. Most amino acid differences between these proteins were located in the interspecies variable domains of malarial MSP-1. Synonymous nucleotide diversity (πS) exceeded nonsynonymous nucleotide diversity (πN) for both PocMSP-1 and PowMSP-1, albeit at a non-significant level. However, when MSP-1 of both these species was considered together, πS was significantly greater than πN (pdiversity at this locus prior to speciation. Phylogenetic analysis based on conserved domains has placed PocMSP-1 and PowMSP-1 in a distinct bifurcating branch that probably diverged from each other around 4.5 million years ago. The MSP-1 sequences support that P. ovale curtisi and P. ovale wallikeri are distinct species. Both species are sympatric in Thailand. The low level of sequence diversity in PocMSP-1 and PowMSP-1 among Thai isolates could stem from persistent low prevalence of these species, limiting the chance of outcrossing at this locus.

  12. Revised Mimivirus major capsid protein sequence reveals intron-containing gene structure and extra domain

    Directory of Open Access Journals (Sweden)

    Suzan-Monti Marie

    2009-05-01

    Full Text Available Abstract Background Acanthamoebae polyphaga Mimivirus (APM is the largest known dsDNA virus. The viral particle has a nearly icosahedral structure with an internal capsid shell surrounded with a dense layer of fibrils. A Capsid protein sequence, D13L, was deduced from the APM L425 coding gene and was shown to be the most abundant protein found within the viral particle. However this protein remained poorly characterised until now. A revised protein sequence deposited in a database suggested an additional N-terminal stretch of 142 amino acids missing from the original deduced sequence. This result led us to investigate the L425 gene structure and the biochemical properties of the complete APM major Capsid protein. Results This study describes the full length 3430 bp Capsid coding gene and characterises the 593 amino acids long corresponding Capsid protein 1. The recombinant full length protein allowed the production of a specific monoclonal antibody able to detect the Capsid protein 1 within the viral particle. This protein appeared to be post-translationnally modified by glycosylation and phosphorylation. We proposed a secondary structure prediction of APM Capsid protein 1 compared to the Capsid protein structure of Paramecium Bursaria Chlorella Virus 1, another member of the Nucleo-Cytoplasmic Large DNA virus family. Conclusion The characterisation of the full length L425 Capsid coding gene of Acanthamoebae polyphaga Mimivirus provides new insights into the structure of the main Capsid protein. The production of a full length recombinant protein will be useful for further structural studies.

  13. A Proteomic Workflow Using High-Throughput De Novo Sequencing Towards Complementation of Genome Information for Improved Comparative Crop Science.

    Science.gov (United States)

    Turetschek, Reinhard; Lyon, David; Desalegn, Getinet; Kaul, Hans-Peter; Wienkoop, Stefanie

    2016-01-01

    The proteomic study of non-model organisms, such as many crop plants, is challenging due to the lack of comprehensive genome information. Changing environmental conditions require the study and selection of adapted cultivars. Mutations, inherent to cultivars, hamper protein identification and thus considerably complicate the qualitative and quantitative comparison in large-scale systems biology approaches. With this workflow, cultivar-specific mutations are detected from high-throughput comparative MS analyses, by extracting sequence polymorphisms with de novo sequencing. Stringent criteria are suggested to filter for confidential mutations. Subsequently, these polymorphisms complement the initially used database, which is ready to use with any preferred database search algorithm. In our example, we thereby identified 26 specific mutations in two cultivars of Pisum sativum and achieved an increased number (17 %) of peptide spectrum matches.

  14. Interactions of rat repetitive sequence MspI8 with nuclear matrix proteins during spermatogenesis

    International Nuclear Information System (INIS)

    Rogolinski, J.; Widlak, P.; Rzeszowska-Wolny, J.

    1996-01-01

    Using the Southwestern blot analysis we have studied the interactions between rat repetitive sequence MspI8 and the nuclear matrix proteins of rats testis cells. Starting from 2 weeks the young to adult animal showed differences in type of testis nuclear matrix proteins recognizing the MspI8 sequence. The same sets of nuclear matrix proteins were detected in some enriched in spermatocytes and spermatids and obtained after fractionation of cells of adult animal by the velocity sedimentation technique. (author). 21 refs, 5 figs

  15. Protein model discrimination using mutational sensitivity derived from deep sequencing.

    Science.gov (United States)

    Adkar, Bharat V; Tripathi, Arti; Sahoo, Anusmita; Bajaj, Kanika; Goswami, Devrishi; Chakrabarti, Purbani; Swarnkar, Mohit K; Gokhale, Rajesh S; Varadarajan, Raghavan

    2012-02-08

    A major bottleneck in protein structure prediction is the selection of correct models from a pool of decoys. Relative activities of ∼1,200 individual single-site mutants in a saturation library of the bacterial toxin CcdB were estimated by determining their relative populations using deep sequencing. This phenotypic information was used to define an empirical score for each residue (RankScore), which correlated with the residue depth, and identify active-site residues. Using these correlations, ∼98% of correct models of CcdB (RMSD ≤ 4Å) were identified from a large set of decoys. The model-discrimination methodology was further validated on eleven different monomeric proteins using simulated RankScore values. The methodology is also a rapid, accurate way to obtain relative activities of each mutant in a large pool and derive sequence-structure-function relationships without protein isolation or characterization. It can be applied to any system in which mutational effects can be monitored by a phenotypic readout. Copyright © 2012 Elsevier Ltd. All rights reserved.

  16. Mapping a nucleolar targeting sequence of an RNA binding nucleolar protein, Nop25

    International Nuclear Information System (INIS)

    Fujiwara, Takashi; Suzuki, Shunji; Kanno, Motoko; Sugiyama, Hironobu; Takahashi, Hisaaki; Tanaka, Junya

    2006-01-01

    Nop25 is a putative RNA binding nucleolar protein associated with rRNA transcription. The present study was undertaken to determine the mechanism of Nop25 localization in the nucleolus. Deletion experiments of Nop25 amino acid sequence showed Nop25 to contain a nuclear targeting sequence in the N-terminal and a nucleolar targeting sequence in the C-terminal. By expressing derivative peptides from the C-terminal as GFP-fusion proteins in the cells, a lysine and arginine residue-enriched peptide (KRKHPRRAQDSTKKPPSATRTSKTQRRRR) allowed a GFP-fusion protein to be transported and fully retained in the nucleolus. When the peptide was fused with cMyc epitope and expressed in the cells, a cMyc epitope was then detected in the nucleolus. Nop25 did not localize in the nucleolus by deletion of the peptide from Nop25. Furthermore, deletion of a subdomain (KRKHPRRAQ) in the peptide or amino acid substitution of lysine and arginine residues in the subdomain resulted in the loss of Nop25 nucleolar localization. These results suggest that the lysine and arginine residue-enriched peptide is the most prominent nucleolar targeting sequence of Nop25 and that the long stretch of basic residues might play an important role in the nucleolar localization of Nop25. Although Nop25 contained putative SUMOylation, phosphorylation and glycosylation sites, the amino acid substitution in these sites had no effect on the nucleolar localization, thus suggesting that these post-translational modifications did not contribute to the localization of Nop25 in the nucleolus. The treatment of the cells, which expressed a GFP-fusion protein with a nucleolar targeting sequence of Nop25, with RNase A resulted in a complete dislocation of the protein from the nucleolus. These data suggested that the nucleolar targeting sequence might therefore play an important role in the binding of Nop25 to RNA molecules and that the RNA binding of Nop25 might be essential for the nucleolar localization of Nop25

  17. Prediction of Carbohydrate-Binding Proteins from Sequences Using Support Vector Machines

    Directory of Open Access Journals (Sweden)

    Seizi Someya

    2010-01-01

    Full Text Available Carbohydrate-binding proteins are proteins that can interact with sugar chains but do not modify them. They are involved in many physiological functions, and we have developed a method for predicting them from their amino acid sequences. Our method is based on support vector machines (SVMs. We first clarified the definition of carbohydrate-binding proteins and then constructed positive and negative datasets with which the SVMs were trained. By applying the leave-one-out test to these datasets, our method delivered 0.92 of the area under the receiver operating characteristic (ROC curve. We also examined two amino acid grouping methods that enable effective learning of sequence patterns and evaluated the performance of these methods. When we applied our method in combination with the homology-based prediction method to the annotated human genome database, H-invDB, we found that the true positive rate of prediction was improved.

  18. Genetic basis of olfactory cognition: extremely high level of DNA sequence polymorphism in promoter regions of the human olfactory receptor genes revealed using the 1000 Genomes Project dataset

    Directory of Open Access Journals (Sweden)

    Elena V. Ignatieva

    2014-03-01

    Full Text Available The molecular mechanism of olfactory cognition is very complicated. Olfactory cognition is initiated by olfactory receptor proteins (odorant receptors, which are activated by olfactory stimuli (ligands. Olfactory receptors are the initial player in the signal transduction cascade producing a nerve impulse, which is transmitted to the brain. The sensitivity to a particular ligand depends on the expression level of multiple proteins involved in the process of olfactory cognition: olfactory receptor proteins, proteins that participate in signal transduction cascade, etc. The expression level of each gene is controlled by its regulatory regions, and especially, by the promoter (a region of DNA about 100–1000 base pairs long located upstream of the transcription start site. We analyzed single nucleotide polymorphisms using human whole-genome data from the 1000 Genomes Project and revealed an extremely high level of single nucleotide polymorphisms in promoter regions of olfactory receptor genes and HLA genes. We hypothesized that the high level of polymorphisms in olfactory receptor promoters was responsible for the diversity in regulatory mechanisms controlling the expression levels of olfactory receptor proteins. Such diversity of regulatory mechanisms may cause the great variability of olfactory cognition of numerous environmental olfactory stimuli perceived by human beings (air pollutants, human body odors, odors in culinary etc.. In turn, this variability may provide a wide range of emotional and behavioral reactions related to the vast variety of olfactory stimuli.

  19. Formation of a Multiple Protein Complex on the Adenovirus Packaging Sequence by the IVa2 Protein▿

    OpenAIRE

    Tyler, Ryan E.; Ewing, Sean G.; Imperiale, Michael J.

    2007-01-01

    During adenovirus virion assembly, the packaging sequence mediates the encapsidation of the viral genome. This sequence is composed of seven functional units, termed A repeats. Recent evidence suggests that the adenovirus IVa2 protein binds the packaging sequence and is involved in packaging of the genome. Study of the IVa2-packaging sequence interaction has been hindered by difficulty in purifying the protein produced in virus-infected cells or by recombinant techniques. We report the first ...

  20. Next Generation Semiconductor Based Sequencing of the Donkey (Equus asinus) Genome Provided Comparative Sequence Data against the Horse Genome and a Few Millions of Single Nucleotide Polymorphisms

    Science.gov (United States)

    Bertolini, Francesca; Scimone, Concetta; Geraci, Claudia; Schiavo, Giuseppina; Utzeri, Valerio Joe; Chiofalo, Vincenzo; Fontanesi, Luca

    2015-01-01

    Few studies investigated the donkey (Equus asinus) at the whole genome level so far. Here, we sequenced the genome of two male donkeys using a next generation semiconductor based sequencing platform (the Ion Proton sequencer) and compared obtained sequence information with the available donkey draft genome (and its Illumina reads from which it was originated) and with the EquCab2.0 assembly of the horse genome. Moreover, the Ion Torrent Personal Genome Analyzer was used to sequence reduced representation libraries (RRL) obtained from a DNA pool including donkeys of different breeds (Grigio Siciliano, Ragusano and Martina Franca). The number of next generation sequencing reads aligned with the EquCab2.0 horse genome was larger than those aligned with the draft donkey genome. This was due to the larger N50 for contigs and scaffolds of the horse genome. Nucleotide divergence between E. caballus and E. asinus was estimated to be ~ 0.52-0.57%. Regions with low nucleotide divergence were identified in several autosomal chromosomes and in the whole chromosome X. These regions might be evolutionally important in equids. Comparing Y-chromosome regions we identified variants that could be useful to track donkey paternal lineages. Moreover, about 4.8 million of single nucleotide polymorphisms (SNPs) in the donkey genome were identified and annotated combining sequencing data from Ion Proton (whole genome sequencing) and Ion Torrent (RRL) runs with Illumina reads. A higher density of SNPs was present in regions homologous to horse chromosome 12, in which several studies reported a high frequency of copy number variants. The SNPs we identified constitute a first resource useful to describe variability at the population genomic level in E. asinus and to establish monitoring systems for the conservation of donkey genetic resources. PMID:26151450

  1. Next Generation Semiconductor Based Sequencing of the Donkey (Equus asinus Genome Provided Comparative Sequence Data against the Horse Genome and a Few Millions of Single Nucleotide Polymorphisms.

    Directory of Open Access Journals (Sweden)

    Francesca Bertolini

    Full Text Available Few studies investigated the donkey (Equus asinus at the whole genome level so far. Here, we sequenced the genome of two male donkeys using a next generation semiconductor based sequencing platform (the Ion Proton sequencer and compared obtained sequence information with the available donkey draft genome (and its Illumina reads from which it was originated and with the EquCab2.0 assembly of the horse genome. Moreover, the Ion Torrent Personal Genome Analyzer was used to sequence reduced representation libraries (RRL obtained from a DNA pool including donkeys of different breeds (Grigio Siciliano, Ragusano and Martina Franca. The number of next generation sequencing reads aligned with the EquCab2.0 horse genome was larger than those aligned with the draft donkey genome. This was due to the larger N50 for contigs and scaffolds of the horse genome. Nucleotide divergence between E. caballus and E. asinus was estimated to be ~ 0.52-0.57%. Regions with low nucleotide divergence were identified in several autosomal chromosomes and in the whole chromosome X. These regions might be evolutionally important in equids. Comparing Y-chromosome regions we identified variants that could be useful to track donkey paternal lineages. Moreover, about 4.8 million of single nucleotide polymorphisms (SNPs in the donkey genome were identified and annotated combining sequencing data from Ion Proton (whole genome sequencing and Ion Torrent (RRL runs with Illumina reads. A higher density of SNPs was present in regions homologous to horse chromosome 12, in which several studies reported a high frequency of copy number variants. The SNPs we identified constitute a first resource useful to describe variability at the population genomic level in E. asinus and to establish monitoring systems for the conservation of donkey genetic resources.

  2. Pooled Enrichment Sequencing Identifies Diversity and Evolutionary Pressures at NLR Resistance Genes within a Wild Tomato Population.

    Science.gov (United States)

    Stam, Remco; Scheikl, Daniela; Tellier, Aurélien

    2016-06-02

    Nod-like receptors (NLRs) are nucleotide-binding domain and leucine-rich repeats containing proteins that are important in plant resistance signaling. Many of the known pathogen resistance (R) genes in plants are NLRs and they can recognize pathogen molecules directly or indirectly. As such, divergence and copy number variants at these genes are found to be high between species. Within populations, positive and balancing selection are to be expected if plants coevolve with their pathogens. In order to understand the complexity of R-gene coevolution in wild nonmodel species, it is necessary to identify the full range of NLRs and infer their evolutionary history. Here we investigate and reveal polymorphism occurring at 220 NLR genes within one population of the partially selfing wild tomato species Solanum pennellii. We use a combination of enrichment sequencing and pooling ten individuals, to specifically sequence NLR genes in a resource and cost-effective manner. We focus on the effects which different mapping and single nucleotide polymorphism calling software and settings have on calling polymorphisms in customized pooled samples. Our results are accurately verified using Sanger sequencing of polymorphic gene fragments. Our results indicate that some NLRs, namely 13 out of 220, have maintained polymorphism within our S. pennellii population. These genes show a wide range of πN/πS ratios and differing site frequency spectra. We compare our observed rate of heterozygosity with expectations for this selfing and bottlenecked population. We conclude that our method enables us to pinpoint NLR genes which have experienced natural selection in their habitat. © The Author 2016. Published by Oxford University Press on behalf of the Society for Molecular Biology and Evolution.

  3. Discovering approximate-associated sequence patterns for protein-DNA interactions

    KAUST Repository

    Chan, Tak Ming

    2010-12-30

    Motivation: The bindings between transcription factors (TFs) and transcription factor binding sites (TFBSs) are fundamental protein-DNA interactions in transcriptional regulation. Extensive efforts have been made to better understand the protein-DNA interactions. Recent mining on exact TF-TFBS-associated sequence patterns (rules) has shown great potentials and achieved very promising results. However, exact rules cannot handle variations in real data, resulting in limited informative rules. In this article, we generalize the exact rules to approximate ones for both TFs and TFBSs, which are essential for biological variations. Results: A progressive approach is proposed to address the approximation to alleviate the computational requirements. Firstly, similar TFBSs are grouped from the available TF-TFBS data (TRANSFAC database). Secondly, approximate and highly conserved binding cores are discovered from TF sequences corresponding to each TFBS group. A customized algorithm is developed for the specific objective. We discover the approximate TF-TFBS rules by associating the grouped TFBS consensuses and TF cores. The rules discovered are evaluated by matching (verifying with) the actual protein-DNA binding pairs from Protein Data Bank (PDB) 3D structures. The approximate results exhibit many more verified rules and up to 300% better verification ratios than the exact ones. The customized algorithm achieves over 73% better verification ratios than traditional methods. Approximate rules (64-79%) are shown statistically significant. Detailed variation analysis and conservation verification on NCBI records demonstrate that the approximate rules reveal both the flexible and specific protein-DNA interactions accurately. The approximate TF-TFBS rules discovered show great generalized capability of exploring more informative binding rules. © The Author 2010. Published by Oxford University Press. All rights reserved.

  4. Large-scale analysis of intrinsic disorder flavors and associated functions in the protein sequence universe.

    Science.gov (United States)

    Necci, Marco; Piovesan, Damiano; Tosatto, Silvio C E

    2016-12-01

    Intrinsic disorder (ID) in proteins has been extensively described for the last decade; a large-scale classification of ID in proteins is mostly missing. Here, we provide an extensive analysis of ID in the protein universe on the UniProt database derived from sequence-based predictions in MobiDB. Almost half the sequences contain an ID region of at least five residues. About 9% of proteins have a long ID region of over 20 residues which are more abundant in Eukaryotic organisms and most frequently cover less than 20% of the sequence. A small subset of about 67,000 (out of over 80 million) proteins is fully disordered and mostly found in Viruses. Most proteins have only one ID, with short ID evenly distributed along the sequence and long ID overrepresented in the center. The charged residue composition of Das and Pappu was used to classify ID proteins by structural propensities and corresponding functional enrichment. Swollen Coils seem to be used mainly as structural components and in biosynthesis in both Prokaryotes and Eukaryotes. In Bacteria, they are confined in the nucleoid and in Viruses provide DNA binding function. Coils & Hairpins seem to be specialized in ribosome binding and methylation activities. Globules & Tadpoles bind antigens in Eukaryotes but are involved in killing other organisms and cytolysis in Bacteria. The Undefined class is used by Bacteria to bind toxic substances and mediate transport and movement between and within organisms in Viruses. Fully disordered proteins behave similarly, but are enriched for glycine residues and extracellular structures. © 2016 The Protein Society.

  5. Representation of protein-sequence information by amino acid subalphabets

    DEFF Research Database (Denmark)

    Andersen, C.A.F.; Brunak, Søren

    2004-01-01

    -sequence information, using machine learning strategies, where the primary goal is the discovery of novel powerful representations for use in AI techniques. In the case of proteins and the 20 different amino acids they typically contain, it is also a secondary goal to discover how the current selection of amino acids...

  6. Characterization and compilation of polymorphic simple sequence repeat (SSR markers of peanut from public database

    Directory of Open Access Journals (Sweden)

    Zhao Yongli

    2012-07-01

    Full Text Available Abstract Background There are several reports describing thousands of SSR markers in the peanut (Arachis hypogaea L. genome. There is a need to integrate various research reports of peanut DNA polymorphism into a single platform. Further, because of lack of uniformity in the labeling of these markers across the publications, there is some confusion on the identities of many markers. We describe below an effort to develop a central comprehensive database of polymorphic SSR markers in peanut. Findings We compiled 1,343 SSR markers as detecting polymorphism (14.5% within a total of 9,274 markers. Amongst all polymorphic SSRs examined, we found that AG motif (36.5% was the most abundant followed by AAG (12.1%, AAT (10.9%, and AT (10.3%.The mean length of SSR repeats in dinucleotide SSRs was significantly longer than that in trinucleotide SSRs. Dinucleotide SSRs showed higher polymorphism frequency for genomic SSRs when compared to trinucleotide SSRs, while for EST-SSRs, the frequency of polymorphic SSRs was higher in trinucleotide SSRs than in dinucleotide SSRs. The correlation of the length of SSR and the frequency of polymorphism revealed that the frequency of polymorphism was decreased as motif repeat number increased. Conclusions The assembled polymorphic SSRs would enhance the density of the existing genetic maps of peanut, which could also be a useful source of DNA markers suitable for high-throughput QTL mapping and marker-assisted selection in peanut improvement and thus would be of value to breeders.

  7. Electrophoretic protein polymorphisms in Kaingang and Guarani Indians of Southern Brazil.

    Science.gov (United States)

    Salzano, Francisco M; Callegari-Jacques, Sidia M; Weimer, Tania A; Franco, Maria H L P; Hutz, Mara H; Petzl-Erler, Maria L

    1997-01-01

    A total of 337 Kaingang and Guarani Indians from two localities were studied in relation to 18 protein genetic loci. In one of the localities, members of these two groups live side by side but show little genetic similarity, emphasizing the influence of cultural factors in the mating behavior of human groups. Integrating the present results with previous ones, it was verified that the genetic relationships among six Kaingang populations do not follow the pattern expected from their geographical distribution. Comparisons made with three other Ge˜-speaking tribes indicate that the Kaingang did not separate well from them. Most (96%) of the variability in the six polymorphic systems considered occur at the intrapopulational level. Am. J. Hum. Biol. 9:505-512, 1997. © 1997 Wiley-Liss, Inc. Copyright © 1997 Wiley-Liss, Inc.

  8. RStrucFam: a web server to associate structure and cognate RNA for RNA-binding proteins from sequence information.

    Science.gov (United States)

    Ghosh, Pritha; Mathew, Oommen K; Sowdhamini, Ramanathan

    2016-10-07

    RNA-binding proteins (RBPs) interact with their cognate RNA(s) to form large biomolecular assemblies. They are versatile in their functionality and are involved in a myriad of processes inside the cell. RBPs with similar structural features and common biological functions are grouped together into families and superfamilies. It will be useful to obtain an early understanding and association of RNA-binding property of sequences of gene products. Here, we report a web server, RStrucFam, to predict the structure, type of cognate RNA(s) and function(s) of proteins, where possible, from mere sequence information. The web server employs Hidden Markov Model scan (hmmscan) to enable association to a back-end database of structural and sequence families. The database (HMMRBP) comprises of 437 HMMs of RBP families of known structure that have been generated using structure-based sequence alignments and 746 sequence-centric RBP family HMMs. The input protein sequence is associated with structural or sequence domain families, if structure or sequence signatures exist. In case of association of the protein with a family of known structures, output features like, multiple structure-based sequence alignment (MSSA) of the query with all others members of that family is provided. Further, cognate RNA partner(s) for that protein, Gene Ontology (GO) annotations, if any and a homology model of the protein can be obtained. The users can also browse through the database for details pertaining to each family, protein or RNA and their related information based on keyword search or RNA motif search. RStrucFam is a web server that exploits structurally conserved features of RBPs, derived from known family members and imprinted in mathematical profiles, to predict putative RBPs from sequence information. Proteins that fail to associate with such structure-centric families are further queried against the sequence-centric RBP family HMMs in the HMMRBP database. Further, all other essential

  9. Domain fusion analysis by applying relational algebra to protein sequence and domain databases.

    Science.gov (United States)

    Truong, Kevin; Ikura, Mitsuhiko

    2003-05-06

    Domain fusion analysis is a useful method to predict functionally linked proteins that may be involved in direct protein-protein interactions or in the same metabolic or signaling pathway. As separate domain databases like BLOCKS, PROSITE, Pfam, SMART, PRINTS-S, ProDom, TIGRFAMs, and amalgamated domain databases like InterPro continue to grow in size and quality, a computational method to perform domain fusion analysis that leverages on these efforts will become increasingly powerful. This paper proposes a computational method employing relational algebra to find domain fusions in protein sequence databases. The feasibility of this method was illustrated on the SWISS-PROT+TrEMBL sequence database using domain predictions from the Pfam HMM (hidden Markov model) database. We identified 235 and 189 putative functionally linked protein partners in H. sapiens and S. cerevisiae, respectively. From scientific literature, we were able to confirm many of these functional linkages, while the remainder offer testable experimental hypothesis. Results can be viewed at http://calcium.uhnres.utoronto.ca/pi. As the analysis can be computed quickly on any relational database that supports standard SQL (structured query language), it can be dynamically updated along with the sequence and domain databases, thereby improving the quality of predictions over time.

  10. Whole genome sequencing options for bacterial strain typing and epidemiologic analysis based on single nucleotide polymorphism versus gene-by-gene-based approaches.

    Science.gov (United States)

    Schürch, A C; Arredondo-Alonso, S; Willems, R J L; Goering, R V

    2018-04-01

    Whole genome sequence (WGS)-based strain typing finds increasing use in the epidemiologic analysis of bacterial pathogens in both public health as well as more localized infection control settings. This minireview describes methodologic approaches that have been explored for WGS-based epidemiologic analysis and considers the challenges and pitfalls of data interpretation. Personal collection of relevant publications. When applying WGS to study the molecular epidemiology of bacterial pathogens, genomic variability between strains is translated into measures of distance by determining single nucleotide polymorphisms in core genome alignments or by indexing allelic variation in hundreds to thousands of core genes, assigning types to unique allelic profiles. Interpreting isolate relatedness from these distances is highly organism specific, and attempts to establish species-specific cutoffs are unlikely to be generally applicable. In cases where single nucleotide polymorphism or core gene typing do not provide the resolution necessary for accurate assessment of the epidemiology of bacterial pathogens, inclusion of accessory gene or plasmid sequences may provide the additional required discrimination. As with all epidemiologic analysis, realizing the full potential of the revolutionary advances in WGS-based approaches requires understanding and dealing with issues related to the fundamental steps of data generation and interpretation. Copyright © 2018 The Authors. Published by Elsevier Ltd.. All rights reserved.

  11. ALGEBRAIC STRUCTURES AND SYNTHESIS PROTEIN: A DESCRIPTION OF GENE CAPN10

    Directory of Open Access Journals (Sweden)

    Obidio Rubio

    2016-06-01

    Full Text Available This article mainly informative, some ways of algebraic modeling of human genome sequences is presented, with special emphasis on describing mutations in the genes, which by modifying protein synthesis involve genetic diseases such as Diabetes Mellitus. Mutations as endomorphisms on an R-module, which consists of a direct sum of groups of sequences 2q37.3 gene, on the rings Z64 and Z125, where el haplotype compound for the polymorphisms SNP43, SNP19 and SNP63 occur.

  12. Genotyping of major histocompatibility complex Class II DRB gene in Rohilkhandi goats by polymerase chain reaction-restriction fragment length polymorphism and DNA sequencing

    Directory of Open Access Journals (Sweden)

    Kush Shrivastava

    2015-10-01

    Full Text Available Aim: To study the major histocompatibility complex (MHC Class II DRB1 gene polymorphism in Rohilkhandi goat using polymerase chain reaction-restriction fragment length polymorphism (PCR-RFLP and nucleotide sequencing techniques. Materials and Methods: DNA was isolated from 127 Rohilkhandi goats maintained at sheep and goat farm, Indian Veterinary Research Institute, Izatnagar, Bareilly. A 284 bp fragment of exon 2 of DRB1 gene was amplified and digested using BsaI and TaqI restriction enzymes. Population genetic parameters were calculated using Popgene v 1.32 and SAS 9.0. The genotypes were then sequenced using Sanger dideoxy chain termination method and were compared with related breeds/species using MEGA 6.0 and Megalign (DNASTAR software. Results: TaqI locus showed three and BsaI locus showed two genotypes. Both the loci were found to be in Hardy–Weinberg equilibrium (HWE, however, population genetic parameters suggest that heterozygosity is still maintained in the population at both loci. Percent diversity and divergence matrix, as well as phylogenetic analysis revealed that the MHC Class II DRB1 gene of Rohilkhandi goats was found to be in close cluster with Garole and Scottish blackface sheep breeds as compared to other goat breeds included in the sequence comparison. Conclusion: The PCR-RFLP patterns showed population to be in HWE and absence of one genotype at one locus (BsaI, both the loci showed excess of one or the other homozygote genotype, however, effective number of alleles showed that allelic diversity is present in the population. Sequence comparison of DRB1 gene of Rohilkhandi goat with other sheep and goat breed assigned Rohilkhandi goat in divergence with Jamanupari and Angora goats.

  13. Adaptive GDDA-BLAST: fast and efficient algorithm for protein sequence embedding.

    Directory of Open Access Journals (Sweden)

    Yoojin Hong

    2010-10-01

    Full Text Available A major computational challenge in the genomic era is annotating structure/function to the vast quantities of sequence information that is now available. This problem is illustrated by the fact that most proteins lack comprehensive annotations, even when experimental evidence exists. We previously theorized that embedded-alignment profiles (simply "alignment profiles" hereafter provide a quantitative method that is capable of relating the structural and functional properties of proteins, as well as their evolutionary relationships. A key feature of alignment profiles lies in the interoperability of data format (e.g., alignment information, physio-chemical information, genomic information, etc.. Indeed, we have demonstrated that the Position Specific Scoring Matrices (PSSMs are an informative M-dimension that is scored by quantitatively measuring the embedded or unmodified sequence alignments. Moreover, the information obtained from these alignments is informative, and remains so even in the "twilight zone" of sequence similarity (<25% identity. Although our previous embedding strategy was powerful, it suffered from contaminating alignments (embedded AND unmodified and high computational costs. Herein, we describe the logic and algorithmic process for a heuristic embedding strategy named "Adaptive GDDA-BLAST." Adaptive GDDA-BLAST is, on average, up to 19 times faster than, but has similar sensitivity to our previous method. Further, data are provided to demonstrate the benefits of embedded-alignment measurements in terms of detecting structural homology in highly divergent protein sequences and isolating secondary structural elements of transmembrane and ankyrin-repeat domains. Together, these advances allow further exploration of the embedded alignment data space within sufficiently large data sets to eventually induce relevant statistical inferences. We show that sequence embedding could serve as one of the vehicles for measurement of low

  14. Generalizing and learning protein-DNA binding sequence representations by an evolutionary algorithm

    KAUST Repository

    Wong, Ka Chun

    2011-02-05

    Protein-DNA bindings are essential activities. Understanding them forms the basis for further deciphering of biological and genetic systems. In particular, the protein-DNA bindings between transcription factors (TFs) and transcription factor binding sites (TFBSs) play a central role in gene transcription. Comprehensive TF-TFBS binding sequence pairs have been found in a recent study. However, they are in one-to-one mappings which cannot fully reflect the many-to-many mappings within the bindings. An evolutionary algorithm is proposed to learn generalized representations (many-to-many mappings) from the TF-TFBS binding sequence pairs (one-to-one mappings). The generalized pairs are shown to be more meaningful than the original TF-TFBS binding sequence pairs. Some representative examples have been analyzed in this study. In particular, it shows that the TF-TFBS binding sequence pairs are not presumably in one-to-one mappings. They can also exhibit many-to-many mappings. The proposed method can help us extract such many-to-many information from the one-to-one TF-TFBS binding sequence pairs found in the previous study, providing further knowledge in understanding the bindings between TFs and TFBSs. © 2011 Springer-Verlag.

  15. Generalizing and learning protein-DNA binding sequence representations by an evolutionary algorithm

    KAUST Repository

    Wong, Ka Chun; Peng, Chengbin; Wong, Manhon; Leung, Kwongsak

    2011-01-01

    Protein-DNA bindings are essential activities. Understanding them forms the basis for further deciphering of biological and genetic systems. In particular, the protein-DNA bindings between transcription factors (TFs) and transcription factor binding sites (TFBSs) play a central role in gene transcription. Comprehensive TF-TFBS binding sequence pairs have been found in a recent study. However, they are in one-to-one mappings which cannot fully reflect the many-to-many mappings within the bindings. An evolutionary algorithm is proposed to learn generalized representations (many-to-many mappings) from the TF-TFBS binding sequence pairs (one-to-one mappings). The generalized pairs are shown to be more meaningful than the original TF-TFBS binding sequence pairs. Some representative examples have been analyzed in this study. In particular, it shows that the TF-TFBS binding sequence pairs are not presumably in one-to-one mappings. They can also exhibit many-to-many mappings. The proposed method can help us extract such many-to-many information from the one-to-one TF-TFBS binding sequence pairs found in the previous study, providing further knowledge in understanding the bindings between TFs and TFBSs. © 2011 Springer-Verlag.

  16. SPiCE : A web-based tool for sequence-based protein classification and exploration

    NARCIS (Netherlands)

    Van den Berg, B.A.; Reinders, M.J.; Roubos, J.A.; De Ridder, D.

    2014-01-01

    Background Amino acid sequences and features extracted from such sequences have been used to predict many protein properties, such as subcellular localization or solubility, using classifier algorithms. Although software tools are available for both feature extraction and classifier construction,

  17. Nucleotide polymorphisms and haplotype diversity of RTCS gene in China elite maize inbred lines.

    Directory of Open Access Journals (Sweden)

    Enying Zhang

    Full Text Available The maize RTCS gene, encoding a LOB domain transcription factor, plays important roles in the initiation of embryonic seminal and postembryonic shoot-borne root. In this study, the genomic sequences of this gene in 73 China elite inbred lines, including 63 lines from 5 temperate heteroric groups and 10 tropic germplasms, were obtained, and the nucleotide polymorphisms and haplotype diversity were detected. A total of 63 sequence variants, including 44 SNPs and 19 indels, were identified at this locus, and most of them were found to be located in the regions of UTR and intron. The coding region of this gene in all tested inbred lines carried 14 haplotypes, which encoding 7 deferring RTCS proteins. Analysis of the polymorphism sites revealed that at least 6 recombination events have occurred. Among all 6 groups tested, only the P heterotic group had a much lower nucleotide diversity than the whole set, and selection analysis also revealed that only this group was under strong negative selection. However, the set of Huangzaosi and its derived lines possessed a higher nucleotide diversity than the whole set, and no selection signal were identified.

  18. GENETIC POLYMORPHISM OF SOME PROTEINS IN THE MILK OF CARPATHIAN GOAT

    Directory of Open Access Journals (Sweden)

    MIHAELA ZAULET

    2008-05-01

    Full Text Available The paper presents some aspects of the polymorphism of some Carpathian goat milk proteins. The Carpathian breed is the main breed of goats reared in Romania. The optimal working conditions were determined for the identification of the casein phenotypes. The technique of the polyacrylamide gel electrophoresis was used. The milk samples were processed to remove the fat and whey and the migration was done in the presence of a standard sample which contained proteins with different molecular weights. The interpretation of the electrophoresis migrations revealed the presence of two genotypes, the homozygous genotype β-Cn BB and the heterozygous β-Cn AB. The homozygous genotype β-Cn AA was not identified in any individual. The heterozygous genotype β-Cn AB displayed a high frequency (59% and it was observed in 10 individuals. The homozygous genotype β-Cn BB was observed in 7 individuals and it had a frequency of 41%. The homozygous genotype β-Cn AA has not been identified in the studied population. The distribution of these genotypes showed that allele β-CnB was predominant (70% over allele β-can (30%.

  19. Synaptosomal-associated protein 25 gene polymorphisms and antisocial personality disorder: association with temperament and psychopathy.

    Science.gov (United States)

    Basoglu, Cengiz; Oner, Ozgur; Ates, Alpay; Algul, Ayhan; Bez, Yasin; Cetin, Mesut; Herken, Hasan; Erdal, Mehmet Emin; Munir, Kerim M

    2011-06-01

    The molecular genetic of personality disorders has been investigated in several studies; however, the association of antisocial behaviours with synaptosomal-associated protein 25 (SNAP25) gene polymorphisms has not. This association is of interest as SNAP25 gene polymorphism has been associated with attention-deficit hyperactivity disorder and personality. We compared the distribution of DdeI and MnII polymorphisms in 91 young male offenders and in 38 sex-matched healthy control subjects. We also investigated the association of SNAP25 gene polymorphisms with severity of psychopathy and with temperament traits: novelty seeking, harm avoidance, and reward dependence. The MnII T/T and DdeI T/T genotypes were more frequently present in male subjects with antisocial personality disorder (APD) than in sex-matched healthy control subjects. The association was stronger when the frequency of both DdeI and MnII T/T were taken into account. In the APD group, the genotype was not significantly associated with the Psychopathy Checklist-Revised scores, measuring the severity of psychopathy. However, the APD subjects with the MnII T/T genotype had higher novelty seeking scores; whereas, subjects with the DdeI T/T genotype had lower reward dependence scores. Again, the association between genotype and novelty seeking was stronger when both DdeI and MnII genotypes were taken into account. DdeI and MnII T/T genotypes may be a risk factor for antisocial behaviours. The association of the SNAP25 DdeI T/T and MnII T/T genotypes with lower reward dependence and higher novelty seeking suggested that SNAP25 genotype might influence other personality disorders, as well.

  20. The Population Genomics of Sunflowers and Genomic Determinants of Protein Evolution Revealed by RNAseq

    Directory of Open Access Journals (Sweden)

    Loren H. Rieseberg

    2012-10-01

    Full Text Available Few studies have investigated the causes of evolutionary rate variation among plant nuclear genes, especially in recently diverged species still capable of hybridizing in the wild. The recent advent of Next Generation Sequencing (NGS permits investigation of genome wide rates of protein evolution and the role of selection in generating and maintaining divergence. Here, we use individual whole-transcriptome sequencing (RNAseq to refine our understanding of the population genomics of wild species of sunflowers (Helianthus spp. and the factors that affect rates of protein evolution. We aligned 35 GB of transcriptome sequencing data and identified 433,257 polymorphic sites (SNPs in a reference transcriptome comprising 16,312 genes. Using SNP markers, we identified strong population clustering largely corresponding to the three species analyzed here (Helianthus annuus, H. petiolaris, H. debilis, with one distinct early generation hybrid. Then, we calculated the proportions of adaptive substitution fixed by selection (alpha and identified gene ontology categories with elevated values of alpha. The “response to biotic stimulus” category had the highest mean alpha across the three interspecific comparisons, implying that natural selection imposed by other organisms plays an important role in driving protein evolution in wild sunflowers. Finally, we examined the relationship between protein evolution (dN/dS ratio and several genomic factors predicted to co-vary with protein evolution (gene expression level, divergence and specificity, genetic divergence [FST], and nucleotide diversity pi. We find that variation in rates of protein divergence was correlated with gene expression level and specificity, consistent with results from a broad range of taxa and timescales. This would in turn imply that these factors govern protein evolution both at a microevolutionary and macroevolutionary timescale. Our results contribute to a general understanding of the

  1. A new polymorphism in the GRP78 is not associated with HBV invasion

    Science.gov (United States)

    Zhu, Xiao; Wang, Yi; Tao, Tao; Li, Dong-Pei; Lan, Fei-Fei; Zhu, Wei; Xie, Dan; Kung, Hsiang-Fu

    2009-01-01

    AIM: To examine the association between -86 bp (T > A) in the glucose-regulated protein 78 gene (GRP78) and hepatitis B virus (HBV) invasion. METHODS: DNA was genotyped for the single-nucleotide polymorphism by polymerase chain reaction followed by sequencing in a sample of 382 unrelated HBV carriers and a total of 350 sex- and age-matched healthy controls. Serological markers for HBV infection were determined by enzyme-linked immunosorbent assay kits or clinical chemistry testing. RESULTS: The distributions of allelotype and genotype in cases were not significantly different from those in controls. In addition, our findings suggested that neither alanine aminotransferase/hepatitis B e antigen nor HBV-DNA were associated with the allele/genotype variation in HBV infected individuals. CONCLUSION: -86 bp T > A polymorphism in GRP78 gene is not related to the clinical risk and acute exacerbation of HBV invasion. PMID:19842229

  2. Rapid detection, classification and accurate alignment of up to a million or more related protein sequences.

    Science.gov (United States)

    Neuwald, Andrew F

    2009-08-01

    The patterns of sequence similarity and divergence present within functionally diverse, evolutionarily related proteins contain implicit information about corresponding biochemical similarities and differences. A first step toward accessing such information is to statistically analyze these patterns, which, in turn, requires that one first identify and accurately align a very large set of protein sequences. Ideally, the set should include many distantly related, functionally divergent subgroups. Because it is extremely difficult, if not impossible for fully automated methods to align such sequences correctly, researchers often resort to manual curation based on detailed structural and biochemical information. However, multiply-aligning vast numbers of sequences in this way is clearly impractical. This problem is addressed using Multiply-Aligned Profiles for Global Alignment of Protein Sequences (MAPGAPS). The MAPGAPS program uses a set of multiply-aligned profiles both as a query to detect and classify related sequences and as a template to multiply-align the sequences. It relies on Karlin-Altschul statistics for sensitivity and on PSI-BLAST (and other) heuristics for speed. Using as input a carefully curated multiple-profile alignment for P-loop GTPases, MAPGAPS correctly aligned weakly conserved sequence motifs within 33 distantly related GTPases of known structure. By comparison, the sequence- and structurally based alignment methods hmmalign and PROMALS3D misaligned at least 11 and 23 of these regions, respectively. When applied to a dataset of 65 million protein sequences, MAPGAPS identified, classified and aligned (with comparable accuracy) nearly half a million putative P-loop GTPase sequences. A C++ implementation of MAPGAPS is available at http://mapgaps.igs.umaryland.edu. Supplementary data are available at Bioinformatics online.

  3. A new polymorphic and multicopy MHC gene family related to nonmammalian class I

    Energy Technology Data Exchange (ETDEWEB)

    Leelayuwat, C.; Degli-Esposti, M.A.; Abraham, L.J. [Univ. of Western Australia, Perth (Australia); Townend, D.C. [Sir Charles Gairdner Hospital, Perth (Australia); Dawkins, R.L. [Royal Perth Hospital, Perth (Australia)]|[Univ. of Western Australia, Perth (Australia)]|[Sir Charles Gairdner Hospital, Perth (Australia)

    1994-12-31

    The authors have used genomic analysis to characterize a region of the central major histocompatibility complex (MHC) spanning {approximately} 300 kilobases (kb) between TNF and HLA-B. This region has been suggested to carry genetic factors relevant to the development of autoimmune diseases such as myasthenia gravis (MG) and insulin dependent diabetes mellitus (IDDM). Genomic sequence was analyzed for coding potential, using two neural network programs, GRAIL and GeneParser. A genomic probe, JAB, containing putative coding sequences (PERB11) located 60 kb centromeric of HLA-B, was used for northern analysis of human tissues. Multiple transcripts were detected. Southern analysis of genomic DNA and overlapping YAC clones, covering the region from BAT1 to HLA-F, indicated that there are at least five copies of PERB11, four of which are located within this region of the MHC. The partial cDNA sequence of PERB11 was obtained from poly-A RNA derived from skeletal muscle. The putative amino acid sequence of PERB11 shares {approximately} 30% identity to MHC class I molecules from various species, including reptiles, chickens, and frogs, as well as to other MHC class I-like molecules, such as the IgG FcR of the mouse and rat and the human Zn-{alpha}2-glycoprotein. From direct comparison of amino acid sequences, it is concluded that PERB11 is a distinct molecule more closely related to nonmammalian than known mammalian MHC class I molecules. Genomic sequence analysis of PERB11 from five MHC ancestral haplotypes (AH) indicated that the gene is polymorphic at both DNA and protein level. The results suggest that the authors have identified a novel polymorphic gene family with multiple copies within the MHC. 48 refs., 10 figs., 2 tabs.

  4. dPORE-miRNA: Polymorphic regulation of microRNA genes

    KAUST Repository

    Schmeier, Sebastian; Schaefer, Ulf; MacPherson, Cameron R.; Bajic, Vladimir B.

    2011-01-01

    Background: MicroRNAs (miRNAs) are short non-coding RNA molecules that act as post-transcriptional regulators and affect the regulation of protein-coding genes. Mostly transcribed by PolII, miRNA genes are regulated at the transcriptional level similarly to protein-coding genes. In this study we focus on human miRNAs. These miRNAs are involved in a variety of pathways and can affect many diseases. Our interest is on possible deregulation of the transcription initiation of the miRNA encoding genes, which is facilitated by variations in the genomic sequence of transcriptional control regions (promoters). Methodology: Our aim is to provide an online resource to facilitate the investigation of the potential effects of single nucleotide polymorphisms (SNPs) on miRNA gene regulation. We analyzed SNPs overlapped with predicted transcription factor binding sites (TFBSs) in promoters of miRNA genes. We also accounted for the creation of novel TFBSs due to polymorphisms not present in the reference genome. The resulting changes in the original TFBSs and potential creation of new TFBSs were incorporated into the Dragon Database of Polymorphic Regulation of miRNA genes (dPORE-miRNA). Conclusions: The dPORE-miRNA database enables researchers to explore potential effects of SNPs on the regulation of miRNAs. dPORE-miRNA can be interrogated with regards to: a/miRNAs (their targets, or involvement in diseases, or biological pathways), b/SNPs, or c/transcription factors. dPORE-miRNA can be accessed at http://cbrc.kaust.edu.sa/dpore and http://apps.sanbi.ac.za/dpore/. Its use is free for academic and non-profit users. © 2011 Schmeier et al.

  5. dPORE-miRNA: Polymorphic regulation of microRNA genes

    KAUST Repository

    Schmeier, Sebastian

    2011-02-04

    Background: MicroRNAs (miRNAs) are short non-coding RNA molecules that act as post-transcriptional regulators and affect the regulation of protein-coding genes. Mostly transcribed by PolII, miRNA genes are regulated at the transcriptional level similarly to protein-coding genes. In this study we focus on human miRNAs. These miRNAs are involved in a variety of pathways and can affect many diseases. Our interest is on possible deregulation of the transcription initiation of the miRNA encoding genes, which is facilitated by variations in the genomic sequence of transcriptional control regions (promoters). Methodology: Our aim is to provide an online resource to facilitate the investigation of the potential effects of single nucleotide polymorphisms (SNPs) on miRNA gene regulation. We analyzed SNPs overlapped with predicted transcription factor binding sites (TFBSs) in promoters of miRNA genes. We also accounted for the creation of novel TFBSs due to polymorphisms not present in the reference genome. The resulting changes in the original TFBSs and potential creation of new TFBSs were incorporated into the Dragon Database of Polymorphic Regulation of miRNA genes (dPORE-miRNA). Conclusions: The dPORE-miRNA database enables researchers to explore potential effects of SNPs on the regulation of miRNAs. dPORE-miRNA can be interrogated with regards to: a/miRNAs (their targets, or involvement in diseases, or biological pathways), b/SNPs, or c/transcription factors. dPORE-miRNA can be accessed at http://cbrc.kaust.edu.sa/dpore and http://apps.sanbi.ac.za/dpore/. Its use is free for academic and non-profit users. © 2011 Schmeier et al.

  6. [Effect of FABP2 gene G54A polymorphism on lipid and glucose metabolism in simple obesity children].

    Science.gov (United States)

    Xu, Yunpeng; Rao, Xiaojiao; Hao, Min; Hou, Lijuan; Zhu, Xiaobo; Chang, Xiaotong

    2016-01-01

    To explore the relationship between intestinal fatty acid binding protein (FABP2) gene G54A polymorphism and simple childhood obesity, the effect of mutant 54A FABP2 gene on serum lipids and glucose metabolism. The total of 83 subjects with overweight/obesity and 100 subjects with healthy/normal weight were involved in this study. The G54A FABP2 gene allele and genotype frequencies between control group and overweight/obesity group were detected using polymerase chain reaction (PCR) -restriction fragment length polymorphism (RFLP) technology, and DNA sequences were confirmed by DNA sequencing. The automatic biochemical analyzer was used to detect fasting blood glucose (FBG), triglyceride (TG), total cholesterol (TC), high-density lipoprotein cholesterol (HDL-C) and low-density lipoprotein cholesterol (LDL-C) levels. Plasma insulin (Ins) was detected by radiation immune method, free fatty acids (FFA) was tested by ELISA method, insulin resistance index ( HOMA-IR ) was also calculated. The correlation between FABP2 G54A polymorphism and the development of children' obesity was analyzed. The relation between FABP2 G54A polymorphism and abnormal blood lipid and insulin resistance was assessed. The results of study on FABP2 gene polymorphism revealed as followed. In overweight/obese groups, the frequencies of GG, GA, AA genotypes was 33.7%, 49.4% and 16.9%, respectively. In control group, the frequencies of GG, GA, AA genotypes was 51. 0% , 40. 0% and 9. 0% , respectively. The differences between two groups was statistically significant (Χ2 = 6.27, P 0.05). The FABP2 gene G54A polymorphism is related to simple children obesity and lipid metabolism abnormality. The allele encoding in FABP2 gene may be a potential factor contributing to promoting lipid metabolism abnormality of and insulin resistance.

  7. The YPLGVG sequence of the Nipah virus matrix protein is required for budding

    Directory of Open Access Journals (Sweden)

    Yan Lianying

    2008-11-01

    Full Text Available Abstract Background Nipah virus (NiV is a recently emerged paramyxovirus capable of causing fatal disease in a broad range of mammalian hosts, including humans. Together with Hendra virus (HeV, they comprise the genus Henipavirus in the family Paramyxoviridae. Recombinant expression systems have played a crucial role in studying the cell biology of these Biosafety Level-4 restricted viruses. Henipavirus assembly and budding occurs at the plasma membrane, although the details of this process remain poorly understood. Multivesicular body (MVB proteins have been found to play a role in the budding of several enveloped viruses, including some paramyxoviruses, and the recruitment of MVB proteins by viral proteins possessing late budding domains (L-domains has become an important concept in the viral budding process. Previously we developed a system for producing NiV virus-like particles (VLPs and demonstrated that the matrix (M protein possessed an intrinsic budding ability and played a major role in assembly. Here, we have used this system to further explore the budding process by analyzing elements within the M protein that are critical for particle release. Results Using rationally targeted site-directed mutagenesis we show that a NiV M sequence YPLGVG is required for M budding and that mutation or deletion of the sequence abrogates budding ability. Replacement of the native and overlapping Ebola VP40 L-domains with the NiV sequence failed to rescue VP40 budding; however, it did induce the cellular morphology of extensive filamentous projection consistent with wild-type VP40-expressing cells. Cells expressing wild-type NiV M also displayed this morphology, which was dependent on the YPLGVG sequence, and deletion of the sequence also resulted in nuclear localization of M. Dominant-negative VPS4 proteins had no effect on NiV M budding, suggesting that unlike other viruses such as Ebola, NiV M accomplishes budding independent of MVB cellular proteins

  8. PPARγ gene polymorphism, C-reactive protein level, BMI and periodontitis in post-menopausal Japanese women.

    Science.gov (United States)

    Wang, Yangming; Sugita, Noriko; Yoshihara, Akihiro; Iwasaki, Masanori; Miyazaki, Hideo; Nakamura, Kazutoshi; Yoshie, Hiromasa

    2016-03-01

    Several studies have reported inconsistent results regarding the association between the PPARγPro12Ala polymorphism and obesity. Obese individuals had higher C-reactive protein (CRP) levels compared with those of normal weight, and PPARγ activation could significantly reduce serum high-sensitive CRP level. We have previously suggested that the Pro12Ala polymorphism represents a susceptibility factor for periodontitis, which is a known risk factor for increased CRP level. The aim was to investigate associations between PPARγ gene polymorphism, serum CRP level, BMI and/or periodontitis among post-menopausal Japanese women. The final sample in this study comprised 359 post-menopausal Japanese women. Periodontal parameters, including PD, CAL and BOP, were measured per tooth. PPARγPro12Ala genotype was determined by PCR-RFLP. Hs-CRP value was measured by a latex nephelometry assay. No significant differences in age, BMI or periodontal parameters were found between the genotypes. The percentages of sites with PD ≥ 4 mm were significantly higher among the hsCRP ≥ 1 mg/l group than the hsCRP periodontitis, serum CRP level or BMI in post-menopausal Japanese women. However, serum hsCRP level correlated with periodontitis in Ala allele carriers, and with BMI in non-carriers. © 2014 John Wiley & Sons A/S and The Gerodontology Society. Published by John Wiley & Sons Ltd.

  9. Comprehensive analysis of Salmonella sequence polymorphisms and development of a LDR-UA assay for the detection and characterization of selected serotypes.

    Science.gov (United States)

    Lauri, Andrea; Castiglioni, Bianca; Mariani, Paola

    2011-07-01

    Salmonella is a major cause of food-borne disease, and Salmonella enterica subspecies I includes the most clinically relevant serotypes. Salmonella serotype determination is important for the disease etiology assessment and contamination source tracking. This task will be facilitated by the disclosure of Salmonella serotype sequence polymorphisms, here annotated in seven genes (sefA, safA, safC, bigA, invA, fimA, and phsB) from 139 S. enterica strains, of which 109 belonging to 44 serotypes of subsp. I. One hundred nineteen polymorphic sites were scored and associated to single serotypes or to serotype groups belonging to S. enterica subsp. I. A diagnostic tool was constructed based on the Ligation Detection Reaction-Universal Array (LDR-UA) for the detection of polymorphic sites uniquely associated to serotypes of primary interest (Salmonella Hadar, Salmonella Infantis, Salmonella Enteritidis, Salmonella Typhimurium, Salmonella Gallinarum, Salmonella Virchow, and Salmonella Paratyphi B). The implementation of promiscuous probes allowed the diagnosis of ten further serotypes that could be associated to a unique hybridization pattern. Finally, the sensitivity and applicability of the tool was tested on target DNA dilutions and with controlled meat contamination, allowing the detection of one Salmonella CFU in 25 g of meat.

  10. [Polymorphisms of mitochondrial DNA hypervariable regions HVR I and HVR II in Changdu Tibetan in China].

    Science.gov (United States)

    Zhao, Jianmin; Kang, Longli; Bian, Liqiang; La, Zong

    2008-10-01

    To analyze the sequence polymorphisms of mitochondrial DNA HVR I and HVR II in Tibetan population in Changdu area of Tibet. mtDNAs obtained from 97 unrelated individuals were amplified and directly sequenced. One hundred and eleven variable sites were identified, including nucleotide transitions, transversions, insertions and deletions. In HVR I region (nt16024-nt16365), sixty-eight polymorphic sites and 92 haplotypes were observed, and the genetic diversity was 0.9985. In HVR II region (nt73-nt340), forty-three polymorphic sites and 91 haplotypes were detected, and the genetic diversity was 0.9882. The random match probability of HVR I and HVR II regions were 0.0120 and 0.0118, respectively. When the sequence analysis of HVR I and HVR II regions were combined, ninety-seven different haplotypes were found. The combined match probability of two unrelated persons having the same sequence was 0.0103. There are some unique polymorphic loci in the Changdu Tibetan population. The results suggest that there are significant difference in the genetic structure in the mitochondrial DNA D-loop region between Changdu Tibetans and other Asian populations and Caucasians. Sequence polymorphism in mitochondrial DNA HVR I and HVR II can be used as a genetic marker for forensic individual identification and genetic analysis.

  11. HLA-G polymorphisms in couples with recurrent spontaneous abortions

    DEFF Research Database (Denmark)

    Hviid, T V; Hylenius, S; Hoegh, A M

    2002-01-01

    % of the RSA women carried the HLA-G*0106 allele compared to 2% of the control women. The 14 bp deletion polymorphism in exon 8 was investigated separately. There were a greater number of heterozygotes for the 14 bp polymorphism in the group of fertile control women than expected, according to Hardy-Weinberg...... equilibrium. Furthermore, the HLA-G alleles without the 14 bp sequence were prominent in the RSA males in contrast to the RSA women in whom alleles including the 14 bp sequence were frequently observed, especially as homozygotes. These results are discussed in relation to two hypotheses concerning HLA...

  12. Exploiting Illumina Sequencing for the Development of 95 Novel Polymorphic EST-SSR Markers in Common Vetch (Vicia sativa subsp. sativa

    Directory of Open Access Journals (Sweden)

    Zhipeng Liu

    2014-05-01

    Full Text Available The common vetch (Vicia sativa subsp. sativa, a self-pollinating and diploid species, is one of the most important annual legumes in the world due to its short growth period, high nutritional value, and multiple usages as hay, grain, silage, and green manure. The available simple sequence repeat (SSR markers for common vetch, however, are insufficient to meet the developing demand for genetic and molecular research on this important species. Here, we aimed to develop and characterise several polymorphic EST-SSR markers from the vetch Illumina transcriptome. A total number of 1,071 potential EST-SSR markers were identified from 1025 unigenes whose lengths were greater than 1,000 bp, and 450 primer pairs were then designed and synthesized. Finally, 95 polymorphic primer pairs were developed for the 10 common vetch accessions, which included 50 individuals. Among the 95 EST-SSR markers, the number of alleles ranged from three to 13, and the polymorphism information content values ranged from 0.09 to 0.98. The observed heterozygosity values ranged from 0.00 to 1.00, and the expected heterozygosity values ranged from 0.11 to 0.98. These 95 EST-SSR markers developed from the vetch Illumina transcriptome could greatly promote the development of genetic and molecular breeding studies pertaining to in this species.

  13. Methylation Sensitive Amplification Polymorphism Sequencing (MSAP-Seq)—A Method for High-Throughput Analysis of Differentially Methylated CCGG Sites in Plants with Large Genomes

    Science.gov (United States)

    Chwialkowska, Karolina; Korotko, Urszula; Kosinska, Joanna; Szarejko, Iwona; Kwasniewski, Miroslaw

    2017-01-01

    Epigenetic mechanisms, including histone modifications and DNA methylation, mutually regulate chromatin structure, maintain genome integrity, and affect gene expression and transposon mobility. Variations in DNA methylation within plant populations, as well as methylation in response to internal and external factors, are of increasing interest, especially in the crop research field. Methylation Sensitive Amplification Polymorphism (MSAP) is one of the most commonly used methods for assessing DNA methylation changes in plants. This method involves gel-based visualization of PCR fragments from selectively amplified DNA that are cleaved using methylation-sensitive restriction enzymes. In this study, we developed and validated a new method based on the conventional MSAP approach called Methylation Sensitive Amplification Polymorphism Sequencing (MSAP-Seq). We improved the MSAP-based approach by replacing the conventional separation of amplicons on polyacrylamide gels with direct, high-throughput sequencing using Next Generation Sequencing (NGS) and automated data analysis. MSAP-Seq allows for global sequence-based identification of changes in DNA methylation. This technique was validated in Hordeum vulgare. However, MSAP-Seq can be straightforwardly implemented in different plant species, including crops with large, complex and highly repetitive genomes. The incorporation of high-throughput sequencing into MSAP-Seq enables parallel and direct analysis of DNA methylation in hundreds of thousands of sites across the genome. MSAP-Seq provides direct genomic localization of changes and enables quantitative evaluation. We have shown that the MSAP-Seq method specifically targets gene-containing regions and that a single analysis can cover three-quarters of all genes in large genomes. Moreover, MSAP-Seq's simplicity, cost effectiveness, and high-multiplexing capability make this method highly affordable. Therefore, MSAP-Seq can be used for DNA methylation analysis in crop

  14. Serum C-reactive protein and C-reactive gene (-717C>T polymorphism are not associated with periodontitis in Indonesian male patients

    Directory of Open Access Journals (Sweden)

    Antonius Winoto Suhartono

    2015-09-01

    Full Text Available Background: Periodontitis is an inflammatory disease caused by periodontal pathogens and influenced by multiple risk factors such as genetics, smoking habit, age and systemic diseases. The inflammatory cascade is characterized by the release of C-reactive protein (CRP. Periodontitis has been reported to have plausible links to increased level of CRP, which in turn has been associated to elevated risk of  cardiovascular disease (CVD. Purpose: The purpose of this study was t o investigate the relationship amongst the severity of periodontitis, CRP level in blood and CRP (-717 C>T gene polymorphism in male Indonesian smokers and non-smokers. Method: The severity of periodontitis was assessed for 97 consenting male Indonesian smokers and non-smokers. The CRP level of the subjects was determined by using immuno-turbidimetric assay performed in PARAHITA Diagnostic Center Laboratory ISO 9001: 2000 Cert No. 15225/2. The rate of CRP (-717C>T gene polymorphism was determined by using PCR-RFLP in Oral Biology Laboratory, Faculty of Dentistry, Universitas Indonesia. Result: The results suggest that the CRP protein level is not significantly associated with the tested CRP gene polymorphism (p>0.05. Also, while the severity of periodontitis increased significantly with subject age, the CRP level in blood serum was not significantly related to the severity of  periodontitis. The genotypes of the tested polymorphism did not show significant association with the severity of periodontitis either in smokers or in the combined population including smokers and non-smokers. The results naturally do not exclude such associations, but suggest that to discern the differences the sample size must be considerably increased. Conclusion: The CRP (-717C>T gene polymorphism and CRP level in blood serum were not found to be associated with the severity of periodontitis in male smokers or in the combined population of smokers and non-smokers.

  15. Harnessing Omics Big Data in Nine Vertebrate Species by Genome-Wide Prioritization of Sequence Variants with the Highest Predicted Deleterious Effect on Protein Function.

    Science.gov (United States)

    Rozman, Vita; Kunej, Tanja

    2018-05-10

    Harnessing the genomics big data requires innovation in how we extract and interpret biologically relevant variants. Currently, there is no established catalog of prioritized missense variants associated with deleterious protein function phenotypes. We report in this study, to the best of our knowledge, the first genome-wide prioritization of sequence variants with the most deleterious effect on protein function (potentially deleterious variants [pDelVars]) in nine vertebrate species: human, cattle, horse, sheep, pig, dog, rat, mouse, and zebrafish. The analysis was conducted using the Ensembl/BioMart tool. Genes comprising pDelVars in the highest number of examined species were identified using a Python script. Multiple genomic alignments of the selected genes were built to identify interspecies orthologous potentially deleterious variants, which we defined as the "ortho-pDelVars." Genome-wide prioritization revealed that in humans, 0.12% of the known variants are predicted to be deleterious. In seven out of nine examined vertebrate species, the genes encoding the multiple PDZ domain crumbs cell polarity complex component (MPDZ) and the transforming acidic coiled-coil containing protein 2 (TACC2) comprise pDelVars. Five interspecies ortho-pDelVars were identified in three genes. These findings offer new ways to harness genomics big data by facilitating the identification of functional polymorphisms in humans and animal models and thus provide a future basis for optimization of protocols for whole genome prioritization of pDelVars and screening of orthologous sequence variants. The approach presented here can inform various postgenomic applications such as personalized medicine and multiomics study of health interventions (iatromics).

  16. Analysis of polymorphisms in the promoter region and protein levels of interleukin-6 gene among gout patients.

    Science.gov (United States)

    Tsai, P-C; Chen, C-J; Lai, H-M; Chang, S-J

    2008-01-01

    To explore the associations between the polymorphisms and protein levels of interleukin-6 (IL-6) gene and gout disease. A total of 120 male gout patients and 184 healthy controls were enrolled. Each patient was matched with 1-2 gout-free controls by age within three years. Four polymorphisms in the promoter of IL-6 gene, including -597G/A, -572C/G, -373A(m)T(n), and -174G/C, and the IL-6 levels were analyzed. The clinical characteristics and biochemical markers in plasma were measured, including age of gout onset, duration of gout history, tophus number, gout attack frequency, uric acid, total cholesterol, triglycerides and creatinine. The mean IL-6 level for gout patients was 9.80 (+/-11.76 pg/ml) which showed no significant difference from the controls (7.06+/-7.58 pg/ml, p=0.230). When the IL-6 levels were dichotomized according to the median value (5 pg/ml), there were significantly higher proportions of the gout patients (59.66%) than controls (44%) with high IL-6 levels (OR=1.88, 95% CI=1.17-3.02, p=0.008). Unique genotype was found at polymorphisms -174G/C and -597G/A. Neither the polymorphisms -572C/G nor -373A(m)T(n) in the genotype or allele distributions showed a significant association related to clinical characteristics, biochemical markers, IL-6 levels or gout disease (all p>0.05). Those with gout disease have greater proportions of high IL-6 levels in plasma than controls, and there is no significant association between the four polymorphisms in the promoter region of IL-6 gene and gout disease.

  17. Bioinformatics Interpretation of Exome Sequencing: Blood Cancer

    Directory of Open Access Journals (Sweden)

    Jiwoong Kim

    2013-03-01

    Full Text Available We had analyzed 10 exome sequencing data and single nucleotide polymorphism chips for blood cancer provided by the PGM21 (The National Project for Personalized Genomic Medicine Award program. We had removed sample G06 because the pair is not correct and G10 because of possible contamination. In-house software somatic copy-number and heterozygosity alteration estimation (SCHALE was used to detect one loss of heterozygosity region in G05. We had discovered 27 functionally important mutations. Network and pathway analyses gave us clues that NPM1, GATA2, and CEBPA were major driver genes. By comparing with previous somatic mutation profiles, we had concluded that the provided data originated from acute myeloid leukemia. Protein structure modeling showed that somatic mutations in IDH2, RASGEF1B, and MSH4 can affect protein structures.

  18. ORFer--retrieval of protein sequences and open reading frames from GenBank and storage into relational databases or text files.

    Science.gov (United States)

    Büssow, Konrad; Hoffmann, Steve; Sievert, Volker

    2002-12-19

    Functional genomics involves the parallel experimentation with large sets of proteins. This requires management of large sets of open reading frames as a prerequisite of the cloning and recombinant expression of these proteins. A Java program was developed for retrieval of protein and nucleic acid sequences and annotations from NCBI GenBank, using the XML sequence format. Annotations retrieved by ORFer include sequence name, organism and also the completeness of the sequence. The program has a graphical user interface, although it can be used in a non-interactive mode. For protein sequences, the program also extracts the open reading frame sequence, if available, and checks its correct translation. ORFer accepts user input in the form of single or lists of GenBank GI identifiers or accession numbers. It can be used to extract complete sets of open reading frames and protein sequences from any kind of GenBank sequence entry, including complete genomes or chromosomes. Sequences are either stored with their features in a relational database or can be exported as text files in Fasta or tabulator delimited format. The ORFer program is freely available at http://www.proteinstrukturfabrik.de/orfer. The ORFer program allows for fast retrieval of DNA sequences, protein sequences and their open reading frames and sequence annotations from GenBank. Furthermore, storage of sequences and features in a relational database is supported. Such a database can supplement a laboratory information system (LIMS) with appropriate sequence information.

  19. Cloning, sequencing and variability analysis of the gap gene from Mycoplasma hominis

    DEFF Research Database (Denmark)

    Mygind, Tina; Jacobsen, Iben Søgaard; Melkova, Renata

    2000-01-01

    The gap gene encodes the glycolytic enzyme glyceraldehyde 3-phosphate dehydrogenase (GAPDH). The gene was cloned and sequenced from the Mycoplasma hominis type strain PG21(T). The intraspecies variability was investigated by inspection of restriction fragment length polymorphism (RFLP) patterns...... after polymerase chain reaction (PCR) amplification of the gap gene from 15 strains and furthermore by sequencing of part of the gene in eight strains. The M. hominis gap gene was found to vary more than the Escherichia coli counterpart, but the variation at nucleotide level gave rise to only a few...... amino acid substitutions. To verify that the gene was expressed in M. hominis, a polyclonal antibody was produced and tested against whole cell protein from 15 strains. The enzyme was expressed in all strains investigated as a 36-kDa protein. All strains except type strain PG21(T) showed reaction...

  20. Neural network and SVM classifiers accurately predict lipid binding proteins, irrespective of sequence homology.

    Science.gov (United States)

    Bakhtiarizadeh, Mohammad Reza; Moradi-Shahrbabak, Mohammad; Ebrahimi, Mansour; Ebrahimie, Esmaeil

    2014-09-07

    Due to the central roles of lipid binding proteins (LBPs) in many biological processes, sequence based identification of LBPs is of great interest. The major challenge is that LBPs are diverse in sequence, structure, and function which results in low accuracy of sequence homology based methods. Therefore, there is a need for developing alternative functional prediction methods irrespective of sequence similarity. To identify LBPs from non-LBPs, the performances of support vector machine (SVM) and neural network were compared in this study. Comprehensive protein features and various techniques were employed to create datasets. Five-fold cross-validation (CV) and independent evaluation (IE) tests were used to assess the validity of the two methods. The results indicated that SVM outperforms neural network. SVM achieved 89.28% (CV) and 89.55% (IE) overall accuracy in identification of LBPs from non-LBPs and 92.06% (CV) and 92.90% (IE) (in average) for classification of different LBPs classes. Increasing the number and the range of extracted protein features as well as optimization of the SVM parameters significantly increased the efficiency of LBPs class prediction in comparison to the only previous report in this field. Altogether, the results showed that the SVM algorithm can be run on broad, computationally calculated protein features and offers a promising tool in detection of LBPs classes. The proposed approach has the potential to integrate and improve the common sequence alignment based methods. Copyright © 2014 Elsevier Ltd. All rights reserved.

  1. An improved classification of G-protein-coupled receptors using sequence-derived features

    Directory of Open Access Journals (Sweden)

    Peng Zhen-Ling

    2010-08-01

    Full Text Available Abstract Background G-protein-coupled receptors (GPCRs play a key role in diverse physiological processes and are the targets of almost two-thirds of the marketed drugs. The 3 D structures of GPCRs are largely unavailable; however, a large number of GPCR primary sequences are known. To facilitate the identification and characterization of novel receptors, it is therefore very valuable to develop a computational method to accurately predict GPCRs from the protein primary sequences. Results We propose a new method called PCA-GPCR, to predict GPCRs using a comprehensive set of 1497 sequence-derived features. The principal component analysis is first employed to reduce the dimension of the feature space to 32. Then, the resulting 32-dimensional feature vectors are fed into a simple yet powerful classification algorithm, called intimate sorting, to predict GPCRs at five levels. The prediction at the first level determines whether a protein is a GPCR or a non-GPCR. If it is predicted to be a GPCR, then it will be further predicted into certain family, subfamily, sub-subfamily and subtype by the classifiers at the second, third, fourth, and fifth levels, respectively. To train the classifiers applied at five levels, a non-redundant dataset is carefully constructed, which contains 3178, 1589, 4772, 4924, and 2741 protein sequences at the respective levels. Jackknife tests on this training dataset show that the overall accuracies of PCA-GPCR at five levels (from the first to the fifth can achieve up to 99.5%, 88.8%, 80.47%, 80.3%, and 92.34%, respectively. We further perform predictions on a dataset of 1238 GPCRs at the second level, and on another two datasets of 167 and 566 GPCRs respectively at the fourth level. The overall prediction accuracies of our method are consistently higher than those of the existing methods to be compared. Conclusions The comprehensive set of 1497 features is believed to be capable of capturing information about amino acid

  2. Sequence variation does not confound the measurement of plasma PfHRP2 concentration in African children presenting with severe malaria

    Directory of Open Access Journals (Sweden)

    Ramutton Thiranut

    2012-08-01

    Full Text Available Abstract Background Plasmodium falciparum histidine-rich protein PFHRP2 measurement is used widely for diagnosis, and more recently for severity assessment in falciparum malaria. The Pfhrp2 gene is highly polymorphic, with deletion of the entire gene reported in both laboratory and field isolates. These issues potentially confound the interpretation of PFHRP2 measurements. Methods Studies designed to detect deletion of Pfhrp2 and its paralog Pfhrp3 were undertaken with samples from patients in seven countries contributing to the largest hospital-based severe malaria trial (AQUAMAT. The quantitative relationship between sequence polymorphism and PFHRP2 plasma concentration was examined in samples from selected sites in Mozambique and Tanzania. Results There was no evidence for deletion of either Pfhrp2 or Pfhrp3 in the 77 samples with lowest PFHRP2 plasma concentrations across the seven countries. Pfhrp2 sequence diversity was very high with no haplotypes shared among 66 samples sequenced. There was no correlation between Pfhrp2 sequence length or repeat type and PFHRP2 plasma concentration. Conclusions These findings indicate that sequence polymorphism is not a significant cause of variation in PFHRP2 concentration in plasma samples from African children. This justifies the further development of plasma PFHRP2 concentration as a method for assessing African children who may have severe falciparum malaria. The data also add to the existing evidence base supporting the use of rapid diagnostic tests based on PFHRP2 detection.

  3. Search for methylation-sensitive amplification polymorphisms in mutant figs.

    Science.gov (United States)

    Rodrigues, M G F; Martins, A B G; Bertoni, B W; Figueira, A; Giuliatti, S

    2013-07-08

    Fig (Ficus carica) breeding programs that use conventional approaches to develop new cultivars are rare, owing to limited genetic variability and the difficulty in obtaining plants via gamete fusion. Cytosine methylation in plants leads to gene repression, thereby affecting transcription without changing the DNA sequence. Previous studies using random amplification of polymorphic DNA and amplified fragment length polymorphism markers revealed no polymorphisms among select fig mutants that originated from gamma-irradiated buds. Therefore, we conducted methylation-sensitive amplified polymorphism analysis to verify the existence of variability due to epigenetic DNA methylation among these mutant selections compared to the main cultivar 'Roxo-de-Valinhos'. Samples of genomic DNA were double-digested with either HpaII (methylation sensitive) or MspI (methylation insensitive) and with EcoRI. Fourteen primer combinations were tested, and on an average, non-methylated CCGG, symmetrically methylated CmCGG, and hemimethylated hmCCGG sites accounted for 87.9, 10.1, and 2.0%, respectively. MSAP analysis was effective in detecting differentially methylated sites in the genomic DNA of fig mutants, and methylation may be responsible for the phenotypic variation between treatments. Further analyses such as polymorphic DNA sequencing are necessary to validate these differences, standardize the regions of methylation, and analyze reads using bioinformatic tools.

  4. Epitope Sequences in Dengue Virus NS1 Protein Identified by Monoclonal Antibodies

    Directory of Open Access Journals (Sweden)

    Leticia Barboza Rocha

    2017-10-01

    Full Text Available Dengue nonstructural protein 1 (NS1 is a multi-functional glycoprotein with essential functions both in viral replication and modulation of host innate immune responses. NS1 has been established as a good surrogate marker for infection. In the present study, we generated four anti-NS1 monoclonal antibodies against recombinant NS1 protein from dengue virus serotype 2 (DENV2, which were used to map three NS1 epitopes. The sequence 193AVHADMGYWIESALNDT209 was recognized by monoclonal antibodies 2H5 and 4H1BC, which also cross-reacted with Zika virus (ZIKV protein. On the other hand, the sequence 25VHTWTEQYKFQPES38 was recognized by mAb 4F6 that did not cross react with ZIKV. Lastly, a previously unidentified DENV2 NS1-specific epitope, represented by the sequence 127ELHNQTFLIDGPETAEC143, is described in the present study after reaction with mAb 4H2, which also did not cross react with ZIKV. The selection and characterization of the epitope, specificity of anti-NS1 mAbs, may contribute to the development of diagnostic tools able to differentiate DENV and ZIKV infections.

  5. Phylogenetic analysis of Gossypium L. using restriction fragment length polymorphism of repeated sequences.

    Science.gov (United States)

    Zhang, Meiping; Rong, Ying; Lee, Mi-Kyung; Zhang, Yang; Stelly, David M; Zhang, Hong-Bin

    2015-10-01

    Cotton is the world's leading textile fiber crop and is also grown as a bioenergy and food crop. Knowledge of the phylogeny of closely related species and the genome origin and evolution of polyploid species is significant for advanced genomics research and breeding. We have reconstructed the phylogeny of the cotton genus, Gossypium L., and deciphered the genome origin and evolution of its five polyploid species by restriction fragment analysis of repeated sequences. Nuclear DNA of 84 accessions representing 35 species and all eight genomes of the genus were analyzed. The phylogenetic tree of the genus was reconstructed using the parsimony method on 1033 polymorphic repeated sequence restriction fragments. The genome origin of its polyploids was determined by calculating the diploid-polyploid restriction fragment correspondence (RFC). The tree is consistent with the morphological classification, genome designation and geographic distribution of the species at subgenus, section and subsection levels. Gossypium lobatum (D7) was unambiguously shown to have the highest RFC with the D-subgenomes of all five polyploids of the genus, while the common ancestor of Gossypium herbaceum (A1) and Gossypium arboreum (A2) likely contributed to the A-subgenomes of the polyploids. These results provide a comprehensive phylogenetic tree of the cotton genus and new insights into the genome origin and evolution of its polyploid species. The results also further demonstrate a simple, rapid and inexpensive method suitable for phylogenetic analysis of closely related species, especially congeneric species, and the inference of genome origin of polyploids that constitute over 70 % of flowering plants.

  6. Monte Carlo simulation of a statistical mechanical model of multiple protein sequence alignment.

    Science.gov (United States)

    Kinjo, Akira R

    2017-01-01

    A grand canonical Monte Carlo (MC) algorithm is presented for studying the lattice gas model (LGM) of multiple protein sequence alignment, which coherently combines long-range interactions and variable-length insertions. MC simulations are used for both parameter optimization of the model and production runs to explore the sequence subspace around a given protein family. In this Note, I describe the details of the MC algorithm as well as some preliminary results of MC simulations with various temperatures and chemical potentials, and compare them with the mean-field approximation. The existence of a two-state transition in the sequence space is suggested for the SH3 domain family, and inappropriateness of the mean-field approximation for the LGM is demonstrated.

  7. Vitamin D-Binding Protein Polymorphisms, 25-Hydroxyvitamin D, Sunshine and Multiple Sclerosis

    Directory of Open Access Journals (Sweden)

    Annette Langer-Gould

    2018-02-01

    Full Text Available Blacks have different dominant polymorphisms in the vitamin D-binding protein (DBP gene that result in higher bioavailable vitamin D than whites. This study tested whether the lack of association between 25-hydroxyvitamin D (25OHD and multiple sclerosis (MS risk in blacks and Hispanics is due to differences in these common polymorphisms (rs7041, rs4588. We recruited incident MS cases and controls (blacks 116 cases/131 controls; Hispanics 183/197; whites 247/267 from Kaiser Permanente Southern California. AA is the dominant rs7041 genotype in blacks (70.0% whereas C is the dominant allele in whites (79.0% AC/CC and Hispanics (77.1%. Higher 25OHD levels were associated with a lower risk of MS in whites who carried at least one copy of the C allele but not AA carriers. No association was found in Hispanics or blacks regardless of genotype. Higher ultraviolet radiation exposure was associated with a lower risk of MS in blacks (OR = 0.06, Hispanics and whites who carried at least one copy of the C allele but not in others. Racial/ethnic variations in bioavailable vitamin D do not explain the lack of association between 25OHD and MS in blacks and Hispanics. These findings further challenge the biological plausibility of vitamin D deficiency as causal for MS.

  8. Fourteen polymorphic microsatellite markers for the fungal banana pathogen Mycosphaerella fijiensis.

    Science.gov (United States)

    Yang, Bao Jun; Zhong, Shao Bin

    2008-07-01

    Fourteen polymorphic microsatellite markers were developed for Mycosphaerella fijiensis, a fungus causing the black sigatoka disease in banana. The sequenced genome of M. fijiensis was screened for sequences with single sequence repeats (SSRs) using a Perl script. Fourteen SSR loci, evaluated on 48 M. fijiensis isolates from Hawaii, were identified to be highly polymorphic. These markers revealed two to 19 alleles, with an average of 6.43 alleles per locus. The estimated gene diversity ranged from 0.091 to 0.930 across the 14 microsatellite loci. The SSR markers developed would be useful for population genetics studies of M. fijiensis. © 2008 The Authors. Journal compilation © 2008 Blackwell Publishing Ltd.

  9. Molecular Identification of Date Palm Cultivars Using Random Amplified Polymorphic DNA (RAPD) Markers.

    Science.gov (United States)

    Al-Khalifah, Nasser S; Shanavaskhan, A E

    2017-01-01

    Ambiguity in the total number of date palm cultivars across the world is pointing toward the necessity for an enumerative study using standard morphological and molecular markers. Among molecular markers, DNA markers are more suitable and ubiquitous to most applications. They are highly polymorphic in nature, frequently occurring in genomes, easy to access, and highly reproducible. Various molecular markers such as restriction fragment length polymorphism (RFLP), amplified fragment length polymorphism (AFLP), simple sequence repeats (SSR), inter-simple sequence repeats (ISSR), and random amplified polymorphic DNA (RAPD) markers have been successfully used as efficient tools for analysis of genetic variation in date palm. This chapter explains a stepwise protocol for extracting total genomic DNA from date palm leaves. A user-friendly protocol for RAPD analysis and a table showing the primers used in different molecular techniques that produce polymorphisms in date palm are also provided.

  10. A sequence-based dynamic ensemble learning system for protein ligand-binding site prediction

    KAUST Repository

    Chen, Peng

    2015-12-03

    Background: Proteins have the fundamental ability to selectively bind to other molecules and perform specific functions through such interactions, such as protein-ligand binding. Accurate prediction of protein residues that physically bind to ligands is important for drug design and protein docking studies. Most of the successful protein-ligand binding predictions were based on known structures. However, structural information is not largely available in practice due to the huge gap between the number of known protein sequences and that of experimentally solved structures

  11. A sequence-based dynamic ensemble learning system for protein ligand-binding site prediction

    KAUST Repository

    Chen, Peng; Hu, ShanShan; Zhang, Jun; Gao, Xin; Li, Jinyan; Xia, Junfeng; Wang, Bing

    2015-01-01

    Background: Proteins have the fundamental ability to selectively bind to other molecules and perform specific functions through such interactions, such as protein-ligand binding. Accurate prediction of protein residues that physically bind to ligands is important for drug design and protein docking studies. Most of the successful protein-ligand binding predictions were based on known structures. However, structural information is not largely available in practice due to the huge gap between the number of known protein sequences and that of experimentally solved structures

  12. Protein backbone angle restraints from searching a database for chemical shift and sequence homology

    Energy Technology Data Exchange (ETDEWEB)

    Cornilescu, Gabriel; Delaglio, Frank; Bax, Ad [National Institutes of Health, Laboratory of Chemical Physics, National Institute of Diabetes and Digestive and Kidney Diseases (United States)

    1999-03-15

    Chemical shifts of backbone atoms in proteins are exquisitely sensitive to local conformation, and homologous proteins show quite similar patterns of secondary chemical shifts. The inverse of this relation is used to search a database for triplets of adjacent residues with secondary chemical shifts and sequence similarity which provide the best match to the query triplet of interest. The database contains 13C{alpha}, 13C{beta}, 13C', 1H{alpha} and 15N chemical shifts for 20 proteins for which a high resolution X-ray structure is available. The computer program TALOS was developed to search this database for strings of residues with chemical shift and residue type homology. The relative importance of the weighting factors attached to the secondary chemical shifts of the five types of resonances relative to that of sequence similarity was optimized empirically. TALOS yields the 10 triplets which have the closest similarity in secondary chemical shift and amino acid sequence to those of the query sequence. If the central residues in these 10 triplets exhibit similar {phi} and {psi} backbone angles, their averages can reliably be used as angular restraints for the protein whose structure is being studied. Tests carried out for proteins of known structure indicate that the root-mean-square difference (rmsd) between the output of TALOS and the X-ray derived backbone angles is about 15 deg. Approximately 3% of the predictions made by TALOS are found to be in error.

  13. C-reactive Protein −717A>G and −286C>T>A Gene Polymorphism and Ischemic Stroke

    Directory of Open Access Journals (Sweden)

    Yan Liu

    2015-01-01

    Full Text Available Background: Inflammation plays a pivotal role in the formation and progression of ischemic stroke. Recently, more and more epidemiological studies have focused on the association between C-reactive protein (CRP −717A > G and −286C > T > A genetic polymorphisms and ischemic stroke. However, the findings of these researches are not conclusive. Methods: We performed a meta-analysis to determine whether these two polymorphisms are associated with the risk of ischemic stroke. Eligible studies were identified from the database of PubMed, Medline, Embase, Web of Science, CNKI, Weipu, and Wanfang. We calculated odds ratios (ORs with 95% confidence intervals (CIs to assess the strength of the association. Results: Four articles were included in our study, including 1926 cases and 2678 controls for −717A > G polymorphism, 652 cases and 1103 controls for −286C > T > A polymorphism. The results of meta-analysis showed that single nucleotide polymorphism (SNP −717A > G was not significantly associated with the risk of ischemic stroke (GG vs. AA, OR = 1.12, 95% CI = 0.83-1.50, P = 0.207; GG + GA vs. AA, OR = 1.04, 95% CI = 0.93-1.17, P = 0.533; GG vs. GA + AA, OR = 1.10, 95% CI = 0.82-1.47, P = 0.220. Meta-analysis of SNP − 286C > T > A also demonstrated no statistical evidence of a significant association with the risk of ischemic stroke (AA vs. CC, OR = 0.86, 95% CI = 0.59-1.25, P = 0.348; AA vs. CC, OR = 0.92, 95% CI = 0.80-1.06, P = 0.609; AA vs. CC, OR = 0.89, 95% CI = 0.62-1.30, P = 0.374. Conclusions: This meta-analysis demonstrated little evidence to support a role of CRP gene −717A > G, −286C > T > A polymorphisms in ischemic stroke predisposition. However, to draw comprehensive and more reliable conclusions, further larger studies are needed to validate the association between CRP gene polymorphisms and ischemic stroke in various ethnic groups.

  14. Combining sequence-based prediction methods and circular dichroism and infrared spectroscopic data to improve protein secondary structure determinations

    Directory of Open Access Journals (Sweden)

    Lees Jonathan G

    2008-01-01

    Full Text Available Abstract Background A number of sequence-based methods exist for protein secondary structure prediction. Protein secondary structures can also be determined experimentally from circular dichroism, and infrared spectroscopic data using empirical analysis methods. It has been proposed that comparable accuracy can be obtained from sequence-based predictions as from these biophysical measurements. Here we have examined the secondary structure determination accuracies of sequence prediction methods with the empirically determined values from the spectroscopic data on datasets of proteins for which both crystal structures and spectroscopic data are available. Results In this study we show that the sequence prediction methods have accuracies nearly comparable to those of spectroscopic methods. However, we also demonstrate that combining the spectroscopic and sequences techniques produces significant overall improvements in secondary structure determinations. In addition, combining the extra information content available from synchrotron radiation circular dichroism data with sequence methods also shows improvements. Conclusion Combining sequence prediction with experimentally determined spectroscopic methods for protein secondary structure content significantly enhances the accuracy of the overall results obtained.

  15. Sequencing and De Novo Transcriptome Assembly of Brachypodium sylvaticum (Poaceae

    Directory of Open Access Journals (Sweden)

    Samuel E. Fox

    2013-03-01

    Full Text Available Premise of the study: We report the de novo assembly and characterization of the transcriptomes of Brachypodium sylvaticum (slender false-brome accessions from native populations of Spain and Greece, and an invasive population west of Corvallis, Oregon, USA. Methods and Results: More than 350 million sequence reads from the mRNA libraries prepared from three B. sylvaticum genotypes were assembled into 120,091 (Corvallis, 104,950 (Spain, and 177,682 (Greece transcript contigs. In comparison with the B. distachyon Bd21 reference genome and GenBank protein sequences, we estimate >90% exome coverage for B. sylvaticum. The transcripts were assigned Gene Ontology and InterPro annotations. Brachypodium sylvaticum sequence reads aligned against the Bd21 genome revealed 394,654 single-nucleotide polymorphisms (SNPs and >20,000 simple sequence repeat (SSR DNA sites. Conclusions: To our knowledge, this is the first report of transcriptome sequencing of invasive plant species with a closely related sequenced reference genome. The sequences and identified SNP variant and SSR sites will provide tools for developing novel genetic markers for use in genotyping and characterization of invasive behavior of B. sylvaticum.

  16. Cloning of soluble alkaline phosphatase cDNA and molecular basis of the polymorphic nature in alkaline phosphatase isozymes of Bombyx mori midgut.

    Science.gov (United States)

    Itoh, M; Kanamori, Y; Takao, M; Eguchi, M

    1999-02-01

    A cDNA coding for soluble type alkaline phosphatase (sALP) of Bombyx mori was isolated. Deduced amino acid sequence showed high identities to various ALPs and partial similarities to ATPase of Manduca sexta. Using this cDNA sequence as a probe, the molecular basis of electrophoretic polymorphism in sALP and membrane-bound type ALP (mALP) was studied. As for mALP, the result suggested that post-translational modification was important for the proteins to express activity and to represent their extensive polymorphic nature, whereas the magnitude of activities was mainly regulated by transcription. On the other hand, sALP zymogram showed poor polymorphism, but one exception was the null mutant, in which the sALP gene was largely lost. Interestingly, the sALP gene was shown to be transcribed into two mRNAs of different sizes, 2.0 and 2.4 Kb. In addition to the null mutant of sALP, we found a null mutant for mALP. Both of these mutants seem phenotypically silent, suggesting that the functional differentiation between these isozymes is not perfect, so that they can still work mutually and complement each other as an indispensable enzyme for B. mori.

  17. DIALIGN: multiple DNA and protein sequence alignment at BiBiServ.

    OpenAIRE

    Morgenstern, Burkhard

    2004-01-01

    DIALIGN is a widely used software tool for multiple DNA and protein sequence alignment. The program combines local and global alignment features and can therefore be applied to sequence data that cannot be correctly aligned by more traditional approaches. DIALIGN is available online through Bielefeld Bioinformatics Server (BiBiServ). The downloadable version of the program offers several new program features. To compare the output of different alignment programs, we developed the program AltA...

  18. Genotyping of human parvovirus B19 in clinical samples from Brazil and Paraguay using heteroduplex mobility assay, single-stranded conformation polymorphism and nucleotide sequencing

    Directory of Open Access Journals (Sweden)

    Marcos César Lima de Mendonça

    2011-06-01

    Full Text Available Heteroduplex mobility assay, single-stranded conformation polymorphism and nucleotide sequencing were utilised to genotype human parvovirus B19 samples from Brazil and Paraguay. Ninety-seven serum samples were collected from individuals presenting with abortion or erythema infectiosum, arthropathies, severe anaemia and transient aplastic crisis; two additional skin samples were collected by biopsy. After the procedure, all clinical samples were classified as genotype 1.

  19. Helicobacter pylori cagL amino acid polymorphism D58E59 pave the way toward peptic ulcer disease while N58E59 is associated with gastric cancer in north of Iran.

    Science.gov (United States)

    Cherati, Mina Rezaee; Shokri-Shirvani, Javad; Karkhah, Ahmad; Rajabnia, Ramzan; Nouri, Hamid Reza

    2017-06-01

    The cagL protein of Helicobacter pylori involving in pathogenesis of gastroduodenal disorders. Therefore, the current study was conducted to determine the cagL amino acid polymorphisms in patients with gastric diseases. One hundred gastric biopsies were collected from gastritis, peptic ulcer (PUD) and gastric cancer (GC) patients and were screened for cagL using polymerase chain reaction (PCR). Also, sequence variations of the cagL were assessed via sequence translation. The cagL geneopositivity was 71.6% in patients were infected with H. pylori. The cagL from PUD indicated a higher rate of D58 amino acid sequence polymorphism than those of the GC and gastritis (P peptic ulcer. However, amino acid N, M, Q and N at the same position alongside V134 increased the risk of gastric cancer. Copyright © 2017 Elsevier Ltd. All rights reserved.

  20. Transcriptome Sequencing and Development of Genic SSR Markers of an Endangered Chinese Endemic Genus Dipteronia Oliver (Aceraceae).

    Science.gov (United States)

    Zhou, Tao; Li, Zhong-Hu; Bai, Guo-Qing; Feng, Li; Chen, Chen; Wei, Yue; Chang, Yong-Xia; Zhao, Gui-Fang

    2016-02-23

    Dipteronia Oliver (Aceraceae) is an endangered Chinese endemic genus consisting of two living species, Dipteronia sinensis and Dipteronia dyeriana. However, studies on the population genetics and evolutionary analyses of Dipteronia have been hindered by limited genomic resources and genetic markers. Here, the generation, de novo assembly and annotation of transcriptome datasets, and a large set of microsatellite or simple sequence repeat (SSR) markers derived from Dipteronia have been described. After Illumina pair-end sequencing, approximately 93.2 million reads were generated and assembled to yield a total of 99,358 unigenes. A majority of these unigenes (53%, 52,789) had at least one blast hit against the public protein databases. Further, 12,377 SSR loci were detected and 4179 primer pairs were designed for experimental validation. Of these 4179 primer pairs, 435 primer pairs were randomly selected to test polymorphism. Our results show that products from 132 primer pairs were polymorphic, in which 97 polymorphic SSR markers were further selected to analyze the genetic diversity of 10 natural populations of Dipteronia. The identification of SSR markers during our research will provide the much valuable data for population genetic analyses and evolutionary studies in Dipteronia.

  1. Oleosome-Associated Protein of the Oleaginous Diatom Fistulifera solaris Contains an Endoplasmic Reticulum-Targeting Signal Sequence

    Directory of Open Access Journals (Sweden)

    Yoshiaki Maeda

    2014-06-01

    Full Text Available Microalgae tend to accumulate lipids as an energy storage material in the specific organelle, oleosomes. Current studies have demonstrated that lipids derived from microalgal oleosomes are a promising source of biofuels, while the oleosome formation mechanism has not been fully elucidated. Oleosome-associated proteins have been identified from several microalgae to elucidate the fundamental mechanisms of oleosome formation, although understanding their functions is still in infancy. Recently, we discovered a diatom-oleosome-associated-protein 1 (DOAP1 from the oleaginous diatom, Fistulifera solaris JPCC DA0580. The DOAP1 sequence implied that this protein might be transported into the endoplasmic reticulum (ER due to the signal sequence. To ensure this, we fused the signal sequence to green fluorescence protein. The fusion protein distributed around the chloroplast as like a meshwork membrane structure, indicating the ER localization. This result suggests that DOAP1 could firstly localize at the ER, then move to the oleosomes. This study also demonstrated that the DOAP1 signal sequence allowed recombinant proteins to be specifically expressed in the ER of the oleaginous diatom. It would be a useful technique for engineering the lipid synthesis pathways existing in the ER, and finally controlling the biofuel quality.

  2. Mannose-binding lectin 2 (Mbl2 gene polymorphisms are related to protein plasma levels, but not to heart disease and infection by Chlamydia

    Directory of Open Access Journals (Sweden)

    M.A.F. Queiroz

    Full Text Available The presence of the single nucleotide polymorphisms in exon 1 of the mannose-binding lectin 2 (MBL2 gene was evaluated in a sample of 159 patients undergoing coronary artery bypass surgery (71 patients undergoing valve replacement surgery and 300 control subjects to investigate a possible association between polymorphisms and heart disease with Chlamydia infection. The identification of the alleles B and D was performed using real time polymerase chain reaction (PCR and of the allele C was accomplished through PCR assays followed by digestion with the restriction enzyme. The comparative analysis of allelic and genotypic frequencies between the three groups did not reveal any significant difference, even when related to previous Chlamydia infection. Variations in the MBL plasma levels were influenced by the presence of polymorphisms, being significantly higher in the group of cardiac patients, but without representing a risk for the disease. The results showed that despite MBL2 gene polymorphisms being associated with the protein plasma levels, the polymorphisms were not enough to predict the development of heart disease, regardless of infection with both species of Chlamydia.

  3. Genetic maps of polymorphic DNA loci on rat chromosome 1

    Energy Technology Data Exchange (ETDEWEB)

    Ding, Yan-Ping; Remmers, E.F.; Longman, R.E. [National Institutes of Health, Bethesda, MD (United States)] [and others

    1996-09-01

    Genetic linkage maps of loci defined by polymorphic DNA markers on rat chromosome 1 were constructed by genotyping F2 progeny of F344/N x LEW/N, BN/SsN x LEW/N, and DA/Bkl x F344/Hsd inbred rat strains. In total, 43 markers were mapped, of which 3 were restriction fragment length polymorphisms and the others were simple sequence length polymorphisms. Nineteen of these markers were associated with genes. Six markers for five genes, {gamma}-aminobutyric acid receptor {beta}3 (Gabrb3), syntaxin 2 (Stx2), adrenergic receptor {beta}3 (Gabrb3), syntaxin 2 (Stx2), adrenergic receptor {beta}1 (Adrb1), carcinoembryonic antigen gene family member 1 (Cgm1), and lipogenic protein S14 (Lpgp), and 20 anonymous loci were not previously reported. Thirteen gene loci (Myl2, Aldoa, Tnt, Igf2, Prkcg, Cgm4, Calm3, Cgm3, Psbp1, Sa, Hbb, Ins1, and Tcp1) were previously mapped. Comparative mapping analysis indicated that the large portion of rat chromosome 1 is homologous to mouse chromosome 7, although the homologous to mouse chromosome 7, although the homologs of two rat genes are located on mouse chromosomes 17 and 19. Homologs of the rat chromosome 1 genes that we mapped are located on human chromosomes 6, 10, 11, 12, 15, 16, and 19. 38 refs., 1 fig., 3 tabs.

  4. The 434(G>C) polymorphism in the eosinophil cationic protein gene and its association with tissue eosinophilia in oral squamous cell carcinomas

    DEFF Research Database (Denmark)

    Pereira, Michele C; Oliveira, Denise T; Olivieri, Eloísa H R

    2010-01-01

    OBJECTIVE: The aim of this study was to investigate the prevalence of the Eosinophil cationic protein (ECP)-gene polymorphism 434(G>C) in oral squamous cell carcinoma (OSCC) patients and its association with tumor-associated tissue eosinophilia (TATE), demographic, clinical, and microscopic...... of ECP-gene polymorphism 434(G>C) with TATE, demographic, clinical, and microscopic variables in OSCC patients. Disease-free survival and overall survival were calculated by the Kaplan-Meier product-limit actuarial method and the comparison of the survival curves were performed using log rank test...

  5. Evaluation of multiple approaches to identify genome-wide polymorphisms in closely related genotypes of sweet cherry (Prunus avium L.

    Directory of Open Access Journals (Sweden)

    Seanna Hewitt

    Full Text Available Identification of genetic polymorphisms and subsequent development of molecular markers is important for marker assisted breeding of superior cultivars of economically important species. Sweet cherry (Prunus avium L. is an economically important non-climacteric tree fruit crop in the Rosaceae family and has undergone a genetic bottleneck due to breeding, resulting in limited genetic diversity in the germplasm that is utilized for breeding new cultivars. Therefore, it is critical to recognize the best platforms for identifying genome-wide polymorphisms that can help identify, and consequently preserve, the diversity in a genetically constrained species. For the identification of polymorphisms in five closely related genotypes of sweet cherry, a gel-based approach (TRAP, reduced representation sequencing (TRAPseq, a 6k cherry SNParray, and whole genome sequencing (WGS approaches were evaluated in the identification of genome-wide polymorphisms in sweet cherry cultivars. All platforms facilitated detection of polymorphisms among the genotypes with variable efficiency. In assessing multiple SNP detection platforms, this study has demonstrated that a combination of appropriate approaches is necessary for efficient polymorphism identification, especially between closely related cultivars of a species. The information generated in this study provides a valuable resource for future genetic and genomic studies in sweet cherry, and the insights gained from the evaluation of multiple approaches can be utilized for other closely related species with limited genetic diversity in the breeding germplasm. Keywords: Polymorphisms, Prunus avium, Next-generation sequencing, Target region amplification polymorphism (TRAP, Genetic diversity, SNParray, Reduced representation sequencing, Whole genome sequencing (WGS

  6. Sequence preservation of osteocalcin protein and mitochondrial DNA in bison bones older than 55 ka

    Science.gov (United States)

    Nielsen-Marsh, Christina M.; Ostrom, Peggy H.; Gandhi, Hasand; Shapiro, Beth; Cooper, Alan; Hauschka, Peter V.; Collins, Matthew J.

    2002-12-01

    We report the first complete sequences of the protein osteocalcin from small amounts (20 mg) of two bison bone (Bison priscus) dated to older than 55.6 ka and older than 58.9 ka. Osteocalcin was purified using new gravity columns (never exposed to protein) followed by microbore reversed-phase high-performance liquid chromatography. Sequencing of osteocalcin employed two methods of matrix-assisted laser desorption ionization mass spectrometry (MALDI-MS): peptide mass mapping (PMM) and post-source decay (PSD). The PMM shows that ancient and modern bison osteocalcin have the same mass to charge (m/z) distribution, indicating an identical protein sequence and absence of diagenetic products. This was confirmed by PSD of the m/z 2066 tryptic peptide (residues 1 19); the mass spectra from ancient and modern peptides were identical. The 129 mass unit difference in the molecular ion between cow (Bos taurus) and bison is caused by a single amino-acid substitution between the taxa (Trp in cow is replaced by Gly in bison at residue 5). Bison mitochondrial control region DNA sequences were obtained from the older than 55.6 ka fossil. These results suggest that DNA and protein sequences can be used to directly investigate molecular phylogenies over a considerable time period, the absolute limit of which is yet to be determined.

  7. Sequence analysis corresponding to the PPE and PE proteins in ...

    Indian Academy of Sciences (India)

    Unknown

    AB repeats; Mycobacterium tuberculosis genome; PE-PPE domain; PPE, PE proteins; sequence analysis; surface antigens. J. Biosci. | Vol. ... bacterium tuberculosis genomes resulted in the identification of a previously uncharacterized 225 amino acid- ...... Vega Lopez F, Brooks L A, Dockrell H M, De Smet K A,. Thompson ...

  8. Protein sequence analysis by incorporating modified chaos game and physicochemical properties into Chou's general pseudo amino acid composition.

    Science.gov (United States)

    Xu, Chunrui; Sun, Dandan; Liu, Shenghui; Zhang, Yusen

    2016-10-07

    In this contribution we introduced a novel graphical method to compare protein sequences. By mapping a protein sequence into 3D space based on codons and physicochemical properties of 20 amino acids, we are able to get a unique P-vector from the 3D curve. This approach is consistent with wobble theory of amino acids. We compute the distance between sequences by their P-vectors to measure similarities/dissimilarities among protein sequences. Finally, we use our method to analyze four datasets and get better results compared with previous approaches. Copyright © 2016 Elsevier Ltd. All rights reserved.

  9. PFP: Automated prediction of gene ontology functional annotations with confidence scores using protein sequence data.

    Science.gov (United States)

    Hawkins, Troy; Chitale, Meghana; Luban, Stanislav; Kihara, Daisuke

    2009-02-15

    Protein function prediction is a central problem in bioinformatics, increasing in importance recently due to the rapid accumulation of biological data awaiting interpretation. Sequence data represents the bulk of this new stock and is the obvious target for consideration as input, as newly sequenced organisms often lack any other type of biological characterization. We have previously introduced PFP (Protein Function Prediction) as our sequence-based predictor of Gene Ontology (GO) functional terms. PFP interprets the results of a PSI-BLAST search by extracting and scoring individual functional attributes, searching a wide range of E-value sequence matches, and utilizing conventional data mining techniques to fill in missing information. We have shown it to be effective in predicting both specific and low-resolution functional attributes when sufficient data is unavailable. Here we describe (1) significant improvements to the PFP infrastructure, including the addition of prediction significance and confidence scores, (2) a thorough benchmark of performance and comparisons to other related prediction methods, and (3) applications of PFP predictions to genome-scale data. We applied PFP predictions to uncharacterized protein sequences from 15 organisms. Among these sequences, 60-90% could be annotated with a GO molecular function term at high confidence (>or=80%). We also applied our predictions to the protein-protein interaction network of the Malaria plasmodium (Plasmodium falciparum). High confidence GO biological process predictions (>or=90%) from PFP increased the number of fully enriched interactions in this dataset from 23% of interactions to 94%. Our benchmark comparison shows significant performance improvement of PFP relative to GOtcha, InterProScan, and PSI-BLAST predictions. This is consistent with the performance of PFP as the overall best predictor in both the AFP-SIG '05 and CASP7 function (FN) assessments. PFP is available as a web service at http

  10. Porcine MYF6 gene: sequence, homology analysis, and variation in the promoter region.

    Science.gov (United States)

    Wyszyńska-Koko, J; Kurył, J

    2004-01-01

    MYF6 gene codes for the bHLH transcription factor belonging to MyoD family. Its expression accompanies the processes of differentiation and maturation of myotubes during embriogenesis and continues on a relatively high level after birth, affecting the muscle phenotype. The porcine MYF6 gene was amplified and sequenced and compared with MYF6 gene sequences of other species. The amino acid sequence was deduced and an interspecies homology analysis was performed. Myf-6 protein shows a high conservation among species of 99 and 97% identity when comparing pig with cow and human, respectively, and of 93% when comparing pig with mouse and rat. The single nucleotide polymorphism (SNP) was revealed within the promoter region, which appeared to be T --> C transition recognized by a MspI restriction enzyme.

  11. Fc receptor gamma subunit polymorphisms and systemic lupus erythematosus

    International Nuclear Information System (INIS)

    Al-Ansari, Aliya; Ollier, W.E.; Gonzalez-Gay, Miguel A.; Gul, Ahmet; Inanac, Murat; Ordi, Jose; Teh, Lee-Suan; Hajeer, Ali H.

    2004-01-01

    To investigate the possible association between Fc receptor gamma polymorphisms and systemic lupus erythematosus (SLE). We have investigated the full FcR gamma gene for polymorphisms using polymerase chain reaction (PCR)-single strand confirmational polymorphisms and DNA sequencing .The polymorphisms identified were genotype using PCR-restriction fragment length polymorphism. Systemic lupus erythematosus cases and controls were available from 3 ethnic groups: Turkish, Spanish and Caucasian. The study was conducted in the year 2001 at the Arthritis Research Campaign, Epidemiology Unit, Manchester University Medical School, Manchester, United Kingdom. Five single nucleotide polymorphisms were identified, 2 in the promoter, one in intron 4 and, 2 in the 3'UTR. Four of the 5 single nucleotide polymorphisms (SNPs) were relatively common and investigated in the 3 populations. Allele and genotype frequencies of all 4 investigated SNPs were not statistically different cases and controls. fc receptor gamma gene does not appear to contribute to SLE susceptibility. The identified polymorphisms may be useful in investigating other diseases where receptors containing the FcR gamma subunit contribute to the pathology. (author)

  12. Single Nucleotide Polymorphism

    DEFF Research Database (Denmark)

    Børsting, Claus; Pereira, Vania; Andersen, Jeppe Dyrberg

    2014-01-01

    Single nucleotide polymorphisms (SNPs) are the most frequent DNA sequence variations in the genome. They have been studied extensively in the last decade with various purposes in mind. In this chapter, we will discuss the advantages and disadvantages of using SNPs for human identification...... of SNPs. This will allow acquisition of more information from the sample materials and open up for new possibilities as well as new challenges....

  13. Rampant adaptive evolution in regions of proteins with unknown function in Drosophila simulans.

    Directory of Open Access Journals (Sweden)

    Alisha K Holloway

    2007-10-01

    Full Text Available Adaptive protein evolution is pervasive in Drosophila. Genomic studies, thus far, have analyzed each protein as a single entity. However, the targets of adaptive events may be localized to particular parts of proteins, such as protein domains or regions involved in protein folding. We compared the population genetic mechanisms driving sequence polymorphism and divergence in defined protein domains and non-domain regions. Interestingly, we find that non-domain regions of proteins are more frequent targets of directional selection. Protein domains are also evolving under directional selection, but appear to be under stronger purifying selection than non-domain regions. Non-domain regions of proteins clearly play a major role in adaptive protein evolution on a genomic scale and merit future investigations of their functional properties.

  14. Association of surfactant protein A polymorphisms with otitis media in infants at risk for asthma

    Directory of Open Access Journals (Sweden)

    Bracken Michael B

    2006-08-01

    Full Text Available Abstract Background Otitis media is one of the most common infections of early childhood. Surfactant protein A functions as part of the innate immune response, which plays an important role in preventing infections early in life. This prospective study utilized a candidate gene approach to evaluate the association between polymorphisms in loci encoding SP-A and risk of otitis media during the first year of life among a cohort of infants at risk for developing asthma. Methods Between September 1996 and December 1998, women were invited to participate if they had at least one other child with physician-diagnosed asthma. Each mother was given a standardized questionnaire within 4 months of her infant's birth. Infant respiratory symptoms were collected during quarterly telephone interviews at 6, 9 and 12 months of age. Genotyping was done on 355 infants for whom whole blood and complete otitis media data were available. Results Polymorphisms at codons 19, 62, and 133 in SP-A1, and 223 in SP-A2 were associated with race/ethnicity. In logistic regression models incorporating estimates of uncertainty in haplotype assignment, the 6A4/1A5haplotype was protective for otitis media among white infants in our study population (OR 0.23; 95% CI 0.07,0.73. Conclusion These results indicate that polymorphisms within SP-A loci may be associated with otitis media in white infants. Larger confirmatory studies in all ethnic groups are warranted.

  15. H-type bovine spongiform encephalopathy associated with E211K prion protein polymorphism: clinical and pathologic features in wild-type and E211K cattle following intracranial inoculation

    Science.gov (United States)

    In 2006 an H-type bovine spongiform encephalopathy (BSE) case was reported in an animal with an unusual polymorphism (E211K) in the prion protein gene. Although the prevalence of this polymorphism is low, cattle carrying the K211 allele are predisposed to rapid onset of H-type BSE when exposed. The ...

  16. Polymorphism of proteins in selected slovak winter wheat genotypes using SDS-PAGE

    Directory of Open Access Journals (Sweden)

    Dana Miháliková

    2016-12-01

    Full Text Available Winter wheat is especially used for bread-making. The specific composition of the grain storage proteins and the representation of individual subunits determines the baking quality of wheat. The aim of this study was to analyze 15 slovak varieties of the winter wheat (Triticum aestivum L. based on protein polymorphism and to predict their technological quality. SDS-PAGE method by ISTA was used to separate glutenin protein subunits. Glutenins were separated into HMW-GS (15.13% and LMW-GS (65.89% on the basis of molecular weight in SDS-PAGE. At the locus Glu-A1 was found allele Null (53% of genotypes and allele 1 (47% of genotypes. The locus Glu-B1 was represented by the HMW-GS subunits 6+8 (33% of genotypes, 7+8 (27% of genotypes, 7+9 (40% of genotypes. At the locus Glu-D1 were detected two subunits, 2+12 (33% of genotypes and 5+10 (67% of genotypes which is correlated with good bread-making properties. The Glu – score was ranged from 4 (genotype Viglanka to 10 (genotypes Viola, Vladarka. According to the representation of individual glutenin subunits in samples, the dendrogram of genetic similarity was constructed. By the prediction of quality the results showed that the best technological quality was significant in the varieties Viola and Vladarka which are suitable for use in food processing.

  17. A Single Nucleotide Polymorphism in the Bax Gene Promoter Affects Transcription and Influences Retinal Ganglion Cell Death

    Directory of Open Access Journals (Sweden)

    Sheila J Semaan

    2010-03-01

    Full Text Available Pro-apoptotic Bax is essential for RGC (retinal ganglion cell death. Gene dosage experiments in mice, yielding a single wild-type Bax allele, indicated that genetic background was able to influence the cell death phenotype. DBA/2J Bax+/− mice exhibited complete resistance to nerve damage after 2 weeks (similar to Bax −/− mice, but 129B6 Bax+/− mice exhibited significant cell loss (similar to wild-type mice. The different cell death phenotype was associated with the level of Bax expression, where 129B6 neurons had twice the level of endogenous Bax mRNA and protein as DBA/2J neurons. Sequence analysis of the Bax promoters between these strains revealed a single nucleotide polymorphism (T129B6 to CDBA/2J at position −515. A 1.5- to 2.5-fold increase in transcriptional activity was observed from the 129B6 promoter in transient transfection assays in a variety of cell types, including RGC5 cells derived from rat RGCs. Since this polymorphism occurred in a p53 half-site, we investigated the requirement of p53 for the differential transcriptional activity. Differential transcriptional activity from either 129B6 or DBA/2J Bax promoters were unaffected in p53−/− cells, and addition of exogenous p53 had no further effect on this difference, thus a role for p53 was excluded. Competitive electrophoretic mobility-shift assays identified two DNA-protein complexes that interacted with the polymorphic region. Those forming Complex 1 bound with higher affinity to the 129B6 polymorphic site, suggesting that these proteins probably comprised a transcriptional activator complex. These studies implicated quantitative expression of the Bax gene as playing a possible role in neuronal susceptibility to damaging stimuli.

  18. Single nucleotide polymorphism barcoding of cytochrome c oxidase I sequences for discriminating 17 species of Columbidae by decision tree algorithm.

    Science.gov (United States)

    Yang, Cheng-Hong; Wu, Kuo-Chuan; Dahms, Hans-Uwe; Chuang, Li-Yeh; Chang, Hsueh-Wei

    2017-07-01

    DNA barcodes are widely used in taxonomy, systematics, species identification, food safety, and forensic science. Most of the conventional DNA barcode sequences contain the whole information of a given barcoding gene. Most of the sequence information does not vary and is uninformative for a given group of taxa within a monophylum. We suggest here a method that reduces the amount of noninformative nucleotides in a given barcoding sequence of a major taxon, like the prokaryotes, or eukaryotic animals, plants, or fungi. The actual differences in genetic sequences, called single nucleotide polymorphism (SNP) genotyping, provide a tool for developing a rapid, reliable, and high-throughput assay for the discrimination between known species. Here, we investigated SNPs as robust markers of genetic variation for identifying different pigeon species based on available cytochrome c oxidase I (COI) data. We propose here a decision tree-based SNP barcoding (DTSB) algorithm where SNP patterns are selected from the DNA barcoding sequence of several evolutionarily related species in order to identify a single species with pigeons as an example. This approach can make use of any established barcoding system. We here firstly used as an example the mitochondrial gene COI information of 17 pigeon species (Columbidae, Aves) using DTSB after sequence trimming and alignment. SNPs were chosen which followed the rule of decision tree and species-specific SNP barcodes. The shortest barcode of about 11 bp was then generated for discriminating 17 pigeon species using the DTSB method. This method provides a sequence alignment and tree decision approach to parsimoniously assign a unique and shortest SNP barcode for any known species of a chosen monophyletic taxon where a barcoding sequence is available.

  19. Visualization of protein sequence features using JavaScript and SVG with pViz.js.

    Science.gov (United States)

    Mukhyala, Kiran; Masselot, Alexandre

    2014-12-01

    pViz.js is a visualization library for displaying protein sequence features in a Web browser. By simply providing a sequence and the locations of its features, this lightweight, yet versatile, JavaScript library renders an interactive view of the protein features. Interactive exploration of protein sequence features over the Web is a common need in Bioinformatics. Although many Web sites have developed viewers to display these features, their implementations are usually focused on data from a specific source or use case. Some of these viewers can be adapted to fit other use cases but are not designed to be reusable. pViz makes it easy to display features as boxes aligned to a protein sequence with zooming functionality but also includes predefined renderings for secondary structure and post-translational modifications. The library is designed to further customize this view. We demonstrate such applications of pViz using two examples: a proteomic data visualization tool with an embedded viewer for displaying features on protein structure, and a tool to visualize the results of the variant_effect_predictor tool from Ensembl. pViz.js is a JavaScript library, available on github at https://github.com/Genentech/pviz. This site includes examples and functional applications, installation instructions and usage documentation. A Readme file, which explains how to use pViz with examples, is available as Supplementary Material A. © The Author 2014. Published by Oxford University Press. All rights reserved. For Permissions, please e-mail: journals.permissions@oup.com.

  20. Implication of the cause of differences in 3D structures of proteins with high sequence identity based on analyses of amino acid sequences and 3D structures.

    Science.gov (United States)

    Matsuoka, Masanari; Sugita, Masatake; Kikuchi, Takeshi

    2014-09-18

    Proteins that share a high sequence homology while exhibiting drastically different 3D structures are investigated in this study. Recently, artificial proteins related to the sequences of the GA and IgG binding GB domains of human serum albumin have been designed. These artificial proteins, referred to as GA and GB, share 98% amino acid sequence identity but exhibit different 3D structures, namely, a 3α bundle versus a 4β + α structure. Discriminating between their 3D structures based on their amino acid sequences is a very difficult problem. In the present work, in addition to using bioinformatics techniques, an analysis based on inter-residue average distance statistics is used to address this problem. It was hard to distinguish which structure a given sequence would take only with the results of ordinary analyses like BLAST and conservation analyses. However, in addition to these analyses, with the analysis based on the inter-residue average distance statistics and our sequence tendency analysis, we could infer which part would play an important role in its structural formation. The results suggest possible determinants of the different 3D structures for sequences with high sequence identity. The possibility of discriminating between the 3D structures based on the given sequences is also discussed.

  1. Population structure of pigs determined by single nucleotide polymorphisms observed in assembled expressed sequence tags.

    Science.gov (United States)

    Matsumoto, Toshimi; Okumura, Naohiko; Uenishi, Hirohide; Hayashi, Takeshi; Hamasima, Noriyuki; Awata, Takashi

    2012-01-01

    We have collected more than 190000 porcine expressed sequence tags (ESTs) from full-length complementary DNA (cDNA) libraries and identified more than 2800 single nucleotide polymorphisms (SNPs). In this study, we tentatively chose 222 SNPs observed in assembled ESTs to study pigs of different breeds; 104 were selected by comparing the cDNA sequences of a Meishan pig and samples of three-way cross pigs (Landrace, Large White, and Duroc: LWD), and 118 were selected from LWD samples. To evaluate the genetic variation between the chosen SNPs from pig breeds, we determined the genotypes for 192 pig samples (11 pig groups) from our DNA reference panel with matrix-assisted laser desorption ionization time-of-flight mass spectrometry. Of the 222 reference SNPs, 186 were successfully genotyped. A neighbor-joining tree showed that the pig groups were classified into two large clusters, namely, Euro-American and East Asian pig populations. F-statistics and the analysis of molecular variance of Euro-American pig groups revealed that approximately 25% of the genetic variations occurred because of intergroup differences. As the F(IS) values were less than the F(ST) values(,) the clustering, based on the Bayesian inference, implied that there was strong genetic differentiation among pig groups and less divergence within the groups in our samples. © 2011 The Authors. Animal Science Journal © 2011 Japanese Society of Animal Science.

  2. Genetic diversity of the merozoite surface protein-3 gene in Plasmodium falciparum populations in Thailand.

    Science.gov (United States)

    Pattaradilokrat, Sittiporn; Sawaswong, Vorthon; Simpalipan, Phumin; Kaewthamasorn, Morakot; Siripoon, Napaporn; Harnyuttanakorn, Pongchai

    2016-10-21

    An effective malaria vaccine is an urgently needed tool to fight against human malaria, the most deadly parasitic disease of humans. One promising candidate is the merozoite surface protein-3 (MSP-3) of Plasmodium falciparum. This antigenic protein, encoded by the merozoite surface protein (msp-3) gene, is polymorphic and classified according to size into the two allelic types of K1 and 3D7. A recent study revealed that both the K1 and 3D7 alleles co-circulated within P. falciparum populations in Thailand, but the extent of the sequence diversity and variation within each allelic type remains largely unknown. The msp-3 gene was sequenced from 59 P. falciparum samples collected from five endemic areas (Mae Hong Son, Kanchanaburi, Ranong, Trat and Ubon Ratchathani) in Thailand and analysed for nucleotide sequence diversity, haplotype diversity and deduced amino acid sequence diversity. The gene was also subject to population genetic analysis (F st ) and neutrality tests (Tajima's D, Fu and Li D* and Fu and Li' F* tests) to determine any signature of selection. The sequence analyses revealed eight unique DNA haplotypes and seven amino acid sequence variants, with a haplotype and nucleotide diversity of 0.828 and 0.049, respectively. Neutrality tests indicated that the polymorphism detected in the alanine heptad repeat region of MSP-3 was maintained by positive diversifying selection, suggesting its role as a potential target of protective immune responses and supporting its role as a vaccine candidate. Comparison of MSP-3 variants among parasite populations in Thailand, India and Nigeria also inferred a close genetic relationship between P. falciparum populations in Asia. This study revealed the extent of the msp-3 gene diversity in P. falciparum in Thailand, providing the fundamental basis for the better design of future blood stage malaria vaccines against P. falciparum.

  3. A two-step recognition of signal sequences determines the translocation efficiency of proteins.

    Science.gov (United States)

    Belin, D; Bost, S; Vassalli, J D; Strub, K

    1996-02-01

    The cytosolic and secreted, N-glycosylated, forms of plasminogen activator inhibitor-2 (PAI-2) are generated by facultative translocation. To study the molecular events that result in the bi-topological distribution of proteins, we determined in vitro the capacities of several signal sequences to bind the signal recognition particle (SRP) during targeting, and to promote vectorial transport of murine PAI-2 (mPAI-2). Interestingly, the six signal sequences we compared (mPAI-2 and three mutated derivatives thereof, ovalbumin and preprolactin) were found to have the differential activities in the two events. For example, the mPAI-2 signal sequence first binds SRP with moderate efficiency and secondly promotes the vectorial transport of only a fraction of the SRP-bound nascent chains. Our results provide evidence that the translocation efficiency of proteins can be controlled by the recognition of their signal sequences at two steps: during SRP-mediated targeting and during formation of a committed translocation complex. This second recognition may occur at several time points during the insertion/translocation step. In conclusion, signal sequences have a more complex structure than previously anticipated, allowing for multiple and independent interactions with the translocation machinery.

  4. Nucleotide sequence of a cDNA for branched chain acyltransferase with analysis of the deduced protein structure

    International Nuclear Information System (INIS)

    Hummel, K.B.; Litwer, S.; Bradford, A.P.; Aitken, A.; Danner, D.J.; Yeaman, S.J.

    1988-01-01

    Nucleotide sequence was determined for a 1.6-kilobase human cDNA putative for the branched chain acyltransferase protein of the branched chain α-ketoacid dehydrogenase complex. Translation of the sequence reveals an open reading frame encoding a 315-amino acid protein of molecular weight 35,759 followed by 560 bases of 3'-untranslated sequence. Three repeats of the polyadenylation signal hexamer ATTAAA are present prior to the polyadenylate tail. Within the open reading frame is a 10-amino acid fragment which matches exactly the amino acid sequence around the lipoate-lysine residue in bovine kidney branched chain acyltransferase, thus confirming the identity of the cDNA. Analysis of the deduced protein structure for the human branched chain acyltransferase revealed an organization into domains similar to that reported for the acyltransferase proteins of the pyruvate and α-ketoglutarate dehydrogenase complexes. This similarity in organization suggests that a more detailed analysis of the proteins will be required to explain the individual substrate and multienzyme complex specificity shown by these acyltransferases

  5. Estimating Genetic Conformism of Korean Mulberry Cultivars Using Random Amplified Polymorphic DNA and Inter-Simple Sequence Repeat Profiling

    Directory of Open Access Journals (Sweden)

    Sunirmal Sheet

    2018-03-01

    Full Text Available Apart from being fed to silkworms in sericulture, the ecologically important Mulberry plant has been used for traditional medicine in Asian countries as well as in manufacturing wine, food, and beverages. Germplasm analysis among Mulberry cultivars originating from South Korea is crucial in the plant breeding program for cultivar development. Hence, the genetic deviations and relations among 8 Morus alba plants, and one Morus lhou plant, of different cultivars collected from South Korea were investigated using 10 random amplified polymorphic DNA (RAPD and 10 inter-simple sequence repeat (ISSR markers in the present study. The ISSR markers exhibited a higher polymorphism (63.42% among mulberry genotypes in comparison to RAPD markers. Furthermore, the similarity coefficient was estimated for both markers and found to be varying between 0.183 and 0.814 for combined pooled data of ISSR and RAPD. The phenogram drawn using the UPGMA cluster method based on combined pooled data of RAPD and ISSR markers divided the nine mulberry genotypes into two divergent major groups and the two individual independent accessions. The distant relationship between Dae-Saug (SM1 and SangchonJo Sang Saeng (SM5 offers a possibility of utilizing them in mulberry cultivar improvement of Morus species of South Korea.

  6. No Effect of the Transforming Growth Factor {beta}1 Promoter Polymorphism C-509T on TGFB1 Gene Expression, Protein Secretion, or Cellular Radiosensitivity

    Energy Technology Data Exchange (ETDEWEB)

    Reuther, Sebastian; Metzke, Elisabeth [Laboratory of Radiobiology and Experimental Radiooncology, University Hospital Hamburg-Eppendorf, Hamburg (Germany); Bonin, Michael [Department of Medical Genetics, University of Tuebingen (Germany); Petersen, Cordula [Clinic of Radiotherapy and Radiooncology, University Hospital Hamburg-Eppendorf, Hamburg (Germany); Dikomey, Ekkehard, E-mail: dikomey@uke.de [Laboratory of Radiobiology and Experimental Radiooncology, University Hospital Hamburg-Eppendorf, Hamburg (Germany); Raabe, Annette [Laboratory of Radiobiology and Experimental Radiooncology, University Hospital Hamburg-Eppendorf, Hamburg (Germany)

    2013-02-01

    Purpose: To study whether the promoter polymorphism (C-509T) affects transforming growth factor {beta}1 gene (TGFB1) expression, protein secretion, and/or cellular radiosensitivity for both human lymphocytes and fibroblasts. Methods and Materials: Experiments were performed with lymphocytes taken either from 124 breast cancer patients or 59 pairs of normal monozygotic twins. We used 15 normal human primary fibroblast strains as controls. The C-509T genotype was determined by polymerase chain reaction-restriction fragment length polymorphism or TaqMan single nucleotide polymorphism (SNP) genotyping assay. The cellular radiosensitivity of lymphocytes was measured by G0/1 assay and that of fibroblasts by colony assay. The amount of extracellular TGFB1 protein was determined by enzyme-linked immunosorbent assay, and TGFB1 expression was assessed via microarray analysis or reverse transcription-polymerase chain reaction. Results: The C-509T genotype was found not to be associated with cellular radiosensitivity, neither for lymphocytes (breast cancer patients, P=.811; healthy donors, P=.181) nor for fibroblasts (P=.589). Both TGFB1 expression and TGFB1 protein secretion showed considerable variation, which, however, did not depend on the C-509T genotype (protein secretion: P=.879; gene expression: lymphocytes, P=.134, fibroblasts, P=.605). There was also no general correlation between TGFB1 expression and cellular radiosensitivity (lymphocytes, P=.632; fibroblasts, P=.573). Conclusion: Our data indicate that any association between the SNP C-509T of TGFB1 and risk of normal tissue toxicity cannot be ascribed to a functional consequence of this SNP, either on the level of gene expression, protein secretion, or cellular radiosensitivity.

  7. Pregnancy-associated osteoporosis with a heterozygous deactivating LDL receptor-related protein 5 (LRP5) mutation and a homozygous methylenetetrahydrofolate reductase (MTHFR) polymorphism.

    Science.gov (United States)

    Cook, Fiona J; Mumm, Steven; Whyte, Michael P; Wenkert, Deborah

    2014-04-01

    Pregnancy-associated osteoporosis (PAO) is a rare, idiopathic disorder that usually presents with vertebral compression fractures (VCFs) within 6 months of a first pregnancy and delivery. Spontaneous improvement is typical. There is no known genetic basis for PAO. A 26-year-old primagravida with a neonatal history of unilateral blindness attributable to hyperplastic primary vitreous sustained postpartum VCFs consistent with PAO. Her low bone mineral density (BMD) seemed to respond to vitamin D and calcium therapy, with no fractures after her next successful pregnancy. Investigation of subsequent fetal losses revealed homozygosity for the methylenetetrahydrofolate reductase (MTHFR) C677T polymorphism associated both with fetal loss and with osteoporosis (OP). Because her neonatal unilateral blindness and OP were suggestive of loss-of-function mutation(s) in the gene that encodes LDL receptor-related protein 5 (LRP5), LRP5 exon and splice site sequencing was also performed. This revealed a unique heterozygous 12-bp deletion in exon 21 (c.4454_4465del, p.1485_1488del SSSS) in the patient, her mother and sons, but not her father or brother. Her mother had a normal BMD, no history of fractures, PAO, ophthalmopathy, or fetal loss. Her two sons had no ophthalmopathy and no skeletal issues. Her osteoporotic father (with a family history of blindness) and brother had low BMDs first documented at ages ∼40 and 32 years, respectively. Serum biochemical and bone turnover studies were unremarkable in all subjects. We postulate that our patient's heterozygous LRP5 mutation together with her homozygous MTHFR polymorphism likely predisposed her to low peak BMD. However, OP did not cosegregate in her family with the LRP5 mutation, the homozygous MTHFR polymorphism, or even the combination of the two, implicating additional genetic or nongenetic factors in her PAO. Nevertheless, exploration for potential genetic contributions to PAO may explain part of the pathogenesis of this

  8. Sequencing of the Chlamydophila psittaci ompA Gene Reveals a New Genotype, E/B, and the Need for a Rapid Discriminatory Genotyping Method

    Science.gov (United States)

    Geens, Tom; Desplanques, Ann; Van Loock, Marnix; Bönner, Brigitte M.; Kaleta, Erhard F.; Magnino, Simone; Andersen, Arthur A.; Everett, Karin D. E.; Vanrompay, Daisy

    2005-01-01

    Twenty-one avian Chlamydophila psittaci isolates from different European countries were characterized using ompA restriction fragment length polymorphism, ompA sequencing, and major outer membrane protein serotyping. Results reveal the presence of a new genotype, E/B, in several European countries and stress the need for a discriminatory rapid genotyping method. PMID:15872282

  9. Experimental Rugged Fitness Landscape in Protein Sequence Space

    Science.gov (United States)

    Hayashi, Yuuki; Aita, Takuyo; Toyota, Hitoshi; Husimi, Yuzuru; Urabe, Itaru; Yomo, Tetsuya

    2006-01-01

    The fitness landscape in sequence space determines the process of biomolecular evolution. To plot the fitness landscape of protein function, we carried out in vitro molecular evolution beginning with a defective fd phage carrying a random polypeptide of 139 amino acids in place of the g3p minor coat protein D2 domain, which is essential for phage infection. After 20 cycles of random substitution at sites 12–130 of the initial random polypeptide and selection for infectivity, the selected phage showed a 1.7×104-fold increase in infectivity, defined as the number of infected cells per ml of phage suspension. Fitness was defined as the logarithm of infectivity, and we analyzed (1) the dependence of stationary fitness on library size, which increased gradually, and (2) the time course of changes in fitness in transitional phases, based on an original theory regarding the evolutionary dynamics in Kauffman's n-k fitness landscape model. In the landscape model, single mutations at single sites among n sites affect the contribution of k other sites to fitness. Based on the results of these analyses, k was estimated to be 18–24. According to the estimated parameters, the landscape was plotted as a smooth surface up to a relative fitness of 0.4 of the global peak, whereas the landscape had a highly rugged surface with many local peaks above this relative fitness value. Based on the landscapes of these two different surfaces, it appears possible for adaptive walks with only random substitutions to climb with relative ease up to the middle region of the fitness landscape from any primordial or random sequence, whereas an enormous range of sequence diversity is required to climb further up the rugged surface above the middle region. PMID:17183728

  10. Experimental rugged fitness landscape in protein sequence space.

    Science.gov (United States)

    Hayashi, Yuuki; Aita, Takuyo; Toyota, Hitoshi; Husimi, Yuzuru; Urabe, Itaru; Yomo, Tetsuya

    2006-12-20

    The fitness landscape in sequence space determines the process of biomolecular evolution. To plot the fitness landscape of protein function, we carried out in vitro molecular evolution beginning with a defective fd phage carrying a random polypeptide of 139 amino acids in place of the g3p minor coat protein D2 domain, which is essential for phage infection. After 20 cycles of random substitution at sites 12-130 of the initial random polypeptide and selection for infectivity, the selected phage showed a 1.7x10(4)-fold increase in infectivity, defined as the number of infected cells per ml of phage suspension. Fitness was defined as the logarithm of infectivity, and we analyzed (1) the dependence of stationary fitness on library size, which increased gradually, and (2) the time course of changes in fitness in transitional phases, based on an original theory regarding the evolutionary dynamics in Kauffman's n-k fitness landscape model. In the landscape model, single mutations at single sites among n sites affect the contribution of k other sites to fitness. Based on the results of these analyses, k was estimated to be 18-24. According to the estimated parameters, the landscape was plotted as a smooth surface up to a relative fitness of 0.4 of the global peak, whereas the landscape had a highly rugged surface with many local peaks above this relative fitness value. Based on the landscapes of these two different surfaces, it appears possible for adaptive walks with only random substitutions to climb with relative ease up to the middle region of the fitness landscape from any primordial or random sequence, whereas an enormous range of sequence diversity is required to climb further up the rugged surface above the middle region.

  11. Experimental rugged fitness landscape in protein sequence space.

    Directory of Open Access Journals (Sweden)

    Yuuki Hayashi

    Full Text Available The fitness landscape in sequence space determines the process of biomolecular evolution. To plot the fitness landscape of protein function, we carried out in vitro molecular evolution beginning with a defective fd phage carrying a random polypeptide of 139 amino acids in place of the g3p minor coat protein D2 domain, which is essential for phage infection. After 20 cycles of random substitution at sites 12-130 of the initial random polypeptide and selection for infectivity, the selected phage showed a 1.7x10(4-fold increase in infectivity, defined as the number of infected cells per ml of phage suspension. Fitness was defined as the logarithm of infectivity, and we analyzed (1 the dependence of stationary fitness on library size, which increased gradually, and (2 the time course of changes in fitness in transitional phases, based on an original theory regarding the evolutionary dynamics in Kauffman's n-k fitness landscape model. In the landscape model, single mutations at single sites among n sites affect the contribution of k other sites to fitness. Based on the results of these analyses, k was estimated to be 18-24. According to the estimated parameters, the landscape was plotted as a smooth surface up to a relative fitness of 0.4 of the global peak, whereas the landscape had a highly rugged surface with many local peaks above this relative fitness value. Based on the landscapes of these two different surfaces, it appears possible for adaptive walks with only random substitutions to climb with relative ease up to the middle region of the fitness landscape from any primordial or random sequence, whereas an enormous range of sequence diversity is required to climb further up the rugged surface above the middle region.

  12. Protein sequences bound to mineral surfaces persist into deep time

    DEFF Research Database (Denmark)

    Demarchi, Beatrice; Hall, Shaun; Roncal-Herrero, Teresa

    2016-01-01

    of Laetoli (3.8 Ma) and Olduvai Gorge (1.3 Ma) in Tanzania. By tracking protein diagenesis back in time we find consistent patterns of preservation, demonstrating authenticity of the surviving sequences. Molecular dynamics simulations of struthiocalcin-1 and -2, the dominant proteins within the eggshell......, reveal that distinct domains bind to the mineral surface. It is the domain with the strongest calculated binding energy to the calcite surface that is selectively preserved. Thermal age calculations demonstrate that the Laetoli and Olduvai peptides are 50 times older than any previously authenticated...

  13. Programming molecular self-assembly of intrinsically disordered proteins containing sequences of low complexity

    Science.gov (United States)

    Simon, Joseph R.; Carroll, Nick J.; Rubinstein, Michael; Chilkoti, Ashutosh; López, Gabriel P.

    2017-06-01

    Dynamic protein-rich intracellular structures that contain phase-separated intrinsically disordered proteins (IDPs) composed of sequences of low complexity (SLC) have been shown to serve a variety of important cellular functions, which include signalling, compartmentalization and stabilization. However, our understanding of these structures and our ability to synthesize models of them have been limited. We present design rules for IDPs possessing SLCs that phase separate into diverse assemblies within droplet microenvironments. Using theoretical analyses, we interpret the phase behaviour of archetypal IDP sequences and demonstrate the rational design of a vast library of multicomponent protein-rich structures that ranges from uniform nano-, meso- and microscale puncta (distinct protein droplets) to multilayered orthogonally phase-separated granular structures. The ability to predict and program IDP-rich assemblies in this fashion offers new insights into (1) genetic-to-molecular-to-macroscale relationships that encode hierarchical IDP assemblies, (2) design rules of such assemblies in cell biology and (3) molecular-level engineering of self-assembled recombinant IDP-rich materials.

  14. Polymorphism of the simple sequence repeat (AAC)5 in the ...

    Indian Academy of Sciences (India)

    2013-12-04

    Dec 4, 2013 ... SSRs could be present in coding and noncoding regions, contributing to genome dynamics and evolution. Previous studies by our research group detected molecular and cytogenetic riboso- mal DNA (rDNA) polymorphisms in Old Portuguese bread and durum wheat cultivars. Considering the rRNA genes.

  15. Genome-wide profiling of DNA-binding proteins using barcode-based multiplex Solexa sequencing.

    Science.gov (United States)

    Raghav, Sunil Kumar; Deplancke, Bart

    2012-01-01

    Chromatin immunoprecipitation (ChIP) is a commonly used technique to detect the in vivo binding of proteins to DNA. ChIP is now routinely paired to microarray analysis (ChIP-chip) or next-generation sequencing (ChIP-Seq) to profile the DNA occupancy of proteins of interest on a genome-wide level. Because ChIP-chip introduces several biases, most notably due to the use of a fixed number of probes, ChIP-Seq has quickly become the method of choice as, depending on the sequencing depth, it is more sensitive, quantitative, and provides a greater binding site location resolution. With the ever increasing number of reads that can be generated per sequencing run, it has now become possible to analyze several samples simultaneously while maintaining sufficient sequence coverage, thus significantly reducing the cost per ChIP-Seq experiment. In this chapter, we provide a step-by-step guide on how to perform multiplexed ChIP-Seq analyses. As a proof-of-concept, we focus on the genome-wide profiling of RNA Polymerase II as measuring its DNA occupancy at different stages of any biological process can provide insights into the gene regulatory mechanisms involved. However, the protocol can also be used to perform multiplexed ChIP-Seq analyses of other DNA-binding proteins such as chromatin modifiers and transcription factors.

  16. Age-related macular degeneration-associated silent polymorphisms in HtrA1 impair its ability to antagonize insulin-like growth factor 1.

    Science.gov (United States)

    Jacobo, Sarah Melissa P; Deangelis, Margaret M; Kim, Ivana K; Kazlauskas, Andrius

    2013-05-01

    Synonymous single nucleotide polymorphisms (SNPs) within a transcript's coding region produce no change in the amino acid sequence of the protein product and are therefore intuitively assumed to have a neutral effect on protein function. We report that two common variants of high-temperature requirement A1 (HTRA1) that increase the inherited risk of neovascular age-related macular degeneration (NvAMD) harbor synonymous SNPs within exon 1 of HTRA1 that convert common codons for Ala34 and Gly36 to less frequently used codons. The frequent-to-rare codon conversion reduced the mRNA translation rate and appeared to compromise HtrA1's conformation and function. The protein product generated from the SNP-containing cDNA displayed enhanced susceptibility to proteolysis and a reduced affinity for an anti-HtrA1 antibody. The NvAMD-associated synonymous polymorphisms lie within HtrA1's putative insulin-like growth factor 1 (IGF-1) binding domain. They reduced HtrA1's abilities to associate with IGF-1 and to ameliorate IGF-1-stimulated signaling events and cellular responses. These observations highlight the relevance of synonymous codon usage to protein function and implicate homeostatic protein quality control mechanisms that may go awry in NvAMD.

  17. Inter-simple sequence repeat (ISSR) markers in the evaluation of ...

    African Journals Online (AJOL)

    shawkat

    2013-02-13

    Feb 13, 2013 ... 666 Afr. J. Biotechnol. Table 1. Number and types of the ISSR bands as well as the total polymorphism percentages generated in six Capsicum hybrids. Primer code. Sequence. Monomorphic band. Polymorphic band. Total band. Polymorphism. (%). Unique. Shared. HB 1. (CAA)5. 4. 0. 1. 5. 20. HB 2. (CAG) ...

  18. The Ising model for prediction of disordered residues from protein sequence alone

    International Nuclear Information System (INIS)

    Lobanov, Michail Yu; Galzitskaya, Oxana V

    2011-01-01

    Intrinsically disordered regions serve as molecular recognition elements, which play an important role in the control of many cellular processes and signaling pathways. It is useful to be able to predict positions of disordered residues and disordered regions in protein chains using protein sequence alone. A new method (IsUnstruct) based on the Ising model for prediction of disordered residues from protein sequence alone has been developed. According to this model, each residue can be in one of two states: ordered or disordered. The model is an approximation of the Ising model in which the interaction term between neighbors has been replaced by a penalty for changing between states (the energy of border). The IsUnstruct has been compared with other available methods and found to perform well. The method correctly finds 77% of disordered residues as well as 87% of ordered residues in the CASP8 database, and 72% of disordered residues as well as 85% of ordered residues in the DisProt database

  19. Molecular Simulations of Sequence-Specific Association of Transmembrane Proteins in Lipid Bilayers

    Science.gov (United States)

    Doxastakis, Manolis; Prakash, Anupam; Janosi, Lorant

    2011-03-01

    Association of membrane proteins is central in material and information flow across the cellular membranes. Amino-acid sequence and the membrane environment are two critical factors controlling association, however, quantitative knowledge on such contributions is limited. In this work, we study the dimerization of helices in lipid bilayers using extensive parallel Monte Carlo simulations with recently developed algorithms. The dimerization of Glycophorin A is examined employing a coarse-grain model that retains a level of amino-acid specificity, in three different phospholipid bilayers. Association is driven by a balance of protein-protein and lipid-induced interactions with the latter playing a major role at short separations. Following a different approach, the effect of amino-acid sequence is studied using the four transmembrane domains of the epidermal growth factor receptor family in identical lipid environments. Detailed characterization of dimer formation and estimates of the free energy of association reveal that these helices present significant affinity to self-associate with certain dimers forming non-specific interfaces.

  20. Role of sequence and structural polymorphism on the mechanical properties of amyloid fibrils.

    Directory of Open Access Journals (Sweden)

    Gwonchan Yoon

    Full Text Available Amyloid fibrils playing a critical role in disease expression, have recently been found to exhibit the excellent mechanical properties such as elastic modulus in the order of 10 GPa, which is comparable to that of other mechanical proteins such as microtubule, actin filament, and spider silk. These remarkable mechanical properties of amyloid fibrils are correlated with their functional role in disease expression. This suggests the importance in understanding how these excellent mechanical properties are originated through self-assembly process that may depend on the amino acid sequence. However, the sequence-structure-property relationship of amyloid fibrils has not been fully understood yet. In this work, we characterize the mechanical properties of human islet amyloid polypeptide (hIAPP fibrils with respect to their molecular structures as well as their amino acid sequence by using all-atom explicit water molecular dynamics (MD simulation. The simulation result suggests that the remarkable bending rigidity of amyloid fibrils can be achieved through a specific self-aggregation pattern such as antiparallel stacking of β strands (peptide chain. Moreover, we have shown that a single point mutation of hIAPP chain constituting a hIAPP fibril significantly affects the thermodynamic stability of hIAPP fibril formed by parallel stacking of peptide chain, and that a single point mutation results in a significant change in the bending rigidity of hIAPP fibrils formed by antiparallel stacking of β strands. This clearly elucidates the role of amino acid sequence on not only the equilibrium conformations of amyloid fibrils but also their mechanical properties. Our study sheds light on sequence-structure-property relationships of amyloid fibrils, which suggests that the mechanical properties of amyloid fibrils are encoded in their sequence-dependent molecular architecture.