21 resultados para GENOME SEQUENCE

em Duke University


Relevância:

70.00% 70.00%

Publicador:

Resumo:

Cryptococcus neoformans is a pathogenic basidiomycetous yeast responsible for more than 600,000 deaths each year. It occurs as two serotypes (A and D) representing two varieties (i.e. grubii and neoformans, respectively). Here, we sequenced the genome and performed an RNA-Seq-based analysis of the C. neoformans var. grubii transcriptome structure. We determined the chromosomal locations, analyzed the sequence/structural features of the centromeres, and identified origins of replication. The genome was annotated based on automated and manual curation. More than 40,000 introns populating more than 99% of the expressed genes were identified. Although most of these introns are located in the coding DNA sequences (CDS), over 2,000 introns in the untranslated regions (UTRs) were also identified. Poly(A)-containing reads were employed to locate the polyadenylation sites of more than 80% of the genes. Examination of the sequences around these sites revealed a new poly(A)-site-associated motif (AUGHAH). In addition, 1,197 miscRNAs were identified. These miscRNAs can be spliced and/or polyadenylated, but do not appear to have obvious coding capacities. Finally, this genome sequence enabled a comparative analysis of strain H99 variants obtained after laboratory passage. The spectrum of mutations identified provides insights into the genetics underlying the micro-evolution of a laboratory strain, and identifies mutations involved in stress responses, mating efficiency, and virulence.

Relevância:

70.00% 70.00%

Publicador:

Resumo:

BACKGROUND: The availability of multiple avian genome sequence assemblies greatly improves our ability to define overall genome organization and reconstruct evolutionary changes. In birds, this has previously been impeded by a near intractable karyotype and relied almost exclusively on comparative molecular cytogenetics of only the largest chromosomes. Here, novel whole genome sequence information from 21 avian genome sequences (most newly assembled) made available on an interactive browser (Evolution Highway) was analyzed. RESULTS: Focusing on the six best-assembled genomes allowed us to assemble a putative karyotype of the dinosaur ancestor for each chromosome. Reconstructing evolutionary events that led to each species' genome organization, we determined that the fastest rate of change occurred in the zebra finch and budgerigar, consistent with rapid speciation events in the Passeriformes and Psittaciformes. Intra- and interchromosomal changes were explained most parsimoniously by a series of inversions and translocations respectively, with breakpoint reuse being commonplace. Analyzing chicken and zebra finch, we found little evidence to support the hypothesis of an association of evolutionary breakpoint regions with recombination hotspots but some evidence to support the hypothesis that microchromosomes largely represent conserved blocks of synteny in the majority of the 21 species analyzed. All but one species showed the expected number of microchromosomal rearrangements predicted by the haploid chromosome count. Ostrich, however, appeared to retain an overall karyotype structure of 2n=80 despite undergoing a large number (26) of hitherto un-described interchromosomal changes. CONCLUSIONS: Results suggest that mechanisms exist to preserve a static overall avian karyotype/genomic structure, including the microchromosomes, with widespread interchromosomal change occurring rarely (e.g., in ostrich and budgerigar lineages). Of the species analyzed, the chicken lineage appeared to have undergone the fewest changes compared to the dinosaur ancestor.

Relevância:

70.00% 70.00%

Publicador:

Resumo:

Ferns are one of the few remaining major clades of land plants for which a complete genome sequence is lacking. Knowledge of genome space in ferns will enable broad-scale comparative analyses of land plant genes and genomes, provide insights into genome evolution across green plants, and shed light on genetic and genomic features that characterize ferns, such as their high chromosome numbers and large genome sizes. As part of an initial exploration into fern genome space, we used a whole genome shotgun sequencing approach to obtain low-density coverage (∼0.4X to 2X) for six fern species from the Polypodiales (Ceratopteris, Pteridium, Polypodium, Cystopteris), Cyatheales (Plagiogyria), and Gleicheniales (Dipteris). We explore these data to characterize the proportion of the nuclear genome represented by repetitive sequences (including DNA transposons, retrotransposons, ribosomal DNA, and simple repeats) and protein-coding genes, and to extract chloroplast and mitochondrial genome sequences. Such initial sweeps of fern genomes can provide information useful for selecting a promising candidate fern species for whole genome sequencing. We also describe variation of genomic traits across our sample and highlight some differences and similarities in repeat structure between ferns and seed plants.

Relevância:

70.00% 70.00%

Publicador:

Resumo:

Photographs from the February 1997 Bermuda meeting. Courtesy of Gert-Jan van Ommen.

Relevância:

60.00% 60.00%

Publicador:

Resumo:

Ongoing Cryptococcus gattii outbreaks in the Western United States and Canada illustrate the impact of environmental reservoirs and both clonal and recombining propagation in driving emergence and expansion of microbial pathogens. C. gattii comprises four distinct molecular types: VGI, VGII, VGIII, and VGIV, with no evidence of nuclear genetic exchange, indicating these represent distinct species. C. gattii VGII isolates are causing the Pacific Northwest outbreak, whereas VGIII isolates frequently infect HIV/AIDS patients in Southern California. VGI, VGII, and VGIII have been isolated from patients and animals in the Western US, suggesting these molecular types occur in the environment. However, only two environmental isolates of C. gattii have ever been reported from California: CBS7750 (VGII) and WM161 (VGIII). The incongruence of frequent clinical presence and uncommon environmental isolation suggests an unknown C. gattii reservoir in California. Here we report frequent isolation of C. gattii VGIII MATα and MATa isolates and infrequent isolation of VGI MATα from environmental sources in Southern California. VGIII isolates were obtained from soil debris associated with tree species not previously reported as hosts from sites near residences of infected patients. These isolates are fertile under laboratory conditions, produce abundant spores, and are part of both locally and more distantly recombining populations. MLST and whole genome sequence analysis provide compelling evidence that these environmental isolates are the source of human infections. Isolates displayed wide-ranging virulence in macrophage and animal models. When clinical and environmental isolates with indistinguishable MLST profiles were compared, environmental isolates were less virulent. Taken together, our studies reveal an environmental source and risk of C. gattii to HIV/AIDS patients with implications for the >1,000,000 cryptococcal infections occurring annually for which the causative isolate is rarely assigned species status. Thus, the C. gattii global health burden could be more substantial than currently appreciated.

Relevância:

60.00% 60.00%

Publicador:

Resumo:

The International Crocodilian Genomes Working Group (ICGWG) will sequence and assemble the American alligator (Alligator mississippiensis), saltwater crocodile (Crocodylus porosus) and Indian gharial (Gavialis gangeticus) genomes. The status of these projects and our planned analyses are described.

Relevância:

60.00% 60.00%

Publicador:

Resumo:

Ferns are the only major lineage of vascular plants not represented by a sequenced nuclear genome. This lack of genome sequence information significantly impedes our ability to understand and reconstruct genome evolution not only in ferns, but across all land plants. Azolla and Ceratopteris are ideal and complementary candidates to be the first ferns to have their nuclear genomes sequenced. They differ dramatically in genome size, life history, and habit, and thus represent the immense diversity of extant ferns. Together, this pair of genomes will facilitate myriad large-scale comparative analyses across ferns and all land plants. Here we review the unique biological characteristics of ferns and describe a number of outstanding questions in plant biology that will benefit from the addition of ferns to the set of taxa with sequenced nuclear genomes. We explain why the fern clade is pivotal for understanding genome evolution across land plants, and we provide a rationale for how knowledge of fern genomes will enable progress in research beyond the ferns themselves.

Relevância:

60.00% 60.00%

Publicador:

Resumo:

Email exchange in 2013 between Kathryn Maxson (Duke) and Kris Wetterstrand (NHGRI), regarding country funding and other data for the HGP sequencing centers. Also includes the email request for such information, from NHGRI to the centers, in 2000, and the aggregate data collected.

Relevância:

60.00% 60.00%

Publicador:

Resumo:

Jean Weissenbach, telephone interview by Kathryn Maxson and Robert Cook-Deegan, conducted from Durham, NC 09 February 2012. Jean Weissenbach, a leader in French genetic mapping, directed the French national sequencing center, Généthon, during the HGP and was instrumental in helping to build agreement to the Bermuda Principles in France.

Relevância:

60.00% 60.00%

Publicador:

Resumo:

Mark Guyer and Jane Peterson, in-person interview with Kathryn Maxson and Robert Cook-Deegan, conducted in Rockville, MD (NIH campus), 18 August 2011. Mark Guyer and Jane Peterson were grants program officers at the NIH during the HGP, and were some of the longest-standing employees in the HGP administrative structure. Both witnessed the transformation of the Office of Genome Research into the National Center for Human Genome Research and, finally, the National Human Genome Research Institute. They were close participants in the history of the Bermuda Principles within the NIH.

Relevância:

30.00% 30.00%

Publicador:

Resumo:

BACKGROUND: The rate of emergence of human pathogens is steadily increasing; most of these novel agents originate in wildlife. Bats, remarkably, are the natural reservoirs of many of the most pathogenic viruses in humans. There are two bat genome projects currently underway, a circumstance that promises to speed the discovery host factors important in the coevolution of bats with their viruses. These genomes, however, are not yet assembled and one of them will provide only low coverage, making the inference of most genes of immunological interest error-prone. Many more wildlife genome projects are underway and intend to provide only shallow coverage. RESULTS: We have developed a statistical method for the assembly of gene families from partial genomes. The method takes full advantage of the quality scores generated by base-calling software, incorporating them into a complete probabilistic error model, to overcome the limitation inherent in the inference of gene family members from partial sequence information. We validated the method by inferring the human IFNA genes from the genome trace archives, and used it to infer 61 type-I interferon genes, and single type-II interferon genes in the bats Pteropus vampyrus and Myotis lucifugus. We confirmed our inferences by direct cloning and sequencing of IFNA, IFNB, IFND, and IFNK in P. vampyrus, and by demonstrating transcription of some of the inferred genes by known interferon-inducing stimuli. CONCLUSION: The statistical trace assembler described here provides a reliable method for extracting information from the many available and forthcoming partial or shallow genome sequencing projects, thereby facilitating the study of a wider variety of organisms with ecological and biomedical significance to humans than would otherwise be possible.

Relevância:

30.00% 30.00%

Publicador:

Resumo:

BACKGROUND: There is considerable interest in the development of methods to efficiently identify all coding variants present in large sample sets of humans. There are three approaches possible: whole-genome sequencing, whole-exome sequencing using exon capture methods, and RNA-Seq. While whole-genome sequencing is the most complete, it remains sufficiently expensive that cost effective alternatives are important. RESULTS: Here we provide a systematic exploration of how well RNA-Seq can identify human coding variants by comparing variants identified through high coverage whole-genome sequencing to those identified by high coverage RNA-Seq in the same individual. This comparison allowed us to directly evaluate the sensitivity and specificity of RNA-Seq in identifying coding variants, and to evaluate how key parameters such as the degree of coverage and the expression levels of genes interact to influence performance. We find that although only 40% of exonic variants identified by whole genome sequencing were captured using RNA-Seq; this number rose to 81% when concentrating on genes known to be well-expressed in the source tissue. We also find that a high false positive rate can be problematic when working with RNA-Seq data, especially at higher levels of coverage. CONCLUSIONS: We conclude that as long as a tissue relevant to the trait under study is available and suitable quality control screens are implemented, RNA-Seq is a fast and inexpensive alternative approach for finding coding variants in genes with sufficiently high expression levels.

Relevância:

30.00% 30.00%

Publicador:

Resumo:

Eukaryotic genomes are mostly composed of noncoding DNA whose role is still poorly understood. Studies in several organisms have shown correlations between the length of the intergenic and genic sequences of a gene and the expression of its corresponding mRNA transcript. Some studies have found a positive relationship between intergenic sequence length and expression diversity between tissues, and concluded that genes under greater regulatory control require more regulatory information in their intergenic sequences. Other reports found a negative relationship between expression level and gene length and the interpretation was that there is selection pressure for highly expressed genes to remain small. However, a correlation between gene sequence length and expression diversity, opposite to that observed for intergenic sequences, has also been reported, and to date there is no testable explanation for this observation. To shed light on these varied and sometimes conflicting results, we performed a thorough study of the relationships between sequence length and gene expression using cell-type (tissue) specific microarray data in Arabidopsis thaliana. We measured median gene expression across tissues (expression level), expression variability between tissues (expression pattern uniformity), and expression variability between replicates (expression noise). We found that intergenic (upstream and downstream) and genic (coding and noncoding) sequences have generally opposite relationships with respect to expression, whether it is tissue variability, median, or expression noise. To explain these results we propose a model, in which the lengths of the intergenic and genic sequences have opposite effects on the ability of the transcribed region of the gene to be epigenetically regulated for differential expression. These findings could shed light on the role and influence of noncoding sequences on gene expression.

Relevância:

30.00% 30.00%

Publicador:

Resumo:

There is great interindividual variability in HIV-1 viral setpoint after seroconversion, some of which is known to be due to genetic differences among infected individuals. Here, our focus is on determining, genome-wide, the contribution of variable gene expression to viral control, and to relate it to genomic DNA polymorphism. RNA was extracted from purified CD4+ T-cells from 137 HIV-1 seroconverters, 16 elite controllers, and 3 healthy blood donors. Expression levels of more than 48,000 mRNA transcripts were assessed by the Human-6 v3 Expression BeadChips (Illumina). Genome-wide SNP data was generated from genomic DNA using the HumanHap550 Genotyping BeadChip (Illumina). We observed two distinct profiles with 260 genes differentially expressed depending on HIV-1 viral load. There was significant upregulation of expression of interferon stimulated genes with increasing viral load, including genes of the intrinsic antiretroviral defense. Upon successful antiretroviral treatment, the transcriptome profile of previously viremic individuals reverted to a pattern comparable to that of elite controllers and of uninfected individuals. Genome-wide evaluation of cis-acting SNPs identified genetic variants modulating expression of 190 genes. Those were compared to the genes whose expression was found associated with viral load: expression of one interferon stimulated gene, OAS1, was found to be regulated by a SNP (rs3177979, p = 4.9E-12); however, we could not detect an independent association of the SNP with viral setpoint. Thus, this study represents an attempt to integrate genome-wide SNP signals with genome-wide expression profiles in the search for biological correlates of HIV-1 control. It underscores the paradox of the association between increasing levels of viral load and greater expression of antiviral defense pathways. It also shows that elite controllers do not have a fully distinctive mRNA expression pattern in CD4+ T cells. Overall, changes in global RNA expression reflect responses to viral replication rather than a mechanism that might explain viral control.

Relevância:

30.00% 30.00%

Publicador:

Resumo:

A precise molecular identification of transmitted hepatitis C virus (HCV) genomes could illuminate key aspects of transmission biology, immunopathogenesis and natural history. We used single genome sequencing of 2,922 half or quarter genomes from plasma viral RNA to identify transmitted/founder (T/F) viruses in 17 subjects with acute community-acquired HCV infection. Sequences from 13 of 17 acute subjects, but none of 14 chronic controls, exhibited one or more discrete low diversity viral lineages. Sequences within each lineage generally revealed a star-like phylogeny of mutations that coalesced to unambiguous T/F viral genomes. Numbers of transmitted viruses leading to productive clinical infection were estimated to range from 1 to 37 or more (median = 4). Four acutely infected subjects showed a distinctly different pattern of virus diversity that deviated from a star-like phylogeny. In these cases, empirical analysis and mathematical modeling suggested high multiplicity virus transmission from individuals who themselves were acutely infected or had experienced a virus population bottleneck due to antiviral drug therapy. These results provide new quantitative and qualitative insights into HCV transmission, revealing for the first time virus-host interactions that successful vaccines or treatment interventions will need to overcome. Our findings further suggest a novel experimental strategy for identifying full-length T/F genomes for proteome-wide analyses of HCV biology and adaptation to antiviral drug or immune pressures.