991 resultados para Text similarity analysis
Resumo:
The aim of this study is to investigate the floristic composition of an Atlantic rain forest fragment located in Cananéia, São Paulo, Brazil, and to contribute to the knowledge on Atlantic forest through the comparative analysis of this and other floristic surveys both on the southern and southeastern Brazil, in different soil and relief types. We surveyed 215 species in 132 genera and 51 families. Classification and ordination analysis were applied to a binary matrix in order to analyze the similarity among 24 surveys, including the present one, of Atlantic forest from the south and southeast coast of Brazil. Higher floristic similarity was observed among this area and the ones where there was marine influence and more rugged relief. The surveys in areas with greater marine influence (sandy soil) were separated from those in other conditions, possibly indicating a species replacement gradient from the steep slopes towards the lowland and were probably related to different edaphic conditions. A latitudinal gradient was found among the surveys apparently confirming a continuous species replacement along the Atlantic forest, related to a restricted distribution of the species. This suggests that it is essential to preserve areas from the whole Atlantic coast. Atlantic forest distribution is quite complex and its composition cannot be adequately represented by small localized areas.
Sequence similarity analysis of Escherichia coli proteins: functional and evolutionary implications.
Resumo:
A computer analysis of 2328 protein sequences comprising about 60% of the Escherichia coli gene products was performed using methods for database screening with individual sequences and alignment blocks. A high fraction of E. coli proteins--86%--shows significant sequence similarity to other proteins in current databases; about 70% show conservation at least at the level of distantly related bacteria, and about 40% contain ancient conserved regions (ACRs) shared with eukaryotic or Archaeal proteins. For > 90% of the E. coli proteins, either functional information or sequence similarity, or both, are available. Forty-six percent of the E. coli proteins belong to 299 clusters of paralogs (intraspecies homologs) defined on the basis of pairwise similarity. Another 10% could be included in 70 superclusters using motif detection methods. The majority of the clusters contain only two to four members. In contrast, nearly 25% of all E. coli proteins belong to the four largest superclusters--namely, permeases, ATPases and GTPases with the conserved "Walker-type" motif, helix-turn-helix regulatory proteins, and NAD(FAD)-binding proteins. We conclude that bacterial protein sequences generally are highly conserved in evolution, with about 50% of all ACR-containing protein families represented among the E. coli gene products. With the current sequence databases and methods of their screening, computer analysis yields useful information on the functions and evolutionary relationships of the vast majority of genes in a bacterial genome. Sequence similarity with E. coli proteins allows the prediction of functions for a number of important eukaryotic genes, including several whose products are implicated in human diseases.
Resumo:
Quantum molecular similarity (QMS) techniques are used to assess the response of the electron density of various small molecules to application of a static, uniform electric field. Likewise, QMS is used to analyze the changes in electron density generated by the process of floating a basis set. The results obtained show an interrelation between the floating process, the optimum geometry, and the presence of an external field. Cases involving the Le Chatelier principle are discussed, and an insight on the changes of bond critical point properties, self-similarity values and density differences is performed
Resumo:
Quantum molecular similarity (QMS) techniques are used to assess the response of the electron density of various small molecules to application of a static, uniform electric field. Likewise, QMS is used to analyze the changes in electron density generated by the process of floating a basis set. The results obtained show an interrelation between the floating process, the optimum geometry, and the presence of an external field. Cases involving the Le Chatelier principle are discussed, and an insight on the changes of bond critical point properties, self-similarity values and density differences is performed
Resumo:
HTLV-1 is endemic in Brazil and HIV/ HTLV-1 coinfection has been detected, mostly in the northeast region. Cosmopolitan HTLV-1a is the main subtype that circulates in Brazil. This study characterized 17 HTLV-1 isolates from HIV coinfected patients of southern (n = 7) and southeastern (n = 10) Brazil. HTLV-1 provirus DNA was amplified by nested PCR (env and LTR) and sequenced. Env sequences (705 bp) from 15 isolates and LTR sequences (731 bp) from 17 isolates showed 99.5% and 98.8% similarity among sequences, respectively. Comparing these sequences with ATK (HTLV-1a) and Mel5 (HTLV-1c) prototypes, similarities of 99% and 97.4%, respectively, for env and LTR with ATK, and 91.6% and 90.3% with Mel5, were detected. Phylogenetic analysis showed that all sequences belonged to the transcontinental subgroup A of the Cosmopolitan subtype, clustering in two Latin American clusters.
Resumo:
Les cadriciels et les bibliothèques sont indispensables aux systèmes logiciels d'aujourd'hui. Quand ils évoluent, il est souvent fastidieux et coûteux pour les développeurs de faire la mise à jour de leur code. Par conséquent, des approches ont été proposées pour aider les développeurs à migrer leur code. Généralement, ces approches ne peuvent identifier automatiquement les règles de modification une-remplacée-par-plusieurs méthodes et plusieurs-remplacées-par-une méthode. De plus, elles font souvent un compromis entre rappel et précision dans leur résultats en utilisant un ou plusieurs seuils expérimentaux. Nous présentons AURA (AUtomatic change Rule Assistant), une nouvelle approche hybride qui combine call dependency analysis et text similarity analysis pour surmonter ces limitations. Nous avons implanté AURA en Java et comparé ses résultats sur cinq cadriciels avec trois approches précédentes par Dagenais et Robillard, M. Kim et al., et Schäfer et al. Les résultats de cette comparaison montrent que, en moyenne, le rappel de AURA est 53,07% plus que celui des autre approches avec une précision similaire (0,10% en moins).
Resumo:
Arguably, the most difficult task in text classification is to choose an appropriate set of features that allows machine learning algorithms to provide accurate classification. Most state-of-the-art techniques for this task involve careful feature engineering and a pre-processing stage, which may be too expensive in the emerging context of massive collections of electronic texts. In this paper, we propose efficient methods for text classification based on information-theoretic dissimilarity measures, which are used to define dissimilarity-based representations. These methods dispense with any feature design or engineering, by mapping texts into a feature space using universal dissimilarity measures; in this space, classical classifiers (e.g. nearest neighbor or support vector machines) can then be used. The reported experimental evaluation of the proposed methods, on sentiment polarity analysis and authorship attribution problems, reveals that it approximates, sometimes even outperforms previous state-of-the-art techniques, despite being much simpler, in the sense that they do not require any text pre-processing or feature engineering.
Resumo:
Drug resistance is one of the major concerns regarding tuberculosis (TB) infection worldwide because it hampers control of the disease. Understanding the underlying mechanisms responsible for drug resistance development is of the highest importance. To investigate clinical data from drug-resistant TB patients at the Tropical Diseases Hospital, Goiás (GO), Brazil and to evaluate the molecular basis of rifampin (R) and isoniazid (H) resistance in Mycobacterium tuberculosis. Drug susceptibility testing was performed on 124 isolates from 100 patients and 24 isolates displayed resistance to R and/or H. Molecular analysis of drug resistance was performed by partial sequencing of the rpoB and katGgenes and analysis of the inhA promoter region. Similarity analysis of isolates was performed by 15 loci mycobacterial interspersed repetitive unit-variable number tandem repeat (MIRU-VNTR) typing. The molecular basis of drug resistance among the 24 isolates from 16 patients was confirmed in 18 isolates. Different susceptibility profiles among the isolates from the same individual were observed in five patients; using MIRU-VNTR, we have shown that those isolates were not genetically identical, with differences in one to three loci within the 15 analysed loci. Drug-resistant TB in GO is caused by M. tuberculosis strains with mutations in previously described sites of known genes and some patients harbour a mixed phenotype infection as a consequence of a single infective event; however, further and broader investigations are needed to support our findings.
Resumo:
The objective of this work was to confirm the hybrids obtained in plants originated from the crossing between the mandarins Citrus deliciosa 'Montenegrina' and C. nobilis 'King', and to estimate the genetic similarity among hybrids, and between each hybrid and its parents. Twenty‑three pairs of microsatellite primers were tested. Fourteen of these pairs showed polymorphic bands between parents. Primers CCSM 129 and CCSME 52 were sufficient to identify the 12 nucellar clones observed in the studied population. Genetic similarity analysis of the population (hybrids and parents) showed 0.56 average similarity. Besides the 12 clones of 'Montenegrina' identified, 25 hybrids were found of which D18, C32, D06, C05 and D09 are the more similar to 'Montenegrina'.
Resumo:
Pesiqta Rabbati is a unique homiletic midrash that follows the liturgical calendar in its presentation of homilies for festivals and special Sabbaths. This article attempts to utilize Pesiqta Rabbati in order to present a global theory of the literary production of rabbinic/homiletic literature. In respect to Pesiqta Rabbati it explores such areas as dating, textual witnesses, integrative apocalyptic meta-narrative, describing and mapping the structure of the text, internal and external constraints that impacted upon the text, text linguistic analysis, form-analysis: problems in the texts and linguistic gap-filling, transmission of text, strict formalization of a homiletic unit, deconstructing and reconstructing homiletic midrashim based upon form-analytic units of the homily, Neusner’s documentary hypothesis, surface structures of the homiletic unit, and textual variants. The suggested methodology may assist scholars in their production of editions of midrashic works by eliminating superfluous material and in their decoding and defining of ancient texts.
Resumo:
A preliminary survey of the spider fauna in natural and artificial forest gap formations at Porto Urucu, a petroleum/natural gas production facility in the Urucu river basin, Coari, Amazonas, Brazil is presented. Sampling was conducted both occasionally and using a protocol composed of a suite of techniques: beating trays (32 samples), nocturnal manual samplings (48), sweeping nets (16), Winkler extractors (24), and pitfall traps (120). A total of 4201 spiders, belonging to 43 families and 393 morphospecies, were collected during the dry season, in July, 2003. Excluding the occasional samples, the observed richness was 357 species. In a performance test of seven species richness estimators, the Incidence Based Coverage Estimator (ICE) was the best fit estimator, with 639 estimated species. To evaluate differences in species richness associated with natural and artificial gaps, samples from between the center of the gaps up to 300 meters inside the adjacent forest matrix were compared through the inspection of the confidence intervals of individual-based rarefaction curves for each treatment. The observed species richness was significantly higher in natural gaps combined with adjacent forest than in the artificial gaps combined with adjacent forest. Moreover, a community similarity analysis between the fauna collected under both treatments demonstrated that there were considerable differences in species composition. The significantly higher abundance of Lycosidae in artificial gap forest is explained by the presence of herbaceous vegetation in the gaps themselves. Ctenidae was significantly more abundant in the natural gap forest, probable due to the increase of shelter availability provided by the fallen trees in the gaps themselves. Both families are identified as potential indicators of environmental change related to the establishment or recovery of artificial gaps in the study area.
Resumo:
Terrestrial isopods are important and dominant component of meso and macrodecomposer soil communities. The present study investigates the diversity and species composition of terrestrial isopods on three forests on the Serra Geral of the state of Rio Grande do Sul, Brazil. The area has two natural formations (Primary Woodland and Secondary Woodland) and one plantation of introduced Pinus. The pitfall traps operated from March 2001 to May 2002, with two summer periods and one winter. There were 14 sampling dates overall. Of the five species found: Alboscia silveirensis Araujo, 1999, Atlantoscia floridana (van Name, 1940), Benthana araucariana Araujo & Lopes, 2003 (Philoscidae), Balloniscus glaber Araujo & Zardo, 1995 (Balloniscidae) and Styloniscus otakensis (Chilton, 1901) (Styloniscidae); only A. floridana is abundant on all environments and B. glaber is nearly exclusive for the native forests. The obtained data made it possible to infer about population characteristics of this species. The Similarity Analysis showed a quantitative difference among the Secondary forest and Pinus plantation, but not a qualitative one. The operational sex ratio (OSR) analysis for A. floridana does not reveal significant differences in male and female proportions among environments. The reproductive period identified in the present study for A. floridana was from spring to autumn in the primary forest and Pinus plantation and during all year for the secondary forest. The OSR analysis for B. glaber reveals no significant differences in abundance between males and females for secondary forest, but the primary forest was a significant difference. The reproductive period for B. glaber extended from summer to autumn (for primary and secondary forest). This is the first record for Brazil of an established terrestrial isopod population in a Pinus sp. plantation area, evidenced by the presence of young, adults and ovigerous females, balanced sex ratio, expected fecundity and reproduction pattern, as compared to populations from native vegetation areas.
Resumo:
Molecular characterization of Paracoccidioides brasiliensis variant strains that had been preserved under mineral oil for decades was carried out by random amplified polymorphic DNA analysis (RAPD). On P. brasiliensis variants in the transitional phase and strains with typical morphology, RAPD produced reproducible polymorphic amplification products that differentiated them. A dendrogram based on the generated RAPD patterns placed the 14 P. brasiliensis strains into five groups with similarity coefficients of 72%. A high correlation between the genotypic and phenotypic characteristics of the strains was observed. A 750 bp-RAPD fragment found only in the wild-type phenotype strains was cloned and sequenced. Genetic similarity analysis using BLASTx suggested that this RAPD marker represents a putative domain of a hypothetical flavin-binding monooxygenase (FMO)-like protein of Neurospora crassa.
Resumo:
In order to obtain a high-resolution Pleistocene stratigraphy, eleven continuouslycored boreholes, 100 to 220m deep were drilled in the northern part of the PoPlain by Regione Lombardia in the last five years. Quantitative provenanceanalysis (QPA, Weltje and von Eynatten, 2004) of Pleistocene sands was carriedout by using multivariate statistical analysis (principal component analysis, PCA,and similarity analysis) on an integrated data set, including high-resolution bulkpetrography and heavy-mineral analyses on Pleistocene sands and of 250 majorand minor modern rivers draining the southern flank of the Alps from West toEast (Garzanti et al, 2004; 2006). Prior to the onset of major Alpine glaciations,metamorphic and quartzofeldspathic detritus from the Western and Central Alpswas carried from the axial belt to the Po basin longitudinally parallel to theSouthAlpine belt by a trunk river (Vezzoli and Garzanti, 2008). This scenariorapidly changed during the marine isotope stage 22 (0.87 Ma), with the onset ofthe first major Pleistocene glaciation in the Alps (Muttoni et al, 2003). PCA andsimilarity analysis from core samples show that the longitudinal trunk river at thistime was shifted southward by the rapid southward and westward progradation oftransverse alluvial river systems fed from the Central and Southern Alps.Sediments were transported southward by braided river systems as well as glacialsediments transported by Alpine valley glaciers invaded the alluvial plain.Kew words: Detrital modes; Modern sands; Provenance; Principal ComponentsAnalysis; Similarity, Canberra Distance; palaeodrainage