7 resultados para ancient Basque texts
em Biblioteca Digital da Produção Intelectual da Universidade de São Paulo
Resumo:
The starting point of this article is the question "How to retrieve fingerprints of rhythm in written texts?" We address this problem in the case of Brazilian and European Portuguese. These two dialects of Modern Portuguese share the same lexicon and most of the sentences they produce are superficially identical. Yet they are conjectured, on linguistic grounds, to implement different rhythms. We show that this linguistic question can be formulated as a problem of model selection in the class of variable length Markov chains. To carry on this approach, we compare texts from European and Brazilian Portuguese. These texts are previously encoded according to some basic rhythmic features of the sentences which can be automatically retrieved. This is an entirely new approach from the linguistic point of view. Our statistical contribution is the introduction of the smallest maximizer criterion which is a constant free procedure for model selection. As a by-product, this provides a solution for the problem of optimal choice of the penalty constant when using the BIC to select a variable length Markov chain. Besides proving the consistency of the smallest maximizer criterion when the sample size diverges, we also make a simulation study comparing our approach with both the standard BIC selection and the Peres-Shields order estimation. Applied to the linguistic sample constituted for our case study, the smallest maximizer criterion assigns different context-tree models to the two dialects of Portuguese. The features of the selected models are compatible with current conjectures discussed in the linguistic literature.
Resumo:
The classification of texts has become a major endeavor with so much electronic material available, for it is an essential task in several applications, including search engines and information retrieval. There are different ways to define similarity for grouping similar texts into clusters, as the concept of similarity may depend on the purpose of the task. For instance, in topic extraction similar texts mean those within the same semantic field, whereas in author recognition stylistic features should be considered. In this study, we introduce ways to classify texts employing concepts of complex networks, which may be able to capture syntactic, semantic and even pragmatic features. The interplay between various metrics of the complex networks is analyzed with three applications, namely identification of machine translation (MT) systems, evaluation of quality of machine translated texts and authorship recognition. We shall show that topological features of the networks representing texts can enhance the ability to identify MT systems in particular cases. For evaluating the quality of MT texts, on the other hand, high correlation was obtained with methods capable of capturing the semantics. This was expected because the golden standards used are themselves based on word co-occurrence. Notwithstanding, the Katz similarity, which involves semantic and structure in the comparison of texts, achieved the highest correlation with the NIST measurement, indicating that in some cases the combination of both approaches can improve the ability to quantify quality in MT. In authorship recognition, again the topological features were relevant in some contexts, though for the books and authors analyzed good results were obtained with semantic features as well. Because hybrid approaches encompassing semantic and topological features have not been extensively used, we believe that the methodology proposed here may be useful to enhance text classification considerably, as it combines well-established strategies. (c) 2012 Elsevier B.V. All rights reserved.
Resumo:
We use a recently developed computerized modeling technique to explore the long-term impacts of indigenous Amazonian hunting in the past, present, and future. The model redefines sustainability in spatial and temporal terms, a major advance over the static "sustainability indices" currently used to study hunting in tropical forests. We validate the model's projections against actual field data from two sites in contemporary Amazonia and use the model to assess various management scenarios for the future of Manu National Park in Peru. We then apply the model to two archaeological contexts, show how its results may resolve long-standing enigmas regarding native food taboos and primate biogeography, and reflect on the ancient history and future of indigenous people in the Amazon.
Resumo:
The use of statistical methods to analyze large databases of text has been useful in unveiling patterns of human behavior and establishing historical links between cultures and languages. In this study, we identified literary movements by treating books published from 1590 to 1922 as complex networks, whose metrics were analyzed with multivariate techniques to generate six clusters of books. The latter correspond to time periods coinciding with relevant literary movements over the last five centuries. The most important factor contributing to the distinctions between different literary styles was the average shortest path length, in particular the asymmetry of its distribution. Furthermore, over time there has emerged a trend toward larger average shortest path lengths, which is correlated with increased syntactic complexity, and a more uniform use of the words reflected in a smaller power-law coefficient for the distribution of word frequency. Changes in literary style were also found to be driven by opposition to earlier writing styles, as revealed by the analysis performed with geometrical concepts. The approaches adopted here are generic and may be extended to analyze a number of features of languages and cultures.
Resumo:
Congenital gonadotropin-releasing hormone (GnRH) deficiency manifests as absent or incomplete sexual maturation and infertility. Although the disease exhibits marked locus and allelic heterogeneity, with the causal mutations being both rare and private, one causal mutation in the prokineticin receptor, PROKR2 L173R, appears unusually prevalent among GnRH-deficient patients of diverse geographic and ethnic origins. To track the genetic ancestry of PROKR2 L173R, haplotype mapping was performed in 22 unrelated patients with GnRH deficiency carrying L173R and their 30 first-degree relatives. The mutations age was estimated using a haplotype-decay model. Thirteen subjects were informative and in all of them the mutation was present on the same approximate to 123 kb haplotype whose population frequency is 10. Thus, PROKR2 L173R represents a founder mutation whose age is estimated at approximately 9000 years. Inheritance of PROKR2 L173R-associated GnRH deficiency was complex with highly variable penetrance among carriers, influenced by additional mutations in the other PROKR2 allele (recessive inheritance) or another gene (digenicity). The paradoxical identification of an ancient founder mutation that impairs reproduction has intriguing implications for the inheritance mechanisms of PROKR2 L173R-associated GnRH deficiency and for the relevant processes of evolutionary selection, including potential selective advantages of mutation carriers in genes affecting reproduction.
Resumo:
Este artigo pretende realçar a importância desempenhada pela obra de Henry James Sumner Maine na formação da antropologia jurídica e da sociologia do direito. Mediante a recuperação da tese central de sua obra Ancient Law, procura ressaltar o papel por ela desempenhado no delineamento de uma nova forma de abordagem da relação entre direito e sociedade. Para tanto, recupera traços gerais das análises de Norbert Rouland e de Niklas Luhmann acerca da importância das ideias do autor.
Resumo:
While the use of statistical physics methods to analyze large corpora has been useful to unveil many patterns in texts, no comprehensive investigation has been performed on the interdependence between syntactic and semantic factors. In this study we propose a framework for determining whether a text (e.g., written in an unknown alphabet) is compatible with a natural language and to which language it could belong. The approach is based on three types of statistical measurements, i.e. obtained from first-order statistics of word properties in a text, from the topology of complex networks representing texts, and from intermittency concepts where text is treated as a time series. Comparative experiments were performed with the New Testament in 15 different languages and with distinct books in English and Portuguese in order to quantify the dependency of the different measurements on the language and on the story being told in the book. The metrics found to be informative in distinguishing real texts from their shuffled versions include assortativity, degree and selectivity of words. As an illustration, we analyze an undeciphered medieval manuscript known as the Voynich Manuscript. We show that it is mostly compatible with natural languages and incompatible with random texts. We also obtain candidates for keywords of the Voynich Manuscript which could be helpful in the effort of deciphering it. Because we were able to identify statistical measurements that are more dependent on the syntax than on the semantics, the framework may also serve for text analysis in language-dependent applications.