932 resultados para medieval texts


Relevância:

20.00% 20.00%

Publicador:

Resumo:

The starting point of this article is the question "How to retrieve fingerprints of rhythm in written texts?" We address this problem in the case of Brazilian and European Portuguese. These two dialects of Modern Portuguese share the same lexicon and most of the sentences they produce are superficially identical. Yet they are conjectured, on linguistic grounds, to implement different rhythms. We show that this linguistic question can be formulated as a problem of model selection in the class of variable length Markov chains. To carry on this approach, we compare texts from European and Brazilian Portuguese. These texts are previously encoded according to some basic rhythmic features of the sentences which can be automatically retrieved. This is an entirely new approach from the linguistic point of view. Our statistical contribution is the introduction of the smallest maximizer criterion which is a constant free procedure for model selection. As a by-product, this provides a solution for the problem of optimal choice of the penalty constant when using the BIC to select a variable length Markov chain. Besides proving the consistency of the smallest maximizer criterion when the sample size diverges, we also make a simulation study comparing our approach with both the standard BIC selection and the Peres-Shields order estimation. Applied to the linguistic sample constituted for our case study, the smallest maximizer criterion assigns different context-tree models to the two dialects of Portuguese. The features of the selected models are compatible with current conjectures discussed in the linguistic literature.

Relevância:

20.00% 20.00%

Publicador:

Resumo:

The classification of texts has become a major endeavor with so much electronic material available, for it is an essential task in several applications, including search engines and information retrieval. There are different ways to define similarity for grouping similar texts into clusters, as the concept of similarity may depend on the purpose of the task. For instance, in topic extraction similar texts mean those within the same semantic field, whereas in author recognition stylistic features should be considered. In this study, we introduce ways to classify texts employing concepts of complex networks, which may be able to capture syntactic, semantic and even pragmatic features. The interplay between various metrics of the complex networks is analyzed with three applications, namely identification of machine translation (MT) systems, evaluation of quality of machine translated texts and authorship recognition. We shall show that topological features of the networks representing texts can enhance the ability to identify MT systems in particular cases. For evaluating the quality of MT texts, on the other hand, high correlation was obtained with methods capable of capturing the semantics. This was expected because the golden standards used are themselves based on word co-occurrence. Notwithstanding, the Katz similarity, which involves semantic and structure in the comparison of texts, achieved the highest correlation with the NIST measurement, indicating that in some cases the combination of both approaches can improve the ability to quantify quality in MT. In authorship recognition, again the topological features were relevant in some contexts, though for the books and authors analyzed good results were obtained with semantic features as well. Because hybrid approaches encompassing semantic and topological features have not been extensively used, we believe that the methodology proposed here may be useful to enhance text classification considerably, as it combines well-established strategies. (c) 2012 Elsevier B.V. All rights reserved.

Relevância:

20.00% 20.00%

Publicador:

Resumo:

The use of statistical methods to analyze large databases of text has been useful in unveiling patterns of human behavior and establishing historical links between cultures and languages. In this study, we identified literary movements by treating books published from 1590 to 1922 as complex networks, whose metrics were analyzed with multivariate techniques to generate six clusters of books. The latter correspond to time periods coinciding with relevant literary movements over the last five centuries. The most important factor contributing to the distinctions between different literary styles was the average shortest path length, in particular the asymmetry of its distribution. Furthermore, over time there has emerged a trend toward larger average shortest path lengths, which is correlated with increased syntactic complexity, and a more uniform use of the words reflected in a smaller power-law coefficient for the distribution of word frequency. Changes in literary style were also found to be driven by opposition to earlier writing styles, as revealed by the analysis performed with geometrical concepts. The approaches adopted here are generic and may be extended to analyze a number of features of languages and cultures.

Relevância:

20.00% 20.00%

Publicador:

Resumo:

The overall aim of the present thesis was to develop and characterise an age assessment method based on incremental lines in dental cementum using contemporary bovine teeth and teeth from archaeological faunal assemblages. The investigations also included two other age assessment methods: tooth wear pattern and macroscopic dental measurements. The first permanent mandibular molar and lower jaws from 70 contemporary cattle of known age and 170 archaeological molar sets from ten different Swedish archaeological sites were used. The following conclusions were drawn: • The number of incremental lines in the dental cementum varied between different parts of the tooth root as well as within one and the same individual. The results from contemporary cattle of known age showed a strong relationship between age and incremental lines in the cementum of the distal part of the mesial root (R2=65.5%) and the known ages of the animals. • With the “best” model variation in age could be explained to 65.5% (R2) by the number of incremental lines. Thus, the remaining age variation (approximately 35%) could not be explained by these lines. Other factors than must thus be responsible. However, with the exception of calves born the present material did not reveal any such significant relationship. • The results from cattle of known age indicate that the method of assessing age on the basis of cemental incremental lines is more reliable than other methods such as tooth wear or tooth measurements. However, by combining counting incremental lines and one variable assessing tooth dimension (tooth height) a slightly stronger relationship could be obtained (R2=74.5%). The results from age assessment of the medieval and post-Reformation cattle emphasize the importance of supplementing any age estimation of archaeological assemblages based on dental indicators with characteristics for the particular assessment model. Furthermore, conclusions based on age assessment with such models can not be drawn with any more detailed time scale than about 2 years leaving at best only 25% (R2) of factors influencing the dental indicator(s) utilized in the model unexplained. The accuracy of the age assessment required by the particular historical context in which the archaeological remains are found should thus decide what level of accuracy should be chosen.

Relevância:

20.00% 20.00%

Publicador: