892 resultados para Newspaper texts


Relevância:

20.00% 20.00%

Publicador:

Resumo:

Coordenação de Aperfeiçoamento de Pessoal de Nível Superior (CAPES)

Relevância:

20.00% 20.00%

Publicador:

Resumo:

The starting point of this article is the question "How to retrieve fingerprints of rhythm in written texts?" We address this problem in the case of Brazilian and European Portuguese. These two dialects of Modern Portuguese share the same lexicon and most of the sentences they produce are superficially identical. Yet they are conjectured, on linguistic grounds, to implement different rhythms. We show that this linguistic question can be formulated as a problem of model selection in the class of variable length Markov chains. To carry on this approach, we compare texts from European and Brazilian Portuguese. These texts are previously encoded according to some basic rhythmic features of the sentences which can be automatically retrieved. This is an entirely new approach from the linguistic point of view. Our statistical contribution is the introduction of the smallest maximizer criterion which is a constant free procedure for model selection. As a by-product, this provides a solution for the problem of optimal choice of the penalty constant when using the BIC to select a variable length Markov chain. Besides proving the consistency of the smallest maximizer criterion when the sample size diverges, we also make a simulation study comparing our approach with both the standard BIC selection and the Peres-Shields order estimation. Applied to the linguistic sample constituted for our case study, the smallest maximizer criterion assigns different context-tree models to the two dialects of Portuguese. The features of the selected models are compatible with current conjectures discussed in the linguistic literature.

Relevância:

20.00% 20.00%

Publicador:

Resumo:

The classification of texts has become a major endeavor with so much electronic material available, for it is an essential task in several applications, including search engines and information retrieval. There are different ways to define similarity for grouping similar texts into clusters, as the concept of similarity may depend on the purpose of the task. For instance, in topic extraction similar texts mean those within the same semantic field, whereas in author recognition stylistic features should be considered. In this study, we introduce ways to classify texts employing concepts of complex networks, which may be able to capture syntactic, semantic and even pragmatic features. The interplay between various metrics of the complex networks is analyzed with three applications, namely identification of machine translation (MT) systems, evaluation of quality of machine translated texts and authorship recognition. We shall show that topological features of the networks representing texts can enhance the ability to identify MT systems in particular cases. For evaluating the quality of MT texts, on the other hand, high correlation was obtained with methods capable of capturing the semantics. This was expected because the golden standards used are themselves based on word co-occurrence. Notwithstanding, the Katz similarity, which involves semantic and structure in the comparison of texts, achieved the highest correlation with the NIST measurement, indicating that in some cases the combination of both approaches can improve the ability to quantify quality in MT. In authorship recognition, again the topological features were relevant in some contexts, though for the books and authors analyzed good results were obtained with semantic features as well. Because hybrid approaches encompassing semantic and topological features have not been extensively used, we believe that the methodology proposed here may be useful to enhance text classification considerably, as it combines well-established strategies. (c) 2012 Elsevier B.V. All rights reserved.

Relevância:

20.00% 20.00%

Publicador:

Resumo:

The use of statistical methods to analyze large databases of text has been useful in unveiling patterns of human behavior and establishing historical links between cultures and languages. In this study, we identified literary movements by treating books published from 1590 to 1922 as complex networks, whose metrics were analyzed with multivariate techniques to generate six clusters of books. The latter correspond to time periods coinciding with relevant literary movements over the last five centuries. The most important factor contributing to the distinctions between different literary styles was the average shortest path length, in particular the asymmetry of its distribution. Furthermore, over time there has emerged a trend toward larger average shortest path lengths, which is correlated with increased syntactic complexity, and a more uniform use of the words reflected in a smaller power-law coefficient for the distribution of word frequency. Changes in literary style were also found to be driven by opposition to earlier writing styles, as revealed by the analysis performed with geometrical concepts. The approaches adopted here are generic and may be extended to analyze a number of features of languages and cultures.

Relevância:

20.00% 20.00%

Publicador:

Resumo:

While the use of statistical physics methods to analyze large corpora has been useful to unveil many patterns in texts, no comprehensive investigation has been performed on the interdependence between syntactic and semantic factors. In this study we propose a framework for determining whether a text (e.g., written in an unknown alphabet) is compatible with a natural language and to which language it could belong. The approach is based on three types of statistical measurements, i.e. obtained from first-order statistics of word properties in a text, from the topology of complex networks representing texts, and from intermittency concepts where text is treated as a time series. Comparative experiments were performed with the New Testament in 15 different languages and with distinct books in English and Portuguese in order to quantify the dependency of the different measurements on the language and on the story being told in the book. The metrics found to be informative in distinguishing real texts from their shuffled versions include assortativity, degree and selectivity of words. As an illustration, we analyze an undeciphered medieval manuscript known as the Voynich Manuscript. We show that it is mostly compatible with natural languages and incompatible with random texts. We also obtain candidates for keywords of the Voynich Manuscript which could be helpful in the effort of deciphering it. Because we were able to identify statistical measurements that are more dependent on the syntax than on the semantics, the framework may also serve for text analysis in language-dependent applications.

Relevância:

20.00% 20.00%

Publicador:

Resumo:

[EN] Comparative of the environmental impacts of a printed newspaper for different impacts categories using the tool of Life Cycle Assessment. The study describes the methodology, the different phases using the usual technology by coldset-offset comparing  to the new digital-inkjet printing.

Relevância:

20.00% 20.00%

Publicador:

Resumo:

In 1936, the African-American intellectual W.E.B. Du Bois visited Nazi Germany for a period of five months. Two years later, the eleven-year-long American exile of the German philosopher Theodor W. Adorno began. From the latter’s perspective, the United States was the “home” of the Culture Industry. One intuitively assumes that these sojourns abroad must have amounted to “hell on earth” for both the civil rights activist W.E.B. Du Bois and the subtle intellectual Adorno. But was this really the case? Or did they perhaps arrive at totally different conclusions? This thesis deals with these questions and attempts to make sense of the experiences of both men. By way of a systematic and comparative analysis of published texts, hitherto unpublished documents and secondary literature, this dissertation first contextualizes Du Bois’s and Adorno’s transatlantic negotiations and then depicts them. The panoply of topics with which both men concerned themselves was diverse. In Du Bois’s case it encompassed Europe, science and technology, Wagner operas, the Olympics, industrial education, race relations, National Socialism and the German Africanist Diedrich Westermann. The opinion pieces which Du Bois wrote for the newspaper “Pittsburgh Courier” during his stay in Germany serve as a major source for this thesis. In his writings on America, Adorno concentrated on what he regarded as the universally victorious Enlightenment and the predominance of mass culture. This investigation also sheds light on the correspondences between the philosopher and Max Horkheimer, Thomas Mann, Walter Benjamin, Siegfried Kracauer and Oskar and Maria Wiesengrund. In these autobiographical texts, Adorno’s thoughts revolve around such diverse topics as the American landscape, his fears as German, Jew and Left-Hegelian as well as the loneliness of the refugee. This dissertation has to refute the intuitive assumption that Du Bois’s and Adorno’s experiences abroad were horrible events for them. Both men judged the foreign countries in which they were staying in an extremely differentiated and subtle manner. Du Bois, for example, was not racially discriminated against in Germany. He was also delighted by the country’s rich cultural offerings. Adorno, for his part, praised the U.S.’s humanity of everyday life and democratic spirit. In short: Although both men partly did have to deal with utterly negative experiences, the metaphor of “hell on earth” is simply untenable as an overall conclusion.