874 resultados para Information retrieval - Australia
Resumo:
Except the article forming the main content most HTML documents on the WWW contain additional contents such as navigation menus, design elements or commercial banners. In the context of several applications it is necessary to draw the distinction between main and additional content automatically. Content extraction and template detection are the two approaches to solve this task. This thesis gives an extensive overview of existing algorithms from both areas. It contributes an objective way to measure and evaluate the performance of content extraction algorithms under different aspects. These evaluation measures allow to draw the first objective comparison of existing extraction solutions. The newly introduced content code blurring algorithm overcomes several drawbacks of previous approaches and proves to be the best content extraction algorithm at the moment. An analysis of methods to cluster web documents according to their underlying templates is the third major contribution of this thesis. In combination with a localised crawling process this clustering analysis can be used to automatically create sets of training documents for template detection algorithms. As the whole process can be automated it allows to perform template detection on a single document, thereby combining the advantages of single and multi document algorithms.
Resumo:
In questo lavoro si introducono i concetti di base di Natural Language Processing, soffermandosi su Information Extraction e analizzandone gli ambiti applicativi, le attività principali e la differenza rispetto a Information Retrieval. Successivamente si analizza il processo di Named Entity Recognition, focalizzando l’attenzione sulle principali problematiche di annotazione di testi e sui metodi per la valutazione della qualità dell’estrazione di entità. Infine si fornisce una panoramica della piattaforma software open-source di language processing GATE/ANNIE, descrivendone l’architettura e i suoi componenti principali, con approfondimenti sugli strumenti che GATE offre per l'approccio rule-based a Named Entity Recognition.
Resumo:
The our reality is characterized by a constant progress and, to follow that, people need to stay up to date on the events. In a world with a lot of existing news, search for the ideal ones may be difficult, because the obstacles that make it arduous will be expanded more and more over time, due to the enrichment of data. In response, a great help is given by Information Retrieval, an interdisciplinary branch of computer science that deals with the management and the retrieval of the information. An IR system is developed to search for contents, contained in a reference dataset, considered relevant with respect to the need expressed by an interrogative query. To satisfy these ambitions, we must consider that most of the developed IR systems rely solely on textual similarity to identify relevant information, defining them as such when they include one or more keywords expressed by the query. The idea studied here is that this is not always sufficient, especially when it's necessary to manage large databases, as is the web. The existing solutions may generate low quality responses not allowing, to the users, a valid navigation through them. The intuition, to overcome these limitations, has been to define a new concept of relevance, to differently rank the results. So, the light was given to Temporal PageRank, a new proposal for the Web Information Retrieval that relies on a combination of several factors to increase the quality of research on the web. Temporal PageRank incorporates the advantages of a ranking algorithm, to prefer the information reported by web pages considered important by the context itself in which they reside, and the potential of techniques belonging to the world of the Temporal Information Retrieval, exploiting the temporal aspects of data, describing their chronological contexts. In this thesis, the new proposal is discussed, comparing its results with those achieved by the best known solutions, analyzing its strengths and its weaknesses.
Resumo:
It has long been known that trypanosomes regulate mitochondrial biogenesis during the life cycle of the parasite; however, the mitochondrial protein inventory (MitoCarta) and its regulation remain unknown. We present a novel computational method for genome-wide prediction of mitochondrial proteins using a support vector machine-based classifier with approximately 90% prediction accuracy. Using this method, we predicted the mitochondrial localization of 468 proteins with high confidence and have experimentally verified the localization of a subset of these proteins. We then applied a recently developed parallel sequencing technology to determine the expression profiles and the splicing patterns of a total of 1065 predicted MitoCarta transcripts during the development of the parasite, and showed that 435 of the transcripts significantly changed their expressions while 630 remain unchanged in any of the three life stages analyzed. Furthermore, we identified 298 alternatively splicing events, a small subset of which could lead to dual localization of the corresponding proteins.
Resumo:
A series of oligodeoxyribonucleotides and oligoribonucleotides containing single and multiple tricyclo(tc)-nucleosides in various arrangements were prepared and the thermal and thermodynamic transition profiles of duplexes with complementary DNA and RNA evaluated. Tc-residues aligned in a non-continuous fashion in an RNA strand significantly decrease affinity to complementary RNA and DNA, mostly as a consequence of a loss of pairing enthalpy DeltaH. Arranging the tc-residues in a continuous fashion rescues T(m) and leads to higher DNA and RNA affinity. Substitution of oligodeoxyribonucleotides in the same way causes much less differences in T(m) when paired to complementary DNA and leads to substantial increases in T(m) when paired to complementary RNA. CD-spectroscopic investigations in combination with molecular dynamics simulations of duplexes with single modifications show that tc-residues in the RNA backbone distinctly influence the conformation of the neighboring nucleotides forcing them into higher energy conformations, while tc-residues in the DNA backbone seem to have negligible influence on the nearest neighbor conformations. These results rationalize the observed affinity differences and are of relevance for the design of tc-DNA containing oligonucleotides for applications in antisense or RNAi therapy.
Resumo:
The synthesis of a caged RNA phosphoramidite building block containing the oxidatively damaged base 5-hydroxycytidine (5-HOrC) has been accomplished. To determine the effect of this highly mutagenic lesion on complementary base recognition and coding properties, this building block was incorporated into a 12-mer oligoribonucleotide for Tm and CD measurements and a 31-mer template strand for primer extension experiments with HIV-, AMV- and MMLV-reverse transcriptase (RT). In UV-melting experiments, we find an unusual biphasic transition with two distinct Tm's when 5-HOrC is paired against a DNA or RNA complement with the base guanine in opposing position. The higher Tm closely matches that of a C-G base pair while the lower is close to that of a C-A mismatch. In single nucleotide extension reactions, we find substantial misincorporation of dAMP and to a lesser extent dTMP, with dAMP almost equaling that of the parent dGMP in the case of HIV-RT. A working hypothesis for the biphasic melting transition does not invoke tautomeric variability of 5-HOrC but rather local structural perturbations of the base pair at low temperature induced by interactions of the 5-HO group with the phosphate backbone. The properties of this RNA damage is discussed in the context of its putative biological function.
Resumo:
The straightforward production and dose-controlled administration of protein therapeutics remain major challenges for the biopharmaceutical manufacturing and gene therapy communities. Transgenes linked to HIV-1-derived vpr and pol-based protease cleavage (PC) sequences were co-produced as chimeric fusion proteins in a lentivirus production setting, encapsidated and processed to fusion peptide-free native protein in pseudotyped lentivirions for intracellular delivery and therapeutic action in target cells. Devoid of viral genome sequences, protein-transducing nanoparticles (PTNs) enabled transient and dose-dependent delivery of therapeutic proteins at functional quantities into a variety of mammalian cells in the absence of host chromosome modifications. PTNs delivering Manihot esculenta linamarase into rodent or human, tumor cell lines and spheroids mediated hydrolysis of the innocuous natural prodrug linamarin to cyanide and resulted in efficient cell killing. Following linamarin injection into nude mice, linamarase-transducing nanoparticles impacted solid tumor development through the bystander effect of cyanide.