85 resultados para CATEGORICAL-DATA
em Biblioteca Digital da Produção Intelectual da Universidade de São Paulo (BDPI/USP)
Resumo:
P>In the context of either Bayesian or classical sensitivity analyses of over-parametrized models for incomplete categorical data, it is well known that prior-dependence on posterior inferences of nonidentifiable parameters or that too parsimonious over-parametrized models may lead to erroneous conclusions. Nevertheless, some authors either pay no attention to which parameters are nonidentifiable or do not appropriately account for possible prior-dependence. We review the literature on this topic and consider simple examples to emphasize that in both inferential frameworks, the subjective components can influence results in nontrivial ways, irrespectively of the sample size. Specifically, we show that prior distributions commonly regarded as slightly informative or noninformative may actually be too informative for nonidentifiable parameters, and that the choice of over-parametrized models may drastically impact the results, suggesting that a careful examination of their effects should be considered before drawing conclusions.Resume Que ce soit dans un cadre Bayesien ou classique, il est bien connu que la surparametrisation, dans les modeles pour donnees categorielles incompletes, peut conduire a des conclusions erronees. Cependant, certains auteurs persistent a negliger les problemes lies a la presence de parametres non identifies. Nous passons en revue la litterature dans ce domaine, et considerons quelques exemples surparametres simples dans lesquels les elements subjectifs influencent de facon non negligeable les resultats, independamment de la taille des echantillons. Plus precisement, nous montrons comment des a priori consideres comme peu ou non-informatifs peuvent se reveler extremement informatifs en ce qui concerne les parametres non identifies, et que le recours a des modeles surparametres peut avoir sur les conclusions finales un impact considerable. Ceci suggere un examen tres attentif de l`impact potentiel des a priori.
Resumo:
We review some issues related to the implications of different missing data mechanisms on statistical inference for contingency tables and consider simulation studies to compare the results obtained under such models to those where the units with missing data are disregarded. We confirm that although, in general, analyses under the correct missing at random and missing completely at random models are more efficient even for small sample sizes, there are exceptions where they may not improve the results obtained by ignoring the partially classified data. We show that under the missing not at random (MNAR) model, estimates on the boundary of the parameter space as well as lack of identifiability of the parameters of saturated models may be associated with undesirable asymptotic properties of maximum likelihood estimators and likelihood ratio tests; even in standard cases the bias of the estimators may be low only for very large samples. We also show that the probability of a boundary solution obtained under the correct MNAR model may be large even for large samples and that, consequently, we may not always conclude that a MNAR model is misspecified because the estimate is on the boundary of the parameter space.
Resumo:
When missing data occur in studies designed to compare the accuracy of diagnostic tests, a common, though naive, practice is to base the comparison of sensitivity, specificity, as well as of positive and negative predictive values on some subset of the data that fits into methods implemented in standard statistical packages. Such methods are usually valid only under the strong missing completely at random (MCAR) assumption and may generate biased and less precise estimates. We review some models that use the dependence structure of the completely observed cases to incorporate the information of the partially categorized observations into the analysis and show how they may be fitted via a two-stage hybrid process involving maximum likelihood in the first stage and weighted least squares in the second. We indicate how computational subroutines written in R may be used to fit the proposed models and illustrate the different analysis strategies with observational data collected to compare the accuracy of three distinct non-invasive diagnostic methods for endometriosis. The results indicate that even when the MCAR assumption is plausible, the naive partial analyses should be avoided.
Resumo:
Objective: to describe women`s feelings about mode of birth. Design: exploratory descriptive design. Semi-structured interviews were conducted using a questionnaire that had been developed previously (categorical data and open-and closed-ended questions). Qualitative analysis of the results was performed through a context analysis technique. Setting: the largest public university hospital in Brazil. Participants: 48 women in their third trimester of pregnancy. Findings: most women expressed a preference for vaginal birth, as they perceived that they would have a faster recovery. Women who expressed a preference for caesarean section did so because of lack of pain during the birth and the need for tubal sterilisation. The majority of women considered it important to have experience with a mode of birth in order to choose a preference. Complications associated with maternal illness were very influential in the decision-making process. Key conclusions: these results provide a useful first step towards the identification of aspects of women`s feelings about modes of birth. Most women expressed a preference for vaginal birth. Further exploration of women`s feelings regarding parturition and the decision-making process is required. (C) 2008 Elsevier Ltd. All rights reserved.
Resumo:
Experimental animal studies have shown that nicotine exposure during gestation alters the expression of fetal hypothalamic neuropeptides involved in the control of appetite. We aimed to determine whether the exposure to maternal smoking during gestation in humans is associated with an altered feeding behavior of the adult offspring. A longitudinal prospective cohort study was conducted including all births from Ribeirao Preto (Sao Paulo, Brazil) between 1978 and 1979. At 24 years of age, a representative random sample was re-evaluated and divided into groups exposed (n = 424) or not (n = 1586) to maternal smoking during gestation. Feeding behavior was analyzed using a food frequency questionnaire. Covariance analysis was used for continuous data and the chi(2) test for categorical data. Results were adjusted for birth weight ratio, body mass index, gender, physical activity and smoking, as well as maternal and subjects` schooling. Individuals exposed to maternal smoking during gestation ate more carbohydrates than proteins (as per the carbohydrate-to-protein ratio) than non-exposed individuals. There were no differences in the consumption of the macronutrients themselves. We propose that this adverse fetal life event programs the individual`s physiology and metabolism persistently, leading to an altered feeding behavior that could contribute to the development of chronic diseases in the long term.
Resumo:
This article presents important properties of standard discrete distributions and its conjugate densities. The Bernoulli and Poisson processes are described as generators of such discrete models. A characterization of distributions by mixtures is also introduced. This article adopts a novel singular notation and representation. Singular representations are unusual in statistical texts. Nevertheless, the singular notation makes it simpler to extend and generalize theoretical results and greatly facilitates numerical and computational implementation.
Resumo:
Advances in diagnostic research are moving towards methods whereby the periodontal risk can be identified and quantified by objective measures using biomarkers. Patients with periodontitis may have elevated circulating levels of specific inflammatory markers that can be correlated to the severity of the disease. The purpose of this study was to evaluate whether differences in the serum levels of inflammatory biomarkers are differentially expressed in healthy and periodontitis patients. Twenty-five patients (8 healthy patients and 17 chronic periodontitis patients) were enrolled in the study. A 15 mL blood sample was used for identification of the inflammatory markers, with a human inflammatory flow cytometry multiplex assay. Among 24 assessed cytokines, only 3 (RANTES, MIG and Eotaxin) were statistically different between groups (p<0.05). In conclusion, some of the selected markers of inflammation are differentially expressed in healthy and periodontitis patients. Cytokine profile analysis may be further explored to distinguish the periodontitis patients from the ones free of disease and also to be used as a measure of risk. The present data, however, are limited and larger sample size studies are required to validate the findings of the specific biomarkers.
Resumo:
ABSTRACT Microphysical and thermodynamical features of two tropical systems, namely Hurricane Ivan and Typhoon Conson, and one sub-tropical, Catarina, have been analyzed based on space-born radar PR measurements available on the TRMM satellite. The procedure to classify the reflectivity profiles followed the Heymsfield et al (2000) and Steiner et al (1995) methodologies. The water and ice content have been calculated using a relationship obtained with data of the surface SPOL radar and PR in Rondonia State in Brazil. The diabatic heating rate due to latent heat release has been estimated using the methodology developed by Tao et al (1990). A more detailed analysis has been performed for Hurricane Catarina, the first of its kind in South Atlantic. High water content mean value has been found in Conson and Ivan at low levels and close to their centers. Results indicate that hurricane Catarina was shallower than the other two systems, with less water and the water was concentrated closer to its center. The mean ice content in Catarina was about 0.05 g kg-1 while in Conson it was 0.06 g kg-1 and in Ivan 0.08 g kg-1. Conson and Ivan had water content up to 0.3 g kg-1 above the 0ºC layer, while Catarina had less than 0.15 g kg-1. The latent heat released by Catarina showed to be very similar to the other two systems, except in the regions closer to the center.
Resumo:
Lepidocharax, new genus, and Lepidocharax diamantina and L. burnsi new species from eastern Brazil are described herein. Lepidocharax is considered a monophyletic genus of the Stevardiinae and can be distinguished from the other members of this subfamily except Planaltina, Pseudocorynopoma, and Xenurobrycon by having the dorsal-fin origin vertically aligned with the anal-fin origin, vs. dorsal fin origin anterior or posterior to anal-fin origin. Additionally the new genus can be distinguished from those three genera by not having the scales extending over the ventral caudal-fin lobe modified to form the dorsal border of the pheromone pouch organ or to represent a pouch scale in sexually mature males. In this paper, we describe these two recently discovered species and the ultrastructure of their spermatozoa.
Resumo:
During the exploration and mapping of new caves in Serra do Ramalho karst area, southern Bahia state, cavers from the Grupo Bambuí de Pesquisas Espeleológicas - GBPE (Belo Horizonte) noticed the presence of troglomorphic catfishes (species with reduced eyes and/or melanic pigmentation), which we intensively investigated with regards to their ecology and behavior since 2005. Non-troglomorphic fishes regularly found in the studied caves were included in this investigation. We present here data on the natural history of two troglobitic (exclusively subterranean troglomorphic species) fishes - Rhamdia enfurnada Bichuette & Trajano, 2005 (Heptapteridae; Gruna do Enfurnado) and Trichomycterus undescribed species (Trichomycteridae; Lapa dos Peixes and Gruna da Água Clara), and non-troglomorphic Hoplias cf. malabaricus, probably a troglophile (able to form populations both in epigean and subterranean habitats) in the Gruna do Enfurnado, and Pimelodella sp., a species with a sink population in the Lapa dos Peixes.
Resumo:
Geographic Data Warehouses (GDW) are one of the main technologies used in decision-making processes and spatial analysis, and the literature proposes several conceptual and logical data models for GDW. However, little effort has been focused on studying how spatial data redundancy affects SOLAP (Spatial On-Line Analytical Processing) query performance over GDW. In this paper, we investigate this issue. Firstly, we compare redundant and non-redundant GDW schemas and conclude that redundancy is related to high performance losses. We also analyze the issue of indexing, aiming at improving SOLAP query performance on a redundant GDW. Comparisons of the SB-index approach, the star-join aided by R-tree and the star-join aided by GiST indicate that the SB-index significantly improves the elapsed time in query processing from 25% up to 99% with regard to SOLAP queries defined over the spatial predicates of intersection, enclosure and containment and applied to roll-up and drill-down operations. We also investigate the impact of the increase in data volume on the performance. The increase did not impair the performance of the SB-index, which highly improved the elapsed time in query processing. Performance tests also show that the SB-index is far more compact than the star-join, requiring only a small fraction of at most 0.20% of the volume. Moreover, we propose a specific enhancement of the SB-index to deal with spatial data redundancy. This enhancement improved performance from 80 to 91% for redundant GDW schemas.
Resumo:
Due to the imprecise nature of biological experiments, biological data is often characterized by the presence of redundant and noisy data. This may be due to errors that occurred during data collection, such as contaminations in laboratorial samples. It is the case of gene expression data, where the equipments and tools currently used frequently produce noisy biological data. Machine Learning algorithms have been successfully used in gene expression data analysis. Although many Machine Learning algorithms can deal with noise, detecting and removing noisy instances from the training data set can help the induction of the target hypothesis. This paper evaluates the use of distance-based pre-processing techniques for noise detection in gene expression data classification problems. This evaluation analyzes the effectiveness of the techniques investigated in removing noisy data, measured by the accuracy obtained by different Machine Learning classifiers over the pre-processed data.
Resumo:
OBJECTIVE: To estimate the spatial intensity of urban violence events using wavelet-based methods and emergency room data. METHODS: Information on victims attended at the emergency room of a public hospital in the city of São Paulo, Southeastern Brazil, from January 1, 2002 to January 11, 2003 were obtained from hospital records. The spatial distribution of 3,540 events was recorded and a uniform random procedure was used to allocate records with incomplete addresses. Point processes and wavelet analysis technique were used to estimate the spatial intensity, defined as the expected number of events by unit area. RESULTS: Of all georeferenced points, 59% were accidents and 40% were assaults. There is a non-homogeneous spatial distribution of the events with high concentration in two districts and three large avenues in the southern area of the city of São Paulo. CONCLUSIONS: Hospital records combined with methodological tools to estimate intensity of events are useful to study urban violence. The wavelet analysis is useful in the computation of the expected number of events and their respective confidence bands for any sub-region and, consequently, in the specification of risk estimates that could be used in decision-making processes for public policies.
Resumo:
The mature larva and pupa of Fulgeochlizus bruchi (Candèze, 1896) are described and illustrated. Bioluminescent patterns are also given. Comments, new data on the first instar larva and natural history data are presented. The first instar larvae differ from the mature larvae mainly in their chaetotaxy, which is sparse and more symmetrically distributed.
Resumo:
Apesar dos elevados riscos à saúde e capacidade para o trabalho dos eletricitários, há carência de estudos sobre o tema no Brasil. O objetivo desse estudo é identificar o perfil de saúde e capacidade para o trabalho de eletricitários de São Paulo. Foi feito um estudo transversal junto a 475 trabalhadores de uma empresa do setor eletricitário. A coleta de dados foi por meio de questionários sobre capacidade para o trabalho, estado de saúde, estresse no trabalho, atividade física, dependência ao tabaco e ao álcool. A consistência interna das escalas foi avaliada usando o coeficiente alfa de Cronbach. Foi feita análise descritiva por meio das médias, desvios-padrão, valores mínimos e máximos dos escores e proporções para as variáveis qualitativas. O estado de saúde dos trabalhadores apresentou pontuação elevada nas dimensões analisadas, com médias entre 72,8 a 91,2 (escore de 0,0 a 100,0 pontos). A capacidade para o trabalho teve pontuação elevada, com média de 41,8 (escore de 7,0 a 49,0 pontos). Concluiu-se que os trabalhadores da população de estudo apresentaram elevados padrões do estado de saúde e da capacidade para o trabalho. Sugere-se o desenvolvimento de estudos longitudinais para avaliar relações causais e a existência de efeito do trabalhador sadio.