822 resultados para discriminant analysis and cluster analysis


Relevância:

100.00% 100.00%

Publicador:

Resumo:

Computer vision is increasingly becoming interested in the rapid estimation of object detectors. The canonical strategy of using Hard Negative Mining to train a Support Vector Machine is slow, since the large negative set must be traversed at least once per detector. Recent work has demonstrated that, with an assumption of signal stationarity, Linear Discriminant Analysis is able to learn comparable detectors without ever revisiting the negative set. Even with this insight, the time to learn a detector can still be on the order of minutes. Correlation filters, on the other hand, can produce a detector in under a second. However, this involves the unnatural assumption that the statistics are periodic, and requires the negative set to be re-sampled per detector size. These two methods differ chie y in the structure which they impose on the co- variance matrix of all examples. This paper is a comparative study which develops techniques (i) to assume periodic statistics without needing to revisit the negative set and (ii) to accelerate the estimation of detectors with aperiodic statistics. It is experimentally verified that periodicity is detrimental.

Relevância:

100.00% 100.00%

Publicador:

Resumo:

Traditional nearest points methods use all the samples in an image set to construct a single convex or affine hull model for classification. However, strong artificial features and noisy data may be generated from combinations of training samples when significant intra-class variations and/or noise occur in the image set. Existing multi-model approaches extract local models by clustering each image set individually only once, with fixed clusters used for matching with various image sets. This may not be optimal for discrimination, as undesirable environmental conditions (eg. illumination and pose variations) may result in the two closest clusters representing different characteristics of an object (eg. frontal face being compared to non-frontal face). To address the above problem, we propose a novel approach to enhance nearest points based methods by integrating affine/convex hull classification with an adapted multi-model approach. We first extract multiple local convex hulls from a query image set via maximum margin clustering to diminish the artificial variations and constrain the noise in local convex hulls. We then propose adaptive reference clustering (ARC) to constrain the clustering of each gallery image set by forcing the clusters to have resemblance to the clusters in the query image set. By applying ARC, noisy clusters in the query set can be discarded. Experiments on Honda, MoBo and ETH-80 datasets show that the proposed method outperforms single model approaches and other recent techniques, such as Sparse Approximated Nearest Points, Mutual Subspace Method and Manifold Discriminant Analysis.

Relevância:

100.00% 100.00%

Publicador:

Resumo:

Existing multi-model approaches for image set classification extract local models by clustering each image set individually only once, with fixed clusters used for matching with other image sets. However, this may result in the two closest clusters to represent different characteristics of an object, due to different undesirable environmental conditions (such as variations in illumination and pose). To address this problem, we propose to constrain the clustering of each query image set by forcing the clusters to have resemblance to the clusters in the gallery image sets. We first define a Frobenius norm distance between subspaces over Grassmann manifolds based on reconstruction error. We then extract local linear subspaces from a gallery image set via sparse representation. For each local linear subspace, we adaptively construct the corresponding closest subspace from the samples of a probe image set by joint sparse representation. We show that by minimising the sparse representation reconstruction error, we approach the nearest point on a Grassmann manifold. Experiments on Honda, ETH-80 and Cambridge-Gesture datasets show that the proposed method consistently outperforms several other recent techniques, such as Affine Hull based Image Set Distance (AHISD), Sparse Approximated Nearest Points (SANP) and Manifold Discriminant Analysis (MDA).

Relevância:

100.00% 100.00%

Publicador:

Resumo:

The concentrations of Na, K, Ca, Mg, Ba, Sr, Fe, Al, Mn, Zn, Pb, Cu, Ni, Cr, Co, Se, U and Ti were determined in the osteoderms and/or flesh of estuarine crocodiles (Crocodylus porosus) captured in three adjacent catchments within the Alligator Rivers Region (ARR) of northern Australia. Results from multivariate analysis of variance showed that when all metals were considered simultaneously, catchment effects were significant (P≤0.05). Despite considerable within-catchment variability, linear discriminant analysis (LDA) showed that differences in elemental signatures in the osteoderms and/or flesh of C. porosus amongst the catchments were sufficient to classify individuals accurately to their catchment of occurrence. Using cross-validation, the accuracy of classifying a crocodile to its catchment of occurrence was 76% for osteoderms and 60% for flesh. These data suggest that osteoderms provide better predictive accuracy than flesh for discriminating crocodiles amongst catchments. There was no advantage in combining the osteoderm and flesh results to increase the accuracy of classification (i.e. 67%). Based on the discriminant function coefficients for the osteoderm data, Ca, Co, Mg and U were the most important elements for discriminating amongst the three catchments. For flesh data, Ca, K, Mg, Na, Ni and Pb were the most important metals for discriminating amongst the catchments. Reasons for differences in the elemental signatures of crocodiles between catchments are generally not interpretable, due to limited data on surface water and sediment chemistry of the catchments or chemical composition of dietary items of C. porosus. From a wildlife management perspective, the provenance or source catchment(s) of 'problem' crocodiles captured at settlements or recreational areas along the ARR coastline may be established using catchment-specific elemental signatures. If the incidence of problem crocodiles can be reduced in settled or recreational areas by effective management at their source, then public safety concerns about these predators may be moderated, as well as the cost of their capture and removal. Copyright © 2002 Elsevier Science B.V.

Relevância:

100.00% 100.00%

Publicador:

Resumo:

Time series classification has been extensively explored in many fields of study. Most methods are based on the historical or current information extracted from data. However, if interest is in a specific future time period, methods that directly relate to forecasts of time series are much more appropriate. An approach to time series classification is proposed based on a polarization measure of forecast densities of time series. By fitting autoregressive models, forecast replicates of each time series are obtained via the bias-corrected bootstrap, and a stationarity correction is considered when necessary. Kernel estimators are then employed to approximate forecast densities, and discrepancies of forecast densities of pairs of time series are estimated by a polarization measure, which evaluates the extent to which two densities overlap. Following the distributional properties of the polarization measure, a discriminant rule and a clustering method are proposed to conduct the supervised and unsupervised classification, respectively. The proposed methodology is applied to both simulated and real data sets, and the results show desirable properties.

Relevância:

100.00% 100.00%

Publicador:

Resumo:

Experimental studies have found that when the state-of-the-art probabilistic linear discriminant analysis (PLDA) speaker verification systems are trained using out-domain data, it significantly affects speaker verification performance due to the mismatch between development data and evaluation data. To overcome this problem we propose a novel unsupervised inter dataset variability (IDV) compensation approach to compensate the dataset mismatch. IDV-compensated PLDA system achieves over 10% relative improvement in EER values over out-domain PLDA system by effectively compensating the mismatch between in-domain and out-domain data.

Relevância:

100.00% 100.00%

Publicador:

Resumo:

Diabetic macular edema (DME) is one of the most common causes of visual loss among diabetes mellitus patients. Early detection and successive treatment may improve the visual acuity. DME is mainly graded into non-clinically significant macular edema (NCSME) and clinically significant macular edema according to the location of hard exudates in the macula region. DME can be identified by manual examination of fundus images. It is laborious and resource intensive. Hence, in this work, automated grading of DME is proposed using higher-order spectra (HOS) of Radon transform projections of the fundus images. We have used third-order cumulants and bispectrum magnitude, in this work, as features, and compared their performance. They can capture subtle changes in the fundus image. Spectral regression discriminant analysis (SRDA) reduces feature dimension, and minimum redundancy maximum relevance method is used to rank the significant SRDA components. Ranked features are fed to various supervised classifiers, viz. Naive Bayes, AdaBoost and support vector machine, to discriminate No DME, NCSME and clinically significant macular edema classes. The performance of our system is evaluated using the publicly available MESSIDOR dataset (300 images) and also verified with a local dataset (300 images). Our results show that HOS cumulants and bispectrum magnitude obtained an average accuracy of 95.56 and 94.39 % for MESSIDOR dataset and 95.93 and 93.33 % for local dataset, respectively.

Relevância:

100.00% 100.00%

Publicador:

Resumo:

This paper analyzes the limitations upon the amount of in- domain (NIST SREs) data required for training a probabilistic linear discriminant analysis (PLDA) speaker verification system based on out-domain (Switchboard) total variability subspaces. By limiting the number of speakers, the number of sessions per speaker and the length of active speech per session available in the target domain for PLDA training, we investigated the relative effect of these three parameters on PLDA speaker verification performance in the NIST 2008 and NIST 2010 speaker recognition evaluation datasets. Experimental results indicate that while these parameters depend highly on each other, to beat out-domain PLDA training, more than 10 seconds of active speech should be available for at least 4 sessions/speaker for a minimum of 800 speakers. If further data is available, considerable improvement can be made over solely out-domain PLDA training.

Relevância:

100.00% 100.00%

Publicador:

Resumo:

A low-altitude platform utilising a 1.8-m diameter tethered helium balloon was used to position a multispectral sensor, consisting of two digital cameras, above a fertiliser trial plot where wheat (Triticum spp.) was being grown. Located in Cecil Plains, Queensland, Australia, the plot was a long-term fertiliser trial being conducted by a fertiliser company to monitor the response of crops to various levels of nutrition. The different levels of nutrition were achieved by varying nitrogen application rates between 0 and 120 units of N at 40 unit increments. Each plot had received the same application rate for 10 years. Colour and near-infrared images were acquired that captured the whole 2 ha plot. These images were examined and relationships sought between the captured digital information and the crop parameters imaged at anthesis and the at-harvest quality and quantity parameters. The statistical analysis techniques used were correlation analysis, discriminant analysis and partial least squares regression. A high correlation was found between the image and yield (R2 = 0.91) and a moderate correlation between the image and grain protein content (R2 = 0.66). The utility of the system could be extended by choosing a more mobile platform. This would increase the potential for the system to be used to diagnose the causes of the variability and allow remediation, and/or to segregate the crop at harvest to meet certain quality parameters.

Relevância:

100.00% 100.00%

Publicador:

Resumo:

Identification of major contributors to odour annoyance in areas with multiple emission sources is necessary to address and resolve odour disputes. In an effort to develop an appropriate tool for this task, odour samples were collected on-site at a piggery and an abattoir (the major odour sources in the area) and at surrounding off-site areas, then analysed using a commercial non-specific chemical sensor array to develop an odour fingerprint database. The developed odour fingerprint database was analysed using two pattern recognition algorithms including a partial least squares-discriminant analysis (PLS-DA) and a Kohonen self-organising map (KSOM). The KSOM model could identify odour samples sourced from the piggery shed 15, piggery pond 8, piggery pond 9, abattoir, motel and others with mean percentage values of 77.5, 65.0, 90.2, 75.7, 44.8 and 64.6%, respectively.

Relevância:

100.00% 100.00%

Publicador:

Resumo:

The study examined the potential of Near Infrared Reflectance (NIR) spectroscopy for field diagnosis of hybrids between Corymbia (formerly Eucalyptus) species. NIR profiles were generated by scanning foliage from a total of 383 hybrid and 533 parental seedlings grown in a common garden and partial least squares discriminant analysis was used to test three-way model power to assign individuals to their appropriate taxon; either a parental or F1 hybrid class. Using the optimised conditions, fresh foliage from eight-month-old seedlings and a handheld NIR instrument (950–1800 nm), the mean assignment rates for the three hybrid groups ranged from 76% to 90%. Hybrid-parent contrast of NIR spectra deviated more so than parent–parent contrast. The F1 taxon assignment rates were usually higher than those for parents at 100% and 72%, respectively. Hybrid resolution was even greater for 2nd generation backcross hybrids. Similar to studies of morphology, taxon assignments tended to be more accurate for hybrid groups in which the parental taxa were more divergent. The practical application of this technique for hybrid diagnosis of seedlings in the nursery will require careful attention to control environmental factors because seedling age and storage effects influenced the ability of NIR to identify hybrids. The technique may also necessitate the generation of comparable reference populations, although exclusions approaches to analysis may circumvent the need for reference populations. The application of NIR in field diagnosis will be further complicated by the need to generate global models across environments but such models have been obtained for reliable prediction of chemistries in other situations.

Relevância:

100.00% 100.00%

Publicador:

Resumo:

The DNA polymorphism among 22 isolates of Sclerospora graminicola, the causal agent of downy mildew disease of pearl millet was assessed using 20 inter simple sequence repeats (ISSR) primers. The objective of the study was to examine the effectiveness of using ISSR markers for unravelling the extent and pattern of genetic diversity in 22 S. graminicola isolates collected from different host cultivars in different states of India. The 19 functional ISSR primers generated 410 polymorphic bands and revealed 89% polymorphism and were able to distinguish all the 22 isolates. Polymorphic bands used to construct an unweighted pair group method of averages (UPGMA) dendrogram based on Jaccard's co-efficient of similarity and principal coordinate analysis resulted in the formation of four major clusters of 22 isolates. The standardized Nei genetic distance among the 22 isolates ranged from 0.0050 to 0.0206. The UPGMA clustering using the standardized genetic distance matrix resulted in the identification of four clusters of the 22 isolates with bootstrap values ranging from 15 to 100. The 3D-scale data supported the UPGMA results, which resulted into four clusters amounting to 70% variation among each other. However, comparing the two methods show that sub clustering by dendrogram and multi dimensional scaling plot is slightly different. All the S. graminicola isolates had distinct ISSR genotypes and cluster analysis origin. The results of ISSR fingerprints revealed significant level of genetic diversity among the isolates and that ISSR markers could be a powerful tool for fingerprinting and diversity analysis in fungal pathogens.

Relevância:

100.00% 100.00%

Publicador:

Resumo:

While plants of a single species emit a diversity of volatile organic compounds (VOCs) to attract or repel interacting organisms, these specific messages may be lost in the midst of the hundreds of VOCs produced by sympatric plants of different species, many of which may have no signal content. Receivers must be able to reduce the babel or noise in these VOCs in order to correctly identify the message. For chemical ecologists faced with vast amounts of data on volatile signatures of plants in different ecological contexts, it is imperative to employ accurate methods of classifying messages, so that suitable bioassays may then be designed to understand message content. We demonstrate the utility of `Random Forests' (RF), a machine-learning algorithm, for the task of classifying volatile signatures and choosing the minimum set of volatiles for accurate discrimination, using datam from sympatric Ficus species as a case study. We demonstrate the advantages of RF over conventional classification methods such as principal component analysis (PCA), as well as data-mining algorithms such as support vector machines (SVM), diagonal linear discriminant analysis (DLDA) and k-nearest neighbour (KNN) analysis. We show why a tree-building method such as RF, which is increasingly being used by the bioinformatics, food technology and medical community, is particularly advantageous for the study of plant communication using volatiles, dealing, as it must, with abundant noise.

Relevância:

100.00% 100.00%

Publicador:

Resumo:

In this paper, we give a brief review of pattern classification algorithms based on discriminant analysis. We then apply these algorithms to classify movement direction based on multivariate local field potentials recorded from a microelectrode array in the primary motor cortex of a monkey performing a reaching task. We obtain prediction accuracies between 55% and 90% using different methods which are significantly above the chance level of 12.5%.

Relevância:

100.00% 100.00%

Publicador:

Resumo:

Myopathies are muscular diseases in which muscle fibers degenerate due to many factors such as nutrient deficiency, infection and mutations in myofibrillar etc. The objective of this study is to identify the bio-markers to distinguish various muscle mutants in Drosophila (fruit fly) using Raman Spectroscopy. Principal Components based Linear Discriminant Analysis (PC-LDA) classification model yielding >95% accuracy was developed to classify such different mutants representing various myopathies according to their physiopathology.