980 resultados para Sparse Data
Resumo:
Well-known data mining algorithms rely on inputs in the form of pairwise similarities between objects. For large datasets it is computationally impossible to perform all pairwise comparisons. We therefore propose a novel approach that uses approximate Principal Component Analysis to efficiently identify groups of similar objects. The effectiveness of the approach is demonstrated in the context of binary classification using the supervised normalized cut as a classifier. For large datasets from the UCI repository, the approach significantly improves run times with minimal loss in accuracy.
Resumo:
This paper introduces an area- and power-efficient approach for compressive recording of cortical signals used in an implantable system prior to transmission. Recent research on compressive sensing has shown promising results for sub-Nyquist sampling of sparse biological signals. Still, any large-scale implementation of this technique faces critical issues caused by the increased hardware intensity. The cost of implementing compressive sensing in a multichannel system in terms of area usage can be significantly higher than a conventional data acquisition system without compression. To tackle this issue, a new multichannel compressive sensing scheme which exploits the spatial sparsity of the signals recorded from the electrodes of the sensor array is proposed. The analysis shows that using this method, the power efficiency is preserved to a great extent while the area overhead is significantly reduced resulting in an improved power-area product. The proposed circuit architecture is implemented in a UMC 0.18 [Formula: see text]m CMOS technology. Extensive performance analysis and design optimization has been done resulting in a low-noise, compact and power-efficient implementation. The results of simulations and subsequent reconstructions show the possibility of recovering fourfold compressed intracranial EEG signals with an SNR as high as 21.8 dB, while consuming 10.5 [Formula: see text]W of power within an effective area of 250 [Formula: see text]m × 250 [Formula: see text]m per channel.
Resumo:
In this paper, we propose a new method for fully-automatic landmark detection and shape segmentation in X-ray images. To detect landmarks, we estimate the displacements from some randomly sampled image patches to the (unknown) landmark positions, and then we integrate these predictions via a voting scheme. Our key contribution is a new algorithm for estimating these displacements. Different from other methods where each image patch independently predicts its displacement, we jointly estimate the displacements from all patches together in a data driven way, by considering not only the training data but also geometric constraints on the test image. The displacements estimation is formulated as a convex optimization problem that can be solved efficiently. Finally, we use the sparse shape composition model as the a priori information to regularize the landmark positions and thus generate the segmented shape contour. We validate our method on X-ray image datasets of three different anatomical structures: complete femur, proximal femur and pelvis. Experiments show that our method is accurate and robust in landmark detection, and, combined with the shape model, gives a better or comparable performance in shape segmentation compared to state-of-the art methods. Finally, a preliminary study using CT data shows the extensibility of our method to 3D data.
Resumo:
We present a novel surrogate model-based global optimization framework allowing a large number of function evaluations. The method, called SpLEGO, is based on a multi-scale expected improvement (EI) framework relying on both sparse and local Gaussian process (GP) models. First, a bi-objective approach relying on a global sparse GP model is used to determine potential next sampling regions. Local GP models are then constructed within each selected region. The method subsequently employs the standard expected improvement criterion to deal with the exploration-exploitation trade-off within selected local models, leading to a decision on where to perform the next function evaluation(s). The potential of our approach is demonstrated using the so-called Sparse Pseudo-input GP as a global model. The algorithm is tested on four benchmark problems, whose number of starting points ranges from 102 to 104. Our results show that SpLEGO is effective and capable of solving problems with large number of starting points, and it even provides significant advantages when compared with state-of-the-art EI algorithms.
Resumo:
Ecological network analysis (ENA) was used to study the effects of Pomatoschistus microps on energy transport through the food web, its impact on other compartments and its possible role as a keystone species in the trophic webs of an Arenicola tidal flat ecosystem and a sparse Zostera noltii bed ecosystem. Three ENA models were constructed: (a) model 1 contains data of the original food web from prior research in the investigated area by Baird et al. (2007), (b) an updated model 2 which included biomass and diet data of P. microps from recent sampling, and (c) model 3 simulating a food web without P. microps. A comparison of energy transport between the different models revealed that more energy is transported from lower trophic levels up the food chain, in the presence of P. microps (models 1 and 2) than in its absence (model 3). Calculations of the keystone index (KSi) revealed the high overall impact (measured as eps_i) of this fish species on food webs. In model 1, P. microps was assigned a low KSi in the Arenicola flat and in the sparse Z. noltii bed. Calculations in model 2 ranked P. microps first for keystoneness and eps_i in both communities, the Arenicola flat and the sparse Z. noltii bed. Taken together, our results give insight into the role of P. microps when considering a whole food web and reveal direct and indirect trophic interactions of this small-sized fish species. These results might illustrate the impact and importance of abundant, widespread species in food webs and facilitate further investigations.
Resumo:
Sites 1085, 1086 and 1087 were drilled off South Africa during Ocean Drilling Program (ODP) Leg 175 to investigate the Benguela Current System. While previous studies have focused on reconstructing the Neogene palaeoceanographic and palaeoclimatic history of these sites, palynology has been largely ignored, except for the Late Pliocene and Quaternary. This study presents palynological data from the upper Middle Miocene to lower Upper Pliocene sediments in Holes 1085A, 1086A and 1087C that provide complementary information about the history of the area. Abundant and diverse marine palynomorphs (mainly dinoflagellate cysts), rare spores and pollen, and dispersed organic matter have been recovered. Multivariate statistical analysis of dispersed organic matter identified three palynofacies assemblages (A, B, C) in the most continuous hole (1085A), and they were defined primarily by amorphous organic matter (AOM), and to a lesser extent black debris, structured phytoclasts, degraded phytoclasts, and marine palynomorphs. Ecostratigraphic interpretation based on dinoflagellate cyst, spore-pollen and palynofacies data allowed us to identify several palaeoceanographic and palaeoclimatic signals. First, the late Middle Miocene was subtropical, and sediments contained the highest percentages of land-derived organic matter, even though they are rich in AOM (palynofacies assemblage A). Second, the Late Miocene was cool-temperate and characterized by periods of intensified upwelling, increase in productivity, abundant and diverse oceanic dinoflagellate cysts, and the highest percentages of AOM (palynofacies assemblage C). Third, the Early to early Late Pliocene was warm-temperate with some dry intervals (increase in grass pollen) and intensified upwelling. Fourth, the Neogene "carbonate crash" identified in other southern oceans was recognized in two palynofacies A samples in Hole 1085A that are nearly barren of dinoflagellate cysts: one Middle Miocene sample (590 mbsf, 13.62 Ma) and one Upper Miocene sample (355 mbsf, 6.5 Ma). Finally, the extremely low percentages of pollen suggest sparse vegetation on the adjacent landmass, and Namib desert conditions were already in existence during the late Middle Miocene.
Resumo:
Sediment accretion and subduction at convergent margins play an important role in the nature of hazardous interplate seismicity (the seismogenic zone) and the subduction recycling of volatiles and continentally derived materials to the Earth's mantle. Identifying and quantifying sediment accretion, essential for a complete mass balance across the margin, can be difficult. Seismic images do not define the processes by which a prism was built, and cored sediments may show disturbed magnetostratigraphy and sparse biostratigraphy. This contribution reports the first use of cosmogenic 10Be depth profiles to define the origin and structural evolution of forearc sedimentary prisms. Biostratigraphy and 10Be model ages generally are in good agreement for sediments drilled at Deep Sea Drilling Project Site 434 in the Japan forearc, and support an origin by imbricate thrusting for the upper section. Forearc sediments from Ocean Drilling Program Site 1040 in Costa Rica lack good fossil or paleomagnetic age control above the decollement. Low and homogeneous 10Be concentrations show that the prism sediments are older than 3-4 Ma, and that the prism is either a paleoaccretionary prism or it formed largely from slump deposits of apron sediments. Low 10Be in Costa Rican lavas and the absence of frontal accretion imply deeper sediment underplating or subduction erosion.
Resumo:
In this paper we propose a novel fast random search clustering (RSC) algorithm for mixing matrix identification in multiple input multiple output (MIMO) linear blind inverse problems with sparse inputs. The proposed approach is based on the clustering of the observations around the directions given by the columns of the mixing matrix that occurs typically for sparse inputs. Exploiting this fact, the RSC algorithm proceeds by parameterizing the mixing matrix using hyperspherical coordinates, randomly selecting candidate basis vectors (i.e. clustering directions) from the observations, and accepting or rejecting them according to a binary hypothesis test based on the Neyman–Pearson criterion. The RSC algorithm is not tailored to any specific distribution for the sources, can deal with an arbitrary number of inputs and outputs (thus solving the difficult under-determined problem), and is applicable to both instantaneous and convolutive mixtures. Extensive simulations for synthetic and real data with different number of inputs and outputs, data size, sparsity factors of the inputs and signal to noise ratios confirm the good performance of the proposed approach under moderate/high signal to noise ratios. RESUMEN. Método de separación ciega de fuentes para señales dispersas basado en la identificación de la matriz de mezcla mediante técnicas de "clustering" aleatorio.
Resumo:
Atrial fibrillation (AF) is a common heart disorder. One of the most prominent hypothesis about its initiation and maintenance considers multiple uncoordinated activation foci inside the atrium. However, the implicit assumption behind all the signal processing techniques used for AF, such as dominant frequency and organization analysis, is the existence of a single regular component in the observed signals. In this paper we take into account the existence of multiple foci, performing a spectral analysis to detect their number and frequencies. In order to obtain a cleaner signal on which the spectral analysis can be performed, we introduce sparsity-aware learning techniques to infer the spike trains corresponding to the activations. The good performance of the proposed algorithm is demonstrated both on synthetic and real data. RESUMEN. Algoritmo basado en técnicas de regresión dispersa para la extracción de las señales cardiacas en pacientes con fibrilación atrial (AF).
Resumo:
Improving the knowledge of demand evolution over time is a key aspect in the evaluation of transport policies and in forecasting future investment needs. It becomes even more critical for the case of toll roads, which in recent decades has become an increasingly common device to fund road projects. However, literature regarding demand elasticity estimates in toll roads is sparse and leaves some important aspects to be analyzed in greater detail. In particular, previous research on traffic analysis does not often disaggregate heavy vehicle demand from the total volume, so that the specific behavioral patternsof this traffic segment are not taken into account. Furthermore, GDP is the main socioeconomic variable most commonly chosen to explain road freight traffic growth over time. This paper seeks to determine the variables that better explain the evolution of heavy vehicle demand in toll roads over time. To that end, we present a dynamic panel data methodology aimed at identifying the key socioeconomic variables that explain the behavior of road freight traffic throughout the years. The results show that, despite the usual practice, GDP may not constitute a suitable explanatory variable for heavy vehicle demand. Rather, considering only the GDP of those sectors with a high impact on transport demand, such as construction or industry, leads to more consistent results. The methodology is applied to Spanish toll roads for the 1990?2011 period. This is an interesting case in the international context, as road freight demand has experienced an even greater reduction in Spain than elsewhere, since the beginning of the economic crisis in 2008.
Resumo:
Census data on endangered species are often sparse, error-ridden, and confined to only a segment of the population. Estimating trends and extinction risks using this type of data presents numerous difficulties. In particular, the estimate of the variation in year-to-year transitions in population size (the “process error” caused by stochasticity in survivorship and fecundities) is confounded by the addition of high sampling error variation. In addition, the year-to-year variability in the segment of the population that is sampled may be quite different from the population variability that one is trying to estimate. The combined effect of severe sampling error and age- or stage-specific counts leads to severe biases in estimates of population-level parameters. I present an estimation method that circumvents the problem of age- or stage-specific counts and is markedly robust to severe sampling error. This method allows the estimation of environmental variation and population trends for extinction-risk analyses using corrupted census counts—a common type of data for endangered species that has hitherto been relatively unusable for these analyses.
Resumo:
The Middle Valley segment at the northern end of the Juan de Fuca Ridge is a deep extensional rift blanketed with 200-500 m of Pleistocene turbiditic sediment. Sites 857 and 858 were drilled during Ocean Drilling Program Leg 139 to determine whether these two sites were hydrologically linked end members of an active hydrothermal circulation system. Site 858 was placed in an area of active hydrothermal discharge with fluids up to 270°C venting through anhydrite-bearing mounds on top of altered sediment. The shallow basement of fine-grained basalt that underlies the vents at Site 858 is interpreted as a seamount that was subsequently buried by turbidites. Site 857 was placed 1.6 km south of the Site 858 vents in a zone of high heat flow and numerous seismically imaged ridge-parallel faults. Drilling at Site 857 encountered sediments that are increasingly altered with depth and that overlie a series of mafic sills at depths of 460-940 m below sea floor. Sill margins and adjacent baked sediment are highly altered to magnesian chlorite and crosscut with veins filled with quartz, chlorite, sulfides, epidote, and wairakite. The sill interiors vary from slightly altered, with unaltered plagioclase and clinopyroxene in a mesostasis replaced by chlorite, to local zones of intense alteration and brecciation. In these latter zones, the sill interiors are pervasively replaced by chlorite, epidote, quartz, pyrite, titanite, and rare actinolite. The most complete replacement is associated with brecciated horizons with low recovery and slickensides on fracture surfaces, which we interpret as intersections between faults and the sills. Geochemically, the alteration of the sill complex is reflected in significant whole-rock depletions in Ca, Sr, and Na with corresponding enrichments in Mg, Al, and most metals. The latter results from the formation of conspicuous sulfide poikiloblasts. In contrast, metamorphism of the Site 858 seamount includes incomplete albitization of plagioclase phenocrysts and replacement of sparse mafic phenocrysts. Much of the basement alteration at Site 858 is confined to crosscutting veins except for a highly altered and veined horizon at the contact between basaltic basement and the overlying sediment. The sill complex at Site 857 is more highly depleted in 18O (d18O = 2.4 per mil - 4.7 per mil) and more pervasively replaced by secondary minerals relative to the extrusives at Site 858 (d18O = 4.5 per mil - 5.5 per mil). There is no evidence of significant albitization of the plagioclase at Site 857, suggesting high Ca/Na in the pore fluids. Fluid-inclusion data from hydrothermal minerals in altered mafic rocks and veins at Sites 857 and 858 show a consistency of homogenization temperatures, varying from 245 to 270°C, which is within the range of temperatures observed for the fluids venting at Site 858. The consistency of the fluid inclusion temperatures, the lack of albitization within the Site 857 sills, and the apparently low water/rock ratio collectively suggest that the sill complex at Site 857 is in thermal equilibrium and being altered by a highly evolved Ca-rich fluid similar to the fluids now venting at Site 858. The alteration evident in these two deep crustal drillsites is a result of the ongoing hydrothermal circulation and is consistent with downhole logging results, instrumented borehole results, and hydrothermal fluid chemistry. The pervasive alteration of the laterally extensive sill-sediment complex at Site 857 determines the chemistry of the fluids that are venting at Site 858. The limited alteration of the Site 858 lavas suggests that this basement edifice acts as a penetrator or ventilator for the regional hydrothermal reservoir with much of the flow focussed at the highly altered and veined sediment-basalt contact.
Resumo:
The flow of ice streams, which account for most discharge from large ice sheets, is controlled by processes operating at their bed. Data from modern ice stream beds are difficult to obtain, but where ice advanced onto continental shelves during glacial periods extensive areas of the former bed can be imaged using modern swath sonar tools. We present new multibeam swath bathymetry data analyzed alongside sparse pre-existing data from the Amundsen Sea Embayment. The compilation is the most extensive, continuous area of multibeam data coverage yet obtained on the inner continental shelf of Antarctica. The data reveal streamlined subglacial bedforms that define a zone of paleo-ice stream convergence but, in contrast to previous models, do not show a simple down-flow progression of bedform types along paleo-ice stream troughs. We interpret high spatial variability of bedforms as indicating a complex mechanical and hydrodynamic regime at the former ice stream beds, consistent with observations from some modern ice streams. We conclude that care must be taken when using bedforms to infer paleo-ice stream velocities.
Resumo:
Data on the occurrence of species are widely used to inform the design of reserve networks. These data contain commission errors (when a species is mistakenly thought to be present) and omission errors (when a species is mistakenly thought to be absent), and the rates of the two types of error are inversely related. Point locality data can minimize commission errors, but those obtained from museum collections are generally sparse, suffer from substantial spatial bias and contain large omission errors. Geographic ranges generate large commission errors because they assume homogenous species distributions. Predicted distribution data make explicit inferences on species occurrence and their commission and omission errors depend on model structure, on the omission of variables that determine species distribution and on data resolution. Omission errors lead to identifying networks of areas for conservation action that are smaller than required and centred on known species occurrences, thus affecting the comprehensiveness, representativeness and efficiency of selected areas. Commission errors lead to selecting areas not relevant to conservation, thus affecting the representativeness and adequacy of reserve networks. Conservation plans should include an estimation of commission and omission errors in underlying species data and explicitly use this information to influence conservation planning outcomes.
Resumo:
We develop an approach for a sparse representation for Gaussian Process (GP) models in order to overcome the limitations of GPs caused by large data sets. The method is based on a combination of a Bayesian online algorithm together with a sequential construction of a relevant subsample of the data which fully specifies the prediction of the model. Experimental results on toy examples and large real-world datasets indicate the efficiency of the approach.