391 resultados para Multigrain Parallelism
Resumo:
Non-finite clauses are sentential constituents with a verbal head that lacks a morphological specification for tense and agreement. In this paper I contend that these clauses are defective not only morphologically but also syntactically, in the sense that they all lack some of the functional categories that make up a full sentence. In particular I argue that to-infinitive clauses, gerund(ive) clauses and participial clauses differ among themselves, and with respect to other subordinate clauses, in the degree of structural defectiveness they display, which goes from the almost complete functional structure of the infinitive to the maximal degree of syntactic truncation of participial clauses (analyzed here as verbal small clauses). I also show the significant parallelism that exists in this respect between English and Spanish non-finite clauses, pointing to the implication this may have for a cross-linguistic approach to the cartography of syntactic structures.
Resumo:
Solving linear systems is an important problem for scientific computing. Exploiting parallelism is essential for solving complex systems, and this traditionally involves writing parallel algorithms on top of a library such as MPI. The SPIKE family of algorithms is one well-known example of a parallel solver for linear systems. The Hierarchically Tiled Array data type extends traditional data-parallel array operations with explicit tiling and allows programmers to directly manipulate tiles. The tiles of the HTA data type map naturally to the block nature of many numeric computations, including the SPIKE family of algorithms. The higher level of abstraction of the HTA enables the same program to be portable across different platforms. Current implementations target both shared-memory and distributed-memory models. In this thesis we present a proof-of-concept for portable linear solvers. We implement two algorithms from the SPIKE family using the HTA library. We show that our implementations of SPIKE exploit the abstractions provided by the HTA to produce a compact, clean code that can run on both shared-memory and distributed-memory models without modification. We discuss how we map the algorithms to HTA programs as well as examine their performance. We compare the performance of our HTA codes to comparable codes written in MPI as well as current state-of-the-art linear algebra routines.
Resumo:
Due to the growth of design size and complexity, design verification is an important aspect of the Logic Circuit development process. The purpose of verification is to validate that the design meets the system requirements and specification. This is done by either functional or formal verification. The most popular approach to functional verification is the use of simulation based techniques. Using models to replicate the behaviour of an actual system is called simulation. In this thesis, a software/data structure architecture without explicit locks is proposed to accelerate logic gate circuit simulation. We call thus system ZSIM. The ZSIM software architecture simulator targets low cost SIMD multi-core machines. Its performance is evaluated on the Intel Xeon Phi and 2 other machines (Intel Xeon and AMD Opteron). The aim of these experiments is to: • Verify that the data structure used allows SIMD acceleration, particularly on machines with gather instructions ( section 5.3.1). • Verify that, on sufficiently large circuits, substantial gains could be made from multicore parallelism ( section 5.3.2 ). • Show that a simulator using this approach out-performs an existing commercial simulator on a standard workstation ( section 5.3.3 ). • Show that the performance on a cheap Xeon Phi card is competitive with results reported elsewhere on much more expensive super-computers ( section 5.3.5 ). To evaluate the ZSIM, two types of test circuits were used: 1. Circuits from the IWLS benchmark suit [1] which allow direct comparison with other published studies of parallel simulators.2. Circuits generated by a parametrised circuit synthesizer. The synthesizer used an algorithm that has been shown to generate circuits that are statistically representative of real logic circuits. The synthesizer allowed testing of a range of very large circuits, larger than the ones for which it was possible to obtain open source files. The experimental results show that with SIMD acceleration and multicore, ZSIM gained a peak parallelisation factor of 300 on Intel Xeon Phi and 11 on Intel Xeon. With only SIMD enabled, ZSIM achieved a maximum parallelistion gain of 10 on Intel Xeon Phi and 4 on Intel Xeon. Furthermore, it was shown that this software architecture simulator running on a SIMD machine is much faster than, and can handle much bigger circuits than a widely used commercial simulator (Xilinx) running on a workstation. The performance achieved by ZSIM was also compared with similar pre-existing work on logic simulation targeting GPUs and supercomputers. It was shown that ZSIM simulator running on a Xeon Phi machine gives comparable simulation performance to the IBM Blue Gene supercomputer at very much lower cost. The experimental results have shown that the Xeon Phi is competitive with simulation on GPUs and allows the handling of much larger circuits than have been reported for GPU simulation. When targeting Xeon Phi architecture, the automatic cache management of the Xeon Phi, handles and manages the on-chip local store without any explicit mention of the local store being made in the architecture of the simulator itself. However, targeting GPUs, explicit cache management in program increases the complexity of the software architecture. Furthermore, one of the strongest points of the ZSIM simulator is its portability. Note that the same code was tested on both AMD and Xeon Phi machines. The same architecture that efficiently performs on Xeon Phi, was ported into a 64 core NUMA AMD Opteron. To conclude, the two main achievements are restated as following: The primary achievement of this work was proving that the ZSIM architecture was faster than previously published logic simulators on low cost platforms. The secondary achievement was the development of a synthetic testing suite that went beyond the scale range that was previously publicly available, based on prior work that showed the synthesis technique is valid.
Resumo:
The morphometric relations allow describing dimensions of trees without prior knowledge of the age, it help the forest planning and implementation of silvicultural treatments, especially when needs to make sustainable use of forests. For this purpose, the aim of this study was to model and comparising the morphometric relations araucaria trees in social position dominant, codominant and dominated in native forest remnant, located in Lages, SC. A total of 294 trees distributed in dbh classes were intentionally selected inside of forest. In each tree was measured dbh, total height, bole height, crown diameter by eight radius, as well as the classification of social position. Simple and multiple linear regression models were used to describe the relation h/d, the proportion of the crown and formal crown in function of diameter at breast height with simple transformation, quadratic, cubic, inverse and logarithmic form. The analysis of covariance with dummy variables were used to describe the social position and tested the parallelism and slope of regression indicating need or not of the use independent regressions. The results indicated that even with great variability in the shape and size of the crown due to growth and competition process, the morphometric relations of araucaria can be accurately estimated by regression models. The relation h/d, proportion of the crown and formal crown can be described by individual model for social position dominant, codominant and dominant, or alternatively a single model with the use of dummy variables that differentiate trees group dominated for the relation h/d and formal crown. The proportion of crown presented difference in dimensions of the trees, being necessary to use dummy variable for each social stratus or use the individual models.
Resumo:
No presente relatório da Prática de Ensino Supervisionada são referidas opções de ensino, procedimentos e reações dos alunos ao processo de ensino. É dada uma grande ênfase ao ambiente de aprendizagem baseado na tecnologia e suportado por uma comunidade de aprendizagem, que tem lugar na própria sala de aula ou na sala de informática. A tecnologia é assumida como um recurso constante na maior parte das aulas através do recurso a tarefas escolhidas intencionalmente tendo em vista a possibilidade de introdução da tecnologia na sua resolução. Esta implementação assumiu várias formas, tais como a exploração de calculadoras, a manipulação do GeoGebra ou simplesmente através da apresentação de ficheiros acabados, o que constitui uma forma de obter uma boa visualização dos objetos matemáticos. A aplicação dos recursos tecnológicos foi progressivamente tornada mais intensiva, atingindo o seu culminar no Projeto de Estágio, designação atribuída a duas aulas concebidas explicitamente para a exploração da temática: “Estabelecimento de um Paralelismo entre a Geometria Tridimensional Dinâmica e as Funções”; Abstract: The Use of Technology in the Classroom as an Instrument of Visualization and Algebrization of the Mathematical Objects In this paper we refer to teaching options, procedures, and to students’ reactions to the teaching processes. We give a lot of reinforcement in the learning environment based on technology and supported by a community of learners, which take place in their own classroom or in the Informatics Class. Technology is assumed as a constant resource in most part of the classes through the intentional tasks’ choosing taking into account the possibility of technology introduction in their resolution. This implementation has assumed several forms, like calculators’ exploration, GeoGebra manipulation or simply by presenting finished files, which is a way of getting a great visualization of mathematical objects. The technological resources’ application turned itself progressively more intensive, presenting its center point on Practice Project, name who was gave to two classes conceived explicitly for the thematic exploration: “The establishment of a parallelism between Dynamic Tridimensional Geometry and the Functions”.
Resumo:
A poster of this paper will be presented at the 25th International Conference on Parallel Architecture and Compilation Technology (PACT ’16), September 11-15, 2016, Haifa, Israel.
Resumo:
O estilo de Tucídides foi estudado desde a Antiguidade, sendo a obra mais completa, apesar de não elogiosa, a de Dionísio de Halicarnasso. Este artigo foca alguns dos elementos mais característicos da escrita de Tucídides, estruturando-os em pequenas secções que abrangem a «variatio», os paralelismos, os «hapax legomena», as abstrações, as definições de conceitos, as generalizações e o uso de documentos «verbatim».
Resumo:
In the multi-core CPU world, transactional memory (TM)has emerged as an alternative to lock-based programming for thread synchronization. Recent research proposes the use of TM in GPU architectures, where a high number of computing threads, organized in SIMT fashion, requires an effective synchronization method. In contrast to CPUs, GPUs offer two memory spaces: global memory and local memory. The local memory space serves as a shared scratch-pad for a subset of the computing threads, and it is used by programmers to speed-up their applications thanks to its low latency. Prior work from the authors proposed a lightweight hardware TM (HTM) support based in the local memory, modifying the SIMT execution model and adding a conflict detection mechanism. An efficient implementation of these features is key in order to provide an effective synchronization mechanism at the local memory level. After a quick description of the main features of our HTM design for GPU local memory, in this work we gather together a number of proposals designed with the aim of improving those mechanisms with high impact on performance. Firstly, the SIMT execution model is modified to increase the parallelism of the application when transactions must be serialized in order to make forward progress. Secondly, the conflict detection mechanism is optimized depending on application characteristics, such us the read/write sets, the probability of conflict between transactions and the existence of read-only transactions. As these features can be present in hardware simultaneously, it is a task of the compiler and runtime to determine which ones are more important for a given application. This work includes a discussion on the analysis to be done in order to choose the best configuration solution.
Resumo:
Bilinear pairings can be used to construct cryptographic systems with very desirable properties. A pairing performs a mapping on members of groups on elliptic and genus 2 hyperelliptic curves to an extension of the finite field on which the curves are defined. The finite fields must, however, be large to ensure adequate security. The complicated group structure of the curves and the expensive field operations result in time consuming computations that are an impediment to the practicality of pairing-based systems. The Tate pairing can be computed efficiently using the ɳT method. Hardware architectures can be used to accelerate the required operations by exploiting the parallelism inherent to the algorithmic and finite field calculations. The Tate pairing can be performed on elliptic curves of characteristic 2 and 3 and on genus 2 hyperelliptic curves of characteristic 2. Curve selection is dependent on several factors including desired computational speed, the area constraints of the target device and the required security level. In this thesis, custom hardware processors for the acceleration of the Tate pairing are presented and implemented on an FPGA. The underlying hardware architectures are designed with care to exploit available parallelism while ensuring resource efficiency. The characteristic 2 elliptic curve processor contains novel units that return a pairing result in a very low number of clock cycles. Despite the more complicated computational algorithm, the speed of the genus 2 processor is comparable. Pairing computation on each of these curves can be appealing in applications with various attributes. A flexible processor that can perform pairing computation on elliptic curves of characteristic 2 and 3 has also been designed. An integrated hardware/software design and verification environment has been developed. This system automates the procedures required for robust processor creation and enables the rapid provision of solutions for a wide range of cryptographic applications.
Resumo:
Based on previous research which shows parallelism between the saliva and blood lactate response during incremental exercise, we hypothesized that a "maximum salivary lactate steady state" (saliva-MLSS) might exist. Thus, the aim of the present investigation was to establish 1) which lower limit for the increase in salivary lactate concentration during a constant workload (i.e., from the 10th to the 20th min) test could be used to determine the saliva-MLSS and 2) if the exercise intensity corresponding to the saliva-MLSS is identical to that evoking the (blood) MLSS. Twelve male amateur athletes of mean (+/-SD) age 24+/-5 year were selected for the study. Based on the results of a previous maximal cycle ergometer test for lactate threshold (LT) determination, each subject performed consecutive constant workload tests of 20-min duration on separate days for MLSS determination, Blood and saliva (25 mu l) samples were collected at 0, 10, and 20 min during the tests for lactate determination. A Student's t-test for paired data demonstrated that a salivary lactate increase of 0.8 mM corresponded to the saliva-MLSS. At this value, indeed, no significant differences were observed between the mean (V) over dot O-2, and W values corresponding to the MLSS and the saliva-MLSS. In conclusion, the present findings indicate that 0.8 mM is the lower limit for the increase in saliva lactate concentration during a constant load test and thus is that which might be used as a reference to determine saliva-MLSS. Furthermore, saliva-MLSS might be used as an alternative to MLSS determination in blood samples.
Resumo:
Biomarkers are nowadays essential tools to be one step ahead for fighting disease, enabling an enhanced focus on disease prevention and on the probability of its occurrence. Research in a multidisciplinary approach has been an important step towards the repeated discovery of new biomarkers. Biomarkers are defined as biochemical measurable indicators of the presence of disease or as indicators for monitoring disease progression. Currently, biomarkers have been used in several domains such as oncology, neurology, cardiovascular, inflammatory and respiratory disease, and several endocrinopathies. Bridging biomarkers in a One Health perspective has been proven useful in almost all of these domains. In oncology, humans and animals are found to be subject to the same environmental and genetic predisposing factors: examples include the existence of mutations in BR-CA1 gene predisposing to breast cancer, both in human and dogs, with increased prevalence in certain dog breeds and human ethnic groups. Also, breast feeding frequency and duration has been related to a decreased risk of breast cancer in women and bitches. When it comes to infectious diseases, this parallelism is prone to be even more important, for as much as 75% of all emerging diseases are believed to be zoonotic. Examples of successful use of biomarkers have been found in several zoonotic diseases such as Ebola, dengue, leptospirosis or West Nile virus infections. Acute Phase Proteins (APPs) have been used for quite some time as biomarkers of inflammatory conditions. These have been used in human health but also in the veterinary field such as in mastitis evaluation and PRRS (porcine respiratory and reproductive syndrome) diagnosis. Advantages rely on the fact that these biomarkers can be much easier to assess than other conventional disease diagnostic approaches (example: measured in easy to collect saliva samples). Another domain in which biomarkers have been essential is food safety: the possibility to measure exposure to chemical contaminants or other biohazards present in the food chain, which are sometimes analytical challenges due to their low bioavailability in body fluids, is nowadays a major breakthrough. Finally, biomarkers are considered the key to provide more personalized therapies, with more efficient outcomes and fewer side effects. This approach is expected to be the correct path to follow also in veterinary medicine, in the near future.
Resumo:
Solving a complex Constraint Satisfaction Problem (CSP) is a computationally hard task which may require a considerable amount of time. Parallelism has been applied successfully to the job and there are already many applications capable of harnessing the parallel power of modern CPUs to speed up the solving process. Current Graphics Processing Units (GPUs), containing from a few hundred to a few thousand cores, possess a level of parallelism that surpasses that of CPUs and there are much less applications capable of solving CSPs on GPUs, leaving space for further improvement. This paper describes work in progress in the solving of CSPs on GPUs, CPUs and other devices, such as Intel Many Integrated Cores (MICs), in parallel. It presents the gains obtained when applying more devices to solve some problems and the main challenges that must be faced when using devices with as different architectures as CPUs and GPUs, with a greater focus on how to effectively achieve good load balancing between such heterogeneous devices.
Resumo:
To reduce the amount of time needed to solve the most complex Constraint Satisfaction Problems (CSPs) usually multi-core CPUs are used. There are already many applications capable of harnessing the parallel power of these devices to speed up the CSPs solving process. Nowadays, the Graphics Processing Units (GPUs) possess a level of parallelism that surpass the CPUs, containing from a few hundred to a few thousand cores and there are much less applications capable of solving CSPs on GPUs, leaving space for possible improvements. This article describes the work in progress for solving CSPs on GPUs and CPUs and compares results with some state-of-the-art solvers, presenting already some good results on GPUs.
Resumo:
Embedding intelligence in extreme edge devices allows distilling raw data acquired from sensors into actionable information, directly on IoT end-nodes. This computing paradigm, in which end-nodes no longer depend entirely on the Cloud, offers undeniable benefits, driving a large research area (TinyML) to deploy leading Machine Learning (ML) algorithms on micro-controller class of devices. To fit the limited memory storage capability of these tiny platforms, full-precision Deep Neural Networks (DNNs) are compressed by representing their data down to byte and sub-byte formats, in the integer domain. However, the current generation of micro-controller systems can barely cope with the computing requirements of QNNs. This thesis tackles the challenge from many perspectives, presenting solutions both at software and hardware levels, exploiting parallelism, heterogeneity and software programmability to guarantee high flexibility and high energy-performance proportionality. The first contribution, PULP-NN, is an optimized software computing library for QNN inference on parallel ultra-low-power (PULP) clusters of RISC-V processors, showing one order of magnitude improvements in performance and energy efficiency, compared to current State-of-the-Art (SoA) STM32 micro-controller systems (MCUs) based on ARM Cortex-M cores. The second contribution is XpulpNN, a set of RISC-V domain specific instruction set architecture (ISA) extensions to deal with sub-byte integer arithmetic computation. The solution, including the ISA extensions and the micro-architecture to support them, achieves energy efficiency comparable with dedicated DNN accelerators and surpasses the efficiency of SoA ARM Cortex-M based MCUs, such as the low-end STM32M4 and the high-end STM32H7 devices, by up to three orders of magnitude. To overcome the Von Neumann bottleneck while guaranteeing the highest flexibility, the final contribution integrates an Analog In-Memory Computing accelerator into the PULP cluster, creating a fully programmable heterogeneous fabric that demonstrates end-to-end inference capabilities of SoA MobileNetV2 models, showing two orders of magnitude performance improvements over current SoA analog/digital solutions.
Resumo:
In this thesis, I study the notion of program equivalences, i.e. proving that two programs can be used interchangeably without altering the overall observable behaviour. This definition is highly dependent on the contexts in which these programs can be used; does the context have exceptions, parallelism, etc... So proofs also need to be adapted according to the expressiveness of those contexts. This thesis presents on the pi-calculus – a concurrent programming language – under various typing constraints. Types allows us to impose different disciplines like forcing a sequential execution, or ensuring linearity, meaning an object can be used once. In each case, the bisimulation, a standard proof technique for the pi-calculus, needs to be adapted accordingly to obtain a suitable equivalence. We then test how using the modified bisimulations can be used to reason about a language with higher-order functions and references, which once translated into the pi-calculus satisfies the typing constraints.