944 resultados para Heterogeneous Processors
Resumo:
Breakthrough advances in microprocessor technology and efficient power management have altered the course of development of processors with the emergence of multi-core processor technology, in order to bring higher level of processing. The utilization of many-core technology has boosted computing power provided by cluster of workstations or SMPs, providing large computational power at an affordable cost using solely commodity components. Different implementations of message-passing libraries and system softwares (including Operating Systems) are installed in such cluster and multi-cluster computing systems. In order to guarantee correct execution of message-passing parallel applications in a computing environment other than that originally the parallel application was developed, review of the application code is needed. In this paper, a hybrid communication interfacing strategy is proposed, to execute a parallel application in a group of computing nodes belonging to different clusters or multi-clusters (computing systems may be running different operating systems and MPI implementations), interconnected with public or private IP addresses, and responding interchangeably to user execution requests. Experimental results demonstrate the feasibility of this proposed strategy and its effectiveness, through the execution of benchmarking parallel applications.
Resumo:
Modern embedded systems embrace many-core shared-memory designs. Due to constrained power and area budgets, most of them feature software-managed scratchpad memories instead of data caches to increase the data locality. It is therefore programmers’ responsibility to explicitly manage the memory transfers, and this make programming these platform cumbersome. Moreover, complex modern applications must be adequately parallelized before they can the parallel potential of the platform into actual performance. To support this, programming languages were proposed, which work at a high level of abstraction, and rely on a runtime whose cost hinders performance, especially in embedded systems, where resources and power budget are constrained. This dissertation explores the applicability of the shared-memory paradigm on modern many-core systems, focusing on the ease-of-programming. It focuses on OpenMP, the de-facto standard for shared memory programming. In a first part, the cost of algorithms for synchronization and data partitioning are analyzed, and they are adapted to modern embedded many-cores. Then, the original design of an OpenMP runtime library is presented, which supports complex forms of parallelism such as multi-level and irregular parallelism. In the second part of the thesis, the focus is on heterogeneous systems, where hardware accelerators are coupled to (many-)cores to implement key functional kernels with orders-of-magnitude of speedup and energy efficiency compared to the “pure software” version. However, three main issues rise, namely i) platform design complexity, ii) architectural scalability and iii) programmability. To tackle them, a template for a generic hardware processing unit (HWPU) is proposed, which share the memory banks with cores, and the template for a scalable architecture is shown, which integrates them through the shared-memory system. Then, a full software stack and toolchain are developed to support platform design and to let programmers exploiting the accelerators of the platform. The OpenMP frontend is extended to interact with it.
Resumo:
The advances in low power micro-processors, wireless networks and embedded systems have raised the need to utilize the significant resources of mobile devices. These devices for example, smart phones, tablets, laptops, wearables, and sensors are gaining enormous processing power, storage capacity and wireless bandwidth. In addition, the advancement in wireless mobile technology has created a new communication paradigm via which a wireless network can be created without any priori infrastructure called mobile ad hoc network (MANET). While progress is being made towards improving the efficiencies of mobile devices and reliability of wireless mobile networks, the mobile technology is continuously facing the challenges of un-predictable disconnections, dynamic mobility and the heterogeneity of routing protocols. Hence, the traditional wired, wireless routing protocols are not suitable for MANET due to its unique dynamic ad hoc nature. Due to the reason, the research community has developed and is busy developing protocols for routing in MANET to cope with the challenges of MANET. However, there are no single generic ad hoc routing protocols available so far, which can address all the basic challenges of MANET as mentioned before. Thus this diverse range of ever growing routing protocols has created barriers for mobile nodes of different MANET taxonomies to intercommunicate and hence wasting a huge amount of valuable resources. To provide interaction between heterogeneous MANETs, the routing protocols require conversion of packets, meta-model and their behavioural capabilities. Here, the fundamental challenge is to understand the packet level message format, meta-model and behaviour of different routing protocols, which are significantly different for different MANET Taxonomies. To overcome the above mentioned issues, this thesis proposes an Interoperable Framework for heterogeneous MANETs called IF-MANET. The framework hides the complexities of heterogeneous routing protocols and provides a homogeneous layer for seamless communication between these routing protocols. The framework creates a unique Ontology for MANET routing protocols and a Message Translator to semantically compare the packets and generates the missing fields using the rules defined in the Ontology. Hence, the translation between an existing as well as newly arriving routing protocols will be achieved dynamically and on-the-fly. To discover a route for the delivery of packets across heterogeneous MANET taxonomies, the IF-MANET creates a special Gateway node to provide cluster based inter-domain routing. The IF-MANET framework can be used to develop different middleware applications. For example: Mobile grid computing that could potentially utilise huge amounts of aggregated data collected from heterogeneous mobile devices. Disaster & crises management applications can be created to provide on-the-fly infrastructure-less emergency communication across organisations by utilising different MANET taxonomies.
Resumo:
With security and surveillance, there is an increasing need to process image data efficiently and effectively either at source or in a large data network. Whilst a Field-Programmable Gate Array (FPGA) has been seen as a key technology for enabling this, the design process has been viewed as problematic in terms of the time and effort needed for implementation and verification. The work here proposes a different approach of using optimized FPGA-based soft-core processors which allows the user to exploit the task and data level parallelism to achieve the quality of dedicated FPGA implementations whilst reducing design time. The paper also reports some preliminary
progress on the design flow to program the structure. An implementation for a Histogram of Gradients algorithm is also reported which shows that a performance of 328 fps can be achieved with this design approach, whilst avoiding the long design time, verification and debugging steps associated with conventional FPGA implementations.
Resumo:
In this work an iterative strategy is developed to tackle the problem of coupling dimensionally-heterogeneous models in the context of fluid mechanics. The procedure proposed here makes use of a reinterpretation of the original problem as a nonlinear interface problem for which classical nonlinear solvers can be applied. Strong coupling of the partitions is achieved while dealing with different codes for each partition, each code in black-box mode. The main application for which this procedure is envisaged arises when modeling hydraulic networks in which complex and simple subsystems are treated using detailed and simplified models, correspondingly. The potentialities and the performance of the strategy are assessed through several examples involving transient flows and complex network configurations.
Resumo:
In the present work, cellulose obtained from sisal, which is a source of rapid growth, was used. Cellulose acetates were produced in heterogeneous medium, using acetic anhydride as esterifying agent and iodine as catalyst, to check if the procedure described in the literature for commercial cellulose also is adequate to sisal cellulose. The results indicated that iodine is an excellent catalyst to obtain sisal cellulose acetates, but the reaction is so fast as described in the literature when, instead of sisal, lower average molar weight cellulose (microcrystalline) is used. The crystallinity index (I(c)) of sisal cellulose acetates diminished compared to sisal cellulose, but there was no direct correlation between their degree of substitution (DS) and I(c). Probably acetyl groups were introduced more homogeneously along the short chains of microcrystalline cellulose, when compared to sisal cellulose, and then for microcrystalline cellulose acetates the Ic decreases as DS increases. Using the linear correlation that was found between degree of substitution (DS) and time reaction is possible to control the DS of sisal cellulose acetates, considering a large interval of degrees of substitution (0.3-2.8).
Resumo:
Background: Although the Clock Drawing Test (CDT) is the second most used test in the world for the screening of dementia, there is still debate over its sensitivity specificity, application and interpretation in dementia diagnosis. This study has three main aims: to evaluate the sensitivity and specificity of the CDT in a sample composed of older adults with Alzheimer`s disease (AD) and normal controls; to compare CDT accuracy to the that of the Mini-mental State Examination (MMSE) and the Cambridge Cognitive Examination (CAMCOG), and to test whether the association of the MMSE with the CDT leads to higher or comparable accuracy as that reported for the CAMCOG. Methods: Cross-sectional assessment was carried out for 121 AD and 99 elderly controls with heterogeneous educational levels from a geriatric outpatient clinic who completed the Cambridge Examination for Mental Disorder of the Elderly (CAMDEX). The CDT was evaluated according to the Shulman, Mendez and Sunderland scales. Results: The CDT showed high sensitivity and specificity. There were significant correlations between the CDT and the MMSE (0.700-0.730; p < 0.001) and between the CDT and the CAMCOG (0.753-0.779; p < 0.001). The combination of the CDT with the MMSE improved sensitivity and specificity (SE = 89.2-90%; SP = 71.7-79.8%). Subgroup analysis indicated that for elderly people with lower education, sensitivity and specificity were both adequate and high. Conclusions: The CDT is a robust screening test when compared with the MMSE or the CAMCOG, independent of the scale used for its interpretation. The combination with the MMSE improves its performance significantly, becoming equivalent to the CAMCOG.
Resumo:
The objective of this study was to estimate the first-order intrinsic kinetic constant (k(1)) and the liquid-phase mass transfer coefficient (k(c)) in a bench-scale anaerobic sequencing batch biofilm reactor (ASBBR) fed with glucose. A dynamic heterogeneous mathematical model, considering two phases (liquid and solid), was developed through mass balances in the liquid and solid phases. The model was adjusted to experimental data obtained from the ASBBR applied for the treatment of glucose-based synthetic wastewater with approximately 500 mg L-1 of glucose, operating in 8 h batch cycles, at 30 degrees C and 300 rpm. The values of the parameters obtained were 0.8911 min(-1) for k(1) and 0.7644 cm min(-1) for kc. The model was validated utilizing the estimated parameters with data obtained from the ASBBR operating in 3 h batch cycles, with a good representation of the experimental behavior. The solid-phase mass transfer flux was found to be the limiting step of the overall glucose conversion rate.
Resumo:
A modeling study was completed to develop a methodology that combines the sequencing and finite difference methods for the simulation of a heterogeneous model of a tubular reactor applied in the treatment of wastewater. The system included a liquid phase (convection diffusion transport) and a solid phase (diffusion reaction) that was obtained by completing a mass balance in the reactor and in the particle, respectively. The model was solved using a pilot-scale horizontal-flow anaerobic immobilized biomass (HAIB) reactor to treat domestic sewage, with the concentration results compared with the experimental data. A comparison of the behavior of the liquid phase concentration profile and the experimental results indicated that both the numerical methods offer a good description of the behavior of the concentration along the reactor. The advantage of the sequencing method over the finite difference method is that it is easier to apply and requires less computational time to model the dynamic simulation of outlet response of HAIB.
Resumo:
In this paper, we consider a real-life heterogeneous fleet vehicle routing problem with time windows and split deliveries that occurs in a major Brazilian retail group. A single depot attends 519 stores of the group distributed in 11 Brazilian states. To find good solutions to this problem, we propose heuristics as initial solutions and a scatter search (SS) approach. Next, the produced solutions are compared with the routes actually covered by the company. Our results show that the total distribution cost can be reduced significantly when such methods are used. Experimental testing with benchmark instances is used to assess the merit of our proposed procedure. (C) 2008 Published by Elsevier B.V.
Resumo:
A series of TiO2 samples with different anatase-to-rutile ratios was prepared by calcination, and the roles of the two crystallite phases of titanium(IV) oxide (TiO2) on the photocatalytic activity in oxidation of phenol in aqueous solution were studied. High dispersion of nanometer-sized anatase in the silica matrix and the possible bonding of Si-O-Ti in SiO2/TiO2 interface were found to stabilize the crystallite transformation from anatase to rutile. The temperature for this transformation was 1200 degrees C for the silica-titania (ST) sample, much higher than 700 degrees C for Degussa P25, a benchmarking photocatalyst. It is shown that samples with higher anatase-to-rutile ratios have higher activities for phenol degradation. However, the activity did not totally disappear after a complete crystallite transformation for P25 samples, indicating some activity of the rutile phase. Furthermore, the activity for the ST samples after calcination decreased significantly, even though the amount of anatase did not change much. The activity of the same samples with different anatase-to-rutile ratios is more related to the amount of the surface-adsorbed water and hydroxyl groups and surface area. The formation of rutile by calcination would reduce the surface-adsorbed water and hydroxyl groups and surface area, leading to the decrease in activity.
Resumo:
In this and a preceding paper, we provide an introduction to the Fujitsu VPP range of vector-parallel supercomputers and to some of the computational chemistry software available for the VPP. Here, we consider the implementation and performance of seven popular chemistry application packages. The codes discussed range from classical molecular dynamics to semiempirical and ab initio quantum chemistry. All have evolved from sequential codes, and have typically been parallelised using a replicated data approach. As such they are well suited to the large-memory/fast-processor architecture of the VPP. For one code, CASTEP, a distributed-memory data-driven parallelisation scheme is presented. (C) 2000 Published by Elsevier Science B.V. All rights reserved.
Resumo:
We have previously demonstrated that or-smooth muscle (alpha -SM) actin is predominantly distributed in the central region and beta -non-muscle (beta -NM) actin in the periphery of cultured rabbit aortic smooth muscle cells (SMCs). To determine whether this reflects a special form of segregation of contractile and cytoskeletal components in SMCs, this study systematically investigated the distribution relationship of structural proteins using high-resolution confocal laser scanning fluorescent microscopy. Not only isoactins but also smooth muscle myosin heavy chain, alpha -actinin, vinculin, and vimentin were heterogeneously distributed in the cultured SMCs. The predominant distribution of beta -NM actin in the cell periphery was associated with densely distributed vinculin plaques and disrupted or striated myosin and ol-actinin aggregates, which may reflect a process of stress fiber assembly during cell spreading and focal adhesion formation. The high-level labeling of alpha -SM actin in the central portion of stress fibers was related to continuous myosin and punctate alpha -actinin distribution, which may represent the maturation of the fibrillar structures. The findings also suggest that the stress fibers, in which actin and myosin filaments organize into sar-comere-like units with alpha -actinin-rich dense bodies analogous to Z-lines, are the contractile vimentin structures of cultured SMCs that link to the network of vimentin-containing intermediate alpha -actinin filaments through the dense bodies and dense plaques.
Resumo:
The evolution of event time and size statistics in two heterogeneous cellular automaton models of earthquake behavior are studied and compared to the evolution of these quantities during observed periods of accelerating seismic energy release Drier to large earthquakes. The two automata have different nearest neighbor laws, one of which produces self-organized critical (SOC) behavior (PSD model) and the other which produces quasi-periodic large events (crack model). In the PSD model periods of accelerating energy release before large events are rare. In the crack model, many large events are preceded by periods of accelerating energy release. When compared to randomized event catalogs, accelerating energy release before large events occurs more often than random in the crack model but less often than random in the PSD model; it is easier to tell the crack and PSD model results apart from each other than to tell either model apart from a random catalog. The evolution of event sizes during the accelerating energy release sequences in all models is compared to that of observed sequences. The accelerating energy release sequences in the crack model consist of an increase in the rate of events of all sizes, consistent with observations from a small number of natural cases, however inconsistent with a larger number of cases in which there is an increase in the rate of only moderate-sized events. On average, no increase in the rate of events of any size is seen before large events in the PSD model.