998 resultados para Normal approximation


Relevância:

20.00% 20.00%

Publicador:

Resumo:

We develop in this article the first actor-critic reinforcement learning algorithm with function approximation for a problem of control under multiple inequality constraints. We consider the infinite horizon discounted cost framework in which both the objective and the constraint functions are suitable expected policy-dependent discounted sums of certain sample path functions. We apply the Lagrange multiplier method to handle the inequality constraints. Our algorithm makes use of multi-timescale stochastic approximation and incorporates a temporal difference (TD) critic and an actor that makes a gradient search in the space of policy parameters using efficient simultaneous perturbation stochastic approximation (SPSA) gradient estimates. We prove the asymptotic almost sure convergence of our algorithm to a locally optimal policy. (C) 2010 Elsevier B.V. All rights reserved.

Relevância:

20.00% 20.00%

Publicador:

Resumo:

The reaction of 2-formylbenzenesulfonyl chloride 1 and its pseudo isomer 2 with primary amines give either the corresponding sulfonamido Schiff bases or the corresponding 2-formylbenzenesulfonamide depending on the concentration of the amine used. The derivatives exist as an equilibrium mixture of the corresponding sulfonamide and 2-alkyl-3-hydroxy(or 3-aminoalkyl)-benzisothiazole-1,1-dioxide. Spectroscopic studies suggest that 2-formylbenzenesulfonamides exist as benzisothiazole-1,1-dioxides in the solid state, as a mixture of 2-formylbenzenesulfonamide and the corresponding benzisothiazole-1,1-dioxide in solution and as 2-formyl-benzenesulfonamides in the gas phase.

Relevância:

20.00% 20.00%

Publicador:

Resumo:

We have investigated tunneling conductances in disordered, normally conducting perovskite oxides close to the metal�insulator transition. We show that the normal state tunneling conductance of perovskite oxides can be cast in a general form G(V) = G0[1 + curly logical orV/V*curly logical orn] with 1?n?0.5 and where V* is an intrinsic energy scale. The exponent n graduall y increases from 0.5 to 1 as the metal-insulator (M-I) transition is approached. In the high-Tc Bi(2212) cuprates, the normally observed, linear G(V)(n=1) can be made sub-linear (n<1) by substitution of Ca with Y. From the similarity of the linear conductances, we suggest proximity to the M-I transition as a likely cause for this G(V)logical or, bar below V dependence. In systems showing linear conductances (nreverse similar, equals1), we find that ?G/?Vreverse similar, equalsG?0 with ?reverse similar, equals 1 and the intrinsic energy scale V*reverse similar, equals25�75 meV in the different oxides investigated.

Relevância:

20.00% 20.00%

Publicador:

Resumo:

A two timescale stochastic approximation scheme which uses coupled iterations is used for simulation-based parametric optimization as an alternative to traditional "infinitesimal perturbation analysis" schemes, It avoids the aggregation of data present in many other schemes. Its convergence is analyzed, and a queueing example is presented.

Relevância:

20.00% 20.00%

Publicador:

Resumo:

A two-time scale stochastic approximation algorithm is proposed for simulation-based parametric optimization of hidden Markov models, as an alternative to the traditional approaches to ''infinitesimal perturbation analysis.'' Its convergence is analyzed, and a queueing example is presented.

Relevância:

20.00% 20.00%

Publicador:

Resumo:

We propose, for the first time, a reinforcement learning (RL) algorithm with function approximation for traffic signal control. Our algorithm incorporates state-action features and is easily implementable in high-dimensional settings. Prior work, e. g., the work of Abdulhai et al., on the application of RL to traffic signal control requires full-state representations and cannot be implemented, even in moderate-sized road networks, because the computational complexity exponentially grows in the numbers of lanes and junctions. We tackle this problem of the curse of dimensionality by effectively using feature-based state representations that use a broad characterization of the level of congestion as low, medium, or high. One advantage of our algorithm is that, unlike prior work based on RL, it does not require precise information on queue lengths and elapsed times at each lane but instead works with the aforementioned described features. The number of features that our algorithm requires is linear to the number of signaled lanes, thereby leading to several orders of magnitude reduction in the computational complexity. We perform implementations of our algorithm on various settings and show performance comparisons with other algorithms in the literature, including the works of Abdulhai et al. and Cools et al., as well as the fixed-timing and the longest queue algorithms. For comparison, we also develop an RL algorithm that uses full-state representation and incorporates prioritization of traffic, unlike the work of Abdulhai et al. We observe that our algorithm outperforms all the other algorithms on all the road network settings that we consider.

Relevância:

20.00% 20.00%

Publicador:

Resumo:

This paper investigates the propagation of a strong shock into an inhomogeneous medium using the new theory of shock dynamics. The equations are simple to solve and involve no trial-and-error method commonly used in this case. The results compare favourably with earlier results obtained in the case of self-similar flows, which arise as a special case of this theory.

Relevância:

20.00% 20.00%

Publicador:

Resumo:

The actor-critic algorithm of Barto and others for simulation-based optimization of Markov decision processes is cast as a two time Scale stochastic approximation. Convergence analysis, approximation issues and an example are studied.

Relevância:

20.00% 20.00%

Publicador:

Relevância:

20.00% 20.00%

Publicador:

Resumo:

We consider the problem of wireless channel allocation to multiple users. A slot is given to a user with a highest metric (e.g., channel gain) in that slot. The scheduler may not know the channel states of all the users at the beginning of each slot. In this scenario opportunistic splitting is an attractive solution. However this algorithm requires that the metrics of different users form independent, identically distributed (iid) sequences with same distribution and that their distribution and number be known to the scheduler. This limits the usefulness of opportunistic splitting. In this paper we develop a parametric version of this algorithm. The optimal parameters of the algorithm are learnt online through a stochastic approximation scheme. Our algorithm does not require the metrics of different users to have the same distribution. The statistics of these metrics and the number of users can be unknown and also vary with time. Each metric sequence can be Markov. We prove the convergence of the algorithm and show its utility by scheduling the channel to maximize its throughput while satisfying some fairness and/or quality of service constraints.

Relevância:

20.00% 20.00%

Publicador:

Resumo:

We consider the problem of scheduling a wireless channel among multiple users. A slot is given to a user with a highest metric (e.g., channel gain) in that slot. The scheduler may not know the channel states of all the users at the beginning of each slot. In this scenario opportunistic splitting is an attractive solution. However this algorithm requires that the metrics of different users form independent, identically distributed (iid) sequences with same distribution and that their distribution and number be known to the scheduler. This limits the usefulness of opportunistic splitting. In this paper we develop a parametric version of this algorithm. The optimal parameters of the algorithm are learnt online through a stochastic approximation scheme. Our algorithm does not require the metrics of different users to have the same distribution. The statistics of these metrics and the number of users can be unknown and also vary with time. We prove the convergence of the algorithm and show its utility by scheduling the channel to maximize its throughput while satisfying some fairness and/or quality of service constraints.

Relevância:

20.00% 20.00%

Publicador:

Resumo:

Epitaxial LaNiO3 thin films have been grown on SrTiO3 and several other substrates by pulsed laser deposition. The films are observed to be metallic down to 15 K, and the temperature dependence of resistivity is similar to that of bulk LaNiO3. Epitaxial, c-axis oriented YBa2Cu3O7-x films with good superconducting properties have been grown on the LaNiO3 (100) films. I-V characteristics of the YBa2Cu3O7-x-LaNiO3 junction are linear, indicating ohmic contact between them.

Relevância:

20.00% 20.00%

Publicador:

Resumo:

We explore a pseudodynamic form of the quadratic parameter update equation for diffuse optical tomographic reconstruction from noisy data. A few explicit and implicit strategies for obtaining the parameter updates via a semianalytical integration of the pseudodynamic equations are proposed. Despite the ill-posedness of the inverse problem associated with diffuse optical tomography, adoption of the quadratic update scheme combined with the pseudotime integration appears not only to yield higher convergence, but also a muted sensitivity to the regularization parameters, which include the pseudotime step size for integration. These observations are validated through reconstructions with both numerically generated and experimentally acquired data. (C) 2011 Optical Society of America

Relevância:

20.00% 20.00%

Publicador:

Resumo:

One of the assumptions of the van der Waals and Platteeuw theory for gas hydrates is that the host water lattice is rigid and not distorted by the presence of guest molecules. In this work, we study the effect of this approximation on the triple-point lines of the gas hydrates. We calculate the triple-point lines of methane and ethane hydrates via Monte Carlo molecular simulations and compare the simulation results with the predictions of van der Waals and Platteeuw theory. Our study shows that even if the exact intermolecular potential between the guest molecules and water is known, the dissociation temperatures predicted by the theory are significantly higher. This has serious implications to the modeling of gas hydrate thermodynamics, and in spite of the several impressive efforts made toward obtaining an accurate description of intermolecular interactions in gas hydrates, the theory will suffer from the problem of robustness if the issue of movement of water molecules is not adequately addressed.