Biblioteca Digital

692 resultados para Critic

Distributed multi-agent actor-critic algorithms with applications to stochastic path finding problems

Relevância:

20.00% 20.00%

Publicador:

Veja mais

Natural Belief-Critic: a reinforcement algorithm for parameter estimation in statistical spoken dialogue systems

Relevância:

20.00% 20.00%

Publicador:

Veja mais

Natural actor and belief critic: Reinforcement algorithm for learning parameters of dialogue systems modelled as POMDPs

Relevância:

20.00% 20.00%

Publicador:

Resumo:

This article presents a novel algorithm for learning parameters in statistical dialogue systems which are modeled as Partially Observable Markov Decision Processes (POMDPs). The three main components of a POMDP dialogue manager are a dialogue model representing dialogue state information; a policy that selects the system's responses based on the inferred state; and a reward function that specifies the desired behavior of the system. Ideally both the model parameters and the policy would be designed to maximize the cumulative reward. However, while there are many techniques available for learning the optimal policy, no good ways of learning the optimal model parameters that scale to real-world dialogue systems have been found yet. The presented algorithm, called the Natural Actor and Belief Critic (NABC), is a policy gradient method that offers a solution to this problem. Based on observed rewards, the algorithm estimates the natural gradient of the expected cumulative reward. The resulting gradient is then used to adapt both the prior distribution of the dialogue model parameters and the policy parameters. In addition, the article presents a variant of the NABC algorithm, called the Natural Belief Critic (NBC), which assumes that the policy is fixed and only the model parameters need to be estimated. The algorithms are evaluated on a spoken dialogue system in the tourist information domain. The experiments show that model parameters estimated to maximize the expected cumulative reward result in significantly improved performance compared to the baseline hand-crafted model parameters. The algorithms are also compared to optimization techniques using plain gradients and state-of-the-art random search algorithms. In all cases, the algorithms based on the natural gradient work significantly better. © 2011 ACM.

Veja mais

Reinforcement learning using a continuous time actor-critic framework with spiking neurons.

Relevância:

20.00% 20.00%

Publicador:

Resumo:

Animals repeat rewarded behaviors, but the physiological basis of reward-based learning has only been partially elucidated. On one hand, experimental evidence shows that the neuromodulator dopamine carries information about rewards and affects synaptic plasticity. On the other hand, the theory of reinforcement learning provides a framework for reward-based learning. Recent models of reward-modulated spike-timing-dependent plasticity have made first steps towards bridging the gap between the two approaches, but faced two problems. First, reinforcement learning is typically formulated in a discrete framework, ill-adapted to the description of natural situations. Second, biologically plausible models of reward-modulated spike-timing-dependent plasticity require precise calculation of the reward prediction error, yet it remains to be shown how this can be computed by neurons. Here we propose a solution to these problems by extending the continuous temporal difference (TD) learning of Doya (2000) to the case of spiking neurons in an actor-critic network operating in continuous time, and with continuous state and action representations. In our model, the critic learns to predict expected future rewards in real time. Its activity, together with actual rewards, conditions the delivery of a neuromodulatory TD signal to itself and to the actor, which is responsible for action choice. In simulations, we show that such an architecture can solve a Morris water-maze-like navigation task, in a number of trials consistent with reported animal performance. We also use our model to solve the acrobot and the cartpole problems, two complex motor control tasks. Our model provides a plausible way of computing reward prediction error in the brain. Moreover, the analytically derived learning rule is consistent with experimental evidence for dopamine-modulated spike-timing-dependent plasticity.

Veja mais

Talking Balls: Critic (VWE) Seeks Straight-Talking Man, or 'Don't Put Your Son on the Stage, Mr Worthington'

Relevância:

20.00% 20.00%

Publicador:

Veja mais

Pierre Bourdieu as a Sociologist of the Economy and Critic of 'Globalisation'

Relevância:

20.00% 20.00%

Publicador:

Resumo:

This article examines Pierre Bourdieu's sociology of the economy and his more recent politically engaged interventions on 'globalisation'. Many scholars regard these as not being in the same academic league as his classic studies on taste, academia, and state elites, etc., and, instead, dismiss them as a private matter or even, as the spleen of Pierre Bourdieu, the individual. This paper questions this disjunction of the 'academic' and 'politically engaged' sides of Pierre Bourdieu's work. First, it argues that his most recent interventions against a neo-liberal globalisation were the logical result of a particular definition of intellectual practice that had been outlined before in his sociology of the intellectual field. It then demonstrates that Bourdieu's economic sociology and critique of contemporary capitalism not only does not contradict his earlier research, but that it provides valuable and original insights into the current transformation of the political economy of the advanced capitalist countries. The paper concludes with a suggestion of how to strengthen the theoretical foundation of Bourdieu's analysis of contemporary capitalism by relating it to and making it compatible with alternative approaches in the tradition of critical political economy.

Veja mais