3 The asymptotic value in finite stochastic games

Miquel Barton

3 The asymptotic value in finite stochastic games

Miquel Barton

2012

Sign up for access to the world's latest research

checkGet notified about relevant papers

checkSave papers to use in your research

checkJoin the discussion with peers

checkTrack your impact

Abstract

We provide a direct, elementary proof for the existence of lim λ→0 v λ , where v λ is the value of λ-discounted finite two-person zero-sum stochastic game. 1 Introduction Two-person zero-sum stochastic games were introduced by Shapley [4]. They are described by a 5-tuple (Ω, I, J , q, g), where Ω is a finite set of states, I and J are finite sets of actions, g : Ω × I × J → [0, 1] is the payoff, q : Ω × I × J → ∆(Ω) the transition and, for any finite set X, ∆(X) denotes the set of probability distributions over X. The functions g and q are bilinearly extended to Ω × ∆(I) × ∆(J). The stochastic game with initial state ω ∈ Ω and discount factor λ ∈ (0, 1] is denoted by Γ λ (ω) and is played as follows: at stage m ≥ 1, knowing the current state ω m , the players choose actions (i m , j m) ∈ I × J ; their choice produces a stage payoff g(ω m , i m , j m) and influences the transition: a new state ω m+1 is chosen according to the probability distribution q(•|ω m , i m , j m). At the end of the game, player 1 receives m≥1 λ(1 − λ) m−1 g(ω m , i m , j m) from player 2. The game Γ λ (ω) has a value v λ (ω), and v λ = (v λ (ω)) ω∈Ω is the unique fixed point of the so-called Shapley operator [4], i.e. v λ = Φ(λ, v λ), where for all f ∈ R Ω :

Miquel Barton

2018

In a zero-sum stochastic game, at each stage, two adversary players take decisions and receive a stage payoff determined by them and by a random variable representing the state of nature. The total payoff is the discounted sum of the stage payoffs. Assume that the players are very patient and use optimal strategies. We then prove that, at any point in the game, players get essentially the same expected payoff: the payoff is constant. This solves a conjecture by Sorin, Venel and Vigeral (2010). The proof relies on the semi-algebraic approach for discounted stochastic games introduced by Bewley and Kohlberg (1976), on the theory of Markov chains with rare transitions, initiated by Friedlin and Wentzell (1984), and on some variational inequalities for value functions inspired by the recent work of Davini, Fathi, Iturriaga and Zavidovique (2016).

Log In

3 The asymptotic value in finite stochastic games

Sign up for access to the world's latest research

Abstract

Related papers