Download Robust equilibria and ε-dominance

Survey
yes no Was this document useful for you?
   Thank you for your participation!

* Your assessment is very important for improving the work of artificial intelligence, which forms the content of this project

Document related concepts

Game mechanics wikipedia , lookup

Turns, rounds and time-keeping systems in games wikipedia , lookup

Minimax wikipedia , lookup

Deathmatch wikipedia , lookup

The Evolution of Cooperation wikipedia , lookup

Artificial intelligence in video games wikipedia , lookup

Prisoner's dilemma wikipedia , lookup

Nash equilibrium wikipedia , lookup

Evolutionary game theory wikipedia , lookup

Chicken (game) wikipedia , lookup

Transcript
Robust equilibria and ε-dominance∗
William Geller†and Rachel Hemphill‡
April 11, 2014
Abstract
We propose a resolution of the backward induction paradox and some
other anomalies in game theory based on refinements of Radner’s ε-equilibria.
We avoid the usual very strong assumption of common knowledge of rationality with complete information, and posit instead some inescapable
uncertainty about others’ actions. The central idea is to require a solution
for a noncooperative game to exhibit some degree of robustness. When
ε = 0, our ε-robust equilibria reduce to Nash equilibria, but for positive ε
our solutions in such games as centipede and traveler’s dilemma contrast
sharply with the Nash predictions and fit very well with experiment and
intuition.
Game theory has had great difficulty dealing convincingly with an important
family of games whose equilibria may be found by a long backward induction.
Exemplars of this problematic but central family are the finitely repeated prisoner’s dilemma [19] and Rosenthal’s centipede game [27], [18]; Basu’s traveler’s
dilemma [3] and imperfect price competition games [9] are related examples. On
the one hand, the theory of Nash equilibrium clearly predicts one sort of behavior in these games, which we could call a race to the bottom; on the other hand,
this behavior seems unreasonable, even for rational, self-interested agents, an intuition reflected in a large number of experiments, including those with subjects
well aware of the Nash equilibrium [5]. Many (overlapping) attempts have been
made to address this seeming paradox: for example, the introduction of a small
chance of irrationality of a carefully chosen type [17], the introduction of limits
on the reasoning or computational power [24] of agents, or a cost for the use of
complex strategies, the weakening of maximizing behavior to near-maximizing
behavior [26], [31], and the replacement of maximizing behavior with a sort of
stochastic or smoothed maximization [22]. There are difficulties with each of
these approaches. For example, the first is open to the objection that alternative
choices of the seeding irrationality lead to a wide variety of outcomes [11], the
∗ We
thank Bob Anderson for his helpful comments.
of Mathematical Sciences, Indiana University-Purdue University Indianapolis. [email protected]; corresponding author.
‡ Department of Mathematical Sciences, Indiana University-Purdue University Indianapolis. [email protected].
† Department
1
second seems not to apply well to games with simple structures like centipede
or traveler’s dilemma, especially with sophisticated agents,1 the third typically
lacks specificity in its predictions, and the fourth depends on the choice both of
a type of smoothing function and a diffusion parameter.
Our approach is closest in spirit to the third and fourth lines of attack,
though we believe it possesses significant advantages. While it is much simpler
than the stochastic framework of quantal response equilibria, it involves an often
radical refinement of the set of ε-equilibria despite the introduction of no further
parameters. This allows for quite sharp predictions in the case of centipede
and traveler’s dilemma which are in striking agreement with experiments and
intuition.
We were motivated originally by the much studied paradox of the finitely
repeated prisoner’s dilemma, which is more than half a century old. Considering
a prisoner’s dilemma repeated 100 times, where every Nash equilibrium leads
both players to play tough on every round, Luce and Raiffa [19] state that they
would not play to a Nash equilibrium. In fact, if strategies were restricted to
those which play nice before some round k, 1 ≤ k ≤ 100, as long as the opponent
also plays nice, and after a tough play by the opponent or the arrival of round
k play tough until the end, they write that they would probably play a strategy
k, where k “is some number in the nineties.” We are able to vindicate their
intuition for the restricted strategy game; the unrestricted game is still beyond
our grasp. In the same way, we resolve the paradox of the centipede game and
the traveler’s dilemma. We also address some limiting instances of stag hunt,
the prototypical assurance game whose history dates to Rousseau [32], [4], as
well as some other relevant examples from the literature.
We introduce here a small circle of closely related solution concepts for games
in strategic form centered on the notions of ε-dominance and ε-robustness. Our
aim is to expand the normative and positive scope of noncooperative game
theory with the simplest possible tools.
A Nash equilibrium requires zero regret from each agent if he has correctly
anticipated others’ strategies, but allows massive regret if another’s strategy is
unforeseen. This makes Nash equilibria precariously dependent on very strong
assumptions. We will instead require small regret from each agent if he predicts
correctly, but also impose some robustness on his strategy, i.e. seek to limit his
regret if his prediction is incorrect. This turns out to be surprisingly fruitful.
0.1
Centipede
Centipede is an extensive-form game which alternates between decision nodes
for two players. At each odd-numbered stage 2l + 1 of the game, player 1 may
grab, which ends the game and results in each player receiving a payoff of l + 1,
or pass, which causes the game to continue to player 2. Similarly, on an evennumbered stage 2l of the game, player 2 may pass, causing the game to continue
1 See the traveler’s dilemma played by the Game Theory Society members in Section 4
below.
2
to player 1’s next decision node, or can grab, which results in player 1 receiving
l − 1 and player 2 receiving l + 2. The game ends at round 198, where player 2
makes the choice between passing, which results in each player receiving 100, or
grabbing, which results in player 1 receiving 98 and player 2 receiving 101; see
Figure 1. A rational player 2 should grab at his final opportunity. Furthermore,
if he knows player 2 will grab at round 198, a rational player 1 has no incentive
to pass at round 197, so he will surely grab then. Continuing in this fashion, we
arrive at the conclusion that player 1 must grab at round 1. This is the unique
rationalizable equilibrium outcome, in the sense of Bernheim for the normal
form, since it is obtained by iterated elimination of dominated strategies and is
thus the unique Nash equilibrium. However, it seems counter to intuition about
how rational, self-interested players should play.
1
A
2
A
1
. . .
D
D
D
(1,1)
(0,3)
(2,2)
1
A
D
(98,98)
2
A
1
D
(97,100)
A
D
(99,99)
2
(100,100)
D
(98,101)
Figure 1: Centipede.
1
ε-dominance
Consider a finite game G in normal, i.e. strategic, form. Let I = {1, 2, ..., n}
be the set of players. Let Ai be the finite set of player i’s pure strategies,
for i ∈ I. Let Si = ∆(Ai ), the set of player i’s mixed strategies,
i.e. the
Q
(|A
|
−
1)-dimensional
simplex
of
probability
vectors.
Let
A
=
A
, A−i =
i
i
i∈I
Q
Q
Q
j∈I,j6=i Sj . Then, a (mixed) strategy
i∈I Si , and S−i =
j∈I,j6=i Aj , S =
profile is an element s = (s1 , s2 , ..., sn ) ∈ S. Let ui : S → R be the payoff
function for player i.
We introduce here several definitions.
Definition 1. A strategy s∗i for player i is called ε-dominant if for all si ∈ Si
and for all s−i ∈ S−i ,
ui (si , s−i ) − ui (s∗i , s−i ) ≤ ε.
That is, a strategy is ε-dominant if it never engenders more than ε regret.2 In
particular, a 0-dominant strategy is never regretted. Equivalently, a strategy s∗i
is ε-dominant if, for all ai ∈ Ai and for all a−i ∈ A−i , ui (ai , a−i ) − ui (s∗i , a−i ) ≤
ε. An undominated ε-dominant strategy is called properly ε-dominant.
Definition 2. Let δi = δi (G) = inf{ε ≥ 0|there exists a strategy si which is
ε-dominant.} If si is ε-dominant for ε = δi , then si is called most-dominant.
We call δi the (dominance) defect for i.
2 Abreu
and Matsushima [1] consider the ε-domination of one strategy by another.
3
The dominance defect measures how far a player is from having a dominant
strategy, and a most-dominant strategy is one that is as near to dominant as
available for the player. As noted in Proposition 1 below, δi is just the minimax
regret considered in decision theory by Savage [30], and so a most-dominant
strategy is just a strategy minimizing maximum regret.
The (dominance) defect δ = δ(G) of the game G is the maximum of the
defects for the players, δi .
1.1
Centipede Analysis
In centipede, at every decision node, there exists a single possible history: both
players must have passed at each of their prior decision nodes. A strategy for a
player is then simply the decision to pass or grab at each of his decision nodes.
However, since the game ends once a player grabs, a strategy is completely
described by the earliest stage at which the player would grab.
So, A1 = {1, 3, 5, ..., 197,Never=199} and A2 = {2, 4, ..., 198,Never=200},
where strategy k indicates the earliest node at which the player would grab.
The best response to an opponent’s (pure) strategy k is strategy k − 1, which
results in a payoff of k2 (if the player undercutting is player 1) or k−1
2 + 2 (player
2), with every strategy of player 2 a best response to player 1 playing his strategy
1. However, the payoff to any strategy j > k − 1 is k2 − 1 (player 1) or k−1
2 +1
(player 2).
One can show that the most-dominant strategy for player 1 places weight
1
on
Never, and the weights decay by a factor of 12 except that the weight on
2
strategy 1 is 2199 . This strategy is ε-dominant for player 1 for ε ≥ δ1 = 1 − 2199 .
The most-dominant strategy for player 2 places no weight on his dominated
strategy, Never, and places 12 on 198 and the weights again decay by a factor of
1
2 , except that the weight on strategy 2 is equal to that placed on strategy 4.
This strategy is ε-dominant for player 2 for ε ≥ δ2 = 1 − 2198 .
A player can at most lose 1 by playing a pure strategy which grabs later
than a best response. Consequently, in normal form centipede, strategies 197
and Never are (properly) ε-dominant for player 1 for values of ε ≥ 1. Similarly,
strategies 196, 198, and Never are ε-dominant for player 2 for ε ≥ 1, with the
first two of these being properly ε-dominant.
2
2.1
Robust Equilibria
ε-equilibria
Definition 3. A strategy profile s∗ forms an ε-equilibrium [26] if for all players i and for all si ∈ Si ,
ui (si , s∗−i ) − ui (s∗ ) ≤ ε.
A strategy profile s∗ forms an ε-equilibrium if no player may gain more than
ε by a unilateral deviation from the ε-equilibrium. Equivalently, we can replace
si ∈ Si by ai ∈ Ai in Definition 3. We denote the set of ε-equilibria by E ⊆ S.
4
Fudenberg and Levine [12] study players’ losses in experimental games. They
summarize their observations, saying “if the play in an experiment converges,
the limit should be one of the ε-self-confirming equilibria of the game. The
crude analysis in this paper suggests that the associated ε’s are typically small
compared with the stakes of the game.” For simultaneous-move games, an εself-confirming equilibrium is just an ε-equilibrium. This supports the selection
among ε-equilibria of a game, with ε small compared to the stakes of the game,
rather than the Nash equilibria, which often fail to be selected in experimental
studies.
Clearly, any strategy profile of ε-dominant strategies forms an ε-equilibrium
with the same value of ε. An ε-dominant ε-equilibrium is a strategy profile
of ε-dominant strategies. ε-dominant ε-equilibria exist if and only if ε ≥ δ(G),
the dominance defect of the game.
2.2
Robust Equilibria
A
B
a
N, N
0, 0
b
0, 0
1 1
N, N
c
0, −N 2
N, −N 2
Figure 2: A game without an ε-dominant row strategy for small ε (N >> 1).
The game of Figure 2 is an example where a more sophisticated notion
than ε-dominance is useful. (A, a) and (B, b) are both Nash equilibria. Just
as dominant strategies will not exist for most games, ε-dominant strategies will
not exist for most games for small ε. Strategy A is not ε-dominant for player 1
for ε < N , because of the possibility of player 2 playing his strategy c. However,
we question the risk of player 2 playing c since this would certainly cause player
2 to lose N 2 . This motivates the development of robust equilibria.
Let Ei = Ei (ε) ⊆ Si be the set of player i’s ε-equilibrium strategies, for
i ∈ I. Then, for si ∈ Ei , let
Rε (si ) =
sup
s−i ∈E−i
½
¾
sup {ui (s̃i , s−i )} − ui (si , s−i ) .
(1)
s̃i ∈Ei
Rε (si ) is the most regret that a player i can experience when playing si for
not having played another ε-equilibrium strategy when the other players play
ε-equilibrium strategies. We define R by replacing E with S or equivalently
with A in the above equation. R(si ) is the most regret that a player i can
experience when playing si for not having played any other strategy when the
other players choose any strategies. We clearly see that Rε (si ) ≤ R(si ).
An ε-equilibrium s eliminates another ε-equilibrium s̃ if for all i
ui (s̃i , t) − ui (si , t)
Rε (si )
≤ ε for all t ∈ E−i and
(2)
Rε (s̃i ) or si = s̃i .
(3)
<
5
So, an ε-equilibrium s eliminates another ε-equilibrium s̃ if for all i the
strategy si is never more than ε worse than the strategy s̃i against opponent
ε-equilibrium strategies and Rε (si ) < Rε (s̃i ) or si = s̃i .
If Rε is replaced by R and E is replaced by S or A then we say that s globally
eliminates ε-equilibrium s̃. So, an ε-equilibrium s globally eliminates another
ε-equilibrium s̃ if for all i the strategy si is never more than ε worse than the
strategy s̃i against any opponent strategies and R(si ) < R(s̃i ) or si = s̃i .
Definition 4. An ε-robust equilibrium is an ε-equilibrium which is not eliminated by any other ε-equilibrium.3
For the game in Figure 2, in an ε-equilibrium, player 2 clearly cannot place
more than Nε2 weight on his strategy c, since this results in a certain loss of ε.
It follows that although (A, a) cannot globally eliminate (B, b) for any value of
1
ε, it eliminates (B, b) for all N −1+
≤ ε < N 2.
1
N2
Definition 5. A globally ε-robust equilibrium is an ε-equilibrium which is
not globally eliminated by any other ε-equilibrium.4
If ε is clear, we can refer just to robust or globally robust equilibria.
Theorem 1. For every finite normal form game G and for all ε ≥ 0, there
exists an ε-robust equilibrium and a globally ε-robust equilibrium.
Proof.
Let ε ≥ 0. Since a Nash equilibrium is an ε-equilibrium for all ε ≥ 0, and
the set of Nash equilibria is nonempty, the set E of ε-equilibria is nonempty.
Moreover, E is closed and hence compact. Define the total regret of a strategy
profile as the sum of each player’s regret:
Rε (s1 , . . . , sn ) =
n
X
Rε (si )
i=1
and similarly for the total global regret R(s). Since Rε and R are continuous
on E, they attain minima, say at profiles s∗ and t∗ respectively. Then s∗ is an
ε-robust equilibrium and t∗ is a globally ε-robust equilibrium.
Definition 6. For a game G, let εG = inf{ε ≥ 0|Ai ⊆ Ei (ε) for all i ∈ I}.
εG is the smallest value of ε for which for all players every pure strategy is
an ε-equilibrium strategy. It can be shown that the infimum is attained by an
elementary compactness argument. If εG = 0, every ai is a Nash equilibrium
strategy. ε-robust equilibria and globally ε-robust equilibria coincide for ε ≥ εG .
The following examples in Figures 4-5 have εG = 0.
3 One could also replace the common ε, here and in other definitions, by a vector
(ε1 , . . . , εn ). This might be useful for example if there were large differences among agents’
payoff scales or attributes.
4 Using other regret functions, other variants of ε-robust equilibria can be defined. For
P
example, using R̄(si ) = a−i ∈A−i {maxai ∈Ai {ui (ai , a−i )} − ui (si , a−i )}.
6
Proposition 1. A strategy has smallest global maximum regret, R, for a player
i iff it is most-dominant for i.
Proof.
R(si ) =
=
max { max {ui (ai , a−i )} − ui (si , a−i )}
a−i ∈A−i ai ∈Ai
min{ε ≥ 0| max { max {ui (ai , a−i )}
a−i ∈A−i ai ∈Ai
−ui (si , a−i )} ≤ ε}
= min{ε ≥ 0|for all a−i ∈ A−i ,
(4)
(5)
(6)
max {ui (ai , a−i )} − ui (si , a−i ) ≤ ε}
ai ∈Ai
=
min{ε ≥ 0|for all a ∈ A, ui (ai , a−i )
−ui (si , a−i ) ≤ ε}
(7)
=
min{ε ≥ 0|si is ε-dominant}.
(8)
Therefore, a strategy for player i has smallest R iff it is ε-dominant for the
smallest value of ε for which there exists a strategy which is ε-dominant for
player i (δi ), i.e., if it is most-dominant.
An ε-robust equilibrium composed of most-dominant strategies is called a
most-dominant equilibrium. For ε ≥ δ, all globally ε-robust equilibria must
be most-dominant equilibria. If ε ≥ εG , and therefore Rε = R, for ε ≥ δ, all εrobust equilibria must be most-dominant equilibria. For example, in centipede,
with ε ≥ δ = 1 − 2199 , the unique ε-robust equilibrium is the pair of mostdominant strategies described in Section 1.1.
For games with two players, a most-dominant strategy, which has the minimum R, is easily computable using the standard linear programming methods for solving zero-sum games. Let (pj,k ) be player i’s payoff matrix for G.
Then let (p∗j,k ) = (maxl∈Ai {pl,k } − pj,k ) if i is originally the column player and
(p∗j,k ) = (maxl∈Ai {pl,k } − pj,k )T if i is originally the row player. Let G∗ = G∗ (i)
be the zero-sum game which has row player payoffs (p∗j,k ).
Proposition 2. A strategy has smallest global maximum regret, R, for a player
i iff it is a column player equilibrium strategy of the associated zero-sum game
G∗ .
7
Proof.
R(si ) =
=
=
=
=
max { max {ui (ai , a−i )} − ui (si , a−i )}
X
max{max pj,k −
pj,k si (j)}
a−i ∈A−i ai ∈Ai
j
k
(9)
(10)
j
X
max{ (max pl,k − pj,k )si (j)}
(11)
X
max{ (p∗j,k )si (j)}}
(12)
X
max{ (p∗j,k )si (j)s−i (k)}
(13)
k
j
k
l
j
s−i
j,k
Therefore,




X

,
(p∗j,k )si (j)s−i (k)
min R(si ) = min max
si  s−i 
si

j,k
which specifies column’s equilibrium payoff in the zero-sum game.
Here si (j) is the weight placed by player i’s strategy si on his jth pure
action. This is the method used to compute the most-dominant strategy weights
for the centipede games. Specifically,
the most-dominant strategy weights are
P
found by setting the regrets j (p∗j,k )si (j) for all k equal to one another (except
for the regret of the one dominated strategy, of course). This gives the weights
described earlier, which form the unique interior vertex of the feasible set for
the LP problem in the positive orthant. Any other vertex is on the boundary
and so places 0 weight on some strategy. However, all undominated strategies
are unique best responses to at least one of the other player’s strategies, and
so the regret experienced when the other player plays that strategy must be
greater. Hence, the maximum regret at a boundary vertex is greater.
2.3
Quasidominance
A
B
a
N, N
0, 0
b
0, 0
1
N,N
c
0, −N 2
N, −N 2
Figure 3: A game with a preferable Nash equilibrium for player 1 (N >> 1).
In the game of Figure 3, (B, b) will be a robust equilibrium in addition to
(A, a), for ε < N2 . Furthermore, A is not ε-dominant for small values of ε
because of the possibility of player 2 playing his strategy c. However, A seems
the most intuitive recommendation for player 1.
8
Definition 7. A strategy s∗i for player i is called ε-quasidominant if for all
si ∈ Ei , and for all s−i ∈ E−i , ui (si , s−i ) − ui (s∗i , s−i ) ≤ ε.
A strategy s∗i is ε-quasidominant for player i if it is never more than ε
worse than any other ε-equilibrium strategy si when the other players select
any of their ε-equilibrium strategies. Equivalently, a strategy s∗i for player i is
ε-quasidominant iff Rε (s∗i ) ≤ ε. If (si , s−i ) and (s̃i , s̃−i ) are ε-robust equilibria
and si is ε-quasidominant while s̃i is not ε-quasidominant, then Rε (si ) ≤ ε and
Rε (s̃i ) > ε. Thus, the conditions for elimination expressed in Equations (2) and
(3) are satisfied for player i.
Therefore, when a player has an ε-quasidominant strategy, quasi-dominance
can act as a means to select among multiple robust equilibrium strategies for
the player when there is not a unique robust equilibrium.
1
≤ ε ≤ N 2.
In the game of Figure 3, A is ε-quasidominant for N −1+
1
N2
2.4
R and Risk-Dominance
A
B
A
X, Y
X − x, V − v
B
U − u, Y − y
U, V
Figure 4: A 2 × 2 game with two pure strict Nash equilibria.
Consider a general 2 × 2 game with two pure strict Nash equilibria, (A, A)
and (B, B), as in Figure 4. The game has εG = 0, therefore Rε = R. Suppose
that (A, A) is an ε-dominant Nash equilibrium and (B, B) is not. Then, R(A) ≤
ε < R(B) for both players. However, note that the deviation loss for a player
at (A, A) is R(B) and the deviation loss for a player at (B, B) is R(A). So, if
R(A) < R(B) for both players, then x > u and y > v. However, this says that
the deviation losses at (A, A) are greater than the deviation losses at (B, B).
So, if there exists an ε for which (A, A) is ε-dominant but (B, B) is not, (A, A)
risk-dominates (B, B).
Kandori, Mailath and Rob [16] developed their Nash equilibrium selection,
long run equilibrium, in the spirit of Foster and Young’s [13] stochastically stable
equilibrium. They show that for symmetric 2 × 2 coordination games the long
run equilibrium will agree with the risk-dominant equilibrium. Thus, for 2 × 2
coordination games, if there is a unique ε-dominant Nash equilibrium for some
ε, then this is the risk-dominant equilibrium, as shown above, and so it is the
long-run equilibrium.
Ellison [10] expanded on Kandori, Mailath and Rob’s result to show that
when a p-dominant Nash equilibrium exists, with p = 12 , it is the long run
equilibrium. This is an extension of Kandori, Mailath and Rob’s result since pdominance, with p = 12 , is a refinement of risk-dominance. Here, an equilibrium
(si , s−i ) is p-dominant, as developed by Morris, Rob, and Shin [23], if si is a
strict best response to any strategy that places at least weight p on s−i . In
9
particular, a 1-dominant equilibrium is just a strict Nash equilibrium, and a
0-dominant equilibrium is a pair of strictly dominant strategies. If a strategy is
p-dominant for p = p∗ , then it is p-dominant for all p ≥ p∗ . This means that
any p-dominant equilibrium is 1-dominant, i.e. a strict Nash equilibrium. Thus,
a mixed equilibrium cannot be p-dominant.
U
D
L
1, 1
0, 1.9
R
1.9, 0
2, 2
Figure 5: Abreu & Matsushima’s game.
The game of Figure 5, due to Abreu & Matsushima [1], is an example of
a 2 × 2 game with two strict Nash equilibria, where there exists an ε (for
example, .1) such that one of the pure Nash equilibria, (U, L) is ε-dominant
while the other, (B, D), is not ε-dominant. As stated, this means that (U, L)
is risk-dominant and is the long-run equilibrium. It is also p-dominant for
1
. The robust equilibrium for this game is slightly more conservative. For
p ≥ 11
1
, the pair of most-dominant strategies form the unique robust equiε ≥ δ = 11
10
1
10
1
librium, ( 11
U, 11
D; 11
L, 11
R). ε-dominance and robust equilibria differ from
these refinements more strikingly in some of the 2 × 2 examples given below.
For games like the traveler’s dilemma, with its unsatisfactory unique Nash
equilibrium, refinements of Nash equilibria, such as risk-dominance, p-dominance,
and long run equilibrium, cannot help. In contrast, ε-dominance and robust
equilibria differ from all refinements in traveler’s dilemma, centipede, and related games.
3
Examples
We will look at several games from the literature, including parametrized versions of some common one-shot games. For the parametrized games, we are
interested in the cases where N becomes large.
3.1
Repeated Prisoner’s Dilemma
In Games and Decisions [19], Luce and Raiffa present a repeated game whose
stage game is a version of the prisoner’s dilemma. If the game is played once,
each player should play his dominant strategy. However, if the game is repeated
100 times and the players’ payoffs are the sum of their payoffs from the individual rounds, the situation is not so simple. Backwards induction dictates
that the Nash equilibrium outcome has both players defect in all periods. The
authors comment that this “is not ‘reasonable’ in the sense that we predict that
most intelligent people would not play accordingly.” They describe a restricted
strategy variant where each player must choose the earliest round at which they
10
will defect if their opponent has not yet defected, and they will defect for every round after that round or after the first time their opponent defects. This
forms a 100 × 100 normal form game. The authors state that they would employ a strategy cooperating until sometime in the nineties. A most-dominant
strategy places zero weight on all even strategies, {2, 4, ..., 100}, and it places
8(9n−1 )
on strategies 1 + 2n for 0 < n < 50 and 9149 on strategy 1. The pair of
949
most-dominant strategies form a robust equilibrium for ε ≥ δ = 1 − 9149 . Pure
strategies 98-100 are ε-dominant for ε ≥ 1, with 98 and 99 properly ε-dominant.
α1
α2
β1
5, 5
6, −4
β2
−4, 6
−3, −3
Figure 6: Luce and Raiffa’s prisoner’s dilemma stage game.
3.2
Traveler’s Dilemma
In traveler’s dilemma, two players must pick a (whole) dollar amount between
2 and 100. As payoffs, the players will receive the smaller of the two numbers
except for a reward or penalty of 2. The player who chose the smaller number
receives an extra reward of 2 and the player who chose the larger number will
be charged a penalty of 2. If they choose the same amount, there are no rewards
or penalties and they each receive the amount chosen.
The strategy 100 is dominated by 99 and so clearly should never be chosen.
However, if your opponent will not play 100, we see that you should not choose
99. Continuing in this fashion, we find that both players should choose 2. This
is the unique rationalizable equilibrium and hence the unique Nash equilibrium
outcome. However, it does not seem to be a reasonable equilibrium. Playing
2 can achieve a payoff of at most 4. Playing 95 can only receive less if your
opponent chooses 2-5, and it receives at most 3 less. It will, however, receive
more if your opponent chooses any value 7-100, and substantially more in most
cases.
In traveler’s dilemma, 96-100 are ε-dominant for ε ≥ 3, with 96-99 properly so. The unique ε-robust equilibrium for ε ≥ δ (δ for the game is slightly
less than 3) is the pair of most-dominant strategies for players 1 and 2, the
weights for which are depicted in Figure 7. The most weight is placed on the
highest undominated strategy and lower strategy weights (i.e. below 98) decay
exponentially.
Goeree and Holt [15] study a variant of the traveler’s dilemma game, with
strategies between 180 and 300 and a reward/penalty R. They conduct treatments both with a low R of 5 and a high R of 180. As intuition would suggest,
when the penalty of being undercut is the high R, and so the only ε-equilibrium
and thus robust equilibrium is the unique Nash equilibrium of (180, 180) for
ε < 180, the participants mostly play near the Nash equilibrium (an average of
11
Figure 7: Traveler’s dilemma most-dominant weights.
201). However, when the penalty for being undercut is small and we have nonNash ε-equilibria and robust equilibria for ε ≥ 9, the participants play mostly
high strategies (an average of 280). The most-dominant strategies are as in
Figure 8, with strategies relabeled by adding 100.
Figure 8: Traveler’s dilemma variant most-dominant weights with R=5.
Capra et al. [8] also conducted a traveler’s dilemma experiment which similarly has players conforming to Nash play for high reward/penalty R when there
is nothing ε-dominant and the only robust equilibrium is the Nash equilibrium
for small ε. The strategies are between 80 and 200 and for a low R of 5 the
participants play “at the opposite end of the range of feasible decisions” from
the Nash equilibrium. This is consistent with the set of ε-dominant strategies.
For ε ≥ 4, the pure strategies 195-200 are ε-dominant. As ε grows we have
progressive inclusion of a larger tail of ε-dominant high-end strategies. The
most-dominant strategies are as in Figure 8.
Becker, Carter and Naeve [5] provide an especially interesting study of the
traveler’s dilemma (as before, with strategies 2-100 and R = 2) since their
participants were members of the Game Theory Society. It is likely that the
participants were capable of performing the backward induction and thus finding
the Nash equilibrium. The only pure strategies submitted by multiple entrants
were 94-100 and the Nash equilibrium strategy 2. The authors found that the
best response to the average strategy played by the entrants was 97. Again, the
most-dominant weights for this version of traveler’s dilemma are as in Figure
12
7. Strategies 96-100 are ε-dominant for ε ≥ 3. Entrants were asked to submit
their beliefs about the distribution of strategies that would be played by all
entrants. Notable is that while 36% of entrants played a best response to their
beliefs, 79% played a strategy that was within 1 of a best response to their
beliefs. Also, the authors point out the somewhat perplexing fact that “It is
however a robust feature of all experiments on the traveler’s dilemma, that with
comparable size of the reward/penalty, the cooperative strategy [i.e. 100] is
used quite often.” This may be because 100 remains ε-dominant despite being
dominated and hence neither being properly ε-dominant nor part of a robust
equilibrium.
3.3
Imperfect Price Competition
Capra et al. [9] studied an imperfect price competition game where participants
chose a whole dollar price in the range [60, 160]. The player i with the lower
price pi received pi , while the player with the higher price received αpi where
α was .8 or .2. If p1 = p2 , then each player receives 1+α
2 p1 . When α = .2, the
participants’ actions converged to 70-90, near the Nash equilibrium prediction
of 60. Nothing is ε-dominant for ε < 126.4 and a most-dominant strategy
places over 50% of the weight on 60. However, when α = .8, for ε ≥ 30.8,
strategies 129-159 are ε-dominant and as ε increases again the tail of higher εdominant strategies expands. Also, a most-dominant strategy places over 50%
of the weight on strategies 150-159, and over 80% of the weight on strategies
117-159. This is consistent with the results of the experiment which found that
participants played strategies of 120 and higher.
3.4
Stag Hunt
In stag hunt, each player has the option of attempting to coordinate on the
larger payoff (Stag) or choosing the safe strategy (Hare). For stag hunt with
a large stag payoff, x = N where N >> 1, Stag is an ε-dominant strategy
for ε ≥ 1. If we use ε-dominance to select among the Nash equilibria, then
R(Stag) = 1 < N − 1 = R( Hare), thus Stag is ε-dominant much earlier than
Hare and we select (Stag,Stag). As noted, this implies that the (Stag,Stag) equilibrium risk-dominates the (Hare,Hare) equilibrium and is the long run equilibrium. (Stag,Stag) is p-dominant for p ≥ N1 . The most-dominant strategy is
( NN−1 S, N1 H). Because of the relation between most-dominant strategies and
robust equilibria, we know that the pair of most-dominant strategies for the
players will be the unique robust equilibrium for ε ≥ NN−1 = δ, when it in fact
will eliminate any other ε-equilibrium. Since the game has εG = 0, this is also
the unique globally robust equilibrium.
Battalio, Samuelson, and Van Huyck [4] conducted an experiment using
three generalized stag hunt games which differ in structure from our stag hunt
games described here. In each game the robust equilibrium has each player
select 80% Hare for ε ≥ δ, which depends on the version of the game. Initially,
players selected on average 37% Hare, but, after a learning procedure, the players
13
selected on average 76% Hare, rather close to the robust equilibrium. The
authors were mostly interested in the differences found in play and in learning
convergence among the three versions of the game. They found more players
selected the Pareto dominant equilibrium in versions that had a larger degree
of Pareto dominance but had a smaller penalty for not playing to the Pareto
dominant equilibrium when one’s opponent does so; these versions have a smaller
‘optimization premium.’
Stag
Hare
Stag
x, x
1, 0
Hare
0, 1
1, 1
Figure 9: Stag hunt.
In stag hunt with a small Stag payoff, x = 1 + N1 with N >> 1, there is little
reason for a player to pursue Stag when they can guarantee themselves almost
as high of a payoff by choosing Hare. Hare is an ε-dominant strategy for all
ε ≥ N1 . R(Hare) = N1 < 1 = R(Stag), so ε−dominance selects the (Hare,Hare)
equilibrium and it is risk-dominant and the long run equilibrium. (Hare,Hare)
is p-dominant for p ≥ N1+1 . The most-dominant strategy, which will minimize
1
N
a player’s maximum possible regret, is ( 1+N
S, 1+N
H). Again, since εG = 0,
because the most-dominant strategy has the smallest Rε = R, the pair of mostdominant strategies will be the unique robust equilibrium and globally robust
1
equilibrium for ε ≥ 1+N
= δ.
3.5
Battle of the Sexes
In the battle of the sexes coordination game, each pure Nash equilibrium is preferred by one of the players and mismatching makes everyone unhappy. Choosing one’s preferred event is ε-dominant for ε ≥ N1 . The pair of most-dominant
2
2
strategies, ( NN2 +1 F, N 21+1 B; N 21+1 F, NN2 +1 B) is the unique robust equilibrium
for ε ≥ N N
2 +1 = δ. This is the mixed Nash equilibrium which is not selected by
risk-dominance or p-dominance.
Football
Ballet
Football
N, N1
0, 0
Ballet
0, 0
1
N,N
Figure 10: Battle of the sexes.
3.6
Chicken
Like battle of the sexes, chicken has two pure Nash equilibria, each favored by
a player. However, in chicken, mismatching in one direction is far worse than
14
the other. Swerve is ε-dominant for ε ≥ N1 . The most-dominant strategy is
1
N2
( 1+N
2 S, 1+N 2 D). The pair of most-dominant strategies is the unique robust
equilibrium for ε ≥ N N
2 +1 = δ. As with battle of the sexes, this is the mixed
Nash equilibrium which is not selected by risk-dominance or p-dominance.
Swerve
Drive
Swerve
0, 0
1
N,0
Drive
0, N1
−N, −N
Figure 11: Chicken.
3.7
Samuelson’s Game
Samuelson [29] describes the game of Figure 12, in which player 2 has a dominant
strategy L. However, stating that many people suggest that player 1 nevertheless should not necessarily play T , his best response to L, Samuelson says further
that “Enriching the theory to account for this type of behaviour[playing B] is an
important priority for future research.” B is an ε-dominant strategy for player
1 for ε ≥ 1. So, (B, L) is an ε-dominant ε-equilibrium for ε ≥ 1. For ε ≥ .99,
1
99
( 100
T + 100
B, L) is the unique robust equilibrium. Since (T, L) is the unique
Nash equilibrium, any refinement of Nash must select (T, L).
T
B
L
100, 1
99, 1
R
0, 0
99, 0
Figure 12: Samuelson’s game.
3.8
Kreps’ Nonstory
Kreps’[18] “nonstory” base game in Figure 13 is a simple coordination game.
However, in addition to choosing s1 or s2 and t1 or t2, the players also choose a
positive integer. Then, if player 1 chooses s2 or player 2 chooses t2, the payoffs
are as in the 2×2 payoff matrix and the integers chosen are irrelevant. However,
if the players choose (s1,t1), then the player who chooses the greatest integer
receives a bonus of 1 in addition to their 100 payoff while the player with the
smaller integer receives 100. If the players happen to choose the same integer
then they each are penalized by 1 and receive 99. Without this last condition,
every s1 or t1 strategy would be dominated. This is because (s1, N ) cannot
perform better than (s1, N + 1), unless player 2’s strategy is (t1, N + 1) and
we apply the penalty. As a result of the greatest integer addition to the game,
any Nash equilibrium must have (s2, t2) as the outcome. However, (s2, t2) does
not seem to be a reasonable outcome to expect. Kreps himself states that he
15
believes the game has no “solution,” but that he believes it does have what he
calls a “partial solution,” which is that the players will choose (s1, t1). Even if
the penalty for matching integers were not applied and every s1 or t1 strategy
was dominated, (s2, t2) is still not a convincing solution. The ε-dominant εequilibria for ε ≥ 1 are the (s1, t1) pairs with any integers.
s1
s2
t1
100, 100
0, 0
t2
0, 0
1, 1
Figure 13: Kreps’ “nonstory” base game.
4
Conclusion
We have introduced here the solution concept of ε-robust equilibrium for games
in normal form. This solution concept captures the idea that we should hedge,
if it costs no more than some sufficiently small ε, against unforeseen strategies
of others, at least if those other strategies occur in some ε-equilibrium. When
ε = 0, we recover the Nash equilibria, but for positive ε we often get well
specified robust equilibria which are sharply at odds with Nash but in striking
harmony with experiment and intuition.
For example in traveler’s dilemma, which is a good simple representative
of games with a unique rationalizable (hence unique Nash) equilibrium found
via a long backward induction, that equilibrium is (2, 2). For ε ≥ 1, Radner’s
ε-equilibria include the entire diagonal ∆ = {(x, x)|2 ≤ x ≤ 100} ⊂ [2, 100]2 .
In contrast, the unique robust equilibrium for all ε ≥ 3 is the almost-dominant
equilibrium pictured in Figure 7, which has the bulk of its weight concentrated
in the upper nineties, in close agreement with the Game Theory Society experiment.
The solution concepts presented here emphasize the necessity for our strategic choices to have some degree of robustness in the face of our uncertainty
about others’ actions. We expect that these simple solution concepts will find
wider applicability.
References
[1] Abreu, D., and Matsushima, H. (1992) Econometrica 60 1439-1442. A response to Glazer and Rosenthal.
[2] Aumann, R. J. (1995) Games and Economic Behavior 8 6-19. Backward
induction and common knowledge of rationality.
[3] Basu, K. (1994) The American Economic Review 84 391-395. The traveler’s
dilemma: Paradoxes of rationality in game theory.
16
[4] Battalio, R., Samuelson, L., & Van Huyck, J. (2001) Econometrica 69 749764. Optimization incentives and coordination failure in laboratory stag
hunt.
[5] Becker, T., Carter, M., & Naeve, J. (2005) Discussion Paper 252, Institut for Economics, University of Hohenheim. Experts playing the traveler’s
dilemma.
[6] Camerer, C. (2003) Behavioral Game Theory (Princeton University Press,
Princeton, NJ).
[7] Camerer, C. (1997) The Journal of Economic Perspectives 11 167-188.
Progress in behavioral game theory.
[8] Capra, C.M, Goeree, J.K., Gomez, R., & Holt, C.A. (1999) The American
Economic Review 89 678-690. Anomolous behavior in a traveler’s dilemma?
[9] Capra, C.M, Goeree, J.K., Gomez, R., & Holt, C.A. (2002) International
Economic Review 43 613-636. Learning and noisy equilibrium behavior in
an experimental study of imperfect price competition.
[10] Ellison, G. (2000) The Review of Economic Studies 67 17-45. Basins of
attraction, long-run stochastic stability, and the speed of step-by-step evolution.
[11] Fudenberg, D., & Maskin, E. (1986) Econometrica 54 533-554. The folk
theorem in repeated games with discounting or with incomplete information.
[12] Fudenberg, D., & Levine, D. (1997) The Quarterly Journal of Economics
112 507-536. Measuring players’ losses in experimental games.
[13] Foster, D., & Young, P. (1990) Theoretical Population Biology 38 219-232.
Stochastic evolutionary dynamic games.
[14] Goeree, J.K. & Holt, C.A. (1999) Proceedings of the National Academy of
Science 96 10564-10567. Stochastic game theory: For playing games, not
just for doing theory.
[15] Goeree, J.K., & Holt, C.A. (2001) The American Economic Review 91
1402-1422. Ten little treasures of game theory and ten intuitive contradictions.
[16] Kandori, M., Mailath, G., & Rob, R. (1993) Econometrica 61 29-56. Learning, Mutation, and long run equilibria in games.
[17] Kreps, D., Milgrom, P., Roberts, J., & Wilson, R. (1982) Journal of Economic Theory 27 245-252. Rational cooperation in the finitely repeated prisoners’ dilemma.
[18] Kreps, D. (1990) A Course in Microeconomic Theory (Princeton University
Press, Princeton, NJ).
17
[19] Luce, R.D. & Raiffa, H. (1957) Games and Decisions (John Wiley & Sons,
Inc., New York, NY).
[20] Mailath, G., Postlewaite, A., & Samuelson, L. (2005) Games and Economic
Behavior 53 126-140. Contemporaneous perfect epsilon-equilibria.
[21] McKelvey, R.D., & Palfrey, T.R. (1992) Econometrica 60 803-836. An experimental study of the centipede game.
[22] McKelvey, R.D., & Palfrey, T.R. (1995) Games and Economic Behavior 10
6-38. Quantal response equilibrium for normal form games.
[23] Morris, S., Rob, R., & Shin, H. (1995) Econometrica 63 145-157. pdominance and belief potential.
[24] Neyman, A. (1997) in Cooperation: Game-Theoretic Approaches, NATO
ASI Series F, Vol. 155, S. Hart and A. Mas-Colell (eds.), (Springer-Verlag,
Berlin), 233-255. Cooperation, repetition and automata.
[25] Nash, J. (1950) Proc. Natl. Acad. Sci. USA 36 48-49. Equilibrium points
in n-person games.
[26] Radner, R. (1980) Journal of Economic Theory 22, 136-154. Collusive behavior in noncooperative epsilon-equilibria of oligopolies with long but finite
lives.
[27] Rosenthal, R. (1981) Journal of Economic Theory 25, 92-100. Games of
perfect information, predatory pricing, and the chain store paradox.
[28] Rubinstein, A. (2006) Econometrica 74 865-883. Dilemmas of an economic
theorist.
[29] Samuelson, L. (1992), in Recent Developments in Game theory, (Edward
Elgar Publishing Limited), 1-41.
[30] Savage, L. J. (1972) The Foundations of Statistics (Dover Publications,
New York, NY), second edition.
[31] Simon, H. (1955) Quarterly Journal of Economics 69 99-118. A behavioral
model of rational choice.
[32] Skyrms, B. (2004) The Stag Hunt and the Evolution of Social Structure
(Cambridge University Press, Cambridge, UK).
18