* Your assessment is very important for improving the work of artificial intelligence, which forms the content of this project
Download Robust equilibria and ε-dominance
Survey
Document related concepts
Game mechanics wikipedia , lookup
Turns, rounds and time-keeping systems in games wikipedia , lookup
The Evolution of Cooperation wikipedia , lookup
Artificial intelligence in video games wikipedia , lookup
Prisoner's dilemma wikipedia , lookup
Nash equilibrium wikipedia , lookup
Transcript
Robust equilibria and ε-dominance∗ William Geller†and Rachel Hemphill‡ April 11, 2014 Abstract We propose a resolution of the backward induction paradox and some other anomalies in game theory based on refinements of Radner’s ε-equilibria. We avoid the usual very strong assumption of common knowledge of rationality with complete information, and posit instead some inescapable uncertainty about others’ actions. The central idea is to require a solution for a noncooperative game to exhibit some degree of robustness. When ε = 0, our ε-robust equilibria reduce to Nash equilibria, but for positive ε our solutions in such games as centipede and traveler’s dilemma contrast sharply with the Nash predictions and fit very well with experiment and intuition. Game theory has had great difficulty dealing convincingly with an important family of games whose equilibria may be found by a long backward induction. Exemplars of this problematic but central family are the finitely repeated prisoner’s dilemma [19] and Rosenthal’s centipede game [27], [18]; Basu’s traveler’s dilemma [3] and imperfect price competition games [9] are related examples. On the one hand, the theory of Nash equilibrium clearly predicts one sort of behavior in these games, which we could call a race to the bottom; on the other hand, this behavior seems unreasonable, even for rational, self-interested agents, an intuition reflected in a large number of experiments, including those with subjects well aware of the Nash equilibrium [5]. Many (overlapping) attempts have been made to address this seeming paradox: for example, the introduction of a small chance of irrationality of a carefully chosen type [17], the introduction of limits on the reasoning or computational power [24] of agents, or a cost for the use of complex strategies, the weakening of maximizing behavior to near-maximizing behavior [26], [31], and the replacement of maximizing behavior with a sort of stochastic or smoothed maximization [22]. There are difficulties with each of these approaches. For example, the first is open to the objection that alternative choices of the seeding irrationality lead to a wide variety of outcomes [11], the ∗ We thank Bob Anderson for his helpful comments. of Mathematical Sciences, Indiana University-Purdue University Indianapolis. [email protected]; corresponding author. ‡ Department of Mathematical Sciences, Indiana University-Purdue University Indianapolis. [email protected]. † Department 1 second seems not to apply well to games with simple structures like centipede or traveler’s dilemma, especially with sophisticated agents,1 the third typically lacks specificity in its predictions, and the fourth depends on the choice both of a type of smoothing function and a diffusion parameter. Our approach is closest in spirit to the third and fourth lines of attack, though we believe it possesses significant advantages. While it is much simpler than the stochastic framework of quantal response equilibria, it involves an often radical refinement of the set of ε-equilibria despite the introduction of no further parameters. This allows for quite sharp predictions in the case of centipede and traveler’s dilemma which are in striking agreement with experiments and intuition. We were motivated originally by the much studied paradox of the finitely repeated prisoner’s dilemma, which is more than half a century old. Considering a prisoner’s dilemma repeated 100 times, where every Nash equilibrium leads both players to play tough on every round, Luce and Raiffa [19] state that they would not play to a Nash equilibrium. In fact, if strategies were restricted to those which play nice before some round k, 1 ≤ k ≤ 100, as long as the opponent also plays nice, and after a tough play by the opponent or the arrival of round k play tough until the end, they write that they would probably play a strategy k, where k “is some number in the nineties.” We are able to vindicate their intuition for the restricted strategy game; the unrestricted game is still beyond our grasp. In the same way, we resolve the paradox of the centipede game and the traveler’s dilemma. We also address some limiting instances of stag hunt, the prototypical assurance game whose history dates to Rousseau [32], [4], as well as some other relevant examples from the literature. We introduce here a small circle of closely related solution concepts for games in strategic form centered on the notions of ε-dominance and ε-robustness. Our aim is to expand the normative and positive scope of noncooperative game theory with the simplest possible tools. A Nash equilibrium requires zero regret from each agent if he has correctly anticipated others’ strategies, but allows massive regret if another’s strategy is unforeseen. This makes Nash equilibria precariously dependent on very strong assumptions. We will instead require small regret from each agent if he predicts correctly, but also impose some robustness on his strategy, i.e. seek to limit his regret if his prediction is incorrect. This turns out to be surprisingly fruitful. 0.1 Centipede Centipede is an extensive-form game which alternates between decision nodes for two players. At each odd-numbered stage 2l + 1 of the game, player 1 may grab, which ends the game and results in each player receiving a payoff of l + 1, or pass, which causes the game to continue to player 2. Similarly, on an evennumbered stage 2l of the game, player 2 may pass, causing the game to continue 1 See the traveler’s dilemma played by the Game Theory Society members in Section 4 below. 2 to player 1’s next decision node, or can grab, which results in player 1 receiving l − 1 and player 2 receiving l + 2. The game ends at round 198, where player 2 makes the choice between passing, which results in each player receiving 100, or grabbing, which results in player 1 receiving 98 and player 2 receiving 101; see Figure 1. A rational player 2 should grab at his final opportunity. Furthermore, if he knows player 2 will grab at round 198, a rational player 1 has no incentive to pass at round 197, so he will surely grab then. Continuing in this fashion, we arrive at the conclusion that player 1 must grab at round 1. This is the unique rationalizable equilibrium outcome, in the sense of Bernheim for the normal form, since it is obtained by iterated elimination of dominated strategies and is thus the unique Nash equilibrium. However, it seems counter to intuition about how rational, self-interested players should play. 1 A 2 A 1 . . . D D D (1,1) (0,3) (2,2) 1 A D (98,98) 2 A 1 D (97,100) A D (99,99) 2 (100,100) D (98,101) Figure 1: Centipede. 1 ε-dominance Consider a finite game G in normal, i.e. strategic, form. Let I = {1, 2, ..., n} be the set of players. Let Ai be the finite set of player i’s pure strategies, for i ∈ I. Let Si = ∆(Ai ), the set of player i’s mixed strategies, i.e. the Q (|A | − 1)-dimensional simplex of probability vectors. Let A = A , A−i = i i i∈I Q Q Q j∈I,j6=i Sj . Then, a (mixed) strategy i∈I Si , and S−i = j∈I,j6=i Aj , S = profile is an element s = (s1 , s2 , ..., sn ) ∈ S. Let ui : S → R be the payoff function for player i. We introduce here several definitions. Definition 1. A strategy s∗i for player i is called ε-dominant if for all si ∈ Si and for all s−i ∈ S−i , ui (si , s−i ) − ui (s∗i , s−i ) ≤ ε. That is, a strategy is ε-dominant if it never engenders more than ε regret.2 In particular, a 0-dominant strategy is never regretted. Equivalently, a strategy s∗i is ε-dominant if, for all ai ∈ Ai and for all a−i ∈ A−i , ui (ai , a−i ) − ui (s∗i , a−i ) ≤ ε. An undominated ε-dominant strategy is called properly ε-dominant. Definition 2. Let δi = δi (G) = inf{ε ≥ 0|there exists a strategy si which is ε-dominant.} If si is ε-dominant for ε = δi , then si is called most-dominant. We call δi the (dominance) defect for i. 2 Abreu and Matsushima [1] consider the ε-domination of one strategy by another. 3 The dominance defect measures how far a player is from having a dominant strategy, and a most-dominant strategy is one that is as near to dominant as available for the player. As noted in Proposition 1 below, δi is just the minimax regret considered in decision theory by Savage [30], and so a most-dominant strategy is just a strategy minimizing maximum regret. The (dominance) defect δ = δ(G) of the game G is the maximum of the defects for the players, δi . 1.1 Centipede Analysis In centipede, at every decision node, there exists a single possible history: both players must have passed at each of their prior decision nodes. A strategy for a player is then simply the decision to pass or grab at each of his decision nodes. However, since the game ends once a player grabs, a strategy is completely described by the earliest stage at which the player would grab. So, A1 = {1, 3, 5, ..., 197,Never=199} and A2 = {2, 4, ..., 198,Never=200}, where strategy k indicates the earliest node at which the player would grab. The best response to an opponent’s (pure) strategy k is strategy k − 1, which results in a payoff of k2 (if the player undercutting is player 1) or k−1 2 + 2 (player 2), with every strategy of player 2 a best response to player 1 playing his strategy 1. However, the payoff to any strategy j > k − 1 is k2 − 1 (player 1) or k−1 2 +1 (player 2). One can show that the most-dominant strategy for player 1 places weight 1 on Never, and the weights decay by a factor of 12 except that the weight on 2 strategy 1 is 2199 . This strategy is ε-dominant for player 1 for ε ≥ δ1 = 1 − 2199 . The most-dominant strategy for player 2 places no weight on his dominated strategy, Never, and places 12 on 198 and the weights again decay by a factor of 1 2 , except that the weight on strategy 2 is equal to that placed on strategy 4. This strategy is ε-dominant for player 2 for ε ≥ δ2 = 1 − 2198 . A player can at most lose 1 by playing a pure strategy which grabs later than a best response. Consequently, in normal form centipede, strategies 197 and Never are (properly) ε-dominant for player 1 for values of ε ≥ 1. Similarly, strategies 196, 198, and Never are ε-dominant for player 2 for ε ≥ 1, with the first two of these being properly ε-dominant. 2 2.1 Robust Equilibria ε-equilibria Definition 3. A strategy profile s∗ forms an ε-equilibrium [26] if for all players i and for all si ∈ Si , ui (si , s∗−i ) − ui (s∗ ) ≤ ε. A strategy profile s∗ forms an ε-equilibrium if no player may gain more than ε by a unilateral deviation from the ε-equilibrium. Equivalently, we can replace si ∈ Si by ai ∈ Ai in Definition 3. We denote the set of ε-equilibria by E ⊆ S. 4 Fudenberg and Levine [12] study players’ losses in experimental games. They summarize their observations, saying “if the play in an experiment converges, the limit should be one of the ε-self-confirming equilibria of the game. The crude analysis in this paper suggests that the associated ε’s are typically small compared with the stakes of the game.” For simultaneous-move games, an εself-confirming equilibrium is just an ε-equilibrium. This supports the selection among ε-equilibria of a game, with ε small compared to the stakes of the game, rather than the Nash equilibria, which often fail to be selected in experimental studies. Clearly, any strategy profile of ε-dominant strategies forms an ε-equilibrium with the same value of ε. An ε-dominant ε-equilibrium is a strategy profile of ε-dominant strategies. ε-dominant ε-equilibria exist if and only if ε ≥ δ(G), the dominance defect of the game. 2.2 Robust Equilibria A B a N, N 0, 0 b 0, 0 1 1 N, N c 0, −N 2 N, −N 2 Figure 2: A game without an ε-dominant row strategy for small ε (N >> 1). The game of Figure 2 is an example where a more sophisticated notion than ε-dominance is useful. (A, a) and (B, b) are both Nash equilibria. Just as dominant strategies will not exist for most games, ε-dominant strategies will not exist for most games for small ε. Strategy A is not ε-dominant for player 1 for ε < N , because of the possibility of player 2 playing his strategy c. However, we question the risk of player 2 playing c since this would certainly cause player 2 to lose N 2 . This motivates the development of robust equilibria. Let Ei = Ei (ε) ⊆ Si be the set of player i’s ε-equilibrium strategies, for i ∈ I. Then, for si ∈ Ei , let Rε (si ) = sup s−i ∈E−i ½ ¾ sup {ui (s̃i , s−i )} − ui (si , s−i ) . (1) s̃i ∈Ei Rε (si ) is the most regret that a player i can experience when playing si for not having played another ε-equilibrium strategy when the other players play ε-equilibrium strategies. We define R by replacing E with S or equivalently with A in the above equation. R(si ) is the most regret that a player i can experience when playing si for not having played any other strategy when the other players choose any strategies. We clearly see that Rε (si ) ≤ R(si ). An ε-equilibrium s eliminates another ε-equilibrium s̃ if for all i ui (s̃i , t) − ui (si , t) Rε (si ) ≤ ε for all t ∈ E−i and (2) Rε (s̃i ) or si = s̃i . (3) < 5 So, an ε-equilibrium s eliminates another ε-equilibrium s̃ if for all i the strategy si is never more than ε worse than the strategy s̃i against opponent ε-equilibrium strategies and Rε (si ) < Rε (s̃i ) or si = s̃i . If Rε is replaced by R and E is replaced by S or A then we say that s globally eliminates ε-equilibrium s̃. So, an ε-equilibrium s globally eliminates another ε-equilibrium s̃ if for all i the strategy si is never more than ε worse than the strategy s̃i against any opponent strategies and R(si ) < R(s̃i ) or si = s̃i . Definition 4. An ε-robust equilibrium is an ε-equilibrium which is not eliminated by any other ε-equilibrium.3 For the game in Figure 2, in an ε-equilibrium, player 2 clearly cannot place more than Nε2 weight on his strategy c, since this results in a certain loss of ε. It follows that although (A, a) cannot globally eliminate (B, b) for any value of 1 ε, it eliminates (B, b) for all N −1+ ≤ ε < N 2. 1 N2 Definition 5. A globally ε-robust equilibrium is an ε-equilibrium which is not globally eliminated by any other ε-equilibrium.4 If ε is clear, we can refer just to robust or globally robust equilibria. Theorem 1. For every finite normal form game G and for all ε ≥ 0, there exists an ε-robust equilibrium and a globally ε-robust equilibrium. Proof. Let ε ≥ 0. Since a Nash equilibrium is an ε-equilibrium for all ε ≥ 0, and the set of Nash equilibria is nonempty, the set E of ε-equilibria is nonempty. Moreover, E is closed and hence compact. Define the total regret of a strategy profile as the sum of each player’s regret: Rε (s1 , . . . , sn ) = n X Rε (si ) i=1 and similarly for the total global regret R(s). Since Rε and R are continuous on E, they attain minima, say at profiles s∗ and t∗ respectively. Then s∗ is an ε-robust equilibrium and t∗ is a globally ε-robust equilibrium. Definition 6. For a game G, let εG = inf{ε ≥ 0|Ai ⊆ Ei (ε) for all i ∈ I}. εG is the smallest value of ε for which for all players every pure strategy is an ε-equilibrium strategy. It can be shown that the infimum is attained by an elementary compactness argument. If εG = 0, every ai is a Nash equilibrium strategy. ε-robust equilibria and globally ε-robust equilibria coincide for ε ≥ εG . The following examples in Figures 4-5 have εG = 0. 3 One could also replace the common ε, here and in other definitions, by a vector (ε1 , . . . , εn ). This might be useful for example if there were large differences among agents’ payoff scales or attributes. 4 Using other regret functions, other variants of ε-robust equilibria can be defined. For P example, using R̄(si ) = a−i ∈A−i {maxai ∈Ai {ui (ai , a−i )} − ui (si , a−i )}. 6 Proposition 1. A strategy has smallest global maximum regret, R, for a player i iff it is most-dominant for i. Proof. R(si ) = = max { max {ui (ai , a−i )} − ui (si , a−i )} a−i ∈A−i ai ∈Ai min{ε ≥ 0| max { max {ui (ai , a−i )} a−i ∈A−i ai ∈Ai −ui (si , a−i )} ≤ ε} = min{ε ≥ 0|for all a−i ∈ A−i , (4) (5) (6) max {ui (ai , a−i )} − ui (si , a−i ) ≤ ε} ai ∈Ai = min{ε ≥ 0|for all a ∈ A, ui (ai , a−i ) −ui (si , a−i ) ≤ ε} (7) = min{ε ≥ 0|si is ε-dominant}. (8) Therefore, a strategy for player i has smallest R iff it is ε-dominant for the smallest value of ε for which there exists a strategy which is ε-dominant for player i (δi ), i.e., if it is most-dominant. An ε-robust equilibrium composed of most-dominant strategies is called a most-dominant equilibrium. For ε ≥ δ, all globally ε-robust equilibria must be most-dominant equilibria. If ε ≥ εG , and therefore Rε = R, for ε ≥ δ, all εrobust equilibria must be most-dominant equilibria. For example, in centipede, with ε ≥ δ = 1 − 2199 , the unique ε-robust equilibrium is the pair of mostdominant strategies described in Section 1.1. For games with two players, a most-dominant strategy, which has the minimum R, is easily computable using the standard linear programming methods for solving zero-sum games. Let (pj,k ) be player i’s payoff matrix for G. Then let (p∗j,k ) = (maxl∈Ai {pl,k } − pj,k ) if i is originally the column player and (p∗j,k ) = (maxl∈Ai {pl,k } − pj,k )T if i is originally the row player. Let G∗ = G∗ (i) be the zero-sum game which has row player payoffs (p∗j,k ). Proposition 2. A strategy has smallest global maximum regret, R, for a player i iff it is a column player equilibrium strategy of the associated zero-sum game G∗ . 7 Proof. R(si ) = = = = = max { max {ui (ai , a−i )} − ui (si , a−i )} X max{max pj,k − pj,k si (j)} a−i ∈A−i ai ∈Ai j k (9) (10) j X max{ (max pl,k − pj,k )si (j)} (11) X max{ (p∗j,k )si (j)}} (12) X max{ (p∗j,k )si (j)s−i (k)} (13) k j k l j s−i j,k Therefore, X , (p∗j,k )si (j)s−i (k) min R(si ) = min max si s−i si j,k which specifies column’s equilibrium payoff in the zero-sum game. Here si (j) is the weight placed by player i’s strategy si on his jth pure action. This is the method used to compute the most-dominant strategy weights for the centipede games. Specifically, the most-dominant strategy weights are P found by setting the regrets j (p∗j,k )si (j) for all k equal to one another (except for the regret of the one dominated strategy, of course). This gives the weights described earlier, which form the unique interior vertex of the feasible set for the LP problem in the positive orthant. Any other vertex is on the boundary and so places 0 weight on some strategy. However, all undominated strategies are unique best responses to at least one of the other player’s strategies, and so the regret experienced when the other player plays that strategy must be greater. Hence, the maximum regret at a boundary vertex is greater. 2.3 Quasidominance A B a N, N 0, 0 b 0, 0 1 N,N c 0, −N 2 N, −N 2 Figure 3: A game with a preferable Nash equilibrium for player 1 (N >> 1). In the game of Figure 3, (B, b) will be a robust equilibrium in addition to (A, a), for ε < N2 . Furthermore, A is not ε-dominant for small values of ε because of the possibility of player 2 playing his strategy c. However, A seems the most intuitive recommendation for player 1. 8 Definition 7. A strategy s∗i for player i is called ε-quasidominant if for all si ∈ Ei , and for all s−i ∈ E−i , ui (si , s−i ) − ui (s∗i , s−i ) ≤ ε. A strategy s∗i is ε-quasidominant for player i if it is never more than ε worse than any other ε-equilibrium strategy si when the other players select any of their ε-equilibrium strategies. Equivalently, a strategy s∗i for player i is ε-quasidominant iff Rε (s∗i ) ≤ ε. If (si , s−i ) and (s̃i , s̃−i ) are ε-robust equilibria and si is ε-quasidominant while s̃i is not ε-quasidominant, then Rε (si ) ≤ ε and Rε (s̃i ) > ε. Thus, the conditions for elimination expressed in Equations (2) and (3) are satisfied for player i. Therefore, when a player has an ε-quasidominant strategy, quasi-dominance can act as a means to select among multiple robust equilibrium strategies for the player when there is not a unique robust equilibrium. 1 ≤ ε ≤ N 2. In the game of Figure 3, A is ε-quasidominant for N −1+ 1 N2 2.4 R and Risk-Dominance A B A X, Y X − x, V − v B U − u, Y − y U, V Figure 4: A 2 × 2 game with two pure strict Nash equilibria. Consider a general 2 × 2 game with two pure strict Nash equilibria, (A, A) and (B, B), as in Figure 4. The game has εG = 0, therefore Rε = R. Suppose that (A, A) is an ε-dominant Nash equilibrium and (B, B) is not. Then, R(A) ≤ ε < R(B) for both players. However, note that the deviation loss for a player at (A, A) is R(B) and the deviation loss for a player at (B, B) is R(A). So, if R(A) < R(B) for both players, then x > u and y > v. However, this says that the deviation losses at (A, A) are greater than the deviation losses at (B, B). So, if there exists an ε for which (A, A) is ε-dominant but (B, B) is not, (A, A) risk-dominates (B, B). Kandori, Mailath and Rob [16] developed their Nash equilibrium selection, long run equilibrium, in the spirit of Foster and Young’s [13] stochastically stable equilibrium. They show that for symmetric 2 × 2 coordination games the long run equilibrium will agree with the risk-dominant equilibrium. Thus, for 2 × 2 coordination games, if there is a unique ε-dominant Nash equilibrium for some ε, then this is the risk-dominant equilibrium, as shown above, and so it is the long-run equilibrium. Ellison [10] expanded on Kandori, Mailath and Rob’s result to show that when a p-dominant Nash equilibrium exists, with p = 12 , it is the long run equilibrium. This is an extension of Kandori, Mailath and Rob’s result since pdominance, with p = 12 , is a refinement of risk-dominance. Here, an equilibrium (si , s−i ) is p-dominant, as developed by Morris, Rob, and Shin [23], if si is a strict best response to any strategy that places at least weight p on s−i . In 9 particular, a 1-dominant equilibrium is just a strict Nash equilibrium, and a 0-dominant equilibrium is a pair of strictly dominant strategies. If a strategy is p-dominant for p = p∗ , then it is p-dominant for all p ≥ p∗ . This means that any p-dominant equilibrium is 1-dominant, i.e. a strict Nash equilibrium. Thus, a mixed equilibrium cannot be p-dominant. U D L 1, 1 0, 1.9 R 1.9, 0 2, 2 Figure 5: Abreu & Matsushima’s game. The game of Figure 5, due to Abreu & Matsushima [1], is an example of a 2 × 2 game with two strict Nash equilibria, where there exists an ε (for example, .1) such that one of the pure Nash equilibria, (U, L) is ε-dominant while the other, (B, D), is not ε-dominant. As stated, this means that (U, L) is risk-dominant and is the long-run equilibrium. It is also p-dominant for 1 . The robust equilibrium for this game is slightly more conservative. For p ≥ 11 1 , the pair of most-dominant strategies form the unique robust equiε ≥ δ = 11 10 1 10 1 librium, ( 11 U, 11 D; 11 L, 11 R). ε-dominance and robust equilibria differ from these refinements more strikingly in some of the 2 × 2 examples given below. For games like the traveler’s dilemma, with its unsatisfactory unique Nash equilibrium, refinements of Nash equilibria, such as risk-dominance, p-dominance, and long run equilibrium, cannot help. In contrast, ε-dominance and robust equilibria differ from all refinements in traveler’s dilemma, centipede, and related games. 3 Examples We will look at several games from the literature, including parametrized versions of some common one-shot games. For the parametrized games, we are interested in the cases where N becomes large. 3.1 Repeated Prisoner’s Dilemma In Games and Decisions [19], Luce and Raiffa present a repeated game whose stage game is a version of the prisoner’s dilemma. If the game is played once, each player should play his dominant strategy. However, if the game is repeated 100 times and the players’ payoffs are the sum of their payoffs from the individual rounds, the situation is not so simple. Backwards induction dictates that the Nash equilibrium outcome has both players defect in all periods. The authors comment that this “is not ‘reasonable’ in the sense that we predict that most intelligent people would not play accordingly.” They describe a restricted strategy variant where each player must choose the earliest round at which they 10 will defect if their opponent has not yet defected, and they will defect for every round after that round or after the first time their opponent defects. This forms a 100 × 100 normal form game. The authors state that they would employ a strategy cooperating until sometime in the nineties. A most-dominant strategy places zero weight on all even strategies, {2, 4, ..., 100}, and it places 8(9n−1 ) on strategies 1 + 2n for 0 < n < 50 and 9149 on strategy 1. The pair of 949 most-dominant strategies form a robust equilibrium for ε ≥ δ = 1 − 9149 . Pure strategies 98-100 are ε-dominant for ε ≥ 1, with 98 and 99 properly ε-dominant. α1 α2 β1 5, 5 6, −4 β2 −4, 6 −3, −3 Figure 6: Luce and Raiffa’s prisoner’s dilemma stage game. 3.2 Traveler’s Dilemma In traveler’s dilemma, two players must pick a (whole) dollar amount between 2 and 100. As payoffs, the players will receive the smaller of the two numbers except for a reward or penalty of 2. The player who chose the smaller number receives an extra reward of 2 and the player who chose the larger number will be charged a penalty of 2. If they choose the same amount, there are no rewards or penalties and they each receive the amount chosen. The strategy 100 is dominated by 99 and so clearly should never be chosen. However, if your opponent will not play 100, we see that you should not choose 99. Continuing in this fashion, we find that both players should choose 2. This is the unique rationalizable equilibrium and hence the unique Nash equilibrium outcome. However, it does not seem to be a reasonable equilibrium. Playing 2 can achieve a payoff of at most 4. Playing 95 can only receive less if your opponent chooses 2-5, and it receives at most 3 less. It will, however, receive more if your opponent chooses any value 7-100, and substantially more in most cases. In traveler’s dilemma, 96-100 are ε-dominant for ε ≥ 3, with 96-99 properly so. The unique ε-robust equilibrium for ε ≥ δ (δ for the game is slightly less than 3) is the pair of most-dominant strategies for players 1 and 2, the weights for which are depicted in Figure 7. The most weight is placed on the highest undominated strategy and lower strategy weights (i.e. below 98) decay exponentially. Goeree and Holt [15] study a variant of the traveler’s dilemma game, with strategies between 180 and 300 and a reward/penalty R. They conduct treatments both with a low R of 5 and a high R of 180. As intuition would suggest, when the penalty of being undercut is the high R, and so the only ε-equilibrium and thus robust equilibrium is the unique Nash equilibrium of (180, 180) for ε < 180, the participants mostly play near the Nash equilibrium (an average of 11 Figure 7: Traveler’s dilemma most-dominant weights. 201). However, when the penalty for being undercut is small and we have nonNash ε-equilibria and robust equilibria for ε ≥ 9, the participants play mostly high strategies (an average of 280). The most-dominant strategies are as in Figure 8, with strategies relabeled by adding 100. Figure 8: Traveler’s dilemma variant most-dominant weights with R=5. Capra et al. [8] also conducted a traveler’s dilemma experiment which similarly has players conforming to Nash play for high reward/penalty R when there is nothing ε-dominant and the only robust equilibrium is the Nash equilibrium for small ε. The strategies are between 80 and 200 and for a low R of 5 the participants play “at the opposite end of the range of feasible decisions” from the Nash equilibrium. This is consistent with the set of ε-dominant strategies. For ε ≥ 4, the pure strategies 195-200 are ε-dominant. As ε grows we have progressive inclusion of a larger tail of ε-dominant high-end strategies. The most-dominant strategies are as in Figure 8. Becker, Carter and Naeve [5] provide an especially interesting study of the traveler’s dilemma (as before, with strategies 2-100 and R = 2) since their participants were members of the Game Theory Society. It is likely that the participants were capable of performing the backward induction and thus finding the Nash equilibrium. The only pure strategies submitted by multiple entrants were 94-100 and the Nash equilibrium strategy 2. The authors found that the best response to the average strategy played by the entrants was 97. Again, the most-dominant weights for this version of traveler’s dilemma are as in Figure 12 7. Strategies 96-100 are ε-dominant for ε ≥ 3. Entrants were asked to submit their beliefs about the distribution of strategies that would be played by all entrants. Notable is that while 36% of entrants played a best response to their beliefs, 79% played a strategy that was within 1 of a best response to their beliefs. Also, the authors point out the somewhat perplexing fact that “It is however a robust feature of all experiments on the traveler’s dilemma, that with comparable size of the reward/penalty, the cooperative strategy [i.e. 100] is used quite often.” This may be because 100 remains ε-dominant despite being dominated and hence neither being properly ε-dominant nor part of a robust equilibrium. 3.3 Imperfect Price Competition Capra et al. [9] studied an imperfect price competition game where participants chose a whole dollar price in the range [60, 160]. The player i with the lower price pi received pi , while the player with the higher price received αpi where α was .8 or .2. If p1 = p2 , then each player receives 1+α 2 p1 . When α = .2, the participants’ actions converged to 70-90, near the Nash equilibrium prediction of 60. Nothing is ε-dominant for ε < 126.4 and a most-dominant strategy places over 50% of the weight on 60. However, when α = .8, for ε ≥ 30.8, strategies 129-159 are ε-dominant and as ε increases again the tail of higher εdominant strategies expands. Also, a most-dominant strategy places over 50% of the weight on strategies 150-159, and over 80% of the weight on strategies 117-159. This is consistent with the results of the experiment which found that participants played strategies of 120 and higher. 3.4 Stag Hunt In stag hunt, each player has the option of attempting to coordinate on the larger payoff (Stag) or choosing the safe strategy (Hare). For stag hunt with a large stag payoff, x = N where N >> 1, Stag is an ε-dominant strategy for ε ≥ 1. If we use ε-dominance to select among the Nash equilibria, then R(Stag) = 1 < N − 1 = R( Hare), thus Stag is ε-dominant much earlier than Hare and we select (Stag,Stag). As noted, this implies that the (Stag,Stag) equilibrium risk-dominates the (Hare,Hare) equilibrium and is the long run equilibrium. (Stag,Stag) is p-dominant for p ≥ N1 . The most-dominant strategy is ( NN−1 S, N1 H). Because of the relation between most-dominant strategies and robust equilibria, we know that the pair of most-dominant strategies for the players will be the unique robust equilibrium for ε ≥ NN−1 = δ, when it in fact will eliminate any other ε-equilibrium. Since the game has εG = 0, this is also the unique globally robust equilibrium. Battalio, Samuelson, and Van Huyck [4] conducted an experiment using three generalized stag hunt games which differ in structure from our stag hunt games described here. In each game the robust equilibrium has each player select 80% Hare for ε ≥ δ, which depends on the version of the game. Initially, players selected on average 37% Hare, but, after a learning procedure, the players 13 selected on average 76% Hare, rather close to the robust equilibrium. The authors were mostly interested in the differences found in play and in learning convergence among the three versions of the game. They found more players selected the Pareto dominant equilibrium in versions that had a larger degree of Pareto dominance but had a smaller penalty for not playing to the Pareto dominant equilibrium when one’s opponent does so; these versions have a smaller ‘optimization premium.’ Stag Hare Stag x, x 1, 0 Hare 0, 1 1, 1 Figure 9: Stag hunt. In stag hunt with a small Stag payoff, x = 1 + N1 with N >> 1, there is little reason for a player to pursue Stag when they can guarantee themselves almost as high of a payoff by choosing Hare. Hare is an ε-dominant strategy for all ε ≥ N1 . R(Hare) = N1 < 1 = R(Stag), so ε−dominance selects the (Hare,Hare) equilibrium and it is risk-dominant and the long run equilibrium. (Hare,Hare) is p-dominant for p ≥ N1+1 . The most-dominant strategy, which will minimize 1 N a player’s maximum possible regret, is ( 1+N S, 1+N H). Again, since εG = 0, because the most-dominant strategy has the smallest Rε = R, the pair of mostdominant strategies will be the unique robust equilibrium and globally robust 1 equilibrium for ε ≥ 1+N = δ. 3.5 Battle of the Sexes In the battle of the sexes coordination game, each pure Nash equilibrium is preferred by one of the players and mismatching makes everyone unhappy. Choosing one’s preferred event is ε-dominant for ε ≥ N1 . The pair of most-dominant 2 2 strategies, ( NN2 +1 F, N 21+1 B; N 21+1 F, NN2 +1 B) is the unique robust equilibrium for ε ≥ N N 2 +1 = δ. This is the mixed Nash equilibrium which is not selected by risk-dominance or p-dominance. Football Ballet Football N, N1 0, 0 Ballet 0, 0 1 N,N Figure 10: Battle of the sexes. 3.6 Chicken Like battle of the sexes, chicken has two pure Nash equilibria, each favored by a player. However, in chicken, mismatching in one direction is far worse than 14 the other. Swerve is ε-dominant for ε ≥ N1 . The most-dominant strategy is 1 N2 ( 1+N 2 S, 1+N 2 D). The pair of most-dominant strategies is the unique robust equilibrium for ε ≥ N N 2 +1 = δ. As with battle of the sexes, this is the mixed Nash equilibrium which is not selected by risk-dominance or p-dominance. Swerve Drive Swerve 0, 0 1 N,0 Drive 0, N1 −N, −N Figure 11: Chicken. 3.7 Samuelson’s Game Samuelson [29] describes the game of Figure 12, in which player 2 has a dominant strategy L. However, stating that many people suggest that player 1 nevertheless should not necessarily play T , his best response to L, Samuelson says further that “Enriching the theory to account for this type of behaviour[playing B] is an important priority for future research.” B is an ε-dominant strategy for player 1 for ε ≥ 1. So, (B, L) is an ε-dominant ε-equilibrium for ε ≥ 1. For ε ≥ .99, 1 99 ( 100 T + 100 B, L) is the unique robust equilibrium. Since (T, L) is the unique Nash equilibrium, any refinement of Nash must select (T, L). T B L 100, 1 99, 1 R 0, 0 99, 0 Figure 12: Samuelson’s game. 3.8 Kreps’ Nonstory Kreps’[18] “nonstory” base game in Figure 13 is a simple coordination game. However, in addition to choosing s1 or s2 and t1 or t2, the players also choose a positive integer. Then, if player 1 chooses s2 or player 2 chooses t2, the payoffs are as in the 2×2 payoff matrix and the integers chosen are irrelevant. However, if the players choose (s1,t1), then the player who chooses the greatest integer receives a bonus of 1 in addition to their 100 payoff while the player with the smaller integer receives 100. If the players happen to choose the same integer then they each are penalized by 1 and receive 99. Without this last condition, every s1 or t1 strategy would be dominated. This is because (s1, N ) cannot perform better than (s1, N + 1), unless player 2’s strategy is (t1, N + 1) and we apply the penalty. As a result of the greatest integer addition to the game, any Nash equilibrium must have (s2, t2) as the outcome. However, (s2, t2) does not seem to be a reasonable outcome to expect. Kreps himself states that he 15 believes the game has no “solution,” but that he believes it does have what he calls a “partial solution,” which is that the players will choose (s1, t1). Even if the penalty for matching integers were not applied and every s1 or t1 strategy was dominated, (s2, t2) is still not a convincing solution. The ε-dominant εequilibria for ε ≥ 1 are the (s1, t1) pairs with any integers. s1 s2 t1 100, 100 0, 0 t2 0, 0 1, 1 Figure 13: Kreps’ “nonstory” base game. 4 Conclusion We have introduced here the solution concept of ε-robust equilibrium for games in normal form. This solution concept captures the idea that we should hedge, if it costs no more than some sufficiently small ε, against unforeseen strategies of others, at least if those other strategies occur in some ε-equilibrium. When ε = 0, we recover the Nash equilibria, but for positive ε we often get well specified robust equilibria which are sharply at odds with Nash but in striking harmony with experiment and intuition. For example in traveler’s dilemma, which is a good simple representative of games with a unique rationalizable (hence unique Nash) equilibrium found via a long backward induction, that equilibrium is (2, 2). For ε ≥ 1, Radner’s ε-equilibria include the entire diagonal ∆ = {(x, x)|2 ≤ x ≤ 100} ⊂ [2, 100]2 . In contrast, the unique robust equilibrium for all ε ≥ 3 is the almost-dominant equilibrium pictured in Figure 7, which has the bulk of its weight concentrated in the upper nineties, in close agreement with the Game Theory Society experiment. The solution concepts presented here emphasize the necessity for our strategic choices to have some degree of robustness in the face of our uncertainty about others’ actions. We expect that these simple solution concepts will find wider applicability. References [1] Abreu, D., and Matsushima, H. (1992) Econometrica 60 1439-1442. A response to Glazer and Rosenthal. [2] Aumann, R. J. (1995) Games and Economic Behavior 8 6-19. Backward induction and common knowledge of rationality. [3] Basu, K. (1994) The American Economic Review 84 391-395. The traveler’s dilemma: Paradoxes of rationality in game theory. 16 [4] Battalio, R., Samuelson, L., & Van Huyck, J. (2001) Econometrica 69 749764. Optimization incentives and coordination failure in laboratory stag hunt. [5] Becker, T., Carter, M., & Naeve, J. (2005) Discussion Paper 252, Institut for Economics, University of Hohenheim. Experts playing the traveler’s dilemma. [6] Camerer, C. (2003) Behavioral Game Theory (Princeton University Press, Princeton, NJ). [7] Camerer, C. (1997) The Journal of Economic Perspectives 11 167-188. Progress in behavioral game theory. [8] Capra, C.M, Goeree, J.K., Gomez, R., & Holt, C.A. (1999) The American Economic Review 89 678-690. Anomolous behavior in a traveler’s dilemma? [9] Capra, C.M, Goeree, J.K., Gomez, R., & Holt, C.A. (2002) International Economic Review 43 613-636. Learning and noisy equilibrium behavior in an experimental study of imperfect price competition. [10] Ellison, G. (2000) The Review of Economic Studies 67 17-45. Basins of attraction, long-run stochastic stability, and the speed of step-by-step evolution. [11] Fudenberg, D., & Maskin, E. (1986) Econometrica 54 533-554. The folk theorem in repeated games with discounting or with incomplete information. [12] Fudenberg, D., & Levine, D. (1997) The Quarterly Journal of Economics 112 507-536. Measuring players’ losses in experimental games. [13] Foster, D., & Young, P. (1990) Theoretical Population Biology 38 219-232. Stochastic evolutionary dynamic games. [14] Goeree, J.K. & Holt, C.A. (1999) Proceedings of the National Academy of Science 96 10564-10567. Stochastic game theory: For playing games, not just for doing theory. [15] Goeree, J.K., & Holt, C.A. (2001) The American Economic Review 91 1402-1422. Ten little treasures of game theory and ten intuitive contradictions. [16] Kandori, M., Mailath, G., & Rob, R. (1993) Econometrica 61 29-56. Learning, Mutation, and long run equilibria in games. [17] Kreps, D., Milgrom, P., Roberts, J., & Wilson, R. (1982) Journal of Economic Theory 27 245-252. Rational cooperation in the finitely repeated prisoners’ dilemma. [18] Kreps, D. (1990) A Course in Microeconomic Theory (Princeton University Press, Princeton, NJ). 17 [19] Luce, R.D. & Raiffa, H. (1957) Games and Decisions (John Wiley & Sons, Inc., New York, NY). [20] Mailath, G., Postlewaite, A., & Samuelson, L. (2005) Games and Economic Behavior 53 126-140. Contemporaneous perfect epsilon-equilibria. [21] McKelvey, R.D., & Palfrey, T.R. (1992) Econometrica 60 803-836. An experimental study of the centipede game. [22] McKelvey, R.D., & Palfrey, T.R. (1995) Games and Economic Behavior 10 6-38. Quantal response equilibrium for normal form games. [23] Morris, S., Rob, R., & Shin, H. (1995) Econometrica 63 145-157. pdominance and belief potential. [24] Neyman, A. (1997) in Cooperation: Game-Theoretic Approaches, NATO ASI Series F, Vol. 155, S. Hart and A. Mas-Colell (eds.), (Springer-Verlag, Berlin), 233-255. Cooperation, repetition and automata. [25] Nash, J. (1950) Proc. Natl. Acad. Sci. USA 36 48-49. Equilibrium points in n-person games. [26] Radner, R. (1980) Journal of Economic Theory 22, 136-154. Collusive behavior in noncooperative epsilon-equilibria of oligopolies with long but finite lives. [27] Rosenthal, R. (1981) Journal of Economic Theory 25, 92-100. Games of perfect information, predatory pricing, and the chain store paradox. [28] Rubinstein, A. (2006) Econometrica 74 865-883. Dilemmas of an economic theorist. [29] Samuelson, L. (1992), in Recent Developments in Game theory, (Edward Elgar Publishing Limited), 1-41. [30] Savage, L. J. (1972) The Foundations of Statistics (Dover Publications, New York, NY), second edition. [31] Simon, H. (1955) Quarterly Journal of Economics 69 99-118. A behavioral model of rational choice. [32] Skyrms, B. (2004) The Stag Hunt and the Evolution of Social Structure (Cambridge University Press, Cambridge, UK). 18