The evolution of cooperation
Two people can each do better by betraying the other, and both do worse if both do. Play that once and betrayal is the answer. Play it again and again, with no known last round, and the answer changes. Set the payoffs, choose the players, and run it.
The prisoner's dilemma is the smallest situation in which doing the sensible thing makes everybody worse off. Two players choose at the same time to cooperate or to defect. Whatever the other does, each is better off defecting. So both defect, and both end up with less than if both had cooperated. There is no cleverness that gets you out of it, because the trap is in the payoffs.
Play it once and there is nothing more to say. Play it repeatedly against the same opponent, with neither of you knowing which round is the last, and something enters that was not there before: what you do now changes what they do later. That is the whole subject of this essay, and it is worth building the machine before making any claims about it.
There is a playable version of this same argument, if you would rather find it out than read it. The Long Game puts you on one side of a fence in a small village and brings a neighbour to it every season. You decide whether the harvest goes over or stays, and so do they. It runs the same payoffs, the same strategies and the same evolutionary tournament as this page, without naming any of them until you have already felt them. This page is the formal version, with the numbers exposed and every source dated. It is one of six games on this site, each one a different situation written down in as few numbers as it takes.
01The payoffs
Four numbers define the game. You get the reward when you both cooperate, the temptation when you defect against a cooperator, the punishment when you both defect, and the sucker's payoff when you cooperate against a defector. The dilemma exists only when two conditions hold at once: the temptation beats the reward beats the punishment beats the sucker, and the reward is worth more than alternating between exploiting and being exploited. Break either and it stops being a dilemma.
| they cooperate | they defect | |
|---|---|---|
| you cooperate | 3 to you 3 to them |
0 to you 5 to them |
| you defect | 5 to you 0 to them |
1 to you 1 to them |
02The players
Each strategy is a rule for choosing the next move from what has happened so far. None of them can see the opponent's code and none of them knows how long the match will last. Tick the ones you want in the tournament. The last entry is yours: four dropdowns that compose a rule out of what happened in the previous round or two. Setting them to cooperate, cooperate, defect, defect gives tit for tat.
Your rule
03The tournament
Every strategy plays every strategy, including a copy of itself, five times over. Each match runs until the coin says stop, which it does with probability one minus the continuation chance after every round. The five lengths are drawn once and then used for every pairing, so that no strategy is helped by having been handed a longer game than its rival. Scores are reported as the average payoff per round, so that a long match and a short one can be compared. Nothing is scored against a fixed number of rounds, because a known last round would let everybody unravel backwards from it, and no player in the tournament is ever told which round is the last.
The head to head table is where the argument actually lives. Green marks a cell where the row strategy scored more than the column strategy did against it, red marks the reverse. Read along the row for tit for tat with the noise control at zero and you will find no green in it at all: against a defector it loses by exactly one round of being taken in, and against anything cooperative it draws. It cannot beat anybody. Whether that is enough to come first depends entirely on who else is in the pool, which is the qualification section six returns to.
04The shadow of the future
The continuation chance is the control that matters most, and it is the one the rest of this essay turns on. If the game continues with a chance of nine in ten, a match lasts ten rounds on average, so a defection today costs you nine rounds of retaliation. If it continues with a chance of one in two, it lasts two rounds, and there is almost no future to punish you in.
expected length of a match = 1 / (1 minus the continuation chance) At a continuation chance of 0.99 that is a hundred rounds. At 0.50 it is two.
Pull the continuation chance down and run the tournament again. Somewhere on the way down, always defect stops being punished for its behaviour and starts being rewarded for it, and the ranking inverts. Nothing about the strategies changed. What changed is how much of the future there was to lose. That is the actual claim of this essay: cooperation is not a moral quality that some rules have and others lack, it is a behaviour that pays when the shadow of the future is long enough, and does not when it is not.
The other control in that rack is worth a minute of your time as well. It flips a small fraction of moves in transmission, so that a player who meant to cooperate is seen to defect. Turn it up to eight per cent and run the tournament again. Strict retaliation now punishes accidents: two copies of tit for tat fall into a run of mutual recrimination that neither of them chose and neither of them can end, and a grudger condemns for ever a defection that was never intended. Watch what this does to the population run in particular, where the rule that waits for a second defection before believing the first tends to end up holding the ground that tit for tat held when the channel was clean. Forgiveness is not softness in a noisy world. It is error correction.
05What survives
A tournament ranks strategies once. It does not say what happens to a population in which the successful ones become more common, which changes who everybody is playing against, which changes who is successful. That second question is the interesting one, and it is answered by running the population forward.
Every strategy starts with an equal share. In each generation, each one is scored against the mix currently present, and its share in the next generation moves in proportion to how it did. A strategy that does well against what is common becomes common, and then everybody has to do well against it. This is the rule Axelrod used for what he called the ecological tournament.
Watch what happens to the exploiters. Always defect does well early, while there are cooperators to take advantage of, and then eats its own food supply: once the naive are gone it is left playing against other defectors and scoring the punishment payoff for ever. A strategy that is retaliatory but not vindictive does badly against nobody and ends up surrounded by copies of itself, all cooperating.
06Axelrod, accurately
Robert Axelrod ran two computer tournaments in 1980. The first had fourteen entries plus a program that moved at random, making fifteen, and each pair played two hundred rounds against every other. The payoffs were five for the temptation, three for the reward, one for the punishment and nothing for the sucker, which is what the controls on this page open with. The winner was the shortest program submitted: tit for tat, four lines of it, entered by the psychologist Anatol Rapoport.
The second tournament was run after the results of the first were published, so every entrant knew what had won and had the chance to build something that would beat it. Sixty two people entered, sixty three programs with random included, and this time the length of each match was decided by a coin rather than fixed, with a chance of 0.00346 of ending after any given round, so that nobody could reason backwards from a known ending. Rapoport submitted tit for tat again. It won again.
Axelrod then ran what he called an ecological tournament: a thousand generations in which the number of copies of each program in the next generation was set in proportion to the payoff it had won in the last. Tit for tat rose. That is the run you can reproduce in figure five, at a smaller scale and with a pool of your choosing.
The honest qualification, which matters: a tournament result is a statement about a pool, not about the world. Rapoport, Seale and Colman re examined the tournaments in 2015 and argued that the victory of tit for tat depended heavily on the presence of particular weak programs which fed it points without taking any from it, so that changing the pool changes the winner. You can produce that effect on this page in about ten seconds by removing always cooperate from the pool. That is not a reason to throw the result away. It is a reason to state it as what it is: in these payoffs, with a long enough future, and against a pool containing a mixture of the naive and the hostile, a rule that is nice, retaliatory, forgiving and predictable does well.
07Where the model stops
Assumptions and limits
- The pool decidesEverything on this page is conditional on which strategies you ticked. There is no such thing as the best strategy in the abstract, only the best against a particular mixture. Any ranking you produce here is a fact about your pool.
- Small samplesEach pair plays five matches, and match length is random, so the ranking moves a little between draws. Use the redraw button a few times before believing any close result. The seed is stored in the link so that a shared configuration reproduces exactly.
- No population structureIn the evolutionary run everybody meets everybody with equal probability. Real populations have neighbours, clusters and reputations, and cooperation is very much easier to sustain when a cooperator is more likely to meet another cooperator than chance would suggest.
- Two players, two movesThe dilemma here is between two parties who each have exactly two options. Most of the situations people reach for this model to explain, from climate agreements to price wars, have many parties, continuous choices and outside enforcement.
- Nothing here is a claim about peopleThese are programs playing an abstract game for points. Human beings in laboratory prisoner's dilemmas behave in ways these rules do not capture, and no result on this page is evidence about what any person will do.
- Replicator dynamics, not biologyThe population update is proportional fitness, following the rule Axelrod used. It has no mutation, no drift and no finite population effects, all of which change outcomes in real evolutionary models.
08Sources
- The first tournament: fourteen submitted programs plus a random one, two hundred rounds per pairing, payoffs of five, three, one and zero for temptation, reward, punishment and sucker. Tit for tat, submitted by Anatol Rapoport, finished first. The tournament was run five times to smooth out the random element. Anatol Rapoport, Darryl A. Seale and Andrew M. Colman, Is Tit for Tat the Answer? On the Conclusions Drawn from Axelrod's Tournaments, PLOS ONE 10(7), 2015. journals.plos.org/plosone/article?id=10.1371/journal.pone.0134128 Corroborated by Leigh Tesfatsion's course notes on the Axelrod tournaments, faculty.sites.iastate.edu. Checked 2026 08 21.
- The second tournament: sixty two entries, sixty three programs including random, entrants knowing the results of the first. Match length was probabilistic rather than fixed, with a chance of 0.00346 of ending after any given round. Tit for tat, entered again by Rapoport, won again. Rapoport, Seale and Colman, as above, for the entry count and the probabilistic match length. The ending probability of 0.00346 and Rapoport's authorship are given in the Theory, Evolution and Games Group summary of the tournaments, egtheory.wordpress.com/2015/03/02/ipd. Checked 2026 08 21.
- The ecological tournament: one thousand generations, in which the number of copies of a strategy at the start of a generation was set equal to the total payoff that strategy won in the previous one. Tit for tat rose. Leigh Tesfatsion, Notes on Axelrod's Iterated Prisoner's Dilemma Tournaments, Iowa State University. faculty.sites.iastate.edu/tesfatsi/archive/econ308/tesfatsion/axeltmts.pdf Checked 2026 08 21.
- The qualification: the authors argue that the success of tit for tat in Axelrod's tournaments owed much to the presence of a small number of weak programs, so that conclusions drawn from the tournaments are conditional on the pool of entrants rather than general. Rapoport, Seale and Colman, Is Tit for Tat the Answer? On the Conclusions Drawn from Axelrod's Tournaments, PLOS ONE 10(7), 2015. journals.plos.org/plosone/article?id=10.1371/journal.pone.0134128 Checked 2026 08 21.