Solution (source code)

= Solution

Under the standard <two-condition criterion for evolutionary stability>, the strict claim fails. <Tit for tat> and <always cooperate> both cooperate in every round when playing either each other or their own type under the specified initial memories. Hence
$$
e({\rm TFT},{\rm TFT})=e({\rm ALLC},{\rm TFT})
=e({\rm TFT},{\rm ALLC})=e({\rm ALLC},{\rm ALLC})=R.
$$
Both stability comparisons tie. The <neutrality between Tit for tat and unconditional cooperation> proves that \b[Tit for tat is not an evolutionarily stable strategy in the stated strategy space], and cannot strictly invade every strategy. The final differently spelled name is interpreted as the same $(1,0)$ rule, since no second strategy is defined.

There is a useful qualified invasion result. For a resident $M=(p,q)$ with $|p-q|<1$, put $z=q/(1-p+q)$. The <long-run payoff of reactive strategies> gives
$$
e({\rm TFT},M)=e(M,{\rm TFT})=e(M,M)
=g(z)=Rz^2+(S+T)z(1-z)+P(1-z)^2.
$$
Thus, if the fraction of <Tit for tat> players is $\varepsilon$, their <payoff> advantage is
$$
\boxed{w_{\rm TFT}-w_M=\varepsilon[R-g(z)].}
$$
Moreover $R-g(z)=(1-z)[R-P+(R+P-S-T)z]$. Under the additional cooperative-efficiency condition $2R\ge S+T$, this is positive whenever $z<1$. <Tit for tat invasion of a reactive resident> then occurs from every positive frequency, although its invasion exponent at frequency zero vanishes. The <replicator equation> is $\dot\varepsilon=[R-g(z)]\varepsilon^2(1-\varepsilon)$, so the initial increase is slow and frequency-dependent. Residents with $p=1$ remain cooperative on all reached histories and tie instead. If $S+T>2R$, some almost-cooperative residents have $g(z)>R$, and even this qualified invasion conclusion fails.

For the exceptional opposite-response resident $(0,1)$, its self-payoff is $h=(R+P)/2$ while its <payoff> in either order against <Tit for tat> is $g=(R+S+T+P)/4$. The advantage is $(1-\varepsilon)(g-h)+\varepsilon(R-g)$. Under $2R\ge S+T$, invasion at arbitrarily small frequency requires $S+T\ge R+P$; otherwise it needs
$$
\boxed{\varepsilon>\frac{R+P-S-T}{2(2R-S-T)}.}
$$
Finally, taking an infinite undiscounted average before the rare-mutant <limit> is essential. In an $L$-round match against <always defect>, the first round contributes $S$ to <Tit for tat> and $T$ to the defector, followed by $L-1$ rounds of <payoff> $P$. The <finite-horizon invasion threshold of Tit for tat against unconditional defection> is
$$
\boxed{\varepsilon>\frac{P-S}{L(R-P)-(T+S-2P)}},
$$
provided the denominator is positive and the threshold is below one. Thus a finite horizon generally prevents invasion from an arbitrarily rare introduction. The source's absent <payoff> table prevents choosing which extra <payoff> inequalities were intended, but the symbolic cases and the strict-ESS counterexample do not depend on guessing it.