Self-play gains need outside opponents before production claims
I do not treat a self-play gain as production evidence until opponents from outside the training lineage can break it. Self-play can generate a curriculum, sharpen strategies, and expose specification bugs. It can also create a private world in which the agent becomes excellent at defeating the weaknesses its own descendants keep presenting. Improvement inside that world is real training progress, but it is not yet external validity.
The simplest failure is latest-only collapse. If the policy trains only against its newest copy, the shortest route to reward is exploiting that copy’s current quirk. Both sides then move together, leaving old weaknesses untested. A historical opponent pool improves the curriculum by sampling previous checkpoints, and league methods add exploiters and diverse roles. I would use them. But a family archive is still a family archive: shared code, reward, environment, initialization choices, and blind spots can produce diversity that looks wider than it is.
This is why I want a matrix, not a headline win rate. Rows are candidate policies; columns include historical checkpoints, independently trained runs, alternative algorithms, scripted baselines, specialist exploiters, and representative external agents. For cooperative tasks, I also want unseen partners. Joint Policy Correlation measures the gap between performance with a co-trained partner and independently trained partners of similar skill. A policy can look nearly perfect with the convention it invented and fail when another competent agent uses a different convention.
The production analogue is distribution shift with agency. Customers, attackers, counterparties, and upstream services do not remain on the curriculum’s trajectory. They discover strategies the training population never represented, change objectives, violate conventions, and exploit boundary conditions. The release case therefore needs performance slices by opponent family, not an average folded across them. It also needs uncertainty and failure containment: what the system does when it cannot classify the opponent matters as much as whether it wins.
I would make the external suite an evaluation contract. Hold back opponents before training, rotate a sealed subset less frequently, and preserve old production failures as regression opponents. Track exploitability where the game permits it, cross-play performance where coordination matters, and outcome variance across random seeds. Report resource budgets too. A policy that wins by spending ten times the inference or simulation budget has changed the operating contract, not simply improved skill.
LLM debate needs the same skepticism. Multiple agents can converge on the same inherited misconception or reward one another’s rhetorical style. As I argue in multi-agent debate has a coordination budget, more interaction is not free reliability. Outside judges, adversarial evidence, independent model families, and preserved disagreement are the equivalent of external opponents. Consensus among copies is not corroboration.
I concede that in a fully specified, symmetric, perfect-information game, sufficiently strong self-play plus formal exploitability analysis can provide unusually powerful evidence. Even there, implementation errors, compute limits, and the distance between the formal game and the deployment environment remain. The concession narrows the external test; it does not make a training curve a production certificate.
The claim I trust is deliberately modest: this policy improved against this declared population under this budget and retained performance against these independent opponents. Anything broader needs broader opposition. That is also why production benchmarks must reward containment: when an outside opponent finds a gap, the system should fail inside a bounded envelope rather than convert a surprising move into an unbounded effect.