For a voluntary AI safety agreement to be self-enforcing, complying has to be a best response: a participant should not expect to do better by breaking it, given how everyone else would react. I think this is the right test for commitments that become costly near a release.
Consider two labs deciding whether to finish an agreed safety evaluation before deploying. Assume both would be better off if both completed the evaluation than if both rushed to release. The difficulty is that either lab can gain an advantage by releasing while the other waits. And if one lab releases early, the other would rather follow than fall behind. For this example, assume these competitive gains outweigh the extra safety risk from skipping the evaluation.
Under those assumptions, skipping is each lab's best response to either choice by the other. Both skipping is the unique Nash equilibrium, despite both preferring mutual restraint -- a pretty simple prisoner's dilemma. A promise to wait leaves the incentive to deviate. Armstrong, Bostrom, and Shulman's AI race model develops the related point that concern about disaster can coexist with competitive pressure to cut precautions.
Repeated interaction could change the calculation. A lab that breaks an agreement might lose future access to shared research, infrastructure, or commercial partnerships. The basic repeated-game condition is that the gain from deviating now must be outweighed by the expected loss of future cooperation. Basically detection has to be likely enough, and that future has to matter enough. Merely expecting to interact again does not guarantee cooperation.
There is a second incentive problem also: someone must carry out the threatened response. Suppose a compute provider promises to suspend a lab after a verified breach. Once the breach happens, suspension could mean losing a major customer. Would the provider still choose it? The game-theoretic requirement is sequential credibility: enforcement must make sense when the decision arrives, accounting for its future consequences. A severe penalty that the provider would waive may deter less than a narrower restriction it has reason to enforce ???
Monitoring therefore needs to identify an actionable breach. Discovering a dangerous capability and concealing a required evaluation are different events. Punishing the discovery itself could discourage honest reporting; sanctioning concealment requires evidence that concealment occurred. I would want the agreement to specify what participants must disclose, how disputed findings are reviewed, and what follows a verified violation.
The complication I find most important is that AI progress could change the value of future cooperation. An agreement might work while every lab needs outside compute and shared expertise. If a leading lab becomes less dependent on those resources, exclusion becomes a weaker deterrent. Its gain from unilateral action might also grow. The same agreement could then stop being self-enforcing without anyone changing their stated values.
I would stress-test a proposed agreement against exactly that scenario: a leading participant with better alternatives, a profitable opportunity to defect, and partners for whom enforcement is expensive. A convincing proposal should explain why compliance remains worthwhile for the lab and enforcement remains worthwhile for its partners. Those are two separate incentive constraints. A commitment is much more credible when it addresses both. Probably some game-theoretic model out there to do this.