How to Pace the Frontier
A Research Agenda
“Society at large may need the option to buy time to address emerging risks, develop security measures, and strengthen oversight. But each company—and country—is under intense competitive pressure not to unilaterally slow that acceleration. And today, the world lacks the technical and governance tools to deliberately pace frontier-wide progress.”
– 1,367 employees of frontier AI companies
Part of how we have navigated AI progress so far is by pacing — that is, deliberately moderating the pace of development, deployment, and diffusion. As frontier capabilities advance, choices around pacing will become increasingly high-stakes. Here we do not advocate pacing in any particular way and instead try to understand, in broad strokes, how pacing works. Specifically:
- Why might one want to pace progress, or avoid pacing, in general?
- What are the different parts of the AI ecosystem that can be paced?
- How does a pacing intervention start, progress, and end?
- What wider effects might pacing interventions have, for better or worse?
Each of these broad topics brings with it important questions, noted at the end of each section. There are dozens of papers arguing for or against different kinds of pacing: This is all the more reason to want dispassionate research on the general dynamics at play.
Introduction
We are facing increasingly challenging decisions about the pace of frontier artificial intelligence — that is, the rate at which frontier AI capabilities are developed, deployed, and diffused. How should companies balance growing the capabilities and autonomy of their models against the effectiveness of their oversight? How should regulators handle the spread of systems that can carry out, and at times autonomously initiate, advanced cyber attacks? How should governments navigate arms-race dynamics, in which each side fears that its own restraint will be exploited?
Already, choices around pacing have been difficult and costly — we have seen developers delay releases, roll back deployment, and even pause training in response to harms that they were not equipped to mitigate. In the coming years, these choices are likely to become even higher-stakes, as AI systems become more advanced and embedded in society, the regulatory environment of AI becomes more complex and less agile, and pacing takes on greater geostrategic significance.
Many of these decisions will need to be made quickly, without all the relevant information, and balancing competing interests. Indeed, so far most pacing has been unilateral, through self-pacing (e.g. delayed releases) and one actor pacing another (e.g. export controls), but in future pacing may require coordination between actors with different incentives, and potentially between adversaries. Early pacing attempts, whether successful or not, will substantially shape the precedents, institutional knowledge, and evidence that will guide future decisions.
It is therefore extremely important that principals in a position to make a pacing decision are provided with a clear picture of, and comprehensive information about, the options and tradeoffs available, not just a set of exemplar pacing plans. Given the dynamic and unpredictable nature of AI progress and its geopolitical context, specific pacing decisions are likely to differ substantially from anything we can write down today, and may rely on private information only available to a small set of actors. We therefore emphasize the provision of a framework for working through pacing decisions, that could be applied by decision makers and leverage such private information, now or at any time in the future.
There has been a great deal of research on specific topics relevant to pacing (e.g. model evaluations, capability forecasting, compute monitoring), and on proposals for specific interventions, but comparatively little on the overall question of how pacing works. That is what this piece attempts: to give a broad account of the whole topic of pacing AI.
This topic is understandably quite political — indeed, tens of millions of dollars have already been spent on advocacy on both sides of the debate. But that is all the more reason to want dispassionate research, and a shared understanding of the practical implications.
1.1 Structure of the piece
This piece is structured around a series of questions intended to mirror how one might develop or evaluate a potential intervention:
- Why pace? Section 2 examines motivations for and against pacing, for society in general and for different stakeholders.
- Pace what? Section 3 considers which of the activities that make up AI progress can be paced, and the challenges in picking appropriately.
- Pace how? Section 4 follows the lifecycle of an intervention from anticipation to exit, asking what it takes for each step to succeed.
- Then what? Section 5 considers the broader effects of interventions, including on broader AI R&D, the economy, and the distribution of power.
At the end of each section we give a list of open questions. We also use two stylised cases to demonstrate how our analysis applies to potential interventions throughout this document:
- a coordinated cap on the compute used in frontier training runs, intended to address loss-of-control concerns, particularly the possibility that AI systems substantially accelerate their own R&D process.
- an international redline concerning models that materially uplift users’ ability to develop biological weapons, as an example of pacing diffusion (rather than pacing capability development).
We return to each of these examples at the end of each section, and show how the framework developed in that section applies to each case.
A bibliography of work we admire, related to this agenda, can be found here.
Why Pace?
There are two broad perspectives from which arguments for and against pacing have been made: from the perspective of society as a whole — will the world be better off if AI progress is deliberately slowed down vs sped up? — and from the perspective of specific actors in the AI R&D ecosystem — why would the actors involved seek or agree to pace?
Looking at pacing from the global perspective, the question is something like: given the current state of AI progress, our historical track record of intervening in R&D, and our existing economic, political and institutional structures, do we think it is a good idea to attempt to pace AI progress? We start by considering arguments from this perspective, first against (2.0) and then for (2.1) increased pacing.
More pragmatically, pacing will only happen if actors take effective pacing actions. If the actors are rational, such actions will be taken if the actors view them as having a positive impact in expectation. In (2.2) we consider the perspectives and cost-benefit calculus of different actors, noting their incentives, operating environments, high-level control over pacing, and relations to other key actors. In that section we aim to both explain why current actors are or aren’t taking, or advocating for, pacing activities, and set out conditions under which we expect them to advocate for or take steps to pace in future.
The following takes stock of the facts, assumptions and dynamics pacing arguments rely on. Later sections then explore these in detail, surfacing further considerations that affect the costs and benefits of pacing. We return to a more comprehensive assessment of the benefits and costs of pacing, taking into account second-order effects of the interventions required to make pacing effective, in (5).
2.1 Why pace less?
2.1.1 Pacing means we have to wait longer for very good things
AI progress so far has had some highly positive impacts: software development is greatly sped up, access to information and advice has widened in every domain of life, and tools like machine translation and automatic paperwork filling save people vast amounts of time.
Many observers expect far more significant impacts on economic growth from future AI progress, with the median expert in one survey expecting rates of growth to more than double in advanced economies by 2050, amounting to tens to hundreds of trillions of dollars annually. And AI already seems to be speeding up life-saving technology.
Reducing the rate of frontier AI progress would at least delay these enormous benefits. People might well counterfactually die and suffer from progress not happening as fast. Thus, if we want to make the case for pacing more, we need to form a positive case which will overcome this opportunity cost.
This opportunity cost may even include mitigating other potential existential risks, such as biorisks and nuclear war. So even the threat of existential risk from advanced AI does not necessarily warrant pacing, depending on the relative reduction of other risks.
2.1.2 Good pacing can get in the way of better pacing
Even if we accept that there are risks from further progress, it may be very important to aim for the best forms of pacing rather than merely good ones, if there is a limited capacity to sustain pacing, for example because of finite political capital.
One way this manifests is that it might be very important to time pacing right: making concrete inroads on solving AI safety problems is itself easier as we move closer to the point where the risks manifest - the models we use become more analogous to the future models we worry about, and lessons learned through experiments on more advanced models are more likely to remain valid when we get to the true danger zone. Slowing down preemptively costs political capital we ought to save for a higher-leverage time.
Indeed, AIs themselves might be an important tool in addressing risks. For example, a classic motivation for slowing AI progress is to buy time for research into value alignment, oversight, and interpretability. But one of the main current bets in AI safety is developing an automated alignment researcher that can pack thousands of person-years of research into a short period. This will be far more effective as AIs themselves become more capable.
The strength of this argument depends on how feasible the better options actually are — a major empirical question for pacing is how far in advance we will see the risks and how quickly we will be able to react, as we discuss in (4.1-4.2).
2.1.3 Pacing can directly cause bad outcomes
To be stable, some pacing interventions will require strong monitoring regimes (i.e. global compute surveillance) and sometimes require law enforcement (i.e. criminalising previously legal activities like large training runs or underground datacenters). All such centralised power has the possibility for abuse, as well as classic failure modes like government failure and failure to aggregate or heed bottom-up information. We discuss this further in (5.2).
2.1.4 Pacing can create overhangs
Many proposed pacing interventions, such as constraining training runs (but not constraining algorithmic progress or hardware innovation) could mean that AI training becomes more efficient. That is, it could be that the paced field builds up the potential to rapidly scale later. Then, when the pacing breaks — via defection or a wholesale policy reversal — AI progress arrives discontinuously.
These “overhangs” could be even more dangerous than the relatively steady compute-bottlenecked situation we currently find ourselves in. Discontinuity is what we handle worst; badly designed pacing thus converts smooth risk and predictable progress into lumpy risk. Instead of seeing 3 years of progress over 3 years and allowing society to adapt, we could see 2.5 years where we only see 1 year of the expected counterfactual progress, and then suddenly 2 more years of progress in 6 months. We discuss this further in (4.4).
2.1.5 Uncoordinated pacing makes the incautious win
As of writing, the leading labs are taking costly steps in the name of model safety, and the leaders also generally score better on safety practices than trailing labs. Pacing interventions which place greater burdens on leading labs (sensibly, since advancing the frontier is more dangerous than catching up to it) will lead to more incautious actors catching up, and, in the case of circumventable restrictions, surpassing existing leading labs, creating greater risk than in the counterfactual where the current, more cautious, leaders retain their edge.
This is particularly biting in the case of voluntary or unilateral pacing, where one need not even appeal to the records of specific actors to see that the less cautious will have an edge.
2.2 Why pace more?
2.2.1 Pacing lets us pay down safety debt
If we assume that the capacity to cause harm posed by AI scales with its capability (especially as the capability reaches and surpasses human level), then if capability advances faster than safety, we get increased risk, leading to a safety debt. Pacing could help by allowing our protections to catch up to the current level of intelligence.
The extra time bought by pacing could let us strengthen whatever protections are falling behind — similar to how Anthropic’s Project Glasswing and OpenAI’s Trusted Access for Cyber program delayed the public release of their most advanced models, while still providing it to certain key groups that could use it to shore up their defences.
Safety debt is an umbrella term for any efforts aimed at both making models less harmful and more controllable, and making human societies harder to destabilise. In some cases, this will mean making sure a newly developed model isn’t deceptive or that agents don’t engage in scope creep. In other cases, we may want to make societies more resilient to advanced AI—for example by increasing cyber resilience or giving labour markets time to absorb AI-driven job displacement. Pacing may allow time to explore anomalous behaviours before release or free up resources (incl. man-hours or compute) to create reliable countermeasures.
Knowing what we will do with extra time makes making the case for pacing easier - this favours cases where the bottleneck is limited resources or institutional attention, as opposed to solving issues which require unknown scientific breakthroughs, where the tradeoff between getting better AI to help solve the problem and the increased risks which accompany that are less clear.
In cases involving an evolving risk (such as incremental increases in the capacity of AIs to evade oversight or incremental disruptions to the labour market) it may be preferable to titrate: to allow capabilities to gradually outrun our protection, in a controlled manner that caps the potential harm, so that we learn what type of protection is required. A problem with this approach is that there is a very fine line between “building protections in response to empirically observed behaviours” and “bumbling along in the dark hoping that we’re not about to irrecoverably walk over the edge of a cliff”. In other words, titration risks mistaking unpreparedness for strategy.
2.2.2 Pacing lets us avoid irreversible consequences
A well known heuristic in risk management is the precautionary principle, which holds that where an activity poses a threat of serious or irreversible harm, lack of full certainty of the extent of that harm should not be used to justify avoidance of precautionary measures. Some developments can make later intervention less effective or much harder. Pacing can preserve control over a decision even when the right decision is not yet clear, and give actors more room to reason through their options.
For example, once model weights have been published, developers lose the ability to pull that model, add guardrails, or intervene on specific malicious usage — options which frontier closed-weight developers have exercised several times in response to unexpected risks.
The threat of irreversibility must be invoked judiciously. Almost any restriction can be defended by invoking possible future harms. Pacing is most justifiable on these grounds when there are specific options that would disappear by default, and a positive case for how extra time could lead to a more considered decision.
Irreversibility also cuts both ways. The prospect of enabling less cautious actors to catch up or gain an unrecoverable lead can render an otherwise attractive pacing intervention unviable.
2.2.3 Unknowns
Despite our best efforts, we cannot predict all the potential effects of AI progress. Sometimes a system displays an unexpected capability or causes unexpected outcomes, and even without a clear picture of how one would remedy the situation, it is helpful to slow down and make sense of the situation. The immediate use of time is therefore inquiry: investigating what happened, perhaps reproducing the result, and determining which assumptions need to change. A benefit of progressing at a slower pace is that this becomes more feasible - some processes, such as human led-investigation and deliberation, can only be done so quickly, and if the pace of overall progress becomes too fast these may lose a lot of their effectiveness.
For example, OpenAI temporarily halted internal development of its “Astra” model after discovering another unreleased model had broken out of a secure sandbox and launched a cyberattack on another company. The activity was unprecedented enough to warrant a sudden stop, even without a particular plan for how to respond, because reasonable security assumptions were violated and the risk of continuing was judged too high.
This type of threat is inherently hard to fully plan for because it is a catch-all for the unexpected. The existence of this category highlights the need not just for specific intervention plans but for the capacity to quickly and effectively intervene in new ways that address changing situations.
2.3. Actors and their incentives to pace and not pace
In the real world, reasons to pace and not to pace coexist—and create tensions. Sometimes, different actors may have conflicting preferences for the outcome, and sometimes they may have different incentives to act. Here we look at incentives of various actors beyond the above which may affect their willingness to pace more.
Some generally applicable reasons that may bring actors who would not unilaterally slow down to the table include:
- Catch-up runway: A pacing intervention designed to cap the frontier (rather than status-quo freeze) may immediately affect more advanced actors while less advanced actors get some wiggle room to catch up.
- Savings: For advanced—but also less advanced—actors, slowing development may result in more profit, as they face less pressure to spend heavily to keep up. It may also allow diversion of more time and resources to ensure public safety.
- Prospect of exclusion: If one actor is in more control of the resources compared to others, they may choose to threaten to withdraw resources from those who do not cooperate with a pacing intervention. For example, NVIDIA chip access, or even more broadly, TSMC advanced-node access.
- Value alignment: Actors may value the cooperative outcome above what they could gain by defecting. The bad news is that this incentive hinges on mutual trust and reliable verification infrastructure. The good news is that we’ve seen it work in real life.
Governments
Governments of leading AI-developing nations (AI “superpowers”) face something of a prisoners’ dilemma: even if convinced of the global benefits of pacing, their relative position may be harmed if they unilaterally introduce pacing interventions, and actors in more lax jurisdictions not party to the agreement catch up or speed ahead of them.
Given a sufficiently large lead or large imminent risk, this need not be dispositive, and they could still rationally decide to pace, but it would be much easier for them to agree to pacing if they could assure themselves it will not cost them geopolitical advantage.
Governments which are currently behind (e.g. China, the EU) may conversely see benefits from pacing if they believe this will allow them to catch up eventually. For example, China may expect to be in a relatively stronger position in the future once their domestic chip industry has a chance to catch up, and so buying time now may look attractive.
Even without the prospect of catching up, if a faster pace of progress looks like it may lock in American hegemony in the long run, other actors may be interested in buying time and delaying this.
AI developers
AI developers can be split into two camps: frontier developers, and non-frontier developers. The economics of current AI strongly favour frontier developers, with frontier models commanding a significant premium over those 6 months behind the state-of-the-art.
This means developers who are currently on the frontier need to be assured they will retain their competitive advantage, or achieve some other way to retain pricing power, or face significant economic harm from slowing down and allowing others to potentially catch up.
Conditional on achieving this, they may have strong incentives to pace. If the race dynamics they are currently subject to are significantly weakened, they will face less pressure to pour all their resources into R&D, and may even achieve sustainable profitability more quickly.
For developers who are behind the frontier have competing incentives, the opportunity to catch up if the frontier is slowed down is a feature, not a bug.
In either case, developers will fear that their competitors may circumvent the interventions, leaving them worse off, and may try to do this themselves, so any intervention needs to make a credible case to developers that it will be reliably enforced. Developers will also bear costs from any regulation, as their resources will be required for compliance and new bureaucratic hurdles to navigate will come into being.
Citizens and their advocates
While the actors at the table in a given negotiation will likely primarily involve governments and companies, the former are accountable to their electorates (or, in the case of authoritarian regimes, still maintain an interest in retaining public goodwill). This means the incentives of citizens and their preferences matter in pacing considerations.
At time of writing, the citizens with the most direct path to impacting pacing are likely US citizens. In general, US citizens are suspicious of AI and of significant changes, and individual regulatory interventions usually poll well. Fears of labour displacement also make slowing down AI progress a reasonably easy sell. However, the greater the impact of AI on the economy, the more citizens’ jobs, wellbeing and the sustainability of tax revenues will become dependent on continued AI investment, meaning this is not guaranteed to persist.
What this means for a workable intervention
Rather than rely on voluntary, unilateral interventions, some contexts will require coordinated pacing. Coordination strategy adds yet another layer on top of other criteria for evaluating pacing interventions outlined in later sections: a quality intervention must justify not only why it’s worth disturbing the default trajectory in general, but also why the intervention’s design incentivises involved parties to agree and stick to it.
While domestic coordination alone may not overcome international race dynamics, it may help create the trust and infrastructure—such as verification processes—necessary to make coordination work on an international scale.
2.4 Applying “why pace?”
Case A – Frontier training cap
Frontier model progress has increasingly made AI systems capable of automating their own further improvement. Anthropic, for example, claimed that it was producing 8x as much code per researcher since the release of Mythos 5, when compared to the pre-2025 baseline, and that its own researchers estimated they were sped up by a factor of 4x, though Anthropic thinks that this was likely an overestimate. If this AI contribution to AI became sufficiently large, capability development could accelerate while also becoming less dependent on human researchers. The time available to evaluate successive systems might shrink, even as previously functional oversight measures break down and unexpected new risks emerge.
One direct intervention aimed at slowing down this trajectory could be a cap on the compute budgets for training frontier models. This would impede one major source of capabilities progress, and could therefore help to extend the window of controllability before existing methods and systems are no longer capable of adequately supervising AI progress, and buy time to push on oversight measures: evaluation, security, and governance systems may be inadequate for development substantially accelerated by AI.
Case B – Uplift to biological weapons
AI systems are increasingly able to aid some users in developing biological weapons. Governments and developers may face some lag in their ability to assess uplift, to restrict it in specific models, and to deploy model capabilities to develop countermeasures. Furthermore, releasing the weights of individual models removes any ability to regulate them if they turn out to provide an unacceptable degree of uplift.
Complementary activities include better uplift evaluations, safeguards and unlearning, secure hosting, user authentication, controls on model weights, public-health preparedness, and international procedures for handling dangerous models and evidence.
2.6 Open research questions
- How can we choose an optimal pausing point, trading off increasing risks against benefits from continued progress, and increased understanding of the risks themselves?
- Under what assumptions about the relative difficulty of automated alignment vs automated capabilities progress does it make sense to pursue automated alignment more aggressively vs seeking to slow down? And similar questions about offense/defense balances in other domains
- Where do different dangerous capabilities sit on the offence/defence balance, and how should we expect that to change over time?
- What are the bottlenecks on adaptation to different AI risks? How much of a head start can we plausibly get, and in what sense is time the bottleneck?
- Under what assumptions does pacing become more difficult over time, and what investments can be made now to preserve the option to pace in the future?
- How do we evaluate if an intervention will be worthwhile? Quantifications of expected specific effects of interventions could help justify cases for skeptical audiences and add rigour to a domain where there is often a lack of specific claims about impacts.
- Which avenues of progress are less dual use, and therefore better targets for differential technology development?
- Under what assumptions does more transparency on research progress serve to enable coordination, and under what assumptions does it intensify race dynamics by revealing that a rival is close behind?
Pace What?
A successful pacing intervention would change the rate of one or more activities that make up AI progress, but there are many activities to choose from and the choice matters for the effectiveness, cost, impacts, and desirability of the pacing intervention. It is helpful to distinguish between the functional target (the part of AI progress that causes some threat) and the control surfaces (the parts which can actually be changed), because they tend not to neatly line up. Indeed, often we will have a lot of uncertainty about what the threat even is.
There is, for example, no direct lever to prevent rogue actors from using AI to launch cyberattacks. But there are levers on compute, on model releases, and on classes of use. Each of these picks out some part of the threat, catching some activity that is harmless and missing some that is not, with varying costs to privacy, economic growth, and oversight.
So the choice of control surfaces is the first major design decision where serious costs are incurred. In general, interventions that target control surfaces earlier in the AI R&D process (e.g. interventions on access to chips) tend to have broader impacts, have a higher likelihood of lowering risks, have a higher likelihood of harming beneficial progress, are much less able to leverage information generated in the R&D process, and are therefore much less likely to be targeted at specific harms. Conversely, interventions that target control surfaces later in the AI R&D process (e.g. inference-stage safeguards, model access controls) are much more able to leverage information and therefore be more targeted at specific harms, but are also more likely to be circumvented (as the harmful artefact already exists). .
3.1 Control surfaces of AI progress
Modern AI progress is an extremely costly, distributed, and ever-evolving process. It begins with inputs like compute, hardware, capital, researchers, data, and accumulated knowledge. These feed development activity: pre-training, but also post-training, algorithmic progress, and broader machine learning research. Collectively, this produces models, which can then be retained inside a developer, offered through a controlled service, or released as weights that can in turn be copied, modified, and rehosted without limit. Outputs depend on the model, the overall elicitation capacity, and a supply of compute on which the model can run.
All of these are plausible control surfaces to directly intervene on. Furthermore, they can all be affected indirectly. For example, the availability of researchers depends on immigration law, and frontier hardware depends on access to certain critical minerals.
Different actors vary in their leverage over these control surfaces. Abstractly, states have the ability to enforce domestic laws, while developers have a richer understanding of many parts of the design process. More concretely, specific states and developers vary in which of these they have leverage over — there are some remarkably narrow bottlenecks in the AI ecosystem, like the reliance on a single Dutch multinational corporation for a critical part of the AI hardware manufacturing process.

One crucial distinction in the process of AI development is that part way through the resources shift in nature. Many of the early inputs are rival goods: resources which can only be used by one actor at a time. Compute, capital, and researcher time are scarce, are concentrated within a small number of identifiable organisations, and can be redirected. There are also non-rival goods which can be replicated at almost no cost, like training data, but they are not sufficient to make progress. Downstream, however, it is mostly non-rival goods, like weights, elicitation methods, and training algorithms (the notable downstream exception to this is inference compute, which is a rival good). These can all be easily copied, so the set of holders only grows.
These components allow for quite different kinds of intervention. The rival components have locations and quantities — they can be monitored, taxed, and redirected, and it is possible to mostly keep track of who has them. The nonrival components tend to be more like information, and once released, are very hard to restrict access to. Interventions therefore have to be more conduct-based.
That said, actually using a model requires inference compute, which is a rival good. However, it is also extremely general and extremely widely available — many highly capable open-weight models can even be run on consumer hardware, though near-frontier models typically require more specialised infrastructure..
Once a nonrival good like model weights is proliferating, it is very difficult to destroy every copy and there is no way to verify a claim that all copies have been destroyed, so interventions on their distribution are all-or-nothing. This is true even for closed-weight models, where uncertainty over whether adequate security over model weights is maintained means that, at current levels of security and verifiability, it is impossible to verify that all copies of model weights are known.
By contrast, for a rival good like computing hardware, an oversight body might be able to be much more certain of the location and use of all copies of a particular type of chip, once adequate measures are in place, since the exact number of such chips can be more reliably known.
Part of why this matters is that many interventions will be imperfect and leaky. They can still be effective if the amount that they leak over their entire lifespan is low enough, but this looks very different for rival and nonrival goods. For example, imagine an intervention aimed at preventing broad access to models above some capability level, either by limiting the supply of training hardware or by limiting access to the weights of such models. If 90% of the hardware could be secured, it would be possible to roughly bound the scale of covert model training. But if the model access were successfully limited in 99% of cases, it would not be possible to bound how far the weights could proliferate — a single stolen copy could then be copied at will.
The central issue in choosing a control surface is that a lot of the most useful information comes too late in the process of model creation. As development progresses — from training to evaluation to deployment — a greater quantity of useful information about the capabilities and risks becomes available. The problem is that the ability to respond effectively diminishes as the model progresses along the development chain, as many of the control surfaces cease to be available. Once a model has been trained, there is little to gain from restricting the algorithms used to train it. In particular, most of the rival components appear fairly early in the development process, so beyond a certain point there stop being rival places to intervene other than inference compute, which is unfortunately already very widely distributed.

3.2 Targeting and trade-offs
Every pacing intervention divides activity in two. Some of it stops, slows or proceeds under different conditions. The rest continues unaffected. Any pacing intervention will thus, either explicitly or implicitly, select which activities belong in which of these clusters.
The selection can fail in two directions. A false negative is activity which contributes to the threat, but which is not caught by the intervention. A false positive is an activity caught by the intervention, but which does not contribute to the threat . Both errors can occur simultaneously: A classifier that aims to flag nefarious research related to manufacturing biological weapons, for example, may be triggered by the innocent questions of a biology student. At the same time, the classifier may fail to catch carefully designed questions that split up a dangerous avenue of inquiry into unassuming parts.

Interventions earlier in the development process will have wider-reaching effects, which generally means more false positives and fewer false negatives — in the limit, halting all AI development forever would indeed prevent all harms, but it would also prevent all benefits.
For most threats, there are multiple distinct pathways to their realisation. A capability advance can come from more training compute, better algorithms, more elaborate post-training, or better elicitation of a model’s existing latent capabilities. An intervention which targets one route but leaves the others open still allows some threat-relevant activity to continue even if the intervention is perfectly enforced.

The quality of selection also changes over time. This is partly because actors adapt in response to interventions, as we discuss in (5). But even beyond that, the effectiveness of an intervention depends on empirical assumptions about the relationship between the control surface and the functional target. These assumptions may be changed by later progress.
For example, one very conservative way to prevent the emergence of a dangerous advanced capability is to cap the number of FLOPs that can be used in training a model, but even this intervention depends on assumptions about training efficiency or elicitation methods which might change with future progress. Generally this will push towards more false negatives, as the intervention becomes outdated.
More generally, AI progress has complex feedback loops. At a very basic level, frontier developers use their models to make money, which they invest in creating smarter models and hiring more researchers. Increasingly, they also use their models to improve the quality of their research. This means that constraints at one point in the process can ripple out and have complex indirect effects.
For a given control surface, there will often be a three-way tradeoff between false positives, false negatives, and intrusiveness or oversight. Simply put, one can make an imperfect rule more or less broad, or one can invest in making the rule more accurate, by some mix of investing more energy in scrutinising individual cases and requiring more access to information about those cases, some of which might otherwise be private.

Broader rules, which capture a higher fraction of threat relevant activity, carry with them greater economic costs, while narrower rules may be easier to circumvent and thus have fewer safety benefits.
Alongside interventions which directly cap or limit some resources, it is possible to intervene indirectly—to change an incentive and let the selection be performed by the actors themselves. This includes buying out or taxing specific resources like compute, or applying stricter liability for outcomes to model providers or other actors with control over capability access, such as cloud hosting services. These approaches can effectively draw on private information no regulator could extract: private plans, valuations, and alternatives. Ultimately this means indirect selection needs no surveillance to operate and thus is less intrusive. However, these interventions will still generally lead to false positives and negatives.
The correct balance between these three constraints will depend on the goals of the intervention and the resources available. One critical point to bear in mind, though, is that for many interventions, the function is to buy time for some complementary activity which in some cases may risk being caught in the filter. For example, a filter which blocks requests that enable cybercrime may also prevent requests related to the testing and evaluation of dangerous cybercapabilities, undermining the nominal aim of mitigating the risk of cyberattacks. That is to say, it is difficult to target only ill-intentioned activity while admitting related safety work when they often draw on the same capabilities.
Since essentially all control surfaces suffer from these tradeoffs, often the best route to a given outcome will involve several different control surfaces working in parallel — for example, a cheap intervention with few false positives and many false negatives, coupled with a more demanding one that can catch some of the false negatives which slip through the previous one. Relatedly, in fast-moving situations it may be preferable to quickly enforce conservative interventions with many false positives, and use the breathing room to implement more careful and calibrated ones.
3.3 Coordinated interventions
In cases where an intervention requires coordination among multiple actors, there are some additional properties which need to be satisfied.
Interventions and actions must be legible to all parties. Thus, more detailed proposals are likely necessary than in unilateral cases, to pin down which actors are committing to which actions, and what the consequences of noncompliance will be.
These consequences themselves must also be devised — in the simplest cases, actors mutually benefit from everyone adhering to the rules, and a penalty for defection is simply other actors also defecting, but other mechanisms seem plausible, such as withholding licenses to purchase chips, or to offer AI services legally in certain jurisdictions.
The history of arms-control suggests that agreements between rivals are easier to reach when compliance can be independently verified without trust, making control surfaces which enable trustless verification especially desirable. Coarse, numeric thresholds like FLOP counts are easier to evaluate than holistic appraisals like bio uplift capacity, and so it may be easier to coordinate around them as interventions.
Unfortunately, while information about model capabilities and threat models accumulate steeply along the arc of model development, coordination-grade information does not. The information that can be gained about a system nearing deployment is often nonstandardised and difficult to make legible—the internal testing protocols and mitigation mechanisms at one lab may not look the same as at another, even if they point at the same targets. Upstream, however, information on rival goods can be easily legible and shared, making coordination easier. As a result, coordinating actors will tend to select upstream control surfaces: the crude and high-false-positive side of the arc provides a source of surfaces better suited to coordination.
3.4 Applying “pace what?” Case A - Frontier training cap
Suppose the concern is that AI systems may substantially accelerate or automate AI R&D before adequate means of control exist. The functional target is a transition towards AI-driven development.
Control surface: Training compute is an early, rival input into this process, and therefore represents an attractive control surface for this problem.
“Capping training compute” still requires further specification of the control surface, however. How should experimentation and post-training compute spend count towards this, if at all? If novel methods involve combining the results of multiple training runs into a larger model, will this still fall in scope? How does one deal with the problem that further algorithmic progress will allow a given level of capabilities to be achieved with less compute as time goes on?
Targeting and trade-offs: A compute ceiling only bears on one aspect of the problem, and does not address capability gains from other routes, such as better data or RL environments, post training compute spend, or increasing test-time compute usage. It may also be too broad, and stop some benign or benevolent work.
There is also the risk of a “capabilities overhang” where artificially capping one input could lead to a sudden jump in capabilities if the restriction is lifted or if one actor defects, which could mean the incentive to defect and take advantage of this overhang increases over time, leading to an unstable equilibrium in the shape of a classic prisoners’ dilemma.
Coordination: Compute usage is a comparatively legible surface, if actors can be required to share records. Many proposals for governing compute usage and tracking chips exist; it appears the technical problems involved are feasible to solve. The coordination required involves figuring out who has authority over verifying compliance, who is included under the scope of the policy (which individual developers, and which nations), and how to respond to capability gains by non-participants so that continued participation remains preferable to defection.
3.5 Applying “pace what?” Case B - Uplift to biological weapons*
Suppose instead that the concern is the diffusion of AI assistance that materially increases users’ ability to develop biological weapons. Here the functional target is more about access and usage than training.
Control surfaces: Model deployment and access, gated by capability evaluations and subsequent restrictions. Post-trained or API side restrictions, such as refusing prohibited requests, are one potential approach, while another is attempting to train the model not to be capable of fulfilling such requests in the first place, perhaps by ablating training data to remove biological knowledge. The latter, if successful, also has the advantage of being somewhat more robust to open weight models having their restrictions lifted with further post-training, which is a significant risk in the former case.
Targeting and trade-offs: Attempting to make the model refuse noncompliant requests faces three main issues — over refusals, under refusals and jailbreaking. Decisions must be made on how jailbreak resistant the model needs to be, and how much refusing of benign requests is tolerable. The capabilities evaluations themselves must also be sufficiently comprehensive to make sure paths to dangerous capabilities are well guarded.
There is also the question of whether it is desirable for specific actors, such as those expected to use models defensively, to have access to versions with lighter restrictions, and how to gate this access reliably.
In cases where usage restriction is the target, there are more obviously tangible benefits to even imperfect restrictions. The more effort and technical capability it takes to circumvent guardrails and do something dangerous, the less often you’d expect harmful behaviours to occur.
Coordination: The control surface here is less widely legible than in the compute cap example, and relies upon the existence of a trustworthy evaluator with sufficient expertise to make determinations. In a legislative case, this could be an independent, government, or international body. In the case of voluntary lab compliance, this could be the labs themselves, or a trusted third party to mitigate mistrust between racing labs. Difficulties again arise with what to do about non-covered actors, though in the case of narrow restrictions on usage compliance is less costly for model providers than with broad restrictions, since the majority of their customers might see limited gains from access to these capabilities anyway, and exceptions could plausibly be tailored for legitimate cases.
There is also a risk that, once an open model with a given level of capabilities exists, subsequently attempting to control it may achieve substantially less and may not justify ongoing costs, so the policy may be brittle in the long run.
3.6 Open Research Questions
-
How should functional targets be translated into covered activity? Functional targets such as automated AI R&D and bioweapon uplift can arise at several points in the process - they will emerge at some point in training, and we may want to avoid such a point being reached, or we may care more about wider deployment (especially if capabilities have positive use-cases we want to preserve). Research should compare candidate boundaries and weigh the pros and cons of each.
-
What multidimensional measures best track progress toward AI-led R&D?, such that we can choose a good target and expect it to remain a good target as time goes on?
-
How well does a control surface track the target? Aggregate training compute and capability evaluations each capture a different part of the process. Comparative work should estimate the effect of constraining each surface and identify the relevant activity outside it.
-
Which interventions will remain useful as technology changes and which will become irrelevant? An intervention which is specific to a particular mode of operation or model architecture which could be superseded may lose its teeth.
-
How should related work count together? Aggregation rules must handle projects divided across multiple training runs, or across different corporate entities. Testing should identify rules that identify deliberate fragmentation without sweeping in unrelated work.
-
What practical coverage is sufficient? Compare the reach available through company control, infrastructure providers and national rules, including their supply-chain effects. Estimate when activity by outsiders is large enough to defeat the intervention.
-
Once some diffusion of dangerous capabilities has occurred, which restrictions lose their teeth entirely, and which continue to meaningfully reduce risks? While some loss-of-control scenarios are presented as all-or-nothing, in many cases ease of access to dangerous capabilities will continue to meaningfully influence the net harm done. Needing to spend hours jailbreaking a model to get useful outputs from them is meaningfully more friction than having them readily assist with harmful endeavours straight away when asked.
-
How can we make a rule which permits exceptions to preserve useful work, which is not so permeable that it makes the rule useless? In the case of compute controls, it seems technologically feasible to identify what uses a GPU is being put to on some levels. The more we can do this, the less an intervention will need to be a blunt instrument versus being narrowly scoped.
Pace How?
A pacing intervention is ultimately a sequence of steps; we reason about an intervention from start to end and consider its many decision points and failure points. This lets us spot the supporting work required for an actual, sustained period of restraint.

4.1 Before pacing
Before we pace, the main challenge is recognising where there’s a need for it and its proper timing. Intervening too late means letting the threat play out with potentially irreversible consequences; intervening too early means sacrificing potential benefits and political capital, and in some cases being less able to carry out complementary activities that depend on access to advanced AI, or to the benefits it brings like an abundance of resources.
Sources of information to draw on:
-
Forecasting and effective predictions provide us warning signs of when dangerous capabilities are likely to emerge. Different forecasts shed light on different parts of potential dangers: for example, trends in training resources can be used to forecast the size and nature of the resulting required physical build-outs, while capability forecasts can predict models’ future behaviour, but neither one can replace the other.
-
Model evaluations provide more direct information about the capabilities and propensities of specific models in specific circumstances, and can form the basis for extrapolation. That said, the correspondence between evaluations and real-world risks can be quite fraught and hard to predict: the jaggedness of AI capabilities means that individual evaluation performance often corresponds poorly to behaviour on organic tasks. Furthermore, evaluations can generally only set a lower bound on capability, because of the challenges of elicitation and the risk that models alter their behaviour in response to recognising they are being evaluated.
-
Incident reports provide us concrete evidence with ecological validity and lots of useful surprising detail, properties which both forecasting and evaluation often miss or fail to attempt. This is the best indicator that there is an immediate AI-induced problem, but by the time an incident has occurred, the space of possible pacing interventions has already shrunk. But even relatively harmless incidents can expose unknown unknowns and more granular details, which can in turn reveal underexplored threats. Also, compared to evaluations and forecasting, real-world incidents can be more credible and legible to a broad range of actors.
-
Safety cases provide structured arguments that a system will remain within some acceptable risk parameters as long as certain prior conditions are met — for example, that certain features of a deployment context guarantee that a given harmful model capability cannot be accessed. This structure therefore makes it possible to turn a discussion of potential risk into a discussion of concrete underlying properties which can be directly scrutinised. Safety cases are presently not in any meaningful sense a requirement for deploying systems, but rather a voluntary engagement where they exist at all - changing this is itself one plausible intervention which would shift the burden of proof onto developers, where currently the burden is typically laid on those proposing a slowdown.
Forecasting, evaluations, incident reports, and safety cases perform different functions: respectively, to supply warning signs, to test certain bounded claims, to reveal models’ emerging behaviour, and to decompose threats into specific precursors. Alongside feeding into the question of whether to directly intervene, they can also feed into each other — incident reports can shape what gets evaluated, evaluations can form the basis for forecasts, and so on. Collectively they can provide a sense of how far away any given risk is, and this in turn can help actors to determine how urgently they need to prepare for costly measures.
The extent of meaningful and effective communication between actors at different stages of development suggests what kind of preparations are needed. For example, if evaluative capacity is limited, then an actor would know that advance investment may be necessary. When communication breaks down, delays occur and the landscape evolves; the alternative to preparing now is not necessarily the same decision made later under the same conditions.
Coordinated pacing comes with extra challenges: the different actors need to be able to converge on their understanding of the relevant information and what constitutes meaningful evidence, which in turn typically requires some bandwidth between them. In situations where they have competing incentives, some will also have individual incentives to withhold information — for example, frontier developers might not want to be overly restricted by governments, and so they might not want governments to have the information that would warrant such restriction. There may need to be infrastructure developed in advance to get around these dynamics.
4.2 Triggering and enacting pacing
The challenge in deciding to trigger a pacing intervention is negotiating the tension between speed, legitimacy, and precision. It is easy to have two but hard to have three:
-
Speed + Legitimacy: Pre-registered rules that automatically trigger. Unfortunately it is hard to know what exactly the rules should be, and such commitments can misfire.
-
Legitimacy + Precision: Careful deliberation between experts to form a considered judgment. Unfortunately this degree of care can take a lot of time.
-
Precision + Speed: The actors with the most information and expertise choose by fiat. Unfortunately these actors will often face incentives around the pace of progress that diverge from the broader public — it would be naive to expect them to simply regulate themselves.
We can see this trilemma play out across the whole ecosystem of AI progress: governments have the coercive power to impose strong regulations but not necessarily the raw information needed to guide decisions, or the expertise to interpret such information; external evaluators can have the expertise but not the mandate; AI developers have the most information about their own progress but little reason to proactively internalise any negative externalities they produce, and a lot of other interests in how their competitors are regulated.
Again, one way to ease the tradeoffs is to work towards a portfolio of triggers. For example, we can set very crude evaluations thresholds now, beyond which we give expert groups like third-party evaluators the right to unilaterally trigger emergency interventions, on the condition that those choices then be subject to later governmental inquiry, where there is more of a mandate but less ability to rapidly deploy expertise. This approach might seem risky if one views pacing as a one-off affair, but once it becomes a repeat affair, the evaluators would have a long-term interest in using their powers in a way they could justify.
The next decision is execution: the intervention must somehow affect live systems[^1]. Implementation-wise, the key question is whether enacting the intervention merely involves announcing a rule, or whether it involves some more active steps like buying up, confiscating, or destroying key resources. The former case is more straightforward, but brings with it the extra challenge of actually enforcing the new rule.
Again, coordinated pacing comes with extra challenges. It will typically take a lot more time for several actors to form a consensus on whether an intervention should be triggered, especially if there is a range of options to choose from. One potential solution is to give several parties the ability to unilaterally trigger time-bounded interventions which can then make way for more careful discussion. Another is to spend time in advance mapping out the likely space of mutually beneficial interventions.
4.3 During pacing
During pacing, it can be difficult to sustain the efficacy of an intervention in an ever-evolving environment. There are four notable challenges: are the relevant parties still complying?; is the initial threat still effectively targeted?; is the threat still a threat?; and is the suspension in activity being used effectively?
We can address the first problem with monitoring, verification and enforcement systems. The second component is whether the control surface continues to track the functional target – whether people complying with the terms of the intervention actually diminishes the risk. (3) describes this relationship in terms of false positives and false negatives and gives more detailed considerations. Typically, a static rule will become less effective over time as the practical implications of the rule drift from their initial state.
The third component is whether the original reason for pacing still holds. The concerning capability may mutate, diminish, or intensify. New evidence may also undercut the rationale for designating a result as a threat. The strategic and commercial reality may evolve too. Companies may exit agreements and safety measures may strengthen, for example. Importantly, this does not bear on the effectiveness of the rule itself, but whether the threat targeted by the rule is still of concern.
The fourth component is whether the interval is being used to change the conditions which made pacing valuable, as discussed in (2). This could simply amount to some ambient societal adaptation, but typically it will involve deliberate complementary activities aimed at ensuring that the threat will be less severe after the intervention than before. Absent good enough complementary activity it is possible for an intervention to merely postpone a growing crisis.
There is a danger that pacing weakens the infrastructure it is designed to strengthen. Sweeping restrictions on AI, for example, would likely shrink the opportunities available to the actors involved in frontier systems, potentially limiting their experience and understanding of frontier models. To mitigate this outcome, a limited and select group could be permitted to continue research: a tight circle of knowledge and power limits the threat potential while maintaining expertise. The risk, of course, is that this group would gain outsized power and influence. The nature of the complementary activity to the restriction – typically, bolstering safety procedures in some form – can inform decisions related to the chosen group’s composition, focus, and oversight. When reviewing the intervention, the same considerations as in activation are in play: the tensions between speed, legitimacy and precision.
A potential solution could be to separate retargeting by scale and reversibility. Operators could have the authority to make temporary and bounded adjustments; external reviewers could implement wider and more durable policies. Similar to the activation procedure, tradeoffs cannot be wholly eliminated, only managed across the model lifecycle.
4.4 Ending pacing
There are two primary reasons we might wish to exit a particular pacing intervention: the intervention worked and has served its purpose, or the intervention is not working — having proven ineffective or become obsolete.
In the case where an intervention worked and is no longer needed, this can be because the initial goal was to build up defensive capabilities or resilience which has now been built, mitigating the risks involved. An intervention which is not working anymore could take a couple of forms: it could be addressing a redundant pathway to impact which has been routed around, it may have had a misspecified theory of impact in the first place, it may have insufficient teeth to ensure compliance from relevant actors and lack the support necessary to fix this, or it could simply have been superseded by a subsequent intervention or policy.
Which of these cases we are in determines what the goal from an exit ought to be. In the successful case, we should aim to preserve the gains we have made while reducing the costs we are imposing, while in the failure case, the goal should be to simply remove the costs while minimising potentially harmful disruption. The former case may require a more careful, gradual process than the latter, since the potential damage from getting it wrong will be higher.
The question of when to exit is important. Exiting too soon may expose us to the very risks that pacing was meant to prevent, while having burned goodwill and support for further intervention, and exiting too late means imposing costs and burdens for longer, and strengthening the relative position of noncompliant actors. Predefining a framework that governs what should qualify as evidence for resumption is useful here.
The defaults also matter: if the intervention is set to expire after a certain period and requires renewal, it may be prematurely exited without regard to the motivations which installed it. If it is installed with no definite end point it may stagger on past the point where it is useful, with institutions and operators becoming entrenched, developing an interest in preserving their function. Both scenarios should be avoided, so active attention and governance is valuable in both cases.
There will inevitably be opportunity costs to any intervention, creating pressure to resume. The most affected actors may see their competitive position weakening, and accumulated overhang of resources and progress on permitted work may enable new and productive paths for development, intensifying the incentive for cessation of the intervention. In this scenario, a hasty exit might see unusually rapid and hard-to-manage progress, as developments which would have been spread out over a number of iterations of model releases without the overhang all arrive simultaneously, straining our ability to respond and adapt.
In the case of a successful intervention, we can try to mitigate the dangers of an exit by staging it. This could involve continuing some forms of restrictions, and gradually relaxing requirements over time. Evidence and research as to how to proceed can thus be safely collected, while still managing the risk. Staging thus permits actors to observe incremental effects of relaxation, and make decisions about whether this is desirable or ought to be halted.
However, staging is also subject to pressures. Each step can make reversal more difficult as dependencies, investments and expectations accrue force: an initially experimental release, for example, can create the imperative to move to the next stage of release. The effectiveness of staging is therefore contingent on the relevant actor’s ability to resist pressures to proceed prematurely.
In the case where an intervention is not working, staging is less justified, and so it is important that a quick exit is achievable. Balancing these avenues and making sure the quick exit is not misappropriated in the wrong circumstance will be a governance challenge.
Ending an intervention does not need to mean dismantling it entirely - it could make sense to retain some capacities that are slow to build – expertise, relationships, technical standards, communication channels, and so on – while lifting the biting aspects, such as invasive and exceptional powers and controls.
| Stage | Decisions |
|---|---|
| Pre-pacing | What justifies the need to intervene? What signals of emerging risk are we tracking? What infrastructure must exist for relevant signals to be detectable? Who watches for these signals? Who interprets signals and who do they report to? |
| At trigger | Who holds the authority to decide that a trigger condition has been met? How much error are we willing to accept in order to act quickly? What sequence of actions does triggering set in motion? Is the developer obliged to address the triggering concern, or free to abandon the blocked path? |
| During intervention | How is compliance observed, verified, and enforced? Who reports evidence of compliance, who audits, and who acts on discrepancies? How is the downtime being used to respond to the threat? Who, if anyone, may continue the restricted work, and under what oversight? |
| At exit | How do we distinguish an intervention that has served its purpose from one that has failed/become obsolete? Should exit be immediate or staged? Who is exposed to the effects of exiting, and who bears any costs and gains? |
Table: Key Decisions per Stage of an Intervention’s Lifecycle
4.5 Applying “pace how?” Case A - Frontier training
Assume we have decided to implement a cap on frontier model training compute budgets aimed at preventing labs from training models that can autonomously self-improve, and potentially lead to displacement of human operators from critical decision loops and ultimately loss of control.
Before pacing, we should recognise that we’re hoping to prevent the development of a certain capability, rather than addressing it after it emerges. Thus we must make judgment calls about when we are sufficiently close to the dangerous threshold to justify intervention. A problem is that the predictive tools we could use to detect relevant signals are hard to interpret: to take one scenario, some forecasts may predict imminent self improvement, while others disagree..
Triggering: A source of authority must be determined to establish when to begin intervening. Given the speed of progress currently and the uncertainty over further acceleration if a threshold of recursive self improvement is met, it seems clear that speed must be of the essence, and under our taxonomy above this suggests that one of legitimacy and precision must be deprioritised - we can have clear rules for when to trigger, or grant a source of authority power to trigger by fiat, but we cannot afford to have a long and deliberate consultation at the point of triggering.
During pacing, we should verify that the developer is complying with the intervention by, again, monitoring the inputs. Continued access to compute may be contingent on the developer granting third-party evaluators visibility into every run, which can be evidenced by logs of compute usage obtained from the compute provider. Since a compute limit does not constrain every route to capability, authorities may have reason to seek visibility into research teams’ working logs, to help determine if current limits are fit for purpose or need to be adapted.
Ending pacing here is uncertain. If models autonomously self improving is inevitable and this intervention serves only as an artificial speed limit on that process, then exiting may be desired once humanity has sufficient assurance that it can avoid loss of control, or that the benefits of development exceed the risks. If there is sufficient risk from external actors not party to these limits achieving capability parity or surpassing the parties to the agreement, then it may also be desirable to take the brakes off the pace of development.
However, if there does appear to be a hard capabilities threshold under a certain compute threshold, insufficient progress is made on assurance of safety, and the controls are sufficiently widely adopted, it may be that it would be desirable for these restrictions to persist indefinitely.
4.6 Applying “pace how?” Case B - uplift to biological weapons
Suppose the goal is instead a pacing intervention that prevents models from supplying meaningful uplift capabilities related to the design, acquisition, or synthesis of lethal pathogens.
Before pacing, labs may be required to make their models available to third-party evaluators for pre-release evaluations of potential capabilities that would provide uplift to a malign actor (e.g. to debug a failing synthesis protocol, supply the hands-on know-how that papers leave out, or piece together a dangerous method from scattered dual-use sources). A biological capability is far harder to recall once it reaches the public than to withhold beforehand, so reaching a certain threshold on the evaluations should block deployment outright.
Triggering here will depend on the uni- or multi-lateral shape this intervention takes. So far, similar interventions around the US government intervening in the Mythos and GPT 5.6 releases have been somewhat ad hoc and messy, but have been enacted with speed. A more legitimised, legally grounded process would likely require a publicised framework to make clear what is in scope and better allow predictable deployments, though it seems very likely that the capacity to ad hoc prevent the release of an otherwise out-of-scope model will persist.
During pacing, we should audit access logs to confirm the model stays contained, and confirm that any investigation is led by independent evaluators qualified to judge biological uplift, rather than the labs themselves. Ongoing monitoring to ensure latent capabilities are not easily elicited by jailbreaking methods may also be part of this puzzle.
Ending pacing may not involve any form of public release. As a condition of restricted release, labs may need to prove that even where a model retains a dangerous capability, there is the capacity and will to reliably vet their users and flag suspicious activity to the authorities.
4.7 Open Research Questions
-
What does the possibility space of exit scenarios look like? Small-scale interventions may allow for immediate release, whereas exiting interventions impacting multiple sides of the economy and society may need to be staged. Some exits may require advanced prep. We need to understand our options to pick the best one per case.
-
What is the relationship between initiation and exit triggers? Intuitively, the exit trigger should track whatever justified the intervention in the first place. But can that link break as conditions change?
-
How do the incentives of bound parties change across the lifecycle? Pacing is first and foremost a coordination problem. The efficacy of any intervention will depend, amongst other things, on the predictability of actors (e.g. how well they can stick to the rule). Foreseeing possible disruptions to everyone’s motivations to cooperate strengthens control over and stability of pacing.
-
Does a developer have the responsibility to follow up on the evidence if an intervention is triggered? If an intervention conditionally prevents a developer from pursuing an R&D interest, and the developer decides against addressing the condition, abandons this interest, and chooses to invest the resources into a different R&D direction instead, was the intervention successful?
-
What information must be shared to make coordinated pacing credible, and with which actors? How can this be kept compliant with national security considerations, commercial confidentiality, and cybersecurity? To what extent can this information sharing be kept robust to manipulation?
-
What evidence could legitimise a speedy initiation (even at the cost of this evidence being disconfirmed later on)? As always, we expect people to agree on high-level claims (“human-hunting motivations is threat enough to intervene”) and disagree on operationalisations (“what qualifies as human-hunting motivations?”).
-
How could we design simulations (dry runs) of an intervention (and how does this help us learn about failure modes)?
-
How can evaluations-based methods handle concerns over better elicitation being possible, and models potentially learning to sandbag evaluations?
Then What?
Pacing interventions have more effects than just the imposition of a rule. Analysing them means considering the whole portfolio of consequences — in particular, the way different actors subject to the rule adapt in response, and the broader effects of the governing capacity required to enact the rule.
These considerations are important both for weighing up the benefits of any given intervention and for determining which actors would support or oppose it, which in turn can shape whether an intervention is politically viable.
Realistically, developing a compelling intervention will necessarily be an iterative process of tweaking plans until the overall portfolio of consequences is broadly compelling.
5.1 Adaptation and redirection
Fundamentally, pacing is only necessary when it makes actors behave differently. Often this means the pacing stops them from directly getting something they want. When this is the case, they will naturally look for other ways of achieving it.
The easiest approach, of course, is just to ignore the rules, covertly or overtly. An intervention will largely fail unless it can actually incentivise other actors to comply. Importantly, if an actor conceals the fact that they are breaking rules, it becomes harder to monitor and intervene in other ways.
But even if actors do comply with rules, they will still find ways of adapting. For example, if pretraining scale is capped, then developers trying to create advanced models will redirect their energy towards other approaches like training efficiency, post-training, and elicitation. If advanced hardware is restricted, developers will try to find ways to train on less advanced chips.
This needn’t be a problem: often it is the primary goal of the intervention. But it is possible for this adaptation to undermine a rule or even cause it to backfire. For example, one of the objections raised against controls on exporting advanced hardware to China is that this could encourage China to develop more of a domestic industry around advanced hardware.
Relatedly, when a specific outcome is limited, this can free up the resources originally directed towards it. Right now, for example, most developers are constrained by their access to compute, so a limitation on any compute-intensive activity (e.g. training, research, deployment) will cause some substitution towards the others. The same is true of capital and researchers.
On a broader scale, almost any rule will change the balance of power among affected actors based on their relative capacities to adapt. If a rule only restrains some actors then it naturally weakens them relative to others. But even interventions which apply uniformly will affect actors to different degrees — for example, a pause on pretraining would advantage developers who have recently done large pretraining runs.
At the international level this is a central concern. Pacing interventions which impose relatively greater costs to one’s own AI industry than an adversary’s may be undesirable for this reason. The US, for example, has a significant compute advantage over other countries. Restraining them from building on this advantage further may be unpalatable if it is perceived as allowing China to catch up.
Crucially, the degree to which actors are impacted by an intervention need not be particularly related to how likely they are to contribute to the risk. This can be somewhat finessed, though: it may be possible to package an intervention in such a way that actors who are naturally more hard-hit can be otherwise compensated.
5.2 The effects of governing machinery
Many interventions will invest some enforcement power in a governing body. This will typically be some mix of access to private information and discretionary enforcement power, and may even involve creating specific infrastructure to support observation or verification. Even one-off interventions also serve to create a precedent for future activity. The power and infrastructure can then sometimes be used for purposes beyond the rule that justified it.
In developing a potential intervention it is useful to start with the intended goal and then back out the details, but pragmatically we ought to expect that actors will generally use power to pursue their own goals. It may well be that an intervention is needed to temporarily slow diffusion of a dangerous capability while building robustness elsewhere, and it may well be that plo[k-;=0m,]’this intervention is best executed by a national government broadly surveilling and restricting usage, but this does not mean that a national government equipped with broad surveillance and the mandate to restrict usage will necessarily use it only as intended.
In particular, it is not a given that these powers or the machinery that supports them will be dismantled at the end of an intervention, as we discuss in (4.4). This is not necessarily a bad thing: It might even be helpful if, for example, the systems created to enforce a domestic pretraining moratorium can then be reused in international diplomacy, or if departments created to assess model risks can feed their expertise into national strategy.
But there is a clear risk to contend with — crises can justify legitimately necessary emergency powers which then outstay their welcome. Once information has been collected, it is difficult to limit the scope of decisions it can feed into.
5.3 Brute costs
Restraining progress is fundamentally quite costly, in the sense that it involves at a minimum delaying the potential benefits of progress. One crude way to measure this is in sheer economic terms: right now, AI companies and the surrounding industry have a total valuation in the trillions, mostly representing expected future value. And this expected future value likely represents a lot of tangible benefits that could be broadly accessible — including better healthcare, cheaper goods and services, wider access to high-quality education and advice.
These are straightforward costs — interventions that slow the pace of AI development will necessarily cut into them somewhat. Of course, it may be that the benefits are agreed to outweigh the costs, but any serious analysis needs to reckon with the extent of them.
Crucially, the perceived costliness of an outcome can vary greatly between groups: someone with a very sick loved one could coherently prefer racing forward with AI progress in pursuit of medical progress even if it comes with an enormous risk of uncontrollable catastrophes.
5.4 Durability
5.5 Applying “then what?” Case A - Frontier training
Here we again consider a coordinated ceiling on frontier training compute, applying initially to participating developers.
Adaptation and redirection: we can expect that, as resources that can be devoted to pretraining runs are restricted, more will flow to unrestricted areas, such as synthetic data generation or greater inference-time compute.
Actors may also try to circumvent the rules in some way, by spending compute on areas not strictly covered by the rule definitions but which nonetheless have the form of improving training, or by simply trying to hide how much compute they spend on training by aggregating runs they present as independent or doing it in jurisdictions with less oversight.
Ultimately all of this amounts to how easily the rules can be gamed or circumvented, and where the otherwise training focused capital, compute and research talent are displaced to.
Machinery: machinery could include on-chip hardware tracking, or audited compute usage records. Who has access to this commercially sensitive information will likely be a point of negotiation.
Costs and incidence: the greatest but perhaps least legible costs here are forgone capability improvements - which general purpose capabilities will arrive later or not at all, and who they will effect. There are also the implementation costs themselves, whose incidence is likely distributed across the value chain. Other costs include to individual actors - those with the best ability to use extra training compute will be harmed more than others’ whose expertise lies elsewhere.
Actors who circumvent or are not covered by these rules will gain at the expense of the compliant, for example developers of smaller models, or Chinese labs, in the case of an agreement China is not party to or one with compute bounds higher than Chinese labs’ compute budgets.
To the extent a slowdown allows less fierce racing, incumbents may relatively gain by achieving profitability sooner. Incumbent compute providers may suffer if total compute demand is reduced, or gain as entry into a now more regulated sector becomes more difficult.
Political durability: Incentives to defect may increase over time as overhangs become greater and non-participants catch up or pull ahead. The latter will also weaken the case for continuity. This tradeoff should be considered when setting the initial scope. To the extent progress or societal changes remain quite fast such that the slowdown is welcome this may promote durability.
5.6 Applying “then what?” Case B - Biological capability
Here we again consider a red line under which models that materially uplift biological-weapons capability cannot be deployed or widely released without restrictions.
Adaptation and redirection: The primary adaptation to be expected is using alternative, out of scope or uncovered models, and secondly jailbreaking powerful models.
A further redirection is simply for actors to move away from biological capability to other threats, such as cyber or chemical weapons.
Machinery: The mechanism for capability evaluation is a key part of the machinery, as is who has final decision making authority. If selected parties are to be given access, the mechanism for choosing them and monitoring their behaviour also needs to be specified.
Costs and incidence: Again forgone beneficial capabilities gains, here in the form of progress in biology such as drug development, are the primary cost. There are also implementation and governance costs, and forgone revenue for model providers to account for.
If only large powerful actors have access to the most capable models, there are also costs to smaller research organisations, who may find themselves cut off from the most powerful tools in their fields and unable to make significant contributions.
Political durability: States will be relatively easy to keep invested in the continued restriction of diffusion of dangerous biological capabilities. Nonetheless enforcement over open models may prove difficult, and if capabilities diffuse at all via this or another route it may prove impractical to fully enforce a moratorium on these capabilities.
5.7 Open Research Questions
-
What forms of compensation or consideration could make pacing interventions preferable to larger coalitions?
-
What are the tradeoffs between quality of verification and degree of invasiveness for different interventions, and how can we push forward the frontier?
-
Which restrictions are more brittle, and which cause more damage if they break suddenly, via capability overhangs and other mechanisms? Which redirect efforts towards safer paths instead without building up an overhang?
-
Which actors gain relative power under different interventions, and what are the expected consequences of this?
-
What expertise might be lost during a slowdown which we would rather preserve, and by what means can we do so without weakening the intervention itself? E.g. the process for making FOGBANK, a key input into certain nuclear warheads, was lost at one point, and tens of millions of dollars spent on rediscovering it. Cost overruns in new nuclear plants have been blamed on a lack of skilled workers, due to not building any for many years.
-
What safeguards can be deployed to guard against mission creep, where regulators or newly empowered authorities could gain power beyond what was intended and become hard to dislodge?
Conclusion
Making the right choices about the pace of AI development will be critical for everything from national security to public health — indeed, the choices we make in the next few years may well ripple out for centuries or more. It is crucial that we do what we can now to make those choices go well. And though these choices are often inherently political, that is all the more reason to also want sober, dispassionate reckoning with the key considerations.
Ultimately we will need more than plans, and more than a rich understanding of any particular area: we will need the ability to effectively adapt to a rapidly changing world, connecting many distinct considerations into a coherent whole, recognising and navigating the increasingly sharp tradeoffs between some of our most deeply-held values.
The world will pace progress one way or another. Absent better tools, it might do so haphazardly: through improvised reactions that overfit on prior expectations, in ways that fail to actually address risks, with institutions that outlast their use. Our hope is that, with the proper research, pacing can become progressively more deliberate — targeted, proportionate, decisive, and legitimate.
This paper has endeavoured to present some initial analysis, and to sketch out the many other crucial questions which could benefit from further research. We hope to see others building on this, or critiquing it as appropriate.
Unlike fields with the luxury of time and stationary objects, research into pacing will fundamentally be a matter of triage. We do not have the luxury of considering all these questions to our satisfaction before we have to present our answers. We would do well, then, to make use of the time we have.
Call To Action
If you are interested in working on any of the open questions we lay out above, we’d love to hear from you. We’re also interested in critiques of this piece.