I. The Limits of Automation in Foreign Policy
This section examines why applying AI to statecraft is harder than applying it to games like Go or chess. It identifies three challenges that efforts must confront: representation, inference, and evaluation –– and addresses the risks of both inaction and premature automation. It concludes by outlining the argument for augmenting statecraft at the task-level, designing and evaluating tools for specific parts of the process, rather than monolithic systems that attempt to automate diplomacy end-to-end.
For this whitepaper, statecraft refers to the construction and execution of strategies and tactics through which political actors shape their external environment and influence other actors in pursuit of political ends (Kaplan 1952; Schelling 1960). Strategies tie political objectives to instruments and theories of change, while tactics are the moves through which these are deployed (von Clausewitz 1832). In practice, statecraft operates through available repertoires of practices (Adler and Pouliot 2011; Goddard et al. 2019). Recent work has shown how statecraft is increasingly conducted as much through institutions and transnational networks as it is with direct interstate bargaining. Diplomacy is then best understood as an institutionalized repertoire within statecraft, rather than its synonym.
Statecraft is treated as multiple repertoires of practices rather than a single problem (Goddard et al. 2019). This makes computational augmentation most defensible when aimed at supporting a range of specific practices, as opposed to the whole process.
Large language models (LLMs) may construct coherent logics and strategies of statecraft, but they are also subject to the considerations outlined above –– this means they cannot be verified, or ‘trained’, on correctness. These models are strong at synthesizing information and proposing linguistically coherent strategies, but proposals must be verified against some ground truth. If outputs are produced faster than institutions can verify them, then there is a risk from both processing bottlenecks and the integration of unchecked analysis into decisions.
A Field in Formation
AI is already entering the practices of statecraft. Ministries and foreign offices are moving quickly to deployment, with government agencies building tools: the U.S. State Department has deployed StateChat, and the U.S. Army has built CamoGPT (Satter et al. 2025; Ringquist 2025). A 2026 survey of 485 negotiation practitioners found AI tools embedded in preparation, analysis, and drafting; practitioners are already reporting concerns over confidentiality, over-reliance, and a loss of judgment (Bruderlein 2026).
Policy attention, by contrast, has focused on AI as an object of statecraft –– chips, diffusion, and governance (Sullivan and Feldman 2026). Its arrival as an instrument of statecraft has proceeded piecemeal, largely without systematic evaluation, but research is converging on this space along three main tracks.
First, model behavior is evaluated in closed-ended foreign policy scenarios or compared against the choices of national security experts (Rivera et al. 2024; Lamparth et al. 2024; Payne 2026). These studies reveal the tendencies of language models placed in strategic decision environments, but do not theorize which tasks should or should not be undertaken by models.
A second branch of research extends the game-playing tradition, building agents that can participate in multi-party bargaining games and negotiation testbeds, most prominently the board game Diplomacy (Meta Fundamental AI Research Diplomacy Team (FAIR)† et al. 2022). Relatedly, AI-mediated deliberation has been found to produce group statements that participants prefer to those of human mediators (Tessler et al. 2024). A final branch studies institutions: digital diplomacy scholarship has started mapping AI-driven practice and practice-oriented work, particularly in peace mediation and negotiation (Manor 2026; Hirblinger 2022).
The older empirical tradition of geopolitical forecasting shares structural components with this whitepaper’s proposal. Forecasting tournaments, such as those conducted under IARPA’s Aggregative Contingent Estimation program, scored probabilistic estimates of political events against explicit resolution criteria and demonstrated that training, teaming, and aggregation could improve accuracy, producing measurable gains over unaided expert judgment (Mellers et al. 2014; Tetlock and Gardner 2015). This work supports a central premise of this whitepaper: that subtasks of strategic judgment can be effectively isolated and evaluated against external criteria, supporting improved team performance.
Despite this innovative work, the field lacks a connective layer linking task-level activities to computational methods and to standards of evaluation. Benchmarks have relied on stylized, high-level scenarios that lack the complexity of diplomatic work, while agents have mastered games whose mastery does not transfer to statecraft –– as this whitepaper will explore (Jensen et al. 2025). The field is rich in promising findings, but those findings lack placement within institutional practice and shared criteria of evaluation.
This paper outlines that connective layer, grounding technical research in real-world practice. The taxonomy is designed to map decomposed tasks onto method classes and evaluation criteria, providing a common architecture in which technical tools can be placed within the repertoires of statecraft. As the capabilities of AI systems continue to evolve, the practice and study of world politics need methods to establish the value of each new tool.
The Boundary of Formalization
Game theory and bargaining theory provide the most mature formal tools for strategic analysis in institutional settings. Foundational work on bargaining, deterrence, incomplete information, and two-level games has shaped the study of diplomatic strategy and conflict (von Neumann and Morgenstern 1944; Nash 1950; Harsanyi 1967). Decision analysis complements these traditions by focusing on how actors should make choices under uncertainty, using subjective probabilities, expected utilities, and explicit value tradeoffs, allowing actors to generate practical guidance amidst complexity (Lax and Sebenius 1986; Keeney and Raiffa 1993; Raiffa et al. 2002). In contrast to approaches oriented mainly toward equilibrium characterization or descriptive explanation, decision analysis is explicitly prescriptive. It breaks complex strategic problems into tractable components, including objectives, uncertainties, payoffs, interdependence, repeated interaction, and institutional design under specified constraints.
The Automated Negotiating Agents Competition has advanced research on bilateral multi-issue negotiation since 2010, producing increasingly sophisticated agents in structured settings (Baarslag et al. 2015). Work on the Human-Agent League and related platforms has clarified the gap between agent-agent benchmarks and negotiation with human counterparts — a gap that has narrowed but still remains informative (Mell et al. 2018).
Recent AI achievements illuminate where the boundary of reliable formalization lies. Meta’s CICERO system achieved human-level performance in the game Diplomacy by combining a dialogue model with strategic planning and reinforcement learning (RL), providing us with evidence that language-conditioned strategic agents can operate in structured negotiation games (Meta Fundamental AI Research Diplomacy Team (FAIR)† et al. 2022).
While CICERO was a significant achievement, the gap between board games and statecraft is instructive. The game Diplomacy has fixed rules, known players, a discrete action space, and a well-defined objective function; CICERO did not have to grapple with the central problems of formalizing actual negotiations (representation, inference, and evaluation). Real diplomatic settings superficially resemble well-structured games, but their complexity lies in the features that board games hold constant.
As LLMs produce increasingly sophisticated analyses, the central question becomes verification: systematically interrogating whether a given contribution supports an actor’s goals. In closed domains, the correctness of broad strategic recommendations can be operationalized ex ante through externally fixed objectives and rules and computed positional advantage. For example, a chess engine’s recommended move can be evaluated using search and self-play, and, in some endgames, against ground truth. These evaluations are imperfect but stable because the rules and the objectives are fixed, allowing players to see whether a certain move brings the player to a more favorable position. Verification is possible because the evaluative criteria are both stable and external to the system.
Holistic, end-to-end diplomatic outputs (“automated” negotiators) currently resist definitive ex-ante verification for the reasons this section will explore. Strategic assessments cannot be checked against ground truth, preferences are strategically concealed, and the relevant counterfactuals are unobservable. The evaluative criteria for a “correct” recommendation are themselves politically contested and strategically manipulable.
Despite the limits of holistic strategic recommendations, many constituent sub-components can be verified. Components such as factual premises, logical coherence, coverage, and consistency can be validated before action is taken. This task-based decomposition can recover verifiability for individual components of a process whose aggregate output remains unverifiable in advance.
Task-based decomposition
The logic of decomposition has roots in systems theory: Complex systems are amenable to analysis precisely to the extent that they are decomposable into semi-independent modules (Simon 1962). Decomposing the processes of statecraft into its constituent tasks allows for verification (ex-ante and ex-post) at the subcomponent level. A taxonomic decomposition can recover partial verifiability by isolating the subtasks that have stable and external evaluation criteria. For example:
- A system that retrieves and synthesizes background research can be evaluated for accuracy or coverage against a documented range of sources;
- A system that generates negotiation options can be reviewed for range and plausibility by domain experts –– a weaker check than verification;
- A system that monitors compliance can be scored against specified, observable indicators.
In each of these examples, tractable evaluations are only possible because decomposition has sufficiently separated each task from the contested core or interconnected assumptions. If an AI system tries to take multiple steps at once, then assumptions over contested values, ground truth, or forecasting models can be carried through and compound over time. Each of these ambiguities requires accountable human judgment to resolve. The holistic ‘correctness’ of a broad end-to-end strategy cannot be evaluated like this.
Well-evaluated components can have errors that interact and propagate when recompiled at the system-level, particularly in the absence of human interventions. The main risk here is an interface error: a subtask may be locally well-evaluated while still producing outputs that distort a broader strategic process. In systems theory, decomposition is faithful when the interactions between subtasks are explicitly specified, allowing for each locally valid output to compose without distortion (C. Y. Baldwin and Clark 2000). The relationships between subtasks should be preserved during recombination, whether each input is pooled, sequentially continued, or reciprocally interdependent (Thompson 1967). The problem of faithful recombination is a central open challenge of the approach that we outline. As there is no guarantee that each verifiable sub-task will aggregate up to a verifiable system, it is essential that we have human teams both executing on tasks beyond machine capability and recombining each of the decomposed subtasks. Human oversight can provide risk mitigation, but no guarantees to correctness. Any designed system, however, should ensure consistent interdependencies across the (loosely coupled) network (Brusoni et al. 2001).
As models improve, AI systems will likely be able to competently handle a broader chunk of relevant subtasks, narrowing the remaining areas of strategic judgment. As such, institutions at the heart of statecraft should ensure that their systems can adequately preserve human judgment, visibility into the system, and mechanisms of democratic accountability amidst the increasing capability. Participants in world politics will need to maintain accountability over the irreducibly strategic, while building trust and institutional capacity incrementally.
Not all mechanisms for model evaluation are equal, with several distinct tiers of evaluability. This includes:
- externally checkable outputs, or indicators (this captures most research or monitoring tasks, with some scope for assessing the correctness of analysis),
- expert-auditable outputs, where reviewers can assess validity but without a ground truth to compare (this includes most analysis tasks or option generation processes),
- non-verifiable judgment calls involving contested value or unknowable predictions (this captures most strategic judgment calls, or live execution tasks).
Evaluability alone does not defend the decision to delegate a task. A system may produce accurate and auditable outputs while still making some judgment over contested values or objectives. Deployment decisions should balance both the evaluability of the output and the institutional setting within which the decision will be made.
Three Challenges
This section examines the resulting three challenges from the computer science and computational social science literature:
- The representation problem asks whether we can model the strategic environment with enough faithfulness to the underlying complexities of reality: whether a model’s states, transitions, and action sets can be formally captured in ways that support computation.
- The inference problem asks whether we can find out what we need to know: whether preferences, constraints, and beliefs can be uncovered from observable behavior.
- The evaluation problem asks whether we can define success: whether objectives can be formalized without sacrificing the contested normative and political dimensions at the heart of diplomatic practices.
The Representation Problem
To meaningfully apply computational methods to diplomatic settings, we must be able to represent strategic environments in formal terms –– that is, as variables, relationships and rules that AI models and systems can process. This requires specifying states (who the actors are, what they believe, what constraints they face), actions (what moves are available), and transitions (how actions change states, which states are likely to follow others).
Institutional endogeneity
The lack of consistent constraints in world politics means that game-playing computational approaches are often inapplicable; these methods treat institutional structures as fixed, like the rules of a game, as opposed to dynamic and non-binding. Institutions structure political and economic interaction by lowering transaction costs, reducing uncertainty, and constraining feasible actions (Coase 1937; North 1991; Keohane 1988). The rational design literature explains why states create particular rules as strategic responses to cooperation problems (Koremenos et al. 2001). This scholarship also demonstrates that institutional design is itself strategic as actors anticipate how rules will constrain future bargaining and design accordingly.
The rules governing diplomatic interaction are endogenous variables, with negotiations creating, modifying, and dissolving the institutions within which they occur. For example, the 2015 Paris Agreement's ratchet mechanism for progressively tightening climate commitments was itself a negotiated institutional innovation that changed the structure of future climate diplomacy. The evolution of WTO dispute settlement procedures, the creation of ad hoc coalitions outside formal treaty structures, and the erosion of arms control regimes all illustrate the same dynamic: Institutional rules are part of what is being contested and (re)constructed.
The WTO Appellate Body crisis illustrates how international institutions can have their own rules endogenously transformed by participating states. Since 2019, the WTO Appellate Body has been unable to properly function due to losing its three-member quorum.Although there are formal rules for the system, the scope of those rules has evolved through strategic non-cooperation. This is a move that sits outside of the institution’s founding “rules” but has transformed the space of potential subsequent moves. An AI system that searches over allowable institutional moves at its initiation, like in a board game, would not have discovered this.
Action space complexity
Unlike games with externally specified legal moves, statecraft has no bounded action set. Systems that combine reinforcement learning, search, and language modules have reached human-level performance in complex games such as chess and Go, as well as video games like Atari and StarCraft II (Mnih et al. 2015; Silver et al. 2018; Vinyals et al. 2019). But the legal moves in these games are discrete and externally specified.
In diplomatic settings, a negotiator can propose novel treaty language, invoke new precedents, or take actions outside the formal negotiation. The action space is not usefully enumerable ex ante. Institutional context, political feasibility, and cognitive constraints impose practical bounds that formal models can only imperfectly approximate.
This complexity of the action space has formal consequences. Pruning the action space to a manageable set of appropriate moves is possible, but it must be informed by domain-specific institutional knowledge rather than derived from game-tree statistics alone. Computational systems do not have to enumerate all possible actions. Instead, they should be able to generate a diverse subset of strategies that includes unconventional options an experienced team might overlook.
In world politics, unconventional moves sometimes succeed because they violate expectations (Jervis 1976; Handel 1981). AlphaGo’s celebrated move 37 was valuable because it lay outside the space of moves humans would consider but still within the space of possible moves (Silver et al. 2016). Any such pruning must distinguish moves that are genuinely unpermitted from those that are surprising but feasible.
Multiparty complexity
Much of the canonical game-theoretic work on crisis bargaining begins with two-actor models, with multilateral negotiations making the strategic structure more fluid (Schelling 1960; Fearon 1995; Zartman and Berman 1982; Odell 2000). Complexity comes from the contested and evolving strategic structure, or “non-stationarity”, as well as the increased number of players. Changing issue-based alignments are often contested and evolving; computational techniques, like multi-agent reinforcement learning, handle such fluidity poorly (Littman 1994; Shoham et al. 2007; Busoniu et al. 2008; Hernandez-Leal et al. 2019; Chalkiadakis et al. 2022).
Limited stationarity and path dependence
Even when institutional rules and actors are fixed, negotiation arrangements may differ due to the intervening path dependencies. Relationships accumulate history over time, with reputations and past concessions constraining future credibility (Axelrod 1984). Repeated game models capture how the shadow of the future shapes present choices, and reputation models formalize how past behavior constrains beliefs about future conduct. But these models typically assume stable preferences and fixed game structures. Many standard reinforcement learning and off-policy evaluation methods assume that transition and reward distributions are stable enough for past data to inform future decisions. When preferences, institutions, relationships, and issues evolve together, diplomatic settings challenge these stationarity assumptions.
Any formal model of a strategic diplomatic environment will either respect these properties and pay the resulting complexity costs or violate them and pay in fidelity. There is no modeling strategy that avoids both costs simultaneously. The practical implication is that formalization is possible but inherently partial. Formal models will cover some subset of the state space with high fidelity and degrade outside that subset. We will not build a complete model for world politics. We can, however, characterize the boundary of each tool’s validity with sufficient precision for practitioners. This is a problem that asks for calibrated incompleteness: models that know what they don’t know.
This representation problem is a domain-specific instance of the general knowledge representation problem: how to represent a complex, changing world in a form that supports automated reasoning (Davis et al. 1993). The specific difficulties of representation in international relations correspond to known hard problems in computer science. Each of these has a substantial technical literature that this project can draw upon and extend.
The Inference Problem
Even if it were possible to adequately model the strategic environment, we would face a second challenge: inferring the parameters –– the key factors –– that shape diplomatic behavior and outcomes. To provide useful decision support, a system should represent and update key beliefs, for example, about what actors want, the constraints they face, or what they believe about each other (Harsanyi 1967; Lake and Powell 1999). In principle, these can be inferred from observed behavior. In practice, diplomatic behavior is often designed to both reveal and conceal, communicating credibility and resolve while maintaining bargaining advantage and strategic ambiguity.
Strategic misrepresentation
The foundational puzzle of rationalist theories of conflict is that war is costly and inefficient but still occurs frequently (Fearon 1995). Explanations for this puzzle include private information, incentives to misrepresent, commitment problems, and issue indivisibility, with later work incorporating time horizons (Toft 2006). Importantly, the existence of a nonempty bargaining range is not automatic; if the stakes are perceived as existential or values genuinely incompatible, then mutually acceptable settlements may not exist at all (Fearon 1995). For this paper, the key implication is that the behaviors an AI system must interpret in the conduct of statecraft are often strategically constructed to obscure the information that inference would require.
Preferences can be inferred computationally from observed behavior, but current algorithms are narrow in statecraft-relevant applications (Ng and Russell 2000). This is done through techniques like inverse reinforcement learning (IRL), which depends on strong assumptions about the environment, actors, and their rewards. These techniques assume that observed actions reveal underlying reward functions. In canonical IRL problems, the goal is to recover a reward function under which observed behavior is optimal or near optimal. This is often underspecified as the same behavior could be optimal for many objectives. In the conduct of statecraft, particularly during diplomatic bargaining, this identification problem is compounded because the observed behavior is itself strategic (Kydd 2005; Sartori 2005; Jervis 1976). Because actions are selected to influence future outcomes, observed behavior is mediated by incentives and constraints, rather than directly revealing beliefs or preferences. Understanding these complications is central for building effective systems that support the practices of statecraft.
The dynamics of statecraft can be more accurately characterized through multi-agent IRL or inverse game theory under asymmetric information. In cooperative settings, the informed actor may have incentives to teach or reveal the relevant reward function, allowing for the function to be inferred (Hadfield-Menell et al. 2016). Diplomatic bargaining often has the opposite structure: actors may have incentives to mask or manipulate beliefs through ambiguity. Non-cooperative settings are more difficult, but reward functions can sometimes be partially recovered in stylized cases, even when one agent obscures its objective from another agent (Zhang et al. 2019).
During the Iran nuclear negotiations that culminated in the 2015 JCPOA, the parties had to infer genuine red lines from stated positions. Outside observers and counterparties had to distinguish genuine constraints from negotiable positions. Iran’s insistence on enrichment capacity could be a constraint reflecting domestic politics and technical path-dependence, or a bargaining position designed to extract concessions elsewhere (Samore 2015). Computational systems cannot resolve this ambiguity by collecting more passive data because the ambiguity is deliberately generated by the actors themselves (Jervis 1976). The difficulty in deciphering intent arises because actors have incentives to distort signals.
As AI capabilities increase, new mechanisms will be required that adapt to new technical capabilities to incentivize cooperation. This shifts the problem from statistical inference to mechanism and institutional design. Instead of attempting to recover information from more data alone, parties need arrangements under which informative signals are costly to fake or are beneficial to reveal.
On defined forecasting benchmarks, data-driven models can outperform physics-based numerical models (Lam et al. 2023). Commitment credibility, however, faces obstacles that weather does not: Actors strategically manipulate the signals a prediction system would need to learn from. Additionally, the deployment of such systems would alter signaling behavior, and the relevant cases to learn from number in the hundreds rather than billions.
Endogenous win-sets
Putnam’s two-level games framework is foundational to diplomatic analysis, capturing how international negotiations are constrained by domestic ratification requirements (Putnam 1988). What Putnam characterizes as the “win-set” is the set of international agreements that would be domestically ratifiable and, therefore, determines what deals are feasible. In practice, leaders also try to shape the sets through narratives, meaning that win-sets are themselves endogenous, strategic instruments that leaders expand or contract based on bargaining incentives.
Leaders may invoke domestic constraints to extract concessions at the negotiating table, or manufacture flexibility to enable deals that serve their interests. For example, the U.S. movement in and out of the Paris Climate Agreement illustrates the role of domestic political constraints. The 2017 announcement, 2020 withdrawal, 2021 reentry, and 2025-26 withdrawal each shifted beliefs about the durability of American climate commitments across administrations. Any estimation procedure must contend with the fact that observed boundaries of win sets are also strategic signals.
Data constraints
The underlying data for effective inference are uneven: rich in some domains (trade negotiations or UN voting) and thin in others (crisis diplomacy or backchannel communications). High-stakes negotiations occur rarely, with few cases comparable to the Cuban Missile Crisis, Camp David, or the Iran nuclear talks. The available data is also selection biased. Observers see only negotiations that reached public stages and have little evidence from failed negotiations or exploratory contacts.
In our preliminary interviews conducted for this project, senior practitioners suggested that negotiation success often depends on personal relationships that operate beneath the visible surface of state interests (Ancheva, forthcoming). This whitepaper focuses on process-level augmentation that underpins both the substantive analysis and the relational work — though the latter remains human-executed. The interpersonal dimension is the highest-risk area for automation and is where personal accountability and relationship management will remain critical.
Practitioners also have examples from their own experience that never entered the documentary record. This knowledge cannot be readily retrieved because it was never institutionally stored (Ancheva, forthcoming). The unrecorded judgments of diplomats include a range of subtle interpersonal and inferential skills. These skills might include recognizing a bluff, or an intuition on a shift in political momentum. Interpersonal expertise is consequential but invisible to systems that learn from documents. The experience of each individual diplomat –– every win or rebuke over the course of a career–– is a far richer learning signal than the outcome-level data or transcripts that we might be able to access after the fact.
The Evaluation Problem
Even if one solves the representation and inference problems, defining and evaluating success across multiple participants remains difficult. Classic optimization techniques need some objective that can be represented mathematically, such as profit or time, to then maximize or minimize. Modern optimization methods are sophisticated, but diplomatic outcomes resist evaluation in these terms for reasons that are normative and political, as well as technical.
Defining Success
A government may have many overlapping, often competing, objectives in a negotiation. Security, economic welfare, international status, domestic political survival, and normative legitimacy may all feature simultaneously. These may be weighted differently by different actors within the same government, and the weights are constantly shifting as circumstances evolve (Allison and Zelikow 1999).
Allison and Zelikow’s analysis of the Cuban Missile Crisis demonstrated that even a single government’s preferences reflect multiple lenses. The “rational actor” model treats the government as a unitary optimizing actor; the “organizational process” model captures how routines shape options; the “bureaucratic politics” model captures how competing internal interests produce policy (Allison and Zelikow 1999). Because no single interest or objective is complete, synthesizing them into a comprehensible strategy remains a matter of qualitative judgment. There is no single objective function for an AI system to optimize as a holistic strategy, because there is no single actor, but locally specified metrics for scoped tasks remain possible.
Quantitative tools can serve many valuable functions in this process but cannot determine which contested political values should dominate when facing conflicting objectives. In this setting, tools may clarify tradeoffs or expose inconsistencies, but there are other constraints that prevent them from providing concrete answers. These include normative constraints, issues of legitimacy, and multiple audiences with strong —often conflicting — preferences.
Normative constraints and legitimacy
In a negotiation, some options are off-limits for reasons beyond strategic costs, particularly in peace agreements and legalization (Chayes and Chayes 1995; Fortna 2004; Abbott et al. 2000). Violations of sovereignty, human rights, and international law may be strategically advantageous and still be impermissible for a given actor. Across the table, one side’s moral sin might be another’s calculated trade-off. Within a single team, one person’s deepest value might be another’s bargaining chip. Legitimacy is normatively constructed and concerns both outcomes and their justifications. Any AI system will need to make these distinctions explicit as opposed to aggregating all normative concerns into a single number or bucket.
When facing normative trade-offs, the role of an AI system may be like Isaiah Berlin’s role of the moral philosopher: outlining costs and tradeoffs in each situation as opposed to giving a judgment. Berlin held that the moral philosopher's role is to lay out the values, issues, and forms of life in collision with one another, as opposed to adjudicating between them (Magee and Berlin 1978). This left both the choice and responsibility to each individual. When facing contested values, this should be seen as a best-case scenario for AI integration –– though better models may help make the facts clearer, or consequences sharper, the choice and the responsibility are non-delegable.
Multiple audiences
Diplomatic communications address multiple audiences, publicly and privately. A message that signals resolve to one audience may signal stubbornness to another; a concession that builds trust with a counterpart may provoke domestic backlash. With multiple audiences, diplomats handle this complexity with real-time judgment under uncertainty (Putnam 1988). Reducing it to an objective function would require artificially fixing contestable political choices.
These problems with evaluation must be seriously considered with any augmentation design. AI tools may not be able to fully specify contested values but must make the structure of the problem transparent. This involves identifying the tradeoff frontiers between competing objectives and flagging cases where a proposed strategy performs well on one metric only by performing poorly on another. This provides decision support in the literal sense: supporting the human act of deciding what to value by clarifying what each choice of values entails.
Potential Benefits and Costs
Scoped computational tools can deliver genuine value to augmenting the existing practices of statecraft; we posit four dimensions to assess this value –– speed, quality, range, and novelty. Achieving and validating these benefits will require systems designed with sufficient modularity and transparency to enable comparisons to benchmarks of existing activities.
- Speed. Research that currently takes weeks could be accomplished in hours with existing language model capabilities. Faster preparation increases capacity for analysis and strategy, conferring early advantages to parties that adopt appropriate tools (Dell’Acqua et al. 2023; Noy and Zhang 2023).
- Quality. Broader source coverage and systematic analyses could reduce the gaps and inconsistencies that impact preparation under time constraints. Realizing quality gains will require disciplined institutional processes, but early evidence suggests that human-AI collaboration can increase creative problem-solving (Boussioux et al. 2024).
- Range. Computational tools can evaluate more scenarios and surface more possibilities than human teams alone working under time pressure. This expansion of known options is where augmentation may add significant value.
- Novelty. AI systems may expand the option space by exploring precedent from prior negotiations, constraints faced, and future paths in ways underexplored by human teams. The possibility of “move 37” moments, where unconventional options are found outside of human experience, is real, albeit currently rare. We see decision-relevant novelty as the longest-horizon payoff of this research agenda. Although current AI methods may present unconventional options, systems cannot reliably certify specific options as better or worse when it would be most consequential. Identifying the value of such novelty depends on strong evaluation mechanisms and strategic judgment under immense uncertainty.
These dimensions of improvement should be scoped and measured against baseline performance of real-world human teams. Without benchmarking, factors like speed or accuracy cannot be accurately assessed, and claims of AI-augmented improvement may remain unfalsifiable. Baselines provide value to institutions, potentially revealing inefficiencies and inconsistencies that can be addressed independently of AI. Moreover, baseline data enables comparative evaluation across tools and institutions, allowing the field to distinguish genuine future advances from exaggerated vendor claims or confirmation bias.
The risks of premature deployment
The competitive dynamics of contemporary geopolitics create pressure toward deployment of AI tools. If adversaries or counterparts are automating diplomatic processes, the instinct to match their capabilities becomes difficult to resist (Horowitz 2018; Allen and Chan 2017). But deploying systems that appear competent while lacking substantive reliability creates specific harms, some of which are predictable (Pozniak and Sania 2026). The cause for caution is found throughout the adjacent empirical evidence. For example, off-the-shelf LLMs display erratic escalation patterns when making strategic decisions in simulated wargames, even in scenarios without conflict seeded (Rivera et al. 2024).
Mistakes in diplomatic contexts are costly. Errors, like misunderstandings or missteps, carry consequences with tangible human costs that can compound over time. Unlike consumer applications where errors are frequent and low-cost, diplomatic errors can permanently damage relationships or escalate conflicts.
The verification asymmetry expounded above compounds these risks. Current systems are trained to perform the surface features of competent analysis with linguistic fluency, presenting clear structure and appropriate citations. Outputs, therefore, may pass superficial review while potentially misjudging strategic dynamics in ways only domain expertise would detect.
Models are becoming incredibly capable at a range of qualitative knowledge tasks, such as complex reasoning, formal writing, and searching over large databases, though limitations still remain (Wei et al. 2023; Dell’Acqua et al. 2023; Liu et al. 2023). These capabilities are, however, uneven and failure modes can be difficult to predict. A model’s jagged competence profile is potentially dangerous if insufficiently understood and indiscriminately deployed. Systems that perform well on writing and summarization tasks may fail on adjacent tasks. Interpreting strategic signals and drafting commitments look mechanically similar to the summarization and formal writing that models do well, but they are far harder to generate, and to verify. Excess faith in a model’s output due to adjacent competencies could introduce blind spots and risk.
The disempowerment of human teams is an important long-term risk with early evidence from a range of literatures, but no data to date to support a strong conclusion. Across industries from healthcare to law, evidence on the benefits and risks of decision delegation and automation is conflicting (Bainbridge 1983; Parasuraman and Manzey 2010). There is a risk that practitioners may become dependent on AI-generated analyses. Without rigorous analytical work, individuals may lose the capacity to evaluate outputs, and eventually to perform the analysis themselves, as the judgment that made delegation safe erodes through disuse (Endsley and Kiris 1995). Organizations that automate extensively may find, when systems fail, that human expertise has been eroded. Such organizations may find they then lack the institutional ability to detect and correct errors.
When multiple parties rely on similar AI systems, sufficiently correlated errors can induce systemic risk. Adjacent domains have explored how similar AI systems can produce simultaneous misreadings, coordinated escalations, or collective blindness to possibilities outside the training distribution (Daníelsson et al. 2022). The diversity of human decisions, for all its inefficiencies, provides a hedge against correlated failures that homogeneous AI systems may not be able to replicate (Hong and Page 2004).
Adversarial pressure on the technology stack.
The analysis of this whitepaper treats counterpart strategic misrepresentation primarily as an inference problem, but it is important to note that AI deployments can function as an additional surface for adversarial counterparts to attack. AI systems may be deliberately targeted through malicious cyberattacks. With decision-support systems, these attacks might attempt to subvert agents themselves, potentially through ‘poisoning’ training data or open-source information to steer agents off-course. These approaches are well-established the adversarial ML literature (Vassilev et al. 2025). In this domain, however, undetected attacks could be devastating — akin to having a ‘double agent’ implanted as a key advisor. As such, any AI risk assessment should treat cybersecurity deployment as a primary consideration.
The risks of inaction
Despite these concerns, the risks of inaction are material. Parties that use AI tools to improve their analytical capabilities, expand their strategic range, and accelerate their response times may hold advantages over those that do not. This could confer benefits across a range of dimensions in adversarial settings.
Beyond competitive pressure, existing bureaucratic processes in statecraft are often cumbersome, slow, and prone to their own systematic biases. Groupthink, availability heuristics, anchoring on familiar precedents, and the clustering of analyses around salient examples are documented challenges for human teams (Janis 1972; Tversky and Kahneman 1973; Welch 2000; Clement and Tse 2005). Choosing not to implement appropriate AI tools means accepting the limitations of current practice while forgoing potential improvements provided by advanced computational tools.
There is an opportunity cost if AI tools are not actively studied for applications in diplomatic settings. If scholars and institutions abstain from developing AI augmentation, the field will be defined by actors with less concern for rigor, safety, or the preservation of democratic accountability. Responsible development requires engagement, not withdrawal. Internal pressures to modernize from an ill-considered perspective may result in premature adoption with institutional downsides.
These system-level constraints motivate the task taxonomy developed in Section II.