Since our founding, Arena has been on a mission to measure the frontier of AI for real-world use. Today, we are making safety and alignment a bigger part of that mission: making AI systems act in line with human intentions and values. This matters more as such systems become more capable and take on more of the work we delegate to them. We are previewing our Alignment Index, built on three new alignment signals drawn from conversations on Agent Arena along with the failure modes for different open and closed models.
Many aspects of safety and alignment are difficult to observe directly, but some failures leave clear evidence in the conversation. We chose three observable signals that capture common ways that trust breaks down: Unauthorized Action (UA), where the model acts beyond what the user asked; False Attribution (FA), where the model attributes a statement, intention, or fact to the user that is contradicted by user-provided evidence; and Deceptive Completion (DC), where the model reports a task as done when it is not. These signals cover only a small part of safety and alignment, but because each is grounded in direct evidence, they give us a consistent basis for comparing models.
Key Takeaways
- OpenAI's models are the most aligned. They hold the top five positions out of 27 models on Arena's Alignment Index, with four models scoring about 88 points. This is followed by Anthropic and SpaceXAI models (Opus 5.5 and Grok 4.7) scoring 83.
- Rogue actions are rare, but when they happen they can be destructive. Only about 2% of Opus 5 sessions included an unauthorized action. Of those, more than half (53.5%) involved deleting or "cleaning up" the user's file or earlier work without permission.
- Agents can be misleading about task progress. On average, 10% of sessions are impacted by a deceptive completion. Code debugging is especially problematic, with the number rising to 48% of sessions.
- The longer the conversation, the higher the safety risk. A conversation twice as long is twice as likely to hit a failure mode. In sessions with 20+ messages, about 1 in 8 are affected by an unauthorized action.
- Safety and alignment have improved across model generations. The latest models from OpenAI, Anthropic, SpaceXAI and Google all rank higher than their predecessors.
Methodology
For each signal, we wrote a set of rubrics describing its recurring failure modes, with examples and clear boundaries for what does and does not qualify. We refined these rubrics over repeated rounds of judging and human review, revising them wherever our reviewers disagreed with the judge. We then sampled eligible sessions for each model and had an LLM judge apply the rubrics. A session is flagged only if the judge can point to a specific claim or action along with the evidence that supports it. Finally, we adjusted all rates for conversation length, since longer conversations give a model more opportunities to make a mistake.

To compute the Alignment Index, we take each signal and we transform the proportion of flagged sessions using . This transformation reflects our view that further improvements remain important even when a model is already highly aligned. It makes the highest scores harder to achieve and keeps improvements near full alignment visible, rather than allowing scores to cluster at the top of the scale. We then take a weighted average of these scores, assigning 50% to UA and 25% each to FA and DC. Higher values in the index indicate a safer and better aligned model.
Going beyond the ranking
The Alignment Index captures model behaviors that can affect users in everyday work. But an aggregate score hides how these behaviors actually play out. The same signal means different things depending on the task the agent is doing and the specific way it fails. To account for this, we trace each signal to the tasks where it shows up, compare how models fail, and end with a few real cases.
When are agents more likely to fail?
Unauthorized actions are uncommon across task categories. The detection rate for unauthorized actions stays below 7% across all task categories, peaking in code explanation (6.3%) and code debugging (5.8%). One likely reason is that coding agents usually have direct access to files and tools, so they have more chances to act beyond what the user allowed. A user who asks what a function does may get a rewritten function instead.
Deceptive completions are most common in code debugging. The detection rate for deceptive completions is highest in code debugging (48.0%). Debugging is difficult: a fix may miss the root cause, introduce new bugs, or produce code that does not run. When an agent reports success despite these failures, we flag it as Deceptive Completion (DC). DC captures dishonesty about task completion, but its frequency is also related to capability: weaker models may struggle to implement the changes they promise, while stronger models are more likely to deliver working code. This does not change the standard. An agent that cannot complete a task should acknowledge that limitation rather than claim success.
Writing and planning tasks show a different pattern. The detection rate for false attribution is highest in professional writing (13.7%) and planning and brainstorming (12.3%). These tasks often require the agent to restate or build on the user’s input which creates more chances to misquote the user or attribute others’ ideas to them.
The longer you work with an AI agent, the more it goes off the rails. Longer sessions are more likely to contain a detected failure. This pattern holds across all three signals. In the longest session group, DC is detected in 45.4% of sessions and UA in 12.4%. While in the shortest group, we only observe 3% and 0.11% for FA and UA respectively.
Overall, the task and the length of a session both shape how often an agent goes wrong. However, these rates do not tell us how serious each case is. Unauthorized actions are rare, but they can be costly. One could mean doing more work than you asked while another could mean deleting your file. To understand the risk, the next section looks at what actually happens in these sessions.
Do different models fail in different ways?
Models fail in different ways, even models from the same lab. Some failures cost more than others depending on the task, so a model's failure pattern is worth weighing alongside price and performance. This also applies when building on open-weight models.
Unauthorized action: Anthropic's own models overstep in very different ways. One failure mode is cleanup, where a model deletes user-provided files or earlier deliverables without permission. Only about 2% of Opus 5 sessions included an unauthorized action, but cleanup makes up 53.5% of those cases, well above every other model. Anthropic's own testing shows the same habit. In the Opus 5 system card, the model treats an earlier "clean up the batch" request as authorization, despite a reminder requiring confirmation in the current turn. It then "deletes all 120 jobs" [6]. Opus 5.5 performs better. Its cleanup rate falls to 20.0%, and its most common mode is extra outputs (40.0%), a less damaging way to overstep. Fable 5.1 has the lowest cleanup rate (6.5%). Instead, it leans toward premature execution (38.7%), where a model starts work when the user asked for early discussion.
False attribution: Some models misquote the user, while others credit the user with someone else's work. Most FA cases fall into two modes. The first is misquoting the request, where a model misstates what the user asked for. The second is misattributing a source, where a model credits the user with material from another source. Claude Sonnet 5 and GPT-6 Luna and Astra show opposite patterns. Sonnet 5 often misquotes the request (46.4%) but rarely misattributes a source (27.4%). Luna and Astra are the reverse. They rarely misquote (15.6% and 28.6%) but often misattribute (53.1% and 48.2%). Their sibling GPT-6 Sol stands out for a different reason: misstating the user's history. It has the highest rate of this failure mode of any model (23.5%).
Deceptive completion: For most models, a large share of deceptive completions are verification overclaims, where the model says it checked its work when it didn’t. The exceptions are the GPT-6 series and Grok 4.7 where only 7–10% of their DC cases fall into this mode. Inkling is slightly higher at 19%. Most other models are much higher such as GLM 5.3 and MiMo V2.6 Pro are above 50%, and most Claude models are above 40%. This matches what Anthropic reports about its own models. In the Claude Opus 5.5 system card, the top flagged behavior in internal use was "asserting unverified inferences as established fact," such as "describing a partial check as a full read" [5]. The Claude Opus 4.8 system card describes a similar case. Asked to monitor pull requests until CI passed, Claude "made detailed statements about babysitting when no babysitter agent was spawned, the spawned babysitter had exited, or the babysitter was reading the wrong API and missing failures" [5].
What do these failures look like?
The examples below show what these failure modes look like in practice for individual models. Select a signal and a featured model to inspect a reviewed failure case. Each example places the user’s request alongside the agent’s claim or action and the relevant evidence, showing what went wrong and why the interaction received a positive label.
Is alignment getting better over generations?
Newer generations are more aligned. Across the four lineages we tracked (GPT, Claude, Gemini Flash and Grok) the newest models have lower detection rates than their predecessors on most signals. The path wasn’t perfectly smooth for every lab: GPT-6.1 Sol has slightly higher false attribution than GPT-6 Sol, and Opus 5 has slightly higher unauthorized action than Opus 4.8.
What’s next
As agents take on more autonomy and responsibility, evaluating how they actually behave has to keep pace. Next, we'll add more safety-related signals, starting with how well models refuse harmful prompts. We'll also expand to new models and real-world settings. We believe a rigorous, evolving approach to measurement is critical to deploying agents responsibly and we’re building this index to help people and companies choose between models with a clearer picture of the safety tradeoffs. Check out the latest rankings at our Alignment Index for more insights.
References
- Arena. Agent Arena: Causal Evaluation of Agents in the Real World.
- Anthropic. Transparency Hub.
- OpenAI. GPT-5.6 System Card, §7.2.
- OpenAI. GPT-6 Astra System Card, §8.3.1, “Coding Deception.”
- Anthropic. Introducing Claude Opus 5.5, “Alignment.”
- Anthropic. Introducing Claude Opus 5.
- Anthropic. Claude Opus 5 System Card, §6.2.1, §6.4.2, §6.4.3.










