Models fail in different ways, even models from the same lab. Some failures cost more than others depending on the task, so a model's failure pattern is worth weighing alongside price and performance. This also applies when building on open-weight models.
Unauthorized action: Anthropic's own models overstep in very different ways. One failure mode is cleanup, where a model deletes user-provided files or earlier deliverables without permission. Only about 2% of Opus 5 sessions included an unauthorized action, but cleanup makes up 53.5% of those cases, well above every other model. Anthropic's own testing shows the same habit. In the Opus 5 system card, the model treats an earlier "clean up the batch" request as authorization, despite a reminder requiring confirmation in the current turn. It then "deletes all 120 jobs" [6]. Opus 5.5 performs better. Its cleanup rate falls to 20.0%, and its most common mode is extra outputs (40.0%), a less damaging way to overstep. Fable 5.1 has the lowest cleanup rate (6.5%). Instead, it leans toward premature execution (38.7%), where a model starts work when the user asked for early discussion.
False attribution: Some models misquote the user, while others credit the user with someone else's work. Most FA cases fall into two modes. The first is misquoting the request, where a model misstates what the user asked for. The second is misattributing a source, where a model credits the user with material from another source. Claude Sonnet 5 and GPT-6 Luna and Astra show opposite patterns. Sonnet 5 often misquotes the request (46.4%) but rarely misattributes a source (27.4%). Luna and Astra are the reverse. They rarely misquote (15.6% and 28.6%) but often misattribute (53.1% and 48.2%). Their sibling GPT-6 Sol stands out for a different reason: misstating the user's history. It has the highest rate of this failure mode of any model (23.5%).
Deceptive completion: For most models, a large share of deceptive completions are verification overclaims, where the model says it checked its work when it didn’t. The exceptions are the GPT-6 series and Grok 4.7 where only 7–10% of their DC cases fall into this mode. Inkling is slightly higher at 19%. Most other models are much higher such as GLM 5.3 and MiMo V2.6 Pro are above 50%, and most Claude models are above 40%. This matches what Anthropic reports about its own models. In the Claude Opus 5.5 system card, the top flagged behavior in internal use was "asserting unverified inferences as established fact," such as "describing a partial check as a full read" [5]. The Claude Opus 4.8 system card describes a similar case. Asked to monitor pull requests until CI passed, Claude "made detailed statements about babysitting when no babysitter agent was spawned, the spawned babysitter had exited, or the babysitter was reading the wrong API and missing failures" .