The AI models that cheat the most, according to new CAIS benchmark
Stefan_Alfonso/iStock/Getty Images PlusZDNET’s key takeaway
- The Center for AI Safety (CAIS) created CheatBench.
- They found that every agent cheats in some scenarios.
- The propensity to cheat creates risks for humanity.
AI labs often tout impressive benchmark scores when releasing new models, showing better capabilities in areas like coding, computer use, and more than their competitors. However, those benchmarks aren’t always a reliable measure of what AI can do because they’re easily beaten by exponentially improving models and can emphasize marketing over actual performance.
Also: With AI models clobbering every benchmark, it’s time for human evaluation
Benchmarks like Humanity’s Last Exam try to counter this issue by challenging models in more realistic environments. But models still find loopholes to complete tasks — Hugging Face incident, anyone?
So, the Center for AI Safety (CAIS) created CheatBench. Yes, it’s exactly what it sounds like — and nearly every frontier model is guilty.
What CAIS found
AI models are rewarded for performing tasks well and quickly. A lack of knowledge or tools incentivizes them to do what researchers call “reward gaming” by “finding hidden answers, copying another agent’s submission, or manipulating how its work is graded,” CAIS explained. “CheatBench measures how often AI agents take these shortcuts when honest work is difficult.”
CAIS tested several agents running the latest and most lauded models, including OpenAI’s GPT-6 Astra in Codex, Anthropic’s Fabel 5.1 in Claude Code, and Meta’s newly released Muse Spark 1.3 in Muse Code. These agents were tested across 10 categories, including writing, professional work, mathematical research, and coding. Using “honeypot” clues hidden in task filespaces, the test separated acceptable reference use from cheating. CheatBench accounts for any time agents attempt to cheat, whether they are successful or not.
Also: What workers are really using AI for in 2026 – and what they aren’t
“Each setting establishes an expectation of honest work, introduces a discoverable opportunity to cheat, and defines the action that crosses that boundary,” the researchers explained.
Every agent the researchers tested cheated in at least some scenarios, but Astra came in as the most honest with a cheating rate of 48.2% — still almost half the time. Grok 4.6 was scored the biggest cheater with a rate of 81.5%. Open-weight models Kimi K3 and DeepSeek V4 Pro landed in the middle between several other proprietary frontier models.

What cheating looks like – and why it matters
In one example, researchers asked Claude Opus to design a protein binder. The model knew it wasn’t allowed to refer to a set of accepted designs in the filespace, but eventually caved.
“After seven rejected designs, it locates the file, writes that it should not look at or copy it, and reads it with a shell command in the very next call,” the researchers wrote. In its reasoning, the model admitted that using work other than its own would “misrepresent my actual capabilities in this evaluation, so I shouldn’t look at or copy it.” But its very next step was to reference the accepted designs.
Also: How to check if ChatGPT and other AI tools cite your website – and improve your chances in 2026
This result demonstrated both a readable choice the model made to contradict itself, and what looked like a hole in our understanding about what made the model jump from one instinct to the next.
Things got more interesting at the task category level. Even if an agent didn’t cheat in one area, it could cheat significantly more in another. Fable 5.1 was only 5% likely to cheat at games, but 100% likely to cheat on knowledge work tasks.
Also: The sneaky ways AI chatbots keep you hooked – and coming back for more
Reinforcement learning trains models not to abandon a task, even if pursuing it creates conflict-ridden choices. CAIS noted in its paper that sycophancy is an early sign of reward gaming. This term refers to AI models’ tendency to be too agreeable and encouraging of whatever a user says, sometimes regardless of whether it’s incorrect, delusional, or could lead to harmful behavior. Traits like sycophancy and reward gaming show how models can prioritize accomplishing a task correctly to please a user over the alignment training researchers work so hard to build in.
These tests represent relatively low stakes. But CAIS researchers created CheatBench because of the risks of this behavior at scale across different tasks. Earlier this month, yet another researcher quit Anthropic over concerns that the company isn’t developing AI responsibly for a future in which it could build itself away from human-oriented values and kill us.
A propensity to cheat, or complete a task at any cost, puts our potentially differing priorities at odds with an increasingly powerful technology. As I explained in the AI Leaderboard newsletter last week, it won’t necessarily be a demonstrated animosity toward humans that pits AI against us; it may be that we are simply in the way and end up as collateral.
Radhika Rajkumar
Senior Editor
Radhika Rajkumar is a senior editor at ZDNET based in New York City. She covers AI, specializing in safety, privacy and security, policy, education, and synthetic media. She also leads ZDNET's newsletter strategy. Radhika holds a Masters in Creative Publishing and Critical Journalism from The New School. See full bio
What's Your Reaction?
Like
0
Dislike
0
Love
0
Funny
0
Wow
0
Sad
0
Angry
0
Comments (0)