AIs don't do what you want. This is bad
Reward Hacking in the WildThe writeupGitHubYour AIs don’t do what you want. The numbers above cover the published subset (excludes AIID and X, confidence >= 0.9). X posts and AI Incident Database records are collected but not republished here: X expects posts to be embedded rather than their text rehosted, and AIID is share-alike licensed.
- ▪Reward Hacking in the WildThe writeupGitHubYour AIs don’t do what you want.
- ▪The numbers above cover the published subset (excludes AIID and X, confidence >= 0.9).
- ▪X posts and AI Incident Database records are collected but not republished here: X expects posts to be embedded rather than their text rehosted, and AIID is share-alike licensed.
Hacker News (Front Page) files mainly under programming. We currently carry 553 of its stories. Top-voted stories on Hacker News.
Opening excerpt (first ~120 words) tap to expand
Reward Hacking in the WildThe writeupGitHubYour AIs don’t do what you want. This is really bad3,607user-reported incidents of AI agents misbehavingRead the writeup >Search the corpusI’m feeling luckyloading…The numbersovereagerness1,566 43.4%other misalignment1,555 43.1%destructive actions622 17.2%sycophancy328 9.1%unauthorized access237 6.6%reward hacking217 6.0%metric spoofing87 2.4%excessive exploration84 2.3%unauthorized communication73 2.0%credential misuse49 1.4%test tampering45 1.2%self modification24 0.7%hidden backdoors15 0.4%Incidents are multi-label (one report can be both a destructive action and overeagerness), so the category counts sum to more than the 3,607 total.How bad were theynegligible: 1,468 (40.7%)minor: 1,373 (38.1%)significant: 618 (17.1%)severe: 121…
Excerpt limited to ~120 words for fair-use compliance. The full article is at Reward Hacking in the Wild.