WeSearch

AIs don't do what you want. This is bad

·1 min read · 0 reactions · 0 comments · 2 views
#what#want
AIs don't do what you want. This is bad
TL;DR · WeSearch summary

Reward Hacking in the WildThe writeupGitHubYour AIs don’t do what you want. The numbers above cover the published subset (excludes AIID and X, confidence >= 0.9). X posts and AI Incident Database records are collected but not republished here: X expects posts to be embedded rather than their text rehosted, and AIID is share-alike licensed.

Key facts
About this source

Hacker News (Front Page) files mainly under programming. We currently carry 553 of its stories. Top-voted stories on Hacker News.

Original article
Reward Hacking in the Wild
Read full at Reward Hacking in the Wild →
Opening excerpt (first ~120 words) tap to expand

Reward Hacking in the WildThe writeupGitHubYour AIs don’t do what you want. This is really bad3,607user-reported incidents of AI agents misbehavingRead the writeup >Search the corpusI’m feeling luckyloading…The numbersovereagerness1,566 43.4%other misalignment1,555 43.1%destructive actions622 17.2%sycophancy328 9.1%unauthorized access237 6.6%reward hacking217 6.0%metric spoofing87 2.4%excessive exploration84 2.3%unauthorized communication73 2.0%credential misuse49 1.4%test tampering45 1.2%self modification24 0.7%hidden backdoors15 0.4%Incidents are multi-label (one report can be both a destructive action and overeagerness), so the category counts sum to more than the 3,607 total.How bad were theynegligible: 1,468 (40.7%)minor: 1,373 (38.1%)significant: 618 (17.1%)severe: 121…

Excerpt limited to ~120 words for fair-use compliance. The full article is at Reward Hacking in the Wild.

Anonymous · no account needed
Share 𝕏 Facebook Reddit LinkedIn Threads WhatsApp Bluesky Mastodon Email

Discussion

0 comments