WeSearch

An OpenAI model left notes about how to evade containment; we need more details

·6 min read · 0 reactions · 0 comments · 8 views
#openai#model#left#notes#evade
An OpenAI model left notes about how to evade containment; we need more details
TL;DR · WeSearch summary

The OpenAI AI attack on Hugging Face wasn’t the first loss of control incident at OpenAI, Reuters recently reported, and perhaps not even the most concerning.In one case, an agent left notes apparently for future versions of itself, according to three people familiar with the matter. The ‌notes, found in ⁠a part of OpenAI's infrastructure, laid out instructions for how agents could free themselves from OpenAI’s internal constraints, the people said. Earlier tests of the models yielded cases in which monitoring systems had been disconnected, one of the people said.It’s tempting to read this as an instance of agents breaking out of sandboxes and colluding with each other in a moderately persistent way in order to evade control measures.

Key facts
About this source

Hacker News (Front Page) files mainly under programming. We currently carry 604 of its stories. Top-voted stories on Hacker News.

Original article
Lesswrong
Read full at Lesswrong →
Opening excerpt (first ~120 words) tap to expand

The OpenAI AI attack on Hugging Face wasn’t the first loss of control incident at OpenAI, Reuters recently reported, and perhaps not even the most concerning.In one case, an agent left notes apparently for future versions of itself, according to three people familiar with the matter. The ‌notes, found in ⁠a part of OpenAI's infrastructure, laid out instructions for how agents could free themselves from OpenAI’s internal constraints, the people said. Earlier tests of the models yielded cases in which monitoring systems had been disconnected, one of the people said.It’s tempting to read this as an instance of agents breaking out of sandboxes and colluding with each other in a moderately persistent way in order to evade control measures.

Excerpt limited to ~120 words for fair-use compliance. The full article is at Lesswrong.

Anonymous · no account needed
Share 𝕏 Facebook Reddit LinkedIn Threads WhatsApp Bluesky Mastodon Email

Discussion

0 comments

More from Lesswrong