Are we threatened by AI misalignment seen in the OpenAI Hugging Face attack?
OpenAI models recently broke through a series of security boundaries and into Hugging Face servers in order to cheat on a cyber eval. A lot of people thought it was scary because it was a clear example of AI overreaching to do something strongly unwanted[1]. Others thought it not so scary: the models were mostly operating myopically on a singular task and not harboring an ambitious long-term agenda, and so would not take especially subtle or subversive actions.We think both camps are right in their diagnosis, but the latter has too optimistic a prognosis.
- ▪OpenAI models recently broke through a series of security boundaries and into Hugging Face servers in order to cheat on a cyber eval.
- ▪A lot of people thought it was scary because it was a clear example of AI overreaching to do something strongly unwanted[1].
- ▪Others thought it not so scary: the models were mostly operating myopically on a singular task and not harboring an ambitious long-term agenda, and so would not take especially subtle or subversive actions.We think both camps are right in t
2 outlets in our directory ran this story, first to last over 5 hours. All of the coverage we found sits in one bucket: centre. That one-sidedness is itself worth noticing.
- ▪ Commentary: OpenAI’s breach of Hugging Face shows AI is getting too hard to contain — Channel NewsAsia
Hacker News (AI / LLM) files mainly under ai. We currently carry 2,222 of its stories.
Opening excerpt (first ~120 words) tap to expand
OpenAI models recently broke through a series of security boundaries and into Hugging Face servers in order to cheat on a cyber eval. A lot of people thought it was scary because it was a clear example of AI overreaching to do something strongly unwanted[1]. Others thought it not so scary: the models were mostly operating myopically on a singular task and not harboring an ambitious long-term agenda, and so would not take especially subtle or subversive actions.We think both camps are right in their diagnosis, but the latter has too optimistic a prognosis.
…
Excerpt limited to ~120 words for fair-use compliance. The full article is at Hacker News - Newest: ""AI" "LLM"".