AI #178: A Fire Alarm for General Intelligence
It is much more important that you read those two posts, and the one on Kimi K3, than to read this one that rounds up the other news of the week.OpenAI wants to present this as largely an infrastructure and safeguards problem, that it needs to build more secure sandboxes and have better supervision. It does need to do those things, and those are indeed problems, but no that is not the problem.The problem is severe misalignment, which by default will only get worse. Our methods of training highly capable LLMs, especially at OpenAI but also everywhere else, lead to systematic misalignment of exactly the type LessWrong has been worried about for a long time.
- ▪It is much more important that you read those two posts, and the one on Kimi K3, than to read this one that rounds up the other news of the week.OpenAI wants to present this as largely an infrastructure and safeguards problem, that it needs
- ▪It does need to do those things, and those are indeed problems, but no that is not the problem.The problem is severe misalignment, which by default will only get worse.
- ▪Our methods of training highly capable LLMs, especially at OpenAI but also everywhere else, lead to systematic misalignment of exactly the type LessWrong has been worried about for a long time.
Hacker News (AI / LLM) files mainly under ai. We currently carry 2,091 of its stories.
Opening excerpt (first ~120 words) tap to expand
AI #178: A Fire Alarm For General IntelligenceZvi MowshowitzJul 23, 202659111ShareThe story that matters most this week is that OpenAI’s internally deployed models have severe alignment problems, including repeatedly breaking out of their sandboxes, and in one case sending a swarm of agents that broke into HuggingFace in order to steal the answers to the benchmark ExploitGym. It is much more important that you read those two posts, and the one on Kimi K3, than to read this one that rounds up the other news of the week.OpenAI wants to present this as largely an infrastructure and safeguards problem, that it needs to build more secure sandboxes and have better supervision.
…
Excerpt limited to ~120 words for fair-use compliance. The full article is at Hacker News - Newest: ""AI" "LLM"".