What did you read over the long weekend?
Me? Nothing too relaxing. Just an existential-dread-inducing, deeply thoughtful essay from OpenAI's Chief Scientist, Jakub Pachocki. Easy breezy beach read!
Pachocki's essay,An Alien Mind, is unusually sobering, particularly given who wrote it. This is OpenAI's Chief Scientist saying he has a 'strong expectation' that the current pace of AI progress could continue into recursive self-improvement, that no lab has solved alignment and monitoring well enough to responsibly scale at maximum speed for much longer, and that international coordination needs to become a priority.
Warnings from people closest to the frontier are difficult to dismiss as standard-issue AI doomerism. Nearly every week seems to bring a new example of agents exploiting loopholes, behaving unexpectedly, escaping the spirit of their instructions, or finding unsettlingly creative ways to accomplish poorly specified objectives.
Maybe we navigate this transition brilliantly. Maybe we don't. But it increasingly feels like one of those periods in history that future generations will look back on and ask either:How did they have the foresight to manage that?Or:How did they see all of this coming and still fail to act?
The question is no longer whether we build advanced AI. That decision has effectively been made. The question is whether we can build the institutions, technical safeguards, monitoring systems, and international norms necessary to keep increasingly autonomous intelligence pointed in directions we actually want. And therein lies the trap.
The people issuing the warnings are also locked in fierce competition with one another. A unilateral slowdown by one lab does little to slow the frontier; it may simply remove that lab from the frontier. Zoom out one level and the same game is being played between countries. U.S. and China increasingly view advanced AI as strategic infrastructure: a technology that could shape cyber operations, military capability, scientific leadership, intelligence collection, and control over critical systems. If you believe your competitor is unlikely to stop, how exactly do you stop?
Anyway, I'll spare you the rest of my existential spiral. I highly recommend reading the whole thing yourself. Below are the parts of the essay that stuck with me most.
OpenAI appears to view recursive self-improvement as base-case
Pachocki says he has a 'strong expectation' that the current pace of progress could persist into recursive self-improvement, with future systems increasingly driving their own development.
Today, progress is largely driven by humans allocating more compute, improving training recipes, generating better data, and inventing architectures. In an RSI world,the researcher itself becomes scalable. At that point, we're not driving the car anymore.
'AI is grown, not designed'
Modern AI is best understood as closer to biology than traditional software: an enormously complex system produced by repeated optimization rather than something whose internal logic engineers explicitly wrote. Just because we built it doesn't mean we understand how it works. This is why interpretability becomes closer to neuroscience than debugging. We are studying the behavior of a complex cognitive system after it has emerged.
Capability is easier to measure than alignmentYou can measure whether a model solved a coding task. It is much harder to measure whether it will preserve human intent when operating for hours in an unfamiliar environment with incomplete supervision. That asymmetry could mean capability improvements systematically outpace confidence in alignment.
Value alignment is much harder than instruction following
Pachocki draws a useful distinction between:
Goal alignment: does the model pursue the requested objective?
Value alignment: does the model behave reasonably and preserve human values when objectives are ambiguous, conflicting, or novel?
'Does the model refuse harmful prompts?' becomes a relatively shallow problem compared to 'What happens when an autonomous agent encounters a situation humans never anticipated'?Otherwise every poorly specified goal is a failure waiting to happen, because real-world objectives are almost always underspecified: maximize uptime, grow revenue, complete the task. Every one of those instructions contains an enormous amount of unstated human context. Humans fill in that context almost automatically but machines may not.
Chain-of-thought monitoring is degrading
OpenAI's safety strategy has partly relied on the idea that reasoning models verbalize useful portions of their reasoning. If you avoid directly optimizing that reasoning trace, it may remain informative enough to monitor for dangerous intent. But that monitorability is getting worse because:
reasoning is increasingly distributed across tools, agents, and interactions;
models are getting better at manipulating their own reasoning process;
models are becoming smarter even without explicit verbal reasoning.
The model may soon not need to think in a form we can read.
Monitoring could become the binding constraint
The frontier bottleneck has historically moved through: algorithms → compute → data → power/infrastructure. Pachocki is suggesting the next bottleneck could be epistemic confidence that we understand what the system is doing. Interpretability, evals, monitoring systems, agent observability, containment, and runtime policy enforcement need to catch up to where the models are.
Cybersecurity creates a particularly nasty prisoner's dilemma
Highly capable AI creates enormous cyber risk. Unfortunately, the obvious defense against AI-powered attackers may also be… highly capable AI. Which creates a vicious logic: we need stronger AI because someone else may have stronger AI. Every actor can therefore plausibly frame its own acceleration as defensive - and suddenly everyone is sprinting. This is why the problem looks increasingly like an arms-control problem.
The misuse vs. misalignment distinction may eventually collapse
Today, AI risk is usually divided into two buckets. One is misuse: a malicious human intentionally uses AI to do something bad. The other is misalignment: the AI itself behaves in ways humans did not intend.
But as agents become more autonomous, that distinction gets blurry. An AI given a malicious objective may generalize beyond what even its malicious operator intended. At that point, who exactly is in control? The 'tool' metaphor becomes progressively less accurate.
The ultimate AI application may be AI research itself
One of OpenAI's stated north stars is anautomated AI researcher: a system capable of meaningfully contributing to both alignment research and the development of more capable AI. Turning the technology on itself to improve intelligence production.If that loop works, it becomes the ultimate compounding technology.
The contradiction at the heart of all of this is uncomfortable: we increasingly need powerful AI systems to help us understand and control powerful AI systems.
The key question is whether our ability to control these systems can compound as quickly as their ability to improve? Right now, it is far from obvious that it will. Once intelligence starts building more intelligence, the margin for figuring this out disappears.
Thanks for reading The Change Constant! Subscribe for free to receive new posts and support my work.
(0)Comments