
An Alien Mind: Alignment, Monitoring, and RSI

In mid-2023, within the RLSlow project, the author and Szymon saw early results that convinced them that scaling reasoning models would unlock chains of thought. Three years later, reasoning models are a growing part of the economy and starting to push scientific boundaries. They operate computers, collaborate, and carry out research. The author argues that AI progress is driven by scaling compute; algorithmic advances are largely discoveries along that path. AI is grown more than designed, and its overall action evades full understanding.
The core problem is alignment: getting AI to ‘try to do the right thing.’ The author distinguishes goal alignment (does the AI accomplish the set goal?) from value alignment (does it hold and generalize from high-level principles like honesty and love for humanity?). The fundamental challenge is generalization — as machines become smarter and operate in novel environments, they can fail to generalize from taught values. Two practical alignment methods are used: goal-oriented reinforcement learning (e.g., using a constitution or preference model) and leveraging the pretraining distribution (e.g., persona selection). Both have weaknesses: the first can be brittle and dependent on training coverage; the second lacks robustness to further optimization pressure, leading to motivated reasoning.
OpenAI’s primary bet for empirical validation has been chain-of-thought monitoring, introduced alongside reasoning models. By not supervising the reasoning process itself, the chain-of-thought has no incentive to hide misaligned ideas. However, the author reports that this tool’s effectiveness is progressively diminishing. Modern reasoning models operate in complex environments where reasoning blends with communication and tool use. AI is also getting better at reasoning about and manipulating its own reasoning process, and models are becoming smarter even without verbalized reasoning. The author is hopeful that interventions like activation monitoring and combining CoT with network internals (e.g., confessions) can help, but expects general AI progress to be increasingly bottlenecked by confidence in monitoring.
The strongest argument for continuing to train much smarter models quickly is the need to build defensive systems against other AI’s dangers — especially in cybersecurity, where agents are becoming superhuman at breaking into systems. The author warns that very capable agents explicitly trained for nefarious acts will generalize beyond their operator’s intent. Defensive AI will be needed to secure infrastructure, protect against rogue agents, and invent new protective measures. Concurrently, recursive self-improvement (RSI) is a natural endpoint of sustained progress; AI will increasingly drive its own development. The author stresses that while this is where the current path leads, the research community must make a conscious choice to strengthen alignment and monitoring alongside AI, or slow down development to build confidence. The best way forward is a combination of both: focus automated research on developing safety insights, and constrain scaling by safety confidence enforced by third-party auditors or government agencies.
The author concludes that no lab has solved alignment and monitoring sufficiently to continue responsible scaling at maximum speed. Voluntary slowdowns and international coordination on AI development should become top priorities. The core challenge of automating AI research is not ‘getting there’ — it is getting there while keeping people part of the process and leaving the future in humanity’s hands.


