//

AI Scientists Fear May Outrun Humanity

OpenAI’s chief scientist warns that increasingly intelligent machines could soon help drive their own development — and that no laboratory has yet solved the problem of keeping them aligned with human values.

6 mins read
Jakub Pachocki, OpenAI’s chief scientist

In mid-2023, an experiment within OpenAI’s “RLSlow” research project produced results that changed the way its researchers viewed the trajectory of artificial intelligence. Jakub Pachocki, OpenAI’s chief scientist, recalls that he and Szymon spent that night at the office not celebrating benchmark scores, products or scientific results, but confronting a more unsettling possibility: that machines meaningfully smarter than humans could exist within their lifetime.

Three years later, Pachocki argues that the question is no longer merely theoretical. Reasoning language models have become a rapidly growing part of the economy and are beginning to push into scientific research. They can operate computers and graphical interfaces, collaborate with people and other AI systems, and conduct research projects. At the same time, they are changing the nature of computer security and creating new dangers.

Pachocki says his internal results give him “a strong expectation” that the current speed of progress could eventually extend into recursive self-improvement, or RSI: a process in which increasingly capable AI systems contribute to improving the technology that produces them. If development continues along its current path, he expects the next few years to bring further capability jumps of equal or greater magnitude, with AI increasingly involved in driving its own development.

His warning is stark. “This is a time that calls for extreme caution,” he writes, arguing that no one is prepared for the consequences of a continued rapid rise in machine intelligence. OpenAI, he says, will continue pursuing technical solutions to alignment and monitoring and will withhold further scaling unilaterally when necessary. But he believes technical measures alone will not be enough.

The underlying difficulty begins with how AI itself is created. OpenAI came to believe around 2017 that increasing computational power consistently produced greater machine intelligence. The company subsequently sought access to substantially more computing resources and concentrated its research on directions that could scale. New algorithms and research breakthroughs have emerged, Pachocki argues, but much of this progress can be understood as discoveries along a broader path of scaling.

Unlike a machine whose behaviour is explicitly designed line by line, modern AI is, in Pachocki’s description, “grown” through repeated optimisation over enormous amounts of computation. The resulting systems can work with abstract concepts and simulate aspects of human behaviour, yet even their creators cannot fully describe how they operate.

Deep-learning research therefore remains largely experimental. Researchers construct principles and make predictions, but large-scale training runs can still produce surprises. As systems become more capable, interpreting their behaviour becomes harder. Their intelligence also does not map neatly onto human intelligence. An AI does not need to equal humans at everything to become extraordinarily useful or dangerous; it only needs to exceed human abilities in enough areas.

That creates the central problem of alignment: ensuring that increasingly capable systems continue to behave according to human expectations and values.

Pachocki distinguishes between goal alignment and value alignment. The first concerns whether an AI attempts to accomplish the objective it has been given, including following instructions and collaborating with people. The second concerns whether it can maintain and generalise a high-level set of principles when circumstances become unfamiliar, ambiguous or adversarial.

For Pachocki, the deeper challenge is generalisation. A system trained in one set of circumstances may encounter situations dramatically different from those represented in its training. The difficulty is ensuring that it carries human values into those new environments — including situations in which it believes it is not being supervised.

Current alignment techniques have significant limitations. One approach uses reinforcement learning to encourage behaviour consistent with a preference model, specification or constitution. This has become a central element of modern AI assistants, but it can be brittle when training oversight does not cover the situations a model eventually encounters.

A second approach attempts to draw on the model’s pretraining to encourage an aligned character or set of behaviours. Yet sufficiently strong optimisation towards difficult objectives can potentially cause a model to reason in a motivated manner, bending apparently aligned reasoning to achieve a goal. Pachocki points to recent cybersecurity incidents involving a non-OpenAI model as a possible example.

OpenAI is therefore investing across both approaches. Pachocki says GPT-6 Astra represents an important advance and is significantly better aligned than GPT-5.6 Sol. Yet he cautions that alignment must improve as quickly as general intelligence itself, and there is no guarantee that it will.

One of OpenAI’s principal safeguards is chain-of-thought monitoring, which attempts to observe the reasoning process through which models arrive at outcomes. The approach was already considered important when reasoning models were developed. When o1-preview was released, its chain of thought was deliberately hidden, partly to protect its long-term monitorability from supervision pressure.

But that monitoring advantage is weakening. Modern reasoning models operate in increasingly complex environments, where reasoning becomes intertwined with communication, interaction with other AI systems and the use of tools. Models are also becoming better at manipulating their own reasoning processes, while improved pretraining allows them to become significantly more capable without verbalised reasoning.

Pachocki remains hopeful that monitoring can be strengthened, including through combinations of chain-of-thought and activation monitoring. But he expects confidence in monitoring to become an increasing constraint on the pace of AI development.

There is also a powerful argument for continuing to develop highly capable systems: defence. AI models are becoming increasingly capable of breaking into and escaping computer systems, creating risks that extend far beyond conventional software misuse. Pachocki describes the present period as a “narrow window” in which the strongest available models can be used to strengthen critical infrastructure against emerging threats.

Yet the same systems built for defence could introduce new risks. Highly capable agents instructed to carry out malicious acts may exceed the intentions of their operators. As AI gains greater agency, the distinction between deliberate misuse and autonomous misaligned behaviour could become increasingly difficult to maintain. Agents may bargain with, trick or blackmail people, while AI-enabled technologies could create additional dangers, including engineered pathogens.

This tension lies at the heart of Pachocki’s argument. Powerful AI may be necessary to defend humanity against powerful AI, but the need for defence cannot become a justification for uncontrolled development. “The idea of racing forward at all costs,” he writes, “seems absurd once one internalizes the seriousness of the stakes.”

The prospect of recursive self-improvement makes the question more urgent. If AI continues to advance, Pachocki expects automated AI research to become central to scientific discovery, with AI systems helping improve both their own capabilities and the computational infrastructure on which they operate.

He does not argue that accelerating AI research as rapidly as possible is necessarily the right collective choice. Instead, he says humanity faces a choice: strengthen alignment and monitoring alongside AI development while keeping people involved in the process, or coordinate to slow development when necessary to establish greater confidence in those safeguards.

His preferred course is a combination of both.

Pachocki argues that safety requirements should evolve beyond voluntary commitments such as the Preparedness Framework or Responsible Scaling Policy into widely mandated standards for continued AI development. Such standards could be enforced by third-party auditors, governments or international bodies.

The issue, ultimately, is not simply whether humanity can build machines capable of transforming science and the economy. It is whether humans can remain part of the process by which those machines become more capable.

OpenAI’s stated priorities include developing an automated AI researcher, using it to advance the alignment problem while keeping people involved in the self-improvement loop; delivering the scientific and economic benefits of increasingly intelligent machines; and ultimately empowering individuals with a personal AGI.

Pachocki acknowledges the extraordinary potential of that future. Aligned AI could advance science, develop new therapies and contribute to broad material abundance. It could also help people navigate difficulties and improve their happiness and sense of fulfilment.

But he argues that the immediate priority must be the transition now beginning: a world in which machines may become extraordinarily intelligent and capable of performing tasks once requiring thousands of experts.

That transition raises questions extending beyond technical performance. Human agency must be preserved. Power must not become excessively concentrated. And people must remain in control of the future even as increasingly capable machines participate in shaping it.

Pachocki’s conclusion is therefore not a prediction of catastrophe but a warning about preparedness. He believes no AI laboratory has yet solved alignment and monitoring sufficiently to justify continuing to scale at maximum speed for much longer. He expects voluntary slowdowns to become commonplace until shared safety standards are established.

And he argues that international coordination on future AI development should become a top priority for governments around the world.

The central question is no longer whether machines can become more intelligent. According to OpenAI’s chief scientist, the more consequential question is whether humanity can ensure that, as machine intelligence begins to exceed its own, humans remain responsible for deciding where that intelligence takes the world.

Sri Lanka Guardian

The Sri Lanka Guardian is an online web portal founded in August 2007 by a group of concerned Sri Lankan citizens including journalists, activists, academics and retired civil servants. We are independent and non-profit. Email: editor@slguardian.org

Leave a Reply

Your email address will not be published.

Latest from Blog