US AI safety researcher Paul Christiano, who played a key role in the introduction of reinforcement learning from human feedback (RLHF), is coming to Toronto next month to discuss the future of the technology and how to ensure it stays aligned with human interests.
The AI Safety Foundation announced on Monday that Christiano has been chosen as the distinguished speaker at the third annual edition of The Hinton Lectures, a series co-founded by Turing Award winner and Nobel laureate Geoffrey Hinton, one of the so-called “godfathers” of AI.
Paul Christiano, ARC
“We are not very good at directing the AI systems we create.”
Ahead of the event, which takes place from Nov. 9 to 11, Christiano connected with BetaKit and other members of the media to discuss his work, the state of AI, and some of the existential challenges he believes it could create in the very near future. “It’s an extremely important moment in history in the field of AI in particular, but for the world more broadly,” he argued.
A decade ago, AI couldn’t string together a sentence, Christiano said. Over the past month, however, AI systems have executed a decade’s worth of progress in mathematics and demonstrated potential to carry out sophisticated real-world plans over time, including in a series of hacks, while also maxing out the tests used to measure their performance.
These factors, coupled with the fact that “we are not very good at directing the AI systems we create,” have created “a particularly dangerous moment,” Christiano said. The OpenAI-Hugging Face cybersecurity incident is an early example of how this could happen. He said this and similar events have caused “very small-scale harms.”
“If you scaled those up to very powerful AI systems trained on a broader distribution of tasks, you might see irreversible harms,” Christiano added.
RELATED: AI godfather Geoffrey Hinton says we must convince AI that it’s our mother
Given the current pace of AI progress and the potential for AI to automate and accelerate the AI research and development process, Christiano said, “There is a real chance that over something like a six-to-18-month horizon, we will be dealing with AI systems such that if they wanted to escape human control, undermine human control, or fight openly against humans, it would be very hard for humanity to prevail.”
Christiano predicted that soon, large parts of the AI development process will be “extremely hard” for humans to follow, let alone understand, requiring us to rely upon AI to measure and review other AI systems—a premise that carries additional risks.
As an early researcher at OpenAI, Christiano helped introduce RLHF, a post-training technique designed to align these systems with human preferences. In his introductory remarks as part of this session, Hinton said RLHF enabled the public rollout of the AI chatbots we know today and gave OpenAI a huge advantage over Google.
“If you scaled those up to very powerful AI systems trained on a broader distribution of tasks, you might see irreversible harms.”
Today, Christiano works as senior technical advisor at NIST’s Center for AI Standards and Innovation, founder and executive director of the Alignment Research Center (ARC), while also sitting on the OpenAI Foundation’s board and serving on OpenAI’s Safety and Security Committee. During Wednesday’s discussion, Christiano did not discuss the decisions of specific governments or companies like OpenAI, but rather spoke to the issue of AI safety more broadly.
Christiano does not believe RLHF is the ultimate answer to making these systems safe. While it has fuelled major breakthroughs in generative AI, he noted that it is “extremely difficult to scale” and “will probably break down at some point.” Different parties are pursuing various alternatives for training AI as it grows, but he said it is unclear which will bear fruit or when.
ARC is one of them. He puts the non-profit’s chance of success at closer to 10 percent than 50 percent, noting its research could become useful at some point but likely won’t fundamentally change our ability to control this tech.
Unless political outcomes slow the pace of AI progress—something many are calling for, including Hinton and fellow Canadian AI pioneer Yoshua Bengio (the latter of whom penned an opinion piece today calling for frontier AI firm employees to leave if they prioritize safety)—Christiano is skeptical he will be conducting any relevant research in five to 10 years.
Establishing a consensus and standards around how to measure the performance of these systems will be key. Amid reports that some companies are restricting access to certain models outside the US, this is an issue that Christiano thinks “cannot really be settled well behind closed doors,” and would be best addressed by the global scientific community, from researchers south of the border to experts in Canada and abroad.
Feature image courtesy the AI Safety Foundation.
