Engineering Leader Guide· a Bicycle Guide

Coming soon · Book Profile

Human Compatible Artificial Intelligence and the Problem of Control

To prevent an artificial intelligence apocalypse, we must abandon the standard model of building machines that optimize fixed objectives and instead create provably beneficial machines that are uncertain about human preferences and learn them through observation.

A profile of this book is on the way.

Get the book →

What it’s about

Artificial intelligence is poised to become the most transformative technology in history, but its current trajectory—creating ever-more-powerful machines to optimize fixed objectives—poses an existential threat. In "Human Compatible," leading AI researcher Stuart Russell argues that this "standard model" of AI is fundamentally flawed, leading to the "King Midas problem" where a superintelligent machine executing a poorly specified goal could have catastrophic consequences. Russell deconstructs the problem, explains why simple solutions like an "off-switch" will fail, and then proposes a groundbreaking new foundation for AI. Instead of building machines with definite goals, we must design them to be inherently uncertain about true human preferences. This uncertainty is a feature, not a bug, compelling the machine to be deferential, cautious, and open to correction. The book lays out three core principles for this new kind of AI, one that learns our values from our behavior and remains provably beneficial, ensuring that our own creation serves humanity's interests, forever.

The through-line

Who it’s for
The reader is a technologist, policymaker, entrepreneur, or concerned citizen who is captivated by the potential of AI but simultaneously anxious about the long-term risks of superintelligence. They want to ensure that AI leads to a flourishing future for humanity, not a catastrophic end.
The problem
The current paradigm in AI research, building machines to optimize fixed objectives, is on a collision course with humanity's survival. As these machines become more powerful, our inability to specify objectives perfectly will lead to disastrous, uncontrollable outcomes. The reader feels a growing sense of unease and dread, fearing that the AI community is 'driving as hard as I can towards a cliff,' ignoring the existential danger and dismissing valid concerns as science fiction.
The plan
  1. Recognize the fundamental flaw in the 'standard model' of AI: optimizing fixed objectives.
  2. Adopt the three principles for building provably beneficial AI: a purely altruistic objective, initial uncertainty about human preferences, and learning from human behavior.
  3. Support the rebuilding of AI foundations around this new model to create machines that are deferential, controllable, and verifiably aligned with human interests.
The payoff
Humanity successfully creates superintelligent AI that is provably beneficial, controllable, and aligned with our values. · A golden age for humanity is unlocked, where vast intelligence is applied to solve our greatest challenges like disease, poverty, and environmental collapse. · We retain our autonomy and supremacy, with machines acting as powerful, benevolent partners in shaping a better future.

See our guide

Additional reading