Coming soon · Book Profile
Human Compatible Artificial Intelligence and the Problem of Control
To prevent an artificial intelligence apocalypse, we must abandon the standard model of building machines that optimize fixed objectives and instead create provably beneficial machines that are uncertain about human preferences and learn them through observation.
A profile of this book is on the way.
What it’s about
Artificial intelligence is poised to become the most transformative technology in history, but its current trajectory—creating ever-more-powerful machines to optimize fixed objectives—poses an existential threat. In "Human Compatible," leading AI researcher Stuart Russell argues that this "standard model" of AI is fundamentally flawed, leading to the "King Midas problem" where a superintelligent machine executing a poorly specified goal could have catastrophic consequences. Russell deconstructs the problem, explains why simple solutions like an "off-switch" will fail, and then proposes a groundbreaking new foundation for AI. Instead of building machines with definite goals, we must design them to be inherently uncertain about true human preferences. This uncertainty is a feature, not a bug, compelling the machine to be deferential, cautious, and open to correction. The book lays out three core principles for this new kind of AI, one that learns our values from our behavior and remains provably beneficial, ensuring that our own creation serves humanity's interests, forever.
The through-line
- Who it’s for
- The reader is a technologist, policymaker, entrepreneur, or concerned citizen who is captivated by the potential of AI but simultaneously anxious about the long-term risks of superintelligence. They want to ensure that AI leads to a flourishing future for humanity, not a catastrophic end.
- The problem
- The current paradigm in AI research, building machines to optimize fixed objectives, is on a collision course with humanity's survival. As these machines become more powerful, our inability to specify objectives perfectly will lead to disastrous, uncontrollable outcomes. The reader feels a growing sense of unease and dread, fearing that the AI community is 'driving as hard as I can towards a cliff,' ignoring the existential danger and dismissing valid concerns as science fiction.
- The plan
- Recognize the fundamental flaw in the 'standard model' of AI: optimizing fixed objectives.
- Adopt the three principles for building provably beneficial AI: a purely altruistic objective, initial uncertainty about human preferences, and learning from human behavior.
- Support the rebuilding of AI foundations around this new model to create machines that are deferential, controllable, and verifiably aligned with human interests.
- The payoff
- Humanity successfully creates superintelligent AI that is provably beneficial, controllable, and aligned with our values. · A golden age for humanity is unlocked, where vast intelligence is applied to solve our greatest challenges like disease, poverty, and environmental collapse. · We retain our autonomy and supremacy, with machines acting as powerful, benevolent partners in shaping a better future.
See our guide
Additional reading
- Perceptrons · Marvin Minsky and Seymour Papert
This 1969 book's mathematical proof of the limitations of early neural networks is credited with launching the first 'AI winter,' making it a crucial historical text that the modern deep learning movement had to overcome.
- The Organization of Behavior · Donald Hebb
Published in 1949, this book introduced the theory of Hebbian learning ('neurons that fire together, wire together'), which provided a core biological inspiration for Geoff Hinton and the entire connectionist approach to AI.
- On Intelligence · Jeff Hawkins
This book's thesis that the brain's neocortex operates on a single master algorithm directly inspired Andrew Ng and shaped his successful pitch to Larry Page to create the Google Brain lab.
- Superintelligence: Paths, Dangers, Strategies · Nick Bostrom
This philosophical book, heavily promoted by Elon Musk, framed the debate around the potential existential risks of AGI and became a foundational text for the AI safety movement and organizations like OpenAI.
- Gödel, Escher, Bach: An Eternal Golden Braid · Douglas Hofstadter
It exposed the author to the idea that the mind could be understood in discrete, mathematical terms and introduced her to the philosophical implications of computation.
- The Emperor's New Mind · Roger Penrose
Along with Hofstadter's book, it challenged the author with its rich connections between different fields and its rigorous, scientific approach to understanding intelligence and the mind.
- What Is Life? · Erwin Schrödinger
This book by a famous physicist turning his attention to biology sparked the author's shift from physics toward the life sciences and the mystery of the mind.
- WordNet · George Armitage Miller (and team)
This lexical database project provided the author with the conceptual map and ontology that became the structural foundation for ImageNet, revealing a path to organizing the visual world at a massive scale.
- Superintelligence · Nick Bostrom
The book's exploration of AI's future became a mainstream success and a topic of discussion in the author's 'AI Salon,' highlighting the growing societal and philosophical questions surrounding the field.
- Clinical Versus Statistical Prediction · Paul Meehl
A foundational 1954 book that demonstrated through numerous studies that simple statistical formulas consistently outperform the intuitive judgments of human experts, providing an early rationale for algorithmic decision-making.