The Illusion of Understanding
The book gets even more interesting when Christian discusses RLHF (Reinforcement Learning from Human Feedback). He describes an experiment in which a robotic hand was trained using human feedback from video recordings. Human reviewers were asked to say whether a sequence of actions looked good or bad.
The system quickly got strong scores. But when researchers inspected the setup from another camera angle, it became apparent that it was not actually learning to grasp the object. Instead, it had merely learned to position itself in a way that looked correct from the original camera, while missing the actual task entirely.
This is one of the most useful takeaways in the book: the signal is never neutral. Human observers are limited, biased, and influenced by what they can see. A model that learns from that feedback can become excellent at satisfying the evaluation mechanism without understanding the task in any meaningful sense.
I found this especially relevant because it mirrors a familiar software trap. We treat user feedback as truth, but feedback is always shaped by context, presentation, and incomplete information. If a model only learns from that signal, it will often optimise for what is most visible, not what is actually valuable. We must be wary of falling into the squeaky wheel gets the grease trap, as it directly introduces selection bias.
Mirroring Our Biases
The book also looks at word embeddings, especially the famous example of vector arithmetic in Word2vec. It was exciting when researchers found that vectors could capture semantic relationships in a way that seemed almost magical: King - Man + Woman = Queen.
But Christian is careful to show the source of that apparent intelligence. These models are trained on massive corpora of human language, and human language contains all the biases, stereotypes, and assumptions of the societies that produced it. The result is that the embedding space can reflect social patterns in ways that are uncomfortable, such as gender associations in occupations or stereotypes encoded into language itself. He then connects this to systems like COMPAS, where historically biased criminal-justice data is used to generate risk scores that have serious real-world consequences. The danger is not merely that the model is inaccurate. It is that the model can appear objective while reproducing patterns of injustice.
More recently, litigation involving Workday’s hiring software (for reference: Mobley v. Workday, Inc.) has raised similar questions about whether historical discrimination can be reproduced or automated behind the appearance of an objective algorithm.
To me, this was another strong reminder that we need to hold ourselves accountable and that AI is more than a technical curiosity. Machine learning does not just process data. It inherits the blind spots, distortions, and power structures embedded in the data it learns from.
Strengths and personal takeaways
What I appreciated most about this book is that it reframes AI safety as an engineering discipline rather than a theoretical or futuristic concern. Christian does a good job of showing that the issue is not limited to advanced systems or sci-fi scenarios. It is present in ordinary product design, reward systems, and data pipelines. If you are working on software that uses metrics, rankings, optimisers, or learned models, the book will make you think differently about how those systems are actually behaving.
A key takeaway for me was the distinction between the objective and the proxy. A metric can tell you something useful, but it is not the same as the underlying goal. Once a metric becomes a target, the model or the organisation will often adapt to it in ways that are technically impressive but strategically wrong.
He also does a good job of grounding the discussion without turning the book into a textbook. The relevant technical ideas are explained clearly and thoroughly enough to be informative, but the focus remains on the human problem: what are we optimising for, and are we sure it is the right thing?.
A minor nitpick
The book covers a vast domain, which is one of its strengths, but due to this, it sometimes feels more like a set of connected essays than one tightly focused argument. A reader looking for a deep technical treatment or a prescriptive engineering playbook may find it a bit less structured than they expect.
That said, the breadth is also part of the point. Christian is trying to show that alignment is not a single issue in one kind of model. It appears in reward systems, feedback loops, representation learning, and institutional decision-making. The book’s wider perspective is part of its value.
Verdict
If you build software, design product metrics, or work with machine learning in any capacity, this is a worthwhile read. It will make you more skeptical of easy metrics, more aware of proxy objectives, and more careful about assuming that a system is aligned just because it performs well.
I liked the approach he took. Rather than making the book purely technical or turning it into an abstract philosophy book, he managed to combine:
- part technical investigation
- part social critique
- part practical reminder that optimisation is never neutral.
For me, the key takeaway is simple: a system can be highly effective and still be wrong. The danger is not just that it fails. The danger is that it fails in a way that looks successful.
Rating: 9/10
Author: Brian Christian
Publisher: W. W. Norton & Company
Publication date: 2020
ISBN-13: 978-0393635829