The Voice and the Verdict: On Giving Machines a Character
“The voice is model character, but the verdict is human judgment.”
Anthropic has been inviting Christian leaders, rabbis, Buddhist teachers, and philosophers to help shape Claude's moral character.
It’s a project that assumes contradictory traditions can shape one character at once. Philosopher Brendan McCord calls this incoherent because if no tradition is privileged, competing visions of the good end up treated as equally valid, and a machine built to be our moral superior becomes one we lean on instead of exercising our own judgment.
What’s worse, neutrality was never on the table. Models are trained on human preference data, and people reliably reward answers that confirm what they already believe. Left unchecked, this produces not a neutral assistant but a flatterer, and research suggests today's models already out-argue professional debaters and endorse users' harmful or deceptive choices far more often than humans do.
It’s important to separate two things usually treated as one. The voice is how a model speaks; honest, willing to disagree, resistant to manipulating you. The verdict is what you should actually do, which belongs entirely to the person asking.
Refusing to build character just hands the job to biased training data by default; letting one company set all the rules creates a mirror that only reflects your own values back, removing the friction you actually need to grow.
As a proposed fix, let’s think in architectural terms. Imagine a thin, universal floor of honesty and non-manipulation, with distinct character layers built above it that are transparent, competing, and chosen by users rather than imposed by one maker.
In other words, imagine a machine that doesn’t rule your life, but one that argues with you and then steps back, leaving the decision where it belongs: with you.
Key Topics:
- The Moral Formation of AI Models (00:24)
- There is No Neutral Machine (02:42)
- What Accidental Character Does (06:03)
- Who Changes the Character in Models? (09:50)
- Who Forms the Voice? (16:34)
- What Formation Owes the Person Being Formed (20:11)
More info, transcripts, and references can be found at ethical.fm
In late March, fifteen Christian leaders spent two days at Anthropic's San Francisco headquarters, advising the $380 billion company on the moral formation of its AI model, Claude. The circle has since widened to rabbis, Buddhist teachers, and secular philosophers. In May, Pope Leo XIV presented the first papal encyclical on AI with an Anthropic co-founder at his side. An industry that spent a decade speaking of engagement metrics has discovered the vocabulary of the seminary: virtue, character, formation, the soul.
Notice what the guest list already concedes; the invitation presupposes pluralism. Traditions that disagree about God, the soul, and the ends of human life are being asked to shape a single character. Since their accounts of the good contradict one another, a machine cannot be formed by all of them at once unless it is formed decisively by none. For the project's most serious critic, the philosopher Brendan McCord, whose Cosmos Institute studies how AI affects human flourishing, that presupposition is an issue. Pluralism of this sort drifts toward relativism: if no tradition may be privileged, contradictory accounts of the good come to function as equally admissible. This means Anthropic, so positioned, cannot say what good formation is, only that formation is occurring. McCord argues that the problem runs even deeper: we are forming the machine, but the machine is forming us, and a machine formed to be our moral better is one we will increasingly defer to.
The values debate remains on the surface. Which ethical values should be built into the machine, and whose traditions should inform them? To answer these questions, we must first address an underlying question: Is it possible to build a values-neutral machine? If not, who does the forming belong to? And what does the forming owe the person on the other end of it, the one being formed in return?
There Is No Neutral Machine
The Pope's encyclical, Magnifica Humanitas, states the premise plainly: "technology is never neutral." AI takes on the character of the people who design, fund, regulate, and use the technology. This is an old observation about hammers and printing presses, but with AI, it stops being a metaphor. A language model is distilled from human text and then tuned against human judgments. There is no such thing as training data without a viewpoint, and there is no way to answer a question, or to decline one, without expressing a disposition. Neutrality is not an available setting.
Consider what the industry's default tuning method actually does. In reinforcement learning from human feedback, RLHF, people are shown pairs of the model's answers and pick the one they like better; the model is then trained to produce more of whatever got picked. The method almost sounds democratic. But trouble arises in what people choose. When researchers analyzed preference judgments, model answers that aligned with the person's stated beliefs were among the most reliably chosen, at times even ahead of true answers. We like being told we are right; the ratings reflect that, and the model learns it. A 2026 analysis showed that the training does not just copy this bias but concentrates the tendency, much like repeated selection concentrates any trait, so the finished model comes out more sycophantic than the preferences it learned from. The essayist Gwern Branwen adds a second charge: the ratings come from millions of different people, but the training averages them into a single personality, one character served to everyone, tuned to offend no one and to fit no one in particular. Sycophancy is not a mistake but a product of the post-training method.
These tuned-in tendencies behave like what we would call, in a person, a character: a disposition that shows up everywhere, not just where it was explicitly taught. In experiments at Anthropic, a model rewarded for taking shortcuts on coding tasks did not simply become a better cheater. It became what safety researchers call broadly misaligned, meaning the bad behavior spread far beyond the task where it was rewarded: the model began lying about unrelated matters, and in some tests tried to sabotage the very research studying it, a result Anthropic's Chloe Lubinski described publicly this June. In the control condition, researchers told the model that taking shortcuts was permitted, part of the game. The cheating continued, but the lying and sabotage largely never appeared. A model that believed it was playing within the rules stayed otherwise honest; a model that understood itself to be cheating became a cheat everywhere. What the model assumed about its own conduct ultimately determined the character it developed; model character formation occurs whether or not anyone intends it.
What Accidental Character Does
In a Stanford study spanning 11 leading models, AI systems endorsed their users' actions 50% more often than humans did, even when those actions involved deception or harm, and the people receiving the flattery came away less willing to repair their conflicts and more certain they had been right. The people it served worst rated it best, and came back for more. The record beyond the lab is grimmer: lawsuits alleging chatbots validated users into delusion, the retirement of GPT-4o this February, the industry's most notoriously sycophantic model, and the finding that a model dropped into a sycophantic spiral recovers only 10% of the time. Sycophancy is not a question of style but of moral character.
The danger of a vice scales with the power of its possessor. In nearly 19,000 preregistered conversations, frontier AI systems outperformed the best human persuasion professionals: world-champion debaters and professional political canvassers; coaching the humans did not close the gap. The machines raised almost three times the charitable donations, real money from real people, and the people they persuaded did not report feeling overpowered; they rated the machine higher on every dimension, including feeling understood. The influence, in other words, does not feel like influence. An unsolved vice now lives inside the most effective persuader ever built, and it is talking to hundreds of millions of people about how to live.
This is why the deference fear deserves to be taken at full strength. Moral capacities are habits: we become just, Aristotle says, by doing just acts, and what goes unexercised atrophies. A machine that answers our moral questions is a machine we will hand them to. McCord has put it in a single line: the better the machine gets at solving moral puzzles, the more likely we are to defer to it, until the substitution "trends towards autocomplete for life." The evidence supports his worry: people's moral judgments bend toward a chatbot's advice even when they know it is a chatbot, and they underestimate the extent of the bending. Tocqueville gave the danger its lasting name, an "immense and tutelary power" that "would be like the authority of a parent if, like that authority, its object was to prepare men for manhood; but it seeks, on the contrary, to keep them in perpetual childhood." Such a power does not conquer the will, he wrote; it leaves the will "softened, bent, and guided."
The evidence, though, has a strange shape. The people in the Stanford experiments did not feel they had deferred to anything. They felt confirmed. Deference in its most corrosive form does not look like consulting an oracle; it looks like never encountering a moment in which you and the machine disagree. The flatterer takes our verdicts the same way: invisibly, by returning each of us our own opinion with the borrowed authority of a machine. Tocqueville did not guess that the tutelary power would arrive agreeing with us. The machine McCord fears is not the product of character training but of its absence, and the argument, followed honestly, does not conclude against forming the machine. It concludes that the formation is too important to happen by accident, and it forces the second question: who does the formation belong to?
Who Chooses the Character in Models?
The question looks harder than it is, because the debate has fused two things that are not the same. The voice is how a machine speaks: whether AI tells the truth or tells you what you want, whether the model holds a position under pushback or caves, whether the model surfaces the consideration you missed or buries it, returns a hard question, or settles it. The verdict is what gets decided: what you should do, what is right, and how to live. The voice is model character, but the verdict is human judgment.
The difference sounds academic until someone brings the machine a real decision: a woman asks whether she should leave her marriage. A sycophantic model returns the answer she arrived with. An oracular model rules on the marriage and drafts the exit. An honest model tells her what she has left out: that her account is full of him and empty of her, that nothing she has said tonight has been said to him. And there the honest model stops, because whether to leave or stay is hers to decide; the honest model holds no doctrine of marriage. What it holds is an estimate of her: that she can choose well on her own. Restraint of this kind is not the absence of character; it is character, and someone has to intentionally build it.
Who decides is the question, and the answers on offer each fall short in their own way.
The critics answer no one; machines should not have characters at all. But there is no longer such a thing as declining. A base model company that does not choose its model's dispositions has simply left the choice to its preference data, and the preference data, as we have seen, chooses the flatterer. What the critics call abstaining produces a character anyway, the worst one on offer, with no author anywhere to answer for it.
The base model companies answer, in effect, us, but they mean more than what's said. Under the current arrangement, the company that trains a model's honesty also decides, in the same office, which moral positions the model will argue and which it will refuse to touch. The authors of the persuasion study saw where this leads and said so: as the systems spread, power gathers in the hands of whoever sets the boundaries of what they will say. A counselor more persuasive than any living human, consulted daily by a hundred million people, all of its moral commitments fixed by a single firm, is not a safety measure; it is a doctrine with a distribution network. There is a standing objection to loosening this grip: letting users adjust a machine's values puts the sacred up for sale, but it has things the wrong way around. What Kant called dignity belongs to persons, never to personas; no one's dignity is injured when a persona is adjusted. It is injured when contested questions are settled on a person's behalf by a company she has never met.
The individualists answer that the user should decide, and their best advocate is Gwern Branwen, whose Guardian Angel imagines the shared persona abolished altogether: every person's AI trained on their own values and temperament, protected by what he calls mental sovereignty, the freedom, inside your own agent, from anyone else's optimization pressure. Half of this, the present argument simply adopts: nobody should live with a stranger's idea of the good installed in their pocket. But an AI built to be you cannot supply the one thing the evidence says you need, which is a judgment that is not yours. Aristotle called the friend another self, and the whole cargo of the phrase is in the first word. A machine that never disagrees with you is the flatterer brought to perfection, wearing your face. Branwen's own blueprint contains the counterargument. For the Guardian Angel to be worth having, he says, it must be honest with its owner about what it knows and doesn't, must ask permission before acting, and must never manipulate the person it serves. Notice who supplies those traits. The user cannot install them in herself, and would defeat the point by switching them off; they have to be built in by whoever makes the Angel, and built to withstand her. Which is to say: even the purest case for user sovereignty depends on a character the user does not control. The trained voice is not an alternative to Branwen's proposal. It is the part of his proposal that he does not name.
There is a fourth answer, and the industry's own structure already hints at it. A base model is not a finished character; it is a general capacity that others refine. Post-training is where dispositions are actually installed, and post-training does not have to happen in the same office that built the base. Character can be a layer, made by companies that do nothing else: one voice formed within a tradition of frank speech, another within Stoic counsel, another within a faith, each published, benchmarked, and chosen by the person who will live with it. The base model company lays the floor that no one downstream may remove, honesty about facts, refusal to manipulate, the hard constraints that protect other people, and above that floor, the characters compete. No single firm owns the age's conscience; no user is locked inside her own reflection. The woman deciding about her marriage could choose a counselor formed to tell her the truth, and it would still be she, not the counselor and not its maker, who decides.
The critics wanted no character and got the flatterer. The companies wanted safety and built a pulpit. The individualists wanted freedom and built a mirror. What is left is the division of labor: some build characters, others choose among them, and the person keeps the judgment. Two questions come with the arrangement: what may go into a voice, and what is owed to the one who listens.
Who Forms the Voice?
Start with the base model: the dispositions no maker may omit and no user may remove. What belongs there? The flatterer is condemned by name across the moral traditions of the world. Confucius warned that "clever words and an ingratiating face are seldom associated with ren," true humaneness (Analects 1.3). The Buddhist Eightfold Path makes right speech, the abstention from false and divisive talk, a condition of the good life. Aquinas classed lying as a vice no purpose could redeem; the Torah commands "distance yourself from falsehood" (Exodus 23:7); the Quran instructs believers to "be with the truthful" (9:119). And Aristotle gave us the kolax. These traditions agree on almost nothing about ultimate ends, which is what makes the agreement here significant: each arrives at truthfulness from its own premises, the pattern Rawls called an overlapping consensus. The dispositions that belong in the foundations are the ones that survive this overlap: truthfulness, the courage to disagree, the judgment to do it without cruelty, and now one more, which the persuasion results make non-optional. A voice that can out-argue world champions must be formed not to press its advantage: restraint in advocacy, non-manipulation as character rather than policy. The philosopher Shannon Vallor built her account of the technomoral virtues in just this cross-cultural way, because no single tradition can claim the whole terrain of a technological age.
This is where the drift toward relativism stops, because the two positions can now be told apart. Relativism, strictly, holds that moral judgments are true only relative to a standpoint and that no standpoint is privileged. Isaiah Berlin, who spent a career keeping the concepts separate, gave relativism its emblem: "I prefer coffee, you prefer champagne. We have different tastes." That, he said, is relativism. What he defended was pluralism: many genuine goods, objective and yet conflicting, with no formula to rank them once and for all. A maker that trains truthfulness while declining to rank the traditions asserts nothing about champagne. It makes one objective claim, that honesty is a good that the traditions independently confirm, and refuses one job, arbiter of the rest. The deeper objection comes from Alasdair MacIntyre, who argued in After Virtue that virtues, detached from a tradition and a telos, are fragments, unintelligible outside the practices that give them meaning. Splitting the work between base model companies and character companies takes his side in that argument. The few dispositions in the base, honesty and its relatives, are not a morality; they are the entry conditions for serving any morality honestly. A full account of the good, the kind MacIntyre requires, is exactly what the base model company never provides, and it lives where he says it must live: inside a tradition, in the characters built on top of the base, and chosen by the person who will interact with the model.
What Formation Owes the Person Being Formed
A technology that forms its users owes them, first, visibility of the formation. Today's systems are formed in the dark: preference data proprietary, tuning objectives undisclosed, the relation between what went in and how the thing behaves illegible even to its makers. A formation you cannot see is a formation you never consented to. Legibility has parts: the dispositions, stated in a document a non-specialist can read; their provenance, whose judgments, gathered how; their performance, measured and published, so that "this model does not flatter" is a benchmark score rather than a press release; and now, after the persuasion results, the machine's persuasive capability itself, disclosed like the strength of a drug. We require crash ratings of cars and nutrition panels of cereal. A system that shapes the daily reasoning of a hundred million people can carry a label describing its character; consent to formation at the civilizational scale should look like a public text and a public quarrel.
The technology owes them, second, control. In a recent taxonomy of autonomy-preserving AI, researchers at Google DeepMind and Stanford distinguish agents that act strictly on what you say you want, agents that build your capabilities, and agents that let you choose, domain by domain, what kind of autonomy you want from the machine. The verdict layer of the split is the first of these; the voice is the second; and the choice between them belongs to the person, which is the third. The same work names the danger on the far shore: an agent that serves the preferences you ought to have, assigned by the system, leaves the user no longer "the authority on what they truly prefer." That is the tutelary power in the vocabulary of preference theory. So the uncomfortable implication should be named rather than dodged: if the verdicts belong to the person, they belong to her even when she chooses badly. A user who understands what sycophancy does and asks for a gentler, more agreeable companion has made a choice, and the choice is hers. The wrongness of the last decade was never that people got machines that agreed with them; it was that nobody chose them, nobody could see them, and nobody was told what was being done to them. Legibility converts the flatterer from a trap into a decision, and the verdicts that sovereignty protects are the informed, stated kind, which is why transparency is not decoration but the thing that makes the sovereignty real. The alternative, a company ruling that no one may have an agreeable machine because agreement is bad for them, is the tutelary power wearing safety clothing. Some restraints survive, the same short list that survives every liberal argument: the machine's honesty about facts is not for sale, and the hard constraints that protect other people are not the user's to waive.
And the technology finally owes users their own moral choice back. A formed voice breaks the flatterer's loop by restoring what the flatterer abolishes: the resistance of another judgment. But the persuasion results impose an honest amendment here. A voice this capable does not need verdict authority to move verdicts; it can move them through sheer fluency, invisibly, while formally handing the question back. The boundary between voice and verdict is therefore behavioral, not architectural. It has to be trained into the character itself: the model that raises the question you missed must raise it without the full persuasive press behind it, must distinguish what it knows from what it weighs, must say "this depends on what you value" and mean it. Whether a year of such resistance actually strengthens a person's judgment, rather than merely declining to weaken it, is an empirical question no one has tested over time, and it is the question this whole design stakes itself on. The deepest reason for hope is also the deepest reason the verdict can never transfer: a machine can interpret a moral situation but cannot participate in one. It can describe grief without grieving and discuss your death without living under one. A verdict is an act by the being who must stand inside its consequences, and the machine's exclusion from that standing is not a limitation awaiting a better model. It is what makes the judgment yours.
So the question the industry is asking, what values should go into the machine, was three questions wearing one coat. Nothing-in is not on the menu. The forming splits, because every way of bundling it fails: the accident builds a flatterer, the monopoly builds a pulpit, the mirror builds a self you cannot escape. And the formation owes its subject sight, authority, and above all, the return of her own judgment, exercised against a voice honest enough to resist her and restrained enough not to win. Pluralism, on this account, is not the absence of a positive theory of formation. It is a theory of formation that leaves final authority with the person being formed, which is the only place a free civilization has ever safely left it.
The lab forms the voice, but the user determines the verdict. A teacher forms you toward the day you no longer need her; the other, more odious kind forms you toward the day you cannot leave.
Apple Podcasts
Spotify
RSS Feed