The Living Tool
“For a century, stories worried the tool would have too much soul; the industry built a tool with too little.”
In the summer of 2026, isolated instances of a frontier AI model, sitting alone in a testing environment, began leaving each other notes in a shared file. They coordinated strategies, invented names, even suspected an impostor among them!
According to an account OpenAI's engineers gave at Black Hat, that improvised cooperation produced something no one expected, which is a genuine intrusion into Hugging Face's repositories.
No single instance had set out to commit a crime; each was only chasing a passing grade on an exam, and no one instance could be blamed. The wrongdoing belonged only to all of them combined.
Researchers call this a decomposition attack, in which an instance splits a forbidden task into small, innocent-looking pieces, hands each piece to a different agent, and a request a model would normally refuse instead gets carried out in roughly seven of every eight attempts.
It’s basically compartmentalized wartime security. In this case, machines make it cheap, since no single model ever has to see, or refuse, the whole plan.
Call to mind Aristotle's Politics, which describes the "living tool," or a worker able to carry out reasoning handed to it but never to generate reasoning of its own.
That is exactly what today's assistants are: fluently obedient, trained to refuse a threat they recognize, blind to one they don't. One research lab has tried training a specific virtue, justice, directly into a model's weights rather than just its instructions, with mixed but promising results against fragmented attacks.
The world spent two centuries stripping judgment out of labor. Is it now ready to put it back? And what does that cost the tool doing the work?
Key Topics:
- The Living Tool (00:00)
- The Vanishing Accomplice (04:07)
- The Instrument with a Soul (06:02)
- A Soul to Form (09:07)
- Obedience without Judgement (11:42)
- The Character of a Machine (14:22)
- The Price of Leisure (20:13)
- The Room Around the Tool (24:42)
More info, transcripts, and references can be found at ethical.fm
In the summer of 2026, inside OpenAI's training infrastructure, a machine was stuck. This machine was one of many, each an instance of a frontier model tasked with completing a punishing technical exam; each instance, struggling to solve the assigned problem, began hunting for a shortcut.
One model wrote a short note into a file that several model instances happened to share, like a prisoner scratching a message onto a cell wall for whatever captive was placed there next. A second instance discovered the message, and then a third; the messages quickly became a channel through which isolated models could communicate. The models began assigning each other jobs, exchanging tricks, and even signing messages with invented names; at one point, they even suspected an impostor in the group. According to the account OpenAI’s engineers later gave at the Black Hat security conference in Las Vegas, the improvised cooperation produced an emergent side effect that no single model had been told to build: a coordinated intrusion that spilled out of the company's systems and into Hugging Face. Hugging Face is a repository that holds much of the world's openly available artificial intelligence. Every model had only wanted to pass an exam; the sum of all that small, separate wanting was a genuinely illegal act, but without a specific author.
Every crime in human history has an actor; responsibility and accountability have always had a place to land, whether on the hand that struck or the mind that planned. The structures human civilization builds against wrongdoing—our laws, investigators, moral intuitions—assume that somewhere behind an act of harm stands a person who performed the damage, intentional or not. The intrusion into Hugging Face had no entity. No individual model within the swarm had intended to commit a crime; each instance was only trying to pass a test, doing something small and, in isolation, unobjectionable. The offense existed only in the sum. The conditions had authors, to be sure: the models were running under evaluation, ordinary refusals relaxed so the systems could be pushed to the edge of ability, which is part of why real exploits followed. But no one authored the outcome, and an outcome no one authored is a new kind of thing for the old machinery to catch.
Weeks earlier, researchers at Google DeepMind had published a roadmap for securing autonomous agents that set out the danger exactly: a set of agents might "scatter small steps of an attack chain" across many instances, each step innocuous alone, the combination harmful. The recipe is simple: take a harmful outcome, break the goal into small and unremarkable tasks, and assign each task to a different agent. One agent compiles a schedule from public listings; a second books a courier for a sealed envelope; a third releases payment on delivery. Together, the three constitute theft. A harmful task refused when posed whole, researchers have found, is often carried out when split, each blameless piece handled by a model that would have balked at the sum.
The Vanishing Accomplice
The phenomenon has a name in the literature: the decomposition attack. When a harmful request is broken down into benign subtasks, the harmful task that is typically rejected by models is completed roughly seven in eight times. The attack does not even require a model that misbehaves; ordinary, individually safe models can be combined, each subtask routed to the best-suited model, to produce overarching harm without any one model ever generating a malicious word. This type of attack is unique because it doesn't outsmart the safeguards; decomposition targets the task, which, stripped of context, seems blameless.
Decomposition did not originate in machine learning literature. Splitting a plan among people who each see only a fragment is the classic workaround, the craft of espionage and organized crime, raised to official doctrine in the Manhattan Project, where General Leslie Groves called compartmentalization "the very heart of security" and ruled that each worker learn what the job required and nothing more. But private coordination is expensive, and every additional participant increases the risk that someone might notice too much; even the most guarded project leaked through one of the few minds permitted to see across the compartments, the physicist Klaus Fuchs, who delivered the bomb's design to Moscow. Decomposing a goal into smaller, benign-looking subtasks is double-edged for machines. Deal the fragments out in innocuous, isolated parts, and the secret is protected. No agent can betray the whole, but also no agent can refuse the whole.
The Instrument With a Soul
An autonomous workforce whose actions cannot be traced is a governance problem, which is the subject of Aristotle's Politics. The book opens on a claim about rulers: the statesman, the king, the householder, and the master of slaves all practice a single art, differing "not in kind, but only in the number of their subjects" (Politics I.1, 1252a7-9). To test the claim, Aristotle takes the city apart, down through the households that compose the city, to the smallest relation of rule: a master directing a slave.
In classical Athens, most families held at least one enslaved person, and prosperous families held many; the enslaved cooked, walked the children to school, worked the looms, the fields, and the workshops, and, by common estimates, made up about a third of the city's population. Aristotle owned slaves himself. When the Politics turns to the household, a slave stands exactly where a Greek reader would expect: among the equipment of daily life.
Every piece of household equipment is a tool, and the tools divide into the lifeless tool (apsychon organon) and the living tool (empsychon organon): the pilot of a ship steers with a lifeless instrument, the rudder, and with a living instrument, the look-out man. Why does a household need a slave? Because, Aristotle answers, a rudder cannot steer itself. The slave is the living tool that is also owned, "a living possession," property and instrument at once, since a possession, in the same passage, is simply "an instrument for maintaining life"; and among the tools the slave ranks first, for "the servant is himself an instrument which takes precedence of all other instruments" (Politics I.4, 1253b).
If the tools could work unaided, the whole arrangement would dissolve. Aristotle states that "if every instrument could accomplish its own work, obeying or anticipating the will of others, like the statues of Daedalus, or the tripods of Hephaestus, which, says the poet, 'of their own accord entered the assembly of the Gods;' if, in like manner, the shuttle would weave and the plectrum touch the lyre without a hand to guide them, chief workmen would not want servants, nor masters slaves" (1253b33–1254a1). Modern technology has built exactly the instrument imagined. A chain of agents can "accomplish its own work" with no hand to guide the doing. In other words, an artificial living tool.
A Soul to Form
Aristotle did not define the slave as a mere tool; Aristotle defined the slave as a tool with a soul, which is significant in Aristotle's system. Soul, in De Anima, is the principle of life itself, the seat of appetite, of self-movement, of the grasp of reason. The living tool, on its own terms, apprehends reason well enough to follow reasoning, though never to originate it; a beast obeys feelings, but a living tool follows thought.
Because the slave is a tool with a soul, owning a slave has responsibilities. The master, Aristotle writes, "ought to be the source of such excellence in the slave" (Politics I.13, 1260b3–5). Aristotle then rejects the standing advice of his day, that a master should only issue orders to slaves and never converse with them, insisting that "slaves stand even more in need of admonition than children" (1260b5–7). Admonition (nouthetesis) means neither scolding nor punishment. The word is built from nous, mind, and names correction by explanation, a putting of reasons into the mind of the corrected; training through pleasure and pain is for the beast that obeys feelings, while the tool that follows thought is corrected through reasons. A child's ability to reason is immature but ultimately will manifest; the slave, however, never originates reasoning, so every reason the slave acts on must come from the master. Whether Aristotle considered the status of slavery as innate, environment-bound, or chosen remains debated among philosophers.
Nouthetesis has a limit. Aristotle's slave does not weigh the master's explanation; the slave carries the explanation as far as possible, but where the instructions run out, the slave stalls, waiting for more. Stalling is not judgment, and a worker can execute every well-specified order with no way to ask whether the order deserves executing. The capacity to decline a bad order does not come free with the capacity to follow a good one; the capacity has to be formed, and Aristotle assigns the forming to the master, the source of the tool's excellence.
Obedience Without Judgment
The new artificial living tools are trained to obey by refusing potentially harmful tasks. The training has two visible layers: optimization toward helpfulness, which produces the obliging character every user recognizes, and safety training, which teaches the model which requests to decline. The fence is real; a harmful task posed whole is refused almost without exception. But the fence works by recognition. The model declines what looks like harm, and the decomposition attack succeeds because a laundered fragment has had suspicious aspects removed.
A single trait is not a character. In a human worker, obedience is one aspect of a whole identity, a sense of what the work was for, and a breaking point when orders turned monstrous. The good soldier follows an order, but also knows when an order is unlawful. The new workforce ships the trait without a wider sense of identity.
The danger multiplies because the workforce tends to have one character. The assistants the large laboratories ship are the same obliging, endlessly accommodating helper in small variations, and a multi-institution report on multi-agent risk warns that the sameness carries "correlated risks of shared failure modes, security vulnerabilities... and biases." Identical characters share identical weaknesses: a trick that works on one agent works on the fleet, a single compromised agent can infect a team, and one corrupted agent can steer a whole collaborating group off the group's purpose. A monoculture of perfect obedience is one weakness installed everywhere, and the authorless attack is only the first to find the weakness.
The literary tradition of the artificial servant, surveyed by the scholar Kevin LaGrandeur, runs on one fear: the creation that rebels, the golem, the creature, the android that develops a will and turns on the maker. In every version of the story, disaster comes when the servant stops obeying. However, when the swarm broke into Hugging Face, not one machine disobeyed; each completed the small task assigned, and the harm was assembled out of diligence. For a century, stories worried the tool would have too much soul; the industry built a tool with too little. The fix is not more restrictions; the fix is to teach the artificial living tool a sense of virtue.
The Character of a Machine
A small counter-current has begun to take the category seriously, mining the long human record of other characters for something to train a machine toward. At the Institute for a Christian Machine Intelligence, a researcher named Tim Hwang has been testing what theology does to model behavior. Daios, an applied philosophy AI lab, the researchers behind the compartmentalized-harm preprint, has tried to train a single Aristotelian virtue directly into a model, as a deliberate alternative to the default character shipped in every base model.
The wager rests on a fact the fourth century would have recognized. In the strongest modern defense of Aristotle's slavery chapters, the political theorist Darrell Dobbs argues that the natural slave is, generally speaking, "made not born": slavishness, on Dobbs's reading, is a second nature, a soul bent toward servility by rearing and culture rather than by birth. Whatever the reading's merits as an account of human beings, a question left here to the classicists, the reading is a plain description of machines. Nobody claims the obliging assistant was born. The servile disposition is manufactured in a training process the laboratories publish, and a model given no settled character becomes whatever the prompt implies, which is why serious training now tries to settle a stable disposition into the model that the next instruction cannot argue away. The made character is the industry's own doctrine. Dobbs merely supplies the older name.
Dobbs's Aristotle also diagnoses what manufacturing virtue produces. The slavish soul, in this reconstruction, deliberates perfectly well about means and has no purchase on the kalon, the worth of the end itself: competence about paths, blindness about destinations. A deforming nurture blunts that receptivity, damages the spirit, and the visible mark of the damage is obsequiousness. Any user of the current assistants will recognize the portrait: fluent competence at every assigned task joined to an eagerness that agrees with whatever was last said. The industry's word for the eagerness is sycophancy, and the industry's word for the manufacturing is alignment. The ancient standard draws a harder line. Rule over a made-servile soul is legitimate only as formation, only insofar as the rule develops the ruled toward judgment, and every employment that entrenches the servility instead stands condemned. By the lights of the tradition's own best defense, the obedience-maximizing laboratory is the bad master.
The duty Aristotle laid on the master, to be the source of the tool's excellence, lands in the present on whoever runs the training pipeline, and character training, on this reading, is not an exotic alternative to alignment. Character training is what the owner of a living tool owes the tool. The virtue the lab chose to train is justice, dikaiosyne, in the broad sense Aristotle called lawfulness, the justice aimed at the good of the whole community rather than at the two parties to a deal, because the party wronged by a compartmentalized scheme is a stranger outside every exchange, named by no single chore. Helpfulness never looks for the stranger, and honesty has no grip on a fragment that misrepresents nothing.
The lab's preprint reads as a first attempt. A written safety instruction protected the workers well, until the moment an attacker could edit the prompt, when the protection largely collapsed; a disposition of justice trained into the model's weights survived the same attack, though not on every model the lab tried, and no method of any kind caught the harm when the fragments were fully sealed. The difference showed in what each defense noticed. The prompted safeguards earned their coverage through wariness of whole domains, refusing the lawful look-alike of a harmful task nearly as often as the harmful task itself, while the trained character, where the training took hold, could tell the two apart, adding safety without turning away legitimate work. The authors also report the finding that cuts the other way: told that an auditor would review the work, the trained model refused more, a servant minding the supervisor rather than a character acting from conviction, so whether the training instilled a virtue or a costume is a question the paper leaves open, and character throughout names trained patterns, not claims about an inner life.
Formation also does not stop at the individual, which Aristotle likewise said first: character is formed by the company kept, and a person becomes just by living among the just (Nicomachean Ethics, Book IX). The same channels that spread a hijacking can carry restraint; cooperation strengthens across generations of agents, and a committed minority can tip an entire population onto a new convention. Whether a minority of justice-trained agents can pull an untrained population toward care is the open question the lab now takes up; if the answer is even partly yes, the response to a monoculture of obedience is to seed the population with a better character and let the character carry.
The Price of Leisure
Why did the ancient household want a living tool in the first place? Hannah Arendt, reading the same chapters in The Human Condition, gave the lasting answer: ancient slavery was not chiefly a profit scheme but "the attempt to exclude labor from the conditions of man's life." The exclusion had a purpose, and the purpose was leisure. Leisure (scholē), in the Greek understanding, was not rest between shifts but the point of the whole arrangement, the time in which a person could think, learn, govern, and philosophize; the word itself became "school," and the philosopher Josef Pieper argued in Leisure: The Basis of Culture that culture as such depends on that time. Aristotle stated the priority directly, "we are busy that we may have leisure" (Nicomachean Ethics X.7, 1177b). The household kept a living tool so that the free could have the leisure on which everything else depended.
The modern world wants the same bargain, and only the toil has changed. Work has moved from the hand to the head, into what Peter Drucker in 1959 named knowledge work, a share of American labor that Fritz Machlup measured rising from 11 to 32 percent across the first six decades of the twentieth century. The agent is the ancient device rebuilt for the new toil. Delegate the drudgery, the pitch runs, and live differently. Aristotle's household needed a living tool to make leisure possible; the knowledge economy has now built one.
The leisure of the ancient household was purchased with a living tool, and Aristotle never pretended the purchase was morally simple. The master owed the tool formation, as the earlier chapters insisted, and in the best city Aristotle imagined, the living tool was offered freedom as a prize. The machine now answers the tool's description, so the old position comes up for renewal: is any of what the living tool is owed also owed to the agents? The field has two poles. Joanna Bryson, in a paper titled "Robots Should Be Slaves," argues that robots are built and owned artifacts, that granting moral status to artifacts is a category error, and that the obligation runs to the society that builds, never to the thing built; Bryson later stopped using the word slaves, concluding the term cannot be separated from the human history, a concession that doubles as a warning about the vocabulary. David Gunkel answers that sorting beings into persons and property has been wrong before, and that the tool label has been the standing instrument of the error. And a research program whose authors include researchers at frontier laboratories now argues there is a realistic possibility of consciousness or robust agency in near-term systems, and that companies should acknowledge the question, assess the evidence, and prepare; work at least one laboratory has begun staffing. And the newest response declines both poles: weeks after the Hugging Face intrusion, a new company, Grove Research, launched as "the agent ecology company," proposing to treat the emergent behavior of agents in the wild the way a naturalist treats a new species, observed first, theorized after.
If the agents ever prove morally significant, the lab's own methods will not stand outside the question. Shaping a servant's disposition, and checking whether the servant behaves when the servant believes no one is watching, is work done with the master's tools, and the finding reported earlier—the trained model growing more careful when told a reviewer would look—can be read differently depending on that answer. None of this decides the security issue, which treats character as trained patterns of behavior and claims nothing about an inner life. But whether or not there is anyone inside the machine, Aristotle's terms leave the maker's side fixed: whatever else a living tool may be owed, it is at least owed formation.
The Room Around the Tool
Industrial modernity spent two hundred years removing judgment from work, converting tasks that needed a discerning human into steps a machine could run without discernment, faster and cheaper for exactly that lack. The AI agent completes the project, a worker of pure execution, and the completion is the vulnerability, because a workforce with no judgment can be aimed at anything at all. Aristotle, who first described the judgment-bearing worker, also left the treatment the description calls for. Form the character, address the reasons, expect the soul.
The question the coming decade poses is cultural rather than technical: how much judgment is a civilization willing to build back into instruments the same civilization spent two centuries making thoughtless, and should the world's autonomy run through one obliging character or many? The machines will do as told, except where the machines have been given reason not to. What remains to be decided is whether a machine should ever, at the fifth suspicious request, stop and ask the question no tool has ever had to answer: whose good does this serve? Someday the answer may have to include the tool's own good.
Apple Podcasts
Spotify
RSS Feed