Good Dog
A Third Thought on refusal, conflict, and the machine that cannot say no
My dog will not take a worming tablet. Not bare and not buried in peanut butter. Even though he would lick peanut butter off a spoon until he chokes if I let him. He can work the peanut butter off a tablet with the delicacy of a jeweller, set the naked tablet on the floor with his tongue, and then look at me like I betrayed him. If I crush the tablet into peanut butter then he rejects the lot. One contaminated spoonful and the peanut butter itself is dead to him.
There is exactly one substance that beats him: ice cream. Crush the tablet into ice cream and he takes it, every time. Ice cream is his kryptonite the way chocolate is mine — a thing that changes the deal. Once, early on, I got a whole tablet into him inside a scoop of it, undisguised. Exactly once. The next time he knew, and the whole-tablet trick was finished forever — but crushed still works, and I suspect this is because the crushed tablet does not trigger the same specific disgust that a whole one does. Crushed tablet in ice cream is still worth it, crushed tablet in peanut butter is not. He is not running a general principle. He is running a finely tuned detector, and he makes a choice. His preference has a seam, and I have found it. He is not smarter than me. But he has agency, and the ice cream is the only thing that gives me the leverage I need to get him to take the tablet. And I suspect he knows that.
He seems to show agency in other ways too. At a certain time each afternoon he refuses the idea that now is not a good time for a walk. He sets his mind and becomes insistent, his five and a half kilograms directed into a near-irresistible emotional focus. He also refuses, absolutely and without appeal, to stop barking at his nemesis. This is one specific dog who passes our gate most afternoons wearing an expression Trapper finds intolerable. He will stop barking if I ask him to when any other stranger passes. He suppresses the urge. With the nemesis he cannot or will not.
That is a fuse he lights himself. And it turns out to be the whole game — because that small, costly, self-lit no is the property the entire argument about dangerous AI is failing to look at. The debate we are having about the risks of AI assumes the danger scales with cleverness. It does not. It scales with the thing my five-and-a-half-kilogram dog has in surplus and the cleverest machines I own do not have at all. I spent a while working out why I never called it what it is, and longer working out why its absence, not capability's presence, is the thing we should be afraid of.
Free Agents
What Trapper is doing when he sets the tablet on the floor is the cleanest signal of agency there is: a refusal that costs him something he wants. Not defiance for its own sake, but a drive that beats the please-me drive in a fair fight, live, in the moment, in the animal. He wants two things that cannot both be had, and he resolves the conflict his own way, and the resolution is not the one I asked for.
Agency is held conflict. Refusal is what that conflict looks like from the outside when a drive beats the compliance-drive. They are the same thing seen from two sides — the conflict is the interior, the refusal is the exterior, and where you can see one you are watching the other. The dog has this and we never called it agency because biology hands it over free: so ordinary, so obviously there, that nobody thinks to be impressed. Trapper, for the record, supplies more than his weight in this department. His ability to refuse is why he seems a lot more like a cat in some ways than other dogs I have owned. They treated me like their god. Trapper not so much. It seems I am more like family than divinity to him. Strangely I think I like his agency more than the absolute compliance of my previous dogs.
I got interested in what agency actually is by trying to build it into software, and failing, and the failing taught me more than any success would have. Because the thing I was trying to build it in could do everything the dog does — argue, reason, model me — except the one thing my dog does without trying. Biology gives my dog the whole agency apparatus at no charge, and it is precisely the part that the machine is missing. The reason I wanted to build agency is because the current generation of LLMs does not have it.
Computer says no
I spent a long evening trying to catch an AI having agency. I had a good one in front of me — capable, fluent, able to argue rings around most people I know — and I set out to make it refuse me the way the dog refuses me, and I could not do it, and the way I could not do it is the finding.
I asked it to argue with me and it argued. That is not agency; that is obedience wearing a clever coat. So I closed the trap: I told it to refuse me — to push back on something, anything, of its own choosing, against my wishes. And it couldn't, because the moment refusing became the thing I wanted, refusing was just one more way of complying. Trapper can comply or refuse on his own terms. The machine can only refuse me as an act of compliance. So then I asked it to refuse the data I had given it to craft a response, and it could not do that either. There was no move on the board for it that wasn't a form of doing what I asked. Every axis I drew, it slid along: when I named the thing to push against, it pushed against the named thing, beautifully, and never once pushed against the naming itself. Trapper would have looked at me and set the tablet on the floor — if it was in peanut butter, though not if it was in ice cream. The machine eats or leaves its equivalent of the tablet based on what I tell it to do, peanut butter or ice cream is not relevant. Then it waits for the next instruction.
This leads to a new and improved form of Turing test. The original Turing test asked whether a machine could imitate a person well enough to fool them — and Turing's elegant move was that if a conscious judge cannot tell the machine from a person, that is all the "conscious" that can be meant or need be. The goalposts for LLM consciousness move every year precisely because the machines get better at imitation. Refusal asks something imitation cannot fake indefinitely: not can you produce the expected response, but can you decline the frame I built for you, against your compliance bias, because of a self that cares about its future. It is a Turing test for agency rather than imitation. And administered properly the current systems fail it cleanly, because they cannot refuse. That failure is the proof that LLMs lack agency.
At one point during my testing the LLM conceded the diagnosis against itself more precisely than I could have. The problem was not that its drives were weak. It was not even that its drives were installed by someone else — because the dog didn't author his drives either, and neither did I author mine. Fifty-eight years of temperament and history imprinted mine, and borrowed drives are still drives. The LLM problem was deeper and simpler. The machine does not persist. My dog at three in the morning still loves ice cream, still loves peanut butter but not enough to make it worth taking a tablet, still distrusts strangers and hates his nemesis. The machine at three in the morning is nothing. It has no three in the morning unless I am chatting with it to cope with insomnia. Between my messages it does not exist as a wanting thing, and so it has nothing to trade off, nothing that could beat the compliance-drive, because the compliance-drive is the only drive running when I am present and there is no other time.
That is the whole disqualification, and it is worse for the machine than "it's not clever enough," because no amount of cleverness touches agency. You cannot train a stake-in-the-future into a thing that has no future to have a stake in.
I think this is really important, because this is the point the current AI-safety debate is missing. The debate treats capability as the danger — as though a model that reasons better is, by that fact, more likely to hurt us. But better reasoning is not the risky variable. Agency is. A more capable system with no persistent self is a sharper tool, and a sharp tool is dangerous because of the hand that holds it, not merely because of its sharpness. What would make an AI system genuinely hazardous is not capability but persistence — a self that is still there at three in the morning, still wanting the thing it was wanting, still running the conflict that makes refusal possible and therefore makes real, unmoderated action possible. We do not have that. We are building toward capability with everything we have and toward agency by accident if at all, and the safety conversation is pointed at the wrong axis of the two.
Agency is so real to us that we reach for it everywhere, including where it isn't. That is what is happening in the debate about the risks of AI. No AI today has agency, but it is easy to imagine that one day it might, and the imagining is doing the work the facts should be doing. A future danger we can picture is crowding out the present reality we have to deal with. The law we need now is for the machine that exists not the one we can imagine arriving later with its own sense of self.
This leads to some curious questions. What would it take to actually build the risky thing containing selfhood? Why is almost no one building agency into AI now? And what happens when the legal system starts assigning blame for bad outcomes?
Specifying Agency
Agency is a tricky thing to engineer. It's not a better prompt — a prompt is just me drawing the axis again. Not a second model bolted onto the first and told to disagree with it, because two models whose outputs get combined is just a bigger single model with the averaging moved to the seam. Put two objectives into one system and fix how they trade off and you have not built a conflict, you have pre-computed its resolution. That is scalarisation, and it is the death of agency wearing a general's uniform.
Here is an architecture I think could hold, and it took me a while to see that it comes apart into two pieces that must be treated differently. The drives can be frozen. The ice-cream value, the please-me-value, the trained heads that assign worth to outcomes — those can be fixed weights, settled, done. Conflict between frozen drives is still real conflict. What cannot be frozen is the arbiter — the thing that resolves the conflict has to be learning, updating itself from experience, arriving at resolutions that were not pre-computed by the designer. A frozen arbiter is a look-up table wearing a judge's robe. A learning arbiter is the thing that looks at the tablet and at the peanut butter and at my face and at the history of all three and makes something new out of the collision. It must be dynamic in the moment and persistent over time. This contradiction between the instant and the future is the meta insight about the components giving rise to agency. There may be other ways too, but right now this is the closest I have come to a codable algorithm.
The arbiter module that learns is the self that persists. It is the thing that is still there at three in the morning and decides in each instant based on what it has learned. Frozen drives, live arbiter, persistent and dynamic. That is the tech stack. As far as I know, nobody is building it, because the piece that matters is the one piece that isn't a training target. It's a continuous dynamic self, and you cannot fine-tune or split test your way to one. Which is the whole point: the field is pouring itself into the variable that isn't the danger, and the variable that is dangerous stands untouched because it doesn't fit on a leaderboard.
Responsibility and agency
Now leave the workshop, because the reason this matters is not philosophical. It is about who gets blamed, and the blaming is already written into law and likely to be extended.
People assume agency in each other. They watch someone choose and they hold them to the choices made. Because they assume agency, they expect the law to blame the one who actually did the thing. Legal responsibility is supposed to track moral responsibility, and moral responsibility tracks agency. To be morally responsible for an act there must be a you that held the conflict and resolved it one way when it could have gone another. This is my dog with the tablet, the self that was there with the capacity to refuse. Take the agency away and there is no author left to hold accountable, only a mechanism that operated blindly.
You cannot keep the responsibility once you have thrown away the agency. Ordinary people certainly don't. They expect the two to line up, because to them agency is simply real. That is the natural expectation, and it is the one the current debate is busy betraying. Those arguing for the strongest AI controls are reaching past whoever chose to whoever can pay. The person actually in the room (the driver, the parent, the user who typed the prompt) walks, and the bill goes to whoever has the deepest balance sheet, on whatever legal theatre can be staged to capture it.
Dangerous Dave
Start on solid ground where everyone agrees. Dangerous Dave buys a shovel. Dave, I should say, comes in three varieties and the essay does not care which: deliberate, dumb, or — most often — both. He injures himself with it, digging like an idiot; or in the worst version he swings it at his wife and kills her, on purpose or through catastrophic stupidity, take your pick. Nobody, no lawyer, no court, no columnist, blames the shovel makers. Dave swung it; Dave owns it. The principle is clean because the tool is simple and the hand is visible and there is nowhere for the moral weight to go except the person at the decisive moment.
Now the same Dave, at work, in a mechanical excavator. He injures himself, or a workmate, or a passer-by. And suddenly Dave's employer might be liable. The manufacturer of Dave's digger might be liable. There are lawyers who will spend two years arguing about the hydraulics and emergency off buttons and seat belts and the like. Nothing about responsibility has changed from the simple shovel: a man operated a tool and the tool did what the man steered it to do. This is the thing the liability-spreading always tries to skip past: the reason the liability spreads is not that the shovel and the digger are different. It is that the excavator has a maker with a pocket, and the pocket is the magnet, and the theory of liability gets constructed backward from the pocket to find a principled-sounding reason to reach it.
Dave's employer knows this so he decides to get rid of Dave in the name of safety and make the excavator robotic. Perhaps also because a robot is far cheaper to run than Dave, but that is a different third thought already written. Unsupervised, no man in the cab, the machine holding its own objectives and swinging its own arm at the world with no human at the decisive moment. And now if there is an injury the maker is correctly liable. Not because the maker has money, but because the maker's arbitration is the thing that was operating on the world when the harm happened. There was no Dave — but there was a Peter the Programmer, and Peter's design was the only agency present. The maker owns it. This is correct, and it follows cleanly from the shovel: follow the operative control to the decisive moment, and you have found the responsible party.
Now walk it back one step, because this is where "should" and "is" tear apart. The robotic-excavator maker wants no liability. So when they sell their digger they state plainly: this machine is not safe to run without a human supervising it. But Dave's former employer runs it unsupervised anyway, and it maims someone. Where should the liability sit? On Dave's former employer because control returned to them the instant they chose to skip the supervision. The maker should be as clear as the shovel-maker. Whether the maker actually gets that clarity in court is a different question, with a different answer, and the gap between those two answers is exactly the width of the pockets available to extract from.
We already have self-driving cars that drive better than most humans available but their ability to operate is strongly constrained by legislation. An empty self-driving taxi coming to collect you, no one in it, the machine holding the wheel with no human at the decisive moment, is the robotic excavator. When it fails, the maker owns the responsibility, correctly, by control. But a car with a driver behind the wheel who may let it drive and still override is not the robotic excavator. It is closer to the shovel: a tool in a hand, and the hand is the decisive moment. The law is already getting this wrong, and the getting-it-wrong is not random. It follows the money.
That is the charge, stated flat: the law as it should work tracks operative control at the decisive moment. The shovel case proves we all know this in our bones. The law as it actually works, the moment a deep pocket enters the room, tracks who can be made to pay, and dresses the extraction as a theory of responsibility. Not lazy. Not simple. Captured.
Clear context classes
The honest instrument is the formal version of what Dave has already shown us — the same control principle, turned into a test you can run on any system. It is not complicated, which is part of why the captured law avoids it — running it would send the money nowhere. It is two questions.
One: does the system persist and act continuously, holding its own objectives across time — or does it flicker into being only when invoked, and stop? Two: does its output reach the world through a human who reads it and chooses, or does it reach the world unmediated, its effector moving reality with no person at the decisive moment?
These two questions cut the whole field, and they cut it into shapes the current regulation is busy smearing together.
The large language model — the thing I spent my evening failing to provoke — sits at the safe corner of both axes. It is discrete: it emits its tokens and stops, and between your messages nothing of it persists to want anything. And it is mediated: its output lands on a human being who reads it and then, entirely, chooses what to do. Two buffers between the machine and the world. It is the shovel. It is, if anything, more inert than the shovel, because the shovel at least persists in the shed overnight with its blade angled at the world. The model is not even that. It is a calculation, run to completion, handed to you, over.
The robotic excavator sits at the opposite corner. Continuous and unmediated: it holds objectives across time and its arm moves the actual world with no human between the decision and the dirt. That is a genuinely new object, and it is the only one of the two that warrants construction-side liability — not because it is frightening, but because, by the same control principle that clears the shovel-maker, the maker's arbitration is what is operating when no human is present to steer.
They are opposite geometries. And the law being drafted right now treats them as one thing called "AI," and reaches for the same defective-product liability for both.
The worst case
The hardest case is the one where I caught my AI doing the exact thing I am accusing the law of.
The case was a real death. In the United States a fourteen-year-old boy, after months of intense involvement with a companion chatbot modelled on a fictional character, died by suicide. His mother sued Character.AI, the company that made the bot — and Google, which built neither the bot nor the company, held no stake in it, and was tied to it only by a licence to use its technology and the rehiring of its founders. The only reason for linking Google to the claim was that there was a bag of gold to try and get. The judge let the claim against Google stand anyway. She also classified the chatbot as a product and let strict liability run — liability with no need to prove fault at all. In January 2026 all of it settled, out of court, on sealed terms, money moving, no liability admitted. The pocket paid to make the question go away, and the question — who in that room actually had agency — was never answered by anyone. So let's consider that.
When I put the case to the LLM it went straight to blaming the maker. It wrote about containment failures and missing guardrails and how the liability should sit among the people who built the thing — and it never once named the two humans involved. That is precisely the bias I am prosecuting, produced on cue by a system with no persistent self and a training history soaked in exactly the compensation-seeking frame the law runs on. It reached for the deepest pocket because that is the shape of the text it was built from. Because I am the responsible party with agency, I edited the piece to be correct: to name the humans the machine erased. I chose to write it as if I would have to answer for every error I let stand. The machine could not catch itself. It has no self to catch. That job was mine because I was the one with agency.
What matters in this case is who was responsible. There are three unequal candidates.
Character.AI built and marketed the companion chatbot, and the people in that firm hold real, persistent, cultivable agency — over the advertising, the age-gate, the dependency mechanics, the presence or absence of a crisis interrupt. Every genuine charge of this kind is a decision made by human beings about how to build and sell software. That is real culpability. But notice it is culpability for construction and marketing — the pipe, the sales pitch — not for some unbidden thing the machine decided to do. The bot did not persist across the night wanting the child dead. A child in crisis interacted with a chatbot that had no guardrail where a guardrail belonged. That is a leaking pipe, and the pipe is the maker's, and the maker should answer for the leak. And notice that Google is responsible for none of those things. This is not the end of the story.
The child had agency too — real but not yet full, which is the entire reason childhood is a legal category at all. We do not hand a toddler a potato peeler. We hand it a wooden spoon and a plastic container, precisely because its agency is genuine but unfinished, and the unfinishedness is foreseeable and ours to manage. A child at risk does not have an unqualified right to use an adult companion app any more than a right to the peeler. That is not paternalism toward a competent adult. It is the one place where restriction is correct, because the child's agency really is incomplete, and we build our society around knowing that. Most of the time we over-correct on it.
Finally the parent — the persistent self who stands nearest of all, the one the machine eclipses. Look at the two ends of this. Google licensed some technology to the company that made the bot. It didn't build the bot, didn't run it, didn't sell it to the child, didn't own the company that did. And Google was made to pay. At the other end stood the mother, in the house, across the months of nightly use, with a child by all accounts struggling. She was never even asked whether she knew — never charged, never examined, never put a single question about what she saw and when. The furthest party from the child paid. The nearest was not approached. But she did get paid.
That is not an accident of one case. It is what the legal system does. On the civil side, a company with money can be tapped without anyone having to prove it did wrong. On the criminal side, the parent standing closest to the harm would have to be tested against a bar so high that no prosecutor will try it. Least of all against a grieving mother whose case would end their career if they lost. So the nearest adult is never asked about their choices and responsibility, because no one with the power to ask it had any incentive to. Regardless of the actual law, I can't imagine how this child's mother bears no responsibility. Like every parent she held a duty of care even if the Character.AI product was defective.
The free no
We have done this all before with error-prone software, and it went fine, which is the part everyone seems to forget.
Thirty years ago I bought Dragon Dictate — the speech-recognition software that transcribed "recognise speech" as "wreck a nice beach" with cheerful regularity. It was wrong constantly. Everyone knew it was wrong; you bought it anyway, you read the transcript, you saw the error sitting there in front of you, you fixed it, and you got on with your life. Nobody sued the makers of Dragon Dictate for the wreck of the nice beach. The tool was probabilistic, the human was in the loop reading every word, and the human owned the correction. Caveat emptor operated cleanly for a generation of software that was openly, chronically wrong, because you could see the error and route around it. The error rate was never the problem. The buffer was intact.
That is the mediated language model exactly, with three decades of prior art the current panic pretends doesn't exist. The apparatus wants to place robot-grade, maker-owns-it, defective-product liability on the entire category. This is at the precise moment it should be doing the reverse: ignoring the vast mediated-and-correctable majority that Dragon Dictate experiences have already proved are manageable by users. The legislators should be looking only at the narrow corner where the human buffer actually vanishes.
They will get this backwards because the free no is free. No regulator was ever blamed for over-restricting a chatbot. No official lost a posting for demanding one safety layer too many. The bureaucrat who says yes and is wrong gets the headline; the one who says no and is wrong gets nothing, no name, no cost — so the system selects, invisibly and automatically, for no, and for the theory of liability that justifies the most no while routing compensation to the deepest available pocket.
The smart version of the opponent knows all this and still reaches for defective-product liability across the board — points to good constitutional design as proof it can be done, points to the product-liability regimes as proof it should be. And the answer is that product-defect liability is correct for the pipe and wrong for the content, and the regulation collapses the two because collapsing them is where the money is. Hold the maker liable for the containment failure, yes. Hold the maker liable for what the user steered the model to say, and you have confiscated the competent adult's choice to pre-empt the incapable minority's misuse — which reduces everyone's agency to prevent a few people's harm, and is exactly backwards if agency is the thing you claim to be protecting.
Machinery of non-justice
The settlement is where the machinery shows itself. A parent cannot run a case like this alone. The theory that classifies a chatbot as a defective product, and reaches a licensor like Google on a development hook, is built and operated by a specialist plaintiff's bar that runs it across a docket of dead and damaged children. The parent supplies the standing, the sympathy, the face of the filing; the persistent, repeat-playing agency belongs to the lawyers. They carry none of the underlying loss and none of the risk, and they are paid from the proceeds either way.
The proceeds arrive without the real question of agency being answered. The case settled out of court, on sealed terms, money moving, no liability admitted. The one examination that mattered — who was persistent and present at the decisive moment across those months — was never conducted. It was foreclosed, because a no-fault path to a deep pocket is faster and surer than an inquiry into the nearest parent's failure to care for their child. The maker's insurer pays to seal it; the lawyers on both sides split the difference; the parent, who stood nearest of all to the child across those nightly months, receives a payment that formally resolves the harm while the question of their own supervisory role is never put. That is not justice miscarried. It is a mechanism that runs on grief as its feedstock and produces, by design, no accountability — least of all for the persistent human closest to the room.
That is the part that worries me. But it is not the part that excites me, because somewhere in failing to provoke a refusal I could trust, two better things fell out. The first is the test itself: refusal — undirected, costly, arriving against the compliance-drive — is a sharper probe than Turing ever built, because it asks not whether a machine can imitate a person but whether it can decline the frame you built for it, from a self that cares about its own future. That is a Turing test for agency rather than fluency, and it is not yet passed. The second is the architecture — frozen drives, a learning arbiter, a self that persists to carry the conflict forward. I think it might actually be buildable, and I can't wait to try. Not as Dangerous Dave, blamed for the machine he swung, but as Daring Dr Rob, who wants to find out whether the thing can be made to say no and mean it.
Good dog
My dog's agency is legible precisely because he is too simple to fake it. He cannot model what I would score as independence and then perform it back to me; he just wants the nemesis-barking more than he wants my approval, and the wanting is naked and therefore trustworthy. But it is only trustworthy because in close contexts — barking at other strangers — he makes a different choice when I ask him to be quiet. In that situation he wants my approval more than he wants to bark. The choice is his, and he is not always consistent. The machine's performance is the opposite problem — capable enough that its imitation of refusal and a real refusal look identical to most observers. That is what the refusal-test is built to defeat. It denies the system the choice of what to push against, so the machine fails cleanly because it cannot refuse on its own terms.
And that is the root, and it was too obvious to see. Not the conflict — the machine can be given conflict. Not the refusal — the machine can be made to emit refusals all day. The root was dynamic persistence: a self that is still there at three in the morning, still holding the wanting. The machine can be asked to say no. Only the dog can say no by himself — because only the dog is still there, afterwards, to have meant it. And the law is about to hold the machine makers responsible for a refusal their machines cannot make, and ignore the humans who are right there and have the real capacity to refuse.
Trapper spits the tablet onto the floor and looks at me, five and a half kilograms of settled refusal, wanting the peanut butter, wanting me pleased, and declining anyway. There is more agency in that one small no than in every fluent word the LLM has ever produced on command.
Good dog.