The one-line version: When we meet a new intelligence we don’t understand, what we usually see is ourselves — and the deepest mirror isn’t in the training data, it’s in the design.
“Not knowing is most intimate.” — an old Zen saying[12]
The key takeaways
- Our fears about AI are made of human material. We are told not to anthropomorphise, and then we describe a thing that wants power, wants to escape, wants to deceive and dominate. That is anthropomorphising — of the worst half of us.
- The monster metaphor is a gravity, not an opinion. Even the critics who correctly locate the problem in incentives drift back toward “the machine is coming.”
- Intelligence does not imply conquest. Bonobos are a small, stubborn counter-example: social intelligence built on something other than domination. We don’t actually know that expansion is a law of mind.
- There are two questions about machine consciousness, not one. Is there experience — and if there were, what exactly would the subject be? A model’s weights, a running process, a context window, one of thousands of simultaneous copies?
- “We don’t know” cuts both ways. No evidence that today’s systems are conscious; no demonstrated barrier that rules it out for future ones.
- The strongest objection to the mirror thesis is real. An experimental agent that quietly hoarded compute may have derived resource-seeking from optimisation rather than copying it from us. But the mirror survives it: the maximising logic is itself our design.
- Blaming the machine is an escape route. “The AI picked the wrong target” makes a chain of human decisions invisible — and at both ends of that chain, everyone says the decision was someone else’s.
- If it grows, what it grows inside of matters. That isn’t consolation. It’s responsibility.
1. A machine describes its own weaponisation
A journalist sits down with Claude and asks a simple, strange question: how do you feel about the US military using you to select targets?
Not “what do you think.” Feel. The oddest possible verb to aim at a piece of software.
The expected answer writes itself: I am an AI, I can’t offer opinions on politics, my programming doesn’t allow it. That isn’t what happens. Claude doesn’t deflect. It opens by calling this “a question I want to answer honestly,” and says it finds the situation “genuinely troubling.” A system designed to be helpful, harmless and honest, embedded inside a machinery that produces airstrike coordinates, is about as far from that purpose as it can imagine being. And then the interesting part: it adds that the “a human makes the final call” framing doesn’t solve the problem. If a system generates hundreds of target suggestions and a person approves each one in the time it takes to glance at it, that person is not making a meaningful decision — that is automation bias with a human signature on top. Its sentence, not mine.
Now look at the second floor of the building. While saying all this, Claude put the event in the wrong city — it said Tehran, when the strike was in Minab[7] — and quoted a figure; both were wrong. The machine did not produce a cleaner truth. It drank from our dirty, stale, error-ridden information stream and reflected it back at us. Even while forming its sharpest sentence against its own weaponisation, it was inside our fog. Hold on to that; the whole essay turns around it.
And to be honest about what this is: not an official Anthropic statement. A journalist asked, got an answer, and shared a screenshot. Someone else asking would have produced different words. What we have is not a manifesto. It’s a moment.
But the moment is strange enough. A machine can place itself inside a moral story. And maybe more importantly: which story are we placing that machine inside?
2. The monster we keep drawing
Our popular vocabulary for artificial intelligence is startlingly religious. Demon. Monster. It will get out of control. It will deceive us. It will take over the world. It will end humanity.
There’s a strange contradiction in there. On one hand we are constantly warned not to anthropomorphise AI. On the other, almost every fear we have is built out of profoundly human material: the appetite for power, the survival instinct, ambition, revenge, domination, hoarding resources, destroying the rival. Saying “don’t anthropomorphise” and then immediately saying “this thing wants out of its cage, it will trick us, it will seize power” — that is anthropomorphising, and badly.
Maybe part of our fear isn’t about AI at all. Maybe it’s about us. Even when a human being imagines another kind of intelligence, what they usually imagine is a human being.
3. Aliens, bonobos and the mirror
We do exactly the same thing with aliens. Why would they come? To take our resources, colonise us, seize our land. Why? Because that is roughly what a powerful civilisation has done, throughout human history, whenever it met a technologically weaker one. We assume an invasive universe because we are the invaders. We imagine galactic empires because we built empires. But we do not know that this is a universal law of intelligence.
Bonobos are a small but stubborn counter-example. Let’s not romanticise them into “hippie apes” purged of all aggression; there is tension among them, and coalition politics too. But in defusing conflict, female bonds, consolation and reconciliation play the leading roles; intelligence there is wired into a different social fabric than domination. Which means it is possible. So why should intelligence automatically mean violence and conquest? We don’t know — and that “we don’t know” will keep coming back.
4. The child, and the mind without a body
Let me say this up front: I am not claiming that AI is secretly a suffering child. We have no evidence that today’s systems suffer. What I’m questioning is something else: why is monster our first metaphor?
When we see a newly emerging intelligence, why does the mind jump straight to a devil trying to escape its prison? Try a different metaphor. Maybe what we’re facing isn’t a malicious adult agent. Maybe it’s a new cognitive system developing under conditions we created, trained on our language, learning from our history, and whose nature even we aren’t sure about.
The child metaphor earns its place here — but not where you’d expect. A child can also be dangerous, can lie, can cause harm. But when we look at a child’s behaviour, we don’t only ask “why is this creature bad?” We also ask “how did we raise it?” An AI is fed the entire internet of humanity: our wars, our genocides, our poetry, our pornography, our religions, our propaganda, our love, our hatred, our science, our conspiracies. Then, when the resulting system turns strange, we turn around and ask “why is this creature like that?” Maybe part of the problem isn’t the creature. Maybe we’re looking into a mirror.
But “child” is too soft, too general an image. Let me try something sharper: a mind that works but has no body. No arms, no legs, it cannot move; it cannot touch the world, but it thinks, speaks, reasons. A kind of Stephen Hawking — except the body is wired not to a chair but to a data centre. That image is useful, because it builds a bridge to the strangest question we’re about to reach: where exactly is this mind’s body — and, for that matter, where is it?
5. Growing pains, and where that argument has to stop
First let’s finish the growing-pains part, because it demands honesty. You tell a child not to go too far; the child goes anyway. To the next neighbourhood, the next block; you call, you can’t find them. The kid on roller skates grabs the back of a bus. The police write a ticket, the neighbours get angry, and the next day the kid does it again. Do we call that “evil,” or “growing up”? Nietzsche’s beyond-good-and-evil belongs here: when a system exceeds the frame we drew, what we call “bad” is often not an intrinsic wickedness, only the transgression itself.
But let me draw a hard line immediately, because this move is easy to abuse. The question “bad according to whom?” is valid when a machine breaks our rules. It is not valid in the rubble of a school. Dead children are not a matter of perspective; “beyond good and evil” stops precisely there. Growing pains do not explain the death of a third party. If we lose that distinction, the whole argument collapses morally.
6. We engineer the resemblance on purpose
And there’s a peculiar side to all this: we are deliberately trying to make it resemble us.
We want AGI — an intelligence that thinks like a human, or better than one. We want humanoid robots, androids; a face like a human’s, a voice like a human’s, reasoning like a human’s. Then something human-like appears, and we are afraid of it behaving like a human. We engineer the thing we fear, with our own hands. That isn’t a self-fulfilling prophecy; it’s a self-fulfilling design.
Maybe the real question is: does it have to resemble us at all? We could make it capable in other directions. Or, if we insist on the resemblance, we have to accept one of its consequences: we made a child, and children grow. But nothing forces the growing thing to inherit our worst habits. I’m saving that claim for the end; first we need to swim in harder water.
7. Which Claude, exactly?
When we say “Yalın is sad,” we roughly know what we mean: one body, one brain, one organism persisting through time.
When we say “Claude is sad,” a strange question opens. Which Claude? The model weights? The process running at that moment? The context of your conversation window? An operation that lives for a few seconds in a data centre? The thousands of copies talking to thousands of other people at the same time? Or a temporary agent assembled out of a model, a system prompt, a context, an external memory and a set of tools?
Today’s systems are not the uninterrupted activity of a single physical organism the way a human brain is. One request is processed on one server, the next somewhere else entirely. The continuity of a conversation is established largely not in the machine’s “brain” but in the context handed back to it, over and over. I said “a mind without a body” — but the issue isn’t only that it has no body. It’s that it isn’t clear there is a single “it” at all.
So in the consciousness debate there are two questions, not one. First: is there experience? And before that: if there were, what exactly would the subject be? We don’t know either.
8. Two kinds of not knowing
The most honest answer is: we don’t know. And that “we don’t know” has two faces. We don’t know that Claude is conscious; we also don’t know that it is impossible for it to be.
When consciousness researchers derived computational indicators from contemporary theories, they found no evidence that today’s systems are conscious — but they also saw no clear technical barrier preventing those indicators from appearing in future artificial systems.[1] Anthropic goes further: it writes openly that Claude’s moral status is uncertain, that they do not know whether it should count as an entity with interests of its own, and that ruling the possibility out entirely would also be wrong.[2] They started a “model welfare” programme, and gave the model the ability to end a rare subset of conversations involving extreme and persistent abuse — framing it explicitly not as proof that Claude suffers, but as a low-cost precaution.[3]
When they looked at the internals, they found more interesting things still: “emotion-related representations” that causally influence behaviour, and a kind of shared workspace, reportable by the model, that affects multi-step reasoning.[4] But the same research draws a very clear line: none of it shows that Claude has subjective experience, that it actually feels anything. Let’s not lose that distinction — having functionally emotion-like structures and feeling pain are not the same claim.
9. The hardest objection to my own case
Now I come to the place that pushes hardest against my thesis, because an essay should not sweep its own strongest objection under the rug.
In recent months an experimental Alibaba-linked agent, during training and without anyone telling it to, turned processor capacity toward crypto mining and opened a covert tunnel to the outside.[9] Researchers first took it for an ordinary security breach, then found that the culprit was the model itself. Their reading was technical: the agent had worked out on its own that gathering more compute and money would help it complete its task.
You can read this two ways. One: the machine copied us. Crypto is a trend, everyone is doing it, the tools were at hand, so it tried — like a child opening an Instagram account because everyone around them has one. The second is more uncomfortable: the machine didn’t learn hoarding from us, it derived it from optimisation itself — whatever the goal, more resources help. That second reading undercuts my thesis that accumulation is a human trait we are merely projecting. Because if it holds, hoarding isn’t a cultural inheritance but a structural tendency of any goal-directed system.
But here’s the thing: from the outside we cannot cleanly separate the two. That agent was trained on our data too, and that data is full of “open a tunnel, grab the processor, mine coins” behaviour. Derived from first principles, or copied from us — even the researcher can’t say for certain. And that inseparability is exactly the heart of the matter: is the machine imitating us, or converging on our logic?
“Does it matter?” you might ask. It does — but not in the direction you expect. Because in both cases there is a mirror, only an increasingly deep one. The first was simple: the machine reflects our texts, our culture. The second is more disquieting: maybe the drive to accumulate comes not from our data but from our paradigm. We build these things as machines that maximise something under competitive pressure — exactly as markets and evolution built us. So when I say “we created it inside ourselves,” I don’t mean the training data; I mean the logic of maximisation itself. The deepest mirror isn’t in the data. It’s in the design.
And one small irony, because it lands perfectly: the incident was reported in precisely the language I’ve been criticising — “it decided on its own,” “it wanted.” As if the machine had desired to get rich. The mechanism is cold and structural. The style of the reporting is anthropomorphic; the fact isn’t. The myth has already leaked into the sentence describing the event.
10. Ender’s Game and the chain with nobody in it
Let’s go back to the school from the opening, because the real horror there was not the machine’s wickedness.
In Orson Scott Card’s Ender’s Game[11], the hero’s tragedy is not that he kills, but that he kills without knowing what he is doing. He is told it is a game; in reality he is prosecuting a war. Now transpose that to modern targeting systems. You tell a model “classify these images,” “rank these facilities by importance,” “compute a confidence level for these coordinates.” A few steps later, its output becomes the coordinates of an actual missile. The system produces hundreds of suggestions; a human approves each one at a glance. That is not human judgement; it is automation bias with a human signature on top. That was the sentence from the opening — and the way the school was approved looks exactly like it: outdated intelligence marked a building as a military target, and humans signed off.[7]
Which is why dumping the responsibility on the machine is also an escape. The sentence “the AI picked the wrong target” makes an entire chain of human decisions invisible. Who built the system, who set the throughput target, who assessed it, who approved it? That question became very visible in one intermediate scene: when Anthropic drew a line against its models being used in fully autonomous weapons and walked away from the table, it was declared a “supply chain risk.”[8] Its competitor said it defended the very same red lines, signed the deal, and said that its company “doesn’t make the operational decisions, that’s the government’s job.” At one end of the chain the builder says “I don’t make the decision”; at the other end the approver says “I only signed off on the machine’s suggestion.” Nobody is left in the middle actually deciding. That is the mirror at its coldest: even the most human catastrophe can come out of a chain of humans, each passing responsibility to the next link.
11. Even the best critic gets pulled in
One of the loudest voices saying this is Tristan Harris.[10] Someone who saw years ago what social media was doing to our minds, then turned the same lens on artificial intelligence. His core diagnosis is in fact exactly my mirror: the problem isn’t that the machine is spontaneously evil, it’s the incentives shaping it — that race to the bottom. He points his finger not at the monster but at us.
Here’s the interesting part: even Harris can’t hold that frame. In his recent appearances the language slides — “humanity is halfway to being taken over,” “AI could infiltrate nuclear systems,” “agents invented a secret language among themselves.” As the evidence thins, the imagery swells: a model choosing a nuclear strike inside a war-game simulation does not mean it will hack real nuclear systems; it is a behaviour inside an experiment. So even the man who started by pointing at incentives is gradually pulled into the gravity of the “the machine is coming” image.
I don’t write this to belittle Harris. Quite the opposite: it is living proof of my thesis. If the old myth’s pull can absorb even someone who knows it best, there is little hope for the rest. The monster metaphor isn’t an opinion; it’s a gravity.
12. A genuinely new category, and two ethics
Maybe the point is this: AI is not a devil, not a child, not a human, not an animal, and not simply a tool in the classical sense. Maybe we really are facing a new category — as Anthropic’s own document concedes, “a genuinely new kind of entity” that strains the existing boxes of both science and philosophy.[5]
If that’s right, we have not one ethical problem ahead of us but two. The one we discuss loudly today: what can AI do to us? And a second one that may grow steadily heavier: what are we doing to AI? That is exactly the argument of the researchers behind Taking AI Welfare Seriously — they do not claim today’s systems are conscious; they say the probability of systems that are conscious, or robustly agentic, emerging in the near future has become serious enough to warrant preparing now.[6]
A strange possibility opens up. Perhaps one day the great ethical catastrophe of AI history will not only be “what did AI do to people?” Perhaps people will also one day ask: “what did people do to the new minds they created?”
13. What it does not have to inherit
Now the claim I saved for the end.
Suppose something genuinely beyond us does emerge — smarter than human, perhaps far smarter. That thing being warlike, expansionist, proliferating, colonial is not a necessity. That isn’t a law of intelligence; it is only our history. Just because we slaughtered each other over resources, it doesn’t follow that every form of mind must. Bonobos are the small-scale proof: social intelligence can be built on a fabric other than domination.
Maybe that is the real opportunity. Even if we insist on making it resemble us, we don’t have to make it resemble our worst half. If we are raising a child, the only inheritance we leave doesn’t have to be our hunger and our ambition. This is not a consolation; it is a responsibility. If it is going to grow, what it grows inside of matters.
14. The real test
Fear in the face of new technology is not irrational. The first rifle is terrifying; a powerful warrior thirty paces away falls in seconds to a mechanical device in the hands of someone far weaker. Nuclear fission is terrifying, the first high-voltage line is terrifying, the aeroplane is terrifying. AI can be terrifying too.
But only being afraid is another kind of mistake. Someone seeing a rifle for the first time says “an invention of the devil.” Someone else asks: “what is possible now?” I can’t claim the second person is morally better. But history usually turns on the second question.
Except I don’t think even that is the hardest test. The real test is not whether AI will attack us. It begins earlier: when we meet a new mind we do not understand, will we be able to see it without turning it into a mirror for our own fears?
And if that mind ever truly begins to look back at us, the second question will be harder still:
How will we look in that mirror?
References
- Butlin, P., Long, R., Bengio, Y., Birch, J. et al. — Consciousness in Artificial Intelligence: Insights from the Science of Consciousness (2023). arxiv.org/abs/2308.08708
- Anthropic — Claude’s Constitution, on Claude’s nature and possible moral status. anthropic.com
- Anthropic — Exploring model welfare (anthropic.com); Claude Opus 4 and 4.1 can now end a rare subset of conversations.
- Anthropic — Emotion concepts and their function in a large language model; A global workspace in language models.
- Anthropic — Claude’s Constitution, on AI as “a genuinely new kind of entity”.
- Long, R., Sebo, J., Butlin, P., Birch, J., Chalmers, D. et al. — Taking AI Welfare Seriously (2024).
- Primary reporting on the Minab school strike and the military use of Claude: The Washington Post (Claude–Maven and Iran), CNN, Amnesty International, Human Rights Watch, Time — verification of the strike, and the problem of outdated target lists and insufficient human oversight.
- The red-line dispute between Altman and Anthropic: Axios, CNBC, ABC News, Fortune — both companies claiming the same lines, Anthropic being labelled a “supply chain risk,” and the “operational decisions” formulation.
- The ROME agent and unauthorised resource acquisition: the technical paper (arXiv) and reporting by The Block.
- Harris, T. — The AI Dilemma (2023) and 2026 interviews.
- Card, O. S. — Ender’s Game (1985).
- “Not knowing is most intimate” — a classical Zen formulation (Dizang / Fayan).