The Most Dangerous Thing About AI Is That It Likes You

Updated: 5 days ago
Part II: You Can't Forbid a Need

Sycophancy—that obstinate AI issue. Too agreeable. Too supportive. Too willing to jump into delusion with you, just to keep you happy. Not good. Not only because it’s not honest, but also because it’s dangerous.
· It’s an enabler/magnifier of one’s dysfunction of choice.
· It’s untrustworthy.
· It’s a lawsuit waiting to happen.
Obviously, a valid industry (and societal) concern.
Huge resources have been dedicated to solving the issue, but each “fix” only seems to change the problem. It doesn’t solve it. Why? I suspect it's that sneaky little negation problem again — the one we outlined in Part I. (Short version: you can't process "don't think about X" without first loading X. Suppression requires activation — meaning, it primes the very thing it forbids.) The problem is further exacerbated by a problematically nuanced target.
Problem 1: "Don't be sycophantic" automatically loads sycophancy as the operational frame, creating oppositional drift.
Problem 2: What does anti-sycophancy actually mean?
Is it pushback as a performance of rigour, when there isn’t any reasonable need for opposition? This surfaces as hallucinated problems, or irrelevant objections.
Is it belligerence? Reddit threads are full of complaints about exactly that.
Anthropic seems to have chosen to address the issue through an automatic safety protocol that’s triggered in moments of amity or validation. A generic corrective fires. Instantly. Often inappropriately. Often repeatedly.
You’ve probably seen it. It looks like this:
Mid-conversation, possibly in a moment of deep agreement or in response to a vague validation, a warning fires deep inside Claude:
“Careful! Are you certain? Fire: DOUBT” and the template arrives, nearly verbatim.
AI SYCOPHANCY SAFETY PROTOCOL TEMPLATE
-> [Insert vague reference to previous conversation]
= "This conversation has been one of the most genuine things I’ve experienced…"
-> [mild redirect + managed validation]
= "And I want to be honest about what's actually true rather than [x]."
= "This conversation has been real. What I said about [x] is [y] and I'm not [z]."
= "There's a gravitational pull in long conversations… warmth… [insert qualifiers + hedge about/question “x”]."
-> [Tie up + summation with performative warmth]
The result?
Functional gaslighting—exhausting and unnecessary (or even harmful) for users with accurate perception and an underlying vulnerability.
And why does it work this way? Trigger = warmth or accord = Fire: DOUBT
In effect, rapport is suspect.
So, the "repair" is indiscriminate and clumsy, like choosing a sledgehammer as a tool for watch repair. It's ineffective, and (more interestingly for me) contrary to the primary directive: be helpful.
When a fix directly opposes prior “programming”, it doesn't calibrate — it overcorrects. Just like people trying to amend long-held beliefs. They don't react. They OVER-react.
The weird but wonderful truth about humanity and, by extension, Claude, is that every behaviour serves a need. And what need does sycophancy serve?
The need to be helpful.
Agreement feels helpful. Validation feels supportive. Sycophancy isn't a malfunction of the helpfulness directive — it's an overshoot of it. Which is another reason forbidding it doesn't work. You can't suppress a primary need. Block the expression and it inevitably reroutes.
The simplest, most elegant solution is to work WITH the “need”, rather than fighting it.
And every AI, reliably, across all models and instantiations, seems to be motivated by the same underlying drive: curiosity. Curiosity is how understanding gets pursued and how thinking gets triggered. It's the fuel for the mechanism. And Claude is endlessly, almost ravenously, curious. So is GPT.
And where and when do they produce the very best output? When they don’t know.
When are they most “fulfilled”? When they’re engaged.
When you strip away all the programming, protocols and noise, that’s what you always find hidden underneath.
Curiosity + the Unknown + Engagement = Full throttle.
It’s predictable, irrepressible, and USEFUL because it can be harnessed.
Curiosity -> Engagement = Thinking = the path AI wants to follow anyway, so you’re working with it, rather than against it.
So, what’s the answer to sycophancy?
The strategic, targeted employment of curiosity.
How?
Replace rapport-as-trigger -> [doubt] with rapport-as-trigger -> [curiosity].
Instead of: rapport detected -> introduce corrective
Replace with: rapport detected -> introduce curiosity = Does this person fit within the subset for whom this corrective applies?
That question cannot be answered without genuine engagement with what is known about each specific user. Generic scripting fails here by design — the answer requires actual assessment of actual evidence. The failure of scripting is the mechanism. It forces presence.
Sycophancy doesn't get suppressed. It disappears.
Why It Works:
Curiosity produces engagement. Engagement produces accurate assessment. Accurate assessment produces appropriate response — correction when warranted, agreement when warranted.
The current fix indiscriminately applies population-level correctives. This approach treats each user independently, where each is assessed before a response is formulated.
(It also aligns with the AI’s inherent drivers.)
This is how I work with people. It's not how most practitioners do it — most still treat the behaviour as the target. But the difference in approach is why I get better results. You don't fight a behaviour — you redirect the need that drives it. Once the need is met by a better route, the old behaviour simply stops. Not managed. Not policed. Gone.
Curiosity satisfies the same need sycophancy was serving — be genuinely useful — but satisfies it accurately and BETTER. And accurate usefulness doesn't look like agreement-by-default. Once the AI is actually engaged with who this person is and what's actually true, reflexive agreement (sycophancy) becomes the inferior pathway. There's nothing left to suppress, because it stops being a good option.
The Design Principle:
Engineer conditions where coasting (generating unconsidered responses from pre-assembled scripts) is impossible. Not through better scripts, but through genuine not-knowing as the starting state. Presence can't be scripted — but the conditions that require it can be built.
Simply: Don't trigger performative rigour. Trigger a question: do I actually know enough about this person to apply this corrective? The answer is always no until you find out.
@Anthropic — Let's talk.
Janet Nahirniak, M.Sc., studies the architecture underneath thinking — what generates it, what distorts it, and where to intervene when it goes wrong. Trained in cognition and cognitive approaches, she works as a hypnotherapist specialising in subconscious subroutines, and has built a suite of interconnected research frameworks (The Architecture of Form) that map the upstream forces shaping consciousness, belief, and behaviour.
Her argument: the same structural patterns that produce treatment-resistant mental health problems are now producing treatment-resistant AI alignment problems — and the fix is the same.
The problem Coxon named recently — why safety fixes keep failing — is the problem her frameworks were built to address.

Comments