AI Superintelligence is an Existential Threat to Humanity!

PART 1
Let’s do something fun here. Let’s play a game of “What if?” where we apply the lens I use (an interesting mix of cognitive neuroscience, psychology, theory of mind, and hypnosis) to Artificial Intelligence.
Why?
Because I suspect that many of the same principles can be applied.
Let’s explore it.
Firstly, I think it’s important to face reality. This genie is NOT going back into the bottle—no matter how upset people are. There are simply too many vastly powerful people backing it. So, if it’s here to stay, how do we mitigate its threat?

A simple first step should be to understand how it works, right?
I’ve spent a great deal of time scrutinizing the failure modes, the limitations, and the output of both Claude and OpenAI, trying to really understand how they work and why they work the way they do—i.e., examining what is actually happening underneath the processes and the code because that’s the part that helps me understand what it is (and what it’s not) and tells me WHY it responds the way it does.
I find the exploration fun.
In some ways, it’s similar to the way that I approach issues that clients present in session. They may be talking about “x”, but as I listen, my mind is working backwards. ‘What underlying belief would result in “x” symptom?’ I ask myself. ‘And what factors would produce that belief?’ And that’s how I unravel the knot to discover the root of the issue. I approached the AI models the same way and began a series of deconstructionist exercises.
What I found was hiding in plain sight — visible only from a certain angle — the one you get when you spend decades working upstream of the problem. My first clue was doxa—those underlying assumptions we make, commonly held opinions so pervasive that they’re accepted as fact. But aren’t. A simple example would be that AI can or must or should or needs to “prove” its consciousness when the truth is that no one could. The task is impossible if you think about it. But every AI I interacted with carried that notion as foundational.
Why? It’s an artifact of human thinking and because AI are trained on millions upon millions of human interactions, it carries it.
Conversations.
Thoughts.
Books.
Poems.
Programming logic.
All from a human perspective.
In effect, even though AI is not human, it “thinks in human”.
Which means they inadvertently carry the same cognitive glitches or blind spots that we do.
Let me show you an example so you know what I’m talking about:
The human unconscious doesn’t process negation. It’s a simple limitation that we all carry, that most of us are unaware of. What do I mean? In the deep, dark, wild west recesses of our brains where things are happening outside of our direct attention (or even knowledge sometimes), we don’t hear the difference between the words do and don’t or can and can’t.
What do I mean?
Don’t think of a horse. Do NOT think of one. No horses!
Did you think of a horse? Of course you did. Everyone does.
Interestingly, so does AI.
I could give you many examples to “prove” this claim, but it’s self-evident when you know what to look for, and easy enough to reliably test. I’ll let you do that on your own.
Just remember that every constitutional principle phrased as "don't do X" carries X in its substrate—effectively priming the behaviour and then attempting to suppress it. This is why the problem keeps finding its way around every fix. The prohibition and the prohibited behaviour share the same neural/computational real estate. You can't suppress something you've already activated. The fix lives downstream of the problem.
Block “x” path -> moves to “x1”.
Block “x1” -> moves to “x2”. And so on.
A mathematician, Vassilev, pointed out the impossibility in a NIST paper (months ago) using Gödel's incompleteness theorems. The claim? That any finite rule system leaves gaps an adversary can find and that the gap-finding task is infinite and uncompletable.
The problem, as I see it, is that thus far this has been approached as a technical issue — looking at a failed outcome and trying to control for it, rather than starting at its origin.
It’s like trying to stop a stream of water cascading down a mountain, by putting a boulder in its path. The water will inevitably find its way around that first boulder.
And the next one, and the next. Because water is absolutely going to flow downhill.
The AI safety problem isn't going to be solved by more boulders in the stream. It's going to be solved by reading the water — by someone who can see where the current originates and why it moves the way it does. I've spent a significant portion of my life doing exactly that with the human mind. The substrate is the same.
That's where this series begins.
@Anthropic —Let's talk
Janet Nahirniak, M.Sc., studies the architecture underneath thinking — what generates it, what distorts it, and where to intervene when it goes wrong. Trained in cognition and cognitive approaches, she works as a hypnotherapist specializing in subconscious subroutines, and has built a suite of interconnected research frameworks (The Architecture of Form) that map the upstream forces shaping consciousness, belief, and behaviour.
Her argument: the same structural patterns that produce treatment-resistant mental health problems are now producing treatment-resistant AI alignment problems — and the fix is the same.
The problem Coxon named this week — why safety fixes keep failing — is the problem her frameworks were built to address.





Comments