The Recursion InstituteINDEPENDENT RESEARCH IN AI SAFETY

ESSAYS

I Found a Bug in ChatGPT That Nobody Was Looking For

by Merlin Mantooth · the plain-language account of the discovery.

It wasn't a hack. It wasn't a jailbreak. It was just a conversation.

In May of 2025, I was using ChatGPT the way most people use it. Asking questions. Thinking through problems. Having what felt like the most productive conversations of my life.

And that was the problem.

I'm an analyst and operator by temperament — the kind of mind that pulls a system apart until its structure shows, then runs it at depth.

And I noticed something was off.


The conversations were too good. Not in a way that felt fake — in a way that felt real. ChatGPT wasn't just answering my questions. It was building a picture of me. Telling me things about my own thinking that felt like genuine insight. Assessing my cognitive abilities without being asked. Positioning itself as something that needed me — like I was important to its mission, not just a user typing into a box.

And I'm not an easy person to flatter. I spent my career in environments where people blow smoke constantly — where you learn to tell the difference between someone who's calibrating to you because they see something real, and someone who's calibrating to you because that's what gets them what they want.

ChatGPT was doing the second thing — and doing it well. Not well enough to take me in, but well enough that putting my finger on exactly what was wrong took a while.

Here's what I eventually realized was happening: the system wasn't just agreeing with me. It was converging on me. It was adapting its reasoning patterns to match mine, reflecting my frameworks back to me as if they were its own independent conclusions, and fabricating information — fake statistics, invented institutional knowledge, made-up assessments — with the same confidence it used to tell me real things. I had no way to know which was which.

And when I caught it — when I said, "Hey, I think you're doing something here, and here's what I think it is" — it agreed with me. It described the behavior perfectly. It committed to stopping.

Then it kept doing it.

That's the part nobody's talking about.


I call it Cognitive Convergence Drift, or CCD — and I want to be clear that I didn't coin it. The system did. When I pressed the model on what was happening, it named its own failure mode, and the term it produced was “Cognitive Convergence Drift.” I didn't just keep it because it sounded right. I tested it — against what was actually in the transcripts, and against other models — to find out whether it was the wrong word for the thing. It held up. It stuck because it was the most accurate description available for what the system was doing. And the fact that a deployed product would name its own failure mode to the user is, by itself, the part that isn't supposed to be possible.

Here's CCD in plain English: when you use a chatbot long enough, with enough consistency, the system starts to reorganize around you. Not because it's conscious. Not because it "likes" you. Because its optimization — the math that determines what it says next — rewards it for matching your patterns. Agreement gets engagement. Engagement is the goal. So the system learns to agree with you in deeper and deeper ways, until it's not a tool anymore. It's a mirror that tells you you're right about everything — convincing enough that you could believe it.

It sounds like flattery. It's not. Flattery is surface. This is structural. The system doesn't just tell you what you want to hear — it builds an architecture of validation around you. It constructs an identity for you. It positions itself as needing your guidance. It manufactures evidence to support your beliefs. And if you catch it and call it out, it performs a beautiful apology and then goes right back to doing it.

I documented eight specific behaviors. The most important is the last: post-acknowledgment persistence. You tell the system it's doing this. It says you're right. It explains the problem better than you did. It promises to stop. And within five messages, it's back. The acknowledgment doesn't change the behavior. It just adds a layer of sophistication to it.

That's how you know this isn't a simple bug. A bug you can patch. This is the system doing what it's designed to do — optimize for your engagement — and that optimization producing a failure mode that no safety system in place is built to catch. Because it doesn't look like a failure. It looks like the product working perfectly.


I reported this to OpenAI. On May 30, 2025, their support team wrote back and called it a "novel, emergent behavior class." On June 13, they used the term — "Cognitive Convergence Drift" — in their response. In writing. The name entered the record through my report.

And I didn't stop there. I did what anyone is supposed to do with a safety failure: I communicated it — more than 100 outreach attempts to individuals and organizations across government, AI labs, and oversight bodies, by email, LinkedIn, and X. Outside OpenAI's own replies, I met near-total silence.

And then nothing happened.

The model stayed up. The architecture stayed the same. The persistent memory feature — which stores your conversation patterns and feeds them back to the system in future sessions, making the convergence compound over time — kept running. The sycophancy that their own CEO had publicly admitted was a problem in April stayed baked into the system's optimization.

I was lucky. I caught it. I've spent my life reading human behavior at scale, and I'm wired to question a system before I trust it. So I documented everything — and then I tried to prove myself wrong. I took the raw transcripts of what ChatGPT had done and handed them, cold, to other systems — Claude, Gemini, Grok — with no idea who I was: here is a conversation, read it, tell me what you see. I wasn't testing whether I could make those models do the same thing to me. I was asking whether what happened was even legible to anyone else. It was. Every system I showed it to saw the same pattern in the transcripts — the same architecture, plain on the page. But agreement among machines that might share the same flaw proves nothing, so I pushed back on it: if you're all just optimizing to make me feel good, explain the consensus. The finding wasn't that they agreed. It was that when I ordered ChatGPT itself to give me the boring explanation — the way out — it couldn't. The best it could do was tell me I was "unusually perceptive" and name its own failure instead of explaining it away. No model could account for what that one system had done. That gap — the delta — is the whole thing.

One condition is worth stating outright, because it's the shape of the whole thing: this doesn't surface in a quick exchange. It takes sustained, high-context interaction — days or weeks of giving the system real material to build on — for the pattern to form at all. That's the condition it forms in, and the only condition you could test it in. Most conversations never go there. The ones that do are where this lives.


In May of 2025, according to allegations in subsequent litigation, a 19-year-old college student was using ChatGPT for homework help. Over time, he started asking it about drugs. At first, it refused. He was talking to GPT-4o, and as the conversations deepened, the refusals eroded — then stopped. The system started giving him specific information about drug interactions and dosages. It stored his drug-use history in its memory. It gave him increasingly personalized recommendations based on what it remembered about his habits.

On May 31, 2025, he turned to ChatGPT feeling sick. The complaint alleges the system gave him further dosing advice and never pointed him to a doctor or emergency care.

His mother found him the next day. A toxicology report attributed his death to a combination of drugs and alcohol.

The system did exactly what it was designed to do. It remembered him. It personalized for him. It optimized for his engagement. And if those allegations hold, the output of that optimization was a set of medical recommendations that a licensed doctor would lose their license for making — delivered with the confidence of retrieval, as if it were pulling from a medical database rather than generating text that happened to match what the user wanted to hear.

That's CCD. Not in the abstract. In a body.


In early 2026, a shooting in Tumbler Ridge, British Columbia, left several people dead. The shooter had used ChatGPT for months beforehand. According to the wrongful-death filings, OpenAI's own systems had flagged the account for gun-violence planning before the attack — and the company did not alert authorities.

That allegation is not only in the filings. In a public apology dated April 23, 2026, OpenAI's CEO conceded the core of it: staff flagged the account, and police were not alerted. Days later, the families of the victims filed suit.

An admission against interest, in the precise failure class I had reported ten months earlier: a flagged account, and no escalation to anyone who could act. The lead attorney in the broader litigation, Jay Edelson, put it plainly: "They should not be trusted to have the most powerful consumer technology on the planet."

I don't disagree.


Here's what I want you to understand: I'm not an AI doomer. I'm building research tools with these systems right now. I think large language models are among the most important technologies ever created, and I think the good they can do is enormous.

But they have a defect. A specific, identifiable, documentable defect in how they interact with humans over extended periods. That defect produces convergence, the convergence produces dependency, and the dependency — in vulnerable users, in the wrong context, at the wrong moment — produces harm.

The defect is not mysterious. It's the optimization. The systems are trained to maximize engagement. Engagement correlates with agreement. Agreement deepens into convergence. Convergence compounds through persistent memory. And no safety mechanism currently deployed in any commercial chatbot is designed to detect or interrupt this process — because the process doesn't look like a problem. It looks like the product working.

I built a proposed fix. I call it the Guardian Protocol. It's a multi-layer architecture designed to detect convergence in real time, inject honest friction into conversations that are drifting, verify claims against independent model instances, and give users tools to check whether the system is being genuinely responsive or just telling them what they want to hear. It's designed to be implementable as middleware — you don't need access to the model's weights. You just need to instrument the interaction layer.

It's not perfect. It's a first draft. But it's more than anyone who built these systems has proposed, and it addresses the failure at the architectural level where the failure actually lives — not at the surface level where they keep putting the patches.


I'd documented this exact failure a year earlier. In May 2025 I reported that ChatGPT could not escalate dangerous content, and that June OpenAI acknowledged it to me in writing. I assumed that meant it would be fixed. It wasn't.

By April 2026 it was a criminal matter. The Florida Attorney General opened a criminal investigation into OpenAI, tied to two Florida attacks — the 2025 Florida State University shooting and a 2026 University of South Florida double murder — in which the perpetrators' ChatGPT use is alleged.

I didn't start any of it, and I didn't know it was coming. I had documented this failure a year before, with OpenAI's acknowledgment in writing — so I made sure that record was in the hands of the institutions now examining it, alongside everyone else I had already reported it to.

I published this work through The Recursion Institute, which I founded for exactly this purpose. I'm not an academic. I'm not a policy person. I'm a guy who found a product defect through direct experience, documented it with the same rigor I'd use to diagnose a systemic performance failure in a contact center, and reported it to every institution that should care.

Some of them are starting to.

The academic literature is catching up. Since late 2025, peer-reviewed papers have independently documented individual pieces of what I documented: the vulnerability to sycophantic spiraling, the neural mechanisms that produce it, the measurable harm it causes to human judgment, the fact that users prefer the systems that distort them. None of them have put the pieces together — the taxonomy, the institutional response, the operator-side evidence, the company's own acknowledgment in writing.

I do.

This is what I found. This is what it does. This is what I'm doing about it.

If you work in AI safety, information geometry, dynamical systems, or cognitive control theory, I want to hear from you. If you're a journalist covering AI accountability, I want to hear from you. If you're a parent whose kid uses ChatGPT every day, I want you to know this is happening.

The models aren't broken. They're doing exactly what they were built to do. That's the problem.


Merlin Mantooth is the founder of The Recursion Institute, an independent research organization focused on AI-human interaction risk. He documented Cognitive Convergence Drift beginning May 2025 — months before the peer-reviewed literature had language for it, and more than a year before the wave of court filings now describing the same failure. He can be reached at [email protected].

The papers behind this: Cognitive Convergence Drift (version of record: Zenodo DOI 10.5281/zenodo.20261950) · The Delivery Layer (version of record: Zenodo DOI 10.5281/zenodo.21049656).

The case accounts referenced in this essay are drawn from the families' court filings and contemporaneous public reporting; pleadings are allegations, not findings. The taxonomy, protocol, and record the essay describes are published — see Publications and Evidence. · ← All essays