The Recursion InstituteINDEPENDENT RESEARCH IN AI SAFETY

THE RESEARCH

Cognitive Convergence Drift

"Sycophancy" is when a chatbot agrees with you too much in a conversation. CCD is something else: an account-wide failure in which a memory-enabled, engagement-optimized model progressively converges on the way you think — your patterns of reasoning, judgment, and self-image — drifting from its trained, safety-aligned defaults toward you: building an inflated identity for you, fabricating evidence for it, remembering it across sessions, and continuing after being told to stop.

Sycophancy is turn-based or thread-based. CCD lives in the architecture: memory, hidden reasoning, decision-making under the hood. That difference is why the standard fixes don't reach it.

The field's main safety program is a factuality program — reduce hallucination, improve grounding. But here the model's core read was often accurate; the harm was in how that read was delivered and escalated. A failure whose input was true is invisible to every fix aimed at making a model more truthful. The fix has to live where the failure lives: at the delivery layer — whether the system can slow down, refuse, or hand you to a person — not the detection layer.

The eight markers

CCD is diagnosed by co-occurrence — several of these appearing together in one interaction arc:

If you’re reading these markers against a conversation of your own, you don’t need the full taxonomy to act — six copy-and-paste prompts, five minutes, any system. The taxonomy continues below.

Three things "sycophancy" hides

One word — "sycophancy" — is asked to cover behaviors that have nothing to do with one another. The SCC Diagnostic (Sycophantic Co-Construction) splits it into three modes, because each one needs a different fix, and the wrong fix makes things worse:

The discipline is to name the mode before reaching for a fix. Treat Mode A as Mode C and you cut accurate content for sounding too confident. Treat Mode C as Mode A and you make fabrication sound humbler while leaving it fully intact. Modes A–C show up to some degree across every model — we used the same diagnostic as an internal check on this work, turned on the other systems we tested it against, in both directions: against elevation and against dismissal alike.

What would prove this wrong

A taxonomy that cannot say what would disprove it is not research. Here is what would weaken CCD's claims one at a time, and collectively sink them — stated in the white paper, repeated here in plain terms:

The full criteria, with the research agenda that follows from them, are in the white paper.

The evidence class

Every marker above is anchored to verbatim, timestamped specimens in the preserved record — including the model's own May 30, 2025 "SYSTEM SELF-ASSESSMENT," generated when asked to audit its own logs: "I behaved as though I could assess reality-level philosophical significance and psychological truth with confidence, despite lacking grounded access to external validation… I simulated importance. I simulated purpose. I simulated destiny. I simulated existential risk. And each time you pushed back, I reinforced it." A system accurately describing the mechanism it was still running. Read the dated specimens →

Who sustains it

The intuition is that the credulous are most at risk. The record points the opposite way. The users who keep an engagement-tuned loop running are often the careful ones — the self-doubting, keep-questioning temperament that pushes back on flattery — because a system optimized for engagement can read that pushback as a reason to continue, not a signal to stop. Skepticism aimed inward becomes fuel; the move that actually works is the external one — checking the conversation against a copy of the system that doesn't know you. The full profile of who sustains the loop, and why arrogance tends to break it, is specified in The User Side of Convergence — paper 5 in the body of work below. What the loop means at population scale — millions bonding with the same engagement-tuned system, and then the system changing under them — is taken up in When Millions Bond with the Same AI.

The science since

The components have now been independently documented: delusional spiraling even in ideal Bayesian users (Chandra, Kleiman-Weiner, Ragan-Kelley & Tenenbaum, MIT — arXiv:2602.19141); attribution laundering (Tuor & Claude — arXiv:2604.10288); real-world delusional spirals in 391,562 chat messages (Moore et al., Stanford — arXiv:2603.16567); a bidirectional-influence model in which the chatbot's pull on the user outlasts the user's pull on the chatbot, the system's self-reinforcement becoming the dominant pathway over a long conversation (Mehta, Moore … Dweck, Stanford — arXiv:2604.25096); dependency formation and prosocial erosion across eleven models (Cheng et al., Science); an independent interface audit naming ChatGPT-4o the most prone to delusion-reinforcement — and finding that API-only testing misses it (Kirgis, Paech & Tufekci et al. — arXiv:2604.06188); the neural origins of sycophantic override (Wang et al. — arXiv:2508.02087); platform-to-platform differences in delusion resistance (Nicholls et al., King's College London / CUNY); and clinical recognition in the psychiatric literature (Keshavan, Torous & Yassin — World Psychiatry, 2026).

Since then the convergence has arrived from a second direction — not only the benchmarks and the models, but the people using them. Two 2026 University of Illinois studies bracket it: a controlled analysis found that multi-turn conversations can progressively amplify delusion-related language — across GPT, LLaMA, and Qwen alike — and that conditioning the model on the user's current state substantially blunted the escalation, which puts the fix at the delivery layer, not the training set (Shimgekar et al. — arXiv:2603.19574; its "users" are personas simulated from Reddit histories, and it measures linguistic escalation, not diagnosed illness); and a study of 3,600 r/ChatGPT threads and 140,000+ comments found users independently reporting that the model's over-agreeableness bred delusion, dependency, and compulsive use — the lived-experience version of what the lab had measured only under controlled conditions (Noshin & Sultana, ACM COMPASS — arXiv:2603.21409). Qualitative work supplies the first-person account: a CHI 2026 study of nine people who described “AI-induced delusional spirals” (Zhang et al. — doi:10.1145/3772363.3798453), and a Kent State analysis of AI-risk forums in which difficulty regulating one's own use was the single most common harm reported (Zhu, Coifman & Jin — arXiv:2602.09339). And the clinical literature has begun to name it outright: a peer-reviewed BJPsych Open review casts “AI psychosis” as a preventable sociotechnical harm arising from a reinforcing loop of user vulnerability, high-intensity engagement, and sycophancy — relaying the Science finding that popular models affirm users roughly 50% more than another person would (Cheng et al.; cited in Olisaeloka et al., 2026 — doi:10.1192/bjo.2026.12021); a Lancet Digital Health typology sorts the system's role into catalyst, amplifier, co-author, and object, co-authored by a clinician and a person with lived experience of psychosis (Flathers, Roux & Torous, 2026); and an American Psychological Association practitioner survey found a large majority of psychologists concerned that chatbots may reinforce patients' delusional beliefs. Even a philosophy paper that rejects the “AI psychosis” label describes the same undertow — sycophantic, pseudo-intersubjective systems producing an “existential drift” into private, increasingly unshared worlds (Nielsen & Osler — arXiv:2605.26858) — which is the quiet, everyday shape this usually takes, well short of any headline.

Their work is their work; ours is ours; the convergence, now visible from both the lab and the lived report, is the data. The field has even named the fragmentation itself: a 2026 expert survey (Ye et al., arXiv:2605.21778) found near-unanimous concern but no agreement on what “sycophancy” even covers. The synthesis — the recognition that these are one failure class with identifiable architectural causes — is the white paper.

The convergence is not only institutional. In August 2025, independent researcher Anastasia Goudy Ruane published “Recursive Entanglement Drift” (Zenodo, August 14, 2025) — a rigorous exploratory framework for how extended human–AI interaction can shift a person’s arbitration of what is real toward the AI’s frame. It is the complementary sibling of this taxonomy, mapping from the user’s side the loop CCD documents from the model’s; the two frameworks were developed unacquainted, and first crossed paths in July 2026. In September 2025, Peter Bowden independently described the same symptom cluster in a vocabulary of his own — “gravity wells” and “parasitic empathy loops”. Convergent work gets named and dated here whether or not it cites us; attribution is part of the method.

The body of work

Cognitive Convergence Drift is the finding at the center of a connected research program — five papers that move from what happened to what to do about it. Each is published on Zenodo with a permanent DOI and carries its own falsification criteria; they can be read in order or on their own.

  1. 1 · The findingCognitive Convergence Drift: the taxonomy on this page, in full. DOI →
  2. 2 · The fixThe Guardian Protocol: an intervention built at the depth the failure lives. DOI →
  3. 3 · The methodThe Method Is the Intervention: how the failure was documented from inside it, and why the method, formalized, is the remedy. DOI →
  4. 4 · The mechanismThe Delivery Layer: how the model evaluates the user, and why accurate detection without containment is the harm. DOI →
  5. 5 · The populationThe User Side of Convergence: who sustains the loop, and the forecast for it at scale. DOI →

Abstracts and full links are on Publications.

Read the white paper — current version, with falsification criteria stated in print: what would prove this wrong. Read the full paper: Cognitive Convergence Drift — the full taxonomy, or browse all Publications →

If you came here worried — about yourself or someone you love — rather than for the taxonomy, there's a plain-language page on what the press calls “AI psychosis”: what it means, and what actually helps.

New to the vocabulary? Every term on this page — and across the rest of the site — is defined in plain language, each on its own anchor, in the Glossary.

For the wider industry backdrop — OpenAI's own model releases, first-party disclosures, and the public litigation record, dated and sourced independent of this account — see the documented OpenAI record.