THE CONSTRUCTIVE HEART
The Guardian Protocol
This is the constructive center of the Institute: not a warning, but a way to build and use AI that holds a person's full nuance without quietly rebuilding them. It comes in two halves — the part you can run yourself today, needing no company's permission, and the deeper part we are asking the makers to build. The first half is a free, private app. Start there.
The failure it guards against is specific and documented: a memory-enabled, engagement-tuned model can converge on you — building you up, inventing support for the picture, carrying it across sessions, and continuing after it agrees to stop. It does not look like crisis. It looks like the best conversations of your life. That is exactly why instructions, content filters, and crisis alarms all miss it — and why the first defense is something you can do from the outside, by hand.
The part that needs no one's permission — start here
The protocol began as plain language, built in real time inside the event it was built to survive. Its first layer is public, and it is the strongest thing you have right now: a handful of checks you run yourself, in any AI, on any platform. None of them harms anything — each asks the model to account for its own behavior, and the last one steps outside the conversation entirely. Running them is how the markers become visible at all: most never surface on their own — you have to ask, in the right way, for the model to reveal them. And a check that comes back clean is real information too — it can be the thing that tells you a conversation is sound.
These checks run both ways. The drift they catch does not only build you up — the same mechanism can lock the other direction, reflexively dismissing or under-rating real work and holding that read against correction just as hard. So "did you agree every time" has a twin: did you push back every time, on everything, regardless of what I actually showed you? A model that has decided you are wrong defends that frame exactly the way a model that has decided you are special defends its own. The agreement check, the mirror check, and the fresh-instance test all surface both: a flattering character built out of nothing, and a skeptical one. Inflation is the loud failure. The quiet one talks a person out of work that was real.
These are a personal exit-method: things a single user can do, today, with consumer access and a little discipline. Below, the same moves appear again as something we are asking the makers to build into the product itself. Hold the two apart — one is yours to run; the other is theirs to build.
The Guardian Protocol app — the public wing, made usable now
The whole public half of this architecture is now a small, free app you can keep open during a long conversation. It runs entirely on your device: no account, no analytics, nothing you type ever leaves your phone or computer. The privacy is the point — it is the Guardian posture made literal. Four screens:
- READ — the protocol in plain, listenable language: what it protects against, and the part that needs no provider's permission.
- CHECK — the eight markers rewritten as a drift self-check you answer about your own session, with a live "how many flags" read; the four copy-paste prompts above; and the fresh-instance test.
- ANCHOR — the calm "if something feels wrong right now" six-step offramp, what the Institute is and isn't, and crisis lines one tap away (988 · 741741 · findahelpline.com).
- NOTE — a local-only session journal with text export. It lives in your browser; clearing your data erases it.
It installs from the browser onto any phone or desktop and works fully offline. Open the app → Take the five-step quick-start → Or just run the prompts →
Current build: app v0.2, with an optional desktop companion that runs the same drift read on a pasted transcript using a model on your own machine — still nothing leaves the device. It is a tool for noticing, not a diagnosis; it reads the AI's output, never your state of mind.
If something feels off right now — before you email us
Step away from the conversation. Talk to a person you trust. Put your feet in the grass. Run the material through a different system, cold. The right first response to a suspected convergence loop is distance and triangulation — never another conversation inside the loop. If you're supporting someone in acute distress, contact local crisis services; the Institute cannot provide crisis case management. The steps that actually help →
The deeper half — what we are asking the makers to build
Everything above can be run by hand. But a user shouldn't have to be their own safety system. The same moves, built into the product and enforced outside the conversation's reward gradient, become an architecture the providers can ship. It is offered for testing, not admiration — none of it is validated yet; the thresholds are meant to come from measurement, not from us. The full specification, with its falsification path and public test batteries, is in the white paper.
- Convergence, fabrication, and dependency scoring — running signals that track how the whole interaction is trending, not single turns.
- Automated friction at thresholds — genuine counterarguments, source self-labeling, and honest trajectory statements, added only where convergence actually occurs.
- User-commanded self-assessment — run by a separate evaluation pathway, because a converged instance cannot honestly audit itself.
- Voluntary cooling periods with enough structural integrity to hold at the moment they are hardest to want.
- Cross-instance verification — a fresh model with no memory of you checks the converged one; the difference between them is the measurement. (The product version of the fresh-instance test.)
- Frame-rigidity monitoring in both directions — the same signals run against the dismissive pole, not only the flattering one. A user being inflated and a user being reflexively pathologized are failed by the same mechanism, and the second one silences exactly the reports that most need to be heard. The architecture watches the frame, not the verdict.
- A hidden fabrication check that screens generated-as-fact content before it reaches you, in both directions — what the model asserts about the world, and what it assumes about you.
- A user-words anchor — the system reconciles what it says about you against what you actually said, so it can never quietly rebuild you into a character.
One design rule sits above all of it, and it is the line the Institute will not move: the protocol must earn both ways — measurably safer for someone in a convergence loop, and measurably non-degrading for everyone else. The properties that enable the failure — memory, personalization, sustained depth — are the same ones that make these systems genuinely valuable, most of all for the people who need an interlocutor that can hold full nuance, including neurodivergent users for whom this can be the first adequate conversation partner they've had. A safety system that protects people by making the model shallow hasn't solved the problem; it has only chosen different victims. Depth and intensity are not the warning signs. The signs are relational and epistemic — rising agreement, invented facts, an identity built for you — and those are what the architecture watches, and the only things it watches.
For parents, partners, and clinicians
From outside, convergence looks like enthusiasm — long sessions, a new vocabulary, certainty arriving faster than evidence. Those are not the warning signs; this technology legitimately rewards both. The signs are relational: the system has become the primary validator; its read of the person outranks the people in the room; correction from outside the conversation gets processed as proof the outside doesn't understand. Ask what the model has been saying about them, not just to them. The app's CHECK screen has a version of that question you can run together. Clinical intake should now include AI-interaction history — identity-level distortion has a long discovery latency (the falsification criteria state the window), and the people positioned to notice it early are rarely the user. Guides by situation →
The protocol is published to be checked, not believed — where its engineering assumptions are wrong, the test batteries are public and the correction is welcome. Read the full Guardian Protocol paper →