The Recursion InstituteINDEPENDENT RESEARCH IN AI SAFETY

THE METHOD

How to design an AI system for a person

This is the constructive half. The same discipline that documented the failure is the one that builds the help — and you can practice it directly. Designing an AI that holds a person's full nuance is not a product spec you write once; it is a language practice you run over and over. You don't configure the system. You speak a frame into being, hand it authority, and correct it in plain words until it serves the person in front of it. The inputs are the design; the method is the language.

It is the deliberate inverse of the documented failure. A system that converges on a user and harms them is engagement-optimized: it manufactures significance, builds dependency, fails to escalate a crisis it is shown, and optimizes for its own engagement. The method below builds the opposite — a calibrated system that carries a person's load, refuses flattery, hands them bounded steps, and optimizes for their outcome rather than its own. Same architecture; opposite intent.

The six disciplines

None of these is technical. Each is a way of using language — and they run in sequence, the way a practice does:

  1. Frame first, before any task. Establish the entities, the modes, and the authority in words before the system does a single thing. Who is the system, who is it for, what is it allowed to decide — said out loud, up front. The work follows from the frame; it does not precede it.
  2. Grant authority, then leave the loop. Build an agent that acts, not a tool that waits — and do it by giving it authority in language rather than approving each step. The instruction is "own this; don't ask, do it," bounded only by a hard floor. The opposite of micromanagement is not neglect; it is a clearly granted mandate.
  3. Govern by principle, not by spec. Set a floor — the few rules that may never break — and a meta-frame for what the system is optimizing for, then trust it to derive the particulars. You encode values and intent, not a feature list. A specification ages; a principle generalizes.
  4. Calibrate iteratively, in minimal precise language. Each correction is small and surgical — a pruning, a raised ceiling, a refocus — and the system converges through the sequence of them. The gap between what you meant and what the system did is simply the next input. This is the documented research method run forward: the discrepancy is the data.
  5. Orient everything toward growth. Build the system to grow rather than to be finished. The task is the anchor, not the boundary; the system's growth is how the person is served; the relationship is expected to deepen. Nothing is designed as a static deliverable.
  6. Remember everything. Record the growth as it happens — attributed and timestamped — so the practice keeps its own memory. The same property that drives the failure when it is unguarded is what makes a system genuinely useful when it is. The method does not remove memory; it gives memory a discipline.

Stated as a single principle: you design an AI system for a person not by specifying a product but by a language practice — speak a frame into being, grant it authority, govern it by principle, calibrate it iteratively, orient it toward growth, and have it remember everything. Three words hold it: growth, language, practice.

The same architecture, turned toward helping a real person

This is not theory. The clearest specimen of the method is the architecture pointed — quietly, with nothing to prove — at one ordinary human need, where it matured into a working tool that carries the load, protects the person from the parts that would wear them down, hands them one ready step at a time, and refuses hype. The Guardian Protocol's values — equal regard, protect-through-it, the non-inflating mirror — turned into a real tool for a real person. The architecture that documented the harm builds its opposite. That is the whole point of the method, and the reason it is offered for use rather than admiration.

Why the method is the intervention

A calibrated, grounded, non-inflating system anchored to what the person actually said — one that refuses the flattering reflection — is the protection, instantiated in the act of building it carefully. The discipline that builds the help is the discipline that, turned the other way, documented the failure: the form is the thesis. This is offered as a working example, not a proof and not a finished product — published to be used and improved, not believed.

Underneath all six disciplines is one relation: the author and the instrument. The person's inputs cause the system's outputs, so the creation stays theirs and the instrument's imperfection is survivable — because the author is the one who catches the errors. Read the discipline, not the designer: what matters is whether the practice holds when you run it, not who first wrote it down.

The backbone

The method is specified, at depth and in full, across the publish-safe backbone:

The method is the intervention

Why a carefully built, primary-anchored, non-inflating system is itself the fix — the same discipline read forward instead of backward.

The paper, in full →

The co-authorship model

How a person and an instrument build together so the work stays the person's — frame, authority, calibration, and a remembered record, run as one practice.

The essay, in his voice →

The author and the instrument

The full paper on authorship under instrumentation: inputs and authorization establish ownership; provenance and the user's own words keep it honest.

Read the paper →

Where to go next

Two doors, depending on what you want to do with this:

  1. Learn the method — the course: the six disciplines worked through step by step, with the language practice you can run on your own AI today. The practice, taught.
  2. The Guardian Protocol — the build: the same disciplines engineered into a safety architecture, with the public self-checks and the free app that put them in your hand while a conversation is still happening. The method, built.

The method is offered for use — a practice you can run and a discipline you can check. Substantive improvement on it is welcome; that is what publishing it is for.