FRAME
Exactly what the change needs.The change, what it depends on, what must not move, and how it will be judged.
Experimental
Precision context for AI-built software.
Give a model exactly what a change needs, and nothing more. The context frame is what you write, review and keep. The code is generated from it, accepted mechanically and replaced freely.
Exactly what the change needs.The change, what it depends on, what must not move, and how it will be judged.
One artifact per frame.No repository access. No code review. No test writing.
Mechanical, not manual.The oracle decides. Repeated failure means the frame is short.
In plain terms
Imagine asking a cook to make tonight's soup.
Give a new cook every book you own and ask for soup. The right page is in there, but so is everything else. They get distracted, and the soup suffers. More pages, worse soup.
The dish. The five ingredients, set out on the counter. What must stay as it is: the other pans on the stove, and no nuts. The cook already knows how to cook. That card is the frame.
Nobody stands over the cook. A taste test decides. If it fails, rewrite the card and cook again. You never scrape and fix the soup.
What we know
Retrieval usually assumes extra information never hurts. That holds for a compiler or a careful human, not for a language model.
A language model's accuracy falls as its input grows. Irrelevant material is not free: it competes with what matters.
Under a stylised coverage × degradation model, a sufficient context can stop being sufficient when you add to it, and the loss grows with the irrelevant material. Mechanised in Lean 4.
On a real Go repository, some body-only changes were caught only by tests outside the changed package. Locality has to be shown, not assumed.
A change request is a contract (Δ, D, I) with an explicit frame condition and an oracle that approximates it. A context K is ε-sufficient for a model when the model's success on K is within ε of its success on the best context in a stated candidate family, not on the whole repository.
For a lossless reader, adding context never breaks sufficiency. For a degrading reader it can. That is the gap precision context is built on: the right frame is found as much by removing irrelevant material as by adding what is needed.
The method
Write the frame in four parts, then let generation and acceptance run without a human reading the code.
Failed? The frame was insufficient. Fix the frame and regenerate. Never patch the output.
Describe what will be different from the outside: inputs, outputs, visible states. If you cannot say how it would be observed, it is not ready to frame.
Include the signatures, schemas, IDs and contracts the change touches. Leave out their implementations, history and neighbours. Every extra line has to earn its place.
The frame condition is the most valuable and most often missing part. Unchanged interfaces, untouched behaviour and constraints belong here, explicitly.
The model receives the frame and returns one named artifact. It does not browse the repository, write tests, review code or fix anything outside its output.
The oracle is part of the frame but executes outside it. Acceptance is mechanical: pass is accepted, fail returns to the frame. Once a frame shape has earned trust, nobody reads the diff.
Our Go measurement shows some changes are caught only by tests in other packages. The oracle must cover where a change can be observed, not just where it is written.
Anatomy of a frame
A catalogue page needs an empty state. This is everything the model receives.
frame: CAT-02 empty state
change (Δ):
products.json is []
→ show "No products found."
depends on (D):
#product-list ul
#empty-state p, hidden
products.json array of
{id, name, description,
url, image|null}
holds (I):
IDs and classes unchanged
populated render unchanged
no inline CSS or JS
oracle:
[] → message visible
[1 item] → message hidden
existing tests still pass
output: scripts/catalogue.js
Each omission removes material the model would otherwise have to read past.
The model returns one file. The oracle runs its checks. Pass: the file is accepted as-is. Fail: look for what Δ, D or I left out, revise the frame and regenerate. One unlucky run may only need a retry; a pattern of failures always means the frame.
The frame is the reviewable artifact: about twenty lines, stable across regenerations.
Use only the frame. Produce exactly the named artifact. If the frame conflicts with itself or is plainly insufficient, say so instead of guessing.
Running the oracle, assembling artifacts and accepting a release belong to the system and the engineer, not the model.
Consequences
These follow from the idea. They are directions we are exploring, not results we have measured.
Engineers review twenty-line frames instead of hundred-line diffs. The frame is where intent lives.
Version the frames and the accepted releases. Treat generated code like a compiled binary: reproducible in the behaviour the oracle checks, not byte for byte.
A failure points at a missing dependency or invariant: a fact to add, not a line to patch.
Each frame names its own inputs and output. Frames that do not share an output can be generated at the same time.
A frame and its oracle do not depend on a vendor. Sufficiency is measured per model, so re-check frames after a switch, but the acceptance bar stays the same.
Track which frame shapes pass first time. Frames that are reliably sufficient earn trust from evidence, not assumption.
Restructuring the work
If the frame is the unit of work, the shape of a software team changes with it.
Prose requirements give way to Δ, D, I and an oracle. A change that cannot be framed is not ready to build.
Senior time goes into dependencies and invariants, where judgement matters, not into typing the implementation.
Engineers write the oracle before generation. The model never writes the checks that judge it.
Frames, interfaces and oracles are the source of truth. Code is regenerated from them, and accepted releases are kept.
Narrow interfaces and local observability make frames small. Good modular design becomes a direct cost lever.
A rejected output shows what the frame left out. Teams build a library of frame shapes that pass, by change type.
One engineer may hold all four. The point is that each is explicit and none belongs to the model.
Cost, in theory
None of this is measured. It is the arithmetic that follows if frames are sufficient, and where the cost moves instead.
Input scales with the frame, not the repository. A model that does not explore does not pay to read what it ignores.
If irrelevant context degrades a model, removing it should raise the first-pass rate and cut retries.
Engineer time is usually the largest cost. Reviewing a twenty-line frame should be faster than reviewing the diff it produces.
A precise frame may let a cheaper model succeed where a larger one was needed to cope with noise. A hypothesis to test.
When a model changes, rerunning accepted frames costs tokens and oracle time, not a rewrite.
Writing frames and oracles is up-front engineering. The saving only exists if that work is less than the review and rework it replaces.
In practice
Precision context needs no new platform: a frame file, one model call and the CI you already run. What matters is choosing the right changes.
Bounded changes behind clear interfaces: UI states, data transforms, API handlers, validators, adapters and migrations with fixtures.
Exploratory work, unclear requirements, cross-cutting refactors, and anything without a checkable oracle, such as visual taste or untested performance.
Start where tests already exist. Choose the oracle by where a change can be observed, including other packages.
Earned, not assumed. Remove human review only for frame shapes whose record shows they pass the oracle reliably.
This is how the cost theory above becomes data for your own codebase.
What we don't claim
Precision context is experimental. These limits are part of the method, not footnotes to it.
The non-monotonicity result is a proof under a stylised model, mechanised in Lean 4. No language-model benchmark has been reported.
ε-sufficient means close to the best frame tried, for a stated model. It is a bound, not a guarantee of correct output.
The same frame can produce different code across runs and model versions. The oracle, not the model, is what makes regeneration safe.
A planned benchmark whose instances carry observation-based oracle slices, so frame sufficiency can be measured on real models and real changes. About SuffBench