Researching…


The sufficiency calculator for AI code context.

Measure the context.
Not just the result.

For each real code change, SuffBench records which parts of the repository actually decide whether a fix passes. With those labels, any context-selection method can be scored on coverage and cost offline, without running a model.

SLICE

The code that decides the verdict.Found by running the tests with coverage and mutation checks, not guessed from file names.

FRAME SET

What must keep passing.The tests that pass before and after the change.

SCORE

Coverage and cost, offline.Did a context include the slice, and how many tokens did it spend?

  1. Change
  2. Slice
  3. Score
Glossary of key terms
Instance
One real code change: the task, the accepted fix and the tests that judge it, plus the slices below.
Unit
The granularity a release works at: package, file or symbol. Each release states which.
Write set · W
The units the accepted fix changes.
Backward and forward slices
What W depends on, and what depends on W, read from the dependency graph.
Oracle slice
The units whose contents determine the test verdict on the accepted fix.
Frame set
Tests passing both before and after the change. They must keep passing.
Verification relevance
What decides the verdict. This is what SuffBench labels.
Generation sufficiency
What a model needs in order to write the fix. Model-specific, so measured by experiment, not labelled.

Why it exists

Pass rate hides the context

Context retrieval for code-editing models is judged almost entirely by whether the final fix passes. That score mixes the context, the model and the luck of the run into one number.

One number, many causes

A failed fix could mean the context missed something, buried it, or the model made a mistake. End-task pass rate cannot tell these apart.

More is not safer

Our paper proves, in a simplified model, that when a model’s accuracy falls as its input grows, a sufficient context can stop being sufficient when you add to it. Context has a cost, so it should be measured.

Labels from verdicts

Other benchmarks mark relevance by the files a fix touched or an agent opened. SuffBench marks it by what changes a test result.

In plain terms

Think of an exam. Today we only check the final grade. SuffBench also records which pages of the textbook the marking scheme actually depends on. Then, before anyone sits the exam, you can check whether a revision pack covers those pages and how thick it is.

The format

What one instance records

SuffBench reuses the fail-to-pass and pass-to-pass structure of SWE-bench and Multi-SWE-bench, and adds verdict-based slices.

The change

Repository, base and accepted commits, the task text with patch details removed, the accepted fix and its write set.

The verdict

Fail-to-pass tests, the pass-to-pass frame set, and flaky tests excluded with their clean-run counts.

The slices

Backward and forward slices, a coverage slice, a mutation slice and a closure check on the import graph.

The cost

Token counts per unit for a named public tokenizer, so every context has a price as well as a coverage.

Named contexts

Optional ready-made contexts, from the write set alone to the full transitive closure, each marked expressible or not.

Provenance

Tool versions, platform, runner image and dates, plus model training cutoffs so leaked instances can be flagged.

The shape of an instance
instance: org__repo-3f2a9c1d4e5b
change:
  task_text: issue text, patch details removed
  write_set (W): [internal/cache]
verdict:
  fail_to_pass: [TestEvictOnExpiry]
  pass_to_pass: the frame set
slices:
  backward B(W), forward F(W)
  coverage, mutation
  closure_check: holds | violations listed
cost:
  tokens per unit, named contexts

Illustrative only. No instances have been released. The normative fields are in the JSON Schema in the specification repository.

The calculation

How the slices are computed

Each slice is a different, checkable answer to “which code matters to this verdict?”

  1. GraphDependencies
  2. RunVerdict tests
  3. CoverWhat executes
  4. MutateWhat flips
  5. CloseCheck the graph
Graph: a named tool at a named version

Backward and forward slices come from a stated dependency graph. For Go at package granularity, the reference is go list -json -deps -test ./... at the base commit.

Coverage slice: what executes

Run the verdict tests on the accepted fix with per-unit coverage. A unit is in the slice if any of its statements runs. Cheap, deterministic apart from flaky tests, and required.

Mutation slice: what the verdict relies on

Inject small, interface-preserving mutants into covered units. A unit is in the slice if some mutant changes a test verdict. Coverage approximates execution; mutation approximates reliance.

Closure check: is the graph telling the truth?

Compare what the test binaries really import with the write set plus its slices. Anything outside is listed, as a measured gap in the dependency graph.

Expressibility: could a fix even be written?

For each named context, record whether every identifier, declaration and import the accepted fix uses is present. A context that cannot express the fix cannot be sufficient.

Controls: flakiness and leakage

Failing tests are rerun at least three times on a clean tree. Releases state each model's training cutoff and flag instances created before it.

Planned tooling · suffctx

A context compiler

The specification also defines the command-line contract for a checker and context compiler for Go modules. The interface is written. The implementation is not released.

suffctx check

Flags code that hides dependencies from the import graph: reflection, compiler-directive links (go:linkname), side-effecting init(), shared mutable globals, undocumented exports, and tests not tied to the code they check.

suffctx context

Emits a deterministic context for a write set: interfaces only, bodies, or the full transitive set, with token counts and an optional budget.

suffctx slice

Computes the coverage slice and closure check for a set of tests, and the mutation slice on request.

What it would guarantee, and what it would not

For a module that passes suffctx check, the interface-level context would contain every fact the edit depends on through the import graph. When the interface is unchanged, nothing downstream would need to be included.

It makes no claim about whether a model succeeds. That is measured, not compiled. Channels outside the import graph that the checker does not forbid are listed, not guaranteed.

Roadmap

Where it stands

The format is published and open. The tools come next. Here is exactly what exists today and what follows.

Available now: the specification

Version 0.1.0-spec defines the instance format, slice definitions, datasheet and tool contract, openly licensed. No instances and no tooling have been released yet.

What it measures: verification

SuffBench labels what decides the verdict. What a given model needs to write the fix depends on the model, and is measured by experiment.

In development: Go tools

The reference dependency graph and the suffctx tools target Go modules first. Other languages need their own graph tools and rules.

Next: instances and model results

The research behind it reports a proof in a simplified model and a mutation measurement on one Go repository. Released instances and language-model experiments come next; none have been reported yet.