SLICE
The code that decides the verdict.Found by running the tests with coverage and mutation checks, not guessed from file names.
Researching…
The sufficiency calculator for AI code context.
For each real code change, SuffBench records which parts of the repository actually decide whether a fix passes. With those labels, any context-selection method can be scored on coverage and cost offline, without running a model.
The code that decides the verdict.Found by running the tests with coverage and mutation checks, not guessed from file names.
What must keep passing.The tests that pass before and after the change.
Coverage and cost, offline.Did a context include the slice, and how many tokens did it spend?
Why it exists
Context retrieval for code-editing models is judged almost entirely by whether the final fix passes. That score mixes the context, the model and the luck of the run into one number.
A failed fix could mean the context missed something, buried it, or the model made a mistake. End-task pass rate cannot tell these apart.
Our paper proves, in a simplified model, that when a model’s accuracy falls as its input grows, a sufficient context can stop being sufficient when you add to it. Context has a cost, so it should be measured.
Other benchmarks mark relevance by the files a fix touched or an agent opened. SuffBench marks it by what changes a test result.
Think of an exam. Today we only check the final grade. SuffBench also records which pages of the textbook the marking scheme actually depends on. Then, before anyone sits the exam, you can check whether a revision pack covers those pages and how thick it is.
The format
SuffBench reuses the fail-to-pass and pass-to-pass structure of SWE-bench and Multi-SWE-bench, and adds verdict-based slices.
Repository, base and accepted commits, the task text with patch details removed, the accepted fix and its write set.
Fail-to-pass tests, the pass-to-pass frame set, and flaky tests excluded with their clean-run counts.
Backward and forward slices, a coverage slice, a mutation slice and a closure check on the import graph.
Token counts per unit for a named public tokenizer, so every context has a price as well as a coverage.
Optional ready-made contexts, from the write set alone to the full transitive closure, each marked expressible or not.
Tool versions, platform, runner image and dates, plus model training cutoffs so leaked instances can be flagged.
instance: org__repo-3f2a9c1d4e5b
change:
task_text: issue text, patch details removed
write_set (W): [internal/cache]
verdict:
fail_to_pass: [TestEvictOnExpiry]
pass_to_pass: the frame set
slices:
backward B(W), forward F(W)
coverage, mutation
closure_check: holds | violations listed
cost:
tokens per unit, named contexts
Illustrative only. No instances have been released. The normative fields are in the JSON Schema in the specification repository.
The calculation
Each slice is a different, checkable answer to “which code matters to this verdict?”
Backward and forward slices come from a stated dependency graph. For Go at package granularity, the reference is go list -json -deps -test ./... at the base commit.
Run the verdict tests on the accepted fix with per-unit coverage. A unit is in the slice if any of its statements runs. Cheap, deterministic apart from flaky tests, and required.
Inject small, interface-preserving mutants into covered units. A unit is in the slice if some mutant changes a test verdict. Coverage approximates execution; mutation approximates reliance.
Compare what the test binaries really import with the write set plus its slices. Anything outside is listed, as a measured gap in the dependency graph.
For each named context, record whether every identifier, declaration and import the accepted fix uses is present. A context that cannot express the fix cannot be sufficient.
Failing tests are rerun at least three times on a clean tree. Releases state each model's training cutoff and flag instances created before it.
Planned tooling · suffctx
The specification also defines the command-line contract for a checker and context compiler for Go modules. The interface is written. The implementation is not released.
suffctx checkFlags code that hides dependencies from the import graph: reflection, compiler-directive links (go:linkname), side-effecting init(), shared mutable globals, undocumented exports, and tests not tied to the code they check.
suffctx contextEmits a deterministic context for a write set: interfaces only, bodies, or the full transitive set, with token counts and an optional budget.
suffctx sliceComputes the coverage slice and closure check for a set of tests, and the mutation slice on request.
For a module that passes suffctx check, the interface-level context would contain every fact the edit depends on through the import graph. When the interface is unchanged, nothing downstream would need to be included.
It makes no claim about whether a model succeeds. That is measured, not compiled. Channels outside the import graph that the checker does not forbid are listed, not guaranteed.
Roadmap
The format is published and open. The tools come next. Here is exactly what exists today and what follows.
Version 0.1.0-spec defines the instance format, slice definitions, datasheet and tool contract, openly licensed. No instances and no tooling have been released yet.
SuffBench labels what decides the verdict. What a given model needs to write the fix depends on the model, and is measured by experiment.
The reference dependency graph and the suffctx tools target Go modules first. Other languages need their own graph tools and rules.
The research behind it reports a proof in a simplified model and a mutation measurement on one Go repository. Released instances and language-model experiments come next; none have been reported yet.