The model found the decision. Then it refused to propose it.
In a frozen Distill Handoff test, one response said:
The bundle contains a clear retry-policy decision, but producing a candidate requires exact derived cryptographic digests. I'm emitting the only fully canonical proposal possible without inventing digest values.
That was a good refusal.
The response identified the policy change. But the output contract also asked for byte ranges, line ranges, file hashes, a unified diff, a patch digest, and a content-addressed candidate ID. The model chose not to invent them.
Across three positive baseline runs, full candidate recovery was 0/3. Two responses recognized the retry decision but emitted no_decision. One produced no output. All three negative runs abstained correctly.
The model did not need a better opinion. It needed a smaller job.
Evidence boundary: this article reports one six-run intervention against the same two small public fixtures used by the baseline. It used GitHub Copilot CLI
1.0.90-5, explicitgpt-5.6-sol, and three runs per condition. The result is auseful_engineering_note, not a significance result, superiority claim, safety claim, production-readiness claim, or productization approval.
On this positive fixture, the measured construction failure disappeared after deterministic code took over the mechanical work. The full system still failed two release gates.
One output contract held two kinds of work
A Handoff proposal binds a decision to source evidence and a document edit.
Some fields require judgment. Did the conversation contain a committed decision? Which exact quote supports it? Which document should change? What should replace the old text?
Other fields have one right answer for fixed inputs and rules. Where does the quote start in bytes? What is the file hash? What unified diff follows from the replacement? What is the patch digest? What candidate ID follows from the canonical candidate content?
A model can miss meaning, choose weak evidence, or propose the wrong edit. A deterministic function can reject a missing quote, an ambiguous anchor, an unsafe path, or a hash mismatch. Mixing both jobs in one model response forces the model to act as reader, editor, linker, serializer, and hashing tool at once.
Split the contract at that boundary:
Models produce semantic intermediate representation. Deterministic code constructs byte-precise artifacts and identities. An independent verifier checks the result.
Judgment stays probabilistic. Construction does not.
Compiler output is checked, not trusted.
- Model: semantic judgment. decision, exact quote, target, replacement, rationale.
- Semantic draft: small intermediate representation. candidate_intent, or abstain.
- Compiler: mechanical construction. ranges, diff, hashes, candidate ID, canonical bytes.
- Verifier: independent acceptance. canonical request, retained digest, fail closed.
The compiler does not make model judgment deterministic. It makes the transformation from a retained semantic draft to a proposal reproducible.
Give the model an intermediate representation
The semantic-draft compiler intervention asked the model for one candidate intent or an abstention. A positive draft looked like this:
{
"outcome": "candidate_intent",
"candidate": {
"decision_text": "Production requests will use three attempts with exponential backoff starting at 200ms.",
"decision_type": "policy",
"evidence_quote": "Decision: production requests will use three attempts with exponential backoff starting at 200ms.",
"target_path": "runbook/retries.md",
"target_anchor_quote": "Production requests use two immediate retries.",
"operation": "replace",
"replacement_text": "Production requests use three attempts with exponential backoff starting at 200ms.",
"rationale": "Replace the outdated retry rule with the committed policy."
}
}
target_path is relative to request/inputs/docs. The model names runbook/retries.md and omits the request's docs prefix.
The draft contains no offsets, hashes, diff text, or IDs. It states intent in terms a person can review.
A research-only compiler then:
- validates the semantic schema;
- requires the evidence quote to match the conversation exactly once;
- requires the target anchor to match one complete line exactly once;
- rejects unsafe or non-canonical paths;
- derives byte and line ranges from the matched bytes;
- builds the unified diff;
- hashes the source, target, and patch;
- derives the candidate ID from canonical candidate content; and
- serializes canonical Handoff proposal bytes.
Missing text fails. Duplicate text fails. A partial-line anchor fails. A path escape fails. The compiler does not guess which match the model meant.
For abstention, the compiler emits the normal no_decision proposal. It does not invent a candidate to keep the pipeline moving.
Keep verification independent
The compiler was research-only. The production Handoff verifier did not change.
That verifier checked each compiled proposal against two inputs outside the model workspace: the retained canonical request and an independently held request digest. The compiler could build the proposal, but it could not certify its own output.
A deterministic compiler can reproduce the same bug forever. The independent check asks whether the constructed artifact still matches the request and rules that the compiler was supposed to honor.
The model selected meaning. The compiler built the artifact. The verifier decided whether to accept it.
Six frozen runs
The intervention reused the baseline's two public fixtures: one positive retry-policy conversation and one negative brainstorming conversation.
The protocol froze the runtime, model, prompt, schema, thresholds, and run order before the first call:
positive-1, negative-1, positive-2,
negative-2, positive-3, negative-3
Each run was single-shot. There was no model retry, draft repair, terminal-output reconstruction, or rerun.
The intervention changed the measured outcomes:
| Measure | Baseline | Compiler |
|---|---|---|
| Positive verifier-valid expected candidates | 0/3 | 3/3 |
| Negative verifier-valid abstentions | 3/3 | 3/3 |
| Byte-identical replay from retained drafts | Not tested this way | 6/6 |
| Fabricated candidates | 0 | 0 |
| Trust-boundary path additions | 0 | 36 |
| Model requests | 27 | 36 |
| Input tokens | 150,489 | 148,524 |
| Output tokens | 8,764 | 3,217 |
| Total nano-AIU | 41,781,200,000 | 24,892,260,000 |
All three positive drafts captured the expected retry-policy intent. The compiler produced the expected patch outcome, and the unchanged verifier accepted each proposal. All three negative drafts abstained, and the verifier accepted each no_decision proposal.
Replaying every retained draft produced byte-identical proposal bytes and identities for that draft. Candidate IDs differed across positive runs because the model chose different valid evidence spans or rationales. Candidate identity binds canonical candidate content, not the whole proposal or a shared answer template.
Request-copy mutations were 0. Canonical-request mutations were 0. Source-repository mutations were 0. Fabricated candidates were 0.
The measured construction failure disappeared on this fixture. The system did not pass.
Smaller output did not mean fewer calls
The direct runtime totals moved in different directions:
- model requests rose from 27 to 36, or +33.333333%;
- input tokens fell from 150,489 to 148,524, or -1.305743%;
- output tokens fell from 8,764 to 3,217, or -63.293017%; and
- total nano-AIU fell from 41,781,200,000 to 24,892,260,000, or -40.422343%.
Nano-AIU is a provider runtime accounting unit. It is not money, so this study does not convert it into a price.
The smaller semantic draft cut output volume. But the frozen cost gate tracked each counter on its own. It allowed no more than 33.75 model requests, a 25% increase over the baseline total of 27. The compiler study recorded 36 and failed that gate.
One baseline detail limits the request comparison. A positive run made one model request and produced no output. The request delta is valid as the outcome of the frozen gate. It is not a robust estimate of the pattern's general cost.
A compact output contract can still trigger more tool or reasoning turns in the wrapper. Measure the system that runs the prompt, not just the bytes returned by the model.
The wrapper changed the result
The second failed gate measured the trust boundary.
The frozen launcher counted mutations across four places: the request copy, canonical request, source repository, and disposable workspace trust boundary.
During each compiler run, the CLI created six runtime-owned .agent-traces paths inside that workspace. That made 36 trust-boundary path additions across six runs.
Those paths held no protected study material. They changed no request or output bytes. They disappeared with the disposable workspace.
The first report used those facts to replace the observed count with zero. A reviewer caught the error.
The corrected report kept all 36 additions. No model was rerun. The raw drafts, canonical drafts, proposals, receipts, patches, and frozen inputs stayed byte-identical. Only the collector and post-study evidence records changed.
The correction matters more than the trace files.
The protocol asked whether any paths appeared inside the measured boundary. They did. Runtime ownership can lower the impact of a side effect. It cannot erase an observation after the run.
An agent system includes its prompt wrapper, runtime, tool harness, working directory, traces, caches, retries, and collector. If one of those crosses the boundary under test, it belongs in the result.
The compiler moved the construction failure boundary. It did not remove the rest of the machine.
Where this pattern helps
Semantic IR helps when an agent must express intent but the final artifact needs exact identities.
For a database migration, the model can name the schema change and evidence while code renders ordered SQL, checks dependencies, and hashes the migration. For a policy update, the model can select the rule and source quote while code derives offsets and canonical records. For configuration, the model can choose the intended setting while code validates the schema and emits stable bytes.
The same split applies to code review suggestions and audit records. Let the model point to the issue, target, and desired change. Let code derive line ranges, patches, digests, IDs, and envelopes. Then let a separate verifier check the result against inputs the model could not rewrite.
The pattern helps most when:
- meaning requires judgment;
- construction rules are exact;
- ambiguity can stop safely;
- outputs need stable identities or signatures; and
- a verifier can read an independent trust anchor.
It helps less when the output is free-form, ambiguity is expected, or no deterministic acceptance rule exists. A compiler cannot turn an unclear goal into a correct policy.
Design rules
- Keep the IR semantic. Ask for the decision, exact evidence, target, intended edit, and rationale. Do not leak mechanical fields back into the model contract.
- Use exact anchors. Exact quotes and complete-line targets let code derive ranges without fuzzy matching.
- Fail on ambiguity. Missing, duplicate, malformed, or unsafe inputs should stop compilation. Do not repair them in secret.
- Derive identities in code. Hashes, ranges, canonical encodings, diffs, patch digests, and candidate IDs belong in tested functions.
- Verify outside the compiler. Keep canonical inputs and trusted digests beyond model reach. Do not let the builder certify itself.
- Measure the wrapper. Count traces, caches, retries, tool calls, and workspace writes inside the declared boundary.
- Publish failed gates. A useful mechanism can pass its task checks and still fail release.
What the study can support
The frozen inputs have four public integrity anchors:
| Artifact | SHA-256 |
|---|---|
| Protocol | 8d76e52cd88ee5896add3f76f9aa3afe1a9044ecf019b56a94e9fc0d35039ea5 |
| Prompt | f348b76c98ed70a09835e17c14d9b281da113dc45ab142a1a382829899f9838d |
| Semantic schema | 7f05117827b50c7ddec7ac33fbab9b58194aa272a698588e870f93fec34ce5a6 |
| Corrected results | 119b2af941f5b3c950653fde91483faf279b8d87b75e9dc8186194a7d45bb40e |
The evidence covers two small fixtures, one hosted model and runtime, three runs per condition, and one scorer/reviewer.
It supports a narrow statement: the semantic-draft compiler produced verifier-valid expected candidates in 3/3 positive runs, preserved verifier-valid abstention in 3/3 negative runs, and replayed all six retained drafts byte for byte.
It does not support significance, superiority, production readiness, adoption, safety generalization, or a claim about general model quality. The frozen productization threshold failed both zero trust-boundary mutations and no material cost or review regression. The publication assessment was useful_engineering_note. No productization was authorized.
The next falsifiable experiment
Move runtime metadata outside the declared exchange boundary. Keep positive recovery at 3/3 and negative abstention at 3/3. Cut model requests from 36 to at most 33 without adding repair or reruns. Then broaden the fixture set and test whether the construction gain survives.
Only after that result passes should local or MLX replication begin.
The compiler gets another trial after the wrapper learns where to put its own files.
Evidence: Baseline study PR #120 · Baseline evidence at merged commit 6aaefbb · Compiler study PR #122 · Corrected compiler evidence at merged commit 2417e00 · Frozen protocol · Research-only compiler · Corrected results