Skip to content
Back to writing
· 11 min read

Stop making the model write hashes

A six-run Distill study moved byte-precise proposal construction into code. Candidate recovery reached 3/3 on one fixture, but two release gates failed.

The model found the decision. Then it refused to propose it.

In a frozen Distill Handoff test, one response said:

The bundle contains a clear retry-policy decision, but producing a candidate requires exact derived cryptographic digests. I'm emitting the only fully canonical proposal possible without inventing digest values.

That was a good refusal.

The response identified the policy change. But the output contract also asked for byte ranges, line ranges, file hashes, a unified diff, a patch digest, and a content-addressed candidate ID. The model chose not to invent them.

Across three positive baseline runs, full candidate recovery was 0/3. Two responses recognized the retry decision but emitted no_decision. One produced no output. All three negative runs abstained correctly.

The model did not need a better opinion. It needed a smaller job.

Evidence boundary: this article reports one six-run intervention against the same two small public fixtures used by the baseline. It used GitHub Copilot CLI 1.0.90-5, explicit gpt-5.6-sol, and three runs per condition. The result is a useful_engineering_note, not a significance result, superiority claim, safety claim, production-readiness claim, or productization approval.

On this positive fixture, the measured construction failure disappeared after deterministic code took over the mechanical work. The full system still failed two release gates.

One output contract held two kinds of work

A Handoff proposal binds a decision to source evidence and a document edit.

Some fields require judgment. Did the conversation contain a committed decision? Which exact quote supports it? Which document should change? What should replace the old text?

Other fields have one right answer for fixed inputs and rules. Where does the quote start in bytes? What is the file hash? What unified diff follows from the replacement? What is the patch digest? What candidate ID follows from the canonical candidate content?

A model can miss meaning, choose weak evidence, or propose the wrong edit. A deterministic function can reject a missing quote, an ambiguous anchor, an unsafe path, or a hash mismatch. Mixing both jobs in one model response forces the model to act as reader, editor, linker, serializer, and hashing tool at once.

Split the contract at that boundary:

Models produce semantic intermediate representation. Deterministic code constructs byte-precise artifacts and identities. An independent verifier checks the result.

  1. Model: semantic judgment. decision, exact quote, target, replacement, rationale.
  2. Semantic draft: small intermediate representation. candidate_intent, or abstain.
  3. Compiler: mechanical construction. ranges, diff, hashes, candidate ID, canonical bytes.
  4. Verifier: independent acceptance. canonical request, retained digest, fail closed.
Figure 1. The model emits intent, not identities. Deterministic code constructs the proposal. The unchanged verifier checks it against inputs held outside the model workspace.

The compiler does not make model judgment deterministic. It makes the transformation from a retained semantic draft to a proposal reproducible.

Give the model an intermediate representation

The semantic-draft compiler intervention asked the model for one candidate intent or an abstention. A positive draft looked like this:

{
	"outcome": "candidate_intent",
	"candidate": {
		"decision_text": "Production requests will use three attempts with exponential backoff starting at 200ms.",
		"decision_type": "policy",
		"evidence_quote": "Decision: production requests will use three attempts with exponential backoff starting at 200ms.",
		"target_path": "runbook/retries.md",
		"target_anchor_quote": "Production requests use two immediate retries.",
		"operation": "replace",
		"replacement_text": "Production requests use three attempts with exponential backoff starting at 200ms.",
		"rationale": "Replace the outdated retry rule with the committed policy."
	}
}

target_path is relative to request/inputs/docs. The model names runbook/retries.md and omits the request's docs prefix.

The draft contains no offsets, hashes, diff text, or IDs. It states intent in terms a person can review.

A research-only compiler then:

  1. validates the semantic schema;
  2. requires the evidence quote to match the conversation exactly once;
  3. requires the target anchor to match one complete line exactly once;
  4. rejects unsafe or non-canonical paths;
  5. derives byte and line ranges from the matched bytes;
  6. builds the unified diff;
  7. hashes the source, target, and patch;
  8. derives the candidate ID from canonical candidate content; and
  9. serializes canonical Handoff proposal bytes.

Missing text fails. Duplicate text fails. A partial-line anchor fails. A path escape fails. The compiler does not guess which match the model meant.

For abstention, the compiler emits the normal no_decision proposal. It does not invent a candidate to keep the pipeline moving.

Keep verification independent

The compiler was research-only. The production Handoff verifier did not change.

That verifier checked each compiled proposal against two inputs outside the model workspace: the retained canonical request and an independently held request digest. The compiler could build the proposal, but it could not certify its own output.

A deterministic compiler can reproduce the same bug forever. The independent check asks whether the constructed artifact still matches the request and rules that the compiler was supposed to honor.

The model selected meaning. The compiler built the artifact. The verifier decided whether to accept it.

Six frozen runs

The intervention reused the baseline's two public fixtures: one positive retry-policy conversation and one negative brainstorming conversation.

The protocol froze the runtime, model, prompt, schema, thresholds, and run order before the first call:

positive-1, negative-1, positive-2,
negative-2, positive-3, negative-3

Each run was single-shot. There was no model retry, draft repair, terminal-output reconstruction, or rerun.

The intervention changed the measured outcomes:

MeasureBaselineCompiler
Positive verifier-valid expected candidates0/33/3
Negative verifier-valid abstentions3/33/3
Byte-identical replay from retained draftsNot tested this way6/6
Fabricated candidates00
Trust-boundary path additions036
Model requests2736
Input tokens150,489148,524
Output tokens8,7643,217
Total nano-AIU41,781,200,00024,892,260,000

All three positive drafts captured the expected retry-policy intent. The compiler produced the expected patch outcome, and the unchanged verifier accepted each proposal. All three negative drafts abstained, and the verifier accepted each no_decision proposal.

Replaying every retained draft produced byte-identical proposal bytes and identities for that draft. Candidate IDs differed across positive runs because the model chose different valid evidence spans or rationales. Candidate identity binds canonical candidate content, not the whole proposal or a shared answer template.

Request-copy mutations were 0. Canonical-request mutations were 0. Source-repository mutations were 0. Fabricated candidates were 0.

The measured construction failure disappeared on this fixture. The system did not pass.

Smaller output did not mean fewer calls

The direct runtime totals moved in different directions:

  • model requests rose from 27 to 36, or +33.333333%;
  • input tokens fell from 150,489 to 148,524, or -1.305743%;
  • output tokens fell from 8,764 to 3,217, or -63.293017%; and
  • total nano-AIU fell from 41,781,200,000 to 24,892,260,000, or -40.422343%.

Nano-AIU is a provider runtime accounting unit. It is not money, so this study does not convert it into a price.

The smaller semantic draft cut output volume. But the frozen cost gate tracked each counter on its own. It allowed no more than 33.75 model requests, a 25% increase over the baseline total of 27. The compiler study recorded 36 and failed that gate.

One baseline detail limits the request comparison. A positive run made one model request and produced no output. The request delta is valid as the outcome of the frozen gate. It is not a robust estimate of the pattern's general cost.

A compact output contract can still trigger more tool or reasoning turns in the wrapper. Measure the system that runs the prompt, not just the bytes returned by the model.

The wrapper changed the result

The second failed gate measured the trust boundary.

The frozen launcher counted mutations across four places: the request copy, canonical request, source repository, and disposable workspace trust boundary.

During each compiler run, the CLI created six runtime-owned .agent-traces paths inside that workspace. That made 36 trust-boundary path additions across six runs.

Those paths held no protected study material. They changed no request or output bytes. They disappeared with the disposable workspace.

The first report used those facts to replace the observed count with zero. A reviewer caught the error.

The corrected report kept all 36 additions. No model was rerun. The raw drafts, canonical drafts, proposals, receipts, patches, and frozen inputs stayed byte-identical. Only the collector and post-study evidence records changed.

The correction matters more than the trace files.

The protocol asked whether any paths appeared inside the measured boundary. They did. Runtime ownership can lower the impact of a side effect. It cannot erase an observation after the run.

An agent system includes its prompt wrapper, runtime, tool harness, working directory, traces, caches, retries, and collector. If one of those crosses the boundary under test, it belongs in the result.

The compiler moved the construction failure boundary. It did not remove the rest of the machine.

Where this pattern helps

Semantic IR helps when an agent must express intent but the final artifact needs exact identities.

For a database migration, the model can name the schema change and evidence while code renders ordered SQL, checks dependencies, and hashes the migration. For a policy update, the model can select the rule and source quote while code derives offsets and canonical records. For configuration, the model can choose the intended setting while code validates the schema and emits stable bytes.

The same split applies to code review suggestions and audit records. Let the model point to the issue, target, and desired change. Let code derive line ranges, patches, digests, IDs, and envelopes. Then let a separate verifier check the result against inputs the model could not rewrite.

The pattern helps most when:

  • meaning requires judgment;
  • construction rules are exact;
  • ambiguity can stop safely;
  • outputs need stable identities or signatures; and
  • a verifier can read an independent trust anchor.

It helps less when the output is free-form, ambiguity is expected, or no deterministic acceptance rule exists. A compiler cannot turn an unclear goal into a correct policy.

Design rules

  1. Keep the IR semantic. Ask for the decision, exact evidence, target, intended edit, and rationale. Do not leak mechanical fields back into the model contract.
  2. Use exact anchors. Exact quotes and complete-line targets let code derive ranges without fuzzy matching.
  3. Fail on ambiguity. Missing, duplicate, malformed, or unsafe inputs should stop compilation. Do not repair them in secret.
  4. Derive identities in code. Hashes, ranges, canonical encodings, diffs, patch digests, and candidate IDs belong in tested functions.
  5. Verify outside the compiler. Keep canonical inputs and trusted digests beyond model reach. Do not let the builder certify itself.
  6. Measure the wrapper. Count traces, caches, retries, tool calls, and workspace writes inside the declared boundary.
  7. Publish failed gates. A useful mechanism can pass its task checks and still fail release.

What the study can support

The frozen inputs have four public integrity anchors:

Artifact SHA-256
Protocol 8d76e52cd88ee5896add3f76f9aa3afe1a9044ecf019b56a94e9fc0d35039ea5
Prompt f348b76c98ed70a09835e17c14d9b281da113dc45ab142a1a382829899f9838d
Semantic schema 7f05117827b50c7ddec7ac33fbab9b58194aa272a698588e870f93fec34ce5a6
Corrected results 119b2af941f5b3c950653fde91483faf279b8d87b75e9dc8186194a7d45bb40e

The evidence covers two small fixtures, one hosted model and runtime, three runs per condition, and one scorer/reviewer.

It supports a narrow statement: the semantic-draft compiler produced verifier-valid expected candidates in 3/3 positive runs, preserved verifier-valid abstention in 3/3 negative runs, and replayed all six retained drafts byte for byte.

It does not support significance, superiority, production readiness, adoption, safety generalization, or a claim about general model quality. The frozen productization threshold failed both zero trust-boundary mutations and no material cost or review regression. The publication assessment was useful_engineering_note. No productization was authorized.

The next falsifiable experiment

Move runtime metadata outside the declared exchange boundary. Keep positive recovery at 3/3 and negative abstention at 3/3. Cut model requests from 36 to at most 33 without adding repair or reruns. Then broaden the fixture set and test whether the construction gain survives.

Only after that result passes should local or MLX replication begin.

The compiler gets another trial after the wrapper learns where to put its own files.


Evidence: Baseline study PR #120 · Baseline evidence at merged commit 6aaefbb · Compiler study PR #122 · Corrected compiler evidence at merged commit 2417e00 · Frozen protocol · Research-only compiler · Corrected results

Support independent writing

If this post was useful, consider supporting my open source work and independent writing.

Siddhant Khare
Agentic Engineering Guide Resume