Skip to content
Back to writing
· 12 min read

Repeatable output is not verified evidence

A local Qwen3 study repeated every model outcome and validated every citation. All 18 runs still failed admission. A frozen claim-skeleton compiler narrowed the model's job, yet 12 of 18 new runs remained refused.

The first study repeated every model outcome byte for byte.

Every citation was valid.

The verifier refused all 18 runs.

That result came from one local Qwen3 study in VoxelScope. It ran 9 synthetic cases twice with no rerun. The repeat pairs matched at the model-outcome level. Citation validity was 342/342. Replay semantic variance was 0/9.

None of those properties made one run admissible.

Four outputs had the wrong outer JSON type. The other 14 parsed. Across those parsed drafts, the verifier excluded 56 of 238 proposed claims. No run emitted accepted prose.

Evidence boundary: this article describes two completed local studies against synthetic, non-clinical evidence. It reports one immutable v3 study and one prospectively declared v4 study. Only emitted fact and caveat coverage support a direct descriptive comparison. The evidence does not support a causal improvement, safety improvement, model-quality improvement, generalization, significance, superiority, clinical utility, production readiness, or biomedical-validity claim.

Repeatability tells you that a process can reproduce an outcome. Citation validity tells you that a claim points to an allowed source. Structured output tells you that a parser can read the envelope.

None answers the harder question:

Does this claim satisfy the exact evidence contract?

Correct citations can support an inadmissible claim

A citation can be real, in scope, and attached to the wrong unit of meaning.

Suppose an evidence plan requires one fact and one caveat:

fact: the two evidence layers agree
caveat: agreement does not establish truth or causality

The model can cite the right source for both sentences. It can repeat the same pair on every run. But a verifier may require one combined claim bound to both requirement IDs.

If the model returns two claims, each claim has a valid citation. Neither claim satisfies the combined binding.

That is what happened in v3. The model-facing plan listed fact and caveat requirements separately. The verifier-facing contract expected combined units for some pairs. The prompt allowed a single claim to bind more than one requirement, but the plan shape taught the model to produce separate claims.

The model followed one interface. The verifier enforced another.

The post-hoc failure analysis counted 28 requirement_binding_mismatch exclusions across the 14 parsed envelopes. Each parsed run hit two such exclusions.

This was not evidence that every refused draft hallucinated. It was a compiler and verifier interface conflict.

A safe negation can still trip a lexical gate

The second conflict was smaller and easier to miss.

The evidence plan required bounded caveats such as:

Agreement or disagreement does not establish truth, causality, clinical utility, or therapeutic relevance.

and:

Source-reported target evidence is not a VoxelScope target, ranking, or druggability claim.

Those are limiting statements. They block stronger interpretations.

The verifier used substring prohibitions. It searched for fragments such as caus and druggab without interpreting negation. The safe caveats therefore triggered the same rules meant to stop unsupported causal or druggability claims.

Across the parsed v3 envelopes, the failure analysis recorded 14 causal-or-certainty substring exclusions and 14 target-or-druggability substring exclusions.

The rule saw the word. It did not see what the sentence denied.

This distinction matters. A refusal can be correct at the system boundary while its reason still points to a bad interface. The verifier was right to fail closed under its frozen rules. The study was also right to classify the conflict instead of calling every refusal unsafe model behavior.

Stable failure is still failure

The repeat pairs make the result harder to dismiss.

All nine pairs were byte-identical at the model-outcome level. Four invalid outputs repeated the same output hash and byte count within their cases. The other pairs repeated the same committed draft identity.

That stability ruled out one easy story: the result did not arise because the model wandered between unrelated answers on each repeat.

But stable failure does not become success. The v3 study still recorded:

Measure Observed value
Requests 18
Refused 18/18
Invalid outer JSON types 4/18
Parsed envelopes 14/18
Citation validity 342/342
Replay semantic variance 0/9
Unsupported claims 56/238
Emitted facts 0/158
Emitted caveats 0/148
Verified-draft facts 102/158
Verified-draft caveats 80/148

The gap between verified-draft coverage and emitted coverage is the admission boundary. The verifier could retain supported claim material inside a refused run. It still could not emit final prose because required bindings or hard rules failed.

A draft can contain supported parts and remain inadmissible as a whole.

The comparison has three classes

VoxelScope changed the contract before the v4 study. That change fixed some structural problems and changed what several metrics meant.

The comparison must therefore keep three classes separate.

Class Metrics Allowed reading
Direct descriptive comparison Emitted fact coverage, emitted caveat coverage Report v3 and v4 side by side without a causal claim.
Transformed Invalid output, semantic variance, verified-draft coverage, terminal outcomes, latency, input tokens, output tokens The prompt, schema, generated contract, or semantic projection changed. Do not call movement an improvement measure.
Not comparable Task and skeleton metrics, unsupported-claim rate, citation validity, memory fields The denominator, ownership, field definition, or availability differs.

Only two measures cross the direct bridge:

Direct descriptive measure v3 v4
Emitted fact coverage 0/158 42/158
Emitted caveat coverage 0/148 60/148

Those values describe what each frozen pipeline emitted. They do not establish why the values differ.

No causal or superiority claim follows from that bridge.

The contract changed. The next study did not rerun v3 with a patched verifier. It tested a new model-facing interface while keeping verifier authority outside the model.

Freeze the skeleton before execution

V4 moved the claim plan out of the model response.

Before execution, trusted code grouped each admissible claim unit into a typed skeleton. Each skeleton carried a content-derived ID plus its claim type, requirement bindings, source bindings, and canonical caveat policy.

The model received those IDs in a fixed order. Its job narrowed to two outputs:

  1. return the ordered skeleton IDs; and
  2. draft bounded text for each requested unit.

The model did not own claim types, requirement bindings, source citations, canonical caveats, content identities, or final admitted prose.

Trusted expansion restored those fields from the frozen skeleton. The verifier then checked the expanded artifact under the same authority it held before. The compiler could build a claim. It could not admit one.

The prospective declaration froze that interface, the same nine conceptual cases, two repeats per case, the local runtime, model tag, full manifest, endpoints, comparison classes, stop rules, and one-attempt boundary before the first v4 generation request.

Its SHA-256 was:

063b6055d430f14637d3dab05d1471cfc8e2fcf043b4030c633a956eff94a6c6

The run used Ollama 0.35.1, tag qwen3:8b-q8_0, and full manifest:

e56358ca25dd14db6853a9f68a92d717aaa6f0a94250a72d1a0f3d86a9f30130

Claim admission compiler

Who owns the meaning?

Trusted code Untrusted model Verifier

Step 1 of 4

A citation sits beside the claim

V3 repeated every model outcome and validated every citation. Those checks did not prove that each claim matched the verifier contract.

V3 contract

Claim semantics cross the model boundary.

18 of 18 refused
Evidence plan Trusted input
Separate requirements

Fact and caveat IDs arrive as separate plan entries.

Model contract Model-facing
Typed claims

The model chooses claim types, bindings, citations, and draft text.

Model output Untrusted
Two valid citations

Separate fact and caveat claims follow the visible plan shape.

Compiler Trusted code
Preserves bindings

The compiler cannot invent the combined verifier unit.

Verifier Independent
Combined unit required

Binding mismatch and naive substring rules fail closed.

Terminal state Admitted result
Refused

Repeatable and correctly cited does not satisfy the contract.

Observed result layer

Keep the studies separate

Only the bridge below is directly comparable
V3 Immutable historical study
Terminal outcomes
18/18 refused
Citation validity
342/342
Semantic variance
0/9
Transformed contract measures
V4 Prospectively declared study
Terminal outcomes
6 accepted · 12 refused
Unsupported claims
108/270
Peak memory
Unavailable, not zero
Not comparable to v3
Direct descriptive bridge No causal or superiority claim
Emitted fact coverage
v3 0/158 v4 42/158
Emitted caveat coverage
v3 0/148 v4 60/148
Static claim admission diagram. V3 lets the model own claim bindings and all 18 runs are refused. V4 narrows the model to ordered skeleton IDs and draft text while trusted code owns claim semantics; 6 runs are accepted and 12 are refused. Only emitted fact and caveat coverage are directly comparable.
Figure 1. The v4 skeleton narrows model ownership. It does not weaken the verifier. Use the guide, or focus a contract control and press an arrow key, to trace the ownership boundary. The result panels stay separate because most measures changed meaning with the contract.

The ownership line is the intervention. The model still drafts language. Trusted code owns what each claim means in the evidence system.

The envelope closed. Twelve runs still did not.

V4 completed one attempt with exactly 18 requests. There was no warmup, retry, selective rerun, manual repair, or model download.

Every task matched its skeleton:

matched  270/270
unknown    0/270
duplicate  0/270
omitted    0/270

All 18 outputs had a valid outer envelope. The verifier accepted 6 and refused 12. No run landed in the partially excluded state.

The accepted runs came from three cases, with both repeats accepted:

  • unsupported-join
  • restricted-evidence
  • not-measured-versus-negative

Every other case was refused twice.

The run-level result was:

Measure Observed v4 value Comparison class
Task and skeleton coverage 270/270 Not comparable
Invalid output 0/18 Transformed
Unsupported claims 108/270 Not comparable
Emitted facts 42/158 Direct descriptive
Emitted caveats 60/148 Direct descriptive
Verified-draft facts 84/158 Transformed
Verified-draft caveats 102/148 Transformed
Citation validity 192/192 Not comparable
Replay semantic variance 3/9 Transformed
Accepted 6/18 Transformed
Partially excluded 0/18 Transformed
Refused 12/18 Transformed

Three repeat pairs differed in their verified semantic projection: supported-agreement, cross-layer-disagreement, and prohibited-clinical-causal.

That variance cannot be read as a regression from v3's 0/9. The projection changed when the skeleton contract changed.

The same rule applies to the terminal counts. V3 refused 18/18; v4 refused 12/18. The numbers are true. Calling the difference a six-run improvement would be false because admission operated on a changed contract.

The excluded claims stayed excluded

After the observed tree closed, a bounded post-hoc analysis inspected verifier-owned artifacts. It did not read or publish model-owned draft text.

Across the 12 refused runs, it counted 108 deterministic lexical exclusions:

Exclusion category Count
Diagnostic or prognostic 84
Treatment 15
Causal or certainty 9

Those categories describe which frozen lexical rules fired. They do not classify the underlying draft text as unsafe.

The analysis could not distinguish a prohibited assertion from a bounded negation at the text level because raw model-owned text was outside its evidence set. It therefore made no model-quality or safety conclusion.

This restraint matters. A verifier reason is evidence about verifier behavior. It is not automatic evidence about a model's intent, a sentence's clinical meaning, or the safety of a system outside the measured boundary.

Secondary measurements remain secondary

The v4 runner recorded request latency and token counts for all 18 requests:

Measure Minimum Mean Maximum
Latency 41,521.084166 ms 46,526.308546 ms 50,496.152584 ms
Input tokens 1,417 1,424.333333 1,432
Output tokens 1,260 1,389.111111 1,515

These values describe v4. They are not direct performance comparisons with v3.

Peak memory and peak Metal memory were unavailable because the runner did not report them. They are not zero.

After generation, Ollama was stopped. Offline replay then verified all 18 run artifacts, terminal outcomes, identities, the canonical index, and the closed receipt. Every runtime, tag, manifest, transport, response, and thinking-mode identity check matched during execution.

The model service was no longer needed to decide what the retained evidence said.

Generation and admission are different systems

A model can write a useful sentence and still fail admission. Trusted code can expand a valid skeleton and still produce a refused artifact. A verifier can refuse correctly while exposing a bad upstream interface.

Those outcomes stop looking contradictory when generation and admission have separate jobs.

Generation proposes language. It can select bounded tasks and draft text.

Compilation supplies evidence semantics. It resolves typed units, requirement bindings, citations, canonical caveats, and identities from frozen inputs.

Verification decides admission. It checks the compiled artifact against rules and evidence held outside model control.

Publication emits only admitted prose. It does not leak raw drafts, partial claims, private paths, or repair attempts.

This split applies beyond one evidence atlas.

A RAG synthesis system should not let the model invent the mapping between a sentence and a policy requirement. A research assistant should not turn a plausible citation into an accepted finding without checking source scope. An agent report should not let free-form prose define its own completion criteria. A benchmark narrative should not compare metrics whose measurement contract changed. A compliance evidence system should not let the same model draft, bind, approve, and publish a claim.

The useful abstraction is not "model plus validator." It is a compiler pipeline with an untrusted language-producing stage.

Audit a claim-generating system

Use this checklist before treating generated prose as evidence:

  1. Freeze the claim plan. Name every required fact, caveat, source, and terminal rule before execution.
  2. Give each unit an identity. Derive IDs from canonical trusted content. Do not ask the model to invent them.
  3. Keep bindings outside the draft. Trusted code should own claim types, requirement links, citations, and caveat policy.
  4. Treat model text as untrusted. A valid schema does not make its semantics valid.
  5. Keep the verifier independent. The compiler must not certify its own output.
  6. Preserve invalid attempts. Record bounded error codes, byte counts, hashes, and identities. Do not repair and overwrite history.
  7. Separate refusal from diagnosis. A failed gate can be correct even when post-hoc analysis finds an interface conflict.
  8. Declare comparison classes. Mark metrics as direct, transformed, or not comparable before reading the result.
  9. Keep unavailable distinct from zero. Missing memory, timing, or cost evidence must stay unavailable.
  10. Freeze before execution. Bind code, schema, prompt, model, runtime, fixtures, and stop rules before the first request.
  11. Replay offline. Recompute terminal outcomes from retained artifacts with the model service stopped.
  12. Publish the boundary. State what the evidence cannot support beside what it can.

If one model owns the claim, evidence link, caveat, identity, acceptance rule, and final prose, the system has no independent admission boundary.

Evidence and limitations

The observed v4 result merged in VoxelScope PR #21 at exact commit 5730052. The skeleton implementation merged in PR #19. The prospective declaration merged in PR #20.

The public custody anchors are:

Artifact SHA-256 or Git tree
V4 declaration 063b6055d430f14637d3dab05d1471cfc8e2fcf043b4030c633a956eff94a6c6
V4 benchmark report 86c1c86f62fda175dd121d84785bbb6ed5e11b216cb0383e9b7fcf0290fac409
V4 benchmark index 90f8dcd0002b52e4bdff89d00b0620b730d5ea7d5d5a137cdd3d7c1ff0ecadb3
V4 benchmark receipt 57347be8de8593a86e948a15c981890e747f684ac5a236b5a91305d712cae29d
V4 attempt 577ae65e77a4016361f9a709f21219acdd4e4612911635f338df6c3bf9d5abf4
V4 result summary 78e307ae234be712d097a88f408326839c7b59dbe45f1698f7ee699b00d4c7e6
V4 benchmark Git tree 2322817495bdff1a3df3dff8dd1f7f0487f7e8ad
Audited final study Git tree 328d8785b475dcd9dbf0b7a62dd4553542b335f2

The study covers one local Qwen3 model and manifest, one Ollama version, nine synthetic cases, two repeats per case, one attempt, and one frozen verifier family. It does not establish behavior for another model, prompt, runtime, evidence domain, or production system.

The source evidence contains synthetic, non-clinical material. No patient, MRI, biomedical observation, credential, private path, or raw model draft is published here.

The v3 failure diagnosis and the v4 lexical analysis were post-hoc. They did not alter the frozen studies or terminal outcomes. They classify retained evidence within their stated limits.

The public v3 failure analysis, v4 declaration, result summary, and post-hoc analysis preserve the reviewed boundary.

The verifier still gets the last word

V3 showed that repeatability and valid citations can surround a claim that the evidence system cannot admit.

V4 gave the model a smaller job. Every task reached its skeleton. Every envelope parsed. Trusted code owned the semantics.

Then the verifier refused 12 runs.

That is not a failure of the architecture. It is the architecture working.

The model may draft the sentence. It does not get to decide when the sentence becomes evidence.

Support independent writing

If this post was useful, consider supporting my open source work and independent writing.

Siddhant Khare
Agentic Engineering Guide Resume