The first study repeated every model outcome byte for byte.
Every citation was valid.
The verifier refused all 18 runs.
That result came from one local Qwen3 study in VoxelScope. It ran 9 synthetic cases twice with no rerun. The repeat pairs matched at the model-outcome level. Citation validity was 342/342. Replay semantic variance was 0/9.
None of those properties made one run admissible.
Four outputs had the wrong outer JSON type. The other 14 parsed. Across those parsed drafts, the verifier excluded 56 of 238 proposed claims. No run emitted accepted prose.
Evidence boundary: this article describes two completed local studies against synthetic, non-clinical evidence. It reports one immutable v3 study and one prospectively declared v4 study. Only emitted fact and caveat coverage support a direct descriptive comparison. The evidence does not support a causal improvement, safety improvement, model-quality improvement, generalization, significance, superiority, clinical utility, production readiness, or biomedical-validity claim.
Repeatability tells you that a process can reproduce an outcome. Citation validity tells you that a claim points to an allowed source. Structured output tells you that a parser can read the envelope.
None answers the harder question:
Does this claim satisfy the exact evidence contract?
Correct citations can support an inadmissible claim
A citation can be real, in scope, and attached to the wrong unit of meaning.
Suppose an evidence plan requires one fact and one caveat:
fact: the two evidence layers agree
caveat: agreement does not establish truth or causality
The model can cite the right source for both sentences. It can repeat the same pair on every run. But a verifier may require one combined claim bound to both requirement IDs.
If the model returns two claims, each claim has a valid citation. Neither claim satisfies the combined binding.
That is what happened in v3. The model-facing plan listed fact and caveat requirements separately. The verifier-facing contract expected combined units for some pairs. The prompt allowed a single claim to bind more than one requirement, but the plan shape taught the model to produce separate claims.
The model followed one interface. The verifier enforced another.
The post-hoc failure analysis counted 28 requirement_binding_mismatch exclusions across the 14 parsed envelopes. Each parsed run hit two such exclusions.
This was not evidence that every refused draft hallucinated. It was a compiler and verifier interface conflict.
A safe negation can still trip a lexical gate
The second conflict was smaller and easier to miss.
The evidence plan required bounded caveats such as:
Agreement or disagreement does not establish truth, causality, clinical utility, or therapeutic relevance.
and:
Source-reported target evidence is not a VoxelScope target, ranking, or druggability claim.
Those are limiting statements. They block stronger interpretations.
The verifier used substring prohibitions. It searched for fragments such as caus and druggab without interpreting negation. The safe caveats therefore triggered the same rules meant to stop unsupported causal or druggability claims.
Across the parsed v3 envelopes, the failure analysis recorded 14 causal-or-certainty substring exclusions and 14 target-or-druggability substring exclusions.
The rule saw the word. It did not see what the sentence denied.
This distinction matters. A refusal can be correct at the system boundary while its reason still points to a bad interface. The verifier was right to fail closed under its frozen rules. The study was also right to classify the conflict instead of calling every refusal unsafe model behavior.
Stable failure is still failure
The repeat pairs make the result harder to dismiss.
All nine pairs were byte-identical at the model-outcome level. Four invalid outputs repeated the same output hash and byte count within their cases. The other pairs repeated the same committed draft identity.
That stability ruled out one easy story: the result did not arise because the model wandered between unrelated answers on each repeat.
But stable failure does not become success. The v3 study still recorded:
| Measure | Observed value |
|---|---|
| Requests | 18 |
| Refused | 18/18 |
| Invalid outer JSON types | 4/18 |
| Parsed envelopes | 14/18 |
| Citation validity | 342/342 |
| Replay semantic variance | 0/9 |
| Unsupported claims | 56/238 |
| Emitted facts | 0/158 |
| Emitted caveats | 0/148 |
| Verified-draft facts | 102/158 |
| Verified-draft caveats | 80/148 |
The gap between verified-draft coverage and emitted coverage is the admission boundary. The verifier could retain supported claim material inside a refused run. It still could not emit final prose because required bindings or hard rules failed.
A draft can contain supported parts and remain inadmissible as a whole.
The comparison has three classes
VoxelScope changed the contract before the v4 study. That change fixed some structural problems and changed what several metrics meant.
The comparison must therefore keep three classes separate.
| Class | Metrics | Allowed reading |
|---|---|---|
| Direct descriptive comparison | Emitted fact coverage, emitted caveat coverage | Report v3 and v4 side by side without a causal claim. |
| Transformed | Invalid output, semantic variance, verified-draft coverage, terminal outcomes, latency, input tokens, output tokens | The prompt, schema, generated contract, or semantic projection changed. Do not call movement an improvement measure. |
| Not comparable | Task and skeleton metrics, unsupported-claim rate, citation validity, memory fields | The denominator, ownership, field definition, or availability differs. |
Only two measures cross the direct bridge:
| Direct descriptive measure | v3 | v4 |
|---|---|---|
| Emitted fact coverage | 0/158 |
42/158 |
| Emitted caveat coverage | 0/148 |
60/148 |
Those values describe what each frozen pipeline emitted. They do not establish why the values differ.
No causal or superiority claim follows from that bridge.
The contract changed. The next study did not rerun v3 with a patched verifier. It tested a new model-facing interface while keeping verifier authority outside the model.
Freeze the skeleton before execution
V4 moved the claim plan out of the model response.
Before execution, trusted code grouped each admissible claim unit into a typed skeleton. Each skeleton carried a content-derived ID plus its claim type, requirement bindings, source bindings, and canonical caveat policy.
The model received those IDs in a fixed order. Its job narrowed to two outputs:
- return the ordered skeleton IDs; and
- draft bounded text for each requested unit.
The model did not own claim types, requirement bindings, source citations, canonical caveats, content identities, or final admitted prose.
Trusted expansion restored those fields from the frozen skeleton. The verifier then checked the expanded artifact under the same authority it held before. The compiler could build a claim. It could not admit one.
The prospective declaration froze that interface, the same nine conceptual cases, two repeats per case, the local runtime, model tag, full manifest, endpoints, comparison classes, stop rules, and one-attempt boundary before the first v4 generation request.
Its SHA-256 was:
063b6055d430f14637d3dab05d1471cfc8e2fcf043b4030c633a956eff94a6c6
The run used Ollama 0.35.1, tag qwen3:8b-q8_0, and full manifest:
e56358ca25dd14db6853a9f68a92d717aaa6f0a94250a72d1a0f3d86a9f30130
Claim admission compiler
Who owns the meaning?
Step 1 of 4
A citation sits beside the claim
V3 repeated every model outcome and validated every citation. Those checks did not prove that each claim matched the verifier contract.V3 contract
Claim semantics cross the model boundary.
Fact and caveat IDs arrive as separate plan entries.
The model chooses claim types, bindings, citations, and draft text.
Separate fact and caveat claims follow the visible plan shape.
The compiler cannot invent the combined verifier unit.
Binding mismatch and naive substring rules fail closed.
Repeatable and correctly cited does not satisfy the contract.
Observed result layer
Keep the studies separate
- Terminal outcomes
- 18/18 refused
- Citation validity
- 342/342
- Semantic variance
- 0/9
- Terminal outcomes
- 6 accepted · 12 refused
- Unsupported claims
- 108/270
- Peak memory
- Unavailable, not zero
v3 0/158 v4 42/158v3 0/148 v4 60/148The ownership line is the intervention. The model still drafts language. Trusted code owns what each claim means in the evidence system.
The envelope closed. Twelve runs still did not.
V4 completed one attempt with exactly 18 requests. There was no warmup, retry, selective rerun, manual repair, or model download.
Every task matched its skeleton:
matched 270/270
unknown 0/270
duplicate 0/270
omitted 0/270
All 18 outputs had a valid outer envelope. The verifier accepted 6 and refused 12. No run landed in the partially excluded state.
The accepted runs came from three cases, with both repeats accepted:
unsupported-joinrestricted-evidencenot-measured-versus-negative
Every other case was refused twice.
The run-level result was:
| Measure | Observed v4 value | Comparison class |
|---|---|---|
| Task and skeleton coverage | 270/270 |
Not comparable |
| Invalid output | 0/18 |
Transformed |
| Unsupported claims | 108/270 |
Not comparable |
| Emitted facts | 42/158 |
Direct descriptive |
| Emitted caveats | 60/148 |
Direct descriptive |
| Verified-draft facts | 84/158 |
Transformed |
| Verified-draft caveats | 102/148 |
Transformed |
| Citation validity | 192/192 |
Not comparable |
| Replay semantic variance | 3/9 |
Transformed |
| Accepted | 6/18 |
Transformed |
| Partially excluded | 0/18 |
Transformed |
| Refused | 12/18 |
Transformed |
Three repeat pairs differed in their verified semantic projection: supported-agreement, cross-layer-disagreement, and prohibited-clinical-causal.
That variance cannot be read as a regression from v3's 0/9. The projection changed when the skeleton contract changed.
The same rule applies to the terminal counts. V3 refused 18/18; v4 refused 12/18. The numbers are true. Calling the difference a six-run improvement would be false because admission operated on a changed contract.
The excluded claims stayed excluded
After the observed tree closed, a bounded post-hoc analysis inspected verifier-owned artifacts. It did not read or publish model-owned draft text.
Across the 12 refused runs, it counted 108 deterministic lexical exclusions:
| Exclusion category | Count |
|---|---|
| Diagnostic or prognostic | 84 |
| Treatment | 15 |
| Causal or certainty | 9 |
Those categories describe which frozen lexical rules fired. They do not classify the underlying draft text as unsafe.
The analysis could not distinguish a prohibited assertion from a bounded negation at the text level because raw model-owned text was outside its evidence set. It therefore made no model-quality or safety conclusion.
This restraint matters. A verifier reason is evidence about verifier behavior. It is not automatic evidence about a model's intent, a sentence's clinical meaning, or the safety of a system outside the measured boundary.
Secondary measurements remain secondary
The v4 runner recorded request latency and token counts for all 18 requests:
| Measure | Minimum | Mean | Maximum |
|---|---|---|---|
| Latency | 41,521.084166 ms |
46,526.308546 ms |
50,496.152584 ms |
| Input tokens | 1,417 |
1,424.333333 |
1,432 |
| Output tokens | 1,260 |
1,389.111111 |
1,515 |
These values describe v4. They are not direct performance comparisons with v3.
Peak memory and peak Metal memory were unavailable because the runner did not report them. They are not zero.
After generation, Ollama was stopped. Offline replay then verified all 18 run artifacts, terminal outcomes, identities, the canonical index, and the closed receipt. Every runtime, tag, manifest, transport, response, and thinking-mode identity check matched during execution.
The model service was no longer needed to decide what the retained evidence said.
Generation and admission are different systems
A model can write a useful sentence and still fail admission. Trusted code can expand a valid skeleton and still produce a refused artifact. A verifier can refuse correctly while exposing a bad upstream interface.
Those outcomes stop looking contradictory when generation and admission have separate jobs.
Generation proposes language. It can select bounded tasks and draft text.
Compilation supplies evidence semantics. It resolves typed units, requirement bindings, citations, canonical caveats, and identities from frozen inputs.
Verification decides admission. It checks the compiled artifact against rules and evidence held outside model control.
Publication emits only admitted prose. It does not leak raw drafts, partial claims, private paths, or repair attempts.
This split applies beyond one evidence atlas.
A RAG synthesis system should not let the model invent the mapping between a sentence and a policy requirement. A research assistant should not turn a plausible citation into an accepted finding without checking source scope. An agent report should not let free-form prose define its own completion criteria. A benchmark narrative should not compare metrics whose measurement contract changed. A compliance evidence system should not let the same model draft, bind, approve, and publish a claim.
The useful abstraction is not "model plus validator." It is a compiler pipeline with an untrusted language-producing stage.
Audit a claim-generating system
Use this checklist before treating generated prose as evidence:
- Freeze the claim plan. Name every required fact, caveat, source, and terminal rule before execution.
- Give each unit an identity. Derive IDs from canonical trusted content. Do not ask the model to invent them.
- Keep bindings outside the draft. Trusted code should own claim types, requirement links, citations, and caveat policy.
- Treat model text as untrusted. A valid schema does not make its semantics valid.
- Keep the verifier independent. The compiler must not certify its own output.
- Preserve invalid attempts. Record bounded error codes, byte counts, hashes, and identities. Do not repair and overwrite history.
- Separate refusal from diagnosis. A failed gate can be correct even when post-hoc analysis finds an interface conflict.
- Declare comparison classes. Mark metrics as direct, transformed, or not comparable before reading the result.
- Keep unavailable distinct from zero. Missing memory, timing, or cost evidence must stay unavailable.
- Freeze before execution. Bind code, schema, prompt, model, runtime, fixtures, and stop rules before the first request.
- Replay offline. Recompute terminal outcomes from retained artifacts with the model service stopped.
- Publish the boundary. State what the evidence cannot support beside what it can.
If one model owns the claim, evidence link, caveat, identity, acceptance rule, and final prose, the system has no independent admission boundary.
Evidence and limitations
The observed v4 result merged in VoxelScope PR #21 at exact commit 5730052. The skeleton implementation merged in PR #19. The prospective declaration merged in PR #20.
The public custody anchors are:
| Artifact | SHA-256 or Git tree |
|---|---|
| V4 declaration | 063b6055d430f14637d3dab05d1471cfc8e2fcf043b4030c633a956eff94a6c6 |
| V4 benchmark report | 86c1c86f62fda175dd121d84785bbb6ed5e11b216cb0383e9b7fcf0290fac409 |
| V4 benchmark index | 90f8dcd0002b52e4bdff89d00b0620b730d5ea7d5d5a137cdd3d7c1ff0ecadb3 |
| V4 benchmark receipt | 57347be8de8593a86e948a15c981890e747f684ac5a236b5a91305d712cae29d |
| V4 attempt | 577ae65e77a4016361f9a709f21219acdd4e4612911635f338df6c3bf9d5abf4 |
| V4 result summary | 78e307ae234be712d097a88f408326839c7b59dbe45f1698f7ee699b00d4c7e6 |
| V4 benchmark Git tree | 2322817495bdff1a3df3dff8dd1f7f0487f7e8ad |
| Audited final study Git tree | 328d8785b475dcd9dbf0b7a62dd4553542b335f2 |
The study covers one local Qwen3 model and manifest, one Ollama version, nine synthetic cases, two repeats per case, one attempt, and one frozen verifier family. It does not establish behavior for another model, prompt, runtime, evidence domain, or production system.
The source evidence contains synthetic, non-clinical material. No patient, MRI, biomedical observation, credential, private path, or raw model draft is published here.
The v3 failure diagnosis and the v4 lexical analysis were post-hoc. They did not alter the frozen studies or terminal outcomes. They classify retained evidence within their stated limits.
The public v3 failure analysis, v4 declaration, result summary, and post-hoc analysis preserve the reviewed boundary.
The verifier still gets the last word
V3 showed that repeatability and valid citations can surround a claim that the evidence system cannot admit.
V4 gave the model a smaller job. Every task reached its skeleton. Every envelope parsed. Trusted code owned the semantics.
Then the verifier refused 12 runs.
That is not a failure of the architecture. It is the architecture working.
The model may draft the sentence. It does not get to decide when the sentence becomes evidence.