Two sliding-window implementations returned the same start coordinate: 0.
They still fed different values to the model.
The synthetic source axis held 155 values. The model window needed 160. Both implementations enumerated the same half-open window:
[0, 160)
MONAI, short for Medical Open Network for AI, is an open-source, PyTorch-based framework for deep learning in healthcare imaging. Its sliding-window inference API placed two zeros before the source and three after it. The older tiler placed all five zeros after the source.
MONAI: 0 0 | source values | 0 0 0
legacy: | source values | 0 0 0 0 0
The coordinates matched. Coverage matched. The patch bytes did not.
That is the sliding-window bug VoxelScope had to refuse before a model saw one tensor.
In this article, ROI means the fixed region-of-interest window size used by the tiler. It describes tensor dimensions, not a clinical finding. VoxelScope stops before the predictor and qualifies the windows themselves.
Evidence boundary: this article describes public synthetic proofs from VoxelScope milestones 6 and 7. It does not publish a private tensor, private geometry, a real window count, model output, segmentation quality, diagnostic performance, GPU speed, or clinical validity. Model loading and inference remain outside this proof.
The lesson applies beyond 3D MRI. Image tiles, video clips, audio chunks, token windows, and blockwise model inputs all depend on the same rule:
A coordinate ledger proves where a window starts. It does not prove which values entered the window.
Matching coordinates are necessary, not sufficient
A sliding-window pipeline has at least two jobs.
First, it enumerates windows. For each spatial axis, it chooses a start and stop. Overlap, stride, and final-anchor rules determine those coordinates.
Second, it materializes values. It decides which source elements fill the window, where padding goes, how axes map into model positions, and how cropped output maps back.
Those jobs are related. They are not the same.
VoxelScope's historical synthetic engine and MONAI 1.4.0 agreed on the positional start sequence for the tested contract. Both used half overlap. Both anchored the final window at the padded extent. Both traversed the last spatial axis fastest.
That agreement was useful. It proved the coordinate schedule.
But the historical engine padded only the high side of an undersized axis. MONAI split the deficit:
before = floor(deficit / 2)
after = deficit - before
For an axis of length 155 and a window of 160, the deficit is 5:
MONAI before = 2
MONAI after = 3
The older engine used:
legacy before = 0
legacy after = 5
The padded extent is 160 in both cases. The only start is still zero. A test that checks window count, start coordinates, final anchors, or complete coverage can pass while every source value sits two positions apart.
Here, [0, 160) is a coordinate in each pipeline's padded working space. It is not a complete mapping back to source positions. That distinction is where the mismatch hides.
That is why a coordinate match cannot establish input equivalence.
Padding changes the meaning of the same slice
Consider a one-dimensional source:
[A B C]
The model expects a window of five values.
Symmetric padding yields:
[0 A B C 0]
High-side padding yields:
[A B C 0 0]
Both windows have shape (5,). Both are finite. Both contain every source value exactly once. Both use start 0 and stop 5.
They are not the same model input.
In three dimensions, that shift can happen independently on each axis. In a four-channel tensor, it affects every modality. The tensor can keep the correct dtype, shape, channel count, and finite-value status while its content moves inside the model window.
This failure survives many common checks:
| Check | Can it pass while padding placement is wrong? |
|---|---|
| Window count | Yes |
| Start and stop coordinates | Yes |
| Full source coverage | Yes |
| Tensor shape | Yes |
| Dtype and finite values | Yes |
| Channel count | Yes |
| Model execution | Yes |
| Exact patch bytes | No |
The last row changes the standard. Do not stop at "the model accepted the tensor." Prove that the intended values occupied the intended positions.
Source note. This figure is generated from closed synthetic article values. The production gate binds the public VoxelScope milestone 6 report, milestone 7 plan, public bundles, exact padding policies, and no-go boundary. It rejects private shapes, private window counts, tensor hashes, geometry, paths, and medical values.
Axis names are contracts, not decoration
Padding was not the only boundary.
Milestone 5 produced a private channel-first tensor in neutral (C,I,J,K) notation. Those letters mean Nibabel voxel-index positions. They do not claim anatomical directions.
The model accepts (N,C,D,H,W).
The reviewed positional bridge binds:
| Preprocessing position | Window position | Model position |
|---|---|---|
I |
S0 |
D |
J |
S1 |
H |
K |
S2 |
W |
The mapping is positional:
I -> D
J -> H
K -> W
It permits no transpose, reorientation, resampling, affine-derived relabeling, or anatomical claim.
That restraint matters. Code often uses Z,Y,X, D,H,W, or I,J,K as if the letters were interchangeable. They are not. One set may describe array positions. Another may describe model positions. A third may imply physical directions.
Renaming an axis does not prove a mapping. Transposing a tensor until the model accepts it does not prove a mapping either.
The contract must state:
- which source position maps to each model position;
- whether any transpose occurs;
- whether orientation changes;
- whether resampling occurs; and
- which geometry identity the tensor keeps.
If one of those facts is implicit, the tensor's shape can be correct while its meaning is not.
Prove the patch two different ways
One implementation can repeat its own mistake.
VoxelScope therefore materializes each synthetic window through two independent paths.
The first path builds the full symmetric padded tensor, then slices the approved window coordinates:
source
-> full symmetric zero pad
-> positional slice
-> patch A
The second path never builds the full padded tensor. It allocates one zeroed window, computes the source intersection, and copies source values into the derived local offset:
source
-> source intersection
-> local destination offset
-> zero-filled patch
-> patch B
The paths share the contract. They do not share materialization logic.
Every emitted float32 C-order window must satisfy:
bytes(patch A) == bytes(patch B)
This comparison catches more than padding disagreement. It can catch:
- a wrong source intersection;
- a wrong local offset;
- an axis swap;
- an off-by-one stop;
- a final-anchor error;
- a channel-order change;
- a dtype conversion; or
- nonzero data written into a padding region.
The model never needs to run for this check to fail.
That is the right boundary. Once model output exists, a shifted input becomes harder to diagnose. The output may still look plausible. A segmentation may still have the expected shape. Aggregate scores may move without revealing why.
Compare the patch bytes before the model can hide the mistake.
Stream evidence, not windows
Exact comparison does not require a directory full of private patches.
The Milestone 7 design streams one pair at a time. It compares the two windows, updates bounded coordinate and byte-stream digests, records coverage, then releases the pair.
It does not persist each window.
That choice reduces the private surface while retaining the facts needed to verify the adapter:
- per-axis starts;
- final anchors;
- padded extent;
- source intersections;
- local padding placement;
- exact byte agreement;
- complete source coverage;
- crop geometry;
- channel order;
- finite values; and
- label exclusion.
The public bundle is smaller still. It contains synthetic statuses and fixture counts. It does not publish a real shape, real window count, affine, orientation, tensor digest, source locator, or private path.
Evidence should be sufficient for the claim and no broader.
Persisting every intermediate array can feel rigorous. It can also create a second dataset, enlarge the privacy boundary, and make review harder. A bounded ledger is stronger when the claim concerns enumeration, equivalence, coverage, and terminal status.
Why basic reconstruction tests can miss this
An identity reconstruction test sounds decisive:
input -> windows -> blend -> crop -> reconstructed input
But it can become circular.
If extraction and reconstruction use the same wrong padding convention, they can agree with each other. The test proves internal consistency, not equivalence to the intended external runtime.
The same risk appears with coverage maps. A map can show that every source voxel received weight while staying silent about where each voxel landed inside a model patch.
Use independent boundaries:
- derive expected starts from a separate oracle;
- derive padding from the pinned external runtime behavior;
- materialize patches through independent algorithms;
- compare exact bytes;
- verify coverage and crop separately; and
- preserve historical fixtures to detect accidental contract drift.
Each check answers one question. None gets to stand in for the rest.
Audit any tiled inference pipeline
Use this checklist for 3D segmentation, large-image tiling, video windows, audio chunks, or long-context blocks.
| Boundary | What to bind | Failure action |
|---|---|---|
| Layout | Source rank, channel order, axis roles, model positions | Refuse implicit transpose or relabeling |
| Window size | Exact per-axis ROI | Refuse reordered components |
| Overlap | Exact rational or pinned calculation | Refuse rounded policy drift |
| Starts | Per-axis enumeration and final-anchor rule | Refuse missing or duplicate windows |
| Padding | Before and after values per axis, fill value | Refuse policy substitution |
| Materialization | Source intersection and local offset | Compare through an independent path |
| Traversal | Exact window order | Refuse ledger/order drift |
| Coverage | Every source position covered | Refuse gaps |
| Crop | Exact return to original positional extent | Refuse shifted output |
| Blend | Mode and weight-map identity | Refuse silent defaults |
| Bytes | Exact patch equality before execution | Stop before model load |
| Publication | Bounded status and non-claims | Keep private observations private |
The order matters.
Check layout before extraction. Check patch bytes before model execution. Check crop before interpreting output. Check publication after a terminal result.
A later gate cannot repair an earlier ambiguity.
What the public proof establishes
Milestone 6 used four non-cubic synthetic proof cases. It bound the positional axis map, ROI (240,240,160), half overlap, stride (120,120,80), final anchors, K-fastest traversal, constant blend mode, symmetric padding, crop, and complete source coverage.
Its public result was precise:
- positional mapping: GO;
- unpadded legacy enumeration: GO;
- general legacy padded reuse: NO-GO;
- private data accessed: false;
- model loaded: false; and
- inference authorized: false.
Milestone 7 added eight synthetic success fixtures and three synthetic refusal fixtures. It checked the two independent materializers, exact window bytes, bounded output, label exclusion, and the same legacy no-go boundary.
Its committed public status is implementation-ready. Real execution is recorded as not-run in the public bundle.
That is enough for this article. The reusable lesson comes from synthetic inputs because the failure concerns array semantics, not anatomy.
What this still does not prove
Exact synthetic patch equivalence does not prove that a model is correct.
It does not prove segmentation quality, source-label semantics, subject independence, clinical value, GPU performance, or memory efficiency.
It does not prove that every sliding-window library uses MONAI's padding, traversal, crop, or blend policy.
It proves a narrower fact: under the pinned positional contract, the reviewed materializers can agree exactly, while general reuse of a different padded extraction policy remains forbidden.
That boundary is useful because the dangerous failure can look healthy:
- valid tensor shape;
- valid dtype;
- finite values;
- complete coverage;
- matching window coordinates; and
- a model that runs.
The patch can still mean something else.
Do not ask only where the window starts.
Ask which values the model received.
Sources: VoxelScope · Milestone 6 review · Milestone 7 review · Exact reviewed commit f5385fd
Public integrity anchors: milestone 6 plan e50352f949a7d87470adf186f729861c770a99e4bc2b4fdb49eae490b101d048 · milestone 6 report 381083f12182689057fac91c804117e6cd23c7dd6ebbcae0a64f5ff9b07d645f · milestone 6 bundle 76895b2af591364156ac021ad482d2ed72fd96ad62f33042a4545238c587caa4 · milestone 7 plan a9a6d123028269e534b48fd1bfe073972e2d4e4c338cbacde98aa6eb8cb35a52 · milestone 7 bundle 3625597508b326497e31688c73ed1deff481fb198a36923f2fb44c921239b2f6