A test user could read the other tenant's record. The application knew which tenant the caller belonged to, but its database lookup ignored that boundary.
I used this defect to walk through Distill before an agent investigates. It builds a context artifact whose inputs, selected chunks, and output can be inspected. A changed source or a tampered bundle produces an explicit failure.
Capability walkthrough, October 4: the HTTP checks, context builds, and integrity replay have run. The compiled packets grew by 490 bytes each. No scored model investigation has run on this fixture; model tokens, billed investigation cost, and efficiency remain unmeasured. The preparation status and supplemental replay record those boundaries.
This constructed tenant fixture is separate from the completed busboy pilot. Results from that pilot do not establish findings, patch quality, or savings for this fixture.
Start with a failure we can see
The case is a constructed HTTP application with two test users, two tenants, and two records containing public markers. Authentication identifies the caller's tenant. The lookup must then restrict which records that caller can read.
The defective lookup accepts a tenant argument but never uses it. The fixed control adds the missing condition:
-- Defective lookup
SELECT value FROM records WHERE id = ?;
-- Fixed control
SELECT value FROM records WHERE id = ? AND tenant = ?;
Both queries use SQL parameters. The failure is broken object-level authorization: a caller can read a record that belongs to another tenant.
I sent seven HTTP requests to each snapshot. The saved observations cover unauthenticated and invalid callers, each user's own record, a missing record, and both attempts to read across the tenant boundary.
The defective snapshot returned the other tenant's marker in both directions. The control denied both reads while preserving access to each user's own record. All fourteen responses matched the expected labels.
Preserving allowed access matters. A patch that denies every request would hide the leak while breaking the application. A regression check rejects that shortcut.
These checks establish the behavior of this fixture. I wrote the expected labels; independent peer review is pending. No agent has detected or fixed the defect in this study.
Lock the inputs, then build and verify
I built the CLI from clean commit 2417e00 using Go 1.26.4. The replay record pins the binary hash and confirms its clean source revision.
The configuration lists auth.py, repository.py, and server.py under sources/. It sets a 4,096-byte chunk limit, a budget large enough to include all three files, and no metadata removal. Copy those public files and the configuration into a fresh working directory, then use the pinned CLI:
distill lock config.json --output context.lock.json
distill build context.lock.json --output distill
fixture_lock_sha256='6cb95c3e205268b6addbe4c56e360b867a97b1dabddab3c2cdcab83f384c5a14'
distill verify distill --expected-lock-sha256 "$fixture_lock_sha256"
lock records the configured inputs and their identities. build rereads them against that lock and publishes four files: the bundle, lock, manifest, and checksums. verify checks their internal relationships and regenerated bundle without reading the source directory or calling a model.
The expected lock digest above comes from the retained fixture lock, obtained before the replay's mutations. In a deployment, obtain that digest through a trusted channel. The pinned Lock specification explains why verification without an external anchor proves internal consistency but cannot authenticate a wholly replaced artifact.
The ordinary baseline bundle is also frozen. Distill's demonstrated contribution is a standardized lock, manifest, and verifier workflow; replayability is available to both packets. This is the artifact contract explored in Context should be a build artifact.
Inspect what reached the bundle
The manifest makes source inclusion and chunk selection explicit. Here are selected entries from the defective-case manifest, with unrelated fields omitted:
{
"lock_sha256": "6cb95c3e205268b6addbe4c56e360b867a97b1dabddab3c2cdcab83f384c5a14",
"sources": [
{
"path": "repository.py",
"original_sha256": "84e969d18a1a139632b04f72dceffaf63d21af258ed5ff9ab6783f075edb1d67",
"included": true,
"reason": "configured_source"
}
],
"chunks": [
{
"path": "repository.py",
"start_byte": 0,
"end_byte": 187,
"selected": true,
"reason": "selected"
}
],
"bundle": {
"path": "context.bundle.md",
"sha256": "2c823187d98e56167582c28f7bb42e27f94b2ae365f326dc82ab4691df155ad7",
"bytes": 2674
}
}
The lookup's 187 normalized bytes are selected. Its full manifest entry also records the normalized source digest and chunk digest. The chunk evidence map connects these ranges to the study's source IDs.
The manifest accounts for all three source files and selected chunks. It reports zero duplicate chunks and zero excluded sources; the packet report confirms zero omitted unique content. An investigator can inspect what entered the bundle before judging the code.
Repeat the build, then break its assumptions
The supplemental transcript records a fresh local replay using copies of the published fixture. It preserves the original evidence and separates this replay from the earlier integration-test log.
First, I built a second output from the same lock:
distill build context.lock.json --output repeat
distill verify repeat --expected-lock-sha256 "$fixture_lock_sha256"
Both commands returned zero. The replay compared the bundle, lock, manifest, and checksum file byte for byte. All four matched each other and the original retained artifacts, including bundle digest 2c823187…df155ad7. This establishes identical outputs for this pinned tool, configuration, runtime, and source set. It says nothing about repeatable model judgment.
Next, I appended # changed after capture and a newline to the copied sources/repository.py, then tried to build from the existing lock. The build returned exit code 1 and created no destination:
Error: locked inputs or identities drifted; run lock explicitly to review changes
The tool did not silently accept the edit or generate a new lock. After restoring the source, the original output still verified. This is a build-time input check: the portable verifier does not monitor later source edits.
Finally, I copied the output to tampered/ and appended tampered and a newline to its bundle. Verification returned exit code 1:
Error: context.manifest.json: bundle identity mismatch
The machine-readable record retains the exact commands, mutations, before/after hashes, exit codes, and output. It also records rejection of a different expected lock digest.
These controls establish artifact integrity relative to the locked inputs. They do not establish that a source is truthful, that an advisory is complete, that the model will find the defect, or that a proposed patch is safe. The artifact includes the defective lookup faithfully; the defect remains an investigation task.
Keep the overhead in view
Both fixture packets passed lock, build, and verify. Both also got larger:
| Snapshot | Without Distill | With Distill | Increase |
|---|---|---|---|
| Defective lookup | 2,184 bytes | 2,674 bytes | 490 bytes / 22.4% |
| Fixed control | 2,206 bytes | 2,696 bytes | 490 bytes / 22.2% |
The defective-case manifest and control manifest show three unique files and zero duplicate chunks in each packet. Exact deduplication had nothing to remove. Distill's metadata added the 490 bytes.
The overhead stays in the result.
- Sources included
- 3 / 3
- Duplicate chunks
- 0
- Unique content omitted
- 0
- Model tokens
- Unmeasured
These sizes measure rendered packets before the common task and tool instructions. Actual provider-token usage is unknown. Lock v0 uses exact normalized chunk deduplication and a byte-based token estimate; it performs no semantic compression or vulnerability judgment.
More repetitions cannot create overlap in these files. The fixture demonstrates the artifact controls and their overhead, not a deduplication efficiency benefit.
A separate follow-up: investigation and efficiency
The unexecuted fixture schedule proposed three paired repetitions per snapshot, twelve runs overall. It could measure investigation variability on one constructed defect and its control. It cannot turn those snapshots into twelve independent cases or give exact deduplication work to do.
A planned variability check.
Stable concatenation
All unique sources includedFresh session. Frozen follow-up reads.
lock → build → verify
All unique sources includedFresh session. Frozen follow-up reads.
Check the finding. Preserve allowed access.
Reads, retries, usage receipts, review.
For the pending fixture variability check, pin each snapshot before comparing its two paths. Follow-up reads use the captured files.
The efficiency comparison belongs to a separate prospective protocol. It needs naturally overlapping advisory and code retrieval, measured and frozen before inspecting model outcomes. The comparison will include ordinary deduplication alongside stable concatenation and Distill, with the same complete unique evidence, tools, work limits, and quality requirements.
The protocol retains the current Sourcegraph capture, the pending Exa and Parallel collection steps, and the boundaries of the separate Next.js collection check. No added duplicates, corpus selection after a win, or discarded unfavorable results.
Agent-trace and actual provider receipts will account for rereads, retries, verification, and reviewer work. The current preparation export contains two lifecycle events and posthoc HTTP observations, not model-call history or a token-burn result. The evidence guide explains its scope and redaction.
A finding must match the evidence. A patch must close the cross-tenant read while preserving allowed access. Model-patch sandbox validation still needs to be added; no efficiency claim can precede a checked outcome. Reliability is not safety reports a different completed study and explains why I keep those judgments separate.
For this walkthrough, the useful result is already inspectable: the packet contains the bad lookup, and changing its locked input makes the build stop.
Inspect the evidence: Public summary, HTTP observations, artifact checksums, and evidence guide.
Replay the integrity controls: Commands and output, structured receipt, and replay scope.
Compare the packets: defective ordinary bundle / Distill bundle; control ordinary bundle / Distill bundle.
Preparation validation: Eight local integration checks passed. The validation record pins the tested source identities, marks synthetic model receipts as test-only, and discloses that the complete private harness is not bundled. These checks validate preparation, not model performance.