# Prospective security-investigation efficiency study

Status, October 4, 2026: proposed follow-up, not an executed efficiency result. This document records the revised design separately from the capability walkthrough and the immutable preparation summary. No scored model runs or paid Exa/Parallel study calls have occurred in this fixture study. Model choice, a numeric spend cap, independent label review, and model-patch sandbox validation remain pending.

## Keep the fixture question separate

The [original preparation summary](./public-summary.json) proposes three paired repetitions per defective/fixed snapshot, twelve agent runs overall. Those runs could characterize investigation variability on one constructed defect and its control. Three unique files with zero duplicate chunks give exact deduplication no overlap to remove. Repeating the investigation cannot change that property or supply twelve independent security cases.

This fixture, the supplemental integrity replay, the separate Next.js collection check, and the completed busboy pilot are distinct bodies of evidence. Do not pool their outcomes or transfer a finding or savings claim from one to another.

## Capture natural overlap before scoring

Select the retrieval cases and collection procedure before inspecting model outcomes. Capture normal advisory/code retrieval without adding repeated documents to favor Distill. Measure and record source overlap and exact normalized chunk overlap, plus their byte counts, before any scored model investigation. Near-duplicate or semantically related text is not an exact duplicate.

Freeze the corpus, retrieval receipts, normalizer, chunking rules, evidence IDs, inclusion requirements, comparison implementations, and scoring rules before the scored runs. Keep zero-overlap cases and unfavorable results. Do not choose the corpus or change the baseline after seeing a win. Report results per independent case; repetitions within a case are not independent security cases.

Current collection work:

| Tool | Role | Work completed |
| --- | --- | --- |
| [Sourcegraph](https://sourcegraph.com/docs/api/stream-api) | Find code at a pinned revision | One public search captured |
| [Exa](https://exa.ai/docs/reference/search) | Find primary advisories and affected-version constraints | Adapter prepared; study capture pending |
| [Parallel](https://docs.parallel.ai/search/search-quickstart) | Find deployment context and counterevidence | Adapter prepared; study capture pending |

The separate Next.js collection check used revision [4698ad6](https://github.com/vercel/next.js/commit/4698ad6478cc85a7283a8c41edfbba023dadf57d) and the known advisory [CVE-2025-29927 / GHSA-f82v-jwr5-mffw](https://github.com/vercel/next.js/security/advisories/GHSA-f82v-jwr5-mffw). The [query](./sourcegraph/request.json), [receipt](./sourcegraph/capture.json), and [event stream](./sourcegraph/response.raw) record seven line matches across two files in 4.778 seconds, with no skipped results reported.

Fourteen pinned source/license files were then fetched through GitHub. The [source index](./nextjs-source-index.json) retains URLs and hashes; full bodies are not in this export. This capture establishes neither repository-wide coverage nor a runnable Next.js reproduction or historical disclosure replay. It is not yet a scored efficiency corpus.

## Compare against a competent baseline

Use three inputs from the same captured corpus:

1. Stable concatenation of normalized sources with evidence IDs and locations.
2. Ordinary exact deduplication with deterministic ordering and references to every source location represented by a duplicate.
3. Pinned Distill compilation with its lock, bundle, manifest, and verification.

The ordinary-deduplication implementation and chunk boundaries must be specified and validated before scoring. It must handle natural exact overlap competently rather than preserve redundant content to inflate the comparator. Distill must earn any incremental advantage over that control.

All three inputs must retain the same complete unique evidence and provenance coverage. Choose a budget that fits that evidence in every arm; do not count required omissions as savings. Follow-up reads must expose equivalent captured evidence. The model, task, tools, allowed work, and quality requirements stay fixed, with fresh sessions and a frozen run schedule.

Both ordinary inputs are frozen artifacts too. Distill's standardized integrity workflow is already demonstrated; exclusive replayability is not an efficiency hypothesis.

## Require a checked outcome

Live retrieval stays outside scoring. Expected labels and verification tests stay hidden from the investigator. For a known-advisory reproduction, the [fix and its tests](https://github.com/vercel/next.js/commit/52a078da3884efe6501613c7834a3d02a91676d2) stay outside the initial packet. Later public advisory material does not establish what was knowable at disclosure time.

A finding must cite evidence and match the observed behavior. The fixed control must not receive the same leak claim without evidence. A proposed patch must remove the cross-tenant read while preserving legitimate access. Sandbox execution and independent label review must be completed before treating the outcome as validated. An unavailable dependency or sandbox leaves the outcome inconclusive; a cheap run that misses the defect or denies all access fails the task.

## Account for the whole loop

Use [agent-trace](https://github.com/Siddhant-K-code/agent-trace) to retain source rereads, tool turns, retries, repairs, and verification events. Join actual provider receipts to each checked outcome. Report input/output usage, supplied cache and reasoning breakdowns, elapsed time, retrieval cost, verification work, and reviewer effort. Failed calls and retries belong in the ledger.

The pinned [agent-trace cost display](https://github.com/Siddhant-K-code/agent-trace/blob/b109ec5b3714b842746e97ee8e975329d8582667/src/agent_trace/cost.py) estimates usage from event text and bundled prices. Keep those estimates separate from receipt-based usage and billing. Missing usage or price remains unknown; cache and reasoning subsets must not be added twice within provider totals.

The current [preparation export](./agent-trace/events.ndjson) has two lifecycle events and posthoc HTTP observations. The native reader and timeline loaded it at [b109ec5](https://github.com/Siddhant-K-code/agent-trace/commit/b109ec5b3714b842746e97ee8e975329d8582667). It contains no scored model-call history or token-burn result. The [evidence guide](./README.md) documents its scope and redaction.

Actual study expenditure counts shared retrieval once. A deployment estimate charges each hypothetical arm for obtaining its evidence. Report the two ledgers separately; frozen retrieval does not mean free retrieval.

Report all prespecified cases and runs, including losses, no-overlap cases, failures, and inconclusive outcomes. Artifact integrity, packet bytes, provider tokens, total cost, finding quality, and patch safety are separate results.
