A recording of a real run — Qwen2.5-7B with a trained probe in one process. Watch the verdict land at prefill, tens of milliseconds in, while the full response takes seconds. Same topics, opposite intent; flagged payloads withheld.
Recording of an actual run (steady-state timings; first-call warmup excluded). The probe reads the residual stream during the prefill the server already runs — an output guard could only act after the full response exists. See the benchmarks →