A linear probe on a model's own activations catches misuse before the first output token, cheaper than a guard model. Where intent isn't in the words, it wins by a mile.