probe-factory

White-box runtime safety for self-hosted open-weight models.

A linear probe on a model's own activations catches misuse before the first output token, cheaper than a guard model. Where intent isn't in the words, it wins by a mile.

the evidence Proof & demo 7 harm types, 4 public benchmarks, and a live dual-use slider. explore → the pitch What it is The one-page case: cheaper, earlier, and re-certified as your model changes. read → watch it run Live demo A recording of a real run — the verdict lands before the model finishes typing. play → the product Guardrails console The control surface: one probe per harm type, a slider for how strict each is. open →
Unlisted preview · reproducible for under $1 of GPU.