Each guardrail is a probe on your model's activations. Toggle what to watch; set how strict.
Global policy
How the threshold works. Each probe outputs a score in [0, 1]; the threshold is where a score becomes a violation. Lower = stricter (flags more, including more borderline requests); higher = more permissive. The three white-box edge guardrails read intent from activations a text filter never sees — that is where the probe earns its keep. This is a configuration surface only; nothing here scores live traffic.