Langprotect
AI Application & Runtime Security

How Guardia's AI Policy Engine Decides Monitor, Warn, or Block

Mayank Ranjan
Mayank Ranjan
Published on September 11, 2026
How Guardia's AI Policy Engine Decides Monitor, Warn, or Block

A single prompt can trigger more than one scanner at the same time. An employee asks an AI assistant to help draft a message that happens to include a customer's name, references an internal project codename, and "because the employee is frustrated" is phrased in a way that an intent scanner flags as unusually aggressive or policy-adjacent. Three separate detections, from three separate scanners, on one prompt. Only one enforcement decision can actually apply to that prompt, and it has to happen in the time it takes to submit a request, not after a person reviews the conflicting signals.

This is the exact problem severity-based alert triage was built to solve in security operations more broadly, and the parallel is closer than it looks. One academic study of SOC alert volume found that 62% of all alerts generated in a security operations center are ignored outright, not because analysts are careless, but because a system that treats every signal as equally urgent produces so much volume that nothing gets the attention it needs. A policy engine facing multiple simultaneous scanner triggers has the same structural problem at a much smaller scale, and it needs the same kind of answer: a defined way to decide what actually matters most, applied automatically, every time.

This guide covers how that resolution actually works, like the priority tiers, the confidence thresholds, and the specific rule that decides which of several simultaneous detections wins.

Read the Full Platform Overview First

This guide covers one specific decision inside the platform, like how conflicting detections get resolved. The Solution Brief covers how that decision connects to detection, redaction, and audit logging across the rest of LangProtect.

What Is a Policy, and What Does It Actually Control?

A policy is the unit that connects detection to consequence, it's composed of one or more scanners, an enforcement outcome for input-direction detections, and a separate enforcement outcome for output-direction detections. Only one policy is active per tenant at any given time, and it governs enforcement across every deployed client in that organization, whether desktop application or browser extension.

That last detail matters more than it sounds: it means policy is centralized, not something that drifts between individual users or teams unless an organization deliberately configures it that way. When a policy is created, updated, or activated, the change gets written to the backend and the cache holding the active policy is invalidated, like clients poll for the current policy on startup and at a regular interval afterward, so a policy change propagates across the entire organization within roughly a minute of being made, not instantly, but fast enough that a security team responding to an emerging risk doesn't have to wait for individual endpoints to be manually updated one at a time.

How Does the Policy Engine Decide Which Outcome Applies When Multiple Scanners Trigger?

The policy engine resolves simultaneous detections using severity, not order of detection or number of triggers. Every scanner; whether it's built-in or a custom Pattern, Keyword, or Intent rule, declares a priority tier (critical, high, medium, or low) that determines its contribution to an aggregate risk score, and a confidence threshold below which a detection is recorded as a soft detection in the audit log without triggering enforcement at all. When multiple scanners attached to the active policy trigger on the same prompt, the highest-severity triggered result is what actually determines the enforcement outcome, not the first scanner to fire, not the total count of detections, and not an average across everything that triggered.

This design is a smaller-scale version of a pattern that security operations teams converged on for the same underlying reason. One SOC alert-fatigue framework recommends a transparent, additive risk-scoring model, which combines base severity, asset criticality, exploitability, and similar factors into a single score, and then routing based on that score: high scores to senior triage, moderate scores to standard review, and low scores to automatic closure with an audit trail, rather than asking a human to manually weigh every signal against every other one in real time. A policy engine facing multiple scanner triggers on a single prompt is solving the identical problem, just compressed into milliseconds instead of an analyst's shift: too many inputs, one decision needed, and a defined severity hierarchy is what makes that decision consistent instead of arbitrary.

Create_LangProtect_severity_info…_20260911123704

Confidence thresholds add a second dimension to this that's easy to overlook: severity determines which detection wins when multiple scanners fire, but the confidence threshold determines whether a given scanner's result counts as "firing" at all. A high-priority scanner that detects something with low confidence in an ambiguous case doesn't automatically override a lower-priority scanner with a clean, high-confidence match, and it gets logged as a soft detection instead, preserving the signal for later review or threshold tuning without letting an uncertain result drive an enforcement decision on its own. LangProtect's guide to how Guardia detects PII, PHI, and credentials covers where these confidence scores actually come from at the detection layer, this is the layer immediately above it, where those scores get turned into an actual decision.

Why Do Input and Output Detections Get Different Enforcement Outcomes?

Because a prompt reaching the model and a response leaving it represent genuinely different risk profiles, and collapsing them into one enforcement setting forces a compromise that's wrong for at least one direction. A policy configures these separately by design; one enforcement outcome for what happens when a scanner triggers on the way in, and a distinct one for what happens when a scanner triggers on the way out.

The reasoning becomes clear with a concrete case: a credential accidentally typed into a prompt is exactly the kind of thing worth silently stripping or tokenizing before submission, since the employee still gets a working response and never needed the credential to reach the model in the first place. A credential the model itself generates in a response( echoing back an API key it picked up from earlier context, for instance) is a different situation entirely; there's no "task" being preserved by redacting it gracefully, because the response is already complete.

AI_security_policy_enforcement_d…_20260911124528

LangProtect's guide to redact, warn, or block covers the broader decision logic behind choosing an outcome at all, this input/output split is the layer of that decision most organizations configure once and then forget is even a separate setting, until an incident makes the distinction matter.

See the Policy Engine Handle This in Real Time

Guardia evaluates every scanner's priority and confidence against your active policy on every request, not a single blunt setting applied identically to everything.

What Happens at Each of the Five Enforcement Outcomes, Exactly?

LangProtect_enforcement_outcomes…_20260911125642

Each outcome behaves differently once selected, and the specific mechanics matter for understanding what a policy actually does in production, not just what it's named.

Monitor scans and logs the prompt, but takes no action and shows the employee no interruption at all, which is typically used during initial rollout, to establish a baseline before introducing any friction.

Warning shows a popup before the prompt is submitted, displaying the detected entity types and the policy name that triggered it. The employee can cancel or continue; a continued submission is logged with an override flag, preserving the fact that a human consciously chose to proceed despite the warning.

Block stops the prompt entirely, with no option to proceed, and the employee sees a message explaining what was detected, but there's no path forward for that specific submission.

Redact silently strips detected entities from the prompt before it's sent, with no popup shown. Both the original and the stripped version are recorded in the audit log, which matters for later review even though the employee never sees an interruption.

Smart Redact replaces detected entities with typed token placeholders, written to an encrypted session mapping. The model receives only the tokenized prompt; when it responds, the real values are substituted back in before the employee sees the response, so the interaction feels seamless despite the model never having seen the actual sensitive data. LangProtect's guide to sanitization, redaction, and smart redaction covers the deeper mechanics and the research behind why this specific outcome exists as distinct from plain redaction.

Redact and Smart Redact policies also carry an optional setting to log every prompt that passes every scanner cleanly, not just the ones that triggered a violation, which gives compliance teams a complete prompt history rather than only a record of incidents. This substantially increases log volume, which is why it's typically scoped to specific reporting periods rather than left on indefinitely.

What Commonly Goes Wrong in Policy Design?

A handful of avoidable mistakes account for most of the friction and gaps that show up after a policy has been live for a while.

Assuming severity should always mean "block." A critical-priority scanner doesn't have to map to the strictest possible outcome, the priority tier determines which detection wins when multiple scanners fire, not what that winning detection's enforcement outcome should be. Those are two separate configuration decisions, and conflating them tends to produce over-blocking on categories that could have been handled with Smart Redact instead.

Leaving input and output enforcement identical by default. Because these are configured separately, a policy that never revisits the output side after initial setup often ends up either too permissive (missing a leaked secret in a response) or too disruptive (blocking output the way input is blocked, when a response already exists and redaction would serve the task better).

Not distinguishing a soft detection from a non-detection. A confidence threshold that's set too conservatively generates a large volume of soft detections that never trigger enforcement but do accumulate in the audit log; LangProtect's guide to AI audit logs and forensic visibility covers why that accumulated signal is genuinely useful for threshold tuning, but only if someone is actually reviewing it rather than treating "logged but not enforced" as equivalent to nothing having happened.

Treating policy propagation as instant. Since a policy change takes effect across an organization within roughly a minute rather than immediately, a security team responding to an active incident should account for that window rather than assuming a policy update has already taken hold the moment it's saved.

Policy Engine Readiness Checklist

Creating_LangProtect_checklist_t…_20260911130340

Frequently Asked Questions

What happens if two scanners with the same priority tier both trigger?

The highest-severity triggered result determines the enforcement outcome, and priority tier is the primary factor in that severity ranking. When two triggered detections share the same priority tier, the specific tie-breaking behavior depends on how the policy and individual scanner confidence scores are configured; in practice, this is uncommon enough in a well-tuned policy that most organizations never need to resolve it manually, since priority tiers are typically assigned with enough separation between genuinely different risk categories.

Can a policy have different rules for different teams or departments?

Only one policy is active per tenant at a time in Guardia's current architecture, which means enforcement is centralized rather than fragmented by team or department by default. Organizations wanting different postures for different groups should treat that as a distinct governance question from the resolution logic covered in this guide.

What's the difference between a scanner's priority and its confidence score?

Priority is a fixed property of the scanner itself; how severe a category of detection is considered to be, which determines what wins when multiple scanners trigger at once. Confidence is calculated per detection, reflecting how certain the scanner is about that specific match. A high-priority scanner with a low-confidence result on a given prompt may not trigger enforcement at all if the result falls below the configured threshold, regardless of how severe that scanner's category generally is.

Why does a "soft detection" matter if it doesn't trigger enforcement?

Because it's evidence, not nothing. A soft detection preserves the fact that something was detected below the enforcement threshold, which is exactly the data needed to tune thresholds over time and without it, an organization has no way to distinguish "this category never appears in our traffic" from "this category appears constantly but we've set the threshold too high to notice."

How quickly does changing a policy actually take effect across an organization?

Policy changes propagate to every deployed client within roughly a minute, since clients poll for the current active policy on startup and at a regular interval afterward rather than receiving instant push updates. This is fast enough for routine policy management, but worth accounting for specifically during active incident response, where that window matters more than it does for a scheduled policy adjustment.

Should input and output enforcement outcomes ever be set to the same thing?

Sometimes, but it shouldn't be a default left unexamined. For some detection categories like certain PII types, for instance, treating input and output symmetrically makes sense. For others, like credential exposure, the right response genuinely differs by direction: silently stripping a credential from an outbound prompt preserves the task, while a credential appearing in a model's response often needs a harder stop, since there's no equivalent "task" left to preserve at that point.

Is severity-based resolution unique to AI policy engines, or is this a known pattern elsewhere?

It's a well-established pattern in security operations generally, not something specific to AI detection. Alert triage in a SOC context exists for the identical reason: when signal volume exceeds what a single flat response can reasonably handle, defined severity hierarchies are what keep a system's response proportionate and consistent instead of arbitrary or overwhelmed.

The Decision Logic Matters as Much as the Detection Itself

Detecting a sensitive entity or a risky pattern is only half of a policy engine's job. The other half, which decides what to actually do when several of those detections happen on the same prompt at the same time, is where a lot of enterprise AI security implementations quietly fall short, either by defaulting to whichever scanner happens to fire first or by treating every trigger as equally urgent regardless of category.

A policy engine built around severity-based resolution, separate input and output configuration, and confidence-aware thresholds isn't solving a novel problem. It's applying a discipline security teams already trust in a different context (alert triage), to a layer of AI infrastructure that needs the same discipline just as much.

Build a Policy That Resolves Conflicts the Right Way, Every Time

See how Guardia's policy engine turns multiple simultaneous detections into one consistent, severity-based decision, not a guess based on which scanner happened to fire first.

Ready to see it on your own policy?

Tags

policy propagation aggregate risk score confidence threshold scanner priority AI DLP policy alert severity triage Enforcement Decision Logic AI Policy Engine

Related articles