AI Guardrails
Stop Unsafe Content
Before It Reaches Your Models
Stop injection, leaks, and harmful content in both directions, with thresholds and responses your team controls.
AI Guardrails That Catch What Filters Miss
When each team writes its own content filters, security ends up auditing a dozen implementations instead of setting one standard. The gaps are rarely in the user’s prompt. Prompt injection hides in documents, web pages, and tool results, and personal data leaks through answers as easily as through inputs.
Your security team defines the rules once, and Lumenova AI Guardrails apply them to both sides of each model call, with thresholds tuned per use case. They check what’s said, not what’s done.
Actions like using a tool, moving a payment, or changing a customer record are governed by the Policy Engine, which can run a guardrail mid-decision, for example to confirm an export contains no personal data before it’s allowed.
CORE CHECKS
Six Checks for Everything Your Models Read and Write
Turn on the checks each use case needs, set the threshold for each, and choose what happens when one fires: block, mask, redact, or flag.
Prompt Injection
Detect instructions hidden in user input, documents, web pages, and tool results before the model acts on them.
Personal Data
Find personal data in prompts and answers, then block, mask, redact, or tag it. Add custom patterns for identifiers like account or record numbers.
Harmful Content
Screen inputs and outputs against a taxonomy of harmful-content categories, with thresholds set per use case.
Topic Controls
Allow only the subjects a use case covers, or restrict the ones it must avoid, so agents stay on task and out of areas they shouldn’t advise on.
Grounding
Score answers against the sources you provide, and flag claims those sources don’t support before they reach a customer or a decision.
Word Lists and Patterns
Block terms, phrases, and regular expressions tied to legal, brand, or internal rules, and sanitize input before it reaches the model.
How a Check Runs
Input arrives
A prompt, a document, or a tool result, on its way to the model.
Check the input
Prompt injection, personal data, harmful content, and topic checks run with your thresholds.
Model
Receives the input that passed, with any masking or redaction already applied.
Check the answer
The answer is checked for personal data, harmful content, and claims your sources don't support.
Answer delivered
Your user, or the agent's next step, gets the answer that passed.
Blocked, masked, redacted, or flagged
Each check takes the action you set for it.
Guardrail decisions
Each decision is recorded and exports as CSV or JSON, filtered by application and time range.
Frequently Asked Questions
Yes. Prompt injection detection checks documents, web pages, and tool results as well as user input, so instructions hidden in a claim file, application, or email are caught before the model acts on them.
Guardrail decisions are recorded and can be exported as CSV or JSON alongside policy decisions and findings, so reviewers and auditors can see what was caught, when, and under which use case.
Start in flag mode on real traffic and review what each check would catch before switching to block.
Yes. Thresholds and actions are set per use case, so a customer-facing claims agent can run stricter checks than an internal research assistant, while both follow standards security defines centrally.
Yes. Topic controls let a use case cover only the subjects it’s approved for, or restrict areas it must avoid, such as financial, medical, or coverage advice. Word lists and patterns can also block specific phrases your compliance team has flagged.
Grounding scores each answer against the sources you provide, such as policy documents or product terms, and flags claims those sources don’t support before they reach a customer.