For AI Platform and Engineering Teams
Ship Agents with Controls
Your Risk Team Can Verify
Stop waiting on reviews. Define agent limits as code your team tests and versions, and let risk and compliance pull their own evidence.
What Slows Platform Teams Down
Reviews need evidence, controls need to be visible, and new tooling needs to fit the stack your team already runs.
When any one is missing, platform and engineering teams spend more time explaining the release than shipping it.
Reviews Built on Screenshots
Every new agent goes through review. Without test results to show, the reviewer asks for another meeting and the release waits.
Limits Buried in Code
The rule that stops an agent is a threshold in someone's service. Risk can't read it, and no one can say which version was running last week.
Tools That Want Your Stack
Governance products often ask you to replace your tracing or switch frameworks. Your operations team already runs Datadog or Splunk and isn't moving.
AI Governance in Practice
Build Agent Controls into Your Workflow
Instrument agents, route calls through the gateway, test policies and evaluations in CI, and give risk teams direct access to decision records.
Drop-In Adoption
Add controls in a few lines of code, without replacing your tracing or framework. Traces keep flowing to Datadog or Splunk.
One Endpoint for Models & MCP
Route model and MCP traffic through the AI Gateway with a base URL change, and keys, budgets, and permissions follow each use case.
Policy as Code
Write limits on tool calls, budgets, and handoffs as policy, then lint, test, and promote it through GitLab merge requests.
CI Evaluation Gates
Start evaluations from CI with the SDK and fail the build on the scores you choose. You add the step to your own CI system.
Self-Serve Records for Risk
Let reviewers filter and export policy decisions, with the policy, reason, and input, as CSV or JSON, without filing a ticket with your team.
Failure Behavior You Set
Choose how limits behave when a check can’t run. Platform policy fails closed, and the SDK fails open by default but can be set to fail closed.
How It Starts
- Your team
Pick one agent
One that runs in production or is close to it.
- Your team
Connect it
Through the SDK, your collector, or the gateway.
- Together
Agree the criteria
Success criteria and who signs off, written down before work starts.
- Together
Run the proof of value
A paid proof of value on your infrastructure. Your team keeps the policies, decision records, and evaluation results.
Frequently Asked Questions
No. Any agent that already emits OpenTelemetry traces can send them through your Collector with no SDK, whatever language it’s written in. The four-line setup applies to the Python SDK.
Tracing is built on OpenTelemetry, so any framework that emits OpenTelemetry traces works. LangChain and LangGraph agents use a callback handler at each call site, and call-order rules currently apply to those two frameworks.
Yes. Policies can be tested against sample input before they bind, and experiments compare prompts, models, and pipelines on your own datasets before anything ships. Evaluations run from CI, so a regression fails the build rather than reaching production.
Start with the least disruptive settings. New use cases can run in flag mode at the gateway, logging what they would block before anything is stopped. The SDK fails open by default, so a policy check that can’t run won’t take down your agent. Policies can also be tested against sample input before they bind.
Policies combine into stacks, and each binding is pinned to a version, so one team changing a rule doesn’t change another team’s use case. Every version records who changed it, when, and why.
The AI Gateway routes to OpenAI, Anthropic, Azure OpenAI, AWS Bedrock, and Google Vertex AI through one OpenAI-compatible endpoint.