AI Gateway
Keep Agent Access
Within Its Approved Scope
Use the AI Gateway to check requests before they reach models, tools, or other agents – and retain a record of who made each call, even after its key is deleted.
Why Agent Traffic
Needs Its Own Control Point
Shared provider keys work until agents arrive. Then those keys spread into notebooks, CI jobs, and MCP configurations, and no one can say which agent made a call or who approved it. Revoking access means tracking down every copy, and a runaway retry loop only shows up on next month’s invoice.
A router can send that traffic to the right model, but it can’t tell which use case an agent belongs to, what it was approved to reach, or what your risk team will ask for later.
The Lumenova AI Gateway answers those questions before each call goes out, and handles the routing too. Every call is tied to an agent and use case, access ends when you remove someone from a directory group, and budget ceilings stop a retry loop before it reaches the invoice.
Capabilities
One Gateway for Access, Spend, and Attribution
Tie every routed call to an identity and an approved use case, cap what it can spend, and keep the record after the key is gone.
Directory Identity
Sign people and agents in through Entra ID or any OIDC provider, and map groups to use cases so access ends when group membership does.
Use Case Scoping
Limit each key to the models, tools, agents, prompts, and skills its use case allows, starting in flag mode before switching to blocking.
MCP Gateway
Put MCP servers behind one endpoint with tools namespaced per server, and check per-key tool permissions before a request reaches the server.
Budgets and Rate Limits
Set budget ceilings per key, team, and user, and add rate limits where needed. Both are checked on arrival, so retry loops stop at the ceiling.
Unregistered Agent Holds
Hold any agent seen in traffic without a registry entry until it’s registered, and revoke keys that cross your flagged-call threshold.
Durable Attribution
Store the agent and use case on each call record at the moment of the call, so history survives key rotation or deletion.
Inline Guardrails
Run prompt-injection, personal-data, and content checks before and after chat completions. Streaming responses aren’t checked after the call.
Routing and Failover
Route to OpenAI, Anthropic, Azure, Bedrock, and Vertex AI through one OpenAI-compatible endpoint with failover.
How a Request Flows
Your app or agent
Calls with its own key, through one OpenAI-compatible endpoint.
Identify
The gateway resolves the team and the approved use case behind the key.
Check
Scope and budget are checked, and guardrails run on chat completions, before the call goes out.
Model, MCP server, or agent
Receives the calls that pass the checks.
Stopped at the gateway
Calls that fail a check are blocked, or flagged while the use case runs in flag mode. Unregistered agents are held until someone registers them.
Call record
The agent, use case, and outcome are written at call time, so the history survives key rotation and deletion.
Frequently Asked Questions
Policy checks run in under a millisecond, at more than 7,000 requests per second per replica, so enforcement adds almost nothing compared with the model call itself.
Yes. The gateway can sit behind your existing managed edge, and the Azure API Management topology is documented end to end.
The gateway is self-hosted inside your perimeter. Prompts and completions go only to the model providers you choose, and never pass through Lumenova AI.
Yes. Response caching means repeated deterministic requests stop costing money, and multi-key load balancing spreads traffic across keys, deployments, or regions.
The gateway can only scope and check the traffic that passes through it. Agents instrumented with the SDK still get policy checks inside the application. If an agent uses neither, its owner needs to register it. Any unregistered agent that does show up in gateway traffic is held until someone reviews it.