October 8, 2026
Agentic AI in Insurance: Mapping the Controls, Workflow by Workflow

Contents
Key Takeaways
- Agentic AI is live in seven insurance workflows. Claims adjudication and fraud triage carries the most exposure; servicing carries the most volume.
- Regulators now test evidence, not intent. 25 states plus D.C. have adopted the NAIC Model Bulletin, and the NAIC’s AI Risk Evaluation Supplement is slated for adoption at the Fall 2026 National Meeting.
- When a system acts instead of recommending, governance moves from the model to the action: permissions, traces, rollback and override.
- Enforce agent permissions outside the agent, reserve adverse consumer actions for people, and monitor outcomes by segment.
- Traces, guardrail decisions and override logs must be captured as the agent acts. They cannot be rebuilt when the exam letter arrives.
Insurance regulators are converging on one question: can the carrier produce dated evidence, for a specific system, of what it was allowed to do, what it did and who checked it?
Agentic systems widen the gap between what a governance policy says and what a system did. This control map takes each workflow in turn and names the decision, data, failure modes, controls and regulatory exposure.
What Changes When a System Moves from Recommending to Acting
Once an agent acts, the unit of governance shifts from the model to the action. A recommender’s errors pass through a person who can catch them. An agent’s errors land directly in the policy administration system, the claims file or the customer’s inbox, and they chain: a misread document produces a wrong classification, which routes a claim to the wrong queue, which triggers the wrong letter.
AM Best drew the same line in a discussion presentation to the NAIC Big Data and AI Working Group in August 2026. Predictive models raise questions of validation, fairness and drift. Agentic systems add authorization, rollback, logging, human override and task boundaries, and they carry the earlier risks forward.
| Recommending system | Agent | |
|---|---|---|
| Unit of risk | One output | A chain of tool calls and writes |
| Human role | Reviews every output | Reviews samples, exceptions, breaches |
| Explainability | Why this score | What it was allowed to do and did, in order |
| Evidence | Validation report, monitoring | Per-action trace, policy in force, override log |
The last row is where most programs are thin. Many carriers can walk an examiner through their AI Systems (AIS) Program. Fewer can produce the trace of what a named agent did on a named claim on a named date, together with the guardrail policy that was active when it acted.
The Regulatory Baseline Every Workflow Inherits
The six sources below shape the insurance AI governance requirements every workflow on the map inherits, so they are set out once here; each workflow table lists only the exposure specific to it. They have one thing in common: none asks whether a carrier meant to govern its AI well. Each asks it to show, for a specific system and on request, what that system did and how it was controlled.
| Source | What it expects | Status |
|---|---|---|
| NAIC Model Bulletin | A written AIS Program covering governance, risk controls, internal audit and vendors, with documentation ready for regulators on request. | 25 states + D.C. as of Aug 31, 2026 |
| AI Risk Evaluation Supplement | A standard question set across four exhibits: AI use, governance, high-risk systems and model data. Pilot states already apply it in market conduct and financial exams. | 12-state pilot; adoption targeted Fall 2026 |
| Unfair trade practices and discrimination laws | Existing prohibitions apply to AI-made decisions. | All states |
| Colorado SB21-169, Reg. 10-1-1 | Versioned inventory, quantitative bias testing, drift monitoring, vendor process. | Auto and health from Oct 15, 2025; evidence on request from July 1, 2026 |
| Market conduct exams | Reconstruct what happened on sampled files. | Ongoing |
| Third-party oversight | Insurer stays responsible; audit and cooperation rights. | Bulletin and Colorado |
Health payers face a hard limit on autonomy. In at least seven states, an AI system cannot be the only basis for denying, reducing or ending coverage of care, so a qualified person has to make that call. The states are Texas, California, Maryland, Nebraska, Indiana, Washington and, since October 1, 2026, Alabama.
For a claims or utilization agent, that sets where the agent must stop and hand the file to a person. Most states do not yet apply a similar rule to property and casualty claims. Florida tried to require human review of every claim denial across all lines, but the bill died in committee in March 2026. Carriers that design human review into denials now will not have to rebuild their agents if similar bills pass.
The Seven Workflows Where Insurers Deploy Agentic AI
The seven workflows below run in the order a policy moves through a carrier, from the first submission to servicing after the sale. Each one gets the same five rows (the decision the agent influences or takes, the data it touches, its failure modes, the controls that contain them, and the regulatory exposure), so you can compare them side by side or skip straight to the workflows you run.
The map shows where each workflow sits by consumer impact and degree of autonomy:

1. Submission Intake and Triage
Intake agents read broker submissions, extract fields, enrich them with third-party data, check appetite and route or decline. The highest-exposure action is the silent decline: a submission scored out of appetite never reaches an underwriter, so no person sees the error.
| Decision | Appetite, routing, priority, sometimes declines |
| Data | Broker emails, applications, SOVs, loss runs, third-party enrichment |
| Failure modes | Extraction errors that propagate; wrong-entity enrichment; instructions hidden in attachments; declines clustering by segment |
| Controls | Confidence thresholds; read-only system access; no decline without a person or versioned rule; input guardrails; decline rates by segment |
| Exposure | Unfair discrimination; bulletin vendor diligence; Supplement data exhibit |
2. Risk Classification and Pricing Support
Agents in pricing support rarely set the rate. They assemble rating inputs, classify the risk, propose a tier or schedule credit, and draft the rationale. The rating plan is filed; the inputs the agent chose to feed it are not, and that is where AI in insurance underwriting creates exposure.
| Decision | Class, territory, tier, schedule credits, which external variables are used |
| Data | Applications, credit-based insurance scores, telematics, imagery, ECDIS |
| Failure modes | Proxy discrimination; departures from the filed plan; drift after vendor data changes |
| Controls | Fairness testing by segment; approved-variable guardrail; reconciliation to the filed plan; revalidation on vendor change |
| Exposure | Colorado Reg. 10-1-1; NYDFS Circular Letter No. 7; rate filing and unfair discrimination laws |
3. Underwriting Referral
Referral is the control most carriers cite when asked about human oversight, and agents test whether it is real. If the agent decides what gets referred, it also decides what does not.
| Decision | Bind within authority, refer, or refer with a recommendation |
| Data | Submission file, guidelines, authority matrices |
| Failure modes | Thresholds that creep; summaries missing the adverse fact; rubber-stamp approvals; authority routed around |
| Controls | Authority enforced outside the agent; referral and approval rates per underwriter; sampling of non-referred binds; reviewers see sources |
| Exposure | Bulletin human oversight; Supplement governance exhibit; file review |
4. First Notice of Loss
FNOL is where most carriers put their first customer-facing agents. The agent captures facts, opens the claim, sets initial coverage flags and routes it. Errors here are small and frequent, and they shape every later decision on the file.
| Decision | Claim creation, coverage flags, severity, assignment, what the claimant hears |
| Data | Transcripts, photos, police and telematics reports, policy and claims history |
| Failure modes | Statements read as coverage positions; mis-captured facts later used to deny; uneven service by language; over-collection of personal data |
| Controls | Approved-language guardrails; transcripts on file; read-back of key facts; data minimization; handoff on injury or request |
| Exposure | Unfair claims settlement; privacy and health data rules; consumer AI notice |
5. Claims Adjudication Support and Fraud Triage
This is the highest-exposure workflow on the map, and it is where the risks of AI in claims handling concentrate. At this stage of AI claims processing, agents summarize files, check coverage, recommend payment or denial, suggest reserves and score fraud. When they act, they can pay, deny, request documents or refer to the special investigations unit.
| Decision | Coverage support, payment, straight-through pay, SIU referral, health downcoding |
| Data | Claim file, policy wording, medical and billing records, estimates, fraud databases |
| Failure modes | Denials on misread facts; fraud flags that delay some groups; leaking payment thresholds; summaries that drop coverage facts |
| Controls | A person decides every adverse action; fraud scores route to investigation only; denial and time-to-pay by segment; full traces; tested kill switch |
| Exposure | Unfair claims settlement; state health AI-denial statutes; Colorado SB21-169 claims management; Supplement high-risk exhibit |
6. Subrogation
Subrogation agents find recovery opportunities in open and closed claims, draft demand letters and manage inter-company arbitration filings. Consumer exposure is lower than in adjudication, but the agent writes to third parties in the carrier’s name, and the insured’s deductible recovery rides on its work.
| Decision | Recovery targets, demand amounts, arbitration, settlement |
| Data | Claim files, liability findings, recovery ledgers |
| Failure modes | Wrong-party or duplicate demands; settlements outside authority; inconsistent deductible reimbursement |
| Controls | Settlement authority in the guardrail layer; templates; ledger duplicate checks; approval above a dollar threshold |
| Exposure | Deductible reimbursement rules; arbitration agreements; outsourced vendor oversight |
7. Policy Servicing and Customer-Facing Assistants
Servicing agents answer questions, process endorsements, change addresses and vehicles, issue ID cards, take payments and handle cancellations. Each action is minor on its own. At volume, a systematic error becomes a market conduct finding.
| Decision | Endorsements, cancellations, payment changes, coverage answers |
| Data | Policy, billing, identity and conversation records |
| Failure modes | Misstated coverage; cancellations without notice; identity checks bypassed; uneven retention offers |
| Controls | Action allowlists; answers grounded in the policy; identity check before any write; conversation sampling; AI disclosure |
| Exposure | Cancellation notice statutes; misrepresentation; consumer AI notice; privacy laws |
How to Govern Autonomous Agents When Human Review of Every Action Stops Being Viable
Once an agent takes thousands of actions a day, reviewing each one is no longer possible. The person moves from checking outputs to setting the boundaries the agent runs inside, testing that those boundaries hold, and proving afterward that they did. Five controls carry that weight.
- Tier by impact and autonomy. Classify every use case in the AI inventory by consumer impact and by whether the system assists, recommends or executes. The tier sets the controls; the technology does not.
- Constrain before the call is made. Enforce permissions, authority limits and prohibited actions in a guardrail layer outside the agent, checked before a tool call executes. A rule written into the agent’s prompt is an instruction the agent may or may not follow, not a control.
- Reserve consequential actions for people. Adverse actions on consumers (declines, denials, reductions, cancellations) go to a qualified person. Routine actions within limits can run straight through.
- Monitor outcomes by segment. Track decline, denial, referral and time-to-pay rates by policyholder segment and by agent version, with thresholds that trigger a documented response.
- Keep the trace. Log every action with its inputs, tool calls, the guardrail policy version in force, and any override or escalation.
These controls make human oversight measurable: sampling rates, override rates, threshold breaches and the response to each. AM Best’s presentation to the NAIC summed up the distinction this way: “Documentary governance answers with policies. Operational governance answers with evidence”.
Evidence You Generate Continuously vs. Evidence You Assemble on Request
AI governance evidence splits into two kinds. Some of it exist only if it was captured when the system acted and cannot be rebuilt later. The rest can be assembled when a regulator asks, provided the underlying records exist.
| Generate continuously (cannot be rebuilt) | Assemble on request (from current records) |
|---|---|
| Per-action traces | Written AIS Program |
| Guardrail decisions with policy version | Inventory and classification snapshot |
| Override and escalation logs | Dated validation and fairness reports |
| Outcome metrics by segment, threshold breaches | Vendor diligence files and contracts |
| Agent, model and prompt version history | Training records and annual reviews |
| Vendor update notices | Supplement exhibit responses |
The left column is where carriers get caught. A team that treats continuous evidence as something to pull together when the exam letter arrives will find the records were never written.
A 90-Day Gap Assessment for Your AI Agents
Start with one agent in claims and one in servicing; they cover the highest exposure and the highest volume. Each check below either produces evidence or names a gap.
- Pull the AI inventory and flag every system that writes to a system of record, sends external communications or triggers a payment. Those are your agents, whatever the vendor calls them.
- Classify each by consumer impact and autonomy, and confirm the classification is recorded with a date.
- For each agent, list the actions it can take and where each permission is enforced. If the answer is the prompt, it is not enforced.
- Pick three files per agent from last quarter and reconstruct what the agent did from logs alone. Time how long it takes.
- Confirm fairness and bias testing exists for every agent touching underwriting, pricing or claims: by segment, dated, and tied to the version in production.
- Review vendor contracts for audit rights, update notification and cooperation with regulatory inquiries. Where a right was refused, record the compensating assurance.
- Find two examples per consequential workflow where a human reviewer changed the agent’s outcome. If none exist, review the review.
- Draft answers to the four Supplement exhibits for your highest-risk agent and mark every answer that relies on evidence you cannot produce today.
- Map each agent to the states where it operates and the bulletin or regulation that applies there, including Colorado Regulation 10-1-1 for life, private passenger auto and health.
Sort every control the checklist turns up into one of three groups: controls that work and leave a record, controls that work but leave no record, and controls that do not exist. Start with the middle group. To an examiner, a control you cannot prove looks the same as one you never had, and fixing it usually means adding logging rather than building something new.
How Lumenova AI Helps Govern Agentic AI in Insurance
Lumenova AI helps carriers, MGAs and payers produce the evidence above as a byproduct of running their agents, rather than reconstructing it when the exam letter arrives.
- AI inventory and use-case classification keep a dated register of every AI system, including vendor-supplied and agentic systems, tiered by risk. That register answers the Supplement’s first exhibit and scopes everything else.
- Pre-deployment AI evaluations produce the testing record Colorado Regulation 10-1-1 and the Supplement ask for, tied to the version that ships.
- Policy controls constrain what an agent may do before the call is made, with authority limits, prohibited actions, and data boundaries enforced outside the agent.
- Continuous observability records each action with the policy in force, creating the evidence trail an examiner samples.
- The Forward Deploy Team works inside your program to prepare for examinations and close the gaps the assessment surfaces, so you don’t have to staff a new governance function from scratch.
If you want to see how this works on your own agents, from inventory to the trace an examiner would sample, book a discovery call to see the platform in action.
Frequently Asked Questions
Agentic AI insurance use cases cluster in seven workflows: submission intake and triage, risk classification and pricing support, underwriting referral, first notice of loss, claims adjudication support and fraud triage, subrogation, and policy servicing. Claims adjudication carries the most consumer exposure; servicing carries the most volume.
The main risks are adverse decisions based on misread or invented facts, fraud flags that delay legitimate claims unevenly across policyholder segments, statements to claimants that read as coverage positions, and payment errors that compound before anyone notices. Each maps to unfair claims settlement practices law, and in health lines to state statutes restricting AI as the sole basis for adverse determinations.
Tier each agent by consumer impact and autonomy, enforce its permissions in a guardrail layer outside the agent, reserve adverse actions for qualified people, monitor outcomes by segment, and log every action with the policy version in force. Oversight then rests on measurable boundaries instead of reviewing every action.
The NAIC Model Bulletin expects a written AIS Program covering governance, risk management and internal controls, internal audit, and third-party systems and data, proportional to the insurer’s use of AI. Regulators may request the inventory, controls, policies, training and vendor diligence. The AI Risk Evaluation Supplement gives examiners a common set of questions for collecting that evidence.
Treat the vendor’s model as your own. Diligence before purchase, contract terms for audit rights, update notification and cooperation with regulators, and revalidation when the vendor changes the system. Both the bulletin and Colorado’s regulation keep the insurer responsible for vendor-supplied data and models.
Yes. Existing unfair discrimination statutes apply whether a person or a system made the decision. Colorado goes further for life, private passenger auto and health, requiring documented quantitative testing of external data and the models that use it.