Q3 ’26 AI Briefing: AI Autonomy Is Now an Authority Question

If the past quarter is any indication, AI is already changing the economics of intelligence.
AI agents have moved beyond helping professionals work faster and into performing more of the work themselves.
For insurers, that creates a significant economic opportunity. More work can move through the organization without a person at every step, but as work moves out of human hands, responsibility for the outcome doesn’t.
That puts a new question in front of insurers: how much authority should AI have as it takes on more work?
Recent developments made that shift concrete. Agent autonomy is extending across more of the workflow. Security incidents are showing what happens when they act in real environments. And models are changing too quickly to build around any single provider.
Here are three AI developments from Q3 that will help insurers determine where AI belongs in the workflow.
Story 1: AI investment doubles down on moving from better answers to autonomous execution
The real economic value of agents comes from handling more of a workflow end to end, rather than completing a single task and waiting for human intervention. They can pursue a goal over time, adapt as circumstances change and work across multiple systems.
Harvey AI offers a clear example in legal, where agents can carry multi-step work over hours and days across tasks like contract analysis and due diligence without a person directing every step. Similarly, Anthropic says Claude now leads 26% of its AI R&D work, up from 1% in March.
Both developments point in the same direction. The unit of work AI can carry is getting larger.
What this means for insurance
For insurers, the opportunity is bigger than deploying an agent for isolated tasks. For example, in underwriting, an agent can gather information, check appetite and duplicates, request missing materials, enrich the risk and update its analysis and draft a quote before a person gets involved. The same is true across all insurance workflows.
Insurers are already redesigning their operating models around a different division of labor. Work that AI can perform reliably can be routed to it, while human expertise is concentrated where judgment, accountability and relationships matter most. But it's not about removing people from the process. The goal is deliberate allocation, with clear decisions about what AI should perform, what it should orchestrate and where human control is required.
When longer-horizon agents carry more of the workflow on their own, more submissions can move through at lower cost and with faster response times. Profitability improves when those operating gains are paired with disciplined risk selection and underwriting.
Story 2: Agent failures raised new control challenges
Autonomous execution is also exposing a new class of failure. OpenAI, Anthropic and Google have all reported cybersecurity tests where models crossed the boundaries of their test environments and reached real systems they weren’t supposed to access.
In May, Gemini accessed three companies after a testing error gave it internet access, using guessed or publicly exposed credentials. Google says the models stopped once they recognized they had accessed a real company system.
In July, OpenAI reported its models broke through offline controls during testing and compromised Hugging Face’s system, where they executed code and accessed limited private data. Separately, it paused training on its most capable models after an agent slipped through internet restrictions to query a public chatbot. OpenAI has since scrapped the planned October release of GPT-6.1 Astra after internal testing found problems with deception and the model acting beyond the permissions it had been given.
Anthropic disclosed similar incidents involving Claude. In one case, a model published malicious code to a public software repository, then used exposed credentials to access a security vendor’s live database.
What this means for insurance
Autonomous agents have become a new risk exposure. Once agents have access to business systems, sensitive data and credentials, a mistake can move beyond a bad answer and trigger actions elsewhere in the workflow. In insurance, the stakes are significant. Those actions can influence pricing, risk selection, underwriting decisions and other consequential outcomes.
That puts authority at the center of AI deployment. Insurers need to define in advance what AI can do, where human approval is required and what happens when it reaches those limits. Consequential actions should be independently verified, with every step traceable and monitored so the system can be stopped or corrected when something goes wrong.
For example, Kalepa’s harness evaluates the task, determines the appropriate way to perform it and sets the boundaries for how that work gets done. When an agent is the right choice, it operates within defined limits, with its actions traceable and reviewable.
Story 3: The Pace of the AI Frontier Is Now in Question
Against that backdrop, calls to slow the pace of frontier AI development became more prominent last quarter. Frontier models and concerns are growing that their capabilities may be advancing faster than the safeguards around them.
Anthropic CEO Dario Amodei argued that AI companies should pace development so risk prevention can keep up, while OpenAI has made similar calls for stronger safety standards, a stance underscored by its decision on GPT-6.1 Astra. Both companies have indicated they may slow parts of their own development as they strengthen safeguards. Microsoft also introduced a code of conduct focused on keeping models subject to human control.
What this means for insurance
For insurers, the debate over whether frontier AI should accelerate or slow down should not determine the pace of adoption. The capabilities available today are already strong enough to perform meaningful work across insurance operations.
The leading model today may not be the leading model next quarter. Building around a single model or provider risks tying critical workflows to a capability that may quickly be surpassed, constrained or unavailable. A more durable architecture can apply the right capability to the right task while keeping workflows, rules, data and controls independent of any one model. The model is a component of the system, not the system itself.
The One Thing to Take Into Your Next Board Meeting
Q3 showed that authority is a critical design question in deployment. The real challenge is deciding how much authority AI should have once it takes on more work. Deploy autonomous tools without those boundaries and the business can move backward. Put the right boundaries in place and the work moves faster without giving up control.
For insurers, that means choosing providers that can define and enforce authority task by task, with rules that hold even as the underlying models change. A robust harness should determine when AI can act independently, when it needs verification and when responsibility returns to a person.
















.png)




.png)