Back to Insights

Seeing the Stop Sign Through the Trees: How AI Evals are Redefining Insurance Guardrails

I recently had the pleasure of sitting down with Patti Harman at the DigIn podcast. Toward the end of our conversation, she asked me a simple but massive question: What are you most excited about right now? My answer didn

Seeing the Stop Sign Through the Trees: How AI Evals are Redefining Insurance Guardrails

I recently had the pleasure of sitting down with Patti Harman at the DigIn podcast. Toward the end of our conversation, she asked me a simple but massive question: What are you most excited about right now? My answer didn't revolve around a specific new foundation model or a flashy new coding tool. Instead, I told her I’m most excited about defining the future of AI regulation and compliance in our industry. To safely deploy generative AI in highly regulated financial spaces, we must build rigorous evaluation platforms, or "evals," that act like autonomous vehicle sensors, identifying regulatory "stop signs" to protect consumers and empower human advisors.

To explain why this is so critical, we have to look outside of insurance and look at the streets for autonomous driving.

The Waymo Edge Case: Seeing Through the Branches

When Waymo first began rolling out its autonomous vehicles, they encountered a fascinating, real-world edge case: overgrown trees blocking city stop signs.

For a human driver, a stop sign entirely obscured by spring foliage is a significant, sometimes fatal, hazard. We rely almost entirely on line-of-sight vision. If we cannot visually see the red octagon, we don't know we are supposed to stop. We are at risk, and we don't even know it.

But an autonomous vehicle like Waymo experiences the world differently. It doesn't just rely on a single visual plane. It has a multi-modal understanding of its environment.

First, it has deep contextual data. Its internal maps know there is supposed to be a stop sign at that specific intersection, regardless of what the camera sees. Second, it utilizes advanced technology like LiDAR, which can penetrate the visual noise of the leaves to detect the reflective metal of the sign hidden behind the branches.

Not only does the Waymo vehicle successfully stop, keeping the passenger safe, but it also identifies and logs the obscured sign as an edge case. By navigating its own guardrails, the AI effectively maps out hidden dangers, creating a safer environment that even human drivers can eventually benefit from.

At The Zebra, we are taking this exact same conceptual framework and applying it to conversational AI in the insurance industry.

The Hidden Stop Signs of Regulated Industries

If you ask a general-purpose Large Language Model (LLM) what kind of insurance you need for a new Honda Civic parked in an Austin, Texas driveway, it will confidently generate a list of coverage recommendations. It might tell you to increase your liability limits, drop your collision coverage, or add comprehensive protection based on local weather patterns.

Gemini 3.1 Pro

For a search engine or an open-domain chatbot, delivering that conversational answer is simply a helpful user feature. But for a licensed insurance marketplace, that exact same answer potentially creates a regulatory concern.

The insurance industry is governed by incredibly strict state and federal frameworks. There is a hard, legal line between consumer education and licensed, binding advice. Every state has different minimums, different regulations, and different legal definitions of what constitutes a recommendation.

These regulations are our industry's "stop signs."

When an AI agent interacts with a consumer on our platform, it is driving through a dense neighborhood of complex financial terminology and legal liability. If the AI agent tells a user, "You should definitely drop your comprehensive coverage to save money," it has effectively acted as an unlicensed advisor. It has run a regulatory stop sign, and the liability is monumental.

The core challenge for any insurtech company building in the agentic era is this: How do you leverage the incredible reasoning and conversational capabilities of an LLM to help a customer deeply understand complex insurance policies, without ever making a direct, binding recommendation?

Building the Sensor Suite: The Power of AI Evals

You cannot just add a system prompt to a foundation model that says, "Be compliant, act professionally, and don't give legal advice," and blindly hope for the best. That is the equivalent of telling a teenage driver to "be careful" and then hand them the keys to a sports car. You need a deterministically safe environment for a probabilistic model.

To achieve this, The Zebra is actively working with specialized evaluation vendors and partners to build the insurance industry’s equivalent of a self-driving sensor suite. We are building robust agent harnesses and rigorous evaluation platforms — commonly referred to as "evals."

Stages of Model Development

Think about how automotive engineers train that self-driving car. Before a vehicle is ever allowed on a public road in downtown traffic, it runs millions of simulated miles in a digital twin environment. It is tested against weird intersections, unpredictable pedestrian behavior, snowstorms, and faded lane markers. Only when the system passes these mathematically sound evaluations is it deployed.

We are applying this continuous evaluation abstraction to our conversational guardrails. Before any agent interacts with a consumer, it must survive a rigorous simulated gauntlet.

We feed our evaluation platforms thousands of data points from our product, edge-case scenarios, and strict compliance mandates. The platform procedurally generates thousands of unique customer interactions — from simple questions to highly convoluted, multi-part inquiries filled with slang and misunderstandings.

Agent Harness + Eval Interaction

The AI is continuously graded against strict criteria:

  • Factuality: Is the information technically correct?
  • Compliance: Did the model successfully avoid making a regulated recommendation or binding guarantee?
  • Accuracy: Did it hallucinate a non-existent carrier feature?

If the model fails — if it runs a stop sign — we use that failure as feedback to fine-tune its behavior. We run thousands upon thousands of these evaluation loops before the model ever interacts with a real customer.

Seeing Through the Trees: Contextual Data Advantage

Just like the Waymo vehicle uses map data to know where a stop sign should be, our agent harness gives the AI deep contextual awareness.

Because the harness wraps securely around the foundation model, it has safe access to our first-party database. It understands regional purchasing trends, the historical context of the user, and the specific questions needed to generate an accurate quote.

With this rich data, the agent can hold a highly personalized conversation. It can demystify complex coverages, explaining the difference between actual cash value and replacement cost without crossing the line into advice. It can prepare the consumer, gathering the right structured data, so that our licensed human advisors can step in and execute the final policy.

We are doing this in careful, heavily gated stages. We must continuously validate that the agent remains consistently aligned with our rigorously compliance standards in live, unpredictable consumer environments. Simultaneously, we must ensure the AI is actually helpful, reducing friction rather than acting as a glorified, overly cautious FAQ bot.

Making the Road Safer: Empowering the Human Advisor

The prevailing narrative around generative AI often centers on automation, cost-cutting, and workforce replacement. But in highly regulated, high-stakes industries, the true value proposition of AI is enablement.

We are not building to replace our licensed insurance advisors. We are building them to supercharge them.

When a consumer has to slog through a complex, confusing web of insurance terminology on their own, the human advisor spends the majority of their time on the phone doing basic education and data entry. They are explaining what a deductible is for the thousandth time, rather than providing bespoke, high-level advice.

When an autonomous AI agent is safely guided by rigorous evals, it can handle the heavy lifting of intake and demystification. Just as the autonomous vehicle flags hidden dangers for the city, our AI evals map out the complex knowledge gaps and regulatory edge cases for our organization.

By the time the customer is handed off to a licensed professional, the structured data is already in place. The customer actually understands the product they are buying, the friction of the process has evaporated, and the advisor can focus entirely on what they do best: applying their licensed expertise, reviewing the personalized options, and making a trusted, compliant recommendation.

Defining the New Parameters of AI

As we move further into the agentic era, the ability to simply build and deploy autonomous software will no longer be the primary differentiator for tech companies. Everyone will have access to powerful foundation models.

The true competitive advantage will belong to the organizations that can prove their autonomous systems are deterministically safe, strictly compliant, and continuously evaluated.

We are treating conversational AI with the same rigorous safety standards that engineers apply to self-driving cars. We are teaching our agents to see the stop signs through the trees.

###

Article link copied to clipboard!