The Agent changed 3 times.
Inspector needed to change with it.

Designing the observability experience for
an AI Agent that keeps evolving.

Team AI Steering / Project AI Session Review / Lead Product Designer / 2026

 

The Agent changed.
Inspector needed to change with it.

Designing the observability experience for
an AI Agent that keeps evolving.

Team AI Steering / Project AI Session Review / Lead Product Designer / 2026

 

The Agent changed.
Inspector needed to change with it.

Designing the observability experience for an AI Agent that keeps evolving.

Team AI Steering / Project AI Session Review / Lead Product Designer / 2026

 

The Agent changed.
Inspector needed to change with it.

Designing the observability experience for an AI Agent that keeps evolving.

Team AI Steering
Project AI Session Review
Lead Product Designer / 2026

 

Conversation and trace

AI Session Review

To improve an Agent, you have to understand its conversations. When an Agent does something unexpected, AI Session Review is where teams go to understand what happened and why. Ground truth.

It brings the customer conversation and Agent trace together in one place — helping teams understand behavior and improve their Agent.

Designing for a product that keeps changing.

Inspect, our session reviewer, has evolved alongside the Agent itself.

V1 simplified the complexity.

V2 exposed more of it.

Just when we thought we had Inspect in a good spot, the underlying architecture changed again. Now we're on to V3.

If there's anything I've learned from designing AI products, it's that the technology, our understanding of it, and what users need change fast.

Self-service requires understanding.

A core tenant Gladly believes, is that customers should be able to manage their Agents themselves.

When something unexpected happens, Session Reviewer is where they go to answer:

Why did my Agent do that?
What did it know?
What should I change?

My hypothesis If we want customers to manage and improve their own Agents, Inspect has to do more than expose information. It has to help them build a mental model of how their Agent works.

And if it looks so technical their eyes glaze over, we've failed.

V1 — We started by hiding the complexity.

The first instinct was to use narrative as a design element. Make AI behavior understandable without requiring people to understand the technical system underneath it.

The conversation stayed primary. Expand a turn and we'd explain what happened in a more human, narrative way.

It was deliberately simple.

Review details

Then our experts asked for the complexity back.

As our internal AI teams used Session Reviewer to investigate Agent behavior, the abstraction started getting in their way.

They didn't need us to explain what happened. They needed to see what actually happened.

Raw trace — Algorithm Component

Example of a technical trace.

Example of the same technical trace in Gladly. It's a prototype. Try it!

Stay in the conversation. Inspect the trace without losing the customer interaction that prompted the investigation.

V2 — Bring the trace to the conversation.

Our AI experts needed the actual trace: what ran, what went in, what came out, and how those pieces led to the response.

They could get that information from developer tools like LangSmith, but that meant leaving the customer conversation and interpreting the raw trace somewhere else.

We brought that depth into Session Reviewer.

The goal wasn't to simplify the system anymore. It was to make the complexity navigable. 

See the algorithms that ran and how they contributed to the response. Get into the technical detail when you need it.

Just when we thought we had Inspect in a good spot...

The next ask sounded straightforward: make Inspect easier to discover and sessions faster to review.

I prototyped a few options to share ahead of a conversation with our AI research lead. That conversation surfaced something much bigger: we were trying to make the V2 Inspector fit a fundamentally different Agent architecture.

Inspect still needed to answer the same question:

Why did my Agent do that?

But the information someone needed to answer it had changed.

The difference between version 2 and version 3

The difference between how the algorithm works in v2 compared to v3.

Same question. Different model.

V2 showed the sequence of algorithms that ran: what went in, what came out, and how each step contributed to the response.

The new Agent worked differently. It accumulated context as the conversation progressed: instructions, tools, results, and memory that shaped what it knew at each turn.

To understand

Why did my Agent do that?

we first needed to answer:

What did the Agent know when it responded?

Open full screen ↗

Try it! The trace can get more sophisticated without asking the customer to become more technical.

Represent what the Agent knows.

The new model changed what Inspect needed to show.

Instead of a sequence of algorithms, we needed to represent the context the Agent had available at each turn — and how that context accumulated over the conversation.

That led to a simple distinction:

New this turn / Carried over

Get deep. Expect change.

Working on AI products requires getting deep enough into the system to represent it clearly for someone else.

But the system you're designing for today may not be the system you're designing for next month.

V1 wasn't wrong. Neither was V2.
Each reflected how the Agent worked, and what customers needed to understand about it, at that moment.

Get deep enough into the system to represent it clearly, while embracing the reality that today's solution is going to change — maybe sooner than you think.

#gladlyAI

EMAIL         INSTAGRAM         TWITTER