Signals helps CX leads run evals without
knowing what an eval is.

Team AI Steering / Project Signals / Lead Product Designer / 2026

Gladly is an AI cx platform used by some of the world’s most loved ecommerce brands.
Its AI Agents handle customer conversations across every channel.

 

Signals is an LLM-as-judge eval system.

A customer describes what they want to measure in plain language, and every conversation is scored against it.

No eval experience required.

Before Signals, customers had to leave Gladly to understand how their AI Agents were performing.

Now they stay in Gladly, and use what they find to decide what to automate next.

Gladly's own teams use Signals too. The first internal signal split sales conversations from support, a number the company had been estimating and needed more certainty on to support a big bet.

Zooming out

This is Signals end to end. A dashboard, two wizards, a criteria library, a detail page and a run page.

I led design on all of it, across two product versions and a mid-project pricing change.

flow diagram for Signals product

Dark solid lines indicate decisions I made. Dashed lines represent features that were cut.

The product

The process

It was already known customers were leaving Gladly to understand how their AI agents were performing.

At the same time, customer-facing evals were a critical product gap that needed to close.

The opportunity was clear: bring that evaluation workflow into Gladly and make it accessible to CX teams.

Cost estimates changed based on number of sessions.

The first iteration was built around a hypothesis: customers would pay to run evals, so they needed visibility and control over their spend.

It was assumed Gladly would bill per evaluation. Everything followed from that assumption.

An estimate show the expected cost before a run.
Scope controls let customers narrow an eval to a subset of conversations instead of running it across everything.

Manual vs LLM generated criteria.

We gave customers control.
They needed confidence.

It was assumed CX leads would write the evaluation criteria themselves.

The first version gave them maximum control: they defined the criteria, label set, and boundaries.

Customers told us it wasn’t difficult to use. The problem was confidence. They weren’t sure they were structuring the eval correctly, or whether the results would actually tell them what they needed to know.

CX managers describe what they want to measure. Gladly generates the criteria and labels, and they edit anything that looks wrong.

The product needed to shift from giving control to creating confidence.

I pushed the team to make LLM generation the default creation path for both criteria and Signals.

Why? When customers are paying for an evaluation, accuracy matters. If customers can’t trust the results, it’s not just a UX problem. It becomes a product and business problem.

I believed Gladly could help customers create more accurate evals by giving them a strong starting point, rather than asking them to start from scratch.

The pivot

Partway through, the revenue model changed.

Signals would be free, and we wanted customers running evals across all of their conversations.

That fundamentally changed how we thought about the design. Cost didn’t disappear. It moved to Gladly, where the person spending it had no reason to think about it.

We wanted customers to use Signals broadly, but we also didn’t want them generating thousands of low-value data points at our expense.

The challenge became: How do we make Signals easy to use at scale while minimizing waste?

This measures the LLM judge against human ground truth, ensuring customers were measuring what they intended.

We added a quality gate before publishing.

To prevent low-confidence Signals from running at scale, I proposed a test step before a Signal could be published.

Customers review 25 labeled conversations and rate each one thumbs up or down. If the Signal reaches 85%+ agreement, it clears the bar for publishing.

If it doesn’t, we show what’s driving the disagreements and what to change, turning the test from a pass/fail check into a way to improve the Signal.

The result

Customers called Signals one of the most innovative products Gladly was building, turning one of our biggest skeptics into a referenceable customer.

It closed a key competitive gap, strengthening Gladly’s position against other AI companies.

And it created a foundation for using Signals data across the platform to surface deeper insights and new opportunities.

 

Signals home

Signals details

Signals run

Important decisions

  • Made LLM generation the default. We made generated criteria and Signals the primary creation path, then removed the hand-authored path once generation was ready.
  • Moved testing into creation as a gate. A Signal had to demonstrate sufficient agreement before it could be published, preventing untested Signals from running at scale.
  • Designed run details around a focused review queue and clustered insights. Instead of presenting hundreds of sessions as a data dump, we turned review into a task with clear areas of disagreement and opportunity.

What I'd do next

  • Ground the judge in Guides and knowledge. Without the procedure or source of truth, a Signal can measure how a customer felt, but not whether the agent followed the right process.
  • Connect Signals to performance reporting. Tag Signals and criteria so results can roll into the dashboards CX leads already use. Today, Signal results live on individual run pages, disconnected from the rest of their reporting.
  • Decide whether Signals should measure human agents, too. Every customer asked for this. The opportunity is clear, but expanding Signals beyond AI agents would require weighing the value against other priorities for the team platform.


#gladlyAI

EMAIL         INSTAGRAM         TWITTER