Research
The team teaching AI to actually resolve.
Heard gets better because a team is dedicated to making it better. Our research pushes on three fronts that matter most in customer service: resolution quality, safety, and Arabic-language performance. Here is how we think about the work.
Our north star
We measure ourselves on outcomes, not eloquence.
It is easy to make a model sound good. Our research asks a harder question: did the customer’s problem actually get solved, safely, in their own language? Every experiment ladders up to moving that number.
What we push on
Three fronts, measured relentlessly
- How often conversations end with a real outcome and no human
- Resolution
- Staying grounded, on-policy, and honest about uncertainty
- Safety
- Native performance across Gulf dialect and Modern Standard Arabic
- Arabic
How we work
The habits behind the progress
Resolution quality
We study where conversations stall and why, then tune the models and the loop to close the gap — measured on real traffic.
Safety research
We work on keeping the agent grounded, on-policy and honest about what it does not know, so autonomy stays trustworthy.
Arabic performance
We invest specifically in Gulf dialect and Modern Standard Arabic — the hardest and most under-served part of the field.
Grounded by evaluation
Nothing ships on vibes. We hold ourselves to evaluations built on the kinds of conversations customers actually have.
Learning in the loop
Corrections from real deployments feed back into how the models reason, so improvement is continuous, not a yearly event.
Sharing what we learn
Insights from building Heard, written up for the practitioners and teams building the future of customer experience.
“The benchmark that matters is whether the customer got their problem solved. Everything else is a proxy.”
Building the future of customer experience with us?
If resolution quality, safety and Arabic AI are your kind of problem, we should talk.