AI for humans

voice
evaluation
KDD
Published

August 10, 2026

Title slide of the KDD 2026 talk 'AI for Humans', subtitle 'Making conversation with AI as natural as talking to a person', with a green audio waveform dissolving into scattered dots

Tomorrow I am giving one of the invited talks in the Applied Data Science track at KDD 2026 in Jeju, Korea. It’s a builder’s tour of some of the audio stack at Boson AI in three parts: how we train models that can hold a conversation, how we design benchmarks that tell us whether the next model is actually better, and how we decide which benchmarks are worth running at all.

There’ll be a show and tell of the running system: voice cloning and audio understanding first, then a live call in Voice Studio. First audio in under a second (750 ms median in production), a web search issued mid-conversation and folded into the next spoken turn, and a language switch mid-interview. TTS covers 100+ languages and STT covers 94, so the switch happens inside one conversation. There is also a demo of the same audio stack driving a live avatar. That combination has not shipped yet; it is coming soon.

The rest of the talk:

Regular readers will recognize the last two parts: ProactBench, IHBench, and benchmark selection each have their own post with the details. And yes, we are hiring, in Santa Clara and Toronto.

Slides: AI for Humans (PDF, 9 MB) · KDD 2026 Applied Data Science track, Jeju, August 11.