Podcast Episode

Where Voice AI Goes Next: Games, Cars and NPCs | Ramit Pahwa

About this episode

Ramit Pahwa is a Senior Research Engineer at Rivian and Volkswagen Group Technologies, where he works on speech, language and multimodal AI. He led Rivian’s wake word detection programme from research through production, training the model on 20,000 hours of audio. His profile reports a 50 per cent reduction in false negatives and a 90 per cent reduction in false positives. His current work includes speech language models, German speech synthesis and video understanding.Ramit also led the team behind Audio2Tool: Speak, Call, Act, a benchmark of around 30,000 spoken queries for testing how well AI assistants turn speech into actions. The work was accepted at Interspeech 2026. Before Rivian, he worked at Secta Labs and Adobe. He holds a master’s degree in computer science from the University of Texas at Austin and studied mathematics and computer science at IIT Kharagpur.Chapters:00:00:00 Ramit on the evolution of voice AI00:05:13 Building Rivian's voice assistant from scratch00:09:34 Why voice assistants still fail00:15:20 The LLM brain behind voice AI00:19:43 Emotion, memory and the privacy line00:22:53 Personal AI assistants that remember everything00:26:10 Security and the limits of AI agents00:31:24 Voice cloning, watermarks and human verification00:36:51 Designing voices people actually want00:41:09 Voice AI, NPCs and expressive avatars00:45:56 Why game NPCs need small local models00:49:32 Edge AI, GPUs and the shift from cloud00:54:36 Proactive agents, robots and multimodal AI00:57:31 How to start building AI for gameshttps://www.linkedin.com/in/ramitpahwa/