99How do audio and speech LLMs work, and how do discrete audio tokens differ from text tokens?▼hardOpenAIGoogle DeepMindMeta1 replies◆ premiumSpeech LLMs either transcribe to text or model audio directly as discrete tokens. The signal is the semantic-vs-acoustic token split and why end-to-end audio models beat ASR-plus-LLM pipelines on latency and prosody.Open full answer →
28How does speech-to-text (Whisper) work, and what matters when building voice AI (STT + TTS)?▼mediumOpenAIGoogleMicrosoft1 replies◆ premiumVoice is a major modality and a common applied-AI surface. The signal is the audio-to-text pipeline, why Whisper is resilient, and the cumulative latency budget that makes or breaks a real-time voice agent.Open full answer →