Unlike traditional conversational AI setups that rely on a stitched-together cascade of text transcription, large language model generation, and separate voice or video synthesis, Griffin operates as a single, full-duplex video-to-video pipeline.