Polyglot: Multilingual Style Preserving Speech-Driven Facial Animation
Published in IEEE International Conference on Automatic Face and Gesture Recognition (FG)(Oral), 2026
We introduce Polyglot, a single unified diffusion-based architecture for personalized multilingual speech-driven facial animation, using transcript embeddings to encode language-aware information and speaker-style embeddings to capture person-specific habits.
