Google has augmented its live dialogue model Gemini 3.8 Live with a feature called "Live Avatar" that's capable of conjuring animated characters and realistic human simulacra to lip-sync Gemini's machine-generated speech. "Conversation is inherently multimodal: we listen, look, speak, and use facial expressions to communicate," explained Shuo-yiin Chang, research scientist at Google DeepMind, and CJ Zheng, Gemini software engineer, in a blog post. "Live Avatar brings these capabilities to enterprise agents. By processing visual and audio inputs simultaneously, it generates enriching conversations for a more comprehensive experience." Five years ago, DeepMind warned about various risks associated with the use of large language models, among...
Læs hele artiklen hos kilden.
Kommentarer (0)
Ingen kommentarer ennå. Bli den første til å kommentere!