Gemini 3.8 TTS Playground
· Source: Simon Willison
Google has released two new text‑to‑speech models from the Gemini family, called gemini‑3.8‑flash‑tts and gemini‑3.8‑flash‑lite‑tts. Each model ships with a collection of more than 2,000 pre‑built voices and can create a custom voice from just a 30‑second audio sample, provided the user holds the appropriate usage rights. Simon Willison built a “playground” that serves as a web interface where users can paste their own API keys and experiment with the models using Gemini’s open CORS policy. The tool lets developers design conversations among multiple characters, assigning each a distinct voice and specifying intonation styles, which simplifies the creation of complex dialogues without writing code. In a demo, Willison used the gemini‑3.8‑flash‑tts model to produce over a minute of audio featuring a conversation between two pelicans; the generation took about 20 seconds and cost $0.0274.
This development is significant because it demonstrates how high‑quality audio generation is becoming increasingly accessible to developers and content creators, cutting both the time and cost involved in producing synthetic voices. The ease of creating personalized voices opens new opportunities in accessibility, education, and entertainment applications.
Read the original article on Simon Willison
This summary is an informational synthesis produced by dataqbs.com. All rights to the original content belong to its author and the cited media outlet. We act solely as curators of technology news and claim no authorship.