Veo: Google DeepMind's Text-to-Video Model
Veo: Google DeepMind's Text-to-Video Model
Veo (also called Google Veo) is a text-to-video machine learning model from Google DeepMind, first announced in May 2024. As a generative AI system, it produces video from user prompts. The Veo 3 release in May 2025 added the ability to generate accompanying audio.
| Developer | Google DeepMind |
| First release | May 2024 |
| Stable release | Veo 3.1 (15 October 2025) |
| Type | Text-to-video model |
Development
Veo, a multimodal video generation model, was announced at Google I/O 2024 in May of that year. Google said it could generate 1080p clips running over a minute long.
In December 2024, Google shipped Veo 2 via VideoFX, adding 4K video generation and a better grasp of physics. By April 2025, Veo 2 had reached advanced users in the Gemini app.
In May 2025, Veo 3 arrived. Beyond generating video, it creates synchronized audio to match the visuals, including dialogue, sound effects, and ambient noise. Alongside it, Google announced Flow, a video-creation tool powered by Veo and Imagen. DeepMind CEO Demis Hassabis framed the launch as the point at which AI video generation moved past the silent-film era. Flow was later rebranded Google Flow at the 2026 Google I/O keynote, where Google also announced Google Flow Music.
Capabilities
Veo is sold across several subscription tiers and through Google "AI credits." It runs in two front ends: Google Gemini and Google Flow. Gemini, built on the Gemini AI chat model, suits shorter, quicker projects, while Flow functions essentially as a movie editor for longer work, letting users maintain continuity with the same characters and actors across clips. Each individual clip tops out at eight seconds.
Reception has not been uniformly positive. Gizmodo noted that Veo 3 users were steering the model toward low-quality content such as man-on-the-street interviews and product-unboxing haul videos. 404 Media reported that the tool tended to recycle the same joke across different prompts.
Commentators speculated that Google had trained the system on YouTube videos or Reddit posts, though Google itself did not disclose its training sources.
In July 2025, Media Matters for America reported that racist and antisemitic videos made with Veo 3 were being posted to TikTok. Ryan Whitwam of Ars Technica observed that, in an ideal world, Veo 3 would refuse to generate such content, but vague prompts and the model's failure to grasp the subtleties of racist tropes, such as substituting monkeys for humans in some clips, made the guardrails easy to evade.
See also
- Sora (text-to-video model)
- Seedance 2.0
- VideoPoet (Google text-to-video model)
- Dream Machine (text-to-video model)
- LTX (AI model)