Google DeepMind has introduced its most capable generative video model, Veo, that generates high-quality videos with a resolution of 1080p. Veo is designed to accurately interpret text prompts and combine them with relevant visual references to create coherent scenes. This advanced understanding of natural language and visual semantics enables Veo to closely follow the prompt and capture intricate details within complex scenes.
Veo’s capabilities extend beyond video generation. It supports film-making controls, allowing creators to apply editing commands to initial videos and create new, edited videos. Masked editing is also supported, enabling changes to specific areas of the video by adding a mask area. Additionally, Veo can generate videos by using an image as input along with the text prompt, conditioning it to follow the image’s style and user prompt’s instructions.
To ensure visual consistency, Veo utilizes cutting-edge latent diffusion transformers that reduce inconsistencies between video frames. This technology keeps characters, objects, and styles in place, providing a more seamless viewing experience.
Google DeepMind’s responsible approach to technology is evident in Veo’s development. Videos created by Veo are watermarked using SynthID, a tool for watermarking and identifying AI-generated content. They are also subjected to safety filters and memorization checking processes to mitigate privacy, copyright, and bias risks.
Veo’s capabilities are not limited to video generation. Google plans to integrate some of Veo’s features into other products, such as YouTube Shorts.
As Veo is launched, Google DeepMind looks forward to collaborating with creators and filmmakers to improve generative video technologies and serve the wider creative community.