On July 16, 2026, Google Vids announced two major updates, centered on introducing the Gemini Omni model. This model enables users to generate high-quality video clips using natural language text prompts and image references (such as photos or sketches). Unlike earlier versions, Omni can blend multiple input modalities, making outputs more aligned with user intent.
More critically, Omni supports step-by-step editing: after generating an initial draft, users can ask Vids in everyday language to change the background, adjust lighting, or add effects, without starting from scratch. This capability also applies to users' own footage, marking a shift from parameter-based editing to conversational interaction.
Previously, Google opened the Veo 3.1 model to all users in February 2026, driving the adoption of video generation within Vids. Official data shows that over the past year, users have created millions of videos in Vids.