Back to feed
News Story
极客公园
1 sources

Vivix Releases World's First Real-Time Interactive Multimodal Models A1 and W1

On July 24, AI company Vivix launched two real-time interactive multimodal models, A1 and W1, with approximately 30B parameters. They support full-duplex interaction, full-body performance, and object interaction, with latency under 0.6 seconds and significantly reduced inference costs. The models push AI video from static content generation to continuous interactive experiences, applicable in scenarios like virtual tour guides and sales.

SynthePulse Insight · AI deep reading

Vivix Launches Real-Time Interactive Multimodal Models: AI Video Moves from "Generating Content" to "Generating Interaction"

Version 1 · 1 source

On July 24, AI company Vivix released two real-time interactive multimodal models, A1 and W1, with approximately 30B activated parameters, claiming end-to-end first response latency under 1.7 seconds and a two-order-of-magnitude reduction in real-time video inference cost within a year. This marks the transition of real-time interactive multimodal technology from the lab to real-world applications.

  • Vivix released two real-time interactive multimodal models, A1 and W1, with approximately 30B activated parameters, the world's first foundation models integrating multimodal reference, native interaction, and streaming video generation.
  • A1 is designed for interactive AI characters, supporting full-body motion, object interaction, spatial movement, real-time full-duplex interaction, and dynamic image input; characters can listen, speak, and act simultaneously.
  • W1 extends interaction to entire scenes and stories, supporting native joint audio-video generation, multimodal reference, and native multi-shot storytelling, allowing users to change the story direction in real time.
  • The technology foundation includes Vivix-Turbo (ultra-low-step distillation) and Vivix Infra (real-time inference system), enabling streaming generation and low-latency response.
  • Official reports state visible reaction latency under 0.6 seconds and end-to-end TTFR under 1.7 seconds; real-time video inference cost has decreased by two orders of magnitude within a year, approaching the CDN bandwidth cost level of mainstream video platforms.
  • Vivix was founded in early 2025 and has completed Series A funding, with historical investors including IDG, Sequoia, BlueRun, 5Y, Monolith, and strategic investors JD.com and Alibaba.
Open section navigationA1: From a "Talking Face" to an "Acting Character"

A1: From a "Talking Face" to an "Acting Character"

Vivix defines A1 as the first real-time full-duplex model for interactive AI characters. Its core capabilities include: characters can walk, pick up objects, and perform full-body actions; support real-time full-duplex interaction, allowing users to interrupt at any time; natural facial expressions and emotions change in real time with the conversation; new images can be dynamically added during interaction.

These capabilities are achieved through a native multimodal generation loop: a Director Agent acts as a real-time director, determining character behavior and maintaining context; video is generated in a streaming manner, generating short segments each time, maintaining consistency through local continuity and global sequence priors; multi-dimensional joint distillation compresses the inference process to two steps, reducing computational load.

W1: Changing the Entire Story in Real Time

W1 extends interaction from a single character to people, objects, scenes, camera angles, sound, and events. It supports native joint audio-video generation (synchronized generation of visuals and sound), multimodal reference (text, images, audio, and video collectively influence generation), and native multi-shot storytelling (automatic switching of perspectives and camera angles).

W1 also adopts streaming generation; user input such as voice, images, touch, or camera operations can change the subsequent story in real time. Vivix-Turbo reduces inference steps through ultra-low-step distillation, combined with camera planning and temporal continuity optimization to maintain consistency of characters, scenes, and camera angles.

Technology Foundation: Vivix-Turbo and Vivix Infra

Vivix-Turbo improves generation speed and stability, while Vivix Infra enables input processing, audio-video generation, decoding, and streaming to run simultaneously, reducing response time and lowering long-term operational costs. When users change their mind mid-stream, the system directly adjusts subsequent generation, preserving already generated content and context without recalculating from scratch.

Official reports indicate visible reaction latency under 0.6 seconds; after accounting for network transmission and live streaming, end-to-end TTFR is under 1.7 seconds. Real-time video inference cost has decreased by two orders of magnitude over the past year, approaching the hourly CDN bandwidth cost level of mainstream video platforms.

Vivix's Vision and Funding Background

Vivix was founded in early 2025 and has completed Series A funding, with historical investors including IDG, Sequoia, BlueRun, 5Y, Monolith, and strategic investors JD.com and Alibaba. The company targets the gap in real-time interactive experiences, aiming to bridge the gap between real-time human presence and pre-recorded content using real-time interactive models.

Vivix believes interaction must bring significant experience gains, such as stronger immersion and more realistic feedback. A1 handles "people," W1 adds "scenes" and "time," and Infra ensures changes are presented in a timely manner while controlling costs. AI-generated content is no longer just waiting to be watched but can continue to evolve after users enter.

Credibility boundary

This article's information primarily comes from GeekPark's report on Vivix's product launch, where performance data (latency, cost) is from official reports and falls under source_claim. Vivix's funding information comes from the report without independent verification. Model capability descriptions are based on official demos and statements; actual performance awaits third-party evaluation.

Insight takeaway

Vivix's A1 and W1 models mark significant progress in real-time interactive multimodal technology, but performance data still requires independent verification. Their core innovation lies in transforming AI video from one-way generation to two-way interaction, potentially reshaping content creation and user experience.

Primary report

极客公园

Primary source