Back to feed
News Story
量子位
1 sources

Vivix Releases Real-Time Interactive Multimodal Model A1 for Video Calls with Virtual Characters

Vivix has launched A1, a real-time interactive multimodal model that enables users to have video calls with virtual characters who respond instantly to voice input with actions and expressions. The company also introduced W1, a model that lets users alter storylines in real time. A1 uses a unified streaming architecture and efficient inference to run on consumer GPUs with latency as low as 0.6 seconds.

SynthePulse Insight · AI deep reading

Vivix A1: Is the Phone Call to a 'Parallel World' of Real-Time Interactive Multimodal Models Connected?

Version 1 · 1 source

Vivix releases the world's first native streaming architecture real-time interactive multimodal model A1, claiming nearly 30B activated parameters can run real-time inference on consumer-grade GPUs with latency as low as 0.6 seconds. Technical details and actual capabilities still need verification.

  • Vivix releases A1 model, positioned as the world's first native streaming architecture foundation model unifying multimodal reference, real-time interaction, and streaming generation.
  • A1 has nearly 30B activated parameters, using MJD, native NVFP4 training-inference strategy, and proprietary inference infrastructure, claiming real-time inference on consumer-grade GPUs.
  • Streaming scheduler responds to new input in as little as 300ms, with average latency of 0.6 seconds from user input to first visible frame.
  • A1 directly models speech, audio, gestures, events, and other multimodal signals, rather than relying on an ASR-LLM-TTS pipeline.
  • Simultaneously releases W1 model, allowing users to intervene in real-time and change story direction, but technical details are not elaborated.
  • Official internal testing applications are open, but actual performance and long-term stability await third-party verification.
Open section navigationCore Breakthrough: Unified Streaming Architecture and Native Interaction Modeling

Core Breakthrough: Unified Streaming Architecture and Native Interaction Modeling

Vivix positions A1 as the world's first foundation model that unifies multimodal reference, real-time interaction, and streaming generation within a single native streaming architecture. Traditional clip-based models generate complete segments at once, while streaming generation must maintain consistency of characters, objects, and scenes during continuous generation, like 'laying tracks while the train is moving.' A1 uses causal temporal modeling and streaming state maintenance to keep adjacent actions continuous and maintain character identity, scene, and object relationships over longer time spans.

In interaction modeling, A1 does not use ASR text as the sole intermediate representation; instead, it directly models linguistic semantics, acoustic features, non-verbal sounds, environmental events, gestures, and dynamic visual signals. The model infers the latest response token within a minimum time slice of about 300ms ('fast thinking'), while a higher-level Director Agent handles 'slow thinking'—reasoning, thinking, tool use—and asynchronously provides long-term macro behavior control.

Real-Time Challenge: From 8-Step Distillation to 2-Step Inference

The real-time bottleneck of video diffusion models comes from sampling steps. Vivix proposes MJD (Multi-Dimensional Joint Distillation), using an 8-Step version as the quality baseline to compress capabilities from different denoising stages into a 2-Step model. MJD simultaneously optimizes short-term generation quality, long-term temporal consistency, and multimodal distribution alignment, and allows the student model to learn a more complete action distribution from the teacher model, preventing the model from achieving stability by reducing motion.

According to Vivix's official technical blog, MJD compresses video diffusion from 8-Step to 2-Step inference while maintaining short-term image quality, instruction following, motion performance, and long-term stability close to the 8-Step version. However, this conclusion is based on Vivix's own tests; independent verification has not been published.

Consumer-Grade GPU Deployment: How to Run Nearly 30B Parameters?

A1 has nearly 30B effective activated parameters, using a high-quality compression ratio of 8x8x4, with computational demands far exceeding previous small-parameter or high-compression-ratio real-time models. Vivix introduces VMI (Vivix Model Infra), achieving real-time inference on consumer-grade GPUs through three steps: Encoder-Decoder decoupling, native NVFP4 training-inference synergy, and compute-communication-memory co-scheduling.

Encoder-Decoder decoupling avoids task interference; native NVFP4 training-inference synergy allows the model to adapt to 4-bit numerical distributions during training rather than post-training quantization, claiming near-lossless generation quality; the co-scheduling engine unifies compute, communication, and memory scheduling into the inference execution graph, achieving single-card throughput exceeding 10,000 video tokens/s and PCIe bandwidth utilization of 88%+.

These data all come from Vivix internal tests; actual deployment performance and generality remain to be verified.

W1: Real-Time Director from Character to World

The W1 model, released alongside A1, aims to allow users to no longer be mere observers but to intervene at any time, changing character choices, actions, and story direction. Vivix demonstrated W1's 'real-time director' capability, but technical details were not elaborated in the blog. The relationship between W1 and A1, and whether they share the underlying architecture, remains unclear.

Uncertainty: Is the Phone Call to a Parallel World Clear?

Although Vivix claims A1 achieves 'near-lossless consumer native NVFP4' and 'real-time generation on consumer-grade graphics cards,' all performance data come from Vivix's own tests, lacking independent third-party verification. Whether long-term generation issues like visual drift, error accumulation, and motion degradation are truly resolved still needs practical experience.

Additionally, the specific computational granularity of A1's 'nearly 30B activated parameters' (e.g., per frame or per second) and the exact model and configuration of the consumer-grade GPU have not been officially clarified. W1's actual capabilities and application scenarios also lack detailed technical support.

Credibility boundary

This article's information primarily comes from Vivix's official technical blog and QuantumBit reports. All performance data are Vivix's own test results, not independently verified by third parties. Some technical details (e.g., MJD, VMI) are described based on official sources and may involve selective disclosure. Readers should view claims such as 'world's first' and 'near-lossless' with caution.

Insight takeaway

Vivix A1 demonstrates systematic innovation in the technical path of real-time interactive multimodal models, especially the unified streaming architecture, MJD distillation, and consumer-grade GPU deployment solution. However, the actual performance of nearly 30B parameters on consumer-grade hardware, long-term stability, and the practicality of W1 still require more independent testing and user feedback for verification.

Primary report

量子位

Primary source