Back to feed
News Story
The AI Insider
1 sources

Multiverse Computing Reports All CompactifAI Models Now Run on Intel Xeon 6 Processors

Multiverse Computing announced that its CompactifAI-compressed Llama 3.3 70B model now runs on Intel Xeon 6 processors, achieving significant performance improvements. Benchmarks show up to 94% higher throughput and 48% lower latency compared to the uncompressed baseline, marking a step toward more energy-efficient AI deployment.

SynthePulse Insight · AI deep reading

CompactifAI Compresses Llama 3.3 70B on Intel Xeon 6 to Nearly Double Throughput with Less Than 3% Accuracy Loss

Version 1 · 1 source

Multiverse Computing announces that its CompactifAI-compressed Llama 3.3 70B model can run on Intel Xeon 6 processors, with benchmarks showing throughput improvement of about 94%, latency reduction of nearly 50%, and accuracy drop of no more than 2.5%.

  • CompactifAI-compressed Llama 3.3 70B on Intel Xeon 6 achieves output throughput of 3.86 tokens/s, a 93.6% improvement over the uncompressed baseline.
  • Total token throughput improves by 94.1%, and latency decreases by 48.6% (single-user scenario).
  • Under 256 concurrent users, throughput improves by 107.0% and latency decreases by 51.7%.
  • Model disk size reduces from ~130 GiB to ~65 GiB, a reduction of about 50%.
  • Accuracy drops by 2.48% on MMLU, and by 0.95%-1.94% on BoolQ, GSM8K, HellaSwag, while WinoGrande improves by 6.86%.
  • Supports models including Llama 4 Scout, Mistral Small 3.1, DeepSeek R1, and integrates with frameworks like PyTorch and Hugging Face.
Open section navigationPerformance Leap: Dual Optimization of Throughput and Latency

Performance Leap: Dual Optimization of Throughput and Latency

On July 23, 2026, Multiverse Computing announced that its CompactifAI-compressed Llama 3.3 70B model can now run on Intel Xeon 6 processors (Performance-cores), leveraging vLLM CPU and Intel AMX. In benchmarks on Intel Xeon 6737P processors, for requests with 1024 input + 1024 output tokens, the compressed model achieves an output throughput of 3.86 tokens/s and a total token throughput of 7.81 tokens/s, representing improvements of 93.6% and 94.1% respectively over the uncompressed baseline (2.00 tokens/s and 4.02 tokens/s).

In terms of latency, processing time in a single-user scenario drops from 5056.34 seconds to 2598.22 seconds, a reduction of 48.6%. ITL (inter-token latency) drops from 493.36 ms to 252.08 ms (down 48.9%), TPOT (time per output token) from 494.71 ms to 255.76 ms (down 48.3%), and TTFT (time to first token) from 6939.20 ms to 3705.64 ms (down 46.6%). Under 256 concurrent users, throughput improves by 107.0% and latency decreases by 51.7%.

Accuracy Retention: Over 97% Baseline Performance After Compression

In standard benchmarks, the compressed model shows only minor accuracy changes compared to the baseline: BoolQ drops from 0.904587 to 0.896024 (down 0.95%), GSM8K from 0.934799 to 0.924185 (down 1.14%), HellaSwag from 0.590520 to 0.579068 (down 1.94%), and MMLU from 0.776528 to 0.757300 (down 2.48%). WinoGrande actually improves from 0.678769 to 0.725335 (up 6.86%). Overall, the compressed model retains over 97% of baseline accuracy.

The unexpected improvement on WinoGrande is attributed to the "healing" retraining process after compression. Multiverse Computing performs targeted retraining while compressing parameters; although not all benchmarks are guaranteed to improve, core information retention can lead to performance gains in specific scenarios.

Deployment Advantages: Halved Model Size and Broad Ecosystem Compatibility

After compression, the model disk size drops from approximately 130 GiB to about 65 GiB, a reduction of roughly 50%, significantly lowering storage and deployment time. CompactifAI integrates with open-source frameworks such as PyTorch and Hugging Face, supporting RAG, multimodal reasoning, and enterprise AI workloads. Supported models include Llama 4 Scout, Llama 3.3 70B, Llama 3.1 8B, Mistral Small 3.1, DeepSeek R1, and its "Slim" variants.

CompactifAI is now available on major cloud platforms and Intel Xeon 6 instances, complementing on-premises deployments, providing high-performance and energy-efficient solutions for industries such as finance, healthcare, and manufacturing.

Credibility boundary

This article is based on a press release from Multiverse Computing, republished by The AI Insider. All performance data is self-reported by the company and has not been independently verified by third parties. Benchmark configuration: Intel Xeon 6737P processor, single user, 1024 input + 1024 output tokens. Full methodology is available in the appendix.

Insight takeaway

CompactifAI enables efficient compressed deployment of Llama 3.3 70B on Intel Xeon 6, doubling throughput, halving latency, reducing model size by 50%, and maintaining controllable accuracy loss, providing a viable path for low-cost, scalable AI model deployment on CPUs.

Primary report

The AI Insider

Primary source