Multiverse Computing reports a 94.1% increase in total token throughput achieved by its CompactifAI-compressed Llama 3.3 70B model running on Intel Xeon 6 processors. This advancement utilizes vLLM CPU and Intel Advanced Matrix Extensions to deliver more energy-efficient and scalable artificial intelligence. Benchmarks using an Intel Xeon 6737P processor show the compressed model processed a request with 1,024 input and 1,024 output tokens in 2,598.22 seconds, a 48.6% reduction in latency compared to the uncompressed baseline’s 5,056.34 seconds. According to the company, this breakthrough delivers improved throughput and latency while maintaining compatibility with mainstream open-source tooling.
CompactifAI Compression Enables Llama 3.3 70B on Intel Xeon 6
Multiverse Computing has demonstrated a substantial leap in large language model performance; its CompactifAI-compressed version of the Llama 3.3 70B model now operates on Intel Xeon 6 processors. This achievement expands the possibilities for deploying powerful AI on more accessible hardware, promising gains in both energy efficiency and scalability. The company’s proprietary compression technology reduces model size while maintaining a high degree of accuracy, a critical balance for real-world applications. Benchmarking on an Intel Xeon 6737P processor revealed significant improvements over the uncompressed baseline. Specifically, the CompactifAI model achieved a total token throughput of 7.81 tokens per second, a 94.1% increase compared to the baseline’s token throughput of 4.02 tokens per second.
Latency also saw marked reductions; for a request comprising 1,024 input and 1,024 output tokens with one concurrent user, processing took 2,598.22 seconds with the compressed model, a 48.6% reduction from the baseline’s processing time of 5,056.34 seconds. Further analysis showed reductions in Inter-Token Latency (ITL) of 48.9% and Time Per Output Token (TPOT) of 48.3%, as measured during benchmarking on the Intel Xeon 6737P processor. The compressed model’s disk size is approximately 50% smaller, decreasing from around 130 GiB to 65 GiB, which lowers storage demands and speeds up deployment. While some accuracy variations were observed across standard benchmarks, such as a 2.48% decrease on MMLU, the model retained over 97% of the baseline’s performance, as measured against standard benchmarks. Enrique Lizaso, Co-founder & CEO of Multiverse Computing, says, “This isn’t just a technical milestone — it’s a change for AI builders.” CompactifAI integrates with PyTorch software and the Hugging Face platform, streamlining its incorporation into existing AI development pipelines.
This isn’t just a technical milestone – it’s a game-changer for AI builders.
Enrique Lizaso, Co-founder & CEO of Multiverse Computing
Detailed analysis of key metrics further illustrates these gains; mean Inter-Token Latency decreased from 493.36 milliseconds to 252.08 milliseconds (48.9% reduction) when benchmarking on the Intel Xeon 6737P processor, while mean Time Per Output Token fell from 494.71 milliseconds to 255.76 milliseconds (48.3% improvement). The largest gains were observed at high concurrency, with throughput increasing by 107.0% and latency decreasing by 51.7% at 256 concurrent users when benchmarking on the Intel Xeon 6737P processor. These performance boosts do not come at the expense of accuracy. A 6.86% increase in accuracy was observed on the WinoGrande benchmark, attributed to a “healing” process of targeted re-training following compression.
Accuracy Retention Across Standard Benchmarks with Compression
Multiverse Computing’s focus on maintaining model accuracy during compression is yielding quantifiable results, demonstrated by recent benchmarks of its CompactifAI technology. While significant gains in throughput and latency have been reported when running the Llama 3.3 70B model on Intel Xeon 6 processors, the company also assessed the impact of compression on performance across several standard evaluation datasets. These tests reveal a nuanced picture, indicating that substantial efficiency improvements do not necessarily come at the cost of substantial accuracy loss. Specifically, the CompactifAI-compressed model exhibited minor variations in scores compared to the uncompressed baseline, as demonstrated by the benchmarking results. BoolQ testing showed a 0.95% decrease, while GSM8K saw a 1.14% reduction. HellaSwag and MMLU benchmarks registered decreases of 1.94% and 2.48% respectively. Interestingly, the WinoGrande (Template) benchmark actually showed a 6.86% increase in score with the compressed model.
Multiverse Computing attributes this counterintuitive result to the “healing” process of targeted re-training applied after compression, optimizing performance and, in some cases, yielding gains. Overall, the company claims the compressed model retained over 97% of the baseline accuracy, suggesting minimal performance impact for practical applications. The ability to run large language models on standard server hardware is expanding rapidly, and Multiverse Computing’s CompactifAI technology is demonstrably broadening access.
See today’s quantum computing news on Quantum Zeitgeist for the latest breakthroughs in qubits, hardware, algorithms, and industry deals.
