Intel, HPE and Multiverse Computing double speech-to-text speed

A single Hewlett Packard Enterprise ProLiant DL380 server now transcribes approximately 17,000 hours of audio daily using only CPUs, a feat previously requiring specialized hardware. Multiverse Computing, in collaboration with HPE and Intel, achieved this by optimizing the Whisper speech-to-text model to 0.4 billion parameters, half its original size, without sacrificing quality, the company says.

“Transcription is one of the highest-volume AI workloads inside a large organisation, and for years it has been treated as a job that needs dedicated accelerators,” says Multiverse Computing, demonstrating a pathway to significantly reduce costs and enhance data sovereignty for enterprises. The optimized model delivers double the throughput, 1,413 tokens per second, compared to the original, on standard enterprise infrastructure.

Intel Xeon 6 and vLLM Enable Accelerated Inference

The deployment of a 0.4 billion parameter Whisper model on standard server CPUs achieves 1,413 tokens per second, a doubling of throughput compared to the original 0.8 billion parameter version, demonstrating significant gains in speech-to-text inference speed. Multiverse Computing, in collaboration with Hewlett Packard Enterprise and Intel, accomplished this performance on a single HPE ProLiant DL380 server equipped with two Intel Xeon 6 processors, utilizing Intel Advanced Matrix Extensions and the vLLM serving framework.

This configuration bypasses the need for dedicated GPU acceleration, a departure from conventional approaches to high-volume transcription workloads. The optimized model requires only 0.75 GB of memory, a substantial reduction from the original’s 1.5 GB, without compromising transcription accuracy. This reduction in model size and computational demand allows for new on-premise speech-to-text processing options, addressing data sovereignty concerns and potentially lowering operational costs for enterprises.

Organizations generating substantial volumes of audio data, such as contact centers, financial institutions, and healthcare providers, often face significant expenses associated with cloud-based transcription services and the infrastructure required to support them. A contact center processing 20 million calls annually can generate approximately 2 million hours of audio, potentially exceeding one million dollars in transcription costs using traditional GPU-based solutions.

By shifting this workload to existing server infrastructure, companies can dramatically reduce these expenses and maintain complete control over sensitive data. The performance gains extend beyond throughput; time to first token was also nearly halved, decreasing from 2.11 seconds to 1.11 seconds, and the real-time factor improved from 356 to 712.

The optimized model maintains comparable accuracy to the original, achieving a word error rate of 2.75% on the LibriSpeech dataset in both English and Spanish, compared to 2.14% for the original, both comfortably below a 3% threshold. The consistent advantage of the optimized model across varying levels of load suggests scalability and reliability for demanding enterprise applications. This achievement represents the first result of an ongoing collaboration between Multiverse Computing, HPE, and Intel, signaling a broader trend towards efficient, specialized AI models that can run securely on existing hardware.

The ability to deploy and operate these models on standard enterprise servers empowers organizations to unlock previously inaccessible use cases, such as comprehensive transcription of all recorded calls for quality assurance and compliance, creation of searchable archives of meetings and support conversations, and captioning of extensive media libraries, according to the company. When transcription costs are minimized, the potential for using audio data expands exponentially, enabling downstream AI systems to benefit from a larger volume of clean, transcribed content.

The team emphasizes that this approach allows organizations to operate “on their own terms, on the hardware they already have,” fostering greater control and flexibility in AI deployment. This work demonstrates that high-volume AI workloads do not necessarily require specialized silicon.

Stay current

See today’s quantum computing news on Quantum Zeitgeist for the latest breakthroughs in qubits, hardware, algorithms, and industry deals.

Avatar of Rusty Flint

Rusty Flint

Rusty is a quantum science nerd. He's been into academic science all his life, but spent his formative years doing less academic things. Now he turns his attention to write about his passion, the quantum realm. He loves all things Quantum Physics especially. Rusty likes the more esoteric side of Quantum Computing and the Quantum world. Everything from Quantum Entanglement to Quantum Physics. Rusty thinks that we are in the 1950s quantum equivalent of the classical computing world. While other quantum journalists focus on IBM's latest chip or which startup just raised $50 million, Rusty's over here writing 3,000-word deep dives on whether quantum entanglement might explain why you sometimes think about someone right before they text you. (Spoiler: it doesn't, but the exploration is fascinating)

Latest Posts by Rusty Flint: