With roughly 2.9 million large language models now publicly available, businesses struggle to pinpoint the best option for their specific needs. Multiverse Computing is addressing this challenge with Luminary, a new platform that evaluates LLMs across complete, multi-turn conversations, a capability missing from current single-prompt assessments.
Luminary uses simulations to model both customers and backend systems, revealing failure points like lost context or misused tools before deployment. “Every enterprise AI programme starts with the same question, and almost nobody answers it with evidence: which model should we actually run?” says Multiverse Computing, offering Luminary as the solution.
Luminary Evaluates LLMs with Full Multi-Turn Conversation Simulation
Luminary identifies gaps in initial briefs and automatically generates multi-turn user journeys, accounting for both standard scenarios and challenging edge cases, all while adhering to a customer’s established policies. This automated journey creation addresses a critical limitation of current LLM evaluation methods, which often rely on manually crafted prompts that fail to capture the complexity of real-world interactions. The platform’s ability to anticipate adversarial scenarios, unexpected user inputs designed to test the model’s robustness, further strengthens the assessment process, moving beyond simple “happy path” testing.
Each evaluation within Luminary unfolds inside a digital twin, simulating not only the customer but also the backend systems and tool calls the LLM must interact with during a conversation. This approach contrasts sharply with single-turn evaluations, which provide a limited snapshot of performance.
Failure points, such as lost context or incorrect tool usage, become apparent over multiple conversational turns, revealing weaknesses that would remain hidden in isolated prompt-response tests. Luminary moves beyond metrics like cost per million tokens, instead measuring cost per completed task, success rate, and policy compliance across the entire conversation. This shift in focus reflects the platform’s emphasis on real-world outcomes rather than abstract performance indicators. Every test case is fully reproducible, turn by turn, and generates an exportable audit trail for compliance teams, which is helpful for regulated industries.
The platform’s vendor-neutral design allows businesses to compare models from Multiverse alongside those from any OpenAI-compliant provider, ensuring an unbiased assessment. “Luminary evaluates the whole conversation,” the company states, emphasizing its ability to assess agentic and conversational systems, a capability beyond the reach of single-turn evaluation. Teams pick the model with the loudest launch and discover the cost six months later,” said Multiverse Computing. “And if you are building a chatbot or an agent, scoring one prompt at a time tells you nothing, because your product is the conversation, not the reply.
Luminary treats model selection the way aerospace treats airframe design. You put the thing in a wind tunnel before you fly it, across the whole flight envelope.”
Every enterprise AI programme starts with the same question, and almost nobody answers it with evidence: which model should we actually run? Teams pick the model with the loudest launch and discover the cost six months later.
Enrique Lizaso, co-founder & CEO of Multiverse Computing
Intelligent Test Generation & Digital Twin Workflow for LLMs
Luminary distinguishes itself from conventional LLM evaluation through its focus on complete conversational journeys, not isolated prompts. This shift reflects a focus on real-world outcomes, prioritizing practical utility over theoretical benchmarks. This level of transparency is important for organizations operating in regulated industries, enabling them to demonstrate due diligence in their AI deployments, Multiverse Computing says.
Enterprises can also use Luminary for ongoing model monitoring, re-running evaluations as new models and versions become available to assess performance regressions, optimize costs, and compare vendors. The platform is currently available through early access, offering a new approach to LLM evaluation for businesses navigating a complex generative AI environment.
We simulate the customer over every turn, the systems and the tool calls, then hand the business a ranked answer on cost per completed task, not a benchmark score.
Cost-Per-Completed-Task & Policy Compliance Metrics in Luminary
The platform assesses models not on isolated prompts, but across entire conversational journeys, revealing shortcomings that emerge only after several exchanges, specifically at turns four, seven, and nine, where models may lose context or misapply tools. This granular approach contrasts with evaluations limited to single interactions, which fail to capture these nuanced failure modes. Each model undergoes testing within a simulated environment replicating both the customer and the backend systems it would interact with during a multi-turn conversation, including function executions and tool calls.
Founded in 2019 and headquartered in San Sebastián, Spain, Multiverse Computing secured $570 million in Series C funding on July 27, 2026, to compress AI models for edge devices. A partnership with Zylon enables the delivery of AI models running natively on Zylon’s on-premise platform within secure European networks, while collaboration with Axelera AI integrates compressed models with Axelera’s edge device platforms. The platform’s ability to rerun evaluations as new models become available facilitates ongoing regression testing, model upgrades, and vendor comparisons, offering a dynamic approach to maintaining optimal performance and cost efficiency.




See today’s quantum computing news on Quantum Zeitgeist for the latest breakthroughs in qubits, hardware, algorithms, and industry deals.
