Classical Algorithms Replicate Quantum Learning with Sufficient Data Samples

A new classical algorithm called kernelized Fitted Q-Iteration matches the performance of Quantum Q-learning in scenarios utilising uniformly random state-action samples. The method replicates results achieved by Quantum Q-learning, a technique employing parameterised quantum circuits for reinforcement learning. Specific conditions relating data encoding, kernels and problem structure enable replication of quantum results with purely classical computation.

The team from University of the Basque Country UPV/EHU and collaborating institutions created kernelized Fitted Q-Iteration. This development addresses an ongoing debate regarding potential benefits offered by quantum computing in machine learning; it seeks to replicate results achieved using Quantum Q-learning, a technique employing complex mathematical functions on quantum computers. They successfully matched the performance of Quantum Q-learning under specific conditions where data points are randomly selected.

By identifying how information encoding, kernels and problem structure interact, they demonstrated that certain quantum computations can be replicated classically. The researchers /EHU alongside collaborators have created a new classical algorithm called kernelized Fitted Q-Iteration. These findings establish key parameters for determining when classical algorithms might outperform their quantum counterparts. The work originates from researchers affiliated with National Institute for Theoretical and Computational Sciences (NITheCS), Freie Universität Berlin, South Africa National Institute, Basque Centre for Applied Mathematics (BCAM), Spain, African Institute, Stellenbosch University and African Institute for Mathematical Sciences (AIMS).

Classical Algorithm Equivalence Achieved for Quantum Reinforcement Learning Performance

Kernelised Fitted Q-Iteration now matches the performance of Quantum Q-learning; previously replicating this quantum reinforcement learning technique classically was impossible under these conditions. Using uniformly random state-action samples, the team /EHU and collaborating institutions demonstrated this equivalence, a scenario mirroring data collected during AI training through experience replay buffers. This dequantization, creating classical algorithms that mimic quantum behaviour, relies on specific relationships between information encoding, kernel design, and problem structure.

Beyond rigorously proving when replication occurs, Kernelised Fitted Q-Iteration offers a useful set of tools even without complete verification of those parameters; it provides an alternative approach to assessing variational quantum algorithm benefits in practical machine learning applications. A simplified data generation process was used to demonstrate its equivalence to Quantum Q-learning, specifically utilising uniformly random state-action samples common in AI training’s experience replay buffers.

The team established finite sample guarantees for this classical approach paired with kernels designed to mimic parameterised quantum circuits, creating a ‘classical shadow’ of quantum behaviour, and identified sufficient conditions relating information encoding within these circuits, kernel design and problem structure under which dequantization allows accurate result replication.

Kernelised Fitted Q-Iteration provides a classical analogue to assess Quantum Q-learning performance

A fundamental hurdle exists when searching for practical quantum advantages in machine learning; demonstrating genuine speedups requires algorithms that outperform their classical counterparts, but proving this is exceptionally difficult. Kernelised Fitted Q-Iteration can now replicate Quantum Q-learning’s performance under specific conditions, creating a ‘classical shadow’ of the quantum process and offering an alternative benchmark for assessing potential benefits. However, such equivalence relies heavily on uniformly random data samples, a condition rarely met by real-world complexity where information arrives sequentially and environments change over time.

Despite current dependence upon idealised data, uniformly distributed samples are rarely found in changing real-world scenarios, this work retains significant value as it establishes a high bar against which to measure any claimed speedups. This achievement provides an alternative benchmark for evaluating potential benefits from variational quantum algorithms, relying on utilising uniformly random data samples mirroring experience replay buffers common in artificial intelligence training.

The researchers showed that Kernelised Fitted Q-Iteration, a classical machine learning technique, could match the performance of Quantum Q-learning when provided with uniformly random data samples. This finding establishes a means to assess whether any speedups observed using quantum approaches are genuinely due to quantum effects or simply result from equivalent classical methods.

By creating this ‘classical shadow’ of quantum behaviour, scientists now have a higher standard against which to measure claimed advantages offered by variational quantum algorithms. The team focused on scenarios modelling large experience replay buffers and identified conditions relating circuit design and problem structure where accurate replication is possible.

👉 More information
🗞 Towards Surrogate Based Dequantization of Quantum Reinforcement Learning
✍️ Pablo Rodriguez-Grasa, Sofiene Jerbi, Mikel Sanz and Ryan Sweke
🧠 ArXiv: https://arxiv.org/abs/2609.16266

Stay current

See today’s quantum computing news on Quantum Zeitgeist for the latest breakthroughs in qubits, hardware, algorithms, and industry deals.

Avatar of Muhammad Rohail T.

Latest Posts by Muhammad Rohail T.: