An attacker employing reinforcement learning can gain information about a quantum key distribution system by adapting their strategy to changes in signal transmission over time, with this adaptive behaviour previously unquantified. The modelling as a sequential decision process enables attacks to be learnt instead of pre-defined, extending its application to more complex channel conditions including those lacking known templates. An attacker can exploit fluctuations within quantum communication channels to gain more information than previously considered during security assessments.
The research quantifies how much advantage an adaptive eavesdropper gains; increases in Holevo information, a measure of accessible information, of up to 0.348 were found under specific conditions. These findings suggest existing calculations used to guarantee QKD system safety may be unnecessarily conservative, particularly for systems deployed across metropolitan areas or government networks. An attacker employing reinforcement learning can use fluctuations within quantum communication channels to gain more information than previously accounted for in security assessments because researchers quantified this adaptive behaviour until now.
Quantum key distribution (QKD) systems are typically assessed assuming a stable connection between calibrations, but realistically devices drift over time creating opportunities for sophisticated attacks. To understand these threats, the team modelled eavesdropping as a sequential decision process, akin to planning a route across town one intersection at a time, where each attack is learnt rather than pre-defined and adapts to changes in signal transmission.
They quantified how much advantage such an adaptive eavesdropper gains using Holevo information, which measures accessible information about the secret key like assessing the clarity of a distorted photograph; increases up to 0.348 were found under specific conditions.
Adaptive attack learning nears security limits in quantum key distribution
Holevo information, measuring how much data an eavesdropper can obtain about a quantum key, increased from 0.135 to 0.348 using reinforcement learning; this figure approaches the breach of established security bounds previously considered unattainable under realistic conditions. The improvement demonstrates that adaptive attackers outperform those employing static strategies when exploiting fluctuations within quantum communication channels, challenging current assumptions regarding stationary noise levels during QKD system assessments.
This approach frames eavesdropping as a sequential decision process enabling attacks to be ‘learnt’, instead of pre-defined, extending methodology to complex and untemplated noise models such as amplitude damping. Researchers at multiple institutions demonstrated enhanced information gain for device-independent E91 protocols with bilateral depolarising noise via adaptive eavesdropping strategies, increasing Holevo information, quantifying potential data leakage, to 0.348.
Further analysis revealed an increase of 0.024 in fidelity against BB84 utilising a drifting bit-flip channel; the team achieved $99\% of the theoretical upper bound for secure communication by jointly searching both circuit structure and rotation angles instead of relying on fixed templates. Attackers who adjust their methods based on observed fluctuations outperform those using pre-defined tactics, nearly reaching established security limits previously thought insurmountable given realistic channel conditions.
Limitations of simulated adaptive attacks on practical quantum cryptography
Quantifying an attacker’s advantage through adaptation is key to building truly secure communication networks; however, simulations rely on asymptotic detection statistics which may not fully reflect real-world performance with limited data or imperfect noise tracking. Determining the impact of delayed estimation, where the eavesdropper reacts slightly behind actual channel changes, remains an open question and could alter observed benefits. It is vital to acknowledge that these simulations employ simplified models and asymptotic detection statistics because real-world systems face limitations regarding both data quantity and noise estimation accuracy. Researchers modelled attacks that ‘learn’ over time by framing eavesdropping as a sequential decision process akin to planning a route where each step adapts to changing conditions; previously, such drifts were considered merely sources of error reducing key rates rather than opportunities for exploitation. This work provides valuable insight into potential gains achievable by attackers adapting to changing conditions in quantum key distribution networks, strengthening arguments for proactive security measures against increasingly sophisticated adversaries.
The research demonstrated an attacker could increase information gained from a quantum key distribution system by adjusting strategies based on channel fluctuations. Specifically, Holevo information rose to 0.348 and fidelity increased by 0.024$ when employing adaptive attacks compared with fixed approaches under simulated noise conditions. These findings suggest that simply accounting for channel drift as background noise is insufficient; instead, the possibility of active exploitation must be considered during protocol design. The authors indicate further work should investigate how delays in estimating these changes affect observed benefits.
👉 More information
🗞 Learnt Attacks on Quantum Key Distribution under Channel Noise and Device Drift
✍️ Marcel Mordarski, Abdelrahman Shehata and Daniel Budina (Imperial College); Benjamin Gras and Roberto Bondesan (Imperial College London)
🧠 ArXiv: https://arxiv.org/abs/2610.01792




See today’s quantum computing news on Quantum Zeitgeist for the latest breakthroughs in qubits, hardware, algorithms, and industry deals.
