The role of quantum entanglement in competitive reinforcement learning was studied using the game Pong. An 8-qubit parameterized quantum circuit served as a feature extractor in a Proximal Policy Optimization algorithm, enabling comparison between separable circuits and architectures with fixed (CZ) or trainable (IsingZZ) entangling gates. Entangled circuits systematically outperform separable ones at equal parameter counts, and in low-capacity regimes match or exceed classical multi-layer perceptrons. Representation similarity analysis revealed that entangled circuits form structurally distinct features consistent with improved modeling of interacting state variables. These findings establish entanglement as a functional resource for representation learning in competitive reinforcement learning.
In ping-pong, two paddles connected by an invisible thread feel each other's movements, guessing the ball's trajectory. Researchers gave a computer a similar thread — quantum entanglement. Into an ordinary trial-and-error learning program, they embedded a quantum circuit of eight qubits — particles that can spin like a coin on its edge. When they become entangled (John Stewart Bell showed the weirdness of this effect), they turn into one whole: push one, and all respond. This circuit served as the algorithm's "eyes," turning screen pixels into interpretable signals.
The thread's tension was tweaked on the fly during the game, strengthening the bond. And the entangled player beat regular neural networks, especially when resources were scarce. It discerned hidden patterns in the opponent's moves, imperceptible to simple programs. The thread vibrated, transmitting information faster than any gravitational waves, and its pull gathered qubits together like a miniature black hole. It's astonishing that such a connection required cooling the chip to near absolute zero — a temperature lower than interstellar space.
🎯 To entangle qubits in the lab, chips are often cooled to near absolute zero — colder than outer space.
🎬 This work brings to mind the idea from the novel Ender's Game, where an AI learns by analyzing human behavior in strategy games.