We investigate the role of information in active feedback control of quantum many-body systems through reinforcement learning. Using (1+1)-dimensional stabilizer circuits up to 128 qubits, trained agents with partial state information prevent entanglement from spreading. When exceeding a critical information threshold, agents develop non-greedy stochastic strategies that reduce entanglement scaling from a volume law to an area law. The mechanism involves creating "bottlenecks" that give rise to pyramid-like structures in the spatial distribution of entanglement, effectively splitting the system and limiting the maximum achievable entanglement. The discovered strategies are shown to be essentially non-equilibrium and require real-time active feedback; they cannot be replaced by simple human-designed rules. This work lays foundations for classically feasible individual control of quantum degrees of freedom and demonstrates the ability of reinforcement learning to stabilize and reveal new critical properties of non-equilibrium steady states.
Quantum particles are like guests at a noisy party: as soon as they mingle, invisible connections arise — entanglement. Without control, these connections quickly entangle everything, creating chaos. But if you post a bouncer at the entrance who sees only a fraction of the guests, he can intervene and break unnecessary contacts. At first, he acts randomly, but over time he learns. The bouncer finds a clever trick: sometimes he lets bursts of disorder through, only to later create 'bottlenecks' — narrow spots that break the crowd into islands. Then the connections grow only along the edges of these islands, like surface area, rather than filling the entire volume. This resembles the structure of black holes, where information is tied to the horizon. This approach makes it possible to stabilize systems of dozens of particles, paving the way to reliable quantum devices.
🎯 The trained agent deliberately allows a little disorder to later achieve global order — a tactic no human thought of.
🎬 Just as in Asimov's 'Foundation' scientists predicted crowd behavior to avoid chaos, this algorithm learns to manage a quantum society of particles.