Key points are not available for this paper at this time.
Reinforcement learning (RL) shows promise for automated quantum circuit design but often stalls due to a “fidelity trap”: by optimizing only state fidelity, agents overlook entanglement structure and become stranded in suboptimal, overly complex circuits, resulting in significantly reduced search efficiency. In this paper, we propose a scheme that overcomes this barrier by implementing an entanglement-aware learning framework and enhancing the agent's reward function with a direct, quantitative measure of entanglement. This approach offers a more comprehensive physical description of the state space. We demonstrate the efficacy of this principle on three- and four-qubit state-synthesis tasks within an expanded gate set. For this problem, where the fidelity-driven agent systematically fails to discover the minimal-depth circuit, our entanglement-aware agent consistently succeeds. This transformative result is highly robust against variations in initial random seeds and extends to multiqubit systems even in the presence of noise. Our findings establish a generalizable principle that incorporating entanglement as an auxiliary reward can significantly enhance RL-based solutions for a broad class of fidelity-centric tasks in quantum physics and pave the way for scalable, automated discovery on near-term quantum devices.
Yu et al. (Fri,) studied this question.