DreamQAS is a model-based reinforcement-learning framework for quantum architecture search that keeps deterministic circuit transitions and action legality exact while learning only the expensive feedback produced after VQE optimization. Its recurrent randomized-prior ensemble supports multi-step imagined policy training, with uncertainty-aware pessimism, truncation, and selective real-VQE checks. According to the abstract, under a shared 15,000-episode budget, DreamQAS achieved the lowest mean frozen-policy energy error on four of five molecular tasks and required 1.6–2.0× fewer real VQE calls at common fine-error targets, rising to 10.6× fewer on BeH2-8q.
No heat snapshots are available in the last 24 hours.