Paper

2026Public

Evaluating Verified Autonomy in Quantum Engineering

Summary

Before AI agents run quantum hardware on their own, their work has to be verifiable. Quantum-Harbor is a quantum lab built for that verification. Agents work as they would at a real bench. Depending on the task, a grader they cannot see checks answers against hidden ground truth, recomputes results from the recorded experiments, or replays the submission on fresh data.

Media

Quantum-Harbor architecture. (a) The agent environment and its public interfaces. (b) Simulated quantum hardware with hidden device parameters and a private evidence log. (c) The grader, which recomputes results, binds submitted claims to the recorded experiments, and retests calibrations on fresh data.
Quantum-Harbor architecture. (a) The agent environment and its public interfaces. (b) Simulated quantum hardware with hidden device parameters and a private evidence log. (c) The grader, which recomputes results, binds submitted claims to the recorded experiments, and retests calibrations on fresh data.
QIQCBench results for 17 agentic systems. (a) Pass rate on the pass-fail tasks. (b) Breakdown of non-pass trials by failure type. (c) Elo ratings over the 7 racing tasks. (d) Passes per task and system, grouped by scientific category.
QIQCBench results for 17 agentic systems. (a) Pass rate on the pass-fail tasks. (b) Breakdown of non-pass trials by failure type. (c) Elo ratings over the 7 racing tasks. (d) Passes per task and system, grouped by scientific category.

Paper

Authors
Naixu Guo, Changhao Li, Siyu Cheng, Qicheng Tang, Binzhao Luo, Bikun Li, Yuxuan Du, Shihao Ru, Jiaqi Cai
My role
Built the Quantum-Harbor framework. Collected the evaluation trajectories of all 12 Chinese models in QIQCBench.

Approach

The agent reads hardware specs and a lab notebook whose calibrations may be stale. It submits experiment jobs and analyzes the raw measurement records. It never sees the true device and noise parameters, and it cannot touch the evidence log that records every job.

The hardware layer covers transmons, neutral atoms, trapped ions, and other platforms. Replay and live-device backends plug in where supported.

QIQCBench runs 49 expert-authored tasks in the lab, spanning calibration and control, error correction and compilation, and sensing and networking.

Outcome

Seventeen frontier agentic systems from eight vendors took the benchmark. Verified pass rates run from 78% down to 3%, and QIQCBench separates systems that score alike on other agentic evaluations. Quantum-Harbor gives the field a baseline for measuring progress toward verified autonomy in quantum engineering.