Paper
Evaluating Verified Autonomy in Quantum Engineering
Summary
Before AI agents run quantum hardware on their own, their work has to be verifiable. Quantum-Harbor is a quantum lab built for that verification. Agents work as they would at a real bench. Depending on the task, a grader they cannot see checks answers against hidden ground truth, recomputes results from the recorded experiments, or replays the submission on fresh data.
Media
Paper
Approach
The agent reads hardware specs and a lab notebook whose calibrations may be stale. It submits experiment jobs and analyzes the raw measurement records. It never sees the true device and noise parameters, and it cannot touch the evidence log that records every job.
The hardware layer covers transmons, neutral atoms, trapped ions, and other platforms. Replay and live-device backends plug in where supported.
QIQCBench runs 49 expert-authored tasks in the lab, spanning calibration and control, error correction and compilation, and sensing and networking.
Outcome
Seventeen frontier agentic systems from eight vendors took the benchmark. Verified pass rates run from 78% down to 3%, and QIQCBench separates systems that score alike on other agentic evaluations. Quantum-Harbor gives the field a baseline for measuring progress toward verified autonomy in quantum engineering.