This paper introduces quantum-api-drift, a benchmark for testing whether LLM-generated quantum code targets the SDK version requested by the user. Using Qiskit v0.43, v1.3, and v2.0, the study evaluates 17 models across 50 tasks, producing 450 generated samples and 1,350 executions per model. Among 16 matched REST API models, diagonal Pass@1 ranges from 0.02 to 0.85. Claude Opus 4.7 leads on v0.43 and v2.0, while Grok 4.20 leads on v1.3. Repair with documentation succeeds in 0.19–0.59 of attempts and works better for v2.0 migration.
No heat snapshots are available in the last 24 hours.