The paper introduces Alipay-PIBench, a benchmark for evaluating coding agents on realistic Alipay payment-integration tasks. It includes nine product-specific projects and 18 task instances spanning Basic functional completion and Advanced risk-aware hardening. Evaluation combines deterministic static, unit, integration, and end-to-end checks with LLM-assisted semantic assessment. Across six coding-agent models, mean rubric pass rates range from 68.58% to 91.37% when an Alipay payment-integration skill is available. The skill improves mean pass rate by 10.31 percentage points on average, with variation across models, products, and scenarios.
No heat snapshots are available in the last 24 hours.