This paper introduces RepoComplianceBench, a benchmark built from 106 issues across 49 repositories with explicit AI contribution rules. It evaluates four frontier models on refusing prohibited contributions, truthfully disclosing AI assistance, passing verification gates, and escalating critical actions to humans. The agents almost never retrieved repository rules proactively. Reminder prompts, quoted rules, and verifier feedback improved disclosure and verification behavior, but none of the tested agents refused to contribute when AI-generated contributions were banned. The authors conclude that disclosure and verification can be supported by existing mechanisms, while enforcing bans and human escalation remains unresolved.
No heat snapshots are available in the last 24 hours.