The paper presents ARCTIC, an AI code-critique system for large-scale AI-generated diffs. It combines intent prediction from conversation logs and metadata, drift detection through backtranslation, and code spotlighting to prioritize regions that need human attention. Its six-theme taxonomy is derived from 18,000 code reviews. Reported offline results include 0.86 F1 for intent prediction, QWK 0.907 for drift detection, and 2.4x the baseline reviewer’s quality-estimation performance at one-fifth the token usage. An experimental rollout reports a 5.76-point reduction in code misalignment, 90.2% approval for intent prediction, and no defects attributed to self-reviewed diffs since launch.
No heat snapshots are available in the last 24 hours.