Enterprises deploy operational systems rather than raw checkpoints, yet prevailing benchmark suites almost exclusively evaluate advertised model identifiers. The authors introduce IB2, a measurement protocol emphasizing serving routes, precision, and tool-call parsers alongside weights. In experiments spanning eleven enterprise systems, differing serving arms shifted scores for the identical model revision from 77.38 to 82.54. By enforcing capability-binding preflight checks and retaining execution failures in scoring denominators, the protocol exposes real-world reliability gaps obscured by checkpoint-centric leaderboards.
There are 7 persisted snapshots in the last 24 hours. Peak heat was 0 at 9/12, 17:00; latest heat is 0.