MerchantBench introduces a persistent benchmark for evaluating long-term coherence in LLM agents performing seller-side e-commerce operations. Its 365-day, order-level simulation is grounded in 98,843 real e-commerce product records and exposes 26 interaction tools. Agents must make interdependent decisions across product sourcing, listing and pricing control, cash-flow management, and adaptation to feedback arriving at different delays. The study evaluates eight LLMs under two agent frameworks across 48 runs. Unlike bounded task benchmarks, the environment links present actions to future choices and makes incoherence observable through cumulative operational effects.
No heat snapshots are available in the last 24 hours.