Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations

First seen · 7/31/2026, 04:00 AMLatest activity · 7/31/2026, 04:00 AM

MerchantBench introduces a persistent benchmark for evaluating long-term coherence in LLM agents performing seller-side e-commerce operations. Its 365-day, order-level simulation is grounded in 98,843 real e-commerce product records and exposes 26 interaction tools. Agents must make interdependent decisions across product sourcing, listing and pricing control, cash-flow management, and adaptation to feedback arriving at different delays. The study evaluates eight LLMs under two agent frameworks across 48 runs. Unlike bounded task benchmarks, the environment links present actions to future choices and makes incoherence observable through cumulative operational effects.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorHuggingFace Daily Papers7/31, 04:00 AMnot independentRepresentative
    MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations