Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

Request-Level Energy Attribution for Batched LLM Serving

First seen · 7/11/2026, 03:41 PMLatest activity · 7/11/2026, 03:41 PM

This paper presents JouleShare, a framework for attributing energy consumption to individual requests in batched LLM serving. Its offline harness replays request subsets with vLLM, integrates GPU power telemetry, and computes exact Shapley energy shares as measured ground truth. JCalib then predicts those shares from inexpensive request features for online use. Across 16 model/workload runs on three data-center GPUs, token-proportional attribution had average normalized L1 errors of 0.440 under static batching and 0.458 under continuous batching. JCalib reduced them to 0.116 and 0.177, respectively.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/11, 03:41 PMnot independentRepresentative
    Request-Level Energy Attribution for Batched LLM Serving