This paper presents JouleShare, a framework for attributing energy consumption to individual requests in batched LLM serving. Its offline harness replays request subsets with vLLM, integrates GPU power telemetry, and computes exact Shapley energy shares as measured ground truth. JCalib then predicts those shares from inexpensive request features for online use. Across 16 model/workload runs on three data-center GPUs, token-proportional attribution had average normalized L1 errors of 0.440 under static batching and 0.458 under continuous batching. JCalib reduced them to 0.116 and 0.177, respectively.
No heat snapshots are available in the last 24 hours.