This paper profiles five vision-language models across three architecture families, four image resolutions, and two platforms: an NVIDIA RTX 3070 and Jetson Orin NX. It reports that average inference power varies by less than 5% across resolutions, image complexity, and prompt types. Each output token takes 11–39 times more wall-clock time than an input token, making decoding the dominant energy cost. Image complexity can create up to 4.1× energy differences at the same resolution because it changes output length. Visual-token pruning saves at most 10% for fixed-token models, while output-length control saves up to 97%.
No heat snapshots are available in the last 24 hours.