This paper introduces TradeLens, a trace-grounded toolkit for evaluating whether LLM-based trading agents generate enough incremental profit to cover the costs of reasoning, tool use, and continual decisions. It reconstructs trading trajectories, attributes profit and cost to interpretable evidence, and diagnoses intelligence-to-profit conversion. Across backbone models, capital scales, trading frequencies, and system architectures, the authors report distinct failure patterns: poor asset selection for DeepSeek-V3.2 and negative timing for GLM-4.7. The work argues that agent evaluation should move beyond capability or return rankings toward trace-level economic diagnosis.
No heat snapshots are available in the last 24 hours.