This paper introduces EAGLE-GRPO, a post-training method for VLMs that write image-editing prompts for e-commerce creatives. Standard GRPO assigns reward only to the complete prompt, making it difficult to identify whether composition, background, or selling-point presentation caused an outcome. EAGLE-GRPO decomposes group-centered rewards over predefined design elements and formulates element-level credit assignment as kernel ridge regression. The authors derive a closed-form solution that requires neither additional rollouts nor a separate credit-assignment model. According to the abstract, the method maintains gains for more training steps and produces prompts yielding higher-quality e-commerce images than competing VLM prompt-writing baselines.
No heat snapshots are available in the last 24 hours.