This paper introduces TAPR, a Task-Aware Prompt Rewriter that reformulates user prompts into task-optimized instructions. TAPR is trained with reinforcement learning using Group Relative Policy Optimization (GRPO), with rewards produced by LLM-as-judge evaluations of both the rewritten prompt and the resulting task output. Using Phi-4-mini-instruct as the base model, the authors report consistent gains across question answering, summarization, and arithmetic reasoning tasks. The reported benchmarks include Natural Questions and GSM8K. Code is available on GitHub, but the abstract does not provide detailed numerical results, judge-model specifications, or ablation findings.
No heat snapshots are available in the last 24 hours.