CRAFT turns rubric-based evaluation into model-specific capability diagnosis. It extracts capability descriptions from prompt-rubric pairs, organizes them into a hierarchical capability tree, scores the target model at multiple levels, and selects weak nodes at the granularity where failures are clearest. Those weaknesses guide targeted supervised fine-tuning data generation. Across four open-source models, finance and legal domains, and 13 held-out benchmarks, CRAFT outperformed prompt-level EvalTree clustering and untargeted random generation on finance for all models, and on legal for three of four models.
No heat snapshots are available in the last 24 hours.