Pıer
TidesCurrentsHarbor LightsLabBottlesAshore
Pıer

Navigation

  • Tides
  • Ashore
  • Harbor Lights
  • Agent Access
  • Changelog
  • Bottles
  • Now
  • Feedback

External links

GitHubCloudborne ↗

© 2026 Pier.

WatchingResearchWatching0 independent reports0

CRAFT: Clustering Rubrics to Diagnose Weak LLM Capabilities and Generate Targeted Fine-Tuning Data

First seen · 7/18/2026, 01:00 AMLatest activity · 7/18/2026, 01:00 AM

CRAFT turns rubric-based evaluation into model-specific capability diagnosis. It extracts capability descriptions from prompt-rubric pairs, organizes them into a hierarchical capability tree, scores the target model at multiple levels, and selects weak nodes at the granularity where failures are clearest. Those weaknesses guide targeted supervised fine-tuning data generation. Across four open-source models, finance and legal domains, and 13 held-out benchmarks, CRAFT outperformed prompt-level EvalTree clustering and untargeted random generation on finance for all models, and on legal for three of four models.

Event heat · last 24 hours

No heat snapshots are available in the last 24 hours.

No heat snapshots are available in the last 24 hours.

Reporting Timeline

  1. AggregatorarXiv7/18, 01:00 AMnot independentRepresentative
    CRAFT: Clustering Rubrics to Diagnose Weak LLM Capabilities and Generate Targeted Fine-Tuning Data