HalluTruthQA-4K is an Arabic factuality dataset containing 4,000 expert-curated question-answering instances across Islamic knowledge, history, science, and geography. It includes 1,643 hallucinated and 2,357 non-hallucinated model responses, with 1,843 character-level erroneous spans. Each instance pairs a question and generated response with a verified reference answer and five plausible distractors. Hallucinated answers additionally receive human-written explanations and hierarchical error types. The dataset is designated for Track 2 of the HalluScoring 2026 shared task and targets hallucination detection, error localization, explanation generation, and factual verification.
No heat snapshots are available in the last 24 hours.