This paper studies why small language models may still misuse supplied statutes in legal question answering. The authors curate 2,165 bilingual QA records covering six Bangladeshi acts and three schedules, then fine-tune Qwen3.5 models at 0.8B, 2B, and 4B parameters. Evaluation uses the 2022 and 2023 Bangladesh Bar Council exams in Bangla and machine-translated English, under no-retrieval, BM25, and FAISS conditions, with strict consistency across three seeded runs. The 0.8B model’s 2022 English FAISS score rises from 2 to 34 out of 100, while the 4B model shows no detectable net gain.
No heat snapshots are available in the last 24 hours.