This paper argues that Kazakh-Russian code-switching identification is governed largely by annotation policy rather than model choice. Russian loanwords integrated into Kazakh are often mislabeled as switches because both languages use Cyrillic. The authors release a document-level gold LID dataset whose guidelines classify integrated borrowings as Kazakh and reserve mixed labels for clause-level switches. They also provide a mixed-only sentiment pool for a filter-first cascade. On a shared test set, FastText, Lingua, raw and windowed HeLI, character-trigram Naive Bayes, and XLM-R show performance ranging from weak to strong, highlighting the importance of the annotation boundary.
No heat snapshots are available in the last 24 hours.