This paper studies LLM-based code completion for Pharo, a Smalltalk-inspired programming language with limited training data and previously only single-token IDE completion. The authors describe an end-to-end pipeline covering Pharo-specific data curation, continued pre-training, and fine-tuning of open code models. They also introduce benchmarks for Pharo syntax learning and masked completion on real-world GitHub repositories. According to the abstract, Pharo-specialized models outperform their base checkpoints and larger code LLMs, while remaining small enough for real-time in-IDE assistance.
No heat snapshots are available in the last 24 hours.