Mr. Arian Shekarlaban, a member of the laboratory, defended his master’s thesis entitled “Training a Persian Large Language Model with Reduced Resource Requirements” on October 19, 2025 (27 Mehr 1404). The abstract of his thesis is presented below.
In recent years, large language models have emerged as one of the key achievements of deep learning, driving remarkable advances in natural language processing. However, the development of these models has primarily focused on high-resource languages. Persian has received comparatively less attention because of limited data availability and computational resources.
This research therefore aimed to design and implement an efficient language model for Persian under resource-constrained conditions. The primary objective was to develop a smaller model with lower training costs while maintaining performance competitive with that of substantially larger models.
A review of previous studies showed that existing model compression and optimization approaches, although effective for high-resource languages, do not adequately address the specific requirements of Persian. These approaches may also encounter challenges such as considerable performance degradation and catastrophic forgetting.
To address these limitations, an innovative method was proposed. First, multilingual pretrained models were subjected to continual pretraining using fewer than ten billion Persian tokens and minimal hardware resources. This process produced a smaller yet optimized language model for Persian.
The main contribution of this research, referred to as “self-merging,” was then introduced. In this approach, the Persian-trained model is merged with its original pretrained version through a process that requires no additional training. This method reduced catastrophic forgetting and language bias while simultaneously improving the model’s general language understanding and its ability to analyze Persian texts.
The proposed model was evaluated on a range of Persian natural language processing tasks. The results showed that, despite its smaller size and lower training cost, the model achieved performance competitive with models several times larger across many tasks.
Overall, this research demonstrates that developing locally adapted large language models for Persian is feasible even under limited-resource conditions and can contribute to the advancement of artificial intelligence applications for the Persian language.
