STRATEGY, CONTENT AND NEW TECHNOLOGIES FOR TRAINING SPECIALISTS IN THE FIELDS OF INFORMATICS, INFORMATION TECHNOLOGY, ELECTRONICS AND AUTOMATION

Multilingual Email Classification Under Class Imbalance for University Rector’s Office Routing

Authors

H. Ye. Brusentsov ORCID 0000-0002-3346-0164
Lviv Polytechnic National University ROR

Keywords

multilingual email classification class imbalance weak supervision TF-IDF linear SVM

Abstract

This study examines the issue of manual sorting of emails in the Rector’s office of a local university. Many emails received by the Rector’s secretariat require manual review and subsequent forwarding to other university departments. In this case, the emails were not processed directly by the secretariat; they had to be forwarded to seven different recipient departments. Emails intended for the departments of education, education quality, education outreach, the international office, science, the utility department and the vice-rector were analyzed. The initial email database contained 9.842 emails. After removing duplicates, 8.153 unique emails remained. Using an algorithm that combined duplicate removal with a search for similar emails yielded 4.672 unique records. Only 477 of these potential records could be considered sufficiently labelled for supervised model training. The three families of models investigated were: TF-IDF using logistic regression, TF-IDF using a support vector machine (SVM), and XLM-RoBERTa. The combination of TF-IDF and SVM yielded the best results. The combination of TF-IDF and SVM outperformed TF-IDF with logistic regression in terms of the F-measure (macro F1: 0.6421, weighted F1: 0.7920 and accuracy: 0.8021) on a random split and an F-measure (macro F1: 0.5371, weighted F1: 0.6854 and accuracy: 0.6947) on a chronological split. Class analysis demonstrated excellent results for the most common high-frequency classes, such as “science” and “education”. However, results for lower-frequency classes were variable. Recall for “education outreach” was zero for both splits due to the presence of only one test record. Of the 102 total errors made in both TF-IDF combinations, all but one occurred below a confidence threshold of 0.5. Thus, overall, this paper identifies the automatic forwarding of emails for the University Rector’s Office as a multilingual classification task with extreme class imbalance. The lack of sufficient labelled data and limited lexical models is the greatest obstacle to implement the solution for this task.

0 0

How to Cite

[1]
“Multilingual Email Classification Under Class Imbalance for University Rector’s Office Routing”, Вісник ВПІ, no. 4, pp. 211–217, Sep. 2026, doi: 10.31649/.

Author Biography

H. Ye. Brusentsov, Lviv Polytechnic National University

Post-Graduate Student the Chair of Artificial Intelligence Systems

References

[1] S. Scerri, G. Gossen, B. Davis, and S. Handschuh, “Classifying Action Items for Semantic Email,” in Proceedings of the Seventh International Conference on Language Resources and Evaluation (LREC’10), Valletta, Malta, 2010. [Online]. Available: https://aclanthology.org/L10-1018/ .
[2] B. Sandrih Todorovic, K. Josipovic, and J. Kodre, “Three Approaches to Client Email Topic Classification,” in Proceedings of the 14th International Conference on Recent Advances in Natural Language Processing, Varna, Bulgaria, 2023, pp. 1015-1022. [Online]. Available: https://aclanthology.org/2023.ranlp-1.109/ .
[3] M. Marcuzzo, A. Zangari, M. Schiavinato, L. Giudice, A. Gasparetto, and A. Albarelli, “A multi-level approach for hierarchical Ticket Classification,” in Proceedings of the Eighth Workshop on Noisy User-generated Text (W-NUT 2022), Gyeongju, Republic of Korea, 2022, pp. 201-214. [Online]. Available: https://aclanthology.org/2022.wnut-1.22/ .
[4] J. Wang, and C. D. Manning, “Baselines and Bigrams: Simple, Good Sentiment and Topic Classification,” in Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics (ACL), Jeju, Republic of Korea, 2012, pp. 90-94. [Online]. Available: https://aclanthology.org/P12-2018/ .
[5] A. Conneau, et al., “Unsupervised Cross-lingual Representation Learning at Scale,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (ACL), 2020, pp. 8440-8451. [Online]. Available: https://aclanthology.org/2020.acl-main.747/ .
[6] H. He, and E. A. Garcia, “Learning from Imbalanced Data,” IEEE Transactions on Knowledge and Data Engineering, vol. 21, no. 9, pp. 1263-1284, Sep. 2009, https://doi.org/10.1109/TKDE.2008.239 .
[7] G. Salton, A. Wong, and C. S. Yang, “A Vector Space Model for Automatic Indexing,” Communications of the ACM, vol. 18, no. 11, pp. 613-620, Nov. 1975, https://doi.org/10.1145/361219.361220 .