Multilingual Email Classification Under Class Imbalance for University Rector’s Office Routing
Keywords
Abstract
This study examines the issue of manual sorting of emails in the Rector’s office of a local university. Many emails received by the Rector’s secretariat require manual review and subsequent forwarding to other university departments. In this case, the emails were not processed directly by the secretariat; they had to be forwarded to seven different recipient departments. Emails intended for the departments of education, education quality, education outreach, the international office, science, the utility department and the vice-rector were analyzed. The initial email database contained 9.842 emails. After removing duplicates, 8.153 unique emails remained. Using an algorithm that combined duplicate removal with a search for similar emails yielded 4.672 unique records. Only 477 of these potential records could be considered sufficiently labelled for supervised model training. The three families of models investigated were: TF-IDF using logistic regression, TF-IDF using a support vector machine (SVM), and XLM-RoBERTa. The combination of TF-IDF and SVM yielded the best results. The combination of TF-IDF and SVM outperformed TF-IDF with logistic regression in terms of the F-measure (macro F1: 0.6421, weighted F1: 0.7920 and accuracy: 0.8021) on a random split and an F-measure (macro F1: 0.5371, weighted F1: 0.6854 and accuracy: 0.6947) on a chronological split. Class analysis demonstrated excellent results for the most common high-frequency classes, such as “science” and “education”. However, results for lower-frequency classes were variable. Recall for “education outreach” was zero for both splits due to the presence of only one test record. Of the 102 total errors made in both TF-IDF combinations, all but one occurred below a confidence threshold of 0.5. Thus, overall, this paper identifies the automatic forwarding of emails for the University Rector’s Office as a multilingual classification task with extreme class imbalance. The lack of sufficient labelled data and limited lexical models is the greatest obstacle to implement the solution for this task.
