Multilingual Contextual Lemmatization
Overview
This repository brings together research on multilingual contextual lemmatization across languages with varying levels of morphological complexity. In particular, we study how contextual representations, morphological information, edit-script design, and multilingual transfer affect lemmatization performance across languages and domains.
The repository is structured around three publications that form the basis of this line of research.
Publications
On the Role of Morphological Information for Contextual Lemmatization
Olia Toporkov and Rodrigo Agerri. Computational Linguistics, 50(1), 2024, pp. 157-191.
This work investigates the role of morphological information in contextual lemmatization across six languages with different levels of morphological complexity. It evaluates whether explicit morphological features are necessary when using modern contextual representations, considering both in-domain and out-of-domain settings.
Evaluating Shortest Edit Script Methods for Contextual Lemmatization
Olia Toporkov and Rodrigo Agerri. LREC-COLING 2024.
This paper studies how different Shortest Edit Script (SES) representations affect contextual lemmatization. The experiments cover seven languages and compare multilingual and language-specific pretrained encoder models in both in-domain and out-of-domain settings.
Lemma Dilemma: On Lemma Generation Without Domain- or Language-Specific Training Data
This work explores contextual lemmatization when domain- or language-specific training data is unavailable. It compares supervised encoder-based approaches and cross-lingual transfer with direct in-context lemma generation using large language models (LLMs) across 12 languages.
Language Coverage
The three studies cover typologically diverse languages with different morphological profiles. The table below summarizes the language coverage across the three publications.
| Language | Morphological Information | SES | Lemma Dilemma |
|---|---|---|---|
| Basque | ✓ | ✓ | ✓ |
| Czech | ✓ | ✓ | ✓ |
| English | ✓ | ✓ | ✓ |
| Russian | ✓ | ✓ | ✓ |
| Spanish | ✓ | ✓ | ✓ |
| Turkish | ✓ | ✓ | ✓ |
| Polish | — | ✓ | ✓ |
| French | — | — | ✓ |
| German | — | — | ✓ |
| Italian | — | — | ✓ |
| Finnish | — | — | ✓ |
| Icelandic | — | — | ✓ |
| Swedish | — | — | ✓ |
Code and Data
Code and experimental resources are organized by publication.
multilingual-contextual-lemmatization/
├── paper-1-morphology/
│ ├── code/
│ └── results/
├── paper-2-edit-scripts/
│ ├── code/
│ └── results/
├── paper-3-crosslingual-llm/
│ ├── code/
│ └── results/
└── data/
Citation
If you use this repository, please cite the paper(s) relevant to the experiments or models you use. BibTeX entries are available from the ACL Anthology pages linked above.
Authors and affiliation
Developed at HiTZ — Basque Center for Language Technology, University of the Basque Country (UPV/EHU).