Proceedings of the First International Conference on Language Technologies for Low-resource Languages (LaTeLL 2026)

Chairs
Lamsiyah, Salima
Ranasinghe, Tharindu
Ezzini, Saad
Estevanell-Valladares, Ernesto Luis
Mitkov, Ruslan

LaTeLL 2026 Proceedings Home | LaTeLL WEBSITE | Bulgarian Association for Computational Linguistics WEBSITE

Full proceedings volume (PDF)
Schedule and Author index (HTML)
Bibliography (BibTeX)
Live website


pdf bib Front matter pages
pdf bib A Kazakh–Russian Corpus of Child-Directed Language for Low-Resource Languages
Albina Mukusheva, Achille Fusco and Cristiano Chesi
pp. 1‑7
pdf bib Data Curation, Annotation Quality, and Error Patterns in Old English Automatic Lemmatisation
Javier Martín Arista
pp. 8‑17
pdf bib Creating the RozMuz Corpus: Applying Ethics and Technology in Language Research
Alicja Helena Derych, Bartłomiej Alberski, Hubert Jankowski and Paweł Dembowski
pp. 18‑23
pdf bib Towards a Universal Dependencies Treebank for Amazigh: A Tarifit Pilot Study
Azzeddine Afrouni, Fadoua Ataa Allah and Jamal Abarnous
pp. 24‑29
pdf bib Arabic Dialect-to-MSA Translation: A Comparative Evaluation of Large Language Models and Neural Machine Translation
Maram I. Alharbi, Jihad R’baiti and Ruslan Mitkov
pp. 30‑39
pdf bib Cross-Lingual Transfer from Portuguese to Nheengatu: Evidence for Contact-Induced Convergence as a Computational Bridge
Rafael Macario Fernandes
pp. 40‑49
pdf bib Retrieval-Augmented Machine Translation for Bohairic Coptic: A Pilot Case Study
So Miyagawa
pp. 50‑58
pdf bib Automating Mwotlap morphology using formal grammar
Fabio Meroni, Alexandre François and Max Silberztein
pp. 59‑67
pdf bib Dependency Parsing Across the Resource Spectrum: Evaluating Architectures on High and Low-Resource Languages
Kevin Guan, Happy Buzaaba and Christiane Fellbaum
pp. 68‑79
pdf bib A finite-state model of Innu verbal inflection
Loïc Daignault-Pichette and François Lareau
pp. 80‑88
pdf bib Morphological Operations in the Persian Verbal System
Marzieh Rabiei and Max Silberztein
pp. 89‑95
pdf bib A Massive Open-Source Corpus for Valencian: Over 4.7 Billion Tokens for Low-Resource Language Modelling
Yoan Gutiérrez, Juan Pablo Consuegra-Ayala, Robiert Sepúlveda-Torres and Rafael Muñoz Guillena
pp. 96‑101
pdf bib LuxIT: A Luxembourgish Instruction Tuning Dataset from Monolingual Seed Data
Julian Valline, Cedric Lothritz, Siwen Guo and Jordi Cabot Sagrera
pp. 102‑120
pdf bib Topic Modeling for Moroccan Darija: A Comparative Study of Classical Machine Learning Approaches
Salma Mekaoui, Ilham Chaker, Arsalane Zarghili and Nikola S. Nikolov
pp. 121‑130
pdf bib Distil and Evolve: Compact Adaptive Document Classification with Continual Learning for Low-Resource Settings
Mehdi Soufiane, Mourad Hammale, Kawtar Zerhouni, Adil Ahidar-Coutrix and Rachid Benouini
pp. 131‑139
pdf bib Do Large Language Models Style-Shift? Register-Conditioned Stylistic Fronting in AI-Generated Icelandic
Anton Karl Ingason, Johanna Mechler and Lilja Björk Stefánsdóttir
pp. 140‑149
pdf bib Benchmarking Large Language Models on Mandarin Proverb Explanation and Contextual Matching
Xiaojing Zhao, Salima Lamsiyah, Emmanuele Chersoni and Han Xu
pp. 150‑160
pdf bib LUXDIAG-RAG: Diagnostic Evaluation of Retrieval-Augmented Generation for Luxembourgish Reading Comprehension
Keerthana Murugaraj, Hedi Tebourbi, Christophe Friezas Gonçalves and Salima Lamsiyah
pp. 161‑171
pdf bib Assessment of Human-in-the-Loop Multi-LLMs for Low-Resource Educational Data Expansion
Christophe Friezas Gonçalves, Hedi Tebourbi, Christoph Schommer and Salima Lamsiyah
pp. 172‑180
pdf bib Human-Centred Approaches in Low-Resource Languages for Educational Applications and Language Learning: A Systematic Literature Review
Noura El Moussa
pp. 181‑190
pdf bib SimuLe-Lux: A Neuro-Symbolic Teaching System for Diagnosing L2 Reading Comprehension
Christophe Friezas Gonçalves, Christoph Schommer and Salima Lamsiyah
pp. 191‑198
pdf bib Towards Readability Assessment for Under-Resourced Arabic Dialects: A Study of Moroccan Darija
Houdaifa Atou, Nouran Khallaf, Salima Lamsiyah and Ruslan Mitkov
pp. 199‑225
pdf bib Why Is Current XAI Not Enough for Arabic NLP? A Critical Survey of the Explainability Gap
Salima Lamsiyah and Ruslan Mitkov
pp. 226‑236
pdf bib Benchmarking Speech Foundation Models and LLMs for Sentiment Analysis in Low-Resource Najdi Arabic of Saudi Arabia
Nadia Ghezaiel Hammouda, Maram I. Alharbi and Ruslan Mitkov
pp. 237‑245
pdf bib Spatial Entity Extraction Methods from Arabic Texts in the Context of Epidemiological Surveillance: A Comparative Study
Fatima Ezzahra El Houbri, Najlae Idrissi, Mathieu Roche and Sarah Valentin
pp. 246‑256
pdf bib Building an Aligned Speech Corpus for Northern Russian
Kseniia Protonina and Daniil Ignatev
pp. 257‑263
pdf bib Probing Geographic Performance Gaps in Moroccan ASR with ALAMA: Annotated Local Audio of Moroccan Arabic
Avery Cole Kanel, Christian Schuler, Bouazza Laracha, Imrane Lbouhli, Yassine Chaouri, Yusser Al Ghussin and Timo Baumann
pp. 264‑271
pdf bib Generator-Guided Amount Recovery for Voice-Based Financial Record-Keeping in Mooré-French Code-Switched Speech
Maimouna Ouattara, El-Hacen Diallo, Fred Philippy, Abdoul Kader Kaboré, Jacques Klein and Tegawendé F. Bissyandé
pp. 272‑283
pdf bib LëtzCross: A Cross-Lingual Page-Level Benchmark for Multimodal Retrieval over Luxembourgish Documents
Omar El Bachyr, Fred Philippy, Laura Maria Bernardy, Saad Ezzini, Jacques Klein and Tegawendé Bissyandé
pp. 284‑294
pdf bib Improving OCR for a Latvian Pronunciation Dictionary
Viesturs Jūlijs Lasmanis
pp. 295‑303
pdf bib LLM-based conversational systems for Sinhala: Investigating model limitations and exploring a framework for agentic retrieval
Malithi P. Alahapperuma, Andreas Vlachidis and Antonis Bikakis
pp. 304‑308
pdf bib Clinical Entity Recognition from Electronic Health Records and Linking to Biomedical Knowledge Bases in Low-resource Language Settings
Eno-Martin Lotman, David Hübner, Kris Collins, Svetla Boytcheva and Ivelina Nikolova-Koleva
pp. 309‑317
pdf bib Beyond RAG: A Multi-Graph, Multi-Agent, Recursive Retrieval Architecture for Traceable and Source-Grounded Legal Answers
Assia Bouamir, Marie Bonnin, Youssef Al Mouatamid and Jihad Zahir
pp. 318‑325
pdf bib FLICK: Few-Label Incremental Learning for Low-Resource Dialects
Ali Almutairi, Abdullah Alsuhaibani, Shoaib Jameel, Aditya Joshi, Gelareh Mohammadi and Imran Razzak
pp. 326‑338
pdf bib Linguistic Proximity Enables Pivot-Based Machine Translation for an Under-Resourced Tribal Language
Pooja Singh, M Kaab Bin Shahid, Atai Waris Khan, Aryan Kumar Jha and Sandeep Kumar
pp. 339‑349
pdf bib Cross-Lingual Hate Speech Detection in Low-Resource Languages: The Case of Persian and Kurdish
Shahin Yousefi, Ernesto Luis Estevanell-Valladares and Ruslan Mitkov
pp. 350‑360
pdf bib Is Overconfidence Language-Specific? Cross-Lingual Calibration and Recalibration Transfer in Base and Instruction-Tuned LLMs
Divya Gupta and Ruslan Mitkov
pp. 361‑369
pdf bib Knowledge Tracing for Early Childhood Learners with Minimal Telemetry from Low-Connectivity Environments: Evidence from India’s Anganwadi Ecosystem
Badmavasan Kirouchenassamy, Rahul Singh, Vishnu Dev, Meenakshi Garg, Sukhna Sawhney and Sudeep Gowrishankar
pp. 370‑377
pdf bib Kuwain 1.5B: An Arabic SLM via Language Injection
Khalil Hennara, Sara Chrouf, Mohamed Motasim Hamed, Zeina Aldallal and Safwan AlModhayan
pp. 378‑387
pdf bib A Neuro-Symbolic RAG System for Marine Environmental Law: from Domain Ontology to Knowledge Graphs and Context-Enriched Legal Question Answering
Assoumana Souley Hadiza, Youssef Al Mouatamid, Marie Bonnin and Jihad Zahir
pp. 388‑397
pdf bib Evaluating Model-Task Fit in Arabic Word Sense Disambiguation
Yousef Younes, Abdelhalim Hafedh Dahou and Brigitte Mathiak
pp. 398‑407
pdf bib Enhancing Urdu ASR with Whisper v3: Fine-Tuning on Latest Datasets and Realistic Multi-Speaker Evaluation with SLM Post-Processing
Zehra Ahmed, Farah Inayat, Zuha Aqib and Sajjad Haider
pp. 408‑417
pdf bib Baseer: A Vision-Language Model for Arabic Document-to-Markdown OCR
Khalil Hennara, Muhammad Hreden, Mohamed Motasim Hamed, Ahmad Bastati, Zeina Aldallal, Sara Chrouf and Safwan AlModhayan
pp. 418‑428
pdf bib Worldview Annotation for Low-Resource Languages
Tetiana Ilman
pp. 429‑437
pdf bib A Multilingual Sentence-Transformer Baseline for Serbian News Topic Classification
Saša Petalinkar, Milica Ikonić Nešić, Ranka Stanković and Jelena Graovac
pp. 438‑448
pdf bib Wasm: A Pipeline for Constructing Structured Arabic Interleaved Multimodal Corpora
Khalil Hennara, Ahmad Bastati, Muhammad Hreden, Mohamed Motasim Hamed, Zeina Aldallal, Sara Chrouf and Safwan AlModhayan
pp. 449‑458
pdf bib Linguistic Specialization of Arabic in Text Embedding Models via Multi-Teacher Knowledge Distillation
Ahmad Abdelfattah, Zeina Aldallal, Sara Chrouf and Safwan AlModhayan
pp. 459‑467

Last modified on September 29, 2026, 9:59 a.m.