Proceedings of the First International Conference on Language Technologies for Low-resource Languages (LaTeLL 2026)

Chairs
Lamsiyah, Salima
Ranasinghe, Tharindu
Ezzini, Saad
Estevanell-Valladares, Ernesto Luis
Mitkov, Ruslan

LaTeLL 2026 Proceedings Home | LaTeLL WEBSITE | Bulgarian Association for Computational Linguistics WEBSITE

Full proceedings volume (PDF)
Schedule and Author index (HTML)
Bibliography (BibTeX)
Live website


PROGRAM

Monday September 28, 2026

 10:50–11:50 Session 1: Corpora, Annotation and Resource Development # %chair1 Yoan Gutiérrez
10:50–11:10A Kazakh–Russian Corpus of Child-Directed Language for Low-Resource Languages
Albina Mukusheva, Achille Fusco and Cristiano Chesi
11:10–11:30Data Curation, Annotation Quality, and Error Patterns in Old English Automatic Lemmatisation
Javier Martín Arista
11:30–11:40Creating the RozMuz Corpus: Applying Ethics and Technology in Language Research
Alicja Helena Derych, Bartłomiej Alberski, Hubert Jankowski and Paweł Dembowski
11:40–11:50Towards a Universal Dependencies Treebank for Amazigh: A Tarifit Pilot Study
Azzeddine Afrouni, Fadoua Ataa Allah and Jamal Abarnous
 11:50–12:30 Session 2: Machine Translation and Cross-Lingual Transfer # %chair1 Salima Lamsiyah
11:50–12:10Arabic Dialect-to-MSA Translation: A Comparative Evaluation of Large Language Models and Neural Machine Translation
Maram I. Alharbi, Jihad R’baiti and Ruslan Mitkov
12:10–12:20Cross-Lingual Transfer from Portuguese to Nheengatu: Evidence for Contact-Induced Convergence as a Computational Bridge
Rafael Macario Fernandes
12:20–12:30Retrieval-Augmented Machine Translation for Bohairic Coptic: A Pilot Case Study
So Miyagawa
 14:20–15:30 Session 3: Morphology, Syntax and Formal Modelling # %chair1 Javier Martín Arista
14:20–14:40Automating Mwotlap morphology using formal grammar
Fabio Meroni, Alexandre François and Max Silberztein
14:40–15:00Dependency Parsing Across the Resource Spectrum: Evaluating Architectures on High and Low-Resource Languages
Kevin Guan, Happy Buzaaba and Christiane Fellbaum
15:00–15:20A finite-state model of Innu verbal inflection
Loïc Daignault-Pichette and François Lareau
15:20–15:30Morphological Operations in the Persian Verbal System
Marzieh Rabiei and Max Silberztein
 15:50–17:10 Session 4: Training Data and Model Adaptation # %chair1 Saad Ezzini
15:50–16:10A Massive Open-Source Corpus for Valencian: Over 4.7 Billion Tokens for Low-Resource Language Modelling
Yoan Gutiérrez, Juan Pablo Consuegra-Ayala, Robiert Sepúlveda-Torres and Rafael Muñoz Guillena
16:10–16:30LuxIT: A Luxembourgish Instruction Tuning Dataset from Monolingual Seed Data
Julian Valline, Cedric Lothritz, Siwen Guo and Jordi Cabot Sagrera
16:30–16:50Topic Modeling for Moroccan Darija: A Comparative Study of Classical Machine Learning Approaches
Salma Mekaoui, Ilham Chaker, Arsalane Zarghili and Nikola S. Nikolov
16:50–17:10Distil and Evolve: Compact Adaptive Document Classification with Continual Learning for Low-Resource Settings
Mehdi Soufiane, Mourad Hammale, Kawtar Zerhouni, Adil Ahidar-Coutrix and Rachid Benouini

Tuesday September 29, 2026

 09:00–10:00 Session 5: Language Model Behaviour and Evaluation # %chair1 Ernesto Luis Estevanell-Valladares
09:00–09:20Do Large Language Models Style-Shift? Register-Conditioned Stylistic Fronting in AI-Generated Icelandic
Anton Karl Ingason, Johanna Mechler and Lilja Björk Stefánsdóttir
09:20–09:40Benchmarking Large Language Models on Mandarin Proverb Explanation and Contextual Matching
Xiaojing Zhao, Salima Lamsiyah, Emmanuele Chersoni and Han Xu
09:40–10:00LUXDIAG-RAG: Diagnostic Evaluation of Retrieval-Augmented Generation for Luxembourgish Reading Comprehension
Keerthana Murugaraj, Hedi Tebourbi, Christophe Friezas Gonçalves and Salima Lamsiyah
 10:50–11:50 Session 6: Human-Centred NLP and Education # %chair1 Salima Lamsiyah
10:50–11:10Assessment of Human-in-the-Loop Multi-LLMs for Low-Resource Educational Data Expansion
Christophe Friezas Gonçalves, Hedi Tebourbi, Christoph Schommer and Salima Lamsiyah
11:10–11:30Human-Centred Approaches in Low-Resource Languages for Educational Applications and Language Learning: A Systematic Literature Review
Noura El Moussa
11:30–11:50SimuLe-Lux: A Neuro-Symbolic Teaching System for Diagnosing L2 Reading Comprehension
Christophe Friezas Gonçalves, Christoph Schommer and Salima Lamsiyah
 11:50–12:50 Session 7: Arabic and Dialect NLP # %chair1 Salmane Chafik
11:50–12:10Towards Readability Assessment for Under-Resourced Arabic Dialects: A Study of Moroccan Darija
Houdaifa Atou, Nouran Khallaf, Salima Lamsiyah and Ruslan Mitkov
12:10–12:30Why Is Current XAI Not Enough for Arabic NLP? A Critical Survey of the Explainability Gap
Salima Lamsiyah and Ruslan Mitkov
12:30–12:40Benchmarking Speech Foundation Models and LLMs for Sentiment Analysis in Low-Resource Najdi Arabic of Saudi Arabia
Nadia Ghezaiel Hammouda, Maram I. Alharbi and Ruslan Mitkov
12:40–12:50Spatial Entity Extraction Methods from Arabic Texts in the Context of Epidemiological Surveillance: A Comparative Study
Fatima Ezzahra El Houbri, Najlae Idrissi, Mathieu Roche and Sarah Valentin
 14:20–15:30 Session 8: Speech and Spoken-Language Technologies # %chair1 Nadia Ghezaiel
14:20–14:40Building an Aligned Speech Corpus for Northern Russian
Kseniia Protonina and Daniil Ignatev
14:40–15:00Probing Geographic Performance Gaps in Moroccan ASR with ALAMA: Annotated Local Audio of Moroccan Arabic
Avery Cole Kanel, Christian Schuler, Bouazza Laracha, Imrane Lbouhli, Yassine Chaouri, Yusser Al Ghussin and Timo Baumann
15:00–15:20Generator-Guided Amount Recovery for Voice-Based Financial Record-Keeping in Mooré-French Code-Switched Speech
Maimouna Ouattara, El-Hacen Diallo, Fred Philippy, Abdoul Kader Kaboré, Jacques Klein and Tegawendé F. Bissyandé
 15:50–17:00 Session 9: Retrieval, RAG and Document Processing # %chair1 Houdaifa Atou
15:50–16:10LëtzCross: A Cross-Lingual Page-Level Benchmark for Multimodal Retrieval over Luxembourgish Documents
Omar El Bachyr, Fred Philippy, Laura Maria Bernardy, Saad Ezzini, Jacques Klein and Tegawendé Bissyandé
16:10–16:30Improving OCR for a Latvian Pronunciation Dictionary
Viesturs Jūlijs Lasmanis
16:30–16:40LLM-based conversational systems for Sinhala: Investigating model limitations and exploring a framework for agentic retrieval
Malithi P. Alahapperuma, Andreas Vlachidis and Antonis Bikakis
16:40–16:50Clinical Entity Recognition from Electronic Health Records and Linking to Biomedical Knowledge Bases in Low-resource Language Settings
Eno-Martin Lotman, David Hübner, Kris Collins, Svetla Boytcheva and Ivelina Nikolova-Koleva
16:50–17:00Beyond RAG: A Multi-Graph, Multi-Agent, Recursive Retrieval Architecture for Traceable and Source-Grounded Legal Answers
Assia Bouamir, Marie Bonnin, Youssef Al Mouatamid and Jihad Zahir
 Session VID: Video Presentations (on the conference website)
 FLICK: Few-Label Incremental Learning for Low-Resource Dialects
Ali Almutairi, Abdullah Alsuhaibani, Shoaib Jameel, Aditya Joshi, Gelareh Mohammadi and Imran Razzak
 Linguistic Proximity Enables Pivot-Based Machine Translation for an Under-Resourced Tribal Language
Pooja Singh, M Kaab Bin Shahid, Atai Waris Khan, Aryan Kumar Jha and Sandeep Kumar
 Cross-Lingual Hate Speech Detection in Low-Resource Languages: The Case of Persian and Kurdish
Shahin Yousefi, Ernesto Luis Estevanell-Valladares and Ruslan Mitkov
 Is Overconfidence Language-Specific? Cross-Lingual Calibration and Recalibration Transfer in Base and Instruction-Tuned LLMs
Divya Gupta and Ruslan Mitkov
 Knowledge Tracing for Early Childhood Learners with Minimal Telemetry from Low-Connectivity Environments: Evidence from India’s Anganwadi Ecosystem
Badmavasan Kirouchenassamy, Rahul Singh, Vishnu Dev, Meenakshi Garg, Sukhna Sawhney and Sudeep Gowrishankar
 Kuwain 1.5B: An Arabic SLM via Language Injection
Khalil Hennara, Sara Chrouf, Mohamed Motasim Hamed, Zeina Aldallal and Safwan AlModhayan
 A Neuro-Symbolic RAG System for Marine Environmental Law: from Domain Ontology to Knowledge Graphs and Context-Enriched Legal Question Answering
Assoumana Souley Hadiza, Youssef Al Mouatamid, Marie Bonnin and Jihad Zahir
 Evaluating Model-Task Fit in Arabic Word Sense Disambiguation
Yousef Younes, Abdelhalim Hafedh Dahou and Brigitte Mathiak
 Enhancing Urdu ASR with Whisper v3: Fine-Tuning on Latest Datasets and Realistic Multi-Speaker Evaluation with SLM Post-Processing
Zehra Ahmed, Farah Inayat, Zuha Aqib and Sajjad Haider
 Baseer: A Vision-Language Model for Arabic Document-to-Markdown OCR
Khalil Hennara, Muhammad Hreden, Mohamed Motasim Hamed, Ahmad Bastati, Zeina Aldallal, Sara Chrouf and Safwan AlModhayan
 Worldview Annotation for Low-Resource Languages
Tetiana Ilman
 A Multilingual Sentence-Transformer Baseline for Serbian News Topic Classification
Saša Petalinkar, Milica Ikonić Nešić, Ranka Stanković and Jelena Graovac
 Wasm: A Pipeline for Constructing Structured Arabic Interleaved Multimodal Corpora
Khalil Hennara, Ahmad Bastati, Muhammad Hreden, Mohamed Motasim Hamed, Zeina Aldallal, Sara Chrouf and Safwan AlModhayan
 Linguistic Specialization of Arabic in Text Embedding Models via Multi-Teacher Knowledge Distillation
Ahmad Abdelfattah, Zeina Aldallal, Sara Chrouf and Safwan AlModhayan