Chairs
Lamsiyah, Salima
Ranasinghe, Tharindu
Ezzini, Saad
Estevanell-Valladares, Ernesto Luis
Mitkov, Ruslan
| LaTeLL 2026 Proceedings Home | LaTeLL WEBSITE | Bulgarian Association for Computational Linguistics WEBSITE |
|
PROGRAM
Monday September 28, 2026 | |
| 10:50–11:50 Session 1: Corpora, Annotation and Resource Development # %chair1 Yoan Gutiérrez | |
| 10:50–11:10 | A Kazakh–Russian Corpus of Child-Directed Language for Low-Resource Languages Albina Mukusheva, Achille Fusco and Cristiano Chesi |
| 11:10–11:30 | Data Curation, Annotation Quality, and Error Patterns in Old English Automatic Lemmatisation Javier Martín Arista |
| 11:30–11:40 | Creating the RozMuz Corpus: Applying Ethics and Technology in Language Research Alicja Helena Derych, Bartłomiej Alberski, Hubert Jankowski and Paweł Dembowski |
| 11:40–11:50 | Towards a Universal Dependencies Treebank for Amazigh: A Tarifit Pilot Study Azzeddine Afrouni, Fadoua Ataa Allah and Jamal Abarnous |
| 11:50–12:30 Session 2: Machine Translation and Cross-Lingual Transfer # %chair1 Salima Lamsiyah | |
| 11:50–12:10 | Arabic Dialect-to-MSA Translation: A Comparative Evaluation of Large Language Models and Neural Machine Translation Maram I. Alharbi, Jihad R’baiti and Ruslan Mitkov |
| 12:10–12:20 | Cross-Lingual Transfer from Portuguese to Nheengatu: Evidence for Contact-Induced Convergence as a Computational Bridge Rafael Macario Fernandes |
| 12:20–12:30 | Retrieval-Augmented Machine Translation for Bohairic Coptic: A Pilot Case Study So Miyagawa |
| 14:20–15:30 Session 3: Morphology, Syntax and Formal Modelling # %chair1 Javier Martín Arista | |
| 14:20–14:40 | Automating Mwotlap morphology using formal grammar Fabio Meroni, Alexandre François and Max Silberztein |
| 14:40–15:00 | Dependency Parsing Across the Resource Spectrum: Evaluating Architectures on High and Low-Resource Languages Kevin Guan, Happy Buzaaba and Christiane Fellbaum |
| 15:00–15:20 | A finite-state model of Innu verbal inflection Loïc Daignault-Pichette and François Lareau |
| 15:20–15:30 | Morphological Operations in the Persian Verbal System Marzieh Rabiei and Max Silberztein |
| 15:50–17:10 Session 4: Training Data and Model Adaptation # %chair1 Saad Ezzini | |
| 15:50–16:10 | A Massive Open-Source Corpus for Valencian: Over 4.7 Billion Tokens for Low-Resource Language Modelling Yoan Gutiérrez, Juan Pablo Consuegra-Ayala, Robiert Sepúlveda-Torres and Rafael Muñoz Guillena |
| 16:10–16:30 | LuxIT: A Luxembourgish Instruction Tuning Dataset from Monolingual Seed Data Julian Valline, Cedric Lothritz, Siwen Guo and Jordi Cabot Sagrera |
| 16:30–16:50 | Topic Modeling for Moroccan Darija: A Comparative Study of Classical Machine Learning Approaches Salma Mekaoui, Ilham Chaker, Arsalane Zarghili and Nikola S. Nikolov |
| 16:50–17:10 | Distil and Evolve: Compact Adaptive Document Classification with Continual Learning for Low-Resource Settings Mehdi Soufiane, Mourad Hammale, Kawtar Zerhouni, Adil Ahidar-Coutrix and Rachid Benouini |
Tuesday September 29, 2026 | |
| 09:00–10:00 Session 5: Language Model Behaviour and Evaluation # %chair1 Ernesto Luis Estevanell-Valladares | |
| 09:00–09:20 | Do Large Language Models Style-Shift? Register-Conditioned Stylistic Fronting in AI-Generated Icelandic Anton Karl Ingason, Johanna Mechler and Lilja Björk Stefánsdóttir |
| 09:20–09:40 | Benchmarking Large Language Models on Mandarin Proverb Explanation and Contextual Matching Xiaojing Zhao, Salima Lamsiyah, Emmanuele Chersoni and Han Xu |
| 09:40–10:00 | LUXDIAG-RAG: Diagnostic Evaluation of Retrieval-Augmented Generation for Luxembourgish Reading Comprehension Keerthana Murugaraj, Hedi Tebourbi, Christophe Friezas Gonçalves and Salima Lamsiyah |
| 10:50–11:50 Session 6: Human-Centred NLP and Education # %chair1 Salima Lamsiyah | |
| 10:50–11:10 | Assessment of Human-in-the-Loop Multi-LLMs for Low-Resource Educational Data Expansion Christophe Friezas Gonçalves, Hedi Tebourbi, Christoph Schommer and Salima Lamsiyah |
| 11:10–11:30 | Human-Centred Approaches in Low-Resource Languages for Educational Applications and Language Learning: A Systematic Literature Review Noura El Moussa |
| 11:30–11:50 | SimuLe-Lux: A Neuro-Symbolic Teaching System for Diagnosing L2 Reading Comprehension Christophe Friezas Gonçalves, Christoph Schommer and Salima Lamsiyah |
| 11:50–12:50 Session 7: Arabic and Dialect NLP # %chair1 Salmane Chafik | |
| 11:50–12:10 | Towards Readability Assessment for Under-Resourced Arabic Dialects: A Study of Moroccan Darija Houdaifa Atou, Nouran Khallaf, Salima Lamsiyah and Ruslan Mitkov |
| 12:10–12:30 | Why Is Current XAI Not Enough for Arabic NLP? A Critical Survey of the Explainability Gap Salima Lamsiyah and Ruslan Mitkov |
| 12:30–12:40 | Benchmarking Speech Foundation Models and LLMs for Sentiment Analysis in Low-Resource Najdi Arabic of Saudi Arabia Nadia Ghezaiel Hammouda, Maram I. Alharbi and Ruslan Mitkov |
| 12:40–12:50 | Spatial Entity Extraction Methods from Arabic Texts in the Context of Epidemiological Surveillance: A Comparative Study Fatima Ezzahra El Houbri, Najlae Idrissi, Mathieu Roche and Sarah Valentin |
| 14:20–15:30 Session 8: Speech and Spoken-Language Technologies # %chair1 Nadia Ghezaiel | |
| 14:20–14:40 | Building an Aligned Speech Corpus for Northern Russian Kseniia Protonina and Daniil Ignatev |
| 14:40–15:00 | Probing Geographic Performance Gaps in Moroccan ASR with ALAMA: Annotated Local Audio of Moroccan Arabic Avery Cole Kanel, Christian Schuler, Bouazza Laracha, Imrane Lbouhli, Yassine Chaouri, Yusser Al Ghussin and Timo Baumann |
| 15:00–15:20 | Generator-Guided Amount Recovery for Voice-Based Financial Record-Keeping in Mooré-French Code-Switched Speech Maimouna Ouattara, El-Hacen Diallo, Fred Philippy, Abdoul Kader Kaboré, Jacques Klein and Tegawendé F. Bissyandé |
| 15:50–17:00 Session 9: Retrieval, RAG and Document Processing # %chair1 Houdaifa Atou | |
| 15:50–16:10 | LëtzCross: A Cross-Lingual Page-Level Benchmark for Multimodal Retrieval over Luxembourgish Documents Omar El Bachyr, Fred Philippy, Laura Maria Bernardy, Saad Ezzini, Jacques Klein and Tegawendé Bissyandé |
| 16:10–16:30 | Improving OCR for a Latvian Pronunciation Dictionary Viesturs Jūlijs Lasmanis |
| 16:30–16:40 | LLM-based conversational systems for Sinhala: Investigating model limitations and exploring a framework for agentic retrieval Malithi P. Alahapperuma, Andreas Vlachidis and Antonis Bikakis |
| 16:40–16:50 | Clinical Entity Recognition from Electronic Health Records and Linking to Biomedical Knowledge Bases in Low-resource Language Settings Eno-Martin Lotman, David Hübner, Kris Collins, Svetla Boytcheva and Ivelina Nikolova-Koleva |
| 16:50–17:00 | Beyond RAG: A Multi-Graph, Multi-Agent, Recursive Retrieval Architecture for Traceable and Source-Grounded Legal Answers Assia Bouamir, Marie Bonnin, Youssef Al Mouatamid and Jihad Zahir |
| Session VID: Video Presentations (on the conference website) | |
| FLICK: Few-Label Incremental Learning for Low-Resource Dialects Ali Almutairi, Abdullah Alsuhaibani, Shoaib Jameel, Aditya Joshi, Gelareh Mohammadi and Imran Razzak | |
| Linguistic Proximity Enables Pivot-Based Machine Translation for an Under-Resourced Tribal Language Pooja Singh, M Kaab Bin Shahid, Atai Waris Khan, Aryan Kumar Jha and Sandeep Kumar | |
| Cross-Lingual Hate Speech Detection in Low-Resource Languages: The Case of Persian and Kurdish Shahin Yousefi, Ernesto Luis Estevanell-Valladares and Ruslan Mitkov | |
| Is Overconfidence Language-Specific? Cross-Lingual Calibration and Recalibration Transfer in Base and Instruction-Tuned LLMs Divya Gupta and Ruslan Mitkov | |
| Knowledge Tracing for Early Childhood Learners with Minimal Telemetry from Low-Connectivity Environments: Evidence from India’s Anganwadi Ecosystem Badmavasan Kirouchenassamy, Rahul Singh, Vishnu Dev, Meenakshi Garg, Sukhna Sawhney and Sudeep Gowrishankar | |
| Kuwain 1.5B: An Arabic SLM via Language Injection Khalil Hennara, Sara Chrouf, Mohamed Motasim Hamed, Zeina Aldallal and Safwan AlModhayan | |
| A Neuro-Symbolic RAG System for Marine Environmental Law: from Domain Ontology to Knowledge Graphs and Context-Enriched Legal Question Answering Assoumana Souley Hadiza, Youssef Al Mouatamid, Marie Bonnin and Jihad Zahir | |
| Evaluating Model-Task Fit in Arabic Word Sense Disambiguation Yousef Younes, Abdelhalim Hafedh Dahou and Brigitte Mathiak | |
| Enhancing Urdu ASR with Whisper v3: Fine-Tuning on Latest Datasets and Realistic Multi-Speaker Evaluation with SLM Post-Processing Zehra Ahmed, Farah Inayat, Zuha Aqib and Sajjad Haider | |
| Baseer: A Vision-Language Model for Arabic Document-to-Markdown OCR Khalil Hennara, Muhammad Hreden, Mohamed Motasim Hamed, Ahmad Bastati, Zeina Aldallal, Sara Chrouf and Safwan AlModhayan | |
| Worldview Annotation for Low-Resource Languages Tetiana Ilman | |
| A Multilingual Sentence-Transformer Baseline for Serbian News Topic Classification Saša Petalinkar, Milica Ikonić Nešić, Ranka Stanković and Jelena Graovac | |
| Wasm: A Pipeline for Constructing Structured Arabic Interleaved Multimodal Corpora Khalil Hennara, Ahmad Bastati, Muhammad Hreden, Mohamed Motasim Hamed, Zeina Aldallal, Sara Chrouf and Safwan AlModhayan | |
| Linguistic Specialization of Arabic in Text Embedding Models via Multi-Teacher Knowledge Distillation Ahmad Abdelfattah, Zeina Aldallal, Sara Chrouf and Safwan AlModhayan | |