RANLP 2023 Proceedings Home | RANLP 2023 Website

Proceedings of the 14th International Conference on Recent Advances in Natural Language Processing

Chairs
Programme Committee Chair:
Ruslan Mitkov, Lancaster University, UK

Organising Committee Chair:
Galia Angelova, IICT, Bulgarian Academy of Sciences, Bulgaria

Full proceedings volume (PDF)
Program and Author index (HTML)
Bibliography (BibTeX)
Live website


PROGRAM

 Bipol: Multi-Axes Evaluation of Bias with Explainability in Benchmark Datasets
Tosin Adewumi, Isabella Södergren, Lama Alkhaled, sana sabah al-azzawi, Foteini Simistira Liwicki and Marcus Liwicki
 Automatically Generating Hindi Wikipedia Pages Using Wikidata as a Knowledge Graph: A Domain-Specific Template Sentences Approach
Aditya Agarwal and Radhika Mamidi
 Cross-lingual Classification of Crisis-related Tweets Using Machine Translation
Shareefa A. Al Amer, Mark Lee and Phillip Smith
 Lexicon-Driven Automatic Sentence Generation for the Skills Section in a Job Posting
Vera Aleksic, Mona Brems, Anna Mathes and Theresa Bertele
 Multilingual Racial Hate Speech Detection Using Transfer Learning
Abinew Ali Ayele, Skadi Dinter, Seid Muhie Yimam and Chris Biemann
 Exploring Amharic Hate Speech Data Collection and Classification Approaches
Abinew Ali Ayele, Seid Muhie Yimam, Tadesse Destaw Belay, Tesfa Tegegne Asfaw and Chris Biemann
 Bhojpuri WordNet: Problems in Translating Hindi Synsets into Bhojpuri
Imran Ali and Praveen Gatla
 3D-EX: A Unified Dataset of Definitions and Dictionary Examples
Fatemah Yousef Almeman, Hadi Sheikhi and Luis Espinosa Anke
 Are You Not moved? Incorporating Sensorimotor Knowledge to Improve Metaphor Detection
Ghadi Alnafesah, Phillip Smith and Mark Lee
 HAQA and QUQA: Constructing Two Arabic Question-Answering Corpora for the Quran and Hadith
Sarah Alnefaie, Eric Atwell and Mohammad Ammar Alsalka
 ConfliBERT-Arabic: A Pre-trained Arabic Language Model for Politics, Conflicts and Violence
Sultan Alsarra, Luay Abdeljaber, Wooseong Yang, Niamat Zawad, Latifur Khan, Patrick Brandt, Javier Osorio and Vito D’Orazio
 A Review in Knowledge Extraction from Knowledge Bases
Fabio Antonio Yanez, Andrés Montoyo, Yoan Gutierrez, Rafael Muñoz and Armando Suarez
 Evaluating of Large Language Models in Relationship Extraction from Unstructured Data: Empirical Study from Holocaust Testimonies
Isuri Anuradha, Le An Ha, Ruslan Mitkov and Vinita Nahar
 Impact of Emojis on Automatic Analysis of Individual Emotion Categories
Ratchakrit Arreerard and Scott Piao
 Was That a Question? Automatic Classification of Discourse Meaning in Spanish
Santiago Arróniz and Sandra Kübler
 Designing the LECOR Learner Corpus for Romanian
Ana Maria Barbu, Elena Irimia, Carmen Mîrzea Vasile and Vasile Păiș
 Non-Parametric Memory Guidance for Multi-Document Summarization
Florian Baud and Alex Aussem
 Beyond Information: Is ChatGPT Empathetic Enough?
Ahmed Belkhir and Fatiha Sadat
 Using Wikidata for Enhancing Compositionality in Pretrained Language Models
Meriem Beloucif, Mihir Bansal and Chris Biemann
 Multimodal Learning for Accurate Visual Question Answering: An Attention-Based Approach
Jishnu Bhardwaj, Anurag Balakrishnan, Satyam Pathak, Ishan Unnarkar, Aniruddha Gawande and Benyamin Ahmadnia
 Generative Models For Indic Languages: Evaluating Content Generation Capabilities
Savita Bhat, Vasudeva Varma and Niranjan Pedanekar
 Measuring Spurious Correlation in Classification: "Clever Hans" in Translationese
Angana Borah, Daria Pylypenko, Cristina España-Bonet and Josef van Genabith
 WIKITIDE: A Wikipedia-Based Timestamped Definition Pairs Dataset
Hsuvas Borkakoty and Luis Espinosa Anke
 BERTabaporu: Assessing a Genre-Specific Language Model for Portuguese NLP
Pablo Botton Costa, Matheus Camasmie Pavan, Wesley Ramos Santos, Samuel Caetano Silva and Ivandré Paraboni
 Comparison of Multilingual Entity Linking Approaches
Ivelina Bozhinova and Andrey T. Tagarev
 Automatic Extraction of the Romanian Academic Word List: Data and Methods
Ana-Maria Bucur, Andreea Dincă, Madalina Chitez and Roxana Rogobete
 Stance Prediction from Multimodal Social Media Data
Lais Carraro Leme Cavalheiro, Matheus Camasmie Pavan and Ivandré Paraboni
 From Stigma to Support: A Parallel Monolingual Corpus and NLP Approach for Neutralizing Mental Illness Bias
Mason Choey
 BB25HLegalSum: Leveraging BM25 and BERT-Based Clustering for the Summarization of Legal Documents
Leonardo Bonalume de Andrade and Karin Becker
 SSSD: Leveraging Pre-trained Models and Semantic Search for Semi-supervised Stance Detection
André Mediote de Sousa and Karin Becker
 Detecting Text Formality: A Study of Text Classification Approaches
Daryna Dementieva, Nikolay Babakov and Alexander Panchenko
 Developing a Multilingual Corpus of Wikipedia Biographies
Hannah Devinney, Anton Eklund, Igor Ryazanov and Jingwen Cai
 A Computational Analysis of the Voices of Shakespeare’s Characters
Liviu P. Dinu and Ana Sabina Uban
 Source Code Plagiarism Detection with Pre-Trained Model Embeddings and Automated Machine Learning
Fahad Ebrahim and Mike Joy
 Identifying Semantic Argument Types in Predication and Copredication Contexts: A Zero-Shot Cross-Lingual Approach
Deniz Ekin Yavas, Laura Kallmeyer, Rainer Osswald, Elisabetta Jezek, Marta Ricchiardi and Long Chen
 A Review of Research-Based Automatic Text Simplification Tools
Isabel Espinosa-Zaragoza, José Abreu-Salas, Elena Lloret, Paloma Moreda and Manuel Palomar
 Vocab-Expander: A System for Creating Domain-Specific Vocabularies Based on Word Embeddings
Michael Faerber and Nicholas Popovic
 On the Generalization of Projection-Based Gender Debiasing in Word Embedding
Elisabetta Fersini, Antonio Candelieri and Lorenzo Pastore
 Mapping Explicit and Implicit Discourse Relations between the RST-DT and the PDTB 3.0
Nelson Filipe Costa, Nadia Sheikh and Leila Kosseim
 Bigfoot in Big Tech: Detecting Out of Domain Conspiracy Theories
Matthew Fort, Zuoyu Tian, Elizabeth Gabel, Nina Georgiades, Noah Sauer, Daniel Dakota and Sandra Kübler
 Deep Learning Approaches to Detecting Safeguarding Concerns in Schoolchildren’s Online Conversations
Emma Franklin and Tharindu Ranasinghe
 On the Identification and Forecasting of Hate Speech in Inceldom
Paolo Gajo, Arianna Muti, Katerina Korre, Silvia Bernardini and Alberto Barrón-Cedeño
 T2KG: Transforming Multimodal Document to Knowledge Graph
Santiago Galiano, Rafael Muñoz, Yoan Gutiérrez, Andrés Montoyo, Jose Ignacio Abreu and Luis Alfonso Ureña
 !Translate: When You Cannot Cook Up a Translation, Explain
Federico Garcea, Margherita Martinelli, Maja Milicević Petrović and Alberto Barrón-Cedeño
 An Evaluation of Source Factors in Concatenation-Based Context-Aware Neural Machine Translation
Harritxu Gete and Thierry Etchegoyhen
 Lessons Learnt from Linear Text Segmentation: a Fair Comparison of Architectural and Sentence Encoding Strategies for Successful Segmentation
Iacopo Ghinassi, Lin Wang, Chris Newell and Matthew Purver
 Student’s t-Distribution: On Measuring the Inter-Rater Reliability When the Observations are Scarce
Serge Gladkoff, Lifeng Han and Goran Nenadic
 Data Augmentation for Fake News Detection by Combining Seq2seq and NLI
Anna Glazkova
 Exploring Unsupervised Semantic Similarity Methods for Claim Verification in Health Care News Articles
Vishwani Gupta, Astrid Viciano, Holger Wormer and Najmehsadat Mousavinezhad
 AlphaMWE-Arabic: Arabic Edition of Multilingual Parallel Corpora with Multiword Expression Annotations
najet hadj mohamed, Malak Rassem, Lifeng Han and Goran Nenadic
 Performance Analysis of Arabic Pre-trained Models on Named Entity Recognition Task
Abdelhalim Hafedh Dahou, Mohamed Amine Cheragui and Ahmed Abdelali
 Discourse Analysis of Argumentative Essays of English Learners Based on CEFR Level
Blaise Hanel and Leila Kosseim
 Improving Translation Quality for Low-Resource Inuktitut with Various Preprocessing Techniques
Mathias Hans Erik Stenlund, Mathilde Nanni, Micaella Bruton and Meriem Beloucif
 Enriched Pre-trained Transformers for Joint Slot Filling and Intent Detection
Momchil Hardalov, Ivan K. Koychev and Preslav Nakov
 Unimodal Intermediate Training for Multimodal Meme Sentiment Classification
Muzhaffar Hazman, Susan McKeever and Josephine Griffith
 Explainable Event Detection with Event Trigger Identification as Rationale Extraction
Hansi Hettiarachchi and Tharindu Ranasinghe
 Clinical Text Classification to SNOMED CT Codes Using Transformers Trained on Linked Open Medical Ontologies
Anton Hristov, Petar Ivanov, Anna Aksenova, Tsvetan Asamov, Pavlin Gyurov, Todor Primov and Svetla Boytcheva
 Towards a Consensus Taxonomy for Annotating Errors in Automatically Generated Text
Rudali Huidrom and Anya Belz
 Uncertainty Quantification of Text Classification in a Multi-Label Setting for Risk-Sensitive Systems
Jinha Hwang, Carol Gudumotu and Benyamin Ahmadnia
 Pretraining Language- and Domain-Specific BERT on Automatically Translated Text
Tatsuya Ishigaki, Yui Uehara, Goran Topić and Hiroya Takamura
 Categorising Fine-to-Coarse Grained Misinformation: An Empirical Study of the COVID-19 Infodemic
Ye Jiang, Xingyi Song, Carolina Scarton, Iknoor Singh, Ahmet Aker and Kalina Bontcheva
 Bridging the Gap between Subword and Character Segmentation in Pretrained Language Models
Shun Kiyono, Sho Takase, Shengzhe Li and Toshinori Sato
 Evaluating Data Augmentation for Medication Identification in Clinical Notes
Jordan C. Koontz, Maite Oronoz and Alicia Pérez
 Advancing Topical Text Classification: A Novel Distance-Based Method with Contextual Embeddings
Andriy Kosar, Guy De Pauw and Walter Daelemans
 Taxonomy-Based Automation of Prior Approval Using Clinical Guidelines
Saranya Krishnamoorthy and Ayush Singh
 Simultaneous Interpreting as a Noisy Channel: How Much Information Gets Through
Maria Kunilovskaya, Heike Przybyl, Ekaterina Lapshinova-Koltunski and Elke Teich
 Challenges of GPT-3-Based Conversational Agents for Healthcare
Fabian Lechner, Allison Claire Lahnala, Charles Welch and Lucie Flek
 Noisy Self-Training with Data Augmentations for Offensive and Hate Speech Detection Tasks
João A. Leite, Carolina Scarton and Diego Furtado Silva
 A Practical Survey on Zero-Shot Prompt Design for In-Context Learning
Yinheng Li
 Classifying COVID-19 Vaccine Narratives
Yue Li, Carolina Scarton, Xingyi Song and Kalina Bontcheva
 Sign Language Recognition and Translation: A Multi-Modal Approach Using Computer Vision and Natural Language Processing
Jacky Li, Jaren Gerdes, James Gojit, Austin Tao, Samyak Katke, Kate Nguyen and Benyamin Ahmadnia
 Classification-Aware Neural Topic Model Combined with Interpretable Analysis - for Conflict Classification
Tianyu Liang, Yida Mu, Soonho Kim, Darline Kuate, Julie Lang, Rob Vos and Xingyi Song
 Data Augmentation for Fake Reviews Detection
Ming Liu and Massimo Poesio
 Coherent Story Generation with Structured Knowledge
Congda Ma, Kotaro Funakoshi, Kiyoaki Shirai and Manabu Okumura
 Studying Common Ground Instantiation Using Audio, Video and Brain Behaviours: The BrainKT Corpus
Eliot Maës, Thierry Legou, Leonor Becerra-Bonache and Philippe Blache
 Reading between the Lines: Information Extraction from Industry Requirements
Ole Magnus Holter and Basil Ell
 Transformer-Based Language Models for Bulgarian
Iva Marinova, Kiril Simov and Petya Osenova
 Multi-task Ensemble Learning for Fake Reviews Detection and Helpfulness Prediction: A Novel Approach
Alimuddin Melleng, Anna Jurek-Loughrey and Deepak P
 Data Fusion for Better Fake Reviews Detection
Alimuddin Melleng, Anna Jurek-Loughrey and Deepak P
 Dimensions of Quality: Contrasting Stylistic vs. Semantic Features for Modelling Literary Quality in 9,000 Novels
Pascale Feldkamp Moreira and Yuri Bizzoni
 BanglaBait: Semi-Supervised Adversarial Approach for Clickbait Detection on Bangla Clickbait Dataset
Md. Motahar Mahtab, Monirul Haque, Mehedi Hasan and Farig Sadeque
 TreeSwap: Data Augmentation for Machine Translation via Dependency Subtree Swapping
Attila Nagy, Dorina Lakatos, Botond Barta and Judit Ács
 Automatic Assessment Of Spoken English Proficiency Based on Multimodal and Multitask Transformers
Kamel Nebhi and György Szaszák
 Medical Concept Mention Identification in Social Media Posts Using a Small Number of Sample References
Vasudevan Nedumpozhimana, Sneha Rautmare, Meegan Gower, Nishtha Jain, Maja Popović, Patricia Buffini and John D. Kelleher
 Context-Aware Module Selection in Modular Dialog Systems
Jan Nehring, René Marcel Berk and Stefan Hillmann
 Human Value Detection from Bilingual Sensory Product Reviews
Boyu Niu, Céline Manetta and Frédérique Segond
 Word Sense Disambiguation for Automatic Translation of Medical Dialogues into Pictographs
Magali Norré, Rémi Cardon, Vincent Vandeghinste and Thomas François
 A Research-Based Guide for the Creation and Deployment of a Low-Resource Machine Translation System
John E. Ortega and Kenneth Ward Church
 MQDD: Pre-training of Multimodal Question Duplicity Detection for Software Engineering Domain
Jan Pasek, Jakub Sido, Miloslav Konopik and Ondrej Prazak
 Forming Trees with Treeformers
Nilay Patel and Jeffrey Flanigan
 Evaluating Unsupervised Hierarchical Topic Models Using a Labeled Dataset
Judicael Poumay and Ashwin Ittoo
 HTMOT: Hierarchical Topic Modelling over Time
Judicael Poumay and Ashwin Ittoo
 Multilingual Continual Learning Approaches for Text Classification
Karan Praharaj and Irina Matveeva
 Can Model Fusing Help Transformers in Long Document Classification? An Empirical Study
Damith Premasiri, Tharindu Ranasinghe and Ruslan Mitkov
 Deep Learning Methods for Identification of Multiword Flower and Plant Names
Damith Premasiri, Amal Haddad Haddad, Tharindu Ranasinghe and Ruslan Mitkov
 Improving Aspect-Based Sentiment with End-to-End Semantic Role Labeling Model
Pavel Přibáň and Ondrej Prazak
 huPWKP: A Hungarian Text Simplification Corpus
Noémi Prótár and Dávid Márk Nemeskey
 Topic Modeling Using Community Detection on a Word Association Graph
Mahfuzur Rahman Chowdhury, Intesur Ahmed, Farig Sadeque and Muhammad Nur Yanhaona
 Exploring Techniques to Detect and Mitigate Non-Inclusive Language Bias in Marketing Communications Using a Dictionary-Based Approach
Bharathi Raja Chakravarthi, Prasanna Kumar Kumaresan, Rahul Ponnusamy, John P. McCrae, Michaela Comerford, Jay Megaro, Deniz Keles and Last Feremenga
 Does the "Most Sinfully Decadent Cake Ever" Taste Good? Answering Yes/No Questions from Figurative Contexts
Geetanjali Rakshit and Jeffrey Flanigan
 Modeling Easiness for Training Transformers with Curriculum Learning
Leonardo Ranaldi, Giulia Pucci and Fabio Massimo Zanzotto
 The Dark Side of the Language: Pre-trained Transformers in the DarkNet
Leonardo Ranaldi, Aria Nourbakhsh, Elena Sofia Ruzzetti, Arianna Patrizi, Dario Onorati, Michele Mastromattei, Francesca Fallucchi and Fabio Massimo Zanzotto
 PreCog: Exploring the Relation between Memorization and Performance in Pre-trained Language Models
Leonardo Ranaldi, Elena Sofia Ruzzetti and Fabio Massimo Zanzotto
 Publish or Hold? Automatic Comment Moderation in Luxembourgish News Articles
Tharindu Ranasinghe, Alistair Plum, Christoph Purschke and Marcos Zampieri
 Cross-Lingual Speaker Identification for Indian Languages
Amaan Rizvi, Anupam Jamatia, Dwijen Rudrapal, Kunal Chakma and Björn Gambäck
 ‘ChemXtract’ A System for Extraction of Chemical Events from Patent Documents
Pattabhi RK Rao and Sobha Lalitha Devi
 Mind the User! Measures to More Accurately Evaluate the Practical Value of Active Learning Strategies
Julia Romberg
 Event Annotation and Detection in Kannada-English Code-Mixed Social Media Data
Sumukh S, Abhinav Appidi and Manish Shrivastava
 Three Approaches to Client Email Topic Classification
Branislava Šandrih Todorović, Katarina Josipović and Jurij Kodre
 Exploring Abstractive Text Summarisation for Podcasts: A Comparative Study of BART and T5 Models
Parth Saxena and Mo El-Haj
 Exploring the Landscape of Natural Language Processing Research
Tim Schopf, Karim Arabi and Florian Matthes
 Efficient Domain Adaptation of Sentence Embeddings Using Adapters
Tim Schopf, Dennis N. Schneider and Florian Matthes
 AspectCSE: Sentence Embeddings for Aspect-Based Semantic Textual Similarity Using Contrastive Learning and Structured Knowledge
Tim Schopf, Emanuel Gerber, Malte Ostendorff and Florian Matthes
 Tackling the Myriads of Collusion Scams on YouTube Comments of Cryptocurrency Videos
Sadat Shahriar and Arjun Mukherjee
 Exploring Deceptive Domain Transfer Strategies: Mitigating the Differences among Deceptive Domains
Sadat Shahriar, Arjun Mukherjee and Omprakash Gnawali
 Party Extraction from Legal Contract Using Contextualized Span Representations of Parties
Sanjeepan Sivapiran, Charangan Vasantharajan and Uthayasanker Thayasivam
 From Fake to Hyperpartisan News Detection Using Domain Adaptation
Răzvan-Alexandru Smădu, Sebastian-Vasile Echim, Dumitru-Clementin Cercel, Iuliana Marin and Florin Pop
 Prompt-Based Approach for Czech Sentiment Analysis
Jakub Šmíd and Pavel Přibáň
 Measuring Gender Bias in Natural Language Processing: Incorporating Gender-Neutral Linguistic Forms for Non-Binary Gender Identities in Abusive Speech Detection
Nasim Sobhani, Kinshuk Sengupta and Sarah Jane Delany
 LeSS: A Computationally-Light Lexical Simplifier for Spanish
Sanja Stajner, Daniel Ibanez and Horacio Saggion
 Hindi to Dravidian Language Neural Machine Translation Systems
Vijay Sundar Ram and Sobha Lalitha Devi
 Looking for Traces of Textual Deepfakes in Bulgarian on Social Media
Irina Temnikova, Iva Marinova, Silvia Gargova, Ruslana Margova and Ivan Koychev
 Propaganda Detection in Russian Telegram Posts in the Scope of the Russian Invasion of Ukraine
Natalia Vanetik, Marina Litvak, Egor Reviakin and Margarita Tiamanova
 Auto-Encoding Questions with Retrieval Augmented Decoding for Unsupervised Passage Retrieval and Zero-Shot Question Generation
Stalin Varanasi, Muhammad Umer Tariq Butt and Guenter Neumann
 NoHateBrazil: A Brazilian Portuguese Text Offensiveness Analysis System
Francielle Vargas, Isabelle Carvalho, Wolfgang S. Schmeisser-Nieto, Fabrício Benevenuto and Thiago Alexandre Salgueiro Pardo
 Socially Responsible Hate Speech Detection: Can Classifiers Reflect Social Stereotypes?
Francielle Vargas, Isabelle Carvalho, Ali Hürriyetoğlu, Thiago Alexandre Salgueiro Pardo and Fabrício Benevenuto
 Predicting Sentence-Level Factuality of News and Bias of Media Outlets
Francielle Vargas, Kokil Jaidka, Thiago Alexandre Salgueiro Pardo and Fabrício Benevenuto
 Classification of US Supreme Court Cases Using BERT-Based Techniques
Shubham Vatsal, Adam Meyers and John E. Ortega
 Kāraka-Based Answer Retrieval for Question Answering in Indic Languages
Devika A. Verma, Ramprasad S. Joshi, Aiman A. Shivani and Rohan D. Gupta
 Comparative Analysis of Named Entity Recognition in the Dungeons and Dragons Domain
gayashan WAG Weerasundara and Nisansa de Silva
 Comparative Analysis of Anomaly Detection Algorithms in Text Data
Yizhou Xu, Kata Gábor, Jérôme Milleret and Frédérique Segond
 Poetry Generation Combining Poetry Theme Labels Representations
Yingyu Yan, Dongzhen Wen, Liang Yang, Dongyu Zhang and Hongfei LIN
 Evaluating Generative Models for Graph-to-Text Generation
Shuzhou Yuan and Michael Faerber
 Microsyntactic Unit Detection Using Word Embedding Models: Experiments on Slavic Languages
Iuliia Zaitova, Irina Stenger and Tania Avgustinova
 Systematic TextRank Optimization in Extractive Summarization
Morris Zieve, Anthony Gregor, Frederik Juul Stokbaek, Hunter Lewis, Ellis Marie Mendoza and Benyamin Ahmadnia