Ioana Buhnila

Center for Data Science in Humanities, Chosun University, Gwangju, South Korea

prof_pic.jpg

previous website at ATILF

https://perso.atilf.fr/ibuhnila/

As HK Research Professor (post-doctoral position), I work on the Elderly Speech Korea Project (Language Biomarkers and AI in Neurocognitive Disorders), under the supervision of Professor Eon-Suk Ko. My research focuses on analyzing linguistic biomarkers of cognitive impairment in neurocognitive disorders and Alzheimer’s dementia, using open-source NLP and AI tools for Korean long-form speech data.

I am also involved in the research project “Language, Cognition and Society,” where I work on evaluating open-source AI tools, such as Whisper, on recordings of conversations between parents and children in Korean. My research objective is to determine what linguistic measures (lexical, semantic) from the automatic transcripts remain reliable for child language development studies.

Previously, I conducted research on the automatic processing of scientific articles on brain cancer and medical knowlegde simplification using Large Language Models, at ATILF, University of Lorraine-CNRS, under Professor Mathieu Constant’s supervision.

I obtained my PhD in Linguistics and NLP from the University of Strasbourg, with the thesis “An Automatic Method for Building a Paraphrase Corpus” (open-source thesis in French), under the supervision of Professor Amalia Todirascu, University of Strasbourg, and Professor Dan Tufis from the Research Institute for Artificial Intelligence (RACAI).

news

Jun 08, 2026 I gave a talk during the NLP Seminar at NASK National Research Institute, Poland, about the paper “HalluGuard: Evidence-Grounded Small Reasoning Models to Mitigate Hallucinations in Retrieval-Augmented Generation” (ACL Findings 2026). Many thanks to Wojciech Kusa for the invitation!
May 08, 2026 Two papers, “TrackList: Tracing Back Query Linguistic Diversity for Head and Tail Medical Knowledge in Open Large Language Models”, and “Evaluating LLM-as-a-Judge for Medical Term Simplification”, were accepted to BioNLP 2026!🩺
Apr 06, 2026 Our paper “HalluGuard: Evidence-Grounded Small Reasoning Models to Mitigate Hallucinations in Retrieval-Augmented Generation” was accepted to ACL 2026 Findings!🌟
Mar 16, 2026 Our paper “Romance Reflexive Constructions Revisited” was accepted at the Universal Dependencies Workshop 2026 @LREC 2026!🎉
Feb 13, 2026 Our paper, “Confabulations from acl publications (cap): A dataset for scientific hallucination detection” was accepted to LREC 2026!🎊

selected publications

  1. pRAGe
    Retrieve, Generate, Evaluate: A Case Study for Medical Paraphrases Generation with Small Language Models
    Ioana Buhnila, Aman Sinha, and Mathieu Constant
    In Proceedings of the 1st Workshop on Towards Knowledgeable Language Models (KnowLLM 2024), ACL 2024, Bangkok, Thailand, August 2024, 2024
  2. COMW
    Chain-of-MetaWriting: Linguistic and Textual Analysis of How Small Language Models Write Young Students Texts
    Ioana Buhnila, Georgeta Cislaru, and Amalia Todirascu
    In Proceedings of the First Workshop on Writing Aids at the Crossroads of AI, Cognitive Science and NLP (WRAICOGS 2025), COLING 2025, Abu Dhabi, UAE, January 2025, 2025
  3. HalluGuard
    HalluGuard: Evidence-Grounded Small Reasoning Models to Mitigate Hallucinations in Retrieval-Augmented Generation
    Loris Bergeron, Ioana Buhnila, Jérôme François, and Radu State
    Findings of the Association for Computational Linguistics: ACL 2026, San Diego, California, 2026
  4. CAP
    Confabulations from acl publications (cap): A dataset for scientific hallucination detection
    Federica Gamba, Aman Sinha, Timothee Mickus, Raul Vazquez, Patanjali Bhamidipati, Claudio Savelli, Ahana Chattopadhyay, Laura A Zanella, Yash Kankanampati, Binesh Arakkal Remesh, and others
    Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026), 2026
  5. Langages
    Approche interdisciplinaire de la notion de reformulation sous-phrastique médicale: entre la linguistique, le TAL et l’IA générative
    Ioana Buhnila
    Langages, 2025
  6. ChildDirectedSpeech
    Classifying Child-Directed Speech and Shared Book Reading from LENA: Language-Specific Modeling and Temporal Resolution Effects
    Sukhwan Jung, Jun Ho Chai, Ioana Buhnila, and Eon-Suk Ko
    Available at SSRN 6359727, 2026