Conference Presentation
Classifying Word-Level Variants in Biblical Hebrew Manuscripts: New Computational Tools for Textual Criticism
Iglika Nikolova-Stoupak
Research Engineer in Natural Language Processing
LORIA, Université de Lorraine / CNRS / Inria
PhD Candidate in Computational Linguistics, Sorbonne Université
France
Presentation Date & Time
10 November 2026
17:00 CET (Serbia time)
Conference
New Discoveries and New Directions in Biblical, Hebrew, and Theological Studies
European Hebrew Journal International Online Conference
Format
Live online via Zoom
Sign Up for Conference Updates by Email
By signing up, you will receive:
- • announcements of newly confirmed international speakers
- • newly announced presentation topics
- • important programme updates and conference news
- • Early Bird reminders before ticket prices increase
- • important registration information and deadlines
- • once a month, a carefully selected digest of important new scientific discoveries and research developments in 2026, especially in Biblical Studies, Hebrew, archaeology, ancient history, Jewish Studies, theology, linguistics, manuscripts, and related fields
Thank you for signing up!
Check your email for updates on new speakers, presentation topics, and conference news.
Something went wrong
Please try again or contact us if the problem persists.
Abstract
The number of variants (or differences) shared by witnesses of the same textual tradition has traditionally been used in stemmatology to determine the witnesses' relationships (Quentin, 1926; Froger, 1968). Importantly, however, the nature of these variants has also been shown to be highly significant (Barthélemy, 1982; Amphoux, 1989).
This presentation will introduce a method developed by our interdisciplinary team for the automatic annotation of variant categories, such as 'plus/minus', 'inversion', 'morphological' and 'lexical'.
The underlying manually annotated data and annotation protocol will first be presented. The different classifier experiments will then be discussed, focusing on the distinction between neural and non-neural models, the development of synthetic data for each variant category, and the use of additional classifier features, such as morphological annotation.
The strongest classifier obtained (F1 score: 0.80) will be presented. The proposed method will briefly be compared with a rule-based alternative.
Finally, future directions will be discussed, including the classifier's integration into a comprehensive pipeline for witness comparison.
Keywords
Computational Classification of Biblical Hebrew Manuscript Variants
Iglika Nikolova-Stoupak's research applies natural-language processing and computational philology to the study of textual transmission. In a 2025 ACL Findings paper with Maxime Amblard, Sophie Robert-Hayek, and Frédérique Rey, she helped develop a classifier for word-level differences in witnesses of Biblical Hebrew manuscripts.
The research classifies aligned word pairs into categories including plus/minus, inversion, morphological, lexical, and unclassifiable variants. The published system combines lexical information, part-of-speech tags, hand-crafted rules, and synthetically derived data and reported an F1 score of 0.80. The research also evaluated neural approaches related to DictaBERT and compared morphological information from manuscript resources.
In 2026, Nikolova-Stoupak and collaborators also presented BEReshiT, an Ancient Hebrew model based on DictaBERT, including a morphological submodel. Together, these projects illustrate how new computational tools can assist stemmatology, manuscript comparison, morphological analysis, and textual criticism while remaining dependent on philological judgment.
Academic Profile
Iglika Nikolova-Stoupak is a researcher in computational linguistics and natural-language processing. She works full-time as a research engineer / developer-researcher at LORIA and the University of Lorraine, where a major part of her work applies NLP methods to stemmatology and the analysis of Hebrew manuscripts.
She is also pursuing a Ph.D. in Computational Linguistics at Sorbonne Université. Her doctoral project is titled "Production of Abridged Versions of Literary Texts: A Multilingual Approach."
She holds Master's degrees in Literature from the University of Essex and Computing from Staffordshire University, both completed with distinction, and completed a two-year specialization in natural-language processing at Kyoto University under a MEXT scholarship. In 2025 she received the "Culture and Creativity" award at the Study UK Alumni Awards in France.
Her research combines NLP, historical and low-resource languages, digital humanities, computational philology, multilinguality, textual transmission, and manuscript analysis.
Education and Academic Qualifications
PhD Candidate in Computational Linguistics
Sorbonne Université, Paris.
M.A./Master's Degree in Literature
University of Essex — with distinction.
Master's Degree in Computing
Staffordshire University — with distinction.
Two-Year Specialization in Natural Language Processing
Kyoto University, under the MEXT scholarship.
2025
Study UK Alumni Awards, France — "Culture and Creativity" award.
Academic / Research Positions
Research Engineer in Natural Language Processing
LORIA / Université de Lorraine.
PhD Candidate
Sorbonne Université.
Research Involvement
Computational stemmatology and Hebrew manuscript analysis at Université de Lorraine.
Computational Methods for Manuscript Analysis
Nikolova-Stoupak's work exemplifies the growing role of machine learning and computational linguistics in biblical manuscript studies. Related digital humanities research at the Institute includes the BibCrit Ancient Witness Bridge for comparing Hebrew Bible variants, which enables scholars to analyze textual witnesses across multiple traditions computationally.
Research Expertise
- • Natural Language Processing
- • Computational Linguistics
- • Biblical Hebrew Manuscripts
- • Ancient Hebrew NLP
- • Computational Philology
- • Stemmatology
- • Textual Criticism
- • Manuscript Variant Classification
- • Historical and Low-Resource Languages
- • Digital Humanities
- • Multilingual Text Processing
- • Language Models for Ancient Languages
Selected Publications
Nikolova-Stoupak, Iglika, Maxime Amblard, Sophie Robert-Hayek, and Frédérique Rey.
"A Classifier of Word-Level Variants in Witnesses of Biblical Hebrew Manuscripts."
Findings of the Association for Computational Linguistics: ACL 2025, 21313–21329.
Nikolova-Stoupak, Iglika, Maxime Amblard, and Frédérique Rey.
"BEReshiT: an Ancient Hebrew Model based on DictaBERT."
LT4HALA 2026 / LREC 2026.
Nikolova-Stoupak, Iglika, Gaël Lejeune, and Eva Schaeffer-Lacroix.
"Does ChatGPT Adapt Itself to the Language Used and the Audience It Implies?"
2025.
Nikolova-Stoupak, Iglika, Gaël Lejeune, and Eva Schaeffer-Lacroix.
"Compilation of a Synthetic Judeo-French Corpus."
LaTeCH-CLfL 2024.
Academic Profiles and Verified Sources
Personal Academic Site
Iglika Nikolova-StoupakACL Anthology
Iglika Nikolova-Stoupak PublicationsMetz TheoLab / Université de Lorraine
Team ProfileConference Participation
Iglika Nikolova-Stoupak presents at the European Hebrew Journal International Online Conference 2026 on 10 November 2026 at 17:00 CET, Serbia time, with the presentation "Classifying Word-Level Variants in Biblical Hebrew Manuscripts: New Computational Tools for Textual Criticism."
Suggested Citation
Nikolova-Stoupak, Iglika. "Classifying Word-Level Variants in Biblical Hebrew Manuscripts: New Computational Tools for Textual Criticism." Presentation at New Discoveries and New Directions in Biblical, Hebrew, and Theological Studies, European Hebrew Journal International Scholarly Online Zoom Conference, 10 November 2026.
Attend Iglika Nikolova-Stoupak's Presentation
Register for the European Hebrew Journal International Online Conference to attend the live presentation and receive twelve months of access to the complete conference recording library.