Language Resources and Evaluation

Papers
(The TQCC of Language Resources and Evaluation is 3. The table below lists those papers that are above that threshold based on CrossRef citation counts [max. 250 papers]. The publications cover those that have been published in the past four years, i.e., from 2022-08-01 to 2026-08-01.)
ArticleCitations
Strategies for managing time and costs in speech corpus creation: insights from the Slovenian ARTUR corpus53
From LIMA to DeepLIMA: following a new path of interoperability29
Spelling errors made by people with dyslexia28
A survey on geocoding: algorithms and datasets for toponym resolution26
Lahjoita puhetta: a large-scale corpus of spoken Finnish with some benchmarks26
Brazilian Portuguese corpora for teaching and translation: the CoMET project25
Hope speech detection in Spanish25
IIT Delhi Dialogue Corpus: a quantitative analysis of a spoken corpus of Hindi24
Prompting encoder models for zero-shot classification: a cross-domain study in Italian22
The narratives of war (NoW) corpus of written testimonies of the Russia-Ukraine war21
The Visual Language Research Corpus (VLRC): an annotated corpus of comics from Asia, Europe, and the United States21
AC-IQuAD: Automatically Constructed Indonesian Question Answering Dataset by Leveraging Wikidata16
Understanding conversational interaction in multiparty conversations: the EVA Corpus15
A new evaluation method: evaluation data and metrics for Chinese grammatical error correction15
A study on methods for revising dependency treebanks: in search of gold14
Quality assessment of Tibetan–Chinese poetry translation: integrating automated metrics and qualitative insights through a cross-system comparison of dedicated NMT engines and a prompted LLM14
Spontaneous, controlled acts of reference between friends and strangers14
TLEX: an efficient method for extracting exact timelines from TimeML temporal graphs13
The properties of panels in global comics: frequency and size of 76 K panels in 1,030 comics from 144 countries13
Speech recognition in edge environments: an exploration of support and impact of model compression13
Toxic comment classification and rationale extraction in code-mixed text leveraging co-attentive multi-task learning13
Assessing linguistic generalisation in language models: a dataset for Brazilian Portuguese12
Human–machine interaction in building an English reference dataset for natural language processing tasks12
Construction of Amharic information retrieval resources and corpora12
Sentiment analysis in Portuguese tweets: an evaluation of diverse word representation models11
LoNLI: An Extensible Framework for Testing Diverse Logical Reasoning Capabilities for NLI11
adaptNMT: an open-source, language-agnostic development environment for neural machine translation11
Ma’aks: manually-curated parallel dataset for Arabic text sentiment swap10
Automatic readability assessment for sentences: neural, hybrid and large language models10
Utilizing phonetic similarity for cross-source and cross-language toponym matching: a benchmark and prototype10
A comparative analysis of encoder only and decoder only models in intent classification and sentiment analysis: navigating the trade-offs in model size and performance9
UHated: hate speech detection in Urdu language using transfer learning9
Conversion of the Spanish WordNet databases into a Prolog-readable format8
The Sanskrit Sembank8
Perspectivist approaches to natural language processing: a survey8
CORAA ASR: a large corpus of spontaneous and prepared speech manually validated for speech recognition in Brazilian Portuguese8
DoSLex: automatic generation of all domain semantically rich sentiment lexicon8
An integrated framework for emotion and sentiment analysis in Tamil and Malayalam visual content8
Chinese-DiMLex: a lexicon of Chinese discourse connectives8
Correction: The corpus of aggressive language in Polish parliamentary debates7
Developing and mining an underage modern Greek chat corpus: Do students show signs of bullying behavior while working on a project?7
Slovenian parliamentary corpus siParl7
Ulysses Tesemõ: a new large corpus for Brazilian legal and governmental domain7
VeLeSpa: An inflected verbal lexicon of Peninsular Spanish and a quantitative analysis of paradigmatic predictability7
Detecting racism in the digital age: a survey of datasets and algorithms7
Book Review: The Routledge handbook of discourse and disinformation7
A performance analysis of a large language model for Marathi language NLP tasks7
TCMeta: a multilingual dataset of COVID tweets for relation-level metaphor analysis7
Uzbek news corpus for named entity recognition7
Language resources for clinical linguistics: introduction to the special issue6
PolitePEER: does peer review hurt? A dataset to gauge politeness intensity in the peer reviews6
Benchmarking Hindi-to-English direct speech-to-speech translation with synthetic data6
Studying word meaning evolution through incremental semantic shift detection6
Open source platform for Estonian speech transcription6
Multi-task learning for multi-dialect Arabic sentiment classification and sarcasm detection6
Benchmarking and AI-assisted human-like evaluation of retrieval-augmented generation for Arabic and English documents6
Sense through time: diachronic word sense annotations for word sense induction and Lexical Semantic Change Detection6
KurdiSent: a corpus for kurdish sentiment analysis6
A new corpus platform for the Texas German Dialect Project5
kidsNARRATE: a versatile corpus for studying Chinese-english bilingual L2 narrative skills in preschoolers5
Text-Muddler: an advanced adversarial paradigm for disrupting NLP-based neural architectures in sentiment analysis frameworks5
Design and construction of Guayaquil radio speech corpus (CHARG)5
Correction to: Two sepedi‑english code‑switched speech corpora5
Part of speech (POS) tagging in Roman Urdu: datasets and models5
JurisTCU: a Brazilian Portuguese information retrieval dataset with query relevance judgments5
Correction: Cross-linguistically consistent semantic and syntactic annotation of child-directed speech5
Developing and testing syllabification systems for South African Sesotho5
A corpus of English learners with Arabic and Hebrew backgrounds5
Sentiment analysis dataset in Moroccan dialect: bridging the gap between Arabic and Latin scripted dialect5
Darijamachaair: a large-scale multi-dimensional emotion and context dataset for Moroccan dialect social media texts5
Finnish parliament ASR corpus4
OMCD: Offensive Moroccan Comments Dataset4
The taggedPBC: annotating a massive parallel corpus for crosslinguistic investigations4
Creation of a gold standard Dutch corpus of clinical notes for adverse drug event detection: the Dutch ADE corpus4
DILLo: an Italian lexical database for speech-language pathologists4
Using BERT models for breast cancer diagnosis from Turkish radiology reports4
The limitations of irony detection in Dutch social media4
Aratox: a multi-dialect, multi-label arabic dataset and model benchmark for toxicity detection4
FullStop: punctuation and segmentation prediction for Dutch with transformers4
ThaiCoref: Thai coreference resolution dataset4
Multilingual speech representation for the Manipuri automatic speech recognition system4
Do you understand Italian? Evaluating LVLMs on Italian visual question-answering3
Correction to: Resources for Turkish natural language processing: A critical survey3
Assessment of pragmatic abilities and cognitive substrates (APACS) brief remote: a novel tool for the rapid and tele-evaluation of pragmatic skills in Italian3
Building an emotion lexicon for Serbian using curated language resources3
MarIA and BETO are sexist: evaluating gender bias in large language models for Spanish3
Towards a resource for multilingual lexicons: an MT assisted and human-in-the-loop multilingual parallel corpus with multi-word expression annotation3
PARSEME-AR: Arabic reference corpus for multiword expressions using PARSEME annotation guidelines3
Sentiment analysis in low-resource contexts: BERT’s impact on Central Kurdish3
Examining inferred author and textual correlates of harmful language annotation3
Multi-layered semantic annotation and the formalisation of annotation schemas for the investigation of modality in a Latin corpus3
Investigating interoperable event corpora: limitations of reusability of resources and portability of models3
HASTIKA: hate speech and target identification in Kannada-English code-mixed text3
Automating translation checks of financial documents using large language models3
Correction: COLLIE: a broad-coverage ontology and lexicon of verbs in English3
MulCogBench: a multi-modal cognitive benchmark dataset for evaluating Chinese and English computational language models3
OLID-BR: offensive language identification dataset for Brazilian Portuguese3
0.19019603729248