Language Resources and Evaluation

Papers
(The median citation count of Language Resources and Evaluation is 1. The table below lists those papers that are above that threshold based on CrossRef citation counts [max. 250 papers]. The publications cover those that have been published in the past four years, i.e., from 2022-08-01 to 2026-08-01.)
ArticleCitations
Strategies for managing time and costs in speech corpus creation: insights from the Slovenian ARTUR corpus53
From LIMA to DeepLIMA: following a new path of interoperability29
Spelling errors made by people with dyslexia28
Lahjoita puhetta: a large-scale corpus of spoken Finnish with some benchmarks26
A survey on geocoding: algorithms and datasets for toponym resolution26
Brazilian Portuguese corpora for teaching and translation: the CoMET project25
Hope speech detection in Spanish25
IIT Delhi Dialogue Corpus: a quantitative analysis of a spoken corpus of Hindi24
Prompting encoder models for zero-shot classification: a cross-domain study in Italian22
The Visual Language Research Corpus (VLRC): an annotated corpus of comics from Asia, Europe, and the United States21
The narratives of war (NoW) corpus of written testimonies of the Russia-Ukraine war21
AC-IQuAD: Automatically Constructed Indonesian Question Answering Dataset by Leveraging Wikidata16
Understanding conversational interaction in multiparty conversations: the EVA Corpus15
A new evaluation method: evaluation data and metrics for Chinese grammatical error correction15
A study on methods for revising dependency treebanks: in search of gold14
Quality assessment of Tibetan–Chinese poetry translation: integrating automated metrics and qualitative insights through a cross-system comparison of dedicated NMT engines and a prompted LLM14
Spontaneous, controlled acts of reference between friends and strangers14
TLEX: an efficient method for extracting exact timelines from TimeML temporal graphs13
The properties of panels in global comics: frequency and size of 76 K panels in 1,030 comics from 144 countries13
Speech recognition in edge environments: an exploration of support and impact of model compression13
Toxic comment classification and rationale extraction in code-mixed text leveraging co-attentive multi-task learning13
Human–machine interaction in building an English reference dataset for natural language processing tasks12
Construction of Amharic information retrieval resources and corpora12
Assessing linguistic generalisation in language models: a dataset for Brazilian Portuguese12
LoNLI: An Extensible Framework for Testing Diverse Logical Reasoning Capabilities for NLI11
adaptNMT: an open-source, language-agnostic development environment for neural machine translation11
Sentiment analysis in Portuguese tweets: an evaluation of diverse word representation models11
Automatic readability assessment for sentences: neural, hybrid and large language models10
Utilizing phonetic similarity for cross-source and cross-language toponym matching: a benchmark and prototype10
Ma’aks: manually-curated parallel dataset for Arabic text sentiment swap10
UHated: hate speech detection in Urdu language using transfer learning9
A comparative analysis of encoder only and decoder only models in intent classification and sentiment analysis: navigating the trade-offs in model size and performance9
Perspectivist approaches to natural language processing: a survey8
CORAA ASR: a large corpus of spontaneous and prepared speech manually validated for speech recognition in Brazilian Portuguese8
DoSLex: automatic generation of all domain semantically rich sentiment lexicon8
An integrated framework for emotion and sentiment analysis in Tamil and Malayalam visual content8
Chinese-DiMLex: a lexicon of Chinese discourse connectives8
Conversion of the Spanish WordNet databases into a Prolog-readable format8
The Sanskrit Sembank8
Slovenian parliamentary corpus siParl7
Ulysses Tesemõ: a new large corpus for Brazilian legal and governmental domain7
VeLeSpa: An inflected verbal lexicon of Peninsular Spanish and a quantitative analysis of paradigmatic predictability7
Detecting racism in the digital age: a survey of datasets and algorithms7
Book Review: The Routledge handbook of discourse and disinformation7
A performance analysis of a large language model for Marathi language NLP tasks7
TCMeta: a multilingual dataset of COVID tweets for relation-level metaphor analysis7
Uzbek news corpus for named entity recognition7
Correction: The corpus of aggressive language in Polish parliamentary debates7
Developing and mining an underage modern Greek chat corpus: Do students show signs of bullying behavior while working on a project?7
Benchmarking Hindi-to-English direct speech-to-speech translation with synthetic data6
Studying word meaning evolution through incremental semantic shift detection6
Open source platform for Estonian speech transcription6
Multi-task learning for multi-dialect Arabic sentiment classification and sarcasm detection6
Benchmarking and AI-assisted human-like evaluation of retrieval-augmented generation for Arabic and English documents6
Sense through time: diachronic word sense annotations for word sense induction and Lexical Semantic Change Detection6
KurdiSent: a corpus for kurdish sentiment analysis6
Language resources for clinical linguistics: introduction to the special issue6
PolitePEER: does peer review hurt? A dataset to gauge politeness intensity in the peer reviews6
Correction to: Two sepedi‑english code‑switched speech corpora5
Part of speech (POS) tagging in Roman Urdu: datasets and models5
JurisTCU: a Brazilian Portuguese information retrieval dataset with query relevance judgments5
Correction: Cross-linguistically consistent semantic and syntactic annotation of child-directed speech5
Developing and testing syllabification systems for South African Sesotho5
A corpus of English learners with Arabic and Hebrew backgrounds5
Sentiment analysis dataset in Moroccan dialect: bridging the gap between Arabic and Latin scripted dialect5
Darijamachaair: a large-scale multi-dimensional emotion and context dataset for Moroccan dialect social media texts5
A new corpus platform for the Texas German Dialect Project5
kidsNARRATE: a versatile corpus for studying Chinese-english bilingual L2 narrative skills in preschoolers5
Text-Muddler: an advanced adversarial paradigm for disrupting NLP-based neural architectures in sentiment analysis frameworks5
Design and construction of Guayaquil radio speech corpus (CHARG)5
The taggedPBC: annotating a massive parallel corpus for crosslinguistic investigations4
Creation of a gold standard Dutch corpus of clinical notes for adverse drug event detection: the Dutch ADE corpus4
DILLo: an Italian lexical database for speech-language pathologists4
Using BERT models for breast cancer diagnosis from Turkish radiology reports4
The limitations of irony detection in Dutch social media4
Aratox: a multi-dialect, multi-label arabic dataset and model benchmark for toxicity detection4
FullStop: punctuation and segmentation prediction for Dutch with transformers4
ThaiCoref: Thai coreference resolution dataset4
Multilingual speech representation for the Manipuri automatic speech recognition system4
Finnish parliament ASR corpus4
OMCD: Offensive Moroccan Comments Dataset4
MarIA and BETO are sexist: evaluating gender bias in large language models for Spanish3
Towards a resource for multilingual lexicons: an MT assisted and human-in-the-loop multilingual parallel corpus with multi-word expression annotation3
PARSEME-AR: Arabic reference corpus for multiword expressions using PARSEME annotation guidelines3
Sentiment analysis in low-resource contexts: BERT’s impact on Central Kurdish3
Examining inferred author and textual correlates of harmful language annotation3
Multi-layered semantic annotation and the formalisation of annotation schemas for the investigation of modality in a Latin corpus3
Investigating interoperable event corpora: limitations of reusability of resources and portability of models3
HASTIKA: hate speech and target identification in Kannada-English code-mixed text3
Automating translation checks of financial documents using large language models3
Correction: COLLIE: a broad-coverage ontology and lexicon of verbs in English3
MulCogBench: a multi-modal cognitive benchmark dataset for evaluating Chinese and English computational language models3
OLID-BR: offensive language identification dataset for Brazilian Portuguese3
Do you understand Italian? Evaluating LVLMs on Italian visual question-answering3
Correction to: Resources for Turkish natural language processing: A critical survey3
Assessment of pragmatic abilities and cognitive substrates (APACS) brief remote: a novel tool for the rapid and tele-evaluation of pragmatic skills in Italian3
Building an emotion lexicon for Serbian using curated language resources3
NILC-Metrix: assessing the complexity of written and spoken language in Brazilian Portuguese2
A survey and study impact of tweet sentiment analysis via transfer learning in low resource scenarios2
ChavacanoMT: a corpus and evaluation of neural machine translation for Philippine Creole Spanish2
Disfluency annotated corpora for Indian English in technical domains2
Data-driven weakly supervised emotion classification with consistency regularization: Mandarin Chinese as a case2
RastrOS Project: Natural Language Processing contributions to the development of an eye-tracking corpus with predictability norms for Brazilian Portuguese2
Detection of political hate speech in Korean language2
Building a relevance feedback corpus for legal information retrieval in the real-case scenario of the Brazilian Chamber of Deputies2
A sentiment corpus for the cryptocurrency financial domain: the CryptoLin corpus2
Rei Miyata: controlled document authoring in a machine translation age2
The Najdi Arabic Corpus: a new corpus for an underrepresented Arabic dialect2
Beyond accuracy: completeness and relevance metrics for evaluating the quality of long answers2
Comparative performance of ensemble machine learning for Arabic cyberbullying and offensive language detection2
Automatic genre identification: a survey2
A comprehensive evaluation of semantic relation knowledge of pretrained language models and humans2
Incremental imbalance-aware deep learning framework for multilingual spoken language identification2
Neural text sanitization with privacy risk indicators: an empirical analysis2
Czech news dataset for semantic textual similarity2
Infectious risk events and their novelty in event-based surveillance: new definitions and annotated corpus2
DiscoNaija: a discourse-annotated parallel Nigerian Pidgin-English corpus2
Parlamint-it: an 18-karat UD treebank of Italian parliamentary speeches2
Speech emotion recognition for the Urdu language2
Beyond plain toxic: building datasets for detection of flammable topics and inappropriate statements2
The Mandarin Chinese speech database: a corpus of 18,820 auditory neutral nonsense sentences2
COLLIE: a broad-coverage ontology and lexicon of verbs in English2
Human-inspired computational models for European Portuguese: a review2
Normalized dataset for Sanskrit word segmentation and morphological parsing2
RUN-AS: a novel approach to annotate news reliability for disinformation detection2
FinnSentiment: a Finnish social media corpus for sentiment polarity annotation2
Spoken Spanish PoS tagging: gold standard dataset2
Content-free speech activity records: interviews with people with schizophrenia2
Arab music improvisation corpus for research (AMICOR): development and machine translation experiments2
“You’ll be a nurse, my son!” Automatically assessing gender biases in autoregressive language models in French and Italian2
SOLD: Sinhala offensive language dataset2
An aligned corpus of Spanish bibles2
A survey of gender bias mitigation in neural machine translation2
Faux Hate: unravelling the web of fake narratives in spreading hateful stories: a multi-label and multi-class dataset in cross-lingual Hindi-English code-mixed text2
A Chinese natural speech complex emotion dataset based on emotion vector annotation method2
Building a specialised Hebrew textual corpus on construction, planning and architecture2
Aspect-based multimodal sentiment analysis via employing visual-to-emotional-caption translation network using visual-caption pairs2
A rich task-oriented dialogue corpus in Vietnamese2
A comparative study of sentence alignment methods for Spanish text simplification2
Evaluation of a rule-based approach to automatic factual question generation using syntactic and semantic analysis2
Entity normalization in a Spanish medical corpus using a UMLS-based lexicon: findings and limitations2
Evaluation of the Brazilian Portuguese version of linguistic inquiry and word count 2015 (BP-LIWC2015)1
CINWA (database of terminology for cultivated plants in indigenous languages of northwestern South America): introducing a resource for research in ethnobiology, anthropology, historical linguistics, 1
Regionalized models for Spanish language variations based on Twitter1
Label modification and bootstrapping for zero-shot cross-lingual hate speech detection1
Evaluation of end-to-end continuous spanish lipreading in different data conditions1
Correction: Aratox: a multi-dialect, multi-label arabic dataset and model benchmark for toxicity detection1
The corpus of aggressive language in Polish parliamentary debates1
UstanceBR: a social media language resource for stance prediction1
A new methodology for automatic creation of concept maps of Turkish texts1
Umplc: the first longitudinal learner corpus of Portuguese1
Error annotation: a review and faceted taxonomy1
A benchmark dataset and evaluation methodology for Chinese zero pronoun translation1
The link between translation difficulty and the quality of machine translation: a literature review and empirical investigation1
A comparative evaluation for question answering over Greek texts by using machine translation and BERT1
A corpus of Persian literary text1
Cat-related slang and sentiment analysis: the CatSlang-SA dataset and model evaluation using social media posts1
A flexible tool for a qualia-enriched FrameNet: the FrameNet Brasil WebTool1
"Approaches to sentiment analysis of Hungarian political news at the sentence level"1
Maithilimt: Developing Multi-Domain Parallel Corpus for Hindi-Maithili Machine Translation1
A new corpus of geolocated ASR transcripts from Germany1
TANDO+: corpus and baselines for document-level machine translation in Basque–Spanish and Basque–French1
Attention and LoRA-based multimodal emotion detection system1
Human–robot dialogue annotation for multi-modal common ground1
Corpus-based computational frame and construction analysis of motion metaphors1
From extended chunking to dependency parsing using traditional Arabic grammar1
Fake news article detection datasets for Hindi language1
Editorial: LRE updates1
Dataset on sentiment-based cryptocurrency-related news and tweets in English and Malay language1
Semantic evaluation metric conforming to AMR theory (SEMCAT): a new similarity metric for abstract meaning representation1
VeLeRo: an inflected verbal lexicon of standard Romanian and a quantitative analysis of morphological predictability1
Mining culture from professional discourse: a lexicon-based hybrid method1
UFLA-FORMS: an academic forms dataset for information extraction in the Portuguese language1
Bridging the linguistic divide: a survey on leveraging large language models for machine translation1
Correction: A comparative study of sentence alignment methods for Spanish text simplification1
Building the VisSE Corpus of Spanish SignWriting1
Toxicbias-reasoning: a multicultural dataset for social bias detection with human-aligned reasoning1
CsFEVER and CTKFacts: acquiring Czech data for fact verification1
Exploring lexical factors in semantic annotation: insights from the classification of nouns in French1
Parafrasário: a variety-based paraphrasary for Portuguese1
Historical Portuguese corpora: a survey1
Linguistic knowledge injected into large language model for Urdu-English neural machine translation1
Detoxifying language model outputs: combining multi-agent debates and reinforcement learning for improved summarization1
Special issue on language technology platforms1
Parallel Trees: a novel resource with aligned dependency and constituency syntactic representations1
Disfluency processing for cascaded speech translation involving English and Indian languages1
The robotic-surgery propositional bank1
DepreSym: A Depression Symptom Annotated Corpus and the Role of Large Language Models as Assessors of Psychological Markers1
“But why??” Evaluation of user-suggested synonyms in the Thesaurus of Modern Slovene1
Aspect sentiment triplet extraction via integrating contextual semantic relevance and syntactic relevance1
Improving Arabic sentiment analysis across context-aware attention deep model based on natural language processing1
Review of recent emotion-annotated text corpora and resources1
Ngalawan Ujaran Sengit: hate speech detection in indonesian code-mixed social media data1
UniC: a dataset for emotion analysis of videos with multimodal and unimodal labels1
Register identification from the unrestricted open Web using the Corpus of Online Registers of English1
POS tagging of low-resource Pashto language: annotated corpus and BERT-based model1
An empirical evaluation of arabic text formality transfer: a comparative study1
Exa-PSD: a new Persian sentiment analysis dataset on Twitter1
Multilingual prediction of semantic norms with language models: a study on English and Chinese1
Semantic search as extractive paraphrase span detection1
0.36455202102661