Computer Speech and Language

Papers
(The median citation count of Computer Speech and Language is 3. The table below lists those papers that are above that threshold based on CrossRef citation counts [max. 250 papers]. The publications cover those that have been published in the past four years, i.e., from 2022-08-01 to 2026-08-01.)
ArticleCitations
A language-agnostic model of child language acquisition84
KddRES: A Multi-level Knowledge-driven Dialogue Dataset for Restaurant Towards Customized Dialogue System79
Stochastic Data-to-Text Generation Using Syntactic Dependency Information79
Editorial Board74
Editorial Board44
Room impulse response reshaping-based expectation–maximization in an underdetermined reverberant environment43
Speech enhancement approach for body-conducted unvoiced speech based on Taylor–Boltzmann machines trained DNN43
Optimization of modular multi-speaker distant conversational speech recognition42
Seq2Seq dynamic planning network for progressive text generation40
Corpus and unsupervised benchmark: Towards Tagalog grammatical error correction37
Towards privacy-preserving conversation analysis in everyday life: Exploring the privacy-utility trade-off36
Unsupervised question-retrieval approach based on topic keywords filtering and multi-task learning34
Identifying offensive memes in low-resource languages: A multi-modal multi-task approach using valence and arousal33
Editorial Board26
A method of phonemic annotation for Chinese dialects based on a deep learning model with adaptive temporal attention and a feature disentangling structure26
PaSCoNT - Parallel Speech Corpus of Northern-central Thai for automatic speech recognition26
Monotonic Gaussian regularization of attention for robust automatic speech recognition24
Complementary regional energy features for spoofed speech detection22
Multi-branch feature aggregation based on multiple weighting for speaker verification22
Misogynistic attitude detection in YouTube comments and replies: A high-quality dataset and algorithmic models22
Vishing: Detecting social engineering in spoken communication — A first survey & urgent roadmap to address an emerging societal challenge22
Contextual emotion detection using ensemble deep learning21
A transformer-based spelling error correction framework for Bangla and resource scarce Indic languages20
Improving self-supervised learning model for audio spoofing detection with layer-conditioned embedding fusion19
Editorial Board19
A hybrid approach to Natural Language Inference for the SICK dataset18
Enhancing analysis of diadochokinetic speech using deep neural networks18
Editorial Board18
Maximal activation weighted memory for aspect based sentiment analysis17
Augmentative and alternative speech communication (AASC) aid for people with dysarthria17
Predicting accentedness and comprehensibility through ASR scores and acoustic features17
Loanword identification based on web resources: A case study on wikipedia17
MS-Swinformer and DMTL: Multi-scale spatial fusion and dynamic multi-task learning for speech emotion recognition17
Preserving the beamforming effect for spatial cue-based pseudo-binaural dereverberation of a single source16
The use of Active Learning systems for stimulus selection and response modelling in perception experiments16
A lightweight approach based on prompt for few-shot relation extraction16
Under the hood: Phonemic Restoration in transformer-based automatic speech recognition16
Representation learning strategies to model pathological speech: Effect of multiple spectral resolutions16
Combined generative and predictive modeling for speech super-resolution16
English–Assamese neural machine translation using prior alignment and pre-trained language model15
Emotion-guided cross-modal alignment for multimodal depression detection15
Combining replay and LoRA for continual learning in natural language understanding15
Conversations in the wild: Data collection, automatic generation and evaluation15
A mobile application using automatic speech analysis for classifying Alzheimer's disease and mild cognitive impairment15
Editorial Board15
A novel channel estimate for noise robust speech recognition14
Privacy-preserving feature extractor using adversarial pruning for TBI assessment from speech14
Compress, Align, and Transfer: A new method for transferring pre-trained language models knowledge to CTC-based speech recognition14
Meta adversarial learning improves low-resource speech recognition14
SecNLP: An NLP classification model watermarking framework based on multi-task learning13
Editorial Board13
A bias evaluation solution for multiple sensitive attribute speech recognition13
Evidence and Axial Attention Guided Document-level Relation Extraction13
Performance assessment of voice conversion models using speech production-based parameters13
Keyword Mamba: Spoken keyword spotting with state space models13
Effects of cross-cultural language differences on social cognition during human-agent interaction in cooperative game environments13
Modality fusion using auxiliary tasks for dementia detection12
A novel graph kernel algorithm for improving the effect of text classification12
Addressing subjectivity in paralinguistic data labeling for improved classification performance: A case study with Spanish-speaking Mexican children using data balancing and semi-supervised learning12
Tailored design of Audio–Visual Speech Recognition models using Branchformers12
A computational analysis of transcribed speech of people living with dementia: The Anchise 2022 Corpus12
MPSA-DenseNet: A novel deep learning model for English accent classification12
FinD: Fine-grained discrepancy-based fake news detection enhanced by event abstract generation12
Neural multi-task learning for end-to-end Arabic aspect-based sentiment analysis12
A flexible BERT model enabling width- and depth-dynamic inference12
Towards inclusive automatic speech recognition12
Offensive language detection in Tamil YouTube comments by adapters and cross-domain knowledge transfer12
Zero-Shot Strike: Testing the generalisation capabilities of out-of-the-box LLM models for depression detection12
Improved relation extraction through key phrase identification using community detection on dependency trees12
Entrainment detection using DNN12
Effective infant cry signal analysis and reasoning using IARO based leaky Bi-LSTM model12
A tag-based methodology for the detection of user repair strategies in task-oriented conversational agents11
Raw acoustic-articulatory multimodal dysarthric speech recognition11
Improving BERT with local context comprehension for multi-turn response selection in retrieval-based dialogue systems11
Towards detecting the level of trust in the skills of a virtual assistant from the user’s speech11
Towards decoupling frontend enhancement and backend recognition in monaural robust ASR11
GenCeption: Evaluate vision LLMs with unlabeled unimodal data11
Enhancing accuracy and privacy in speech-based depression detection through speaker disentanglement11
A closer look at reinforcement learning-based automatic speech recognition10
Multi-task unified model for Chinese aspect-based sentiment analysis10
Objective and subjective evaluation of speech enhancement methods in the UDASE task of the 7th CHiME challenge10
An experimental study of diffusion-based general speech restoration with predictive-guided conditioning10
Prototypical networks relation classification model based on entity convolution10
Listening to Variation: Human and Whisper Responses to Grammatical Gender Agreement10
Test-retest reliability of acoustic and linguistic measures of speech tasks10
DiffATSM: High quality adaptive time-scale modification using diffusion-based post-processing9
Multiple time-instances features based approach for reference-free speech quality measurement9
Hate speech and offensive language detection in Dravidian languages using deep ensemble framework9
Towards lifelong human assisted speaker diarization9
A neural network approach for speech enhancement and noise-robust bandwidth extension9
Continual End-to-End Speech-to-Text translation using augmented bi-sampler9
Morse wavelet transform-based features for voice liveness detection8
Refining the evaluation of speech synthesis: A summary of the Blizzard Challenge 20238
Adaptive feature extraction for entity relation extraction8
Universal constituency treebanking and parsing: A pilot study8
Deep feature representations and fusion strategies for speech emotion recognition from acoustic and linguistic modalities: A systematic review8
A novel approach to cross-linguistic transfer learning for hope speech detection in Tamil and Malayalam8
End-to-End Speech-to-Text Translation: A Survey8
A physical exertion inspired multi-task learning framework for detecting out-of-breath speech8
Speech self-supervised representations benchmarking: A case for larger probing heads7
Pitch-Aware multi-feature fusion for classifying statements, questions, and exclamations in low-resource languages7
Goal-oriented conditional variational autoencoders for proactive and knowledge-aware conversational recommender system7
V-APA: A Voice-driven Agentic Process Automation System7
Two in One: A multi-task framework for politeness turn identification and phrase extraction in goal-oriented conversations7
Measuring and implementing lexical alignment: A systematic literature review7
Three-stage modular speaker diarization collaborating with front-end techniques in the CHiME-8 NOTSOFAR-1 challenge7
SEBGM: Sentence Embedding Based on Generation Model with multi-task learning7
An automated quality evaluation framework of psychotherapy conversations with local quality estimates7
Significance of chirp MFCC as a feature in speech and audio applications7
Minerva 2 for speech and language tasks7
Classification of stuttering – The ComParE challenge and beyond7
Automatic screening of mild cognitive impairment and Alzheimer’s disease by means of posterior-thresholding hesitation representation7
Channel and channel subband selection for speaker diarization6
A cross-attention augmented model for event-triggered context-aware story generation6
Using Knowledge Induction strategies: LLMs can do better in knowledge-driven dialogue tasks6
Direct enhancement of pre-trained speech embeddings for speech processing in noisy conditions6
GTSO: Gradient tangent search optimization enabled voice transformer with speech intelligibility for aphasia6
Preserving speaker information in direct Speech-to-Speech Translation with non-autoregressive generation and pre-training6
Improving named entity correctness of abstractive summarization by generative negative sampling6
Spoofing countermeasure for fake speech detection using brute force features6
QuAVA: A privacy-aware architecture for conversational desktop Content Retrieval systems6
Multilingual non-intrusive binaural intelligibility prediction based on phone classification6
A knowledge-augmented heterogeneous graph convolutional network for aspect-level multimodal sentiment analysis6
Building a text retrieval system for the Sanskrit language: Exploring indexing, stemming, and searching issues6
C-KGE: Curriculum learning-based Knowledge Graph Embedding6
A new speech corpus of super-elderly Japanese for acoustic modeling5
Editorial Board5
Assessing language models’ task and language transfer capabilities for sentiment analysis in dialog data5
Scale-aware dual-branch complex convolutional recurrent network for monaural speech enhancement5
Speaking to remember: Model-based adaptive vocabulary learning using automatic speech recognition5
On significance of constant-Q transform for pop noise detection5
Neural referential form selection: Generalisability and interpretability5
A potential relation trigger method for entity-relation quintuple extraction in text with excessive entities5
Towards better Chinese-centric neural machine translation for low-resource languages5
Two evaluations on Ontology-style relation annotations5
Accurate speaker counting, diarization and separation for advanced recognition of multichannel multispeaker conversations5
Rep-MCA-former: An efficient multi-scale convolution attention encoder for text-independent speaker verification5
On the use of DiaPer models and matching algorithm for RTVE speaker diarization 2024 dataset5
Cross-lingual multi-speaker speech synthesis with limited bilingual training data5
Uncertainty-aware non-autoregressive neural machine translation5
FE-CFNER: Feature Enhancement-based approach for Chinese Few-shot Named Entity Recognition5
Real-time audio enhancement framework for vocal performances based on LSTM and time-frequency masking algorithm5
Optimizing pipeline task-oriented dialogue systems using post-processing networks5
UniKDD: A Unified Generative model for Knowledge-driven Dialogue5
One-class neural network with hybrid pooling on dual-band frequency for spoofing speech detection5
LRetUNet: A U-Net-based retentive network for single-channel speech enhancement5
Copiously Quote Classics: Improving Chinese Poetry Generation with historical allusion knowledge5
Editorial Board4
What’s so complex about conversational speech? A comparison of HMM-based and transformer-based ASR architectures4
Deep learning based multi-source localization with source splitting and its effectiveness in multi-talker speech recognition4
Multi-task learning neural framework for categorizing sexism4
An experimental review of speaker diarization methods with application to two-speaker conversational telephone speech recordings4
A code-mixed task-oriented dialog dataset for medical domain4
TadaStride: Using time adaptive strides in audio data for effective downsampling4
How to make embeddings suitable for PLDA4
Predicting children’s perceived reading proficiency with prosody modeling4
Editorial Board4
Enhancing Turkish Coreference Resolution: Insights from deep learning, dropped pronouns, and multilingual transfer learning4
Analysis of Instantaneous Frequency Components of Speech Signals for Epoch Extraction4
Knowledge-grounded dialogue modelling with dialogue-state tracking, domain tracking, and entity extraction4
Character expression for spoken dialogue systems with semi-supervised learning using Variational Auto-Encoder4
Time–Frequency Causal Hidden Markov Model for speech-based Alzheimer’s disease longitudinal detection4
A semi-supervised high-quality pseudo labels algorithm based on multi-constraint optimization for speech deception detection4
COMPASS: A creative support system that alerts novelists to the unnoticed missing contents4
Simultaneous speech and background sound recognition in diverse acoustic environments with branched neural networks4
Spectral–temporal saliency masks and modulation tensorgrams for generalizable COVID-19 detection4
Editorial Board4
M-Sim: Multi-level Semantic Inference Model for Chinese short answer scoring in low-resource scenarios4
Demystifying large language models in second language development research4
An analysis of machine learning models for sentiment analysis of Tamil code-mixed data4
Trainable multi-channel front-ends for joint beamforming and speaker embedding extraction4
EMGVox-GAN: A transformative approach to EMG-based speech synthesis, enhancing clarity, and efficiency via extensive dataset utilization4
Speech acoustics to rt-MRI articulatory dynamics inversion with video diffusion model4
Enhanced audio-visual speech enhancement with posterior sampling methods in recurrent variational autoencoders4
New research on monaural speech segregation based on quality assessment3
Sentiment analysis for live video comments with variational residual representations3
Single-channel speech enhancement using colored spectrograms3
Editorial Board3
Supervised speech separation combined with adaptive beamforming3
Incorporating external knowledge for text matching model3
Multimodal laryngoscopic video analysis for assisted diagnosis of vocal fold paralysis3
Automatic offline annotation of turn-taking transitions in task-oriented dialogue3
Editorial Board3
Modelling child comprehension: A case of suffixal passive construction in Korean3
HOTTEST: Hate and Offensive content identification in Tamil using Transformers and Enhanced STemming3
Editorial Board3
AraFastQA: a transformer model for question-answering for Arabic language using few-shot learning3
Deep learning-based speaker-adaptive postfiltering with limited adaptation data for embedded text-to-speech synthesis systems3
Leveraging saliency-based pre-trained foundation model representations to uncover breathing patterns in speech3
Gnowsis: Multimodal multitask learning for oral proficiency assessments3
Speech intelligibility assessment of dysarthria using Fisher vector encoding3
RepSum: A general abstractive summarization framework with dynamic word embedding representation correction3
Exploring the ability of LLMs to classify written proficiency levels3
The use of variable length stimuli for assessing segmental distortion in TTS evaluation3
Train from scratch: Single-stage joint training of speech separation and recognition3
Knowledge-enhanced meta-prompt for few-shot relation extraction3
Self-feeding training method for semi-supervised grammatical error correction3
A speech prediction model based on codec modeling and transformer decoding3
LeBenchmark 2.0: A standardized, replicable and enhanced framework for self-supervised representations of French speech3
Editorial Board3
A generalized decoding method for neural text generation3
Robustness in deepfake speech detection: A survey of failure mechanisms including an experimental case study3
Taking relations as known conditions: A tagging based method for relational triple extraction3
Replay spoof detection using energy separation based instantaneous frequency estimation from quadrature and in-phase components3
Survey of end-to-end multi-speaker automatic speech recognition for monaural audio3
Multi-level context features extraction for named entity recognition3
Towards explainable spoofed speech attribution and detection: A probabilistic approach for characterizing speech synthesizer components3
0.23174285888672