International Journal of Multimedia Information Retrieval

Papers
(The TQCC of International Journal of Multimedia Information Retrieval is 9. The table below lists those papers that are above that threshold based on CrossRef citation counts [max. 250 papers]. The publications cover those that have been published in the past four years, i.e., from 2022-05-01 to 2026-05-01.)
ArticleCitations
Generative adversarial networks and its applications in the biomedical image segmentation: a comprehensive survey93
VERITE: a Robust benchmark for multimodal misinformation detection accounting for unimodal bias74
Video anomaly detection with memory-guided multilevel embedding67
Recent trends in recommender systems: a survey59
Multiple object tracking under occlusions based on the stage-wise association strategy with weak cues58
Enhancing Facial Beauty Prediction via a Dual-Pathway Hybrid Architecture Integrating Vmamba and ViT36
Strengthening attention: knowledge distillation via cross-layer feature fusion for image classification32
DELIGHT-Net: DEep and LIGHTweight network to segment Indian text at word level from wild scenic images31
Enhanced YOLOv10 for small object detection with context-aware and adaptive modules28
VPC-VoxelNet: multi-modal fusion 3D object detection networks based on virtual point clouds28
CSAM: Capsule spatial attention mask network for visual question answering27
Feature-NeuS: Neural Implicit Surface Reconstruction Using Feature Multi-View Consistency Constraint26
Prototype local–global alignment network for image–text retrieval24
Multi-objective reinforcement learning for recommender systems: a comprehensive survey of methods, challenges, and future directions20
Hierarchical multi-modal fusion with vision transformers for robust action recognition in infrared-visible videos19
MMDL: a multi-modal deep learning for video highlight detection in sports19
Similarity-based face image retrieval using sparsely embedded deep features and binary code learning18
Human behavior recognition based on DualBiNet model17
Visual and semantic ensemble for scene text recognition with gated dual mutual attention16
How can users’ comments posted on social media videos be a source of effective tags?16
FiCo-ITR: bridging fine-grained and coarse-grained image-text retrieval for comparative performance analysis16
Multimodal music datasets? Challenges and future goals in music processing15
Semantic-enhanced discriminative embedding learning for cross-modal retrieval14
A Comprehensive Review of Multimodal Visual Representation Learning: Tracing the Evolution from CNNs to Transformers and Beyond14
State of art and emerging trends on group recommender system: a comprehensive review14
CAMIR: fine-tuning CLIP and multi-head cross-attention mechanism for multimodal image retrieval with sketch and text features14
Generative adversarial networks for 2D-based CNN pose-invariant face recognition13
Cross-domain image retrieval: methods and applications13
DAF-Net: dense attention feature pyramid network for multiscale object detection13
MFAFD: a few-shot learning method for cascading models with parameter free attention and finite discrete space12
An emotion-driven, transformer-based network for multimodal fake news detection12
Multi-scale object detection with feature enhancement for traffic scenes12
Ultra fast-inference depth completion with linear attention-based cascaded hourglass network11
Human action recognition using an optical flow-gated recurrent neural network11
Study of Alzheimer’s disease brain impairment and methods for its early diagnosis: a comprehensive survey10
InceptionDepth-wiseYOLOv2: improved implementation of YOLO framework for pedestrian detection10
Multi-view learning for camouflaged object detection with PVTv210
Concept-based and embedding-based models in lifelog retrieval: an empirical comparison of performance10
Weighted semantic feature based self-supervised deep cross-modal hashing10
Image enhancement with bi-directional normalization and color attention-guided generative adversarial networks10
Organ segmentation from computed tomography images using the 3D convolutional neural network: a systematic review10
A voting-based novel spatio-temporal fusion framework for video saliency using transfer learning mechanism9
Optical music recognition for homophonic scores with neural networks and synthetic music generation9
FOF: a fine-grained object detection and feature extraction end-to-end network9
Maximizing mutual information inside intra- and inter-modality for audio-visual event retrieval9
0.088756084442139