International Journal of Multimedia Information Retrieval

Papers
(The median citation count of International Journal of Multimedia Information Retrieval is 3. The table below lists those papers that are above that threshold based on CrossRef citation counts [max. 250 papers]. The publications cover those that have been published in the past four years, i.e., from 2022-08-01 to 2026-08-01.)
ArticleCitations
Video anomaly detection with memory-guided multilevel embedding103
Multiple object tracking under occlusions based on the stage-wise association strategy with weak cues87
Recent trends in recommender systems: a survey70
VERITE: a Robust benchmark for multimodal misinformation detection accounting for unimodal bias40
Strengthening attention: knowledge distillation via cross-layer feature fusion for image classification36
Enhancing Facial Beauty Prediction via a Dual-Pathway Hybrid Architecture Integrating Vmamba and ViT33
Enhanced YOLOv10 for small object detection with context-aware and adaptive modules32
Optimized data-cube search for enhanced video summarization via shot boundary detection32
VPC-VoxelNet: multi-modal fusion 3D object detection networks based on virtual point clouds30
DELIGHT-Net: DEep and LIGHTweight network to segment Indian text at word level from wild scenic images24
CSAM: Capsule spatial attention mask network for visual question answering22
Prototype local–global alignment network for image–text retrieval22
Multi-objective reinforcement learning for recommender systems: a comprehensive survey of methods, challenges, and future directions21
Feature-NeuS: Neural Implicit Surface Reconstruction Using Feature Multi-View Consistency Constraint21
Hierarchical multi-modal fusion with vision transformers for robust action recognition in infrared-visible videos20
MMDL: a multi-modal deep learning for video highlight detection in sports19
Similarity-based face image retrieval using sparsely embedded deep features and binary code learning18
FiCo-ITR: bridging fine-grained and coarse-grained image-text retrieval for comparative performance analysis16
Human behavior recognition based on DualBiNet model16
Visual and semantic ensemble for scene text recognition with gated dual mutual attention14
A Comprehensive Review of Multimodal Visual Representation Learning: Tracing the Evolution from CNNs to Transformers and Beyond14
CAMIR: fine-tuning CLIP and multi-head cross-attention mechanism for multimodal image retrieval with sketch and text features14
Multimodal music datasets? Challenges and future goals in music processing14
DAF-Net: dense attention feature pyramid network for multiscale object detection14
Generative adversarial networks for 2D-based CNN pose-invariant face recognition13
An emotion-driven, transformer-based network for multimodal fake news detection13
State of art and emerging trends on group recommender system: a comprehensive review13
Multi-scale object detection with feature enhancement for traffic scenes13
MFAFD: a few-shot learning method for cascading models with parameter free attention and finite discrete space12
Ultra fast-inference depth completion with linear attention-based cascaded hourglass network12
Weighted semantic feature based self-supervised deep cross-modal hashing11
FOF: a fine-grained object detection and feature extraction end-to-end network11
Concept-based and embedding-based models in lifelog retrieval: an empirical comparison of performance11
Image enhancement with bi-directional normalization and color attention-guided generative adversarial networks11
Multi-view learning for camouflaged object detection with PVTv211
Human action recognition using an optical flow-gated recurrent neural network11
Optical music recognition for homophonic scores with neural networks and synthetic music generation10
Study of Alzheimer’s disease brain impairment and methods for its early diagnosis: a comprehensive survey10
A Reproducibility Study of Multimodal Embeddings for Recommender Systems10
A voting-based novel spatio-temporal fusion framework for video saliency using transfer learning mechanism9
Zero-shot quantization for object detection via scene-aware synthesis and instance-guided alignment9
Improving skeleton-based action recognition with interactive object information8
Style-aware adversarial pairwise ranking for image recommendation systems8
Stratified Graph Indexing for efficient search in deep descriptor databases8
Maximizing mutual information inside intra- and inter-modality for audio-visual event retrieval8
MCDINO: Self-supervised learning of masks based on combination of multi-path channel attention and local feature weighting8
Enhancing multimodal recommendation via contrastive self-supervised modality-preserving learning7
TCKGE: Transformers with contrastive learning for knowledge graph embedding7
Optimising few-shot class-incremental learning for fine-grained visual recognition7
Few-shot and meta-learning methods for image understanding: a survey7
An interactive attribute-preserving fashion recommendation with 3D image-based virtual try-on7
ETG: the graph convolutional network was enhanced with an EA-transformer for aspect sentiment triplet extraction7
FDAM: full-dimension attention module for deep convolutional neural networks7
Who is gambling? Finding cryptocurrency gamblers using multi-modal retrieval methods6
Joint multi-scale information and long-range dependence for video captioning6
Dual-feature collaborative relation-attention networks for visual question answering6
A hierarchical multi-modal injection architecture for synergistic music understanding and generation5
Partial multimodal hashing with multi-level semantics and adversarial learning5
DMFNet: geometric multi-scale pixel-level contrastive learning for video salient object detection5
Enhancing action recognition via dynamic cross-frame differential modeling5
Deep multimodal learning for time series analysis in social computing: a survey5
$$HF^{2}\text {-}Net$$: hybrid fine-tuning heterogeneous fusion network for visible-infrared person Re-identification5
CoCoOpter: Pre-train, prompt, and fine-tune the vision-language model for few-shot image classification4
Sentiment analysis using deep learning techniques: a comprehensive review4
LG-MLFormer: local and global MLP for image captioning4
Multi-modal emotion recognition using tensor decomposition fusion and self-supervised multi-tasking4
Emotion-aware music tower blocks (EmoMTB ): an intelligent audiovisual interface for music discovery and recommendation4
Image forgery classification and localization through vision transformers4
Similar interior coordination image retrieval with multi-view features4
Ornament image retrieval using few-shot learning4
Special Issue on Open-Domain Image Retrieval in the Wild4
Gender classification from face images using central difference convolutional networks4
ANROT-HELANet: adverserially and naturally robust attention-based aggregation network via the hellinger distance for few-shot classification4
Global and local label-constrained alignment for image-text matching3
A survey of multimodal recommender systems: methods, challenges, and future directions3
Special issue on cross-modal retrieval and analysis3
CLIP-based fusion-modal reconstructing hashing for large-scale unsupervised cross-modal retrieval3
Multi-aware coreference relation network for visual dialog3
Parameter-efficient tuning of cross-modal retrieval for a specific database via trainable textual and visual prompts3
Dual-matrix guided reconstruction hashing for unsupervised cross-modal retrieval3
Deep multiple aggregation networks for action recognition3
Enhancing deep learning image classification using data augmentation and genetic algorithm-based optimization3
Cross-modal alignment with synthetic caption for text-based person search3
A novel method for video shot boundary detection using CNN-LSTM approach3
3D skeleton-based human motion prediction using spatial–temporal graph convolutional network3
H-ARN: A holo-attentive relational network for holistic facial beauty prediction via distribution learning3
Remote Sensing Image Change Captioning: A Comprehensive Review3
A new CNN-based semantic object segmentation for autonomous vehicles in urban traffic scenes3
0.13918113708496