IEEE Transactions on Multimedia

Papers
(The TQCC of IEEE Transactions on Multimedia is 16. The table below lists those papers that are above that threshold based on CrossRef citation counts [max. 250 papers]. The publications cover those that have been published in the past four years, i.e., from 2022-08-01 to 2026-08-01.)
ArticleCitations
Improving Vision Anomaly Detection With the Guidance of Language Modality1006
Focusing on Subtle Differences: A Feature Disentanglement Model for Series Photo Selection547
Rethinking Video Sentence Grounding From a Tracking Perspective With Memory Network and Masked Attention398
Rethinking Affine Transform for Efficient Image Enhancement: A Color Space Perspective366
FoodSAM: Any Food Segmentation261
Online Low-Light Sand-Dust Video Enhancement Using Adaptive Dynamic Brightness Correction and a Rolling Guidance Filter228
Simulate, Refocus and Ensemble: An Attention-Refocusing Scheme for Domain Generalization207
SGG-Nets: Generic Rotation-Invariant Plugin Networks for Point Cloud Analysis204
ViDR-GNN: Vision Implicit Discriminative Reorganization Graph Neural Networks202
Dual-Task Mutual Reinforcing Embedded Joint Video Paragraph Retrieval and Grounding196
Few-Shot Generative Model Adaptation via Style-Guided Prompt195
Weakly-Supervised Video Object Grounding via Learning Uni-Modal Associations191
Towards Substation Semantic Segmentation: A benchmark dataset and a cross-attention embedded hierarchical network184
HRVFusion: Video-based Long-Term Heart Rate Variability Measurement with Conditional Diffusion Models168
Revisiting the Adversarial Transferability: Towards a Perspective of Semantic Preservation162
Exploring Kernel Transformations for Implicit Neural Representations160
Posture-Movement-Frequency-Enhanced Graph Convolutional Network for Gait Emotion Recognition158
Mask-Aware Kernel Learning for Action Recognition156
LMAgent: A Large-scale Multimodal Agents Society for Multi-user Simulation156
Bias-Correction Feature Learner for Semi-Supervised Instance Segmentation155
Mix-Based Training Strategies for Learning Implicit Neural Representations154
Bidirectional Translation Between UHD-HDR and HD-SDR Videos153
Optimal Transport-Based Patch Matching for Image Style Transfer151
Neighborhood Contrastive Transformer for Change Captioning151
Robust Multi-Stage Tracking via Multi-Scale and Multi-Level Representation Learning144
Watch Where You Move: Region-Aware Dynamic Aggregation and Excitation for Gait Recognition142
PropMambaSR: Lightweight Image Super-Resolution with Propagation State Space Model139
Adaptive Weight Generator for Multi-Task Image Recognition by Task Grouping Prompt138
Semantic-Aware Triplet Loss for Image Classification136
Semantic Dual-Adversarial Network for Blended-Target Domain Adaptation136
DWSF-Net: A Dynamic Wavelet-based Spatial-frequency Fusion Network for Multispectral Object Detection135
AMS-Net: Adaptive Multi-Scale Network for Image Compressive Sensing135
Late Fusion Multiple Kernel Clustering With Local Kernel Alignment Maximization134
Disaggregation Distillation for Person Search132
Rényi Entropy Induced Efficient and Balanced One-Step Multi-View Clustering132
Vision-Controllable Language Model for Image-Guided Story Ending Generation128
Multi-Level Transitional Contrast Learning for Personalized Image Aesthetics Assessment125
Semi-Supervised Contrastive Learning With Similarity Co-Calibration123
Distributed Deep Point Cloud Feature Compression for Vehicle-to-Vehicle Cooperative Perception120
Guided Image-to-Image Translation by Discriminator-Generator Communication119
Weakly-Supervised 3D Visual Grounding Based on Visual Language Alignment119
One-Shot Human Motion Transfer via Occlusion-Robust Flow Prediction and Neural Texturing118
MHRN: A Multimodal Hierarchical Reasoning Network for Topic Detection115
SCSP: An Unsupervised Image-to-Image Translation Network Based on Semantic Cooperative Shape Perception115
Unsupervised Learning-Based Framework for Deepfake Video Detection114
BMB: Balanced Memory Bank for Long-Tailed Semi-Supervised Learning114
Long Video Understanding With Learnable Retrieval in Video-Language Models113
Efficient Cross-Modal Video Retrieval With Meta-Optimized Frames112
Transferable Backdoor Attack on Any CLIP Model With Any Target Class by Pre-Trained Hack Network111
Quality Assessment for DIBR-Synthesized Views Based on Wavelet Transform and Gradient Magnitude Similarity111
Asymptotics-Aware Multi-View Subspace Clustering110
Structured Graph Reasoning for Traffic Anomaly Detection108
Self-Guided Discriminative Locality Preserving Projections107
Vulnerability of Feature Extractors in 2D Image-Based 3D Object Retrieval105
Dynamic Mosaics: Saliency-Guided Adaptive Masking for Occluded Person Re-Identification104
MGKsite: Multi-Modal Knowledge-Driven Site Selection via Intra and Inter-Modal Graph Fusion104
Beyond Simple Extraction: Unleashing the Potential of Encoder Interaction in Few-Shot Segmentation103
Interpretable Graph Convolutional Network for Multi-View Semi-Supervised Learning102
TFBF: Temporal-Frequency Bidirectional Fusion for Action Quality Assessment100
Outliers Adaptation Exploration and Centroids Matching Label Refinement for Unsupervised Person Re-Identification100
GLCT: A Novel Global-Local Constraint for Unpaired Image-to-Image Translation100
SkyML: A MLaaS Federation Design for Multicloud-Based Multimedia Analytics98
Cross-modal Semantic Relevance is An Efficient Gatekeeper for Audio-Visual Video Parsing98
ICE: Interactive 3D Game Character Facial Editing via Dialogue97
Siamese Alignment Network for Weakly Supervised Video Moment Retrieval96
MVPC-CLIP: Multi-Granularity Visual Prompt Co-Operative for Aerial Video Recognition96
Anomaly-Led Prompting Learning Caption Generating Model and Benchmark96
Ensemble Prototype Networks for Unsupervised Cross-Modal Hashing With Cross-Task Consistency96
Distilling Multi-View Diffusion Models Into 3D Generators94
Skeleton-Based Action Recognition With Select-Assemble-Normalize Graph Convolutional Networks94
Adversarial 3D-to-Real Watermarking: Revealing Invisible Messages Hidden in Complexly Distorted Surfaces93
BASNet: Boundary Assisted Network for Image Splicing Forgery Detection91
Scale Up Composed Image Retrieval Learning via Modification Text Generation91
Self-Mining the Confident Prototypes for Source-Free Unsupervised Domain Adaptation in Image Segmentation90
Pixel Bleach Network for Detecting Face Forgery Under Compression90
$\rm {M}^{2}\rm {C-EvDet}$: Multi-Domain Multi-Order Cross-Modal Knowledge Distillation for Event-based Object Detection88
3D-SceneQ: Empowering 3D LLM With Query-Guided Adaptive Pruning and Multi-Modal Feature Enhancement88
Progressive Local Filter Pruning for Image Retrieval Acceleration87
XMusic: Towards a Generalized and Controllable Symbolic Music Generation Framework86
Disentangled Graph Variational Auto-Encoder for Multimodal Recommendation With Interpretability85
Semi-Supervised Domain Adaptation for Major Depressive Disorder Detection84
Dynamic Contrastive Distillation for Image-Text Retrieval83
Feature First: Advancing Image-Text Retrieval Through Improved Visual Features83
Deep Semantic-Consistent Penalizing Hashing for Cross-Modal Retrieval82
Improving Pre-Trained Model-Based Speech Emotion Recognition From a Low-Level Speech Feature Perspective82
Semi-Supervised Domain Adaptation via Joint Transductive and Inductive Subspace Learning82
Perceptual Image Hashing Using Feature Fusion of Orthogonal Moments80
CVKD-UDA: Cross-View Knowledge Distillation for 3D Unsupervised Domain Adaptive Segmentation79
Hear Me, See Me, Understand Me: Audio-Visual Autism Behavior Recognition79
Hierarchical Equalization Loss for Long-Tailed Instance Segmentation78
Adaptive HEVC Video Steganography With High Performance Based on Attention-Net and PU Partition Modes77
PhotoHelper: Portrait Photographing Guidance Via Deep Feature Retrieval and Fusion77
SLCGC: A lightweight Self-supervised Low-Pass Contrastive Graph Clustering Network for Hyperspectral Images77
Towards Neural Codec-Empowered 360$^\circ$ Video Streaming: A Saliency-Aided Synergistic Approach74
Towards Temporal Event Detection: A Dataset, Benchmarks and Challenges74
Image-Based Structured Vehicle Behavior Analysis Inspired by Interactive Cognition73
ALCER3D: Adaptive Learning Constraints for Enhanced Retrieval of Complex Indoor 3D Scenarios72
Cps-STS: Bridging the Gap Between Content and Position for Coarse-Point-Supervised Scene Text Spotter72
Supervised Contrastive Learning for Indoor Point Cloud Oversegmentation71
DEHand: Deformable Encoding for Photo-Realistic Free-View and Free-Pose Hand Rendering70
RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios70
JPEG AI Compressed Domain Face Detection: A Multi-Scale Bridging Perspective68
Reconstructed Graph Constrained Auto-Encoders for Multi-View Representation Learning68
Cross-Domain Sample Relationship Learning for Facial Expression Recognition68
Semi-Supervised Authentically Distorted Image Quality Assessment With Consistency-Preserving Dual-Branch Convolutional Neural Network68
Depth Map Super-Resolution via Deep Cross-Modality and Cross-Scale Guidance67
Foodfusion: A Novel Approach for Food Image Composition via Diffusion Models67
Motion Direction Awareness: A Biomimetic Dynamic Capture Mechanism for Video Prediction67
Rate-Adaptive Neural Network for Image Compressive Sensing66
High Specificity Guided Cross-Domain Few-Shot Segmentation66
Vulnerabilities in AI-Generated Image Detection: The Challenge of Adversarial Attacks66
RUL: Region Uncertainty Learning for Robust Face Recognition66
Enhanced Context Mining and Filtering for Learned Video Compression65
Investigating the Effective Dynamic Information of Spectral Shapes for Audio Classification65
Improving Out-of-Distribution Generalization on Point Clouds with Cross-Domain Adversarial Distillation65
Boosting Universal Adversarial Attack on Deep Neural Networks65
Video Instance Segmentation by Instance Flow Assembly64
Reliable Multi-View Clustering with Graph Neural Network64
Human-Centric Behavior Description in Videos: New Benchmark and Model63
Cooperative Bargaining Game Based Adaptive Video Multicast Over Mobile Edge Networks63
Saliency-Aware Adversarial Attacks on Visual Trackers63
Personalized Fashion Recommendation With Discrete Content-Based Tensor Factorization62
EPM-Net: Efficient Feature Extraction, Point-Pair Feature Matching for Robust 6-D Pose Estimation62
Denoised Semantic Features for Local Consistent No-Reference Image Quality Assessment62
Multimodal Progressive Modulation Network for Micro-Video Multi-Label Classification61
Spatial-Temporal Saliency Guided Unbiased Contrastive Learning for Video Scene Graph Generation61
A Multidimensional Media Adaptation Framework for Live Holographic Communication61
Enhancing Representation Inversion and Alignment for Zero-Shot Composed Image Retrieval60
Can Machines Generate Personalized Music? A Hybrid Favorite-Aware Method for User Preference Music Transfer60
MMIFN: A Multi-Modal Interactive Fusion Network for Omnidirectional Image Quality Assessment60
CMANet: Context-Aware Mutual Attention Network for Referring Image Segmentation60
SDE2D: Semantic-Guided Discriminability Enhancement Feature Detector and Descriptor60
Exploring Kernel-Based Texture Transfer for Pose-Guided Person Image Generation60
Wavelet-Domain Masked Image Modeling for Color-Consistent HDR Video Reconstruction59
DREAMT: Diversity Enlarged Mutual Teaching for Unsupervised Domain Adaptive Person Re-Identification59
Dynamic Strategy Prompt Reasoning for Emotional Support Conversation58
Exploring Basic Expression Representation for Compound Facial Expression Recognition58
Look&listen: Multi-Modal Correlation Learning for Active Speaker Detection and Speech Enhancement58
Compositional Text-to-Image Synthesis With Training-Free Layout-Guided Diffusion58
Velocity First? Rethinking 3D Object Detection with 4D Millimeter Wave Radar58
Show, Tell and Rephrase: Diverse Video Captioning via Two-Stage Progressive Training58
Bidirectional Prototype-Reward Co-Evolution for Test-Time Adaptation of Vision-Language Models57
Prototypical Bidirectional Adaptation and Learning for Cross-Domain Semantic Segmentation57
Generalizing Beyond Patterns: Dynamic Moment Query Recalibrating for Out-of-Distribution Video Temporal Localization57
TPE-ADE: Thumbnail-Preserving Encryption Based on Adaptive Deviation Embedding for JPEG Images56
Exploring Local and Global Consistent Correlation on Hypergraph for Rotation Invariant Point Cloud Analysis56
HP-C4D: A Fast Camera and 4D Radar Fusion Framework With Height Prediction for 3D Object Detection55
UniCrossGait: Unified Cross-Modal Gait Recognition Based on Knowledge Distillation55
VOLTER: Visual Collaboration and Dual-Stream Fusion for Scene Text Recognition55
Motion Deblur by Learning Residual From Events55
STNet: Scale Tree Network With Multi-Level Auxiliator for Crowd Counting54
HSV-Driven Illumination-Aware Iterative Network for Unsupervised Low-Light Enhancement54
FedSH: Towards Privacy-Preserving Text-Based Person Re-Identification54
CRSOT: Cross-Resolution Object Tracking Using Unaligned Frame and Event Cameras54
Sparse Transformer for Ultra-Sparse Sampled Video Compressive Sensing54
Universal Infrared Image Nonuniformity Correction via Stripe-Aware Attention Network54
RA-SSU: Towards Fine-Grained Audio-Visual Learning with Region-Aware Sound Source Understanding54
A Two-Stream Hybrid Convolution-Transformer Network Architecture for Clothing-Change Person Re-Identification54
RetinexGS: Enhancing 3D Gaussian Splatting for Low-Light Scenes Via Retinex-Guided Decomposition54
Deep Unfolding Network for Image Compressed Sensing by Content-Adaptive Gradient Updating and Deformation-Invariant Non-Local Modeling54
Sounding Depressed? Personalized Deep Learning Model for Depression Detection From Speech and Text53
Primary Code Guided Targeted Attack against Cross-modal Hashing Retrieval53
Sentiment-Enhanced Graph-Based Sarcasm Explanation in Dialogue53
Ocean's Duality: Physics-Driven Dual-Branch Framework for Underwater Vision Enhancement53
Twin Tensor Learning for Consistency and Inconsistency: A Unified Affinity Learning Framework for Multi-View Clustering53
RSNet: Relation Separation Network for Few-Shot Similar Class Recognition53
Multi-View User Preference Modeling for Personalized Text-to-Image Generation52
Pedestrian Trajectory Prediction Based on Social Interactions Learning With Random Weights52
No-Reference Bitstream-Layer Model for Perceptual Quality Assessment of V-PCC Encoded Point Clouds52
Video-to-Music Recommendation Using Temporal Alignment of Segments52
Action-Responsive Contrastive Network for Fine-Grained Skeleton-Based Action Recognition52
RA-SCIC: Region-Aware Screen Content Image Coding Towards Fidelity and Efficiency52
Test-Time Model Adaptation for Visual Question Answering With Debiased Self-Supervisions51
IEIRNet: Inconsistency Exploiting Based Identity Rectification for Face Forgery Detection51
Style-Agnostic Representation Learning for Visible-Infrared Person Re-Identification51
Synthesize Boundaries: A Boundary-Aware Self-Consistent Framework for Weakly Supervised Salient Object Detection51
Probabilistic Temporal Masked Attention for Cross-View Online Action Detection51
C-CTX: Cubic-Checkerboard Context Entropy Model for Learned Image Compression50
Knowledge Distillation-Based Domain-Invariant Representation Learning for Domain Generalization50
REDEditing: Relationship-Driven Precise Backdoor Poisoning on Text-to-Image Diffusion Models50
High Fidelity Face-Swapping With Style ConvTransformer and Latent Space Selection50
SegTrans: Transferable Adversarial Examples for Segmentation Models50
Anchor-Guided Discrete Multi-View Clustering50
FFFN: Frame-By-Frame Feedback Fusion Network for Video Super-Resolution50
Dual Representation Aggregation Network for Blind Image Super-Resolution via Iterative Bi-level Optimization50
Benchmark Dataset and Pair-Wise Ranking Method for Quality Evaluation of Night-Time Image Enhancement49
RD-VTA: Rule-Data Guided Video-to-Audio Generation for Fine-Grained Footstep Sound49
Improving Fine-Grained Image Classification With Multimodal Information49
GLFF: Global and Local Feature Fusion for AI-Synthesized Image Detection49
Multimodal Sentiment Analysis With Image-Text Interaction Network49
Low-Light Image Enhancement via Self-Reinforced Retinex Projection Model49
Exploiting EfficientSAM and Temporal Coherence for Audio-Visual Segmentation49
Underwater Image Enhancement With Cascaded Contrastive Learning49
DA-Net: Density-Aware 3D Object Detection Network for Point Clouds49
FGDNet: Fine-Grained Detection Network Towards Face Anti-Spoofing48
Reordered $k$-Means: A New Baseline for View-Unaligned Multi-View Clustering48
MGHead: Motion-Aware Animated Gaussian Head Avatars With Anchored Skeletal Structures48
Augment One With Others: Generalizing to Unforeseen Variations for Visual Tracking48
OpenSlot: Mixed Open-Set Recognition With Object-Centric Learning48
Towards Region-Aware Finer Self-Supervised Learning for Fine-Grained Visual Recognition48
Rethinking the Role of Vector Quantization for Blind Image Restoration47
FOF-X: Towards Real-time Detailed Human Reconstruction from a Single Image47
DDGA: Domain Distance Guided Active Domain Adaptation for Unpaired Super-Resolution47
Dense Video Captioning With Early Linguistic Information Fusion47
Blind Video Quality Assessment at the Edge47
Unleash the Power of Vision-Language Models by Visual Attention Prompt and Multimodal Interaction47
Flow Guidance Deformable Compensation Network for Video Frame Interpolation47
DIP: Diffusion Learning of Inconsistency Pattern for General DeepFake Detection47
CenterTube: Tracking Multiple 3D Objects With 4D Tubelets in Dynamic Point Clouds47
Bridging the Short-Term and Long-Term Gap: A Cross-Task Continuous Learning Person Re-Identification Problem47
Multi-View Depth Estimation With Uncertainty Constraints for Virtual-Real Occlusion47
Semantics Alternating Enhancement and Bidirectional Aggregation for Referring Video Object Segmentation47
Graph Convolutional Network With Unknown Class Number46
RaFPN: Relation-Aware Feature Pyramid Network for Dense Image Prediction46
Unsupervised Deepfake Detection via Camera Source Clustering and Temporal-Spatial Features46
Exploring Cross-Modal Mutual Prompt Learning for Video Quality Assessment46
VRTNet: Vector Rectifier Transformer for Two-View Correspondence Learning46
Compression of Plenoptic Point Cloud Attributes Using 6-D Point Clouds and 6-D Transforms46
SwimVG: Step-Wise Multimodal Fusion and Adaption for Visual Grounding46
Category-Contrastive Fine-Grained Crowd Counting and Beyond46
Cross-Modality Feature Fusion for Forward-Looking Sonar Image Segmentation in Complex Underwater Environments46
Simultaneously Training and Compressing Vision-and-Language Pre-Training Model46
Inexactly Matched Referring Expression Comprehension With Rationale46
Scene Graph Knowledge Enhanced Hashing with Contrastive Learning for Image-Text Retrieval46
Soundscape Captioning Using Sound Affective Quality Network and Large Language Model46
Underwater Adaptive Video Transmissions Using MIMO-Based Software-Defined Acoustic Modems46
SSPNet: Predicting Visual Saliency Shifts46
Edge-Assisted Massive Video Delivery Over Cell-Free Massive MIMO46
Towards a Multi-Granulated Statistical Framework for Human–Machine Collaboration in Image Classification46
Tensorformer: Normalized Matrix Attention Transformer for High-Quality Point Cloud Reconstruction46
Progressive Learning Model for Big Data Analysis Using Subnetwork and Moore-Penrose Inverse45
Point Cloud Soft Multicast for Untethered XR Users45
Visibility-Based Geometry Pruning of Neural Plenoptic Scene Representations45
Instruction-Driven 3D Facial Expression Generation and Transition45
Noise Aware Audio-Visual Speech Denoising45
MVL-Net: Pairwise Learning for Multi-View Multiple People Labelling45
Tuning-Free High-Resolution Video Diffusion With Spatial-Temporal Latent Grouping45
Interpretable Multi-View Representation Learning Towards Complex Scenes: From Homogeneity to Heterogeneity45
Question Understanding and Temporality Guiding for Video Question Answering45
Progressive Learning of Instance-Level Proxy Semantics for Few-Shot Action Recognition45
CMI-Net: Cross-View Message Token Interaction Network for 3D Shape Recognition45
CNIE: Content-Aware Non-Transferable Information Extraction for Fine-Grained Visual Categorization44
Neural-Enhanced Rate Adaptation and Computation Distribution for Emerging mmWave Multi-User 3D Video Streaming Systems44
CLCT: Complementary Local Consensus Transformer for Two-View Correspondence Pruning44
Bidirectional Maximum Entropy Training With Word Co-Occurrence for Video Captioning44
Indistinguishability Analysis of JPEG Image Encryption Schemes44
Develop Then Rival: A Human Vision-Inspired Framework for Superimposed Image Decomposition44
MPPM: A Mobile-Efficient Part Model for Object re-ID44
Low-Light Image Enhancement With SAM-Based Structure Priors and Guidance44
Face De-Occlusion With Deep Cascade Guidance Learning44
1.2189140319824