International Journal of Computer Vision

Papers
(The H4-Index of International Journal of Computer Vision is 58. The table below lists those papers that are above that threshold based on CrossRef citation counts [max. 250 papers]. The publications cover those that have been published in the past four years, i.e., from 2022-08-01 to 2026-08-01.)
ArticleCitations
Exploring the Semi-Supervised Video Object Segmentation Problem from a Cyclic Perspective980
Guest Editorial: Special Issue on Open-World Visual Recognition468
Bootstrapping Vision-Language Models for Frequency-Centric Self-Supervised Remote Physiological Measurement444
Common Pole–Polar Properties of Central Catadioptric Sphere and Line Images Used for Camera Calibration397
GenKL: An Iterative Framework for Resolving Label Ambiguity and Label Non-conformity in Web Images Via a New Generalized KL Divergence361
Guest Editorial: Special Issue on Large-Scale Generative Models for Content Creation and Manipulation239
MoDA: Modeling Deformable 3D Objects from Casual Videos214
Multi-Modal Few-Shot Object Detection with Meta-Learning-Based Cross-Modal Prompting213
Robust Averaging using Adaptive Annealing202
Exocentric-to-Egocentric Adaptation for Temporal Action Segmentation with Unlabeled Synchronized Video Pairs200
AutoIT: Automated Image Tagging with Random Perturbation196
Correction: Multi-source-free Domain Adaptive Object Detection195
Invert Your Prompt: Editing-Aware Diffusion Inversion176
Image-based Morphological Characterization of Filamentous Biological Structures with Non-constant Curvature Shape Feature172
Learning Extensible Series-Parallel Lookup Tables for Efficient Image Super-Resolution169
Large-Scale Pre-Trained Models Empowering Phrase Generalization in Temporal Sentence Localization169
View Birdification in the Crowd: Ground-Plane Localization from Perceived Movements162
Instance-dependent Label Distribution Estimation for Learning with Label Noise147
SCT: A Simple Baseline for Parameter-Efficient Fine-Tuning via Salient Channels135
Correction: Hearing and Seeing Through CLIP: A Framework for Self-Supervised Sound Source Localization134
A Minimal Solution for Image-Based Sphere Estimation132
PanAf20K: A Large Video Dataset for Wild Ape Detection and Behaviour Recognition131
Learning Text-to-Video Retrieval from Image Captioning130
Learning Accurate Performance Predictors for Ultrafast Automated Model Compression129
Learning with Enriched Inductive Biases for Vision-Language Models129
Conditional Temporal Variational AutoEncoder for Action Video Prediction127
From Open Set to Closed Set: Supervised Spatial Divide-and-Conquer for Object Counting125
RigNet++: Semantic Assisted Repetitive Image Guided Network for Depth Completion123
BioDrone: A Bionic Drone-Based Single Object Tracking Benchmark for Robust Vision114
EAN: Event Adaptive Network for Enhanced Action Recognition111
UniAttack: Unified Physical-Digital Face Attack Detection103
Weakly Supervised Salient Object Detection with Text Supervision99
Image Synthesis Under Limited Data: A Survey and Taxonomy97
Learning Discriminative Features for Visual Tracking via Scenario Decoupling94
MPANet: Motion Pattern Aggregation Network for Gait Recognition93
Delving Deeper into Anti-Aliasing in ConvNets87
Dissecting Out-of-Distribution Detection and Open-Set Recognition: A Critical Analysis of Methods and Benchmarks87
FastComposer: Tuning-Free Multi-subject Image Generation with Localized Attention86
OpenMonkeyChallenge: Dataset and Benchmark Challenges for Pose Estimation of Non-human Primates85
Are Vision Transformers Robust to Spurious Correlations?80
Guest Editorial: Special Issue on the British Machine Vision Conference 202279
FunnyNet-W: Multimodal Learning of Funny Moments in Videos in the Wild78
NAFT and SynthStab: A RAFT-Based Network and a Synthetic Dataset for Digital Video Stabilization78
Relating View Directions of Complementary-View Mobile Cameras via the Human Shadow77
Learning to Generalize Heterogeneous Representation for Cross-Modality Image Synthesis via Multiple Domain Interventions76
In the Eye of Transformer: Global–Local Correlation for Egocentric Gaze Estimation and Beyond76
Learning Latent Part-Whole Hierarchies for Point Clouds73
Correction: Consistent Prompt Tuning for Generalized Category Discovery71
Bi-calibration Networks for Weakly-Supervised Video Representation Learning71
Vision-Language Alignment Learning Under Affinity and Divergence Principles for Few-Shot Out-of-Distribution Generalization68
Learning Accurate Low-bit Quantization towards Efficient Computational Imaging68
Correction: SOTVerse: A User-Defined Task Space of Single Object Tracking67
CAS-AIR-3D: A Large-scale Low-quality Multi-modal Face Database66
OmniDrag: Enabling Motion Control for Omnidirectional Image-to-Video Generation64
Symmetria: A Synthetic Dataset for Learning in Point Clouds61
Dynamic Knowledge Transfer for Mitigating Spurious Correlations in Deep Learning60
Skeleton Ground Truth Extraction: Methodology, Annotation Tool and Benchmarks58
UMSCS: A Novel Unpaired Multimodal Image Segmentation Method Via Cross-Modality Generative and Semi-supervised Learning58
0.22523403167725