Saved in:
| Main Authors: | Biswas, Dipayan, Shah, Shishir, Subhlok, Jaspal |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2506.13657 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Visual Content Detection in Educational Videos with Transfer Learning and Dataset Enrichment
by: Biswas, Dipayan, et al.
Published: (2025)
by: Biswas, Dipayan, et al.
Published: (2025)
VideoOrion: Tokenizing Object Dynamics in Videos
by: Feng, Yicheng, et al.
Published: (2024)
by: Feng, Yicheng, et al.
Published: (2024)
MaskVD: Region Masking for Efficient Video Object Detection
by: Sarkar, Sreetama, et al.
Published: (2024)
by: Sarkar, Sreetama, et al.
Published: (2024)
Show, Don't Tell: Detecting Novel Objects by Watching Human Videos
by: Akl, James, et al.
Published: (2026)
by: Akl, James, et al.
Published: (2026)
From Label Error Detection to Correction: A Modular Framework and Benchmark for Object Detection Datasets
by: Penquitt, Sarina, et al.
Published: (2025)
by: Penquitt, Sarina, et al.
Published: (2025)
Pack and Detect: Fast Object Detection in Videos Using Region-of-Interest Packing
by: Kumar, Athindran Ramesh, et al.
Published: (2018)
by: Kumar, Athindran Ramesh, et al.
Published: (2018)
Object-Centric Diffusion for Efficient Video Editing
by: Kahatapitiya, Kumara, et al.
Published: (2024)
by: Kahatapitiya, Kumara, et al.
Published: (2024)
Object-Centric Cropping for Visual Few-Shot Classification
by: Abdali, Aymane, et al.
Published: (2025)
by: Abdali, Aymane, et al.
Published: (2025)
Reconstruction-Guided Slot Curriculum: Addressing Object Over-Fragmentation in Video Object-Centric Learning
by: Moon, WonJun, et al.
Published: (2026)
by: Moon, WonJun, et al.
Published: (2026)
Correlation of Object Detection Performance with Visual Saliency and Depth Estimation
by: Bartolo, Matthias, et al.
Published: (2024)
by: Bartolo, Matthias, et al.
Published: (2024)
VONet: Unsupervised Video Object Learning With Parallel U-Net Attention and Object-wise Sequential VAE
by: Yu, Haonan, et al.
Published: (2024)
by: Yu, Haonan, et al.
Published: (2024)
Leveraging Transformers for Weakly Supervised Object Localization in Unconstrained Videos
by: Murtaza, Shakeeb, et al.
Published: (2024)
by: Murtaza, Shakeeb, et al.
Published: (2024)
Extending Dataset Pruning to Object Detection: A Variance-based Approach
by: Yagi, Ryota
Published: (2025)
by: Yagi, Ryota
Published: (2025)
VILOD: A Visual Interactive Labeling Tool for Object Detection
by: Holm, Isac
Published: (2025)
by: Holm, Isac
Published: (2025)
GraphKD: Exploring Knowledge Distillation Towards Document Object Detection with Structured Graph Creation
by: Banerjee, Ayan, et al.
Published: (2024)
by: Banerjee, Ayan, et al.
Published: (2024)
Moving Object Proposals with Deep Learned Optical Flow for Video Object Segmentation
by: Shi, Ge, et al.
Published: (2024)
by: Shi, Ge, et al.
Published: (2024)
CMTM: Cross-Modal Token Modulation for Unsupervised Video Object Segmentation
by: Jeon, Inseok, et al.
Published: (2026)
by: Jeon, Inseok, et al.
Published: (2026)
ROSE: Remove Objects with Side Effects in Videos
by: Miao, Chenxuan, et al.
Published: (2025)
by: Miao, Chenxuan, et al.
Published: (2025)
Reasoning-Enhanced Object-Centric Learning for Videos
by: Li, Jian, et al.
Published: (2024)
by: Li, Jian, et al.
Published: (2024)
Learning to Detect Label Errors by Making Them: A Method for Segmentation and Object Detection Datasets
by: Penquitt, Sarina, et al.
Published: (2025)
by: Penquitt, Sarina, et al.
Published: (2025)
Model-Agnostic Open-Set Air-to-Air Visual Object Detection for Reliable UAV Perception
by: Loukovitis, Spyridon, et al.
Published: (2025)
by: Loukovitis, Spyridon, et al.
Published: (2025)
VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM
by: Yuan, Yuqian, et al.
Published: (2024)
by: Yuan, Yuqian, et al.
Published: (2024)
Temporal Object Captioning for Street Scene Videos from LiDAR Tracks
by: Gopinathan, Vignesh, et al.
Published: (2025)
by: Gopinathan, Vignesh, et al.
Published: (2025)
Saliency-Motion Guided Trunk-Collateral Network for Unsupervised Video Object Segmentation
by: Zheng, Xiangyu, et al.
Published: (2025)
by: Zheng, Xiangyu, et al.
Published: (2025)
DDLP: Unsupervised Object-Centric Video Prediction with Deep Dynamic Latent Particles
by: Daniel, Tal, et al.
Published: (2023)
by: Daniel, Tal, et al.
Published: (2023)
Object-Centric Learning for Real-World Videos by Predicting Temporal Feature Similarities
by: Zadaianchuk, Andrii, et al.
Published: (2023)
by: Zadaianchuk, Andrii, et al.
Published: (2023)
P2ANet: A Dataset and Benchmark for Dense Action Detection from Table Tennis Match Broadcasting Videos
by: Bian, Jiang, et al.
Published: (2022)
by: Bian, Jiang, et al.
Published: (2022)
Towards Robust Cross-Dataset Object Detection Generalization under Domain Specificity
by: Chakraborty, Ritabrata, et al.
Published: (2026)
by: Chakraborty, Ritabrata, et al.
Published: (2026)
RefDecoder: Enhancing Visual Generation with Conditional Video Decoding
by: Fan, Xiang, et al.
Published: (2026)
by: Fan, Xiang, et al.
Published: (2026)
Multimodal Object Query Initialization for 3D Object Detection
by: van Geerenstein, Mathijs R., et al.
Published: (2023)
by: van Geerenstein, Mathijs R., et al.
Published: (2023)
CinePile: A Long Video Question Answering Dataset and Benchmark
by: Rawal, Ruchit, et al.
Published: (2024)
by: Rawal, Ruchit, et al.
Published: (2024)
360VFI: A Dataset and Benchmark for Omnidirectional Video Frame Interpolation
by: Lu, Wenxuan, et al.
Published: (2024)
by: Lu, Wenxuan, et al.
Published: (2024)
Video Editing for Audio-Visual Dubbing
by: Manela, Binyamin, et al.
Published: (2025)
by: Manela, Binyamin, et al.
Published: (2025)
DoughNet: A Visual Predictive Model for Topological Manipulation of Deformable Objects
by: Bauer, Dominik, et al.
Published: (2024)
by: Bauer, Dominik, et al.
Published: (2024)
Closing the Visual Sim-to-Real Gap with Object-Composable NeRFs
by: Mishra, Nikhil, et al.
Published: (2024)
by: Mishra, Nikhil, et al.
Published: (2024)
General and Efficient Visual Goal-Conditioned Reinforcement Learning using Object-Agnostic Masks
by: Shahriar, Fahim, et al.
Published: (2025)
by: Shahriar, Fahim, et al.
Published: (2025)
ImpliHateVid: A Benchmark Dataset and Two-stage Contrastive Learning Framework for Implicit Hate Speech Detection in Videos
by: Rehman, Mohammad Zia Ur, et al.
Published: (2025)
by: Rehman, Mohammad Zia Ur, et al.
Published: (2025)
Unifying Specialized Visual Encoders for Video Language Models
by: Chung, Jihoon, et al.
Published: (2025)
by: Chung, Jihoon, et al.
Published: (2025)
Comparing Surface Landmine Object Detection Models on a New Drone Flyby Dataset
by: Agrawal-Chung, Navin, et al.
Published: (2024)
by: Agrawal-Chung, Navin, et al.
Published: (2024)
Coreset Selection for Object Detection
by: Lee, Hojun, et al.
Published: (2024)
by: Lee, Hojun, et al.
Published: (2024)
Similar Items
-
Visual Content Detection in Educational Videos with Transfer Learning and Dataset Enrichment
by: Biswas, Dipayan, et al.
Published: (2025) -
VideoOrion: Tokenizing Object Dynamics in Videos
by: Feng, Yicheng, et al.
Published: (2024) -
MaskVD: Region Masking for Efficient Video Object Detection
by: Sarkar, Sreetama, et al.
Published: (2024) -
Show, Don't Tell: Detecting Novel Objects by Watching Human Videos
by: Akl, James, et al.
Published: (2026) -
From Label Error Detection to Correction: A Modular Framework and Benchmark for Object Detection Datasets
by: Penquitt, Sarina, et al.
Published: (2025)