HierarQ: Task-Aware Hierarchical Q-Former for Enhanced Video Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Azad, Shehreen, Vineet, Vibhav, Rawat, Yogesh Singh |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DisenQ: Disentangling Q-Former for Activity-Biometrics
by: Azad, Shehreen, et al.
Published: (2025)
by: Azad, Shehreen, et al.
Published: (2025)
StreamReady: Learning What to Answer and When in Long Streaming Videos
by: Azad, Shehreen, et al.
Published: (2026)
by: Azad, Shehreen, et al.
Published: (2026)
Understanding Depth and Height Perception in Large Visual-Language Models
by: Azad, Shehreen, et al.
Published: (2024)
by: Azad, Shehreen, et al.
Published: (2024)
Activity-Biometrics: Person Identification from Daily Activities
by: Azad, Shehreen, et al.
Published: (2024)
by: Azad, Shehreen, et al.
Published: (2024)
Robustness Analysis on Foundational Segmentation Models
by: Schiappa, Madeline Chantry, et al.
Published: (2023)
by: Schiappa, Madeline Chantry, et al.
Published: (2023)
On Occlusions in Video Action Detection: Benchmark Datasets And Training Recipes
by: Modi, Rajat, et al.
Published: (2024)
by: Modi, Rajat, et al.
Published: (2024)
Navigating Hallucinations for Reasoning of Unintentional Activities
by: Grover, Shresth, et al.
Published: (2024)
by: Grover, Shresth, et al.
Published: (2024)
A Large-Scale Analysis on Contextual Self-Supervised Video Representation Learning
by: Kumar, Akash, et al.
Published: (2025)
by: Kumar, Akash, et al.
Published: (2025)
OmViD: Omni-supervised active learning for video action detection
by: Rana, Aayush, et al.
Published: (2025)
by: Rana, Aayush, et al.
Published: (2025)
Probing Conceptual Understanding of Large Visual-Language Models
by: Schiappa, Madeline, et al.
Published: (2023)
by: Schiappa, Madeline, et al.
Published: (2023)
CoSPlan: Corrective Sequential Planning via Scene Graph Incremental Updates
by: Grover, Shresth, et al.
Published: (2025)
by: Grover, Shresth, et al.
Published: (2025)
Physics Knowledge in Frontier Models: A Diagnostic Study of Failure Modes
by: Bagdonaviciute, Ieva, et al.
Published: (2025)
by: Bagdonaviciute, Ieva, et al.
Published: (2025)
Stable Mean Teacher for Semi-supervised Video Action Detection
by: Kumar, Akash, et al.
Published: (2024)
by: Kumar, Akash, et al.
Published: (2024)
Coarse Attribute Prediction with Task Agnostic Distillation for Real World Clothes Changing ReID
by: Pathak, Priyank, et al.
Published: (2025)
by: Pathak, Priyank, et al.
Published: (2025)
Contextual Self-paced Learning for Weakly Supervised Spatio-Temporal Video Grounding
by: Kumar, Akash, et al.
Published: (2025)
by: Kumar, Akash, et al.
Published: (2025)
Foundation Models for Video Understanding: A Survey
by: Madan, Neelu, et al.
Published: (2024)
by: Madan, Neelu, et al.
Published: (2024)
Enhancing Diffusion Face Generation with Contrastive Embeddings and SegFormer Guidance
by: Rawat, Dhruvraj Singh, et al.
Published: (2025)
by: Rawat, Dhruvraj Singh, et al.
Published: (2025)
PEEKABOO: Interactive Video Generation via Masked-Diffusion
by: Jain, Yash, et al.
Published: (2023)
by: Jain, Yash, et al.
Published: (2023)
Scaling Open-Vocabulary Action Detection
by: Sia, Zhen Hao, et al.
Published: (2025)
by: Sia, Zhen Hao, et al.
Published: (2025)
Q-Bench-Video: Benchmarking the Video Quality Understanding of LMMs
by: Zhang, Zicheng, et al.
Published: (2024)
by: Zhang, Zicheng, et al.
Published: (2024)
Semi-supervised Active Learning for Video Action Detection
by: Singh, Ayush, et al.
Published: (2023)
by: Singh, Ayush, et al.
Published: (2023)
Grounding Task Assistance with Multimodal Cues from a Single Demonstration
by: Sarch, Gabriel, et al.
Published: (2025)
by: Sarch, Gabriel, et al.
Published: (2025)
Asynchronous Perception Machine For Efficient Test-Time-Training
by: Modi, Rajat, et al.
Published: (2024)
by: Modi, Rajat, et al.
Published: (2024)
MolVision: Molecular Property Prediction with Vision Language Models
by: Adak, Deepan, et al.
Published: (2025)
by: Adak, Deepan, et al.
Published: (2025)
iSafetyBench: A video-language benchmark for safety in industrial environment
by: Abdullah, Raiyaan, et al.
Published: (2025)
by: Abdullah, Raiyaan, et al.
Published: (2025)
Colors See Colors Ignore: Clothes Changing ReID with Color Disentanglement
by: Pathak, Priyank, et al.
Published: (2025)
by: Pathak, Priyank, et al.
Published: (2025)
DIFFER: Disentangling Identity Features via Semantic Cues for Clothes-Changing Person Re-ID
by: Liang, Xin, et al.
Published: (2025)
by: Liang, Xin, et al.
Published: (2025)
ProDiG: Progressive Diffusion-Guided Gaussian Splatting for Aerial to Ground Reconstruction
by: Mitra, Sirshapan, et al.
Published: (2026)
by: Mitra, Sirshapan, et al.
Published: (2026)
Q-Former Autoencoder: A Modern Framework for Medical Anomaly Detection
by: Dalmonte, Francesco, et al.
Published: (2025)
by: Dalmonte, Francesco, et al.
Published: (2025)
EZ-CLIP: Efficient Zeroshot Video Action Recognition
by: Ahmad, Shahzad, et al.
Published: (2023)
by: Ahmad, Shahzad, et al.
Published: (2023)
Task-adaptive Q-Face
by: Sun, Haomiao, et al.
Published: (2024)
by: Sun, Haomiao, et al.
Published: (2024)
MM-GEN: Enhancing Task Performance Through Targeted Multimodal Data Curation
by: Joshi, Siddharth, et al.
Published: (2025)
by: Joshi, Siddharth, et al.
Published: (2025)
OS-Marathon: Benchmarking Computer-Use Agents on Long-Horizon Repetitive Tasks
by: Wu, Jing, et al.
Published: (2026)
by: Wu, Jing, et al.
Published: (2026)
GaitCrafter: Diffusion Model for Biometric Preserving Gait Synthesis
by: Mitra, Sirshapan, et al.
Published: (2025)
by: Mitra, Sirshapan, et al.
Published: (2025)
STPro: Spatial and Temporal Progressive Learning for Weakly Supervised Spatio-Temporal Grounding
by: Garg, Aaryan, et al.
Published: (2025)
by: Garg, Aaryan, et al.
Published: (2025)
Hier-EgoPack: Hierarchical Egocentric Video Understanding with Diverse Task Perspectives
by: Peirone, Simone Alberto, et al.
Published: (2025)
by: Peirone, Simone Alberto, et al.
Published: (2025)
A Skill-augmented Agentic Framework and Benchmark for Multi-Video Understanding
by: Zhang, Yue, et al.
Published: (2026)
by: Zhang, Yue, et al.
Published: (2026)
SkillFormer: Unified Multi-View Video Understanding for Proficiency Estimation
by: Bianchi, Edoardo, et al.
Published: (2025)
by: Bianchi, Edoardo, et al.
Published: (2025)
Q-ARVD: Quantizing Autoregressive Video Diffusion Models
by: Tang, Siao, et al.
Published: (2026)
by: Tang, Siao, et al.
Published: (2026)
Out of Sight, Not Out of Context? Egocentric Spatial Reasoning in VLMs Across Disjoint Frames
by: Ravi, Sahithya, et al.
Published: (2025)
by: Ravi, Sahithya, et al.
Published: (2025)
Similar Items
-
DisenQ: Disentangling Q-Former for Activity-Biometrics
by: Azad, Shehreen, et al.
Published: (2025) -
StreamReady: Learning What to Answer and When in Long Streaming Videos
by: Azad, Shehreen, et al.
Published: (2026) -
Understanding Depth and Height Perception in Large Visual-Language Models
by: Azad, Shehreen, et al.
Published: (2024) -
Activity-Biometrics: Person Identification from Daily Activities
by: Azad, Shehreen, et al.
Published: (2024) -
Robustness Analysis on Foundational Segmentation Models
by: Schiappa, Madeline Chantry, et al.
Published: (2023)