Looking Beyond the Obvious: A Survey on Abstract Concept Recognition for Video Understanding
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Mago, Gowreesh, Mettes, Pascal, Rudinac, Stevan |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
ArtRAG: Retrieval-Augmented Generation with Structured Context for Visual Art Understanding
par: Wang, Shuai, et autres
Publié: (2025)
par: Wang, Shuai, et autres
Publié: (2025)
A Survey on Backbones for Deep Video Action Recognition
par: Tang, Zixuan, et autres
Publié: (2024)
par: Tang, Zixuan, et autres
Publié: (2024)
Hyperbolic Safety-Aware Vision-Language Models
par: Poppi, Tobia, et autres
Publié: (2025)
par: Poppi, Tobia, et autres
Publié: (2025)
A Survey of Video Datasets for Grounded Event Understanding
par: Sanders, Kate, et autres
Publié: (2024)
par: Sanders, Kate, et autres
Publié: (2024)
Look Beyond: Two-Stage Scene View Generation via Panorama and Video Diffusion
par: Kang, Xueyang, et autres
Publié: (2025)
par: Kang, Xueyang, et autres
Publié: (2025)
Prompt Optimizer of Text-to-Image Diffusion Models for Abstract Concept Understanding
par: Fan, Zezhong, et autres
Publié: (2024)
par: Fan, Zezhong, et autres
Publié: (2024)
Looking into Concept Explanation Methods for Diabetic Retinopathy Classification
par: Storås, Andrea M., et autres
Publié: (2024)
par: Storås, Andrea M., et autres
Publié: (2024)
Abstracted Gaussian Prototypes for True One-Shot Concept Learning
par: Zou, Chelsea, et autres
Publié: (2024)
par: Zou, Chelsea, et autres
Publié: (2024)
SV3.3B: A Sports Video Understanding Model for Action Recognition
par: Kodathala, Sai Varun, et autres
Publié: (2025)
par: Kodathala, Sai Varun, et autres
Publié: (2025)
AdaFocus: Adaptive Relevance-Diversity Sampling with Zero-Cache Look-back for Efficient Long Video Understanding
par: Yang, Xiao, et autres
Publié: (2026)
par: Yang, Xiao, et autres
Publié: (2026)
MedObvious: Exposing the Medical Moravec's Paradox in VLMs via Clinical Triage
par: Khan, Ufaq, et autres
Publié: (2026)
par: Khan, Ufaq, et autres
Publié: (2026)
VITATECS: A Diagnostic Dataset for Temporal Concept Understanding of Video-Language Models
par: Li, Shicheng, et autres
Publié: (2023)
par: Li, Shicheng, et autres
Publié: (2023)
Hyperbolic Concept Bottleneck Models
par: Uyterlinde, Daniel, et autres
Publié: (2026)
par: Uyterlinde, Daniel, et autres
Publié: (2026)
A Survey on Generative AI and LLM for Video Generation, Understanding, and Streaming
par: Zhou, Pengyuan, et autres
Publié: (2024)
par: Zhou, Pengyuan, et autres
Publié: (2024)
Handwritten Text Recognition: A Survey
par: Garrido-Munoz, Carlos, et autres
Publié: (2025)
par: Garrido-Munoz, Carlos, et autres
Publié: (2025)
A Case Study on Concept Induction for Neuron-Level Interpretability in CNN
par: Sarma, Moumita Sen, et autres
Publié: (2026)
par: Sarma, Moumita Sen, et autres
Publié: (2026)
Beyond Perception: Evaluating Abstract Visual Reasoning through Multi-Stage Task
par: Jiang, Yanbei, et autres
Publié: (2025)
par: Jiang, Yanbei, et autres
Publié: (2025)
Beyond Task-Specific Reasoning: A Unified Conditional Generative Framework for Abstract Visual Reasoning
par: Shi, Fan, et autres
Publié: (2025)
par: Shi, Fan, et autres
Publié: (2025)
Video Panels for Long Video Understanding
par: Doorenbos, Lars, et autres
Publié: (2025)
par: Doorenbos, Lars, et autres
Publié: (2025)
A Survey of Camouflaged Object Detection and Beyond
par: Xiao, Fengyang, et autres
Publié: (2024)
par: Xiao, Fengyang, et autres
Publié: (2024)
Towards Human-Understandable Multi-Dimensional Concept Discovery
par: Grobrügge, Arne, et autres
Publié: (2025)
par: Grobrügge, Arne, et autres
Publié: (2025)
The CASTLE 2024 Dataset: Advancing the Art of Multimodal Understanding
par: Rossetto, Luca, et autres
Publié: (2025)
par: Rossetto, Luca, et autres
Publié: (2025)
Understanding the Limitations of Diffusion Concept Algebra Through Food
par: Zeng, E. Zhixuan, et autres
Publié: (2024)
par: Zeng, E. Zhixuan, et autres
Publié: (2024)
Continual Hyperbolic Learning of Instances and Classes
par: Ayoughi, Melika, et autres
Publié: (2025)
par: Ayoughi, Melika, et autres
Publié: (2025)
A Comprehensive Survey on Concept Erasure in Text-to-Image Diffusion Models
par: Kim, Changhoon, et autres
Publié: (2025)
par: Kim, Changhoon, et autres
Publié: (2025)
Understanding Video Transformers via Universal Concept Discovery
par: Kowal, Matthew, et autres
Publié: (2024)
par: Kowal, Matthew, et autres
Publié: (2024)
Deep Learning in Palmprint Recognition-A Comprehensive Survey
par: Gao, Chengrui, et autres
Publié: (2025)
par: Gao, Chengrui, et autres
Publié: (2025)
VideoPrism: A Foundational Visual Encoder for Video Understanding
par: Zhao, Long, et autres
Publié: (2024)
par: Zhao, Long, et autres
Publié: (2024)
Why We Look Where We Look: Emergent Human-like Fixations of a Foveated Visual Language Model Maximizing Scene Understanding
par: Murlidaran, Shravan, et autres
Publié: (2026)
par: Murlidaran, Shravan, et autres
Publié: (2026)
MoCA-Video: Motion-Aware Concept Alignment for Consistent Video Editing
par: Zhang, Tong, et autres
Publié: (2025)
par: Zhang, Tong, et autres
Publié: (2025)
Personalized Video Summarization by Multimodal Video Understanding
par: Chen, Brian, et autres
Publié: (2024)
par: Chen, Brian, et autres
Publié: (2024)
Video Diffusion Models Excel at Tracking Similar-Looking Objects Without Supervision
par: Zhang, Chenshuang, et autres
Publié: (2025)
par: Zhang, Chenshuang, et autres
Publié: (2025)
What and When to Look?: Temporal Span Proposal Network for Video Relation Detection
par: Woo, Sangmin, et autres
Publié: (2021)
par: Woo, Sangmin, et autres
Publié: (2021)
On Denoising Walking Videos for Gait Recognition
par: Jin, Dongyang, et autres
Publié: (2025)
par: Jin, Dongyang, et autres
Publié: (2025)
Exploring Explainability in Video Action Recognition
par: Saha, Avinab, et autres
Publié: (2024)
par: Saha, Avinab, et autres
Publié: (2024)
VideoWeave: A Data-Centric Approach for Efficient Video Understanding
par: Durante, Zane, et autres
Publié: (2026)
par: Durante, Zane, et autres
Publié: (2026)
Compositional Entailment Learning for Hyperbolic Vision-Language Models
par: Pal, Avik, et autres
Publié: (2024)
par: Pal, Avik, et autres
Publié: (2024)
VideoLLM Benchmarks and Evaluation: A Survey
par: Kumar, Yogesh
Publié: (2025)
par: Kumar, Yogesh
Publié: (2025)
A Survey: Spatiotemporal Consistency in Video Generation
par: Yin, Zhiyu, et autres
Publié: (2025)
par: Yin, Zhiyu, et autres
Publié: (2025)
Segment Anything for Videos: A Systematic Survey
par: Zhang, Chunhui, et autres
Publié: (2024)
par: Zhang, Chunhui, et autres
Publié: (2024)
Documents similaires
-
ArtRAG: Retrieval-Augmented Generation with Structured Context for Visual Art Understanding
par: Wang, Shuai, et autres
Publié: (2025) -
A Survey on Backbones for Deep Video Action Recognition
par: Tang, Zixuan, et autres
Publié: (2024) -
Hyperbolic Safety-Aware Vision-Language Models
par: Poppi, Tobia, et autres
Publié: (2025) -
A Survey of Video Datasets for Grounded Event Understanding
par: Sanders, Kate, et autres
Publié: (2024) -
Look Beyond: Two-Stage Scene View Generation via Panorama and Video Diffusion
par: Kang, Xueyang, et autres
Publié: (2025)