Why Can't I Open My Drawer? Mitigating Object-Driven Shortcuts in Zero-Shot Compositional Action Recognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ahn, Geo, Lee, Inwoong, Kim, Taeoh, Shim, Minho, Wee, Dongyoon, Choi, Jinwoo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Classification Matters: Improving Video Action Detection with Class-Specific Attention
von: Lee, Jinsung, et al.
Veröffentlicht: (2024)
von: Lee, Jinsung, et al.
Veröffentlicht: (2024)
Video-Oasis: Rethinking Evaluation of Video Understanding
von: Lim, Geuntaek, et al.
Veröffentlicht: (2026)
von: Lim, Geuntaek, et al.
Veröffentlicht: (2026)
Prototypes are Balanced Units for Efficient and Effective Partially Relevant Video Retrieval
von: Moon, WonJun, et al.
Veröffentlicht: (2025)
von: Moon, WonJun, et al.
Veröffentlicht: (2025)
Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs
von: Hyun, Jeongseok, et al.
Veröffentlicht: (2025)
von: Hyun, Jeongseok, et al.
Veröffentlicht: (2025)
DEVIAS: Learning Disentangled Video Representations of Action and Scene
von: Bae, Kyungho, et al.
Veröffentlicht: (2023)
von: Bae, Kyungho, et al.
Veröffentlicht: (2023)
SlotVTG: Object-Centric Adapter for Generalizable Video Temporal Grounding
von: Han, Jiwook, et al.
Veröffentlicht: (2026)
von: Han, Jiwook, et al.
Veröffentlicht: (2026)
Fine-Grained Open-Vocabulary Object Recognition via User-Guided Segmentation
von: Ahn, Jinwoo, et al.
Veröffentlicht: (2024)
von: Ahn, Jinwoo, et al.
Veröffentlicht: (2024)
PCEvE: Part Contribution Evaluation Based Model Explanation for Human Figure Drawing Assessment and Beyond
von: Lee, Jongseo, et al.
Veröffentlicht: (2024)
von: Lee, Jongseo, et al.
Veröffentlicht: (2024)
PoseBridge: Bridging the Skeletonization Gap for Zero-Shot Skeleton-Based Action Recognition
von: Lee, Sanghyeon, et al.
Veröffentlicht: (2026)
von: Lee, Sanghyeon, et al.
Veröffentlicht: (2026)
Decomposed Attention Fusion in MLLMs for Training-Free Video Reasoning Segmentation
von: Han, Su Ho, et al.
Veröffentlicht: (2025)
von: Han, Su Ho, et al.
Veröffentlicht: (2025)
EVIDENT: Routing MLLM Adaptation through Entity-Grounded Visual Evidence for Cross-Domain Video Temporal Grounding
von: Ahn, Geo, et al.
Veröffentlicht: (2026)
von: Ahn, Geo, et al.
Veröffentlicht: (2026)
A Simple Baseline with Single-encoder for Referring Image Segmentation
von: Yu, Seonghoon, et al.
Veröffentlicht: (2024)
von: Yu, Seonghoon, et al.
Veröffentlicht: (2024)
Learning Primitive Relations for Compositional Zero-Shot Learning
von: Lee, Insu, et al.
Veröffentlicht: (2025)
von: Lee, Insu, et al.
Veröffentlicht: (2025)
Why Can't I See My Clusters? A Precision-Recall Approach to Dimensionality Reduction Validation
von: van der Hoorn, Diede P. M., et al.
Veröffentlicht: (2025)
von: van der Hoorn, Diede P. M., et al.
Veröffentlicht: (2025)
Motion-Oriented Compositional Neural Radiance Fields for Monocular Dynamic Human Modeling
von: Kim, Jaehyeok, et al.
Veröffentlicht: (2024)
von: Kim, Jaehyeok, et al.
Veröffentlicht: (2024)
I Can't Patch My OT Systems! A Look at CISA's KEVC Workarounds & Mitigations for OT
von: Huff, Philip, et al.
Veröffentlicht: (2025)
von: Huff, Philip, et al.
Veröffentlicht: (2025)
C2C: Component-to-Composition Learning for Zero-Shot Compositional Action Recognition
von: Li, Rongchang, et al.
Veröffentlicht: (2024)
von: Li, Rongchang, et al.
Veröffentlicht: (2024)
Spot-Compose: A Framework for Open-Vocabulary Object Retrieval and Drawer Manipulation in Point Clouds
von: Lemke, Oliver, et al.
Veröffentlicht: (2024)
von: Lemke, Oliver, et al.
Veröffentlicht: (2024)
CAST: Cross-Attention in Space and Time for Video Action Recognition
von: Lee, Dongho, et al.
Veröffentlicht: (2023)
von: Lee, Dongho, et al.
Veröffentlicht: (2023)
OBSER: Object-Based Sub-Environment Recognition for Zero-Shot Environmental Inference
von: Choi, Won-Seok, et al.
Veröffentlicht: (2025)
von: Choi, Won-Seok, et al.
Veröffentlicht: (2025)
Why AI Can't Simulate Extreme Decision-Making
von: Rosehill, Daniel, et al.
Veröffentlicht: (2026)
von: Rosehill, Daniel, et al.
Veröffentlicht: (2026)
Why I Can't Create a Learning Center
von: Miller, Rosalind
Veröffentlicht: (1975)
von: Miller, Rosalind
Veröffentlicht: (1975)
Why Can't I Ever Find Anything in the Library?
von: Radford, Neil, et al.
Veröffentlicht: (1983)
von: Radford, Neil, et al.
Veröffentlicht: (1983)
Zero-Shot Action Recognition in Surveillance Videos
von: Pereira, Joao, et al.
Veröffentlicht: (2024)
von: Pereira, Joao, et al.
Veröffentlicht: (2024)
CoMoGaussian: Continuous Motion-Aware Gaussian Splatting from Motion-Blurred Images
von: Lee, Jungho, et al.
Veröffentlicht: (2025)
von: Lee, Jungho, et al.
Veröffentlicht: (2025)
Unlocking Transfer Learning for Open-World Few-Shot Recognition
von: Kim, Byeonggeun, et al.
Veröffentlicht: (2024)
von: Kim, Byeonggeun, et al.
Veröffentlicht: (2024)
"We are currently clean on OPSEC": Why JD Can't Encrypt
von: Chiodo, Maurice, et al.
Veröffentlicht: (2026)
von: Chiodo, Maurice, et al.
Veröffentlicht: (2026)
Why the Center Can't Hold: A Diagnosis of Puritanized America
von: O’Neill, Tom
Veröffentlicht: (2019)
von: O’Neill, Tom
Veröffentlicht: (2019)
Why I Can't Read Wallace Stegner, and Other Essays
von: Cook-Lynn, Elizabeth
Veröffentlicht: (2025)
von: Cook-Lynn, Elizabeth
Veröffentlicht: (2025)
Sampling Bag of Views for Open-Vocabulary Object Detection
von: Choi, Hojun, et al.
Veröffentlicht: (2024)
von: Choi, Hojun, et al.
Veröffentlicht: (2024)
Son of Why Johnny Can't Read and What You Do About It, by Hugo Flesch, Son of Rudolf Flesch, Author of Son of Why Johnny Can't Read and...
von: Flesch, Hugo
Veröffentlicht: (1970)
von: Flesch, Hugo
Veröffentlicht: (1970)
Leveraging Temporal Contextualization for Video Action Recognition
von: Kim, Minji, et al.
Veröffentlicht: (2024)
von: Kim, Minji, et al.
Veröffentlicht: (2024)
CoCoGaussian: Leveraging Circle of Confusion for Gaussian Splatting from Defocused Images
von: Lee, Jungho, et al.
Veröffentlicht: (2024)
von: Lee, Jungho, et al.
Veröffentlicht: (2024)
Continual Learning Improves Zero-Shot Action Recognition
von: Gowda, Shreyank N, et al.
Veröffentlicht: (2024)
von: Gowda, Shreyank N, et al.
Veröffentlicht: (2024)
Novel Semantic Prompting for Zero-Shot Action Recognition
von: Iqbal, Salman, et al.
Veröffentlicht: (2026)
von: Iqbal, Salman, et al.
Veröffentlicht: (2026)
Time Blindness: Why Video-Language Models Can't See What Humans Can?
von: Upadhyay, Ujjwal, et al.
Veröffentlicht: (2025)
von: Upadhyay, Ujjwal, et al.
Veröffentlicht: (2025)
Can’t Touch This
Veröffentlicht: (2024)
Veröffentlicht: (2024)
Telling Stories for Common Sense Zero-Shot Action Recognition
von: Gowda, Shreyank N, et al.
Veröffentlicht: (2023)
von: Gowda, Shreyank N, et al.
Veröffentlicht: (2023)
I Can't Join, But I Will Send My Agent: Stand-in Enhanced Asynchronous Meetings (SEAM)
von: Bai, Zhongyi, et al.
Veröffentlicht: (2025)
von: Bai, Zhongyi, et al.
Veröffentlicht: (2025)
ZeSTA: Zero-Shot TTS Augmentation with Domain-Conditioned Training for Data-Efficient Personalized Speech Synthesis
von: Choi, Youngwon, et al.
Veröffentlicht: (2026)
von: Choi, Youngwon, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Classification Matters: Improving Video Action Detection with Class-Specific Attention
von: Lee, Jinsung, et al.
Veröffentlicht: (2024) -
Video-Oasis: Rethinking Evaluation of Video Understanding
von: Lim, Geuntaek, et al.
Veröffentlicht: (2026) -
Prototypes are Balanced Units for Efficient and Effective Partially Relevant Video Retrieval
von: Moon, WonJun, et al.
Veröffentlicht: (2025) -
Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs
von: Hyun, Jeongseok, et al.
Veröffentlicht: (2025) -
DEVIAS: Learning Disentangled Video Representations of Action and Scene
von: Bae, Kyungho, et al.
Veröffentlicht: (2023)