Classification Matters: Improving Video Action Detection with Class-Specific Attention
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lee, Jinsung, Kim, Taeoh, Lee, Inwoong, Shim, Minho, Wee, Dongyoon, Cho, Minsu, Kwak, Suha |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Why Can't I Open My Drawer? Mitigating Object-Driven Shortcuts in Zero-Shot Compositional Action Recognition
von: Ahn, Geo, et al.
Veröffentlicht: (2026)
von: Ahn, Geo, et al.
Veröffentlicht: (2026)
Video-Oasis: Rethinking Evaluation of Video Understanding
von: Lim, Geuntaek, et al.
Veröffentlicht: (2026)
von: Lim, Geuntaek, et al.
Veröffentlicht: (2026)
Prototypes are Balanced Units for Efficient and Effective Partially Relevant Video Retrieval
von: Moon, WonJun, et al.
Veröffentlicht: (2025)
von: Moon, WonJun, et al.
Veröffentlicht: (2025)
Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs
von: Hyun, Jeongseok, et al.
Veröffentlicht: (2025)
von: Hyun, Jeongseok, et al.
Veröffentlicht: (2025)
Planning in 8 Tokens: A Compact Discrete Tokenizer for Latent World Model
von: Kim, Dongwon, et al.
Veröffentlicht: (2026)
von: Kim, Dongwon, et al.
Veröffentlicht: (2026)
Tempered Self-Similarity Alignment for Physically Plausible Video Generation
von: Kim, Manjin, et al.
Veröffentlicht: (2026)
von: Kim, Manjin, et al.
Veröffentlicht: (2026)
Decomposed Attention Fusion in MLLMs for Training-Free Video Reasoning Segmentation
von: Han, Su Ho, et al.
Veröffentlicht: (2025)
von: Han, Su Ho, et al.
Veröffentlicht: (2025)
Part-Aware Bottom-Up Group Reasoning for Fine-Grained Social Interaction Detection
von: Kim, Dongkeun, et al.
Veröffentlicht: (2025)
von: Kim, Dongkeun, et al.
Veröffentlicht: (2025)
Online Temporal Action Localization with Memory-Augmented Transformer
von: Song, Youngkil, et al.
Veröffentlicht: (2024)
von: Song, Youngkil, et al.
Veröffentlicht: (2024)
ActFusion: a Unified Diffusion Model for Action Segmentation and Anticipation
von: Gong, Dayoung, et al.
Veröffentlicht: (2024)
von: Gong, Dayoung, et al.
Veröffentlicht: (2024)
Towards More Practical Group Activity Detection: A New Benchmark and Model
von: Kim, Dongkeun, et al.
Veröffentlicht: (2023)
von: Kim, Dongkeun, et al.
Veröffentlicht: (2023)
CoMoGaussian: Continuous Motion-Aware Gaussian Splatting from Motion-Blurred Images
von: Lee, Jungho, et al.
Veröffentlicht: (2025)
von: Lee, Jungho, et al.
Veröffentlicht: (2025)
CoCoGaussian: Leveraging Circle of Confusion for Gaussian Splatting from Defocused Images
von: Lee, Jungho, et al.
Veröffentlicht: (2024)
von: Lee, Jungho, et al.
Veröffentlicht: (2024)
FreeAction: Training-Free Techniques for Enhanced Fidelity of Trajectory-to-Video Generation
von: Kim, Seungwook, et al.
Veröffentlicht: (2025)
von: Kim, Seungwook, et al.
Veröffentlicht: (2025)
Learning Audio-guided Video Representation with Gated Attention for Video-Text Retrieval
von: Jeong, Boseung, et al.
Veröffentlicht: (2025)
von: Jeong, Boseung, et al.
Veröffentlicht: (2025)
Structured State-Space Regularization for Generation-Friendly Image Tokenization
von: Lee, Jinsung, et al.
Veröffentlicht: (2026)
von: Lee, Jinsung, et al.
Veröffentlicht: (2026)
A Simple Baseline with Single-encoder for Referring Image Segmentation
von: Yu, Seonghoon, et al.
Veröffentlicht: (2024)
von: Yu, Seonghoon, et al.
Veröffentlicht: (2024)
Bootstrapping Top-down Information for Self-modulating Slot Attention
von: Kim, Dongwon, et al.
Veröffentlicht: (2024)
von: Kim, Dongwon, et al.
Veröffentlicht: (2024)
FREST: Feature RESToration for Semantic Segmentation under Multiple Adverse Conditions
von: Lee, Sohyun, et al.
Veröffentlicht: (2024)
von: Lee, Sohyun, et al.
Veröffentlicht: (2024)
Improving Sound Source Localization with Joint Slot Attention on Image and Audio
von: Kim, Inho, et al.
Veröffentlicht: (2025)
von: Kim, Inho, et al.
Veröffentlicht: (2025)
Extreme Point Supervised Instance Segmentation
von: Lee, Hyeonjun, et al.
Veröffentlicht: (2024)
von: Lee, Hyeonjun, et al.
Veröffentlicht: (2024)
PLOT: Text-based Person Search with Part Slot Attention for Corresponding Part Discovery
von: Park, Jicheol, et al.
Veröffentlicht: (2024)
von: Park, Jicheol, et al.
Veröffentlicht: (2024)
SeiT++: Masked Token Modeling Improves Storage-efficient Training
von: Lee, Minhyun, et al.
Veröffentlicht: (2023)
von: Lee, Minhyun, et al.
Veröffentlicht: (2023)
Leveraging Temporal Contextualization for Video Action Recognition
von: Kim, Minji, et al.
Veröffentlicht: (2024)
von: Kim, Minji, et al.
Veröffentlicht: (2024)
Video Summarization with Large Language Models
von: Lee, Min Jung, et al.
Veröffentlicht: (2025)
von: Lee, Min Jung, et al.
Veröffentlicht: (2025)
Robust Promptable Video Object Segmentation
von: Lee, Sohyun, et al.
Veröffentlicht: (2026)
von: Lee, Sohyun, et al.
Veröffentlicht: (2026)
Harnessing the Power of Training-Free Techniques in Text-to-2D Generation for Text-to-3D Generation via Score Distillation Sampling
von: Lee, Junhong, et al.
Veröffentlicht: (2025)
von: Lee, Junhong, et al.
Veröffentlicht: (2025)
GaRA-SAM: Robustifying Segment Anything Model with Gated-Rank Adaptation
von: Lee, Sohyun, et al.
Veröffentlicht: (2025)
von: Lee, Sohyun, et al.
Veröffentlicht: (2025)
Coherent Human-Scene Reconstruction from Multi-Person Multi-View Video in a Single Pass
von: Kim, Sangmin, et al.
Veröffentlicht: (2026)
von: Kim, Sangmin, et al.
Veröffentlicht: (2026)
Improving Text-to-Image Generation with Intrinsic Self-Confidence Rewards
von: Kim, Seungwook, et al.
Veröffentlicht: (2026)
von: Kim, Seungwook, et al.
Veröffentlicht: (2026)
RoDyGS: Robust Dynamic Gaussian Splatting for Casual Videos
von: Lee, Junmyeong, et al.
Veröffentlicht: (2024)
von: Lee, Junmyeong, et al.
Veröffentlicht: (2024)
Improving Text-based Person Search via Part-level Cross-modal Correspondence
von: Park, Jicheol, et al.
Veröffentlicht: (2024)
von: Park, Jicheol, et al.
Veröffentlicht: (2024)
Humans as a Calibration Pattern: Dynamic 3D Scene Reconstruction from Unsynchronized and Uncalibrated Videos
von: Choi, Changwoon, et al.
Veröffentlicht: (2024)
von: Choi, Changwoon, et al.
Veröffentlicht: (2024)
Dual Prototype Attention for Unsupervised Video Object Segmentation
von: Cho, Suhwan, et al.
Veröffentlicht: (2022)
von: Cho, Suhwan, et al.
Veröffentlicht: (2022)
CAST: Cross-Attention in Space and Time for Video Action Recognition
von: Lee, Dongho, et al.
Veröffentlicht: (2023)
von: Lee, Dongho, et al.
Veröffentlicht: (2023)
OFF-CLIP: Improving Normal Detection Confidence in Radiology CLIP with Simple Off-Diagonal Term Auto-Adjustment
von: Park, Junhyun, et al.
Veröffentlicht: (2025)
von: Park, Junhyun, et al.
Veröffentlicht: (2025)
Locality-Aware Zero-Shot Human-Object Interaction Detection
von: Kim, Sanghyun, et al.
Veröffentlicht: (2025)
von: Kim, Sanghyun, et al.
Veröffentlicht: (2025)
ActionSwitch: Class-agnostic Detection of Simultaneous Actions in Streaming Videos
von: Kang, Hyolim, et al.
Veröffentlicht: (2024)
von: Kang, Hyolim, et al.
Veröffentlicht: (2024)
MaskRIS: Semantic Distortion-aware Data Augmentation for Referring Image Segmentation
von: Lee, Minhyun, et al.
Veröffentlicht: (2024)
von: Lee, Minhyun, et al.
Veröffentlicht: (2024)
NVS-Adapter: Plug-and-Play Novel View Synthesis from a Single Image
von: Jeong, Yoonwoo, et al.
Veröffentlicht: (2023)
von: Jeong, Yoonwoo, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Why Can't I Open My Drawer? Mitigating Object-Driven Shortcuts in Zero-Shot Compositional Action Recognition
von: Ahn, Geo, et al.
Veröffentlicht: (2026) -
Video-Oasis: Rethinking Evaluation of Video Understanding
von: Lim, Geuntaek, et al.
Veröffentlicht: (2026) -
Prototypes are Balanced Units for Efficient and Effective Partially Relevant Video Retrieval
von: Moon, WonJun, et al.
Veröffentlicht: (2025) -
Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs
von: Hyun, Jeongseok, et al.
Veröffentlicht: (2025) -
Planning in 8 Tokens: A Compact Discrete Tokenizer for Latent World Model
von: Kim, Dongwon, et al.
Veröffentlicht: (2026)