What and When to Look?: Temporal Span Proposal Network for Video Relation Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Woo, Sangmin, Noh, Junhyug, Kim, Kangil |
|---|---|
| Format: | Preprint |
| Published: |
2021
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Ordering Matters: Rank-Aware Selective Fusion for Blended Emotion Recognition
by: Lee, Junghyun, et al.
Published: (2026)
by: Lee, Junghyun, et al.
Published: (2026)
What Happens When: Learning Temporal Orders of Events in Videos
by: Ahn, Daechul, et al.
Published: (2025)
by: Ahn, Daechul, et al.
Published: (2025)
Beyond Spatial Frequency: Pixel-wise Temporal Frequency-based Deepfake Video Detection
by: Kim, Taehoon, et al.
Published: (2025)
by: Kim, Taehoon, et al.
Published: (2025)
Watch Video, Catch Keyword: Context-aware Keyword Attention for Moment Retrieval and Highlight Detection
by: Um, Sung Jin, et al.
Published: (2025)
by: Um, Sung Jin, et al.
Published: (2025)
Don't Miss the Forest for the Trees: Attentional Vision Calibration for Large Vision Language Models
by: Woo, Sangmin, et al.
Published: (2024)
by: Woo, Sangmin, et al.
Published: (2024)
RITUAL: Random Image Transformations as a Universal Anti-hallucination Lever in Large Vision Language Models
by: Woo, Sangmin, et al.
Published: (2024)
by: Woo, Sangmin, et al.
Published: (2024)
Denoising Task Routing for Diffusion Models
by: Park, Byeongjun, et al.
Published: (2023)
by: Park, Byeongjun, et al.
Published: (2023)
AI-Generated Images: What Humans and Machines See When They Look at the Same Image
by: Poletti, Silvia, et al.
Published: (2026)
by: Poletti, Silvia, et al.
Published: (2026)
Gaussian Mixture Proposals with Pull-Push Learning Scheme to Capture Diverse Events for Weakly Supervised Temporal Video Grounding
by: Kim, Sunoh, et al.
Published: (2023)
by: Kim, Sunoh, et al.
Published: (2023)
Diffusion Model Patching via Mixture-of-Prompts
by: Ham, Seokil, et al.
Published: (2024)
by: Ham, Seokil, et al.
Published: (2024)
Beyond Softmax: Dual-Branch Sigmoid Architecture for Accurate Class Activation Maps
by: Oh, Yoojin, et al.
Published: (2025)
by: Oh, Yoojin, et al.
Published: (2025)
KeyRe-ID: Keypoint-Guided Person Re-Identification using Part-Aware Representation in Videos
by: Kim, Jinseong, et al.
Published: (2025)
by: Kim, Jinseong, et al.
Published: (2025)
AHA -- Predicting What Matters Next: Online Highlight Detection Without Looking Ahead
by: Chang, Aiden, et al.
Published: (2025)
by: Chang, Aiden, et al.
Published: (2025)
MemoryTalker: Personalized Speech-Driven 3D Facial Animation via Audio-Guided Stylization
by: Kim, Hyung Kyu, et al.
Published: (2025)
by: Kim, Hyung Kyu, et al.
Published: (2025)
Multimodal Language Models See Better When They Look Shallower
by: Chen, Haoran, et al.
Published: (2025)
by: Chen, Haoran, et al.
Published: (2025)
LookWise: Knowing When and Where to Look for Fine-Grained Visual Reasoning in Multimodal Large Language Models
by: Shen, Yuxiang, et al.
Published: (2026)
by: Shen, Yuxiang, et al.
Published: (2026)
Exposing and Mitigating Temporal Attack in Deepfake Video Detection
by: Gu, Zheyuan, et al.
Published: (2026)
by: Gu, Zheyuan, et al.
Published: (2026)
Video Diffusion Models Excel at Tracking Similar-Looking Objects Without Supervision
by: Zhang, Chenshuang, et al.
Published: (2025)
by: Zhang, Chenshuang, et al.
Published: (2025)
APT: Improving Diffusion Models for High Resolution Image Generation with Adaptive Path Tracing
by: Han, Sangmin, et al.
Published: (2025)
by: Han, Sangmin, et al.
Published: (2025)
STRIDE: When to Speak Meets Sequence Denoising for Streaming Video Understanding
by: Kim, Junho, et al.
Published: (2026)
by: Kim, Junho, et al.
Published: (2026)
When Looking Is Not Enough: Visual Attention Structure Reveals Hallucination in MLLMs
by: Cao, Fanpu, et al.
Published: (2026)
by: Cao, Fanpu, et al.
Published: (2026)
Efficient Neural Video Representation with Temporally Coherent Modulation
by: Shin, Seungjun, et al.
Published: (2025)
by: Shin, Seungjun, et al.
Published: (2025)
Faster Image2Video Generation: A Closer Look at CLIP Image Embedding's Impact on Spatio-Temporal Cross-Attentions
by: Taghipour, Ashkan, et al.
Published: (2024)
by: Taghipour, Ashkan, et al.
Published: (2024)
Bridging Modalities, Spanning Time: Structured Memory for Ultra-Long Agentic Video Reasoning
by: Li, Jiazheng, et al.
Published: (2026)
by: Li, Jiazheng, et al.
Published: (2026)
FreeTimeGS++: Secrets of Dynamic Gaussian Splatting and Their Principles
by: Lee, Lucas Yunkyu, et al.
Published: (2026)
by: Lee, Lucas Yunkyu, et al.
Published: (2026)
Pinpointing Trigger Moment for Grounded Video QA: Enhancing Spatio-temporal Grounding in Multimodal Large Language Models
by: Seo, Jinhwan, et al.
Published: (2025)
by: Seo, Jinhwan, et al.
Published: (2025)
MambaTAD: When State-Space Models Meet Long-Range Temporal Action Detection
by: Lu, Hui, et al.
Published: (2025)
by: Lu, Hui, et al.
Published: (2025)
Weakly Supervised Video Anomaly Detection and Localization with Spatio-Temporal Prompts
by: Wu, Peng, et al.
Published: (2024)
by: Wu, Peng, et al.
Published: (2024)
Smart Routing for Multimodal Video Retrieval: When to Search What
by: Rosa, Kevin Dela
Published: (2025)
by: Rosa, Kevin Dela
Published: (2025)
Temporal Object-Aware Vision Transformer for Few-Shot Video Object Detection
by: Kumar, Yogesh, et al.
Published: (2025)
by: Kumar, Yogesh, et al.
Published: (2025)
OTT-Vid: Optimal Transport Temporal Token Compression for Video Large Language Models
by: Kang, Minseok, et al.
Published: (2026)
by: Kang, Minseok, et al.
Published: (2026)
Scalp Diagnostic System With Label-Free Segmentation and Training-Free Image Translation
by: Kim, Youngmin, et al.
Published: (2024)
by: Kim, Youngmin, et al.
Published: (2024)
EgoExo-Con: Exploring View-Invariant Video Temporal Understanding
by: Jung, Minjoon, et al.
Published: (2025)
by: Jung, Minjoon, et al.
Published: (2025)
TempGlitch: Evaluating Vision-Language Models for Temporal Glitch Detection in Gameplay Videos
by: Yu, Yakun, et al.
Published: (2026)
by: Yu, Yakun, et al.
Published: (2026)
Human-Centric Video Anomaly Detection Through Spatio-Temporal Pose Tokenization and Transformer
by: Noghre, Ghazal Alinezhad, et al.
Published: (2024)
by: Noghre, Ghazal Alinezhad, et al.
Published: (2024)
SAM-PM: Enhancing Video Camouflaged Object Detection using Spatio-Temporal Attention
by: Meeran, Muhammad Nawfal, et al.
Published: (2024)
by: Meeran, Muhammad Nawfal, et al.
Published: (2024)
Spanning Tree Autoregressive Visual Generation
by: Lee, Sangkyu, et al.
Published: (2025)
by: Lee, Sangkyu, et al.
Published: (2025)
SAIL: Similarity-Aware Guidance and Inter-Caption Augmentation-based Learning for Weakly-Supervised Dense Video Captioning
by: Kim, Ye-Chan, et al.
Published: (2026)
by: Kim, Ye-Chan, et al.
Published: (2026)
BF-STVSR: B-Splines and Fourier-Best Friends for High Fidelity Spatial-Temporal Video Super-Resolution
by: Kim, Eunjin, et al.
Published: (2025)
by: Kim, Eunjin, et al.
Published: (2025)
DETACH : Decomposed Spatio-Temporal Alignment for Exocentric Video and Ambient Sensors with Staged Learning
by: Yoon, Junho, et al.
Published: (2025)
by: Yoon, Junho, et al.
Published: (2025)
Similar Items
-
Ordering Matters: Rank-Aware Selective Fusion for Blended Emotion Recognition
by: Lee, Junghyun, et al.
Published: (2026) -
What Happens When: Learning Temporal Orders of Events in Videos
by: Ahn, Daechul, et al.
Published: (2025) -
Beyond Spatial Frequency: Pixel-wise Temporal Frequency-based Deepfake Video Detection
by: Kim, Taehoon, et al.
Published: (2025) -
Watch Video, Catch Keyword: Context-aware Keyword Attention for Moment Retrieval and Highlight Detection
by: Um, Sung Jin, et al.
Published: (2025) -
Don't Miss the Forest for the Trees: Attentional Vision Calibration for Large Vision Language Models
by: Woo, Sangmin, et al.
Published: (2024)