ViKey: Enhancing Temporal Understanding in Videos via Visual Prompting
Fuente:
arXiv
Saved in:
| Main Authors: | Lee, Yeonkyung, Ju, Dayun, Kim, Youngmin, Kang, Seil, Hwang, Seong Jae |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FALCON: Frequency Adjoint Link with CONtinuous Density Mask for Fast Single Image Dehazing
by: Kim, Donghyun, et al.
Published: (2024)
by: Kim, Donghyun, et al.
Published: (2024)
EAGLE: Eigen Aggregation Learning for Object-Centric Unsupervised Semantic Segmentation
by: Kim, Chanyoung, et al.
Published: (2024)
by: Kim, Chanyoung, et al.
Published: (2024)
See What You Are Told: Visual Attention Sink in Large Multimodal Models
by: Kang, Seil, et al.
Published: (2025)
by: Kang, Seil, et al.
Published: (2025)
Your Large Vision-Language Model Only Needs A Few Attention Heads For Visual Grounding
by: Kang, Seil, et al.
Published: (2025)
by: Kang, Seil, et al.
Published: (2025)
Interpretable Motion-Attentive Maps: Spatio-Temporally Localizing Concepts in Video Diffusion Transformers
by: Jun, Youngjun, et al.
Published: (2026)
by: Jun, Youngjun, et al.
Published: (2026)
Advancing Text-Driven Chest X-Ray Generation with Policy-Based Reinforcement Learning
by: Han, Woojung, et al.
Published: (2024)
by: Han, Woojung, et al.
Published: (2024)
Spatial Transport Optimization by Repositioning Attention Map for Training-Free Text-to-Image Synthesis
by: Han, Woojung, et al.
Published: (2025)
by: Han, Woojung, et al.
Published: (2025)
Distilling Spectral Graph for Object-Context Aware Open-Vocabulary Semantic Segmentation
by: Kim, Chanyoung, et al.
Published: (2024)
by: Kim, Chanyoung, et al.
Published: (2024)
Real-Time Visual Attribution Streaming in Thinking Model
by: Kang, Seil, et al.
Published: (2026)
by: Kang, Seil, et al.
Published: (2026)
CoBra: Complementary Branch Fusing Class and Semantic Knowledge for Robust Weakly Supervised Semantic Segmentation
by: Han, Woojung, et al.
Published: (2024)
by: Han, Woojung, et al.
Published: (2024)
Pathology-Aware Adaptive Watermarking for Text-Driven Medical Image Synthesis
by: Kim, Chanyoung, et al.
Published: (2025)
by: Kim, Chanyoung, et al.
Published: (2025)
PRETI: Patient-Aware Retinal Foundation Model via Metadata-Guided Representation Learning
by: Lee, Yeonkyung, et al.
Published: (2025)
by: Lee, Yeonkyung, et al.
Published: (2025)
Interpreting Attention Heads for Image-to-Text Information Flow in Large Vision-Language Models
by: Kim, Jinyeong, et al.
Published: (2025)
by: Kim, Jinyeong, et al.
Published: (2025)
Towards Continuous Sign Language Conversation from Isolated Signs
by: Kim, Youngmin, et al.
Published: (2026)
by: Kim, Youngmin, et al.
Published: (2026)
ECLIPSE: Efficient Continual Learning in Panoptic Segmentation with Visual Prompt Tuning
by: Kim, Beomyoung, et al.
Published: (2024)
by: Kim, Beomyoung, et al.
Published: (2024)
ViRAC: A Vision-Reasoning Agent Head Movement Control Framework in Arbitrary Virtual Environments
by: Hwang, Juyeong, et al.
Published: (2025)
by: Hwang, Juyeong, et al.
Published: (2025)
Delaunay Canopy: Building Wireframe Reconstruction from Airborne LiDAR Point Clouds via Delaunay Graph
by: Kim, Donghyun, et al.
Published: (2026)
by: Kim, Donghyun, et al.
Published: (2026)
Enhancing Multi-Image Understanding through Delimiter Token Scaling
by: Lee, Minyoung, et al.
Published: (2026)
by: Lee, Minyoung, et al.
Published: (2026)
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding
by: Bao, Xiaoyi, et al.
Published: (2025)
by: Bao, Xiaoyi, et al.
Published: (2025)
STELLA: Continual Audio-Video Pre-training with Spatio-Temporal Localized Alignment
by: Lee, Jaewoo, et al.
Published: (2023)
by: Lee, Jaewoo, et al.
Published: (2023)
VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding
by: Kim, Kangsan, et al.
Published: (2024)
by: Kim, Kangsan, et al.
Published: (2024)
ViP-LLaVA: Making Large Multimodal Models Understand Arbitrary Visual Prompts
by: Cai, Mu, et al.
Published: (2023)
by: Cai, Mu, et al.
Published: (2023)
SurgViVQA: Temporally-Grounded Video Question Answering for Surgical Scene Understanding
by: Drago, Mauro Orazio, et al.
Published: (2025)
by: Drago, Mauro Orazio, et al.
Published: (2025)
Enhancing Spatio-Temporal Zero-shot Action Recognition with Language-driven Description Attributes
by: Kim, Yehna, et al.
Published: (2025)
by: Kim, Yehna, et al.
Published: (2025)
PLATYPUS: Progressive Local Surface Estimator for Arbitrary-Scale Point Cloud Upsampling
by: Kim, Donghyun, et al.
Published: (2024)
by: Kim, Donghyun, et al.
Published: (2024)
State Space Prompting via Gathering and Spreading Spatio-Temporal Information for Video Understanding
by: Zhou, Jiahuan, et al.
Published: (2025)
by: Zhou, Jiahuan, et al.
Published: (2025)
SeViCES: Unifying Semantic-Visual Evidence Consensus for Long Video Understanding
by: Sheng, Yuan, et al.
Published: (2025)
by: Sheng, Yuan, et al.
Published: (2025)
STOP: Integrated Spatial-Temporal Dynamic Prompting for Video Understanding
by: Liu, Zichen, et al.
Published: (2025)
by: Liu, Zichen, et al.
Published: (2025)
A 2-Stage Model for Vehicle Class and Orientation Detection with Photo-Realistic Image Generation
by: Kim, Youngmin, et al.
Published: (2025)
by: Kim, Youngmin, et al.
Published: (2025)
SlumpGuard: An AI-Powered Real-Time System for Automated Concrete Slump Prediction via Video Analysis
by: Kim, Youngmin, et al.
Published: (2025)
by: Kim, Youngmin, et al.
Published: (2025)
ViPCap: Retrieval Text-Based Visual Prompts for Lightweight Image Captioning
by: Kim, Taewhan, et al.
Published: (2024)
by: Kim, Taewhan, et al.
Published: (2024)
MonoWAD: Weather-Adaptive Diffusion Model for Robust Monocular 3D Object Detection
by: Oh, Youngmin, et al.
Published: (2024)
by: Oh, Youngmin, et al.
Published: (2024)
Temporal Alignment-Free Video Matching for Few-shot Action Recognition
by: Lee, SuBeen, et al.
Published: (2025)
by: Lee, SuBeen, et al.
Published: (2025)
ViMU: Benchmarking Video Metaphorical Understanding
by: Li, Qi, et al.
Published: (2026)
by: Li, Qi, et al.
Published: (2026)
UCMNet: Uncertainty-Aware Context Memory Network for Under-Display Camera Image Restoration
by: Kim, Daehyun, et al.
Published: (2026)
by: Kim, Daehyun, et al.
Published: (2026)
Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding
by: Jin, Peng, et al.
Published: (2023)
by: Jin, Peng, et al.
Published: (2023)
Mitigating Mask Prior Drift and Positional Attention Collapse in Large Diffusion Vision-Language Models
by: Hong, Sujung, et al.
Published: (2026)
by: Hong, Sujung, et al.
Published: (2026)
V-CORE: Temporally Consistent Video Understanding for Video-LLM
by: Kang, Zhengjian, et al.
Published: (2026)
by: Kang, Zhengjian, et al.
Published: (2026)
ViSpeak: Visual Instruction Feedback in Streaming Videos
by: Fu, Shenghao, et al.
Published: (2025)
by: Fu, Shenghao, et al.
Published: (2025)
Multi-Context Temporal Consistent Modeling for Referring Video Object Segmentation
by: Choi, Sun-Hyuk, et al.
Published: (2025)
by: Choi, Sun-Hyuk, et al.
Published: (2025)
Similar Items
-
FALCON: Frequency Adjoint Link with CONtinuous Density Mask for Fast Single Image Dehazing
by: Kim, Donghyun, et al.
Published: (2024) -
EAGLE: Eigen Aggregation Learning for Object-Centric Unsupervised Semantic Segmentation
by: Kim, Chanyoung, et al.
Published: (2024) -
See What You Are Told: Visual Attention Sink in Large Multimodal Models
by: Kang, Seil, et al.
Published: (2025) -
Your Large Vision-Language Model Only Needs A Few Attention Heads For Visual Grounding
by: Kang, Seil, et al.
Published: (2025) -
Interpretable Motion-Attentive Maps: Spatio-Temporally Localizing Concepts in Video Diffusion Transformers
by: Jun, Youngjun, et al.
Published: (2026)