Measure Twice, Cut Once: A Semantic-Oriented Approach to Video Temporal Localization with Video LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Pang, Zongshang, Otani, Mayu, Nakashima, Yuta |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Would Deep Generative Models Amplify Bias in Future Models?
by: Chen, Tianwei, et al.
Published: (2024)
by: Chen, Tianwei, et al.
Published: (2024)
VideoAuto-R1: Video Auto Reasoning via Thinking Once, Answering Twice
by: Liu, Shuming, et al.
Published: (2026)
by: Liu, Shuming, et al.
Published: (2026)
LTSim: Layout Transportation-based Similarity Measure for Evaluating Layout Generation
by: Otani, Mayu, et al.
Published: (2024)
by: Otani, Mayu, et al.
Published: (2024)
TemporalVLM: Video LLMs for Temporal Reasoning in Long Videos
by: Fateh, Fawad Javed, et al.
Published: (2024)
by: Fateh, Fawad Javed, et al.
Published: (2024)
EMMA: Concept Erasure Benchmark with Comprehensive Semantic Metrics and Diverse Categories
by: Wei, Lu, et al.
Published: (2025)
by: Wei, Lu, et al.
Published: (2025)
Scanning Only Once: An End-to-end Framework for Fast Temporal Grounding in Long Videos
by: Pan, Yulin, et al.
Published: (2023)
by: Pan, Yulin, et al.
Published: (2023)
Learning Local and Global Temporal Contexts for Video Semantic Segmentation
by: Sun, Guolei, et al.
Published: (2022)
by: Sun, Guolei, et al.
Published: (2022)
FlowCut: Unsupervised Video Instance Segmentation via Temporal Mask Matching
by: Sari, Alp Eren, et al.
Published: (2025)
by: Sari, Alp Eren, et al.
Published: (2025)
V-STaR: Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning
by: Cheng, Zixu, et al.
Published: (2025)
by: Cheng, Zixu, et al.
Published: (2025)
From Global to Local: Social Bias Transfer in CLIP
by: Ramos, Ryan, et al.
Published: (2025)
by: Ramos, Ryan, et al.
Published: (2025)
Lost in Time: A New Temporal Benchmark for VideoLLMs
by: Cores, Daniel, et al.
Published: (2024)
by: Cores, Daniel, et al.
Published: (2024)
VTG-LLM: Integrating Timestamp Knowledge into Video LLMs for Enhanced Video Temporal Grounding
by: Guo, Yongxin, et al.
Published: (2024)
by: Guo, Yongxin, et al.
Published: (2024)
Harnessing the Latent Diffusion Model for Training-Free Image Style Transfer
by: Masui, Kento, et al.
Published: (2024)
by: Masui, Kento, et al.
Published: (2024)
Don't Look Twice: Faster Video Transformers with Run-Length Tokenization
by: Choudhury, Rohan, et al.
Published: (2024)
by: Choudhury, Rohan, et al.
Published: (2024)
Temporal-Oriented Recipe for Transferring Large Vision-Language Model to Video Understanding
by: Nguyen, Thong, et al.
Published: (2025)
by: Nguyen, Thong, et al.
Published: (2025)
Unleashing Vision-Language Semantics for Deepfake Video Detection
by: Zhu, Jiawen, et al.
Published: (2026)
by: Zhu, Jiawen, et al.
Published: (2026)
Weakly Supervised Video Anomaly Detection and Localization with Spatio-Temporal Prompts
by: Wu, Peng, et al.
Published: (2024)
by: Wu, Peng, et al.
Published: (2024)
Think Twice, Act Once: Token-Aware Compression and Action Reuse for Efficient Inference in Vision-Language-Action Models
by: Tan, Xudong, et al.
Published: (2025)
by: Tan, Xudong, et al.
Published: (2025)
LSA: Localized Semantic Alignment for Enhancing Temporal Consistency in Traffic Video Generation
by: Karimov, Mirlan, et al.
Published: (2026)
by: Karimov, Mirlan, et al.
Published: (2026)
Semantic Parsing of Colonoscopy Videos with Multi-Label Temporal Networks
by: Kelner, Ori, et al.
Published: (2023)
by: Kelner, Ori, et al.
Published: (2023)
Propose, Assess, Search: Harnessing LLMs for Goal-Oriented Planning in Instructional Videos
by: Islam, Md Mohaiminul, et al.
Published: (2024)
by: Islam, Md Mohaiminul, et al.
Published: (2024)
Think Twice, Click Once: Enhancing GUI Grounding via Fast and Slow Systems
by: Tang, Fei, et al.
Published: (2025)
by: Tang, Fei, et al.
Published: (2025)
Measure Twice, Click Once: Co-evolving Proposer and Visual Critic via Reinforcement Learning for GUI Grounding
by: Wang, Wenkai, et al.
Published: (2026)
by: Wang, Wenkai, et al.
Published: (2026)
IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video Captioning
by: Qiu, Tianheng, et al.
Published: (2025)
by: Qiu, Tianheng, et al.
Published: (2025)
VideoINSTA: Zero-shot Long Video Understanding via Informative Spatial-Temporal Reasoning with LLMs
by: Liao, Ruotong, et al.
Published: (2024)
by: Liao, Ruotong, et al.
Published: (2024)
Enhancing Temporal Understanding in Video-LLMs through Stacked Temporal Attention in Vision Encoders
by: Rasekh, Ali, et al.
Published: (2025)
by: Rasekh, Ali, et al.
Published: (2025)
Video Self-Stitching Graph Network for Temporal Action Localization
by: Zhao, Chen, et al.
Published: (2020)
by: Zhao, Chen, et al.
Published: (2020)
Semi-Supervised Pipe Video Temporal Defect Interval Localization
by: Huang, Zhu, et al.
Published: (2024)
by: Huang, Zhu, et al.
Published: (2024)
Single-step Diffusion-based Video Coding with Semantic-Temporal Guidance
by: Xue, Naifu, et al.
Published: (2025)
by: Xue, Naifu, et al.
Published: (2025)
Enrich and Detect: Video Temporal Grounding with Multimodal LLMs
by: Pramanick, Shraman, et al.
Published: (2025)
by: Pramanick, Shraman, et al.
Published: (2025)
Spatio-Temporal Attention for Consistent Video Semantic Segmentation in Automated Driving
by: Varghese, Serin, et al.
Published: (2026)
by: Varghese, Serin, et al.
Published: (2026)
Text-Video Retrieval with Global-Local Semantic Consistent Learning
by: Zhang, Haonan, et al.
Published: (2024)
by: Zhang, Haonan, et al.
Published: (2024)
Dynamic-Aware Video Distillation: Optimizing Temporal Resolution Based on Video Semantics
by: Zhao, Yinjie, et al.
Published: (2025)
by: Zhao, Yinjie, et al.
Published: (2025)
Mitigating Hallucination in VideoLLMs via Temporal-Aware Activation Engineering
by: Cai, Jianfeng, et al.
Published: (2025)
by: Cai, Jianfeng, et al.
Published: (2025)
Harnessing Synthetic Preference Data for Enhancing Temporal Understanding of Video-LLMs
by: Vani, Sameep, et al.
Published: (2025)
by: Vani, Sameep, et al.
Published: (2025)
EchoPrune: Interpreting Redundancy as Temporal Echoes for Efficient VideoLLMs
by: Li, Jiameng, et al.
Published: (2026)
by: Li, Jiameng, et al.
Published: (2026)
When and How to Cut Classical Concerts? A Multimodal Automated Video Editing Approach
by: Gonzálbez-Biosca, Daniel, et al.
Published: (2025)
by: Gonzálbez-Biosca, Daniel, et al.
Published: (2025)
SmartSight: Mitigating Hallucination in Video-LLMs Without Compromising Video Understanding via Temporal Attention Collapse
by: Sun, Yiming, et al.
Published: (2025)
by: Sun, Yiming, et al.
Published: (2025)
VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
by: Cheng, Zesen, et al.
Published: (2024)
by: Cheng, Zesen, et al.
Published: (2024)
MR. Video: "MapReduce" is the Principle for Long Video Understanding
by: Pang, Ziqi, et al.
Published: (2025)
by: Pang, Ziqi, et al.
Published: (2025)
Similar Items
-
Would Deep Generative Models Amplify Bias in Future Models?
by: Chen, Tianwei, et al.
Published: (2024) -
VideoAuto-R1: Video Auto Reasoning via Thinking Once, Answering Twice
by: Liu, Shuming, et al.
Published: (2026) -
LTSim: Layout Transportation-based Similarity Measure for Evaluating Layout Generation
by: Otani, Mayu, et al.
Published: (2024) -
TemporalVLM: Video LLMs for Temporal Reasoning in Long Videos
by: Fateh, Fawad Javed, et al.
Published: (2024) -
EMMA: Concept Erasure Benchmark with Comprehensive Semantic Metrics and Diverse Categories
by: Wei, Lu, et al.
Published: (2025)