LensWalk: Agentic Video Understanding by Planning How You See in Videos
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Keliang, Li, Yansong, Shen, Hongze, Liu, Mengdi, Chang, Hong, Shan, Shiguang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HERM: Benchmarking and Enhancing Multimodal LLMs for Human-Centric Understanding
von: Li, Keliang, et al.
Veröffentlicht: (2024)
von: Li, Keliang, et al.
Veröffentlicht: (2024)
Preacher: Paper-to-Video Agentic System
von: Liu, Jingwei, et al.
Veröffentlicht: (2025)
von: Liu, Jingwei, et al.
Veröffentlicht: (2025)
VideoTemp-o3: Harmonizing Temporal Grounding and Video Understanding in Agentic Thinking-with-Videos
von: Liu, Wenqi, et al.
Veröffentlicht: (2026)
von: Liu, Wenqi, et al.
Veröffentlicht: (2026)
Segment Anything for Videos: A Systematic Survey
von: Zhang, Chunhui, et al.
Veröffentlicht: (2024)
von: Zhang, Chunhui, et al.
Veröffentlicht: (2024)
Agentic Video Intelligence: A Flexible Framework for Advanced Video Exploration and Understanding
von: Gao, Hong, et al.
Veröffentlicht: (2025)
von: Gao, Hong, et al.
Veröffentlicht: (2025)
VideoExplorer: Think With Videos For Agentic Long-Video Understanding
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
von: Yuan, Huaying, et al.
Veröffentlicht: (2025)
Deep Video Discovery: Agentic Search with Tool Use for Long-form Video Understanding
von: Zhang, Xiaoyi, et al.
Veröffentlicht: (2025)
von: Zhang, Xiaoyi, et al.
Veröffentlicht: (2025)
UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing
von: Li, Yiheng, et al.
Veröffentlicht: (2024)
von: Li, Yiheng, et al.
Veröffentlicht: (2024)
HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes
von: Li, Keliang, et al.
Veröffentlicht: (2025)
von: Li, Keliang, et al.
Veröffentlicht: (2025)
INFACT: A Diagnostic Benchmark for Induced Faithfulness and Factuality Hallucinations in Video-LLMs
von: Yang, Junqi, et al.
Veröffentlicht: (2026)
von: Yang, Junqi, et al.
Veröffentlicht: (2026)
VideoSEAL: Mitigating Evidence Misalignment in Agentic Long Video Understanding by Decoupling Answer Authority
von: Qiu, Chenhao, et al.
Veröffentlicht: (2026)
von: Qiu, Chenhao, et al.
Veröffentlicht: (2026)
Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding
von: Wang, Ziyang, et al.
Veröffentlicht: (2025)
von: Wang, Ziyang, et al.
Veröffentlicht: (2025)
Visual Agentic Memory: Enabling Online Long Video Understanding via Online Indexing, Hierarchical Memory, and Agentic Retrieval
von: Li, Aiden Yiliu, et al.
Veröffentlicht: (2026)
von: Li, Aiden Yiliu, et al.
Veröffentlicht: (2026)
AnyMo: Scaling Any-Modality Conditional Motion Generation with Masked Modeling
von: Li, Yiheng, et al.
Veröffentlicht: (2026)
von: Li, Yiheng, et al.
Veröffentlicht: (2026)
BIMM: Brain Inspired Masked Modeling for Video Representation Learning
von: Wan, Zhifan, et al.
Veröffentlicht: (2024)
von: Wan, Zhifan, et al.
Veröffentlicht: (2024)
From Static to Dynamic: Adapting Landmark-Aware Image Models for Facial Expression Recognition in Videos
von: Chen, Yin, et al.
Veröffentlicht: (2023)
von: Chen, Yin, et al.
Veröffentlicht: (2023)
VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning
von: Li, Chenglin, et al.
Veröffentlicht: (2026)
von: Li, Chenglin, et al.
Veröffentlicht: (2026)
VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding
von: Wang, Shihao, et al.
Veröffentlicht: (2025)
von: Wang, Shihao, et al.
Veröffentlicht: (2025)
VCA: Video Curious Agent for Long Video Understanding
von: Yang, Zeyuan, et al.
Veröffentlicht: (2024)
von: Yang, Zeyuan, et al.
Veröffentlicht: (2024)
FrameOracle: Learning What to See and How Much to See in Videos
von: Li, Chaoyu, et al.
Veröffentlicht: (2025)
von: Li, Chaoyu, et al.
Veröffentlicht: (2025)
On Denoising Walking Videos for Gait Recognition
von: Jin, Dongyang, et al.
Veröffentlicht: (2025)
von: Jin, Dongyang, et al.
Veröffentlicht: (2025)
Hierarchical Long Video Understanding with Audiovisual Entity Cohesion and Agentic Search
von: Yin, Xinlei, et al.
Veröffentlicht: (2026)
von: Yin, Xinlei, et al.
Veröffentlicht: (2026)
Causality Model for Semantic Understanding on Videos
von: Yicong, Li
Veröffentlicht: (2025)
von: Yicong, Li
Veröffentlicht: (2025)
LEMON: How Well Do MLLMs Perform Temporal Multimodal Understanding on Instructional Videos?
von: Yu, Zhuang, et al.
Veröffentlicht: (2026)
von: Yu, Zhuang, et al.
Veröffentlicht: (2026)
Resolving Evidence Sparsity: Agentic Context Engineering for Long-Document Understanding
von: Liu, Keliang, et al.
Veröffentlicht: (2025)
von: Liu, Keliang, et al.
Veröffentlicht: (2025)
Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos
von: Fei, Jiajun, et al.
Veröffentlicht: (2024)
von: Fei, Jiajun, et al.
Veröffentlicht: (2024)
TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs
von: Zhang, Jun, et al.
Veröffentlicht: (2025)
von: Zhang, Jun, et al.
Veröffentlicht: (2025)
Watching, Reasoning, and Searching: A Video Deep Research Benchmark on Open Web for Agentic Video Reasoning
von: Liu, Chengwen, et al.
Veröffentlicht: (2026)
von: Liu, Chengwen, et al.
Veröffentlicht: (2026)
OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?
von: Li, Yifei, et al.
Veröffentlicht: (2025)
von: Li, Yifei, et al.
Veröffentlicht: (2025)
TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning
von: Zeng, Xiangyu, et al.
Veröffentlicht: (2024)
von: Zeng, Xiangyu, et al.
Veröffentlicht: (2024)
Time Blindness: Why Video-Language Models Can't See What Humans Can?
von: Upadhyay, Ujjwal, et al.
Veröffentlicht: (2025)
von: Upadhyay, Ujjwal, et al.
Veröffentlicht: (2025)
An Efficient Streaming Video Understanding Framework with Agentic Control
von: Liu, Jinming, et al.
Veröffentlicht: (2026)
von: Liu, Jinming, et al.
Veröffentlicht: (2026)
LumiVideo: An Intelligent Agentic System for Video Color Grading
von: Guo, Yuchen, et al.
Veröffentlicht: (2026)
von: Guo, Yuchen, et al.
Veröffentlicht: (2026)
VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding
von: Lin, Kuanwei, et al.
Veröffentlicht: (2026)
von: Lin, Kuanwei, et al.
Veröffentlicht: (2026)
Bridging Modalities, Spanning Time: Structured Memory for Ultra-Long Agentic Video Reasoning
von: Li, Jiazheng, et al.
Veröffentlicht: (2026)
von: Li, Jiazheng, et al.
Veröffentlicht: (2026)
HIS-GPT: Towards 3D Human-In-Scene Multimodal Understanding
von: Zhao, Jiahe, et al.
Veröffentlicht: (2025)
von: Zhao, Jiahe, et al.
Veröffentlicht: (2025)
Uncovering What, Why and How: A Comprehensive Benchmark for Causation Understanding of Video Anomaly
von: Du, Hang, et al.
Veröffentlicht: (2024)
von: Du, Hang, et al.
Veröffentlicht: (2024)
VideoPrism: A Foundational Visual Encoder for Video Understanding
von: Zhao, Long, et al.
Veröffentlicht: (2024)
von: Zhao, Long, et al.
Veröffentlicht: (2024)
Video Panels for Long Video Understanding
von: Doorenbos, Lars, et al.
Veröffentlicht: (2025)
von: Doorenbos, Lars, et al.
Veröffentlicht: (2025)
UVCG: Leveraging Temporal Consistency for Universal Video Protection
von: Li, KaiZhou, et al.
Veröffentlicht: (2024)
von: Li, KaiZhou, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
HERM: Benchmarking and Enhancing Multimodal LLMs for Human-Centric Understanding
von: Li, Keliang, et al.
Veröffentlicht: (2024) -
Preacher: Paper-to-Video Agentic System
von: Liu, Jingwei, et al.
Veröffentlicht: (2025) -
VideoTemp-o3: Harmonizing Temporal Grounding and Video Understanding in Agentic Thinking-with-Videos
von: Liu, Wenqi, et al.
Veröffentlicht: (2026) -
Segment Anything for Videos: A Systematic Survey
von: Zhang, Chunhui, et al.
Veröffentlicht: (2024) -
Agentic Video Intelligence: A Flexible Framework for Advanced Video Exploration and Understanding
von: Gao, Hong, et al.
Veröffentlicht: (2025)