TIR-Flow: Active Video Search and Reasoning with Frozen VLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jin, Hongbo, Xie, Siyi, Ding, Jiayu, Lin, Kuanwei, Li, Ge |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VideoCuRL: Video Curriculum Reinforcement Learning with Orthogonal Difficulty Decomposition
von: Jin, Hongbo, et al.
Veröffentlicht: (2025)
von: Jin, Hongbo, et al.
Veröffentlicht: (2025)
VISTA: Mitigating Semantic Inertia in Video-LLMs via Training-Free Dynamic Chain-of-Thought Routing
von: Jin, Hongbo, et al.
Veröffentlicht: (2025)
von: Jin, Hongbo, et al.
Veröffentlicht: (2025)
VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding
von: Lin, Kuanwei, et al.
Veröffentlicht: (2026)
von: Lin, Kuanwei, et al.
Veröffentlicht: (2026)
VideoTIR: Accurate Understanding for Long Videos with Efficient Tool-Integrated Reasoning
von: Gao, Zhe, et al.
Veröffentlicht: (2026)
von: Gao, Zhe, et al.
Veröffentlicht: (2026)
SignDATA: Data Pipeline for Sign Language Translation
von: Chen, Kuanwei, et al.
Veröffentlicht: (2026)
von: Chen, Kuanwei, et al.
Veröffentlicht: (2026)
Beyond Shortcuts: Mitigating Visual Illusions in Frozen VLMs via Qualitative Reasoning
von: Guo, Hao, et al.
Veröffentlicht: (2026)
von: Guo, Hao, et al.
Veröffentlicht: (2026)
VISD: Enhancing Video Reasoning via Structured Self-Distillation
von: Lin, Hao, et al.
Veröffentlicht: (2026)
von: Lin, Hao, et al.
Veröffentlicht: (2026)
TIR-Bench: A Comprehensive Benchmark for Agentic Thinking-with-Images Reasoning
von: Li, Ming, et al.
Veröffentlicht: (2025)
von: Li, Ming, et al.
Veröffentlicht: (2025)
Towards Compact Sign Language Translation: Frame Rate and Model Size Trade-offs
von: Chen, Kuanwei, et al.
Veröffentlicht: (2026)
von: Chen, Kuanwei, et al.
Veröffentlicht: (2026)
Similarity-as-Evidence: Calibrating Overconfident VLMs for Interpretable and Label-Efficient Medical Active Learning
von: Xie, Zhuofan, et al.
Veröffentlicht: (2026)
von: Xie, Zhuofan, et al.
Veröffentlicht: (2026)
Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning
von: Ge, Yuyao, et al.
Veröffentlicht: (2025)
von: Ge, Yuyao, et al.
Veröffentlicht: (2025)
Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs
von: Zhang, Xintong, et al.
Veröffentlicht: (2025)
von: Zhang, Xintong, et al.
Veröffentlicht: (2025)
VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization
von: Cheng, Junhao, et al.
Veröffentlicht: (2026)
von: Cheng, Junhao, et al.
Veröffentlicht: (2026)
Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding
von: Xie, Yuan, et al.
Veröffentlicht: (2025)
von: Xie, Yuan, et al.
Veröffentlicht: (2025)
Clapper: Compact Learning and Video Representation in VLMs
von: Kong, Lingyu, et al.
Veröffentlicht: (2025)
von: Kong, Lingyu, et al.
Veröffentlicht: (2025)
Linear Scaling Video VLMs for Long Video Understanding
von: Eyzaguirre, Cristobal, et al.
Veröffentlicht: (2026)
von: Eyzaguirre, Cristobal, et al.
Veröffentlicht: (2026)
SIRI-Bench: Challenging VLMs' Spatial Intelligence through Complex Reasoning Tasks
von: Song, Zijian, et al.
Veröffentlicht: (2025)
von: Song, Zijian, et al.
Veröffentlicht: (2025)
On RGB-TIR Stereo Calibration under Extreme Resolution Asymmetry
von: Król, Michał, et al.
Veröffentlicht: (2026)
von: Król, Michał, et al.
Veröffentlicht: (2026)
TIR-Agent: Training an Explorative and Efficient Agent for Image Restoration
von: Zhang, Yisheng, et al.
Veröffentlicht: (2026)
von: Zhang, Yisheng, et al.
Veröffentlicht: (2026)
DIR-TIR: Dialog-Iterative Refinement for Text-to-Image Retrieval
von: Zhen, Zongwei, et al.
Veröffentlicht: (2025)
von: Zhen, Zongwei, et al.
Veröffentlicht: (2025)
Reasmory: 3D Reconstruction as Explicit Memory for VLMs Spatial Reasoning
von: He, Jixuan, et al.
Veröffentlicht: (2026)
von: He, Jixuan, et al.
Veröffentlicht: (2026)
FrozenSeg: Harmonizing Frozen Foundation Models for Open-Vocabulary Segmentation
von: Chen, Xi, et al.
Veröffentlicht: (2024)
von: Chen, Xi, et al.
Veröffentlicht: (2024)
How Auxiliary Reasoning Unleashes GUI Grounding in VLMs
von: Li, Weiming, et al.
Veröffentlicht: (2025)
von: Li, Weiming, et al.
Veröffentlicht: (2025)
CLGRPO: Reasoning Ability Enhancement for Small VLMs
von: Wang, Fanyi, et al.
Veröffentlicht: (2025)
von: Wang, Fanyi, et al.
Veröffentlicht: (2025)
Reinforcing Video Reasoning with Focused Thinking
von: Dang, Jisheng, et al.
Veröffentlicht: (2025)
von: Dang, Jisheng, et al.
Veröffentlicht: (2025)
VideoASMR-Bench: Can AI-Generated ASMR Videos Fool VLMs and Humans?
von: Wang, Jiaqi, et al.
Veröffentlicht: (2025)
von: Wang, Jiaqi, et al.
Veröffentlicht: (2025)
Frozen-DETR: Enhancing DETR with Image Understanding from Frozen Foundation Models
von: Fu, Shenghao, et al.
Veröffentlicht: (2024)
von: Fu, Shenghao, et al.
Veröffentlicht: (2024)
ConMe: Rethinking Evaluation of Compositional Reasoning for Modern VLMs
von: Huang, Irene, et al.
Veröffentlicht: (2024)
von: Huang, Irene, et al.
Veröffentlicht: (2024)
RASR: Retrieval-Augmented Semantic Reasoning for Fake News Video Detection
von: Li, Hui, et al.
Veröffentlicht: (2026)
von: Li, Hui, et al.
Veröffentlicht: (2026)
LMM-Track4D: Eliciting 4D Dynamic Reasoning in LMMs via Trajectory-Grounded Dialogue
von: Li, Chaoyue, et al.
Veröffentlicht: (2026)
von: Li, Chaoyue, et al.
Veröffentlicht: (2026)
[De|Re]constructing VLMs' Reasoning in Counting
von: Alghisi, Simone, et al.
Veröffentlicht: (2025)
von: Alghisi, Simone, et al.
Veröffentlicht: (2025)
Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?
von: Cheng, Junhao, et al.
Veröffentlicht: (2025)
von: Cheng, Junhao, et al.
Veröffentlicht: (2025)
SpinBench: Perspective and Rotation as a Lens on Spatial Reasoning in VLMs
von: Zhang, Yuyou, et al.
Veröffentlicht: (2025)
von: Zhang, Yuyou, et al.
Veröffentlicht: (2025)
Decomposing Visual Classification: Assessing Tree-Based Reasoning in VLMs
von: Elmansoury, Sary, et al.
Veröffentlicht: (2025)
von: Elmansoury, Sary, et al.
Veröffentlicht: (2025)
VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMs
von: Törtei, Brigitta Malagurski, et al.
Veröffentlicht: (2025)
von: Törtei, Brigitta Malagurski, et al.
Veröffentlicht: (2025)
Drive-R1: Bridging Reasoning and Planning in VLMs for Autonomous Driving with Reinforcement Learning
von: Li, Yue, et al.
Veröffentlicht: (2025)
von: Li, Yue, et al.
Veröffentlicht: (2025)
ALSS-YOLO: An Adaptive Lightweight Channel Split and Shuffling Network for TIR Wildlife Detection in UAV Imagery
von: He, Ang, et al.
Veröffentlicht: (2024)
von: He, Ang, et al.
Veröffentlicht: (2024)
Assessing Situational and Spatial Awareness of VLMs with Synthetically Generated Video
von: Benschop, Pascal, et al.
Veröffentlicht: (2026)
von: Benschop, Pascal, et al.
Veröffentlicht: (2026)
VideoMem: Enhancing Ultra-Long Video Understanding via Adaptive Memory Management
von: Jin, Hongbo, et al.
Veröffentlicht: (2025)
von: Jin, Hongbo, et al.
Veröffentlicht: (2025)
Spotlight: Identifying and Localizing Video Generation Errors Using VLMs
von: Chinchure, Aditya, et al.
Veröffentlicht: (2025)
von: Chinchure, Aditya, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
VideoCuRL: Video Curriculum Reinforcement Learning with Orthogonal Difficulty Decomposition
von: Jin, Hongbo, et al.
Veröffentlicht: (2025) -
VISTA: Mitigating Semantic Inertia in Video-LLMs via Training-Free Dynamic Chain-of-Thought Routing
von: Jin, Hongbo, et al.
Veröffentlicht: (2025) -
VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding
von: Lin, Kuanwei, et al.
Veröffentlicht: (2026) -
VideoTIR: Accurate Understanding for Long Videos with Efficient Tool-Integrated Reasoning
von: Gao, Zhe, et al.
Veröffentlicht: (2026) -
SignDATA: Data Pipeline for Sign Language Translation
von: Chen, Kuanwei, et al.
Veröffentlicht: (2026)