Thinking in Frames: How Visual Context and Test-Time Scaling Empower Video Reasoning
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Chengzu, Wang, Zanyi, Li, Jiaang, Xu, Yi, Zhou, Han, Zhang, Huanyu, An, Ruichuan, Jiang, Dengyang, An, Zhaochong, Vulić, Ivan, Belongie, Serge, Korhonen, Anna |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Visual Planning: Let's Think Only with Images
di: Xu, Yi, et al.
Pubblicazione: (2025)
di: Xu, Yi, et al.
Pubblicazione: (2025)
Large Language Models are Miscalibrated In-Context Learners
di: Li, Chengzu, et al.
Pubblicazione: (2023)
di: Li, Chengzu, et al.
Pubblicazione: (2023)
TopViewRS: Vision-Language Models as Top-View Spatial Reasoners
di: Li, Chengzu, et al.
Pubblicazione: (2024)
di: Li, Chengzu, et al.
Pubblicazione: (2024)
Video Understanding: From Geometry and Semantics to Unified Models
di: An, Zhaochong, et al.
Pubblicazione: (2026)
di: An, Zhaochong, et al.
Pubblicazione: (2026)
Imagine while Reasoning in Space: Multimodal Visualization-of-Thought
di: Li, Chengzu, et al.
Pubblicazione: (2025)
di: Li, Chengzu, et al.
Pubblicazione: (2025)
Self-Augmented In-Context Learning for Unsupervised Word Translation
di: Li, Yaoyiran, et al.
Pubblicazione: (2024)
di: Li, Yaoyiran, et al.
Pubblicazione: (2024)
How Well Do Models Follow Visual Instructions? VIBE: A Systematic Benchmark for Visual Instruction-Driven Image Editing
di: Zhang, Huanyu, et al.
Pubblicazione: (2026)
di: Zhang, Huanyu, et al.
Pubblicazione: (2026)
Deforming Videos to Masks: Flow Matching for Referring Video Segmentation
di: Wang, Zanyi, et al.
Pubblicazione: (2025)
di: Wang, Zanyi, et al.
Pubblicazione: (2025)
Beyond the Final Layer: Intermediate Representations for Better Multilingual Calibration in Large Language Models
di: Zhou, Ej, et al.
Pubblicazione: (2025)
di: Zhou, Ej, et al.
Pubblicazione: (2025)
RAVENEA: A Benchmark for Multimodal Retrieval-Augmented Visual Culture Understanding
di: Li, Jiaang, et al.
Pubblicazione: (2025)
di: Li, Jiaang, et al.
Pubblicazione: (2025)
On Bilingual Lexicon Induction with Large Language Models
di: Li, Yaoyiran, et al.
Pubblicazione: (2023)
di: Li, Yaoyiran, et al.
Pubblicazione: (2023)
Quantifying Language Disparities in Multilingual Large Language Models
di: Hu, Songbo, et al.
Pubblicazione: (2025)
di: Hu, Songbo, et al.
Pubblicazione: (2025)
Analyzing and Adapting Large Language Models for Few-Shot Multilingual NLU: Are We There Yet?
di: Razumovskaia, Evgeniia, et al.
Pubblicazione: (2024)
di: Razumovskaia, Evgeniia, et al.
Pubblicazione: (2024)
11Plus-Bench: Demystifying Multimodal LLM Spatial Reasoning with Cognitive-Inspired Analysis
di: Li, Chengzu, et al.
Pubblicazione: (2025)
di: Li, Chengzu, et al.
Pubblicazione: (2025)
Improving Bilingual Lexicon Induction with Cross-Encoder Reranking
di: Li, Yaoyiran, et al.
Pubblicazione: (2022)
di: Li, Yaoyiran, et al.
Pubblicazione: (2022)
Lost in Embeddings: Information Loss in Vision-Language Models
di: Li, Wenyan, et al.
Pubblicazione: (2025)
di: Li, Wenyan, et al.
Pubblicazione: (2025)
SQATIN: Supervised Instruction Tuning Meets Question Answering for Improved Dialogue NLU
di: Razumovskaia, Evgeniia, et al.
Pubblicazione: (2023)
di: Razumovskaia, Evgeniia, et al.
Pubblicazione: (2023)
Polyglot Teachers: Evaluating Language Models for Multilingual Synthetic Data Generation
di: Miranda, Lester James V., et al.
Pubblicazione: (2026)
di: Miranda, Lester James V., et al.
Pubblicazione: (2026)
Emergent Communication Pretraining for Few-Shot Machine Translation
di: Li, Yaoyiran, et al.
Pubblicazione: (2020)
di: Li, Yaoyiran, et al.
Pubblicazione: (2020)
Generalized Few-shot 3D Point Cloud Segmentation with Vision-Language Model
di: An, Zhaochong, et al.
Pubblicazione: (2025)
di: An, Zhaochong, et al.
Pubblicazione: (2025)
Unlocking the Potential of Grounding DINO in Videos: Parameter-Efficient Adaptation for Limited-Data Spatial-Temporal Localization
di: Wang, Zanyi, et al.
Pubblicazione: (2026)
di: Wang, Zanyi, et al.
Pubblicazione: (2026)
Scaling Sparse Fine-Tuning to Large Language Models
di: Ansell, Alan, et al.
Pubblicazione: (2024)
di: Ansell, Alan, et al.
Pubblicazione: (2024)
Multi-Modal Framing Analysis of News
di: Arora, Arnav, et al.
Pubblicazione: (2025)
di: Arora, Arnav, et al.
Pubblicazione: (2025)
Agentic Policy Optimization via Instruction-Policy Co-Evolution
di: Zhou, Han, et al.
Pubblicazione: (2025)
di: Zhou, Han, et al.
Pubblicazione: (2025)
AutoPEFT: Automatic Configuration Search for Parameter-Efficient Fine-Tuning
di: Zhou, Han, et al.
Pubblicazione: (2023)
di: Zhou, Han, et al.
Pubblicazione: (2023)
Translation-Enhanced Multilingual Text-to-Image Generation
di: Li, Yaoyiran, et al.
Pubblicazione: (2023)
di: Li, Yaoyiran, et al.
Pubblicazione: (2023)
ChatMotion: A Multimodal Multi-Agent for Human Motion Analysis
di: Li, Lei, et al.
Pubblicazione: (2025)
di: Li, Lei, et al.
Pubblicazione: (2025)
Evaluation of Cultural Competence of Vision-Language Models
di: Yadav, Srishti, et al.
Pubblicazione: (2025)
di: Yadav, Srishti, et al.
Pubblicazione: (2025)
DIALIGHT: Lightweight Multilingual Development and Evaluation of Task-Oriented Dialogue Systems with Large Language Models
di: Hu, Songbo, et al.
Pubblicazione: (2024)
di: Hu, Songbo, et al.
Pubblicazione: (2024)
Improving Word Translation via Two-Stage Contrastive Learning
di: Li, Yaoyiran, et al.
Pubblicazione: (2022)
di: Li, Yaoyiran, et al.
Pubblicazione: (2022)
VGGRPO: Towards World-Consistent Video Generation with 4D Latent Reward
di: An, Zhaochong, et al.
Pubblicazione: (2026)
di: An, Zhaochong, et al.
Pubblicazione: (2026)
Revisiting the Perception-Distortion Trade-off with Spatial-Semantic Guided Super-Resolution
di: Wang, Dan, et al.
Pubblicazione: (2026)
di: Wang, Dan, et al.
Pubblicazione: (2026)
Multimodality Helps Few-shot 3D Point Cloud Semantic Segmentation
di: An, Zhaochong, et al.
Pubblicazione: (2024)
di: An, Zhaochong, et al.
Pubblicazione: (2024)
Latent Sketchpad: Sketching Visual Thoughts to Elicit Multimodal Reasoning in MLLMs
di: Zhang, Huanyu, et al.
Pubblicazione: (2025)
di: Zhang, Huanyu, et al.
Pubblicazione: (2025)
Fairer Preferences Elicit Improved Human-Aligned Large Language Model Judgments
di: Zhou, Han, et al.
Pubblicazione: (2024)
di: Zhou, Han, et al.
Pubblicazione: (2024)
MMEarth-Bench: Global Model Adaptation via Multimodal Test-Time Training
di: Gordon, Lucia, et al.
Pubblicazione: (2026)
di: Gordon, Lucia, et al.
Pubblicazione: (2026)
Rethinking Few-shot 3D Point Cloud Semantic Segmentation
di: An, Zhaochong, et al.
Pubblicazione: (2024)
di: An, Zhaochong, et al.
Pubblicazione: (2024)
Donors and Recipients: On Asymmetric Transfer Across Tasks and Languages with Parameter-Efficient Fine-Tuning
di: Dymkiewicz, Kajetan, et al.
Pubblicazione: (2025)
di: Dymkiewicz, Kajetan, et al.
Pubblicazione: (2025)
Better Language Models Exhibit Higher Visual Alignment
di: Ruthardt, Jona, et al.
Pubblicazione: (2024)
di: Ruthardt, Jona, et al.
Pubblicazione: (2024)
When Thinking Hurts: Mitigating Visual Forgetting in Video Reasoning via Frame Repetition
di: Sun, Xiaokun, et al.
Pubblicazione: (2026)
di: Sun, Xiaokun, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Visual Planning: Let's Think Only with Images
di: Xu, Yi, et al.
Pubblicazione: (2025) -
Large Language Models are Miscalibrated In-Context Learners
di: Li, Chengzu, et al.
Pubblicazione: (2023) -
TopViewRS: Vision-Language Models as Top-View Spatial Reasoners
di: Li, Chengzu, et al.
Pubblicazione: (2024) -
Video Understanding: From Geometry and Semantics to Unified Models
di: An, Zhaochong, et al.
Pubblicazione: (2026) -
Imagine while Reasoning in Space: Multimodal Visualization-of-Thought
di: Li, Chengzu, et al.
Pubblicazione: (2025)