When Thinking Hurts: Mitigating Visual Forgetting in Video Reasoning via Frame Repetition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sun, Xiaokun, Wang, Yubo, Cao, Haoyu, Xu, Linli |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models
von: Tang, Jianting, et al.
Veröffentlicht: (2025)
von: Tang, Jianting, et al.
Veröffentlicht: (2025)
MVP: Enhancing Video Large Language Models via Self-supervised Masked Video Prediction
von: Sun, Xiaokun, et al.
Veröffentlicht: (2026)
von: Sun, Xiaokun, et al.
Veröffentlicht: (2026)
Break the Visual Perception: Adversarial Attacks Targeting Encoded Visual Tokens of Large Vision-Language Models
von: Wang, Yubo, et al.
Veröffentlicht: (2024)
von: Wang, Yubo, et al.
Veröffentlicht: (2024)
Dynamic Token Compression for Efficient Video Understanding through Reinforcement Learning
von: Wang, Shida, et al.
Veröffentlicht: (2026)
von: Wang, Shida, et al.
Veröffentlicht: (2026)
Thinking with Frames: Generative Video Distortion Evaluation via Frame Reward Model
von: Wang, Yuan, et al.
Veröffentlicht: (2026)
von: Wang, Yuan, et al.
Veröffentlicht: (2026)
Thinking Before Looking: Improving Multimodal LLM Reasoning via Mitigating Visual Hallucination
von: Zheng, Haojie, et al.
Veröffentlicht: (2024)
von: Zheng, Haojie, et al.
Veröffentlicht: (2024)
Generative Frame Sampler for Long Video Understanding
von: Yao, Linli, et al.
Veröffentlicht: (2025)
von: Yao, Linli, et al.
Veröffentlicht: (2025)
Mitigating Visual Forgetting via Take-along Visual Conditioning for Multi-modal Long CoT Reasoning
von: Sun, Hai-Long, et al.
Veröffentlicht: (2025)
von: Sun, Hai-Long, et al.
Veröffentlicht: (2025)
Geometry without Position? When Positional Embeddings Help and Hurt Spatial Reasoning
von: Shi, Jian, et al.
Veröffentlicht: (2026)
von: Shi, Jian, et al.
Veröffentlicht: (2026)
Thinking in Frames: How Visual Context and Test-Time Scaling Empower Video Reasoning
von: Li, Chengzu, et al.
Veröffentlicht: (2026)
von: Li, Chengzu, et al.
Veröffentlicht: (2026)
FrameThinker: Learning to Think with Long Videos via Multi-Turn Frame Spotlighting
von: He, Zefeng, et al.
Veröffentlicht: (2025)
von: He, Zefeng, et al.
Veröffentlicht: (2025)
HRVDA: High-Resolution Visual Document Assistant
von: Liu, Chaohu, et al.
Veröffentlicht: (2024)
von: Liu, Chaohu, et al.
Veröffentlicht: (2024)
When RAG Hurts: Diagnosing and Mitigating Attention Distraction in Retrieval-Augmented LVLMs
von: Zhao, Beidi, et al.
Veröffentlicht: (2026)
von: Zhao, Beidi, et al.
Veröffentlicht: (2026)
Temporal Reasoning Transfer from Text to Video
von: Li, Lei, et al.
Veröffentlicht: (2024)
von: Li, Lei, et al.
Veröffentlicht: (2024)
DiG: Differential Grounding for Enhancing Fine-Grained Perception in Multimodal Large Language Model
von: Tao, Zhou, et al.
Veröffentlicht: (2025)
von: Tao, Zhou, et al.
Veröffentlicht: (2025)
When Thinking Drifts: Evidential Grounding for Robust Video Reasoning
von: Luo, Mi, et al.
Veröffentlicht: (2025)
von: Luo, Mi, et al.
Veröffentlicht: (2025)
Chain-of-Frames: Advancing Video Understanding in Multimodal LLMs via Frame-Aware Reasoning
von: Ghazanfari, Sara, et al.
Veröffentlicht: (2025)
von: Ghazanfari, Sara, et al.
Veröffentlicht: (2025)
FrameMind: Frame-Interleaved Video Reasoning via Reinforcement Learning
von: Ge, Haonan, et al.
Veröffentlicht: (2025)
von: Ge, Haonan, et al.
Veröffentlicht: (2025)
Reinforcing Video Reasoning with Focused Thinking
von: Dang, Jisheng, et al.
Veröffentlicht: (2025)
von: Dang, Jisheng, et al.
Veröffentlicht: (2025)
TwiFF (Think With Future Frames): A Large-Scale Dataset for Dynamic Visual Reasoning
von: Liu, Junhua, et al.
Veröffentlicht: (2026)
von: Liu, Junhua, et al.
Veröffentlicht: (2026)
Relaxing Anchor-Frame Dominance for Mitigating Hallucinations in Video Large Language Models
von: Liu, Zijian, et al.
Veröffentlicht: (2026)
von: Liu, Zijian, et al.
Veröffentlicht: (2026)
Conan: Progressive Learning to Reason Like a Detective over Multi-Scale Visual Evidence
von: Ouyang, Kun, et al.
Veröffentlicht: (2025)
von: Ouyang, Kun, et al.
Veröffentlicht: (2025)
Think, then Score: Decoupled Reasoning and Scoring for Video Reward Modeling
von: Wang, Yuan, et al.
Veröffentlicht: (2026)
von: Wang, Yuan, et al.
Veröffentlicht: (2026)
Thinking with Drafts: Speculative Temporal Reasoning for Efficient Long Video Understanding
von: Hu, Pengfei, et al.
Veröffentlicht: (2025)
von: Hu, Pengfei, et al.
Veröffentlicht: (2025)
MCF-VC: Mitigate Catastrophic Forgetting in Class-Incremental Learning for Multimodal Video Captioning
von: Xiong, Huiyu, et al.
Veröffentlicht: (2024)
von: Xiong, Huiyu, et al.
Veröffentlicht: (2024)
DeepLatent: Think with Images via Parallel Latent Visual Reasoning
von: Lu, Dongchen, et al.
Veröffentlicht: (2026)
von: Lu, Dongchen, et al.
Veröffentlicht: (2026)
VideoAuto-R1: Video Auto Reasoning via Thinking Once, Answering Twice
von: Liu, Shuming, et al.
Veröffentlicht: (2026)
von: Liu, Shuming, et al.
Veröffentlicht: (2026)
Pathological Prior-Guided Multiple Instance Learning For Mitigating Catastrophic Forgetting in Breast Cancer Whole Slide Image Classification
von: Zheng, Weixi, et al.
Veröffentlicht: (2025)
von: Zheng, Weixi, et al.
Veröffentlicht: (2025)
Reinforcing Video Reasoning Segmentation to Think Before It Segments
von: Gong, Sitong, et al.
Veröffentlicht: (2025)
von: Gong, Sitong, et al.
Veröffentlicht: (2025)
VideoChat-A1: Thinking with Long Videos by Chain-of-Shot Reasoning
von: Wang, Zikang, et al.
Veröffentlicht: (2025)
von: Wang, Zikang, et al.
Veröffentlicht: (2025)
Graph-to-Frame RAG: Visual-Space Knowledge Fusion for Training-Free and Auditable Video Reasoning
von: Yang, Songyuan, et al.
Veröffentlicht: (2026)
von: Yang, Songyuan, et al.
Veröffentlicht: (2026)
TACO: Think-Answer Consistency for Optimized Long-Chain Reasoning and Efficient Data Learning via Reinforcement Learning in LVLMs
von: Kan, Zhehan, et al.
Veröffentlicht: (2025)
von: Kan, Zhehan, et al.
Veröffentlicht: (2025)
Finding the Correct Visual Evidence Without Forgetting: Mitigating Hallucination in LVLMs via Inter-Layer Visual Attention Discrepancy
von: Xie, Yutong, et al.
Veröffentlicht: (2026)
von: Xie, Yutong, et al.
Veröffentlicht: (2026)
Mitigating Visual Knowledge Forgetting in MLLM Instruction-tuning via Modality-decoupled Gradient Descent
von: Wu, Junda, et al.
Veröffentlicht: (2025)
von: Wu, Junda, et al.
Veröffentlicht: (2025)
VTimeCoT: Thinking by Drawing for Video Temporal Grounding and Reasoning
von: Zhang, Jinglei, et al.
Veröffentlicht: (2025)
von: Zhang, Jinglei, et al.
Veröffentlicht: (2025)
EDEN: Enhanced Diffusion for High-quality Large-motion Video Frame Interpolation
von: Zhang, Zihao, et al.
Veröffentlicht: (2025)
von: Zhang, Zihao, et al.
Veröffentlicht: (2025)
Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?
von: Cheng, Junhao, et al.
Veröffentlicht: (2025)
von: Cheng, Junhao, et al.
Veröffentlicht: (2025)
MIRROR: Multimodal Iterative Reasoning via Reflection on Visual Regions
von: Zhang, Haoyu, et al.
Veröffentlicht: (2026)
von: Zhang, Haoyu, et al.
Veröffentlicht: (2026)
VideoSTF: Stress-Testing Output Repetition in Video Large Language Models
von: Cao, Yuxin, et al.
Veröffentlicht: (2026)
von: Cao, Yuxin, et al.
Veröffentlicht: (2026)
Beyond Shortcuts: Mitigating Visual Illusions in Frozen VLMs via Qualitative Reasoning
von: Guo, Hao, et al.
Veröffentlicht: (2026)
von: Guo, Hao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
BASIC: Boosting Visual Alignment with Intrinsic Refined Embeddings in Multimodal Large Language Models
von: Tang, Jianting, et al.
Veröffentlicht: (2025) -
MVP: Enhancing Video Large Language Models via Self-supervised Masked Video Prediction
von: Sun, Xiaokun, et al.
Veröffentlicht: (2026) -
Break the Visual Perception: Adversarial Attacks Targeting Encoded Visual Tokens of Large Vision-Language Models
von: Wang, Yubo, et al.
Veröffentlicht: (2024) -
Dynamic Token Compression for Efficient Video Understanding through Reinforcement Learning
von: Wang, Shida, et al.
Veröffentlicht: (2026) -
Thinking with Frames: Generative Video Distortion Evaluation via Frame Reward Model
von: Wang, Yuan, et al.
Veröffentlicht: (2026)