Gespeichert in:
| Hauptverfasser: | Chen, Shuhang, Yuan, Hangjie, Xu, Yunqiu, Liu, Pengwei, Feng, Tao, Cen, Jun, Huang, Zeying, Yang, Yi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2503.16549 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CogFlow: Bridging Perception and Reasoning through Knowledge Internalization for Visual Mathematical Problem Solving
von: Chen, Shuhang, et al.
Veröffentlicht: (2026)
von: Chen, Shuhang, et al.
Veröffentlicht: (2026)
SAMora: Enhancing SAM through Hierarchical Self-Supervised Pre-Training for Medical Images
von: Chen, Shuhang, et al.
Veröffentlicht: (2025)
von: Chen, Shuhang, et al.
Veröffentlicht: (2025)
MC-Bench: A Benchmark for Multi-Context Visual Grounding in the Era of MLLMs
von: Xu, Yunqiu, et al.
Veröffentlicht: (2024)
von: Xu, Yunqiu, et al.
Veröffentlicht: (2024)
Echoes of ownership: Adversarial-guided dual injection for copyright protection in MLLMs
von: Xia, Chengwei, et al.
Veröffentlicht: (2026)
von: Xia, Chengwei, et al.
Veröffentlicht: (2026)
LumosFlow: Motion-Guided Long Video Generation
von: Chen, Jiahao, et al.
Veröffentlicht: (2025)
von: Chen, Jiahao, et al.
Veröffentlicht: (2025)
Knowledge is Power: Advancing Few-shot Action Recognition with Multimodal Semantics from MLLMs
von: Xing, Jiazheng, et al.
Veröffentlicht: (2026)
von: Xing, Jiazheng, et al.
Veröffentlicht: (2026)
ZeroFlow: Overcoming Catastrophic Forgetting is Easier than You Think
von: Feng, Tao, et al.
Veröffentlicht: (2025)
von: Feng, Tao, et al.
Veröffentlicht: (2025)
Open Eyes, Then Reason: Fine-grained Visual Mathematical Understanding in MLLMs
von: Zhang, Shan, et al.
Veröffentlicht: (2025)
von: Zhang, Shan, et al.
Veröffentlicht: (2025)
Lumos-1: On Autoregressive Video Generation with Discrete Diffusion from a Unified Model Perspective
von: Yuan, Hangjie, et al.
Veröffentlicht: (2025)
von: Yuan, Hangjie, et al.
Veröffentlicht: (2025)
DFVEdit: Conditional Delta Flow Vector for Zero-shot Video Editing
von: Cai, Lingling, et al.
Veröffentlicht: (2025)
von: Cai, Lingling, et al.
Veröffentlicht: (2025)
Aesthetic Image Captioning with Saliency Enhanced MLLMs
von: Tao, Yilin, et al.
Veröffentlicht: (2025)
von: Tao, Yilin, et al.
Veröffentlicht: (2025)
GridPrune: From "Where to Look" to "What to Select" in Visual Token Pruning for MLLMs
von: Duan, Yuxiang, et al.
Veröffentlicht: (2025)
von: Duan, Yuxiang, et al.
Veröffentlicht: (2025)
Curriculum Sampling: A Two-Phase Curriculum for Efficient Training of Flow Matching
von: Sun, Pengwei
Veröffentlicht: (2026)
von: Sun, Pengwei
Veröffentlicht: (2026)
Perceptual Flow Network for Visually Grounded Reasoning
von: Li, Yangfu, et al.
Veröffentlicht: (2026)
von: Li, Yangfu, et al.
Veröffentlicht: (2026)
The Perceptual Observatory Characterizing Robustness and Grounding in MLLMs
von: Anvekar, Tejas, et al.
Veröffentlicht: (2025)
von: Anvekar, Tejas, et al.
Veröffentlicht: (2025)
Math Blind: Failures in Diagram Understanding Undermine Reasoning in MLLMs
von: Sun, Yanpeng, et al.
Veröffentlicht: (2025)
von: Sun, Yanpeng, et al.
Veröffentlicht: (2025)
Rethinking the Stability-Plasticity Trade-off in Continual Learning from an Architectural Perspective
von: Lu, Aojun, et al.
Veröffentlicht: (2025)
von: Lu, Aojun, et al.
Veröffentlicht: (2025)
IF-Bench: Benchmarking and Enhancing MLLMs for Infrared Images with Generative Visual Prompting
von: Zhang, Tao, et al.
Veröffentlicht: (2025)
von: Zhang, Tao, et al.
Veröffentlicht: (2025)
Do MLLMs Exhibit Human-like Perceptual Behaviors? HVSBench: A Benchmark for MLLM Alignment with Human Perceptual Behavior
von: Lin, Jiaying, et al.
Veröffentlicht: (2024)
von: Lin, Jiaying, et al.
Veröffentlicht: (2024)
Lifting the Veil on Visual Information Flow in MLLMs: Unlocking Pathways to Faster Inference
von: Yin, Hao, et al.
Veröffentlicht: (2025)
von: Yin, Hao, et al.
Veröffentlicht: (2025)
MathSticks: A Benchmark for Visual Symbolic Compositional Reasoning with Matchstick Puzzles
von: Ji, Yuheng, et al.
Veröffentlicht: (2025)
von: Ji, Yuheng, et al.
Veröffentlicht: (2025)
We-Math 2.0: A Versatile MathBook System for Incentivizing Visual Mathematical Reasoning
von: Qiao, Runqi, et al.
Veröffentlicht: (2025)
von: Qiao, Runqi, et al.
Veröffentlicht: (2025)
VOILA: Evaluation of MLLMs For Perceptual Understanding and Analogical Reasoning
von: Yilmaz, Nilay, et al.
Veröffentlicht: (2025)
von: Yilmaz, Nilay, et al.
Veröffentlicht: (2025)
When Policy Entropy Constraint Fails: Preserving Diversity in Flow-based RLHF via Perceptual Entropy
von: Tan, Xiaofeng, et al.
Veröffentlicht: (2026)
von: Tan, Xiaofeng, et al.
Veröffentlicht: (2026)
MathCanvas: Intrinsic Visual Chain-of-Thought for Multimodal Mathematical Reasoning
von: Shi, Weikang, et al.
Veröffentlicht: (2025)
von: Shi, Weikang, et al.
Veröffentlicht: (2025)
Adapt before Continual Learning
von: Lu, Aojun, et al.
Veröffentlicht: (2025)
von: Lu, Aojun, et al.
Veröffentlicht: (2025)
Why Does RL Generalize Better Than SFT? A Data-Centric Perspective on VLM Post-Training
von: Lu, Aojun, et al.
Veröffentlicht: (2026)
von: Lu, Aojun, et al.
Veröffentlicht: (2026)
Revisiting Neural Networks for Continual Learning: An Architectural Perspective
von: Lu, Aojun, et al.
Veröffentlicht: (2024)
von: Lu, Aojun, et al.
Veröffentlicht: (2024)
Context Tokens are Anchors: Understanding the Repetition Curse in dMLLMs from an Information Flow Perspective
von: Zhao, Qiyan, et al.
Veröffentlicht: (2026)
von: Zhao, Qiyan, et al.
Veröffentlicht: (2026)
VisAidMath: Benchmarking Visual-Aided Mathematical Reasoning
von: Ma, Jingkun, et al.
Veröffentlicht: (2024)
von: Ma, Jingkun, et al.
Veröffentlicht: (2024)
GSM8K-V: Can Vision Language Models Solve Grade School Math Word Problems in Visual Contexts
von: Yuan, Fan, et al.
Veröffentlicht: (2025)
von: Yuan, Fan, et al.
Veröffentlicht: (2025)
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking
von: Meng, Desen, et al.
Veröffentlicht: (2025)
von: Meng, Desen, et al.
Veröffentlicht: (2025)
PriOr-Flow: Enhancing Primitive Panoramic Optical Flow with Orthogonal View
von: Liu, Longliang, et al.
Veröffentlicht: (2025)
von: Liu, Longliang, et al.
Veröffentlicht: (2025)
Branch, or Layer? Zeroth-Order Optimization for Continual Learning of Vision-Language Models
von: Liu, Ziwei, et al.
Veröffentlicht: (2025)
von: Liu, Ziwei, et al.
Veröffentlicht: (2025)
Law of Vision Representation in MLLMs
von: Yang, Shijia, et al.
Veröffentlicht: (2024)
von: Yang, Shijia, et al.
Veröffentlicht: (2024)
ControlGUI: Guiding Generative GUI Exploration through Perceptual Visual Flow
von: Garg, Aryan, et al.
Veröffentlicht: (2025)
von: Garg, Aryan, et al.
Veröffentlicht: (2025)
InfiMM-WebMath-40B: Advancing Multimodal Pre-Training for Enhanced Mathematical Reasoning
von: Han, Xiaotian, et al.
Veröffentlicht: (2024)
von: Han, Xiaotian, et al.
Veröffentlicht: (2024)
Human Cognitive Benchmarks Reveal Foundational Visual Gaps in MLLMs
von: Huang, Jen-Tse, et al.
Veröffentlicht: (2025)
von: Huang, Jen-Tse, et al.
Veröffentlicht: (2025)
UAVBench and UAVIT-1M: Benchmarking and Enhancing MLLMs for Low-Altitude UAV Vision-Language Understanding
von: Zhan, Yang, et al.
Veröffentlicht: (2026)
von: Zhan, Yang, et al.
Veröffentlicht: (2026)
A Faster Path to Continual Learning
von: Li, Wei, et al.
Veröffentlicht: (2026)
von: Li, Wei, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
CogFlow: Bridging Perception and Reasoning through Knowledge Internalization for Visual Mathematical Problem Solving
von: Chen, Shuhang, et al.
Veröffentlicht: (2026) -
SAMora: Enhancing SAM through Hierarchical Self-Supervised Pre-Training for Medical Images
von: Chen, Shuhang, et al.
Veröffentlicht: (2025) -
MC-Bench: A Benchmark for Multi-Context Visual Grounding in the Era of MLLMs
von: Xu, Yunqiu, et al.
Veröffentlicht: (2024) -
Echoes of ownership: Adversarial-guided dual injection for copyright protection in MLLMs
von: Xia, Chengwei, et al.
Veröffentlicht: (2026) -
LumosFlow: Motion-Guided Long Video Generation
von: Chen, Jiahao, et al.
Veröffentlicht: (2025)