ReGATE: Learning Faster and Better with Fewer Tokens in MLLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Chaoyu, Kulkarni, Yogesh, Fazli, Pooyan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EgoVITA: Learning to Plan and Verify for Egocentric Video Reasoning
by: Kulkarni, Yogesh, et al.
Published: (2025)
by: Kulkarni, Yogesh, et al.
Published: (2025)
AVATAR: Reinforcement Learning to See, Hear, and Reason Over Video
by: Kulkarni, Yogesh, et al.
Published: (2025)
by: Kulkarni, Yogesh, et al.
Published: (2025)
VideoPASTA: 7K Preference Pairs That Matter for Video-LLM Alignment
by: Kulkarni, Yogesh, et al.
Published: (2025)
by: Kulkarni, Yogesh, et al.
Published: (2025)
VideoSAVi: Self-Aligned Video Language Models without Human Supervision
by: Kulkarni, Yogesh, et al.
Published: (2024)
by: Kulkarni, Yogesh, et al.
Published: (2024)
VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding
by: Li, Chaoyu, et al.
Published: (2024)
by: Li, Chaoyu, et al.
Published: (2024)
CASHEW: Stabilizing Multimodal Reasoning via Iterative Trajectory Aggregation
by: Li, Chaoyu, et al.
Published: (2026)
by: Li, Chaoyu, et al.
Published: (2026)
VideoA11y: Method and Dataset for Accessible Video Description
by: Li, Chaoyu, et al.
Published: (2025)
by: Li, Chaoyu, et al.
Published: (2025)
ChartQA-X: Generating Explanations for Visual Chart Reasoning
by: Hegde, Shamanthak, et al.
Published: (2025)
by: Hegde, Shamanthak, et al.
Published: (2025)
FrameOracle: Learning What to See and How Much to See in Videos
by: Li, Chaoyu, et al.
Published: (2025)
by: Li, Chaoyu, et al.
Published: (2025)
Re:Verse -- Can Your VLM Read a Manga?
by: Baranwal, Aaditya, et al.
Published: (2025)
by: Baranwal, Aaditya, et al.
Published: (2025)
OSCaR: Object State Captioning and State Change Representation
by: Nguyen, Nguyen, et al.
Published: (2024)
by: Nguyen, Nguyen, et al.
Published: (2024)
HiDrop: Hierarchical Vision Token Reduction in MLLMs via Late Injection, Concave Pyramid Pruning, and Early Exit
by: Wu, Hao, et al.
Published: (2026)
by: Wu, Hao, et al.
Published: (2026)
D$^{3}$ToM: Decider-Guided Dynamic Token Merging for Accelerating Diffusion MLLMs
by: Chang, Shuochen, et al.
Published: (2025)
by: Chang, Shuochen, et al.
Published: (2025)
Linking Perception, Confidence and Accuracy in MLLMs
by: Du, Yuetian, et al.
Published: (2026)
by: Du, Yuetian, et al.
Published: (2026)
The Better You Learn, The Smarter You Prune: Towards Efficient Vision-language-action Models via Differentiable Token Pruning
by: Jiang, Titong, et al.
Published: (2025)
by: Jiang, Titong, et al.
Published: (2025)
CoRe^2: Collect, Reflect and Refine to Generate Better and Faster
by: Shao, Shitong, et al.
Published: (2025)
by: Shao, Shitong, et al.
Published: (2025)
From Text to Pixel: Advancing Long-Context Understanding in MLLMs
by: Lu, Yujie, et al.
Published: (2024)
by: Lu, Yujie, et al.
Published: (2024)
An Empirical Study on Configuring In-Context Learning Demonstrations for Unleashing MLLMs' Sentimental Perception Capability
by: Wu, Daiqing, et al.
Published: (2025)
by: Wu, Daiqing, et al.
Published: (2025)
MMDuet2: Enhancing Proactive Interaction of Video MLLMs with Multi-Turn Reinforcement Learning
by: Wang, Yueqian, et al.
Published: (2025)
by: Wang, Yueqian, et al.
Published: (2025)
Unhackable Temporal Rewarding for Scalable Video MLLMs
by: Yu, En, et al.
Published: (2025)
by: Yu, En, et al.
Published: (2025)
Dyna-Mind: Learning to Simulate from Experience for Better AI Agents
by: Yu, Xiao, et al.
Published: (2025)
by: Yu, Xiao, et al.
Published: (2025)
Latent Sketchpad: Sketching Visual Thoughts to Elicit Multimodal Reasoning in MLLMs
by: Zhang, Huanyu, et al.
Published: (2025)
by: Zhang, Huanyu, et al.
Published: (2025)
LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-Steering
by: Bi, Jinhe, et al.
Published: (2024)
by: Bi, Jinhe, et al.
Published: (2024)
Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No!
by: Imam, Mohamed Fazli, et al.
Published: (2025)
by: Imam, Mohamed Fazli, et al.
Published: (2025)
Human Cognitive Benchmarks Reveal Foundational Visual Gaps in MLLMs
by: Huang, Jen-Tse, et al.
Published: (2025)
by: Huang, Jen-Tse, et al.
Published: (2025)
CameraBench: Benchmarking Visual Reasoning in MLLMs via Photography
by: Fang, I-Sheng, et al.
Published: (2025)
by: Fang, I-Sheng, et al.
Published: (2025)
Leveraging the Power of MLLMs for Gloss-Free Sign Language Translation
by: Kim, Jungeun, et al.
Published: (2024)
by: Kim, Jungeun, et al.
Published: (2024)
RefereeBench: Are Video MLLMs Ready to be Multi-Sport Referees
by: Xu, Yichen, et al.
Published: (2026)
by: Xu, Yichen, et al.
Published: (2026)
Drawing the Line: Enhancing Trustworthiness of MLLMs Through the Power of Refusal
by: Wang, Yuhao, et al.
Published: (2024)
by: Wang, Yuhao, et al.
Published: (2024)
Exploring the Design Space of Visual Context Representation in Video MLLMs
by: Du, Yifan, et al.
Published: (2024)
by: Du, Yifan, et al.
Published: (2024)
Injecting Distributional Awareness into MLLMs via Reinforcement Learning for Deep Imbalanced Regression
by: Du, Yao, et al.
Published: (2026)
by: Du, Yao, et al.
Published: (2026)
Make LVLMs Focus: Context-Aware Attention Modulation for Better Multimodal In-Context Learning
by: Li, Yanshu, et al.
Published: (2025)
by: Li, Yanshu, et al.
Published: (2025)
Faster Training, Fewer Labels: Self-Supervised Pretraining for Fine-Grained BEV Segmentation
by: Busch, Daniel, et al.
Published: (2026)
by: Busch, Daniel, et al.
Published: (2026)
Seeing is Believing: Rich-Context Hallucination Detection for MLLMs via Backward Visual Grounding
by: Guo, Pinxue, et al.
Published: (2025)
by: Guo, Pinxue, et al.
Published: (2025)
VidComposition: Can MLLMs Analyze Compositions in Compiled Videos?
by: Tang, Yolo Y., et al.
Published: (2024)
by: Tang, Yolo Y., et al.
Published: (2024)
Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs
by: Yeh, Chun-Hsiao, et al.
Published: (2025)
by: Yeh, Chun-Hsiao, et al.
Published: (2025)
AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation
by: Wang, Junyang, et al.
Published: (2023)
by: Wang, Junyang, et al.
Published: (2023)
Octavius: Mitigating Task Interference in MLLMs via LoRA-MoE
by: Chen, Zeren, et al.
Published: (2023)
by: Chen, Zeren, et al.
Published: (2023)
Grandes Modelos de Linguagem Multimodais (MLLMs): Da Teoria à Prática
by: da Silva, Neemias, et al.
Published: (2026)
by: da Silva, Neemias, et al.
Published: (2026)
The Side Effects of Being Smart: Safety Risks in MLLMs' Multi-Image Reasoning
by: Chen, Renmiao, et al.
Published: (2026)
by: Chen, Renmiao, et al.
Published: (2026)
Similar Items
-
EgoVITA: Learning to Plan and Verify for Egocentric Video Reasoning
by: Kulkarni, Yogesh, et al.
Published: (2025) -
AVATAR: Reinforcement Learning to See, Hear, and Reason Over Video
by: Kulkarni, Yogesh, et al.
Published: (2025) -
VideoPASTA: 7K Preference Pairs That Matter for Video-LLM Alignment
by: Kulkarni, Yogesh, et al.
Published: (2025) -
VideoSAVi: Self-Aligned Video Language Models without Human Supervision
by: Kulkarni, Yogesh, et al.
Published: (2024) -
VidHalluc: Evaluating Temporal Hallucinations in Multimodal Large Language Models for Video Understanding
by: Li, Chaoyu, et al.
Published: (2024)