Learning to Focus and Precise Cropping: A Reinforcement Learning Framework with Information Gaps and Grounding Loss for MLLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Zhao, Xuanpu, Tan, Zhentao, Sheng, Dianmo, Chen, Tianxiang, Liu, Yao, Wu, Yue, Gong, Tao, Chu, Qi, Yu, Nenghai |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Towards More Unified In-context Visual Understanding
by: Sheng, Dianmo, et al.
Published: (2023)
by: Sheng, Dianmo, et al.
Published: (2023)
TCI-Former: Thermal Conduction-Inspired Transformer for Infrared Small Target Detection
by: Chen, Tianxiang, et al.
Published: (2024)
by: Chen, Tianxiang, et al.
Published: (2024)
Flora: Effortless Context Construction to Arbitrary Length and Scale
by: Chen, Tianxiang, et al.
Published: (2025)
by: Chen, Tianxiang, et al.
Published: (2025)
Llama SLayer 8B: Shallow Layers Hold the Key to Knowledge Injection
by: Chen, Tianxiang, et al.
Published: (2024)
by: Chen, Tianxiang, et al.
Published: (2024)
SAPL: Semantic-Agnostic Prompt Learning in CLIP for Weakly Supervised Image Manipulation Localization
by: Wang, Xinghao, et al.
Published: (2026)
by: Wang, Xinghao, et al.
Published: (2026)
Bootstrapping Audio-Visual Segmentation by Strengthening Audio Cues
by: Chen, Tianxiang, et al.
Published: (2024)
by: Chen, Tianxiang, et al.
Published: (2024)
MiM-ISTD: Mamba-in-Mamba for Efficient Infrared Small Target Detection
by: Chen, Tianxiang, et al.
Published: (2024)
by: Chen, Tianxiang, et al.
Published: (2024)
Transformer based Pluralistic Image Completion with Reduced Information Loss
by: Liu, Qiankun, et al.
Published: (2024)
by: Liu, Qiankun, et al.
Published: (2024)
Multi-spectral Class Center Network for Face Manipulation Detection and Localization
by: Miao, Changtao, et al.
Published: (2023)
by: Miao, Changtao, et al.
Published: (2023)
Mixture-of-Noises Enhanced Forgery-Aware Predictor for Multi-Face Manipulation Detection and Localization
by: Miao, Changtao, et al.
Published: (2024)
by: Miao, Changtao, et al.
Published: (2024)
Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning
by: Zhang, Bob, et al.
Published: (2025)
by: Zhang, Bob, et al.
Published: (2025)
Mean-squared Energy Difference for Exploring Potential Energy Landscapes of Supercooled Liquids
by: Zhang, Dianmo, et al.
Published: (2025)
by: Zhang, Dianmo, et al.
Published: (2025)
Context-Aware Weakly Supervised Image Manipulation Localization with SAM Refinement
by: Wang, Xinghao, et al.
Published: (2025)
by: Wang, Xinghao, et al.
Published: (2025)
SPARROW: Learning Spatial Precision and Temporal Referential Consistency in Pixel-Grounded Video MLLMs
by: Alansari, Mohamad, et al.
Published: (2026)
by: Alansari, Mohamad, et al.
Published: (2026)
GroundingME: Exposing the Visual Grounding Gap in MLLMs through Multi-Dimensional Evaluation
by: Li, Rang, et al.
Published: (2025)
by: Li, Rang, et al.
Published: (2025)
Precision Agriculture: Crop Mapping using Machine Learning and Sentinel-2 Satellite Imagery
by: Zhao, Kui, et al.
Published: (2023)
by: Zhao, Kui, et al.
Published: (2023)
Precision-Focused Reinforcement Learning Model for Robotic Object Pushing
by: Bergmann, Lara, et al.
Published: (2024)
by: Bergmann, Lara, et al.
Published: (2024)
Exploiting Modality-Specific Features For Multi-Modal Manipulation Detection And Grounding
by: Wang, Jiazhen, et al.
Published: (2023)
by: Wang, Jiazhen, et al.
Published: (2023)
Learning Solution-Aware Transformers for Efficiently Solving Quadratic Assignment Problem
by: Tan, Zhentao, et al.
Published: (2024)
by: Tan, Zhentao, et al.
Published: (2024)
Injecting Distributional Awareness into MLLMs via Reinforcement Learning for Deep Imbalanced Regression
by: Du, Yao, et al.
Published: (2026)
by: Du, Yao, et al.
Published: (2026)
UDQL: Bridging The Gap between MSE Loss and The Optimal Value Function in Offline Reinforcement Learning
by: Zhang, Yu, et al.
Published: (2024)
by: Zhang, Yu, et al.
Published: (2024)
TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning
by: Zeng, Xiangyu, et al.
Published: (2024)
by: Zeng, Xiangyu, et al.
Published: (2024)
FishBEV: Distortion-Resilient Bird's Eye View Segmentation with Surround-View Fisheye Cameras
by: Li, Hang, et al.
Published: (2025)
by: Li, Hang, et al.
Published: (2025)
Observe-R1: Unlocking Reasoning Abilities of MLLMs with Dynamic Progressive Reinforcement Learning
by: Guo, Zirun, et al.
Published: (2025)
by: Guo, Zirun, et al.
Published: (2025)
When Agents Look the Same: Quantifying Distillation-Induced Similarity in Tool-Use Behaviors
by: Yang, Chenghao, et al.
Published: (2026)
by: Yang, Chenghao, et al.
Published: (2026)
Robust Losses for Decision-Focused Learning
by: Schutte, Noah, et al.
Published: (2023)
by: Schutte, Noah, et al.
Published: (2023)
Radial 3D Focusing Energy Critical INLS equations with defocusing perturbation: Ground states, Scattering, and Blow-up
by: Gou, Tianxiang, et al.
Published: (2024)
by: Gou, Tianxiang, et al.
Published: (2024)
VideoZoomer: Reinforcement-Learned Temporal Focusing for Long Video Reasoning
by: Ding, Yang, et al.
Published: (2025)
by: Ding, Yang, et al.
Published: (2025)
Video-R1: Reinforcing Video Reasoning in MLLMs
by: Feng, Kaituo, et al.
Published: (2025)
by: Feng, Kaituo, et al.
Published: (2025)
PsychēChat: An Empathic Framework Focused on Emotion Shift Tracking and Safety Risk Analysis in Psychological Counseling
by: Xia, Zhentao, et al.
Published: (2026)
by: Xia, Zhentao, et al.
Published: (2026)
LAKAN: Landmark-assisted Adaptive Kolmogorov-Arnold Network for Face Forgery Detection
by: Jiang, Jiayao, et al.
Published: (2025)
by: Jiang, Jiayao, et al.
Published: (2025)
STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?
by: Li, Yun, et al.
Published: (2025)
by: Li, Yun, et al.
Published: (2025)
ClaHF: A Human Feedback-inspired Reinforcement Learning Framework for Improving Classification Tasks
by: Xu, Tianxiang, et al.
Published: (2026)
by: Xu, Tianxiang, et al.
Published: (2026)
Prediction Loss Guided Decision-Focused Learning
by: Jeon, Haeun, et al.
Published: (2025)
by: Jeon, Haeun, et al.
Published: (2025)
Learning Where, What and How to Transfer: A Multi-Role Reinforcement Learning Approach for Evolutionary Multitasking
by: Zhan, Jiajun, et al.
Published: (2025)
by: Zhan, Jiajun, et al.
Published: (2025)
Training-Free In-Context Forensic Chain for Image Manipulation Detection and Localization
by: Chen, Rui, et al.
Published: (2025)
by: Chen, Rui, et al.
Published: (2025)
GuardTrace-VL: Detecting Unsafe Multimodel Reasoning via Iterative Safety Supervision
by: Xiang, Yuxiao, et al.
Published: (2025)
by: Xiang, Yuxiao, et al.
Published: (2025)
Compositional Learning of Visually-Grounded Concepts Using Reinforcement
by: Lin, Zijun, et al.
Published: (2023)
by: Lin, Zijun, et al.
Published: (2023)
Horizon Reduction as Information Loss in Offline Reinforcement Learning
by: Nidadala, Uday Kumar, et al.
Published: (2025)
by: Nidadala, Uday Kumar, et al.
Published: (2025)
Cost-Effective Cyber-Physical System Prototype for Precision Agriculture with a Focus on Crop Growth
by: Kumar, Pawan, et al.
Published: (2024)
by: Kumar, Pawan, et al.
Published: (2024)
Similar Items
-
Towards More Unified In-context Visual Understanding
by: Sheng, Dianmo, et al.
Published: (2023) -
TCI-Former: Thermal Conduction-Inspired Transformer for Infrared Small Target Detection
by: Chen, Tianxiang, et al.
Published: (2024) -
Flora: Effortless Context Construction to Arbitrary Length and Scale
by: Chen, Tianxiang, et al.
Published: (2025) -
Llama SLayer 8B: Shallow Layers Hold the Key to Knowledge Injection
by: Chen, Tianxiang, et al.
Published: (2024) -
SAPL: Semantic-Agnostic Prompt Learning in CLIP for Weakly Supervised Image Manipulation Localization
by: Wang, Xinghao, et al.
Published: (2026)