Automated Multi-level Preference for MLLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Mengxi, Wu, Wenhao, Lu, Yu, Song, Yuxin, Rong, Kang, Yao, Huanjin, Zhao, Jianbo, Liu, Fanglong, Sun, Yifan, Feng, Haocheng, Wang, Jingdong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Dense Connector for MLLMs
von: Yao, Huanjin, et al.
Veröffentlicht: (2024)
von: Yao, Huanjin, et al.
Veröffentlicht: (2024)
GPT4Vis: What Can GPT-4 Do for Zero-shot Visual Recognition?
von: Wu, Wenhao, et al.
Veröffentlicht: (2023)
von: Wu, Wenhao, et al.
Veröffentlicht: (2023)
CoLoGen: Progressive Learning of Concept-Localization Duality for Unified Image Generation
von: Song, YuXin, et al.
Veröffentlicht: (2026)
von: Song, YuXin, et al.
Veröffentlicht: (2026)
MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI
von: Yao, Huanjin, et al.
Veröffentlicht: (2025)
von: Yao, Huanjin, et al.
Veröffentlicht: (2025)
MonoFormer: One Transformer for Both Diffusion and Autoregression
von: Zhao, Chuyang, et al.
Veröffentlicht: (2024)
von: Zhao, Chuyang, et al.
Veröffentlicht: (2024)
Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search
von: Yao, Huanjin, et al.
Veröffentlicht: (2024)
von: Yao, Huanjin, et al.
Veröffentlicht: (2024)
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders
von: Fang, Bo, et al.
Veröffentlicht: (2025)
von: Fang, Bo, et al.
Veröffentlicht: (2025)
Instruction-Oriented Preference Alignment for Enhancing Multi-Modal Comprehension Capability of MLLMs
von: Wang, Zitian, et al.
Veröffentlicht: (2025)
von: Wang, Zitian, et al.
Veröffentlicht: (2025)
FullAnno: A Data Engine for Enhancing Image Comprehension of MLLMs
von: Hao, Jing, et al.
Veröffentlicht: (2024)
von: Hao, Jing, et al.
Veröffentlicht: (2024)
Spatial Preference Rewarding for MLLMs Spatial Understanding
von: Qiu, Han, et al.
Veröffentlicht: (2025)
von: Qiu, Han, et al.
Veröffentlicht: (2025)
SFTformer: A Spatial-Frequency-Temporal Correlation-Decoupling Transformer for Radar Echo Extrapolation
von: Xu, Liangyu, et al.
Veröffentlicht: (2024)
von: Xu, Liangyu, et al.
Veröffentlicht: (2024)
Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation
von: Chen, Zeyu, et al.
Veröffentlicht: (2026)
von: Chen, Zeyu, et al.
Veröffentlicht: (2026)
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities
von: Liu, Huan, et al.
Veröffentlicht: (2024)
von: Liu, Huan, et al.
Veröffentlicht: (2024)
NavAgent: Multi-scale Urban Street View Fusion For UAV Embodied Vision-and-Language Navigation
von: Liu, Youzhi, et al.
Veröffentlicht: (2024)
von: Liu, Youzhi, et al.
Veröffentlicht: (2024)
Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs
von: Li, Xudong, et al.
Veröffentlicht: (2025)
von: Li, Xudong, et al.
Veröffentlicht: (2025)
RefAlign: Representation Alignment for Reference-to-Video Generation
von: Wang, Lei, et al.
Veröffentlicht: (2026)
von: Wang, Lei, et al.
Veröffentlicht: (2026)
MS-DETR: Efficient DETR Training with Mixed Supervision
von: Zhao, Chuyang, et al.
Veröffentlicht: (2024)
von: Zhao, Chuyang, et al.
Veröffentlicht: (2024)
On the Generalization Capacities of MLLMs for Spatial Intelligence
von: Zhang, Gongjie, et al.
Veröffentlicht: (2026)
von: Zhang, Gongjie, et al.
Veröffentlicht: (2026)
Law of Vision Representation in MLLMs
von: Yang, Shijia, et al.
Veröffentlicht: (2024)
von: Yang, Shijia, et al.
Veröffentlicht: (2024)
Query-Kontext: An Unified Multimodal Model for Image Generation and Editing
von: Song, Yuxin, et al.
Veröffentlicht: (2025)
von: Song, Yuxin, et al.
Veröffentlicht: (2025)
Exploring the Design Space of Visual Context Representation in Video MLLMs
von: Du, Yifan, et al.
Veröffentlicht: (2024)
von: Du, Yifan, et al.
Veröffentlicht: (2024)
SAMA: Factorized Semantic Anchoring and Motion Alignment for Instruction-Guided Video Editing
von: Zhang, Xinyao, et al.
Veröffentlicht: (2026)
von: Zhang, Xinyao, et al.
Veröffentlicht: (2026)
GVA: Reconstructing Vivid 3D Gaussian Avatars from Monocular Videos
von: Liu, Xinqi, et al.
Veröffentlicht: (2024)
von: Liu, Xinqi, et al.
Veröffentlicht: (2024)
XLD: A Cross-Lane Dataset for Benchmarking Novel Driving View Synthesis
von: Li, Hao, et al.
Veröffentlicht: (2024)
von: Li, Hao, et al.
Veröffentlicht: (2024)
Towards Camera-Robust 3D Localization: Equation-Anchored Tool-Use for MLLMs
von: Jiang, Xueying, et al.
Veröffentlicht: (2026)
von: Jiang, Xueying, et al.
Veröffentlicht: (2026)
Explore the Hallucination on Low-level Perception for MLLMs
von: Sun, Yinan, et al.
Veröffentlicht: (2024)
von: Sun, Yinan, et al.
Veröffentlicht: (2024)
Tell Model Where to Look: Mitigating Hallucinations in MLLMs by Vision-Guided Attention
von: Zhao, Jianfei, et al.
Veröffentlicht: (2025)
von: Zhao, Jianfei, et al.
Veröffentlicht: (2025)
X-Stream: Exploring MLLMs as Multiplexers for Multi-Stream Understanding
von: Sun, Peiwen, et al.
Veröffentlicht: (2026)
von: Sun, Peiwen, et al.
Veröffentlicht: (2026)
OmniAlign-V: Towards Enhanced Alignment of MLLMs with Human Preference
von: Zhao, Xiangyu, et al.
Veröffentlicht: (2025)
von: Zhao, Xiangyu, et al.
Veröffentlicht: (2025)
Re-HOLD: Video Hand Object Interaction Reenactment via adaptive Layout-instructed Diffusion Model
von: Fan, Yingying, et al.
Veröffentlicht: (2025)
von: Fan, Yingying, et al.
Veröffentlicht: (2025)
TexRO: Generating Delicate Textures of 3D Models by Recursive Optimization
von: Wu, Jinbo, et al.
Veröffentlicht: (2024)
von: Wu, Jinbo, et al.
Veröffentlicht: (2024)
GIR: 3D Gaussian Inverse Rendering for Relightable Scene Factorization
von: Shi, Yahao, et al.
Veröffentlicht: (2023)
von: Shi, Yahao, et al.
Veröffentlicht: (2023)
Global-Local Dual Perception for MLLMs in High-Resolution Text-Rich Image Translation
von: Lu, Junxin, et al.
Veröffentlicht: (2026)
von: Lu, Junxin, et al.
Veröffentlicht: (2026)
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark
von: Cheng, Ziming, et al.
Veröffentlicht: (2025)
von: Cheng, Ziming, et al.
Veröffentlicht: (2025)
IDPruner: Harmonizing Importance and Diversity in Visual Token Pruning for MLLMs
von: Tan, Yifan, et al.
Veröffentlicht: (2026)
von: Tan, Yifan, et al.
Veröffentlicht: (2026)
ViSS-R1: Self-Supervised Reinforcement Video Reasoning
von: Fang, Bo, et al.
Veröffentlicht: (2025)
von: Fang, Bo, et al.
Veröffentlicht: (2025)
Assessing Model Generalization in Vicinity
von: Liu, Yuchi, et al.
Veröffentlicht: (2024)
von: Liu, Yuchi, et al.
Veröffentlicht: (2024)
HumanSense: From Multimodal Perception to Empathetic Context-Aware Responses through Reasoning MLLMs
von: Qin, Zheng, et al.
Veröffentlicht: (2025)
von: Qin, Zheng, et al.
Veröffentlicht: (2025)
GGRt: Towards Pose-free Generalizable 3D Gaussian Splatting in Real-time
von: Li, Hao, et al.
Veröffentlicht: (2024)
von: Li, Hao, et al.
Veröffentlicht: (2024)
VDG: Vision-Only Dynamic Gaussian for Driving Simulation
von: Li, Hao, et al.
Veröffentlicht: (2024)
von: Li, Hao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Dense Connector for MLLMs
von: Yao, Huanjin, et al.
Veröffentlicht: (2024) -
GPT4Vis: What Can GPT-4 Do for Zero-shot Visual Recognition?
von: Wu, Wenhao, et al.
Veröffentlicht: (2023) -
CoLoGen: Progressive Learning of Concept-Localization Duality for Unified Image Generation
von: Song, YuXin, et al.
Veröffentlicht: (2026) -
MMReason: An Open-Ended Multi-Modal Multi-Step Reasoning Benchmark for MLLMs Toward AGI
von: Yao, Huanjin, et al.
Veröffentlicht: (2025) -
MonoFormer: One Transformer for Both Diffusion and Autoregression
von: Zhao, Chuyang, et al.
Veröffentlicht: (2024)