Empowering Reliable Visual-Centric Instruction Following in MLLMs
Fuente:
arXiv
Salvato in:
| Autori principali: | He, Weilei, Ju, Feng, Fan, Zhiyuan, Min, Rui, Cheng, Minhao, Fung, Yi R. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Scalable Token-Level Hallucination Detection in Large Language Models
di: Min, Rui, et al.
Pubblicazione: (2026)
di: Min, Rui, et al.
Pubblicazione: (2026)
TextSquare: Scaling up Text-Centric Visual Instruction Tuning
di: Tang, Jingqun, et al.
Pubblicazione: (2024)
di: Tang, Jingqun, et al.
Pubblicazione: (2024)
Evaluating Large Language Models at Evaluating Instruction Following
di: Zeng, Zhiyuan, et al.
Pubblicazione: (2023)
di: Zeng, Zhiyuan, et al.
Pubblicazione: (2023)
CREAM: Consistency Regularized Self-Rewarding Language Models
di: Wang, Zhaoyang, et al.
Pubblicazione: (2024)
di: Wang, Zhaoyang, et al.
Pubblicazione: (2024)
Sparsity Forcing: Reinforcing Token Sparsity of MLLMs
di: Chen, Feng, et al.
Pubblicazione: (2025)
di: Chen, Feng, et al.
Pubblicazione: (2025)
Hierarchical Multi-Label Generation with Probabilistic Level-Constraint
di: Chen, Linqing, et al.
Pubblicazione: (2025)
di: Chen, Linqing, et al.
Pubblicazione: (2025)
MoDA: Modulation Adapter for Fine-Grained Visual Grounding in Instructional MLLMs
di: Barrios, Wayner, et al.
Pubblicazione: (2025)
di: Barrios, Wayner, et al.
Pubblicazione: (2025)
Uncovering, Explaining, and Mitigating the Superficial Safety of Backdoor Defense
di: Min, Rui, et al.
Pubblicazione: (2024)
di: Min, Rui, et al.
Pubblicazione: (2024)
Learning Interactive World Model for Object-Centric Reinforcement Learning
di: Feng, Fan, et al.
Pubblicazione: (2025)
di: Feng, Fan, et al.
Pubblicazione: (2025)
Improving Your Model Ranking on Chatbot Arena by Vote Rigging
di: Min, Rui, et al.
Pubblicazione: (2025)
di: Min, Rui, et al.
Pubblicazione: (2025)
Long-horizon Visual Instruction Generation with Logic and Attribute Self-reflection
di: Suo, Yucheng, et al.
Pubblicazione: (2025)
di: Suo, Yucheng, et al.
Pubblicazione: (2025)
M2-Reasoning: Empowering MLLMs with Unified General and Spatial Reasoning
di: AI, Inclusion, et al.
Pubblicazione: (2025)
di: AI, Inclusion, et al.
Pubblicazione: (2025)
Goal Representations for Instruction Following: A Semi-Supervised Language Interface to Control
di: Myers, Vivek, et al.
Pubblicazione: (2023)
di: Myers, Vivek, et al.
Pubblicazione: (2023)
Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs
di: Zhang, Xin, et al.
Pubblicazione: (2026)
di: Zhang, Xin, et al.
Pubblicazione: (2026)
Community-Centric Graph Unlearning
di: Li, Yi, et al.
Pubblicazione: (2024)
di: Li, Yi, et al.
Pubblicazione: (2024)
SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language Models
di: He, Zirui, et al.
Pubblicazione: (2025)
di: He, Zirui, et al.
Pubblicazione: (2025)
Task-Centric Policy Optimization from Misaligned Motion Priors
di: Zheng, Ziang, et al.
Pubblicazione: (2026)
di: Zheng, Ziang, et al.
Pubblicazione: (2026)
Improving Instruction Following in Language Models through Proxy-Based Uncertainty Estimation
di: Lee, JoonHo, et al.
Pubblicazione: (2024)
di: Lee, JoonHo, et al.
Pubblicazione: (2024)
Growing Visual Generative Capacity for Pre-Trained MLLMs
di: Wang, Hanyu, et al.
Pubblicazione: (2025)
di: Wang, Hanyu, et al.
Pubblicazione: (2025)
Input Snapshots Fusion for Scalable Discrete-Time Dynamic Graph Neural Networks
di: Qi, QingGuo, et al.
Pubblicazione: (2024)
di: Qi, QingGuo, et al.
Pubblicazione: (2024)
Environment Scaling for Interactive Agentic Experience Collection: A Survey
di: Huang, Yuchen, et al.
Pubblicazione: (2025)
di: Huang, Yuchen, et al.
Pubblicazione: (2025)
Layer-Aware Task Arithmetic: Disentangling Task-Specific and Instruction-Following Knowledge
di: Chen, Yan-Lun, et al.
Pubblicazione: (2025)
di: Chen, Yan-Lun, et al.
Pubblicazione: (2025)
Multimodal Prompt Optimization: Why Not Leverage Multiple Modalities for MLLMs
di: Choi, Yumin, et al.
Pubblicazione: (2025)
di: Choi, Yumin, et al.
Pubblicazione: (2025)
Agglomerative Federated Learning: Empowering Larger Model Training via End-Edge-Cloud Collaboration
di: Wu, Zhiyuan, et al.
Pubblicazione: (2023)
di: Wu, Zhiyuan, et al.
Pubblicazione: (2023)
SimpleOCR: Rendering Visualized Questions to Teach MLLMs to Read
di: Peng, Yibo, et al.
Pubblicazione: (2026)
di: Peng, Yibo, et al.
Pubblicazione: (2026)
CLIPErase: Efficient Unlearning of Visual-Textual Associations in CLIP
di: Yang, Tianyu, et al.
Pubblicazione: (2024)
di: Yang, Tianyu, et al.
Pubblicazione: (2024)
Chain of Thought Empowers Transformers to Solve Inherently Serial Problems
di: Li, Zhiyuan, et al.
Pubblicazione: (2024)
di: Li, Zhiyuan, et al.
Pubblicazione: (2024)
Off-Policy Selection for Initiating Human-Centric Experimental Design
di: Gao, Ge, et al.
Pubblicazione: (2024)
di: Gao, Ge, et al.
Pubblicazione: (2024)
VideoScaffold: Elastic-Scale Visual Hierarchies for Streaming Video Understanding in MLLMs
di: Zheng, Naishan, et al.
Pubblicazione: (2025)
di: Zheng, Naishan, et al.
Pubblicazione: (2025)
Enhancing and Assessing Instruction-Following with Fine-Grained Instruction Variants
di: Yang, Jiuding, et al.
Pubblicazione: (2024)
di: Yang, Jiuding, et al.
Pubblicazione: (2024)
Learning Variable-Length Tokenization for Generative Recommendation
di: Wang, Minhao, et al.
Pubblicazione: (2026)
di: Wang, Minhao, et al.
Pubblicazione: (2026)
Achieving Constant Regret in Linear Markov Decision Processes
di: Zhang, Weitong, et al.
Pubblicazione: (2024)
di: Zhang, Weitong, et al.
Pubblicazione: (2024)
Beyond Visual Realism: Toward Reliable Financial Time Series Generation
di: Zhang, Fan, et al.
Pubblicazione: (2026)
di: Zhang, Fan, et al.
Pubblicazione: (2026)
Unveiling Instruction-Specific Neurons & Experts: An Analytical Framework for LLM's Instruction-Following Capabilities
di: Zhang, Junyan, et al.
Pubblicazione: (2025)
di: Zhang, Junyan, et al.
Pubblicazione: (2025)
RAR: Retrieving And Ranking Augmented MLLMs for Visual Recognition
di: Liu, Ziyu, et al.
Pubblicazione: (2024)
di: Liu, Ziyu, et al.
Pubblicazione: (2024)
From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities
di: Jiang, Shixin, et al.
Pubblicazione: (2024)
di: Jiang, Shixin, et al.
Pubblicazione: (2024)
Financial Instruction Following Evaluation (FIFE)
di: Matlin, Glenn, et al.
Pubblicazione: (2025)
di: Matlin, Glenn, et al.
Pubblicazione: (2025)
Empower Low-Altitude Economy: A Reliability-Aware Dynamic Weighting Allocation for Multi-modal UAV Beam Prediction
di: Li, Haojin, et al.
Pubblicazione: (2025)
di: Li, Haojin, et al.
Pubblicazione: (2025)
Precision over Diversity: High-Precision Reward Generalizes to Robust Instruction Following
di: Zeng, Yirong, et al.
Pubblicazione: (2026)
di: Zeng, Yirong, et al.
Pubblicazione: (2026)
Codebook-Centric Deep Hashing: End-to-End Joint Learning of Semantic Hash Centers and Neural Hash Function
di: Yin, Shuo, et al.
Pubblicazione: (2025)
di: Yin, Shuo, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Scalable Token-Level Hallucination Detection in Large Language Models
di: Min, Rui, et al.
Pubblicazione: (2026) -
TextSquare: Scaling up Text-Centric Visual Instruction Tuning
di: Tang, Jingqun, et al.
Pubblicazione: (2024) -
Evaluating Large Language Models at Evaluating Instruction Following
di: Zeng, Zhiyuan, et al.
Pubblicazione: (2023) -
CREAM: Consistency Regularized Self-Rewarding Language Models
di: Wang, Zhaoyang, et al.
Pubblicazione: (2024) -
Sparsity Forcing: Reinforcing Token Sparsity of MLLMs
di: Chen, Feng, et al.
Pubblicazione: (2025)