DAMA: Data- and Model-aware Alignment of Multi-modal LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lu, Jinda, Wu, Junkang, Li, Jinghan, Jia, Xiaojun, Wang, Shuo, Zhang, YiFan, Fang, Junfeng, Wang, Xiang, He, Xiangnan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization
von: Lu, Jinda, et al.
Veröffentlicht: (2025)
von: Lu, Jinda, et al.
Veröffentlicht: (2025)
Bridging Perception and Reasoning: Token Reweighting for RLVR in Multimodal LLMs
von: Lu, Jinda, et al.
Veröffentlicht: (2026)
von: Lu, Jinda, et al.
Veröffentlicht: (2026)
Enhancing Multi-Modal LLMs Reasoning via Difficulty-Aware Group Normalization
von: Li, Jinghan, et al.
Veröffentlicht: (2026)
von: Li, Jinghan, et al.
Veröffentlicht: (2026)
Beyond Where to Look: Trajectory-Guided Reinforcement Learning for Multimodal RLVR
von: Lu, Jinda, et al.
Veröffentlicht: (2026)
von: Lu, Jinda, et al.
Veröffentlicht: (2026)
GuardAlign: Test-time Safety Alignment in Multimodal Large Language Models
von: Zhu, Xingyu, et al.
Veröffentlicht: (2026)
von: Zhu, Xingyu, et al.
Veröffentlicht: (2026)
Boosting Few-Shot Learning via Attentive Feature Regularization
von: Zhu, Xingyu, et al.
Veröffentlicht: (2024)
von: Zhu, Xingyu, et al.
Veröffentlicht: (2024)
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs
von: Jiang, Houcheng, et al.
Veröffentlicht: (2026)
von: Jiang, Houcheng, et al.
Veröffentlicht: (2026)
Mitigating Hallucinations in Large Vision-Language Models without Performance Degradation
von: Zhu, Xingyu, et al.
Veröffentlicht: (2026)
von: Zhu, Xingyu, et al.
Veröffentlicht: (2026)
Principled Steering via Null-space Projection for Jailbreak Defense in Vision-Language Models
von: Zhu, Xingyu, et al.
Veröffentlicht: (2026)
von: Zhu, Xingyu, et al.
Veröffentlicht: (2026)
Cracking the Code of Hallucination in LVLMs with Vision-aware Head Divergence
von: He, Jinghan, et al.
Veröffentlicht: (2024)
von: He, Jinghan, et al.
Veröffentlicht: (2024)
Rethinking Visual Content Refinement in Low-Shot CLIP Adaptation
von: Lu, Jinda, et al.
Veröffentlicht: (2024)
von: Lu, Jinda, et al.
Veröffentlicht: (2024)
Causal-HalBench: Uncovering LVLMs Object Hallucinations Through Causal Intervention
von: Xu, Zhe, et al.
Veröffentlicht: (2025)
von: Xu, Zhe, et al.
Veröffentlicht: (2025)
LLM-enhanced Action-aware Multi-modal Prompt Tuning for Image-Text Matching
von: Tian, Mengxiao, et al.
Veröffentlicht: (2025)
von: Tian, Mengxiao, et al.
Veröffentlicht: (2025)
Multi-level Cross-modal Alignment for Image Clustering
von: Qiu, Liping, et al.
Veröffentlicht: (2024)
von: Qiu, Liping, et al.
Veröffentlicht: (2024)
DiffGAD: A Diffusion-based Unsupervised Graph Anomaly Detector
von: Li, Jinghan, et al.
Veröffentlicht: (2024)
von: Li, Jinghan, et al.
Veröffentlicht: (2024)
Multi-modal Instruction Tuned LLMs with Fine-grained Visual Perception
von: He, Junwen, et al.
Veröffentlicht: (2024)
von: He, Junwen, et al.
Veröffentlicht: (2024)
R^2-Mem: Reflective Experience for Memory Search
von: Wang, Xinyuan, et al.
Veröffentlicht: (2026)
von: Wang, Xinyuan, et al.
Veröffentlicht: (2026)
Unified Parameter-Efficient Unlearning for LLMs
von: Ding, Chenlu, et al.
Veröffentlicht: (2024)
von: Ding, Chenlu, et al.
Veröffentlicht: (2024)
Liquid: Language Models are Scalable and Unified Multi-modal Generators
von: Wu, Junfeng, et al.
Veröffentlicht: (2024)
von: Wu, Junfeng, et al.
Veröffentlicht: (2024)
InstructEngine: Instruction-driven Text-to-Image Alignment
von: Lu, Xingyu, et al.
Veröffentlicht: (2025)
von: Lu, Xingyu, et al.
Veröffentlicht: (2025)
Enhancing Tree Type Detection in Forest Fire Risk Assessment: Multi-Stage Approach and Color Encoding with Forest Fire Risk Evaluation Framework for UAV Imagery
von: Zhang, Jinda
Veröffentlicht: (2024)
von: Zhang, Jinda
Veröffentlicht: (2024)
DMPT: Decoupled Modality-aware Prompt Tuning for Multi-modal Object Re-identification
von: Lin, Minghui, et al.
Veröffentlicht: (2025)
von: Lin, Minghui, et al.
Veröffentlicht: (2025)
3D-aware Image Generation and Editing with Multi-modal Conditions
von: Li, Bo, et al.
Veröffentlicht: (2024)
von: Li, Bo, et al.
Veröffentlicht: (2024)
MOCHA: Multi-modal Objects-aware Cross-arcHitecture Alignment
von: Camuffo, Elena, et al.
Veröffentlicht: (2025)
von: Camuffo, Elena, et al.
Veröffentlicht: (2025)
bi-GRPO: Bidirectional Optimization for Jailbreak Backdoor Injection on LLMs
von: Ji, Wence, et al.
Veröffentlicht: (2025)
von: Ji, Wence, et al.
Veröffentlicht: (2025)
Quantile Advantage Estimation: Stabilizing RLVR for LLM Reasoning
von: Wu, Junkang, et al.
Veröffentlicht: (2025)
von: Wu, Junkang, et al.
Veröffentlicht: (2025)
MCAD: Multi-teacher Cross-modal Alignment Distillation for efficient image-text retrieval
von: Lei, Youbo, et al.
Veröffentlicht: (2023)
von: Lei, Youbo, et al.
Veröffentlicht: (2023)
Hierarchical Semantic Alignment for Image Clustering
von: Zhu, Xingyu, et al.
Veröffentlicht: (2025)
von: Zhu, Xingyu, et al.
Veröffentlicht: (2025)
Learning Procedural-aware Video Representations through State-Grounded Hierarchy Unfolding
von: Zhao, Jinghan, et al.
Veröffentlicht: (2025)
von: Zhao, Jinghan, et al.
Veröffentlicht: (2025)
EIMC: Efficient Instance-aware Multi-modal Collaborative Perception
von: Yang, Kang, et al.
Veröffentlicht: (2026)
von: Yang, Kang, et al.
Veröffentlicht: (2026)
Larger or Smaller Reward Margins to Select Preferences for Alignment?
von: Huang, Kexin, et al.
Veröffentlicht: (2025)
von: Huang, Kexin, et al.
Veröffentlicht: (2025)
Multi-modal Generative AI: Multi-modal LLMs, Diffusions, and the Unification
von: Wang, Xin, et al.
Veröffentlicht: (2024)
von: Wang, Xin, et al.
Veröffentlicht: (2024)
Accelerating Diffusion Transformer via Gradient-Optimized Cache
von: Qiu, Junxiang, et al.
Veröffentlicht: (2025)
von: Qiu, Junxiang, et al.
Veröffentlicht: (2025)
Multi-modal Semantic Understanding with Contrastive Cross-modal Feature Alignment
von: Zhang, Ming, et al.
Veröffentlicht: (2024)
von: Zhang, Ming, et al.
Veröffentlicht: (2024)
Reliable Cross-modal Alignment via Prototype Iterative Construction
von: Ma, Xiang, et al.
Veröffentlicht: (2025)
von: Ma, Xiang, et al.
Veröffentlicht: (2025)
SeViCES: Unifying Semantic-Visual Evidence Consensus for Long Video Understanding
von: Sheng, Yuan, et al.
Veröffentlicht: (2025)
von: Sheng, Yuan, et al.
Veröffentlicht: (2025)
End-to-end Open-vocabulary Video Visual Relationship Detection using Multi-modal Prompting
von: Wang, Yongqi, et al.
Veröffentlicht: (2024)
von: Wang, Yongqi, et al.
Veröffentlicht: (2024)
On the Multi-modal Vulnerability of Diffusion Models
von: Yang, Dingcheng, et al.
Veröffentlicht: (2024)
von: Yang, Dingcheng, et al.
Veröffentlicht: (2024)
Do We Really Need Curated Malicious Data for Safety Alignment in Multi-modal Large Language Models?
von: Wang, Yanbo, et al.
Veröffentlicht: (2025)
von: Wang, Yanbo, et al.
Veröffentlicht: (2025)
Steering and Rectifying Latent Representation Manifolds in Frozen Multi-modal LLMs for Video Anomaly Detection
von: Cai, Zhaolin, et al.
Veröffentlicht: (2026)
von: Cai, Zhaolin, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization
von: Lu, Jinda, et al.
Veröffentlicht: (2025) -
Bridging Perception and Reasoning: Token Reweighting for RLVR in Multimodal LLMs
von: Lu, Jinda, et al.
Veröffentlicht: (2026) -
Enhancing Multi-Modal LLMs Reasoning via Difficulty-Aware Group Normalization
von: Li, Jinghan, et al.
Veröffentlicht: (2026) -
Beyond Where to Look: Trajectory-Guided Reinforcement Learning for Multimodal RLVR
von: Lu, Jinda, et al.
Veröffentlicht: (2026) -
GuardAlign: Test-time Safety Alignment in Multimodal Large Language Models
von: Zhu, Xingyu, et al.
Veröffentlicht: (2026)