Through the Lens of Character: Resolving Modality-Role Interference in Multimodal Role-Playing Agent
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tang, Yihong, Chen, Kehai, Bai, Xuefeng, Zhang, Min |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Character-R1: Enhancing Role-Aware Reasoning in Role-Playing Agents via RLVR
von: Tang, Yihong, et al.
Veröffentlicht: (2026)
von: Tang, Yihong, et al.
Veröffentlicht: (2026)
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey
von: Qu, Bingzheng, et al.
Veröffentlicht: (2026)
von: Qu, Bingzheng, et al.
Veröffentlicht: (2026)
Mitigating Multimodal Hallucination via Phase-wise Self-reward
von: Zhang, Yu, et al.
Veröffentlicht: (2026)
von: Zhang, Yu, et al.
Veröffentlicht: (2026)
OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality Interaction
von: Zhang, Haonan, et al.
Veröffentlicht: (2025)
von: Zhang, Haonan, et al.
Veröffentlicht: (2025)
Thinking in Character: Advancing Role-Playing Agents with Role-Aware Reasoning
von: Tang, Yihong, et al.
Veröffentlicht: (2025)
von: Tang, Yihong, et al.
Veröffentlicht: (2025)
Benchmarking and Improving Large Vision-Language Models for Fundamental Visual Graph Understanding and Reasoning
von: Zhu, Yingjie, et al.
Veröffentlicht: (2024)
von: Zhu, Yingjie, et al.
Veröffentlicht: (2024)
The Rise of Darkness: Safety-Utility Trade-Offs in Role-Playing Dialogue Agents
von: Tang, Yihong, et al.
Veröffentlicht: (2025)
von: Tang, Yihong, et al.
Veröffentlicht: (2025)
CRPO: Character-centric Group Relative Policy Optimization for Role-aware Reasoning in Role-playing Agents
von: Tang, Yihong, et al.
Veröffentlicht: (2026)
von: Tang, Yihong, et al.
Veröffentlicht: (2026)
Beyond Rigid: Benchmarking Non-Rigid Video Editing
von: Qu, Bingzheng, et al.
Veröffentlicht: (2026)
von: Qu, Bingzheng, et al.
Veröffentlicht: (2026)
Beyond the Vision Encoder: Identifying and Mitigating Spatial Bias in Large Vision-Language Models
von: Zhu, Yingjie, et al.
Veröffentlicht: (2025)
von: Zhu, Yingjie, et al.
Veröffentlicht: (2025)
Play to Generalize: Learning to Reason Through Game Play
von: Xie, Yunfei, et al.
Veröffentlicht: (2025)
von: Xie, Yunfei, et al.
Veröffentlicht: (2025)
CLEAR: Character Unlearning in Textual and Visual Modalities
von: Dontsov, Alexey, et al.
Veröffentlicht: (2024)
von: Dontsov, Alexey, et al.
Veröffentlicht: (2024)
From Perception to Reasoning: Deep Thinking Empowers Multimodal Large Language Models
von: Zhu, Wenxin, et al.
Veröffentlicht: (2025)
von: Zhu, Wenxin, et al.
Veröffentlicht: (2025)
Video2Roleplay: A Multimodal Dataset and Framework for Video-Guided Role-playing Agents
von: Zhang, Xueqiao, et al.
Veröffentlicht: (2025)
von: Zhang, Xueqiao, et al.
Veröffentlicht: (2025)
Modality Bias in LVLMs: Analyzing and Mitigating Object Hallucination via Attention Lens
von: Zheng, Haohan, et al.
Veröffentlicht: (2025)
von: Zheng, Haohan, et al.
Veröffentlicht: (2025)
Reading, Not Thinking: Understanding and Bridging the Modality Gap When Text Becomes Pixels in Multimodal LLMs
von: Sun, Kaiser, et al.
Veröffentlicht: (2026)
von: Sun, Kaiser, et al.
Veröffentlicht: (2026)
PresentAgent-2: Towards Generalist Multimodal Presentation Agents
von: Wu, Wei, et al.
Veröffentlicht: (2026)
von: Wu, Wei, et al.
Veröffentlicht: (2026)
Plug-and-Play Grounding of Reasoning in Multimodal Large Language Models
von: Chen, Jiaxing, et al.
Veröffentlicht: (2024)
von: Chen, Jiaxing, et al.
Veröffentlicht: (2024)
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding
von: Li, Yun, et al.
Veröffentlicht: (2025)
von: Li, Yun, et al.
Veröffentlicht: (2025)
WorldMemArena: Evaluating Multimodal Agent Memory Through Action-World Interaction
von: Liu, Chengzhi, et al.
Veröffentlicht: (2026)
von: Liu, Chengzhi, et al.
Veröffentlicht: (2026)
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
von: Zhang, Yi-Fan, et al.
Veröffentlicht: (2025)
von: Zhang, Yi-Fan, et al.
Veröffentlicht: (2025)
Evaluating Linguistic Capabilities of Multimodal LLMs in the Lens of Few-Shot Learning
von: Dogan, Mustafa, et al.
Veröffentlicht: (2024)
von: Dogan, Mustafa, et al.
Veröffentlicht: (2024)
Language Plays a Pivotal Role in the Object-Attribute Compositional Generalization of CLIP
von: Abbasi, Reza, et al.
Veröffentlicht: (2024)
von: Abbasi, Reza, et al.
Veröffentlicht: (2024)
Leveraging Entity Information for Cross-Modality Correlation Learning: The Entity-Guided Multimodal Summarization
von: Zhang, Yanghai, et al.
Veröffentlicht: (2024)
von: Zhang, Yanghai, et al.
Veröffentlicht: (2024)
Spatial Knowledge Graph-Guided Multimodal Synthesis
von: Xue, Yida, et al.
Veröffentlicht: (2025)
von: Xue, Yida, et al.
Veröffentlicht: (2025)
DocLens : A Tool-Augmented Multi-Agent Framework for Long Visual Document Understanding
von: Zhu, Dawei, et al.
Veröffentlicht: (2025)
von: Zhu, Dawei, et al.
Veröffentlicht: (2025)
Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception
von: Wang, Junyang, et al.
Veröffentlicht: (2024)
von: Wang, Junyang, et al.
Veröffentlicht: (2024)
Large Vision-Language Model Alignment and Misalignment: A Survey Through the Lens of Explainability
von: Shu, Dong, et al.
Veröffentlicht: (2025)
von: Shu, Dong, et al.
Veröffentlicht: (2025)
Evaluating and Advancing Multimodal Large Language Models in Perception Ability Lens
von: Chen, Feng, et al.
Veröffentlicht: (2024)
von: Chen, Feng, et al.
Veröffentlicht: (2024)
MC-MKE: A Fine-Grained Multimodal Knowledge Editing Benchmark Emphasizing Modality Consistency
von: Zhang, Junzhe, et al.
Veröffentlicht: (2024)
von: Zhang, Junzhe, et al.
Veröffentlicht: (2024)
Is Extending Modality The Right Path Towards Omni-Modality?
von: Zhu, Tinghui, et al.
Veröffentlicht: (2025)
von: Zhu, Tinghui, et al.
Veröffentlicht: (2025)
Ancient but Digitized: Developing Handwritten Optical Character Recognition for East Syriac Script Through Creating KHAMIS Dataset
von: Majeed, Ameer, et al.
Veröffentlicht: (2024)
von: Majeed, Ameer, et al.
Veröffentlicht: (2024)
The Power of Personality: A Human Simulation Perspective to Investigate Large Language Model Agents
von: Duan, Yifan, et al.
Veröffentlicht: (2025)
von: Duan, Yifan, et al.
Veröffentlicht: (2025)
Text Role Classification in Scientific Charts Using Multimodal Transformers
von: Kim, Hye Jin, et al.
Veröffentlicht: (2024)
von: Kim, Hye Jin, et al.
Veröffentlicht: (2024)
VaseVQA: Multimodal Agent and Benchmark for Ancient Greek Pottery
von: Ge, Jinchao, et al.
Veröffentlicht: (2025)
von: Ge, Jinchao, et al.
Veröffentlicht: (2025)
Culture In a Frame: C$^3$B as a Comic-Based Benchmark for Multimodal Culturally Awareness
von: Song, Yuchen, et al.
Veröffentlicht: (2025)
von: Song, Yuchen, et al.
Veröffentlicht: (2025)
Multimodal Prompt Learning with Missing Modalities for Sentiment Analysis and Emotion Recognition
von: Guo, Zirun, et al.
Veröffentlicht: (2024)
von: Guo, Zirun, et al.
Veröffentlicht: (2024)
MANTA: Cross-Modal Semantic Alignment and Information-Theoretic Optimization for Long-form Multimodal Understanding
von: Zhong, Ziqi, et al.
Veröffentlicht: (2025)
von: Zhong, Ziqi, et al.
Veröffentlicht: (2025)
Qalam : A Multimodal LLM for Arabic Optical Character and Handwriting Recognition
von: Bhatia, Gagan, et al.
Veröffentlicht: (2024)
von: Bhatia, Gagan, et al.
Veröffentlicht: (2024)
EmbodiedEval: Evaluate Multimodal LLMs as Embodied Agents
von: Cheng, Zhili, et al.
Veröffentlicht: (2025)
von: Cheng, Zhili, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Character-R1: Enhancing Role-Aware Reasoning in Role-Playing Agents via RLVR
von: Tang, Yihong, et al.
Veröffentlicht: (2026) -
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey
von: Qu, Bingzheng, et al.
Veröffentlicht: (2026) -
Mitigating Multimodal Hallucination via Phase-wise Self-reward
von: Zhang, Yu, et al.
Veröffentlicht: (2026) -
OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality Interaction
von: Zhang, Haonan, et al.
Veröffentlicht: (2025) -
Thinking in Character: Advancing Role-Playing Agents with Role-Aware Reasoning
von: Tang, Yihong, et al.
Veröffentlicht: (2025)