Cross-Modal Coreference Alignment: Enabling Reliable Information Transfer in Omni-LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Hongcheng, Wang, Yuhao, Chen, Zhe, Wang, Pingjie, Zhu, Zhiyuan, Hou, Yixuan, Wang, Yanfeng, Wang, Yu |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
M2K-VDG: Model-Adaptive Multimodal Knowledge Anchor Enhanced Video-grounded Dialogue Generation
by: Liu, Hongcheng, et al.
Published: (2024)
by: Liu, Hongcheng, et al.
Published: (2024)
When Seeing Is not Enough: Revealing the Limits of Active Reasoning in MLLMs
by: Liu, Hongcheng, et al.
Published: (2025)
by: Liu, Hongcheng, et al.
Published: (2025)
Towards Omni-RAG: Comprehensive Retrieval-Augmented Generation for Large Language Models in Medical Applications
by: Chen, Zhe, et al.
Published: (2025)
by: Chen, Zhe, et al.
Published: (2025)
VocalBench-DF: A Benchmark for Evaluating Speech LLM Robustness to Disfluency
by: Liu, Hongcheng, et al.
Published: (2025)
by: Liu, Hongcheng, et al.
Published: (2025)
Drawing the Line: Enhancing Trustworthiness of MLLMs Through the Power of Refusal
by: Wang, Yuhao, et al.
Published: (2024)
by: Wang, Yuhao, et al.
Published: (2024)
HeteroRAG: A Heterogeneous Retrieval-Augmented Generation Framework for Medical Vision Language Tasks
by: Chen, Zhe, et al.
Published: (2025)
by: Chen, Zhe, et al.
Published: (2025)
Selecting Auxiliary Data via Neural Tangent Kernels for Low-Resource Domains
by: Wang, Pingjie, et al.
Published: (2025)
by: Wang, Pingjie, et al.
Published: (2025)
MedCare: Advancing Medical LLMs through Decoupling Clinical Alignment and Knowledge Aggregation
by: Liao, Yusheng, et al.
Published: (2024)
by: Liao, Yusheng, et al.
Published: (2024)
Automatic Interactive Evaluation for Large Language Models with State Aware Patient Simulator
by: Liao, Yusheng, et al.
Published: (2024)
by: Liao, Yusheng, et al.
Published: (2024)
VocalBench: Benchmarking the Vocal Conversational Abilities for Speech Interaction Models
by: Liu, Heyang, et al.
Published: (2025)
by: Liu, Heyang, et al.
Published: (2025)
MM-SAP: A Comprehensive Benchmark for Assessing Self-Awareness of Multimodal Large Language Models in Perception
by: Wang, Yuhao, et al.
Published: (2024)
by: Wang, Yuhao, et al.
Published: (2024)
Decoding Linguistic Representations of Human Brain
by: Wang, Yu, et al.
Published: (2024)
by: Wang, Yu, et al.
Published: (2024)
Med-PMC: Medical Personalized Multi-modal Consultation with a Proactive Ask-First-Observe-Next Paradigm
by: Liu, Hongcheng, et al.
Published: (2024)
by: Liu, Hongcheng, et al.
Published: (2024)
DSVD: Dynamic Self-Verify Decoding for Faithful Generation in Large Language Models
by: Guo, YiQiu, et al.
Published: (2025)
by: Guo, YiQiu, et al.
Published: (2025)
Agentic Active Omni-Modal Perception for Multi-Hop Audio-Visual Reasoning
by: Xu, Ke, et al.
Published: (2026)
by: Xu, Ke, et al.
Published: (2026)
VocalBench-zh: Decomposing and Benchmarking the Speech Conversational Abilities in Mandarin Context
by: Liu, Heyang, et al.
Published: (2025)
by: Liu, Heyang, et al.
Published: (2025)
M$^3$AV: A Multimodal, Multigenre, and Multipurpose Audio-Visual Academic Lecture Dataset
by: Chen, Zhe, et al.
Published: (2024)
by: Chen, Zhe, et al.
Published: (2024)
CS3-Bench: Evaluating and Enhancing Speech-to-Speech LLMs for Mandarin-English Code-Switching
by: Liu, Heyang, et al.
Published: (2025)
by: Liu, Heyang, et al.
Published: (2025)
Multimodal Cross-Document Event Coreference Resolution Using Linear Semantic Transfer and Mixed-Modality Ensembles
by: Nath, Abhijnan, et al.
Published: (2024)
by: Nath, Abhijnan, et al.
Published: (2024)
DICE: Structured Reasoning in LLMs through SLM-Guided Chain-of-Thought Correction
by: Li, Yiqi, et al.
Published: (2025)
by: Li, Yiqi, et al.
Published: (2025)
Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalities
by: Zhou, Ziwei, et al.
Published: (2025)
by: Zhou, Ziwei, et al.
Published: (2025)
Omni-Safety under Cross-Modality Conflict: Vulnerabilities, Dynamics Mechanisms and Efficient Alignment
by: Wang, Kun, et al.
Published: (2026)
by: Wang, Kun, et al.
Published: (2026)
OmniTrace: A Unified Framework for Generation-Time Attribution in Omni-Modal LLMs
by: Yan, Qianqi, et al.
Published: (2026)
by: Yan, Qianqi, et al.
Published: (2026)
Improving LLMs' Learning for Coreference Resolution
by: Gan, Yujian, et al.
Published: (2025)
by: Gan, Yujian, et al.
Published: (2025)
Cross-Document Contextual Coreference Resolution in Knowledge Graphs
by: Dong, Zhang, et al.
Published: (2025)
by: Dong, Zhang, et al.
Published: (2025)
Is Extending Modality The Right Path Towards Omni-Modality?
by: Zhu, Tinghui, et al.
Published: (2025)
by: Zhu, Tinghui, et al.
Published: (2025)
Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning
by: Ou, Siqu, et al.
Published: (2025)
by: Ou, Siqu, et al.
Published: (2025)
Do LLMs and VLMs Share Neurons for Inference? Evidence and Mechanisms of Cross-Modal Transfer
by: Cui, Chenhang, et al.
Published: (2026)
by: Cui, Chenhang, et al.
Published: (2026)
Ola: Pushing the Frontiers of Omni-Modal Language Model
by: Liu, Zuyan, et al.
Published: (2025)
by: Liu, Zuyan, et al.
Published: (2025)
Miner:Mining Intrinsic Mastery for Data-Efficient RL in Large Reasoning Models
by: Jiang, Shuyang, et al.
Published: (2026)
by: Jiang, Shuyang, et al.
Published: (2026)
Reducing Tool Hallucination via Reliability Alignment
by: Xu, Hongshen, et al.
Published: (2024)
by: Xu, Hongshen, et al.
Published: (2024)
OmniSIFT: Modality-Asymmetric Token Compression for Efficient Omni-modal Large Language Models
by: Ding, Yue, et al.
Published: (2026)
by: Ding, Yue, et al.
Published: (2026)
OmniGAIA: Towards Native Omni-Modal AI Agents
by: Li, Xiaoxi, et al.
Published: (2026)
by: Li, Xiaoxi, et al.
Published: (2026)
FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs
by: Chen, Qian, et al.
Published: (2026)
by: Chen, Qian, et al.
Published: (2026)
Can We Trust LLMs? Mitigate Overconfidence Bias in LLMs through Knowledge Transfer
by: Yang, Haoyan, et al.
Published: (2024)
by: Yang, Haoyan, et al.
Published: (2024)
Generating Harder Cross-document Event Coreference Resolution Datasets using Metaphoric Paraphrasing
by: Ahmed, Shafiuddin Rehan, et al.
Published: (2024)
by: Ahmed, Shafiuddin Rehan, et al.
Published: (2024)
CorPipe at CRAC 2026: Empty Nodes and Cross-Lingual Transfer in Multilingual Coreference Resolution
by: Straka, Milan
Published: (2026)
by: Straka, Milan
Published: (2026)
Differentiable Faithfulness Alignment for Cross-Model Circuit Transfer
by: Shao, Shun, et al.
Published: (2026)
by: Shao, Shun, et al.
Published: (2026)
VocalNet-MDM: Accelerating Streaming Speech LLM via Self-Distilled Masked Diffusion Modeling
by: Cheng, Ziyang, et al.
Published: (2026)
by: Cheng, Ziyang, et al.
Published: (2026)
Cross-Modal Knowledge Distillation for Speech Large Language Models
by: Wang, Enzhi, et al.
Published: (2025)
by: Wang, Enzhi, et al.
Published: (2025)
Similar Items
-
M2K-VDG: Model-Adaptive Multimodal Knowledge Anchor Enhanced Video-grounded Dialogue Generation
by: Liu, Hongcheng, et al.
Published: (2024) -
When Seeing Is not Enough: Revealing the Limits of Active Reasoning in MLLMs
by: Liu, Hongcheng, et al.
Published: (2025) -
Towards Omni-RAG: Comprehensive Retrieval-Augmented Generation for Large Language Models in Medical Applications
by: Chen, Zhe, et al.
Published: (2025) -
VocalBench-DF: A Benchmark for Evaluating Speech LLM Robustness to Disfluency
by: Liu, Hongcheng, et al.
Published: (2025) -
Drawing the Line: Enhancing Trustworthiness of MLLMs Through the Power of Refusal
by: Wang, Yuhao, et al.
Published: (2024)