Saved in:
| Main Authors: | Xiao, Teng, Li, Zuchao, Zhang, Lefei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2509.19018 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
NOTA: Multimodal Music Notation Understanding for Visual Large Language Model
by: Tang, Mingni, et al.
Published: (2025)
by: Tang, Mingni, et al.
Published: (2025)
Sparse is Enough in Fine-tuning Pre-trained Large Language Models
by: Song, Weixi, et al.
Published: (2023)
by: Song, Weixi, et al.
Published: (2023)
Teaching Your Models to Understand Code via Focal Preference Alignment
by: Wu, Jie, et al.
Published: (2025)
by: Wu, Jie, et al.
Published: (2025)
Model Hemorrhage and the Robustness Limits of Large Language Models
by: Ma, Ziyang, et al.
Published: (2025)
by: Ma, Ziyang, et al.
Published: (2025)
From AR to Diffusion: Efficiently Adapting Large Language Models with Strictly Causal and Elastic Horizons
by: Ma, Xiangyu, et al.
Published: (2026)
by: Ma, Xiangyu, et al.
Published: (2026)
TimeOmni-VL: Unified Models for Time Series Understanding and Generation
by: Guan, Tong, et al.
Published: (2026)
by: Guan, Tong, et al.
Published: (2026)
Latent Space Translation via Semantic Alignment
by: Maiorca, Valentino, et al.
Published: (2023)
by: Maiorca, Valentino, et al.
Published: (2023)
CyIN: Cyclic Informative Latent Space for Bridging Complete and Incomplete Multimodal Learning
by: Lin, Ronghao, et al.
Published: (2026)
by: Lin, Ronghao, et al.
Published: (2026)
MGIMM: Multi-Granularity Instruction Multimodal Model for Attribute-Guided Remote Sensing Image Detailed Description
by: Yang, Cong, et al.
Published: (2024)
by: Yang, Cong, et al.
Published: (2024)
OmniRetrieval: Unified Retrieval across Heterogeneous Knowledge Sources
by: Baek, Jinheon, et al.
Published: (2026)
by: Baek, Jinheon, et al.
Published: (2026)
Ming-Omni: A Unified Multimodal Model for Perception and Generation
by: AI, Inclusion, et al.
Published: (2025)
by: AI, Inclusion, et al.
Published: (2025)
Exploring Representation-Aligned Latent Space for Better Generation
by: Xu, Wanghan, et al.
Published: (2025)
by: Xu, Wanghan, et al.
Published: (2025)
LatentUM: Unleashing the Potential of Interleaved Cross-Modal Reasoning via a Latent-Space Unified Model
by: Jin, Jiachun, et al.
Published: (2026)
by: Jin, Jiachun, et al.
Published: (2026)
Towards Unified and Lossless Latent Space for 3D Molecular Latent Diffusion Modeling
by: Luo, Yanchen, et al.
Published: (2025)
by: Luo, Yanchen, et al.
Published: (2025)
Latent Space Communication via K-V Cache Alignment
by: Dery, Lucio M., et al.
Published: (2026)
by: Dery, Lucio M., et al.
Published: (2026)
Bottleneck Tokens for Unified Multimodal Retrieval
by: Sun, Siyu, et al.
Published: (2026)
by: Sun, Siyu, et al.
Published: (2026)
FedEHR-Gen: Federated Synthetic Time-Series EHR Generation via Latent Space Alignment and Distribution-Aware Aggregation
by: Bai, Jun, et al.
Published: (2026)
by: Bai, Jun, et al.
Published: (2026)
AR-Omni: A Unified Autoregressive Model for Any-to-Any Generation
by: Cheng, Dongjie, et al.
Published: (2026)
by: Cheng, Dongjie, et al.
Published: (2026)
Unifying Ranking and Generation in Query Auto-Completion via Retrieval-Augmented Generation and Multi-Objective Alignment
by: Yuan, Kai, et al.
Published: (2026)
by: Yuan, Kai, et al.
Published: (2026)
Agent-Omni: Test-Time Multimodal Reasoning via Model Coordination for Understanding Anything
by: Lin, Huawei, et al.
Published: (2025)
by: Lin, Huawei, et al.
Published: (2025)
MolBind: Multimodal Alignment of Language, Molecules, and Proteins
by: Xiao, Teng, et al.
Published: (2024)
by: Xiao, Teng, et al.
Published: (2024)
Heterogeneous-Modal Unsupervised Domain Adaptation via Latent Space Bridging
by: Yang, Jiawen, et al.
Published: (2025)
by: Yang, Jiawen, et al.
Published: (2025)
MRAMG-Bench: A Comprehensive Benchmark for Advancing Multimodal Retrieval-Augmented Multimodal Generation
by: Yu, Qinhan, et al.
Published: (2025)
by: Yu, Qinhan, et al.
Published: (2025)
DeepRTL: Bridging Verilog Understanding and Generation with a Unified Representation Model
by: Liu, Yi, et al.
Published: (2025)
by: Liu, Yi, et al.
Published: (2025)
4D Multimodal Co-attention Fusion Network with Latent Contrastive Alignment for Alzheimer's Diagnosis
by: Wei, Yuxiang, et al.
Published: (2025)
by: Wei, Yuxiang, et al.
Published: (2025)
A Unified Multi-Agent Framework for Universal Multimodal Understanding and Generation
by: Li, Jiulin, et al.
Published: (2025)
by: Li, Jiulin, et al.
Published: (2025)
BrainOmni: A Brain Foundation Model for Unified EEG and MEG Signals
by: Xiao, Qinfan, et al.
Published: (2025)
by: Xiao, Qinfan, et al.
Published: (2025)
Unified Biomolecular Trajectory Generation via Pretrained Variational Bridge
by: Yu, Ziyang, et al.
Published: (2026)
by: Yu, Ziyang, et al.
Published: (2026)
Understanding the Emergence of Multimodal Representation Alignment
by: Tjandrasuwita, Megan, et al.
Published: (2025)
by: Tjandrasuwita, Megan, et al.
Published: (2025)
Latent Diffusion Inversion Requires Understanding the Latent Space
by: Rao, Mingxing, et al.
Published: (2025)
by: Rao, Mingxing, et al.
Published: (2025)
IAM: Efficient Inference through Attention Mapping between Different-scale LLMs
by: Zhao, Yi, et al.
Published: (2025)
by: Zhao, Yi, et al.
Published: (2025)
Towards Bridging the Reward-Generation Gap in Direct Alignment Algorithms
by: Xiao, Zeguan, et al.
Published: (2025)
by: Xiao, Zeguan, et al.
Published: (2025)
Understand What LLM Needs: Dual Preference Alignment for Retrieval-Augmented Generation
by: Dong, Guanting, et al.
Published: (2024)
by: Dong, Guanting, et al.
Published: (2024)
Cross-Domain Diffusion with Progressive Alignment for Efficient Adaptive Retrieval
by: Luo, Junyu, et al.
Published: (2025)
by: Luo, Junyu, et al.
Published: (2025)
TopoPrune: Robust Data Pruning via Unified Latent Space Topology
by: Roy, Arjun, et al.
Published: (2026)
by: Roy, Arjun, et al.
Published: (2026)
OmniISR: A Unified Framework for Centralized and Federated Learning via Intermediate Supervision and Regularization
by: Kou, Wei-Bin, et al.
Published: (2026)
by: Kou, Wei-Bin, et al.
Published: (2026)
Centroid-centered Modeling for Efficient Vision Transformer Pre-training
by: Yan, Xin, et al.
Published: (2023)
by: Yan, Xin, et al.
Published: (2023)
RACER: Retrieval-Augmented Contextual Rapid Speculative Decoding
by: Zhang, Zihong, et al.
Published: (2026)
by: Zhang, Zihong, et al.
Published: (2026)
STARFlow2: Bridging Language Models and Normalizing Flows for Unified Multimodal Generation
by: Shen, Ying, et al.
Published: (2026)
by: Shen, Ying, et al.
Published: (2026)
Latent Space Alignment for Semantic Channel Equalization
by: Hüttebräucker, Tomás, et al.
Published: (2024)
by: Hüttebräucker, Tomás, et al.
Published: (2024)
Similar Items
-
NOTA: Multimodal Music Notation Understanding for Visual Large Language Model
by: Tang, Mingni, et al.
Published: (2025) -
Sparse is Enough in Fine-tuning Pre-trained Large Language Models
by: Song, Weixi, et al.
Published: (2023) -
Teaching Your Models to Understand Code via Focal Preference Alignment
by: Wu, Jie, et al.
Published: (2025) -
Model Hemorrhage and the Robustness Limits of Large Language Models
by: Ma, Ziyang, et al.
Published: (2025) -
From AR to Diffusion: Efficiently Adapting Large Language Models with Strictly Causal and Elastic Horizons
by: Ma, Xiangyu, et al.
Published: (2026)