Representation Forcing for Bottleneck-Free Unified Multimodal Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Yuqing, Lin, Zhijie, Yang, Ceyuan, Zhao, Yang, Xiao, Fei, He, Hao, Zhao, Qi, Ding, Zihan, Wang, Fuyun, Wang, Shuai, Zhang, Youliang, Fan, Haoqi, Liu, Xihui |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Continuous Adversarial Flow Models
von: Lin, Shanchuan, et al.
Veröffentlicht: (2026)
von: Lin, Shanchuan, et al.
Veröffentlicht: (2026)
Adversarial Flow Models
von: Lin, Shanchuan, et al.
Veröffentlicht: (2025)
von: Lin, Shanchuan, et al.
Veröffentlicht: (2025)
Context Unrolling in Omni Models
von: Yang, Ceyuan, et al.
Veröffentlicht: (2026)
von: Yang, Ceyuan, et al.
Veröffentlicht: (2026)
SeedVR: Seeding Infinity in Diffusion Transformer Towards Generic Video Restoration
von: Wang, Jianyi, et al.
Veröffentlicht: (2025)
von: Wang, Jianyi, et al.
Veröffentlicht: (2025)
Cubic Discrete Diffusion: Discrete Visual Generation on High-Dimensional Representation Tokens
von: Wang, Yuqing, et al.
Veröffentlicht: (2026)
von: Wang, Yuqing, et al.
Veröffentlicht: (2026)
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
von: Han, Jiaming, et al.
Veröffentlicht: (2025)
von: Han, Jiaming, et al.
Veröffentlicht: (2025)
Loong: Generating Minute-level Long Videos with Autoregressive Language Models
von: Wang, Yuqing, et al.
Veröffentlicht: (2024)
von: Wang, Yuqing, et al.
Veröffentlicht: (2024)
Graph Bottlenecked Social Recommendation
von: Yang, Yonghui, et al.
Veröffentlicht: (2024)
von: Yang, Yonghui, et al.
Veröffentlicht: (2024)
UniMC: Taming Diffusion Transformer for Unified Keypoint-Guided Multi-Class Image Generation
von: Guo, Qin, et al.
Veröffentlicht: (2025)
von: Guo, Qin, et al.
Veröffentlicht: (2025)
Bottleneck Tokens for Unified Multimodal Retrieval
von: Sun, Siyu, et al.
Veröffentlicht: (2026)
von: Sun, Siyu, et al.
Veröffentlicht: (2026)
Learning Optimal Multimodal Information Bottleneck Representations
von: Wu, Qilong, et al.
Veröffentlicht: (2025)
von: Wu, Qilong, et al.
Veröffentlicht: (2025)
Gemini in Reasoning: Unveiling Commonsense in Multimodal Large Language Models
von: Wang, Yuqing, et al.
Veröffentlicht: (2023)
von: Wang, Yuqing, et al.
Veröffentlicht: (2023)
Multimodal Representation-disentangled Information Bottleneck for Multimodal Recommendation
von: Wang, Hui, et al.
Veröffentlicht: (2025)
von: Wang, Hui, et al.
Veröffentlicht: (2025)
GenArtist: Multimodal LLM as an Agent for Unified Image Generation and Editing
von: Wang, Zhenyu, et al.
Veröffentlicht: (2024)
von: Wang, Zhenyu, et al.
Veröffentlicht: (2024)
LVD-2M: A Long-take Video Dataset with Temporally Dense Captions
von: Xiong, Tianwei, et al.
Veröffentlicht: (2024)
von: Xiong, Tianwei, et al.
Veröffentlicht: (2024)
UniECS: Unified Multimodal E-Commerce Search Framework with Gated Cross-modal Fusion
von: Liang, Zihan, et al.
Veröffentlicht: (2025)
von: Liang, Zihan, et al.
Veröffentlicht: (2025)
Understanding Multimodal Hallucination with Parameter-Free Representation Alignment
von: Wang, Yueqian, et al.
Veröffentlicht: (2024)
von: Wang, Yueqian, et al.
Veröffentlicht: (2024)
UniCom: Unified Multimodal Modeling via Compressed Continuous Semantic Representations
von: Zhao, Yaqi, et al.
Veröffentlicht: (2026)
von: Zhao, Yaqi, et al.
Veröffentlicht: (2026)
FreeBind: Free Lunch in Unified Multimodal Space via Knowledge Fusion
von: Wang, Zehan, et al.
Veröffentlicht: (2024)
von: Wang, Zehan, et al.
Veröffentlicht: (2024)
Scaling Laws For Diffusion Transformers
von: Liang, Zhengyang, et al.
Veröffentlicht: (2024)
von: Liang, Zhengyang, et al.
Veröffentlicht: (2024)
VTok: A Unified Video Tokenizer with Decoupled Spatial-Temporal Latents
von: Wang, Feng, et al.
Veröffentlicht: (2026)
von: Wang, Feng, et al.
Veröffentlicht: (2026)
Difficult Task Yes but Simple Task No: Unveiling the Laziness in Multimodal LLMs
von: Zhao, Sihang, et al.
Veröffentlicht: (2024)
von: Zhao, Sihang, et al.
Veröffentlicht: (2024)
NeuroBind: Towards Unified Multimodal Representations for Neural Signals
von: Yang, Fengyu, et al.
Veröffentlicht: (2024)
von: Yang, Fengyu, et al.
Veröffentlicht: (2024)
Can't See the Forest for the Trees: Benchmarking Multimodal Safety Awareness for Multimodal LLMs
von: Wang, Wenxuan, et al.
Veröffentlicht: (2025)
von: Wang, Wenxuan, et al.
Veröffentlicht: (2025)
Bridging Domain Generalization to Multimodal Domain Generalization via Unified Representations
von: Huang, Hai, et al.
Veröffentlicht: (2025)
von: Huang, Hai, et al.
Veröffentlicht: (2025)
From Static Analysis to Audience Dissemination: A Training-Free Multimodal Controversy Detection Multi-Agent Framework
von: Ding, Zihan, et al.
Veröffentlicht: (2026)
von: Ding, Zihan, et al.
Veröffentlicht: (2026)
SHAPE: Unifying Safety, Helpfulness and Pedagogy for Educational LLMs
von: Zhao, Sihang, et al.
Veröffentlicht: (2026)
von: Zhao, Sihang, et al.
Veröffentlicht: (2026)
Cascade-Free Mandarin Visual Speech Recognition via Semantic-Guided Cross-Representation Alignment
von: Yang, Lei, et al.
Veröffentlicht: (2026)
von: Yang, Lei, et al.
Veröffentlicht: (2026)
Bridging Continuous and Discrete Tokens for Autoregressive Visual Generation
von: Wang, Yuqing, et al.
Veröffentlicht: (2025)
von: Wang, Yuqing, et al.
Veröffentlicht: (2025)
IBCapsNet: Information Bottleneck Capsule Network for Noise-Robust Representation Learning
von: Xiang, Canqun, et al.
Veröffentlicht: (2026)
von: Xiang, Canqun, et al.
Veröffentlicht: (2026)
Research on Edge Computing and Cloud Collaborative Resource Scheduling Optimization Based on Deep Reinforcement Learning
von: Wang, Yuqing, et al.
Veröffentlicht: (2025)
von: Wang, Yuqing, et al.
Veröffentlicht: (2025)
Research on Enhancing Cloud Computing Network Security using Artificial Intelligence Algorithms
von: Wang, Yuqing, et al.
Veröffentlicht: (2025)
von: Wang, Yuqing, et al.
Veröffentlicht: (2025)
Machine Learning-Based Cloud Computing Compliance Process Automation
von: Wang, Yuqing, et al.
Veröffentlicht: (2025)
von: Wang, Yuqing, et al.
Veröffentlicht: (2025)
Intelligent Resource Allocation Optimization for Cloud Computing via Machine Learning
von: Wang, Yuqing, et al.
Veröffentlicht: (2025)
von: Wang, Yuqing, et al.
Veröffentlicht: (2025)
Cloud Computing Energy Consumption Prediction Based on Kernel Extreme Learning Machine Algorithm Improved by Vector Weighted Average Algorithm
von: Wang, Yuqing, et al.
Veröffentlicht: (2025)
von: Wang, Yuqing, et al.
Veröffentlicht: (2025)
Design and implementation of a distributed security threat detection system integrating federated learning and multimodal LLM
von: Wang, Yuqing, et al.
Veröffentlicht: (2025)
von: Wang, Yuqing, et al.
Veröffentlicht: (2025)
Enhancing Multimodal Unified Representations for Cross Modal Generalization
von: Huang, Hai, et al.
Veröffentlicht: (2024)
von: Huang, Hai, et al.
Veröffentlicht: (2024)
End-to-End Training for Autoregressive Video Diffusion via Self-Resampling
von: Guo, Yuwei, et al.
Veröffentlicht: (2025)
von: Guo, Yuwei, et al.
Veröffentlicht: (2025)
BabelDOC: Better Layout-Preserving PDF Translation via Intermediate Representation
von: Yang, Qi, et al.
Veröffentlicht: (2026)
von: Yang, Qi, et al.
Veröffentlicht: (2026)
MACRO: Advancing Multi-Reference Image Generation with Structured Long-Context Data
von: Chen, Zhekai, et al.
Veröffentlicht: (2026)
von: Chen, Zhekai, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Continuous Adversarial Flow Models
von: Lin, Shanchuan, et al.
Veröffentlicht: (2026) -
Adversarial Flow Models
von: Lin, Shanchuan, et al.
Veröffentlicht: (2025) -
Context Unrolling in Omni Models
von: Yang, Ceyuan, et al.
Veröffentlicht: (2026) -
SeedVR: Seeding Infinity in Diffusion Transformer Towards Generic Video Restoration
von: Wang, Jianyi, et al.
Veröffentlicht: (2025) -
Cubic Discrete Diffusion: Discrete Visual Generation on High-Dimensional Representation Tokens
von: Wang, Yuqing, et al.
Veröffentlicht: (2026)