TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Zhiheng, Ren, Weiming, Liu, Haozhe, Zhou, Zijian, Chen, Shoufa, Qiu, Haonan, Huang, Xiaoke, An, Zhaochong, Yang, Fanny, Patel, Aditya, Atliha, Viktar, Ng, Tony, Han, Xiao, Zhu, Chuyan, Zhang, Chenyang, Liu, Ding, Perez-Rua, Juan-Manuel, He, Sen, Schmidhuber, Jürgen, Chen, Wenhu, Luo, Ping, Liu, Wei, Xiang, Tao, Schult, Jonas, Cong, Yuren |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HiStream: Efficient High-Resolution Video Generation via Redundancy-Eliminated Streaming
von: Qiu, Haonan, et al.
Veröffentlicht: (2025)
von: Qiu, Haonan, et al.
Veröffentlicht: (2025)
Scaling Zero-Shot Reference-to-Video Generation
von: Zhou, Zijian, et al.
Veröffentlicht: (2025)
von: Zhou, Zijian, et al.
Veröffentlicht: (2025)
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
von: Liu, Zhiheng, et al.
Veröffentlicht: (2026)
von: Liu, Zhiheng, et al.
Veröffentlicht: (2026)
VecGlypher: Unified Vector Glyph Generation with Language Models
von: Huang, Xiaoke, et al.
Veröffentlicht: (2026)
von: Huang, Xiaoke, et al.
Veröffentlicht: (2026)
CogDoc: Towards Unified thinking in Documents
von: Xu, Qixin, et al.
Veröffentlicht: (2025)
von: Xu, Qixin, et al.
Veröffentlicht: (2025)
OneStory: Coherent Multi-Shot Video Generation with Adaptive Memory
von: An, Zhaochong, et al.
Veröffentlicht: (2025)
von: An, Zhaochong, et al.
Veröffentlicht: (2025)
WavFlow: Audio Generation in Waveform Space
von: Zhou, Feiyan, et al.
Veröffentlicht: (2026)
von: Zhou, Feiyan, et al.
Veröffentlicht: (2026)
RationalRewards: Reasoning Rewards Scale Visual Generation Both Training and Test Time
von: Wang, Haozhe, et al.
Veröffentlicht: (2026)
von: Wang, Haozhe, et al.
Veröffentlicht: (2026)
Starve to Perceive: Taming Lazy Perception in VLMs with Constrained Visual Bandwidth
von: Wu, Yuhuan, et al.
Veröffentlicht: (2026)
von: Wu, Yuhuan, et al.
Veröffentlicht: (2026)
UniPaint: Unified Space-time Video Inpainting via Mixture-of-Experts
von: Wan, Zhen, et al.
Veröffentlicht: (2024)
von: Wan, Zhen, et al.
Veröffentlicht: (2024)
Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning
von: Wang, Haozhe, et al.
Veröffentlicht: (2025)
von: Wang, Haozhe, et al.
Veröffentlicht: (2025)
Beyond Attribution: Unified Concept-Level Explanations
von: Liu, Junhao, et al.
Veröffentlicht: (2024)
von: Liu, Junhao, et al.
Veröffentlicht: (2024)
Flash-Unified: A Training-Free and Task-Aware Acceleration Framework for Native Unified Models
von: Ke, Junlong, et al.
Veröffentlicht: (2026)
von: Ke, Junlong, et al.
Veröffentlicht: (2026)
TransText: Alpha-as-RGB Representation for Transparent Text Animation
von: Zhang, Fei, et al.
Veröffentlicht: (2026)
von: Zhang, Fei, et al.
Veröffentlicht: (2026)
UniVideo: Unified Understanding, Generation, and Editing for Videos
von: Wei, Cong, et al.
Veröffentlicht: (2025)
von: Wei, Cong, et al.
Veröffentlicht: (2025)
CoCo-Fed: A Unified Framework for Memory- and Communication-Efficient Federated Learning at the Wireless Edge
von: Guo, Zhiheng, et al.
Veröffentlicht: (2026)
von: Guo, Zhiheng, et al.
Veröffentlicht: (2026)
PixelWorld: How Far Are We from Perceiving Everything as Pixels?
von: Lyu, Zhiheng, et al.
Veröffentlicht: (2025)
von: Lyu, Zhiheng, et al.
Veröffentlicht: (2025)
Lazy Layers to Make Fine-Tuned Diffusion Models More Traceable
von: Liu, Haozhe, et al.
Veröffentlicht: (2024)
von: Liu, Haozhe, et al.
Veröffentlicht: (2024)
NAG: A Unified Native Architecture for Encoder-free Text-Graph Modeling in Language Models
von: Gong, Haisong, et al.
Veröffentlicht: (2026)
von: Gong, Haisong, et al.
Veröffentlicht: (2026)
Mixture of States: Routing Token-Level Dynamics for Multimodal Generation
von: Liu, Haozhe, et al.
Veröffentlicht: (2025)
von: Liu, Haozhe, et al.
Veröffentlicht: (2025)
UniMC: Taming Diffusion Transformer for Unified Keypoint-Guided Multi-Class Image Generation
von: Guo, Qin, et al.
Veröffentlicht: (2025)
von: Guo, Qin, et al.
Veröffentlicht: (2025)
Temp-R1: A Unified Autonomous Agent for Complex Temporal KGQA via Reverse Curriculum Reinforcement Learning
von: Gong, Zhaoyan, et al.
Veröffentlicht: (2026)
von: Gong, Zhaoyan, et al.
Veröffentlicht: (2026)
Emergent Hierarchical Reasoning in LLMs through Reinforcement Learning
von: Wang, Haozhe, et al.
Veröffentlicht: (2025)
von: Wang, Haozhe, et al.
Veröffentlicht: (2025)
Balancing Knowledge Updates: Toward Unified Modular Editing in LLMs
von: Liu, Jiahao, et al.
Veröffentlicht: (2025)
von: Liu, Jiahao, et al.
Veröffentlicht: (2025)
OneWorld: Taming Scene Generation with 3D Unified Representation Autoencoder
von: Gao, Sensen, et al.
Veröffentlicht: (2026)
von: Gao, Sensen, et al.
Veröffentlicht: (2026)
Unifying Multimodal Retrieval via Document Screenshot Embedding
von: Ma, Xueguang, et al.
Veröffentlicht: (2024)
von: Ma, Xueguang, et al.
Veröffentlicht: (2024)
Quantifying the Gap between Understanding and Generation within Unified Multimodal Models
von: Wang, Chenlong, et al.
Veröffentlicht: (2026)
von: Wang, Chenlong, et al.
Veröffentlicht: (2026)
SWE-Next: Scalable Real-World Software Engineering Tasks for Agents
von: Liang, Jiarong, et al.
Veröffentlicht: (2026)
von: Liang, Jiarong, et al.
Veröffentlicht: (2026)
AEGIS: Exploring the Limit of World Knowledge Capabilities for Unified Mulitmodal Models
von: Lin, Jintao, et al.
Veröffentlicht: (2026)
von: Lin, Jintao, et al.
Veröffentlicht: (2026)
GenTron: Diffusion Transformers for Image and Video Generation
von: Chen, Shoufa, et al.
Veröffentlicht: (2023)
von: Chen, Shoufa, et al.
Veröffentlicht: (2023)
FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editing
von: Cong, Yuren, et al.
Veröffentlicht: (2023)
von: Cong, Yuren, et al.
Veröffentlicht: (2023)
Moment Alignment: Unifying Gradient and Hessian Matching for Domain Generalization
von: Chen, Yuen, et al.
Veröffentlicht: (2025)
von: Chen, Yuen, et al.
Veröffentlicht: (2025)
Unisoma: A Unified Transformer-based Solver for Multi-Solid Systems
von: Tao, Shilong, et al.
Veröffentlicht: (2025)
von: Tao, Shilong, et al.
Veröffentlicht: (2025)
Greed is Good: A Unifying Perspective on Guided Generation
von: Blasingame, Zander W., et al.
Veröffentlicht: (2025)
von: Blasingame, Zander W., et al.
Veröffentlicht: (2025)
Sparse-PGD: A Unified Framework for Sparse Adversarial Perturbations Generation
von: Zhong, Xuyang, et al.
Veröffentlicht: (2024)
von: Zhong, Xuyang, et al.
Veröffentlicht: (2024)
HiDream-O1-Image: A Natively Unified Image Generative Foundation Model with Pixel-level Unified Transformer
von: Cai, Qi, et al.
Veröffentlicht: (2026)
von: Cai, Qi, et al.
Veröffentlicht: (2026)
UniGeo: Taming Video Diffusion for Unified Consistent Geometry Estimation
von: Sun, Yang-Tian, et al.
Veröffentlicht: (2025)
von: Sun, Yang-Tian, et al.
Veröffentlicht: (2025)
Cross-Space Synergy: A Unified Framework for Multimodal Emotion Recognition in Conversation
von: Lyu, Xiaosen, et al.
Veröffentlicht: (2025)
von: Lyu, Xiaosen, et al.
Veröffentlicht: (2025)
VanGogh: A Unified Multimodal Diffusion-based Framework for Video Colorization
von: Fang, Zixun, et al.
Veröffentlicht: (2025)
von: Fang, Zixun, et al.
Veröffentlicht: (2025)
On the Limits of Token Reduction for Efficient Unified Vision Language Training
von: Chen, Siyi, et al.
Veröffentlicht: (2026)
von: Chen, Siyi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
HiStream: Efficient High-Resolution Video Generation via Redundancy-Eliminated Streaming
von: Qiu, Haonan, et al.
Veröffentlicht: (2025) -
Scaling Zero-Shot Reference-to-Video Generation
von: Zhou, Zijian, et al.
Veröffentlicht: (2025) -
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
von: Liu, Zhiheng, et al.
Veröffentlicht: (2026) -
VecGlypher: Unified Vector Glyph Generation with Language Models
von: Huang, Xiaoke, et al.
Veröffentlicht: (2026) -
CogDoc: Towards Unified thinking in Documents
von: Xu, Qixin, et al.
Veröffentlicht: (2025)