Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Zhiheng, Ren, Weiming, Huang, Xiaoke, Chen, Shoufa, Li, Tianhong, Chen, Mengzhao, Ji, Yatai, He, Sen, Schult, Jonas, Zeng, Belinda, Xiang, Tao, Chen, Wenhu, Luo, Ping, Zettlemoyer, Luke, Cong, Yuren |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HiStream: Efficient High-Resolution Video Generation via Redundancy-Eliminated Streaming
by: Qiu, Haonan, et al.
Published: (2025)
by: Qiu, Haonan, et al.
Published: (2025)
PixelWorld: How Far Are We from Perceiving Everything as Pixels?
by: Lyu, Zhiheng, et al.
Published: (2025)
by: Lyu, Zhiheng, et al.
Published: (2025)
TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models
by: Liu, Zhiheng, et al.
Published: (2025)
by: Liu, Zhiheng, et al.
Published: (2025)
WavFlow: Audio Generation in Waveform Space
by: Zhou, Feiyan, et al.
Published: (2026)
by: Zhou, Feiyan, et al.
Published: (2026)
PixelFlow: Pixel-Space Generative Models with Flow
by: Chen, Shoufa, et al.
Published: (2025)
by: Chen, Shoufa, et al.
Published: (2025)
Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning
by: Wang, Haozhe, et al.
Published: (2025)
by: Wang, Haozhe, et al.
Published: (2025)
VISTA: Enhancing Long-Duration and High-Resolution Video Understanding by Video Spatiotemporal Augmentation
by: Ren, Weiming, et al.
Published: (2024)
by: Ren, Weiming, et al.
Published: (2024)
Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
by: Sun, Peize, et al.
Published: (2024)
by: Sun, Peize, et al.
Published: (2024)
WorldWeaver: Generating Long-Horizon Video Worlds via Rich Perception
by: Liu, Zhiheng, et al.
Published: (2025)
by: Liu, Zhiheng, et al.
Published: (2025)
Vamba: Understanding Hour-Long Videos with Hybrid Mamba-Transformers
by: Ren, Weiming, et al.
Published: (2025)
by: Ren, Weiming, et al.
Published: (2025)
ABC: Achieving Better Control of Multimodal Embeddings using VLMs
by: Schneider, Benjamin, et al.
Published: (2025)
by: Schneider, Benjamin, et al.
Published: (2025)
GenTron: Diffusion Transformers for Image and Video Generation
by: Chen, Shoufa, et al.
Published: (2023)
by: Chen, Shoufa, et al.
Published: (2023)
AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks
by: Ku, Max, et al.
Published: (2024)
by: Ku, Max, et al.
Published: (2024)
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation
by: Ma, Wentao, et al.
Published: (2025)
by: Ma, Wentao, et al.
Published: (2025)
Unifying Multimodal Retrieval via Document Screenshot Embedding
by: Ma, Xueguang, et al.
Published: (2024)
by: Ma, Xueguang, et al.
Published: (2024)
FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editing
by: Cong, Yuren, et al.
Published: (2023)
by: Cong, Yuren, et al.
Published: (2023)
Quantifying the Gap between Understanding and Generation within Unified Multimodal Models
by: Wang, Chenlong, et al.
Published: (2026)
by: Wang, Chenlong, et al.
Published: (2026)
WorldAfford: Affordance Grounding based on Natural Language Instructions
by: Chen, Changmao, et al.
Published: (2024)
by: Chen, Changmao, et al.
Published: (2024)
Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models
by: Yasunaga, Michihiro, et al.
Published: (2025)
by: Yasunaga, Michihiro, et al.
Published: (2025)
RationalRewards: Reasoning Rewards Scale Visual Generation Both Training and Test Time
by: Wang, Haozhe, et al.
Published: (2026)
by: Wang, Haozhe, et al.
Published: (2026)
OmniEdit: Building Image Editing Generalist Models Through Specialist Supervision
by: Wei, Cong, et al.
Published: (2024)
by: Wei, Cong, et al.
Published: (2024)
SWE-Next: Scalable Real-World Software Engineering Tasks for Agents
by: Liang, Jiarong, et al.
Published: (2026)
by: Liang, Jiarong, et al.
Published: (2026)
Prompt-A-Video: Prompt Your Video Diffusion Model via Preference-Aligned LLM
by: Ji, Yatai, et al.
Published: (2024)
by: Ji, Yatai, et al.
Published: (2024)
SVTRv2: CTC Beats Encoder-Decoder Models in Scene Text Recognition
by: Du, Yongkun, et al.
Published: (2024)
by: Du, Yongkun, et al.
Published: (2024)
VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks
by: Jiang, Ziyan, et al.
Published: (2024)
by: Jiang, Ziyan, et al.
Published: (2024)
ConsistI2V: Enhancing Visual Consistency for Image-to-Video Generation
by: Ren, Weiming, et al.
Published: (2024)
by: Ren, Weiming, et al.
Published: (2024)
Reconstruction Alignment Improves Unified Multimodal Models
by: Xie, Ji, et al.
Published: (2025)
by: Xie, Ji, et al.
Published: (2025)
From Denoising to Refining: A Corrective Framework for Vision-Language Diffusion Model
by: Ji, Yatai, et al.
Published: (2025)
by: Ji, Yatai, et al.
Published: (2025)
TheoremExplainAgent: Towards Video-based Multimodal Explanations for LLM Theorem Understanding
by: Ku, Max, et al.
Published: (2025)
by: Ku, Max, et al.
Published: (2025)
SafePLUG: Empowering Multimodal LLMs with Pixel-Level Insight and Temporal Grounding for Traffic Accident Understanding
by: Sheng, Zihao, et al.
Published: (2025)
by: Sheng, Zihao, et al.
Published: (2025)
OmniEncoder: See, Hear, and Feel Continuous Motion Like Humans With One Encoder
by: Bai, Detao, et al.
Published: (2026)
by: Bai, Detao, et al.
Published: (2026)
Understanding Transformer Encoder-Decoder Representations through Bernoulli Dropout
by: Chen, Xuanzhou
Published: (2026)
by: Chen, Xuanzhou
Published: (2026)
Scaling Zero-Shot Reference-to-Video Generation
by: Zhou, Zijian, et al.
Published: (2025)
by: Zhou, Zijian, et al.
Published: (2025)
Radical‐Induced Hour‐Level Afterglow and Efficient Circularly Polarized Luminescence From Metal Halide Hybrid Glasses
by: Tianhong Chen, et al.
Published: (2026)
by: Tianhong Chen, et al.
Published: (2026)
Thermal‐Induced Hole‐Current Reinforcement in Quantum Dot Light‐Emitting Diodes
by: Tianhong Zhou, et al.
Published: (2025)
by: Tianhong Zhou, et al.
Published: (2025)
Radical‐Induced Hour‐Level Afterglow and Efficient Circularly Polarized Luminescence From Metal Halide Hybrid Glasses
by: Tianhong Chen, et al.
Published: (2026)
by: Tianhong Chen, et al.
Published: (2026)
IDA-VLM: Towards Movie Understanding via ID-Aware Large Vision-Language Model
by: Ji, Yatai, et al.
Published: (2024)
by: Ji, Yatai, et al.
Published: (2024)
Bridging Pixels and Words: Mask-Aware Local Semantic Fusion for Multimodal Media Verification
by: Chen, Zizhao, et al.
Published: (2026)
by: Chen, Zizhao, et al.
Published: (2026)
Why Parenthood Strains Relationships: Investigating the Mechanisms Behind Declining Relationship Satisfaction
by: Matthias Pollmann‐Schult
Published: (2026)
by: Matthias Pollmann‐Schult
Published: (2026)
PixelBytes: Catching Unified Embedding for Multimodal Generation
by: Furfaro, Fabien
Published: (2024)
by: Furfaro, Fabien
Published: (2024)
Similar Items
-
HiStream: Efficient High-Resolution Video Generation via Redundancy-Eliminated Streaming
by: Qiu, Haonan, et al.
Published: (2025) -
PixelWorld: How Far Are We from Perceiving Everything as Pixels?
by: Lyu, Zhiheng, et al.
Published: (2025) -
TUNA: Taming Unified Visual Representations for Native Unified Multimodal Models
by: Liu, Zhiheng, et al.
Published: (2025) -
WavFlow: Audio Generation in Waveform Space
by: Zhou, Feiyan, et al.
Published: (2026) -
PixelFlow: Pixel-Space Generative Models with Flow
by: Chen, Shoufa, et al.
Published: (2025)