PixelBytes: Catching Unified Representation for Multimodal Generation
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Furfaro, Fabien |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PixelBytes: Catching Unified Embedding for Multimodal Generation
von: Furfaro, Fabien
Veröffentlicht: (2024)
von: Furfaro, Fabien
Veröffentlicht: (2024)
Unified Multimodal Understanding via Byte-Pair Visual Encoding
von: Zhang, Wanpeng, et al.
Veröffentlicht: (2025)
von: Zhang, Wanpeng, et al.
Veröffentlicht: (2025)
Enhancing Multimodal Unified Representations for Cross Modal Generalization
von: Huang, Hai, et al.
Veröffentlicht: (2024)
von: Huang, Hai, et al.
Veröffentlicht: (2024)
From Pixels to Tokens: Byte-Pair Encoding on Quantized Visual Modalities
von: Zhang, Wanpeng, et al.
Veröffentlicht: (2024)
von: Zhang, Wanpeng, et al.
Veröffentlicht: (2024)
UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning
von: Liu, Ye, et al.
Veröffentlicht: (2025)
von: Liu, Ye, et al.
Veröffentlicht: (2025)
Catch-Up Mix: Catch-Up Class for Struggling Filters in CNN
von: Kang, Minsoo, et al.
Veröffentlicht: (2024)
von: Kang, Minsoo, et al.
Veröffentlicht: (2024)
UniEval: Unified Holistic Evaluation for Unified Multimodal Understanding and Generation
von: Li, Yi, et al.
Veröffentlicht: (2025)
von: Li, Yi, et al.
Veröffentlicht: (2025)
Semantic Generative Tuning for Unified Multimodal Models
von: Yu, Songsong, et al.
Veröffentlicht: (2026)
von: Yu, Songsong, et al.
Veröffentlicht: (2026)
GLaMM: Pixel Grounding Large Multimodal Model
von: Rasheed, Hanoona, et al.
Veröffentlicht: (2023)
von: Rasheed, Hanoona, et al.
Veröffentlicht: (2023)
Pixel-Grounded Retrieval for Knowledgeable Large Multimodal Models
von: Kim, Jeonghwan, et al.
Veröffentlicht: (2026)
von: Kim, Jeonghwan, et al.
Veröffentlicht: (2026)
The Thinking Pixel: Recursive Sparse Reasoning in Multimodal Diffusion Latents
von: Sun, Yuwei, et al.
Veröffentlicht: (2026)
von: Sun, Yuwei, et al.
Veröffentlicht: (2026)
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
von: Qu, Liao, et al.
Veröffentlicht: (2024)
von: Qu, Liao, et al.
Veröffentlicht: (2024)
Discrete Diffusion Models with MLLMs for Unified Medical Multimodal Generation
von: Mao, Jiawei, et al.
Veröffentlicht: (2025)
von: Mao, Jiawei, et al.
Veröffentlicht: (2025)
Steering Visual Generation in Unified Multimodal Models with Understanding Supervision
von: Liu, Zeyu, et al.
Veröffentlicht: (2026)
von: Liu, Zeyu, et al.
Veröffentlicht: (2026)
Towards a Multimodal Large Language Model with Pixel-Level Insight for Biomedicine
von: Huang, Xiaoshuang, et al.
Veröffentlicht: (2024)
von: Huang, Xiaoshuang, et al.
Veröffentlicht: (2024)
OmniCam: Unified Multimodal Video Generation via Camera Control
von: Yang, Xiaoda, et al.
Veröffentlicht: (2025)
von: Yang, Xiaoda, et al.
Veröffentlicht: (2025)
NextFlow: Unified Sequential Modeling Activates Multimodal Understanding and Generation
von: Zhang, Huichao, et al.
Veröffentlicht: (2026)
von: Zhang, Huichao, et al.
Veröffentlicht: (2026)
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation
von: Xiao, Yicheng, et al.
Veröffentlicht: (2025)
von: Xiao, Yicheng, et al.
Veröffentlicht: (2025)
Archon: A Unified Multimodal Model for Holistic Digital Human Generation
von: Bao, Chong, et al.
Veröffentlicht: (2026)
von: Bao, Chong, et al.
Veröffentlicht: (2026)
CodeBind: Decoupled Representation Learning for Multimodal Alignment with Unified Compositional Codebook
von: Chen, Zeyu, et al.
Veröffentlicht: (2026)
von: Chen, Zeyu, et al.
Veröffentlicht: (2026)
RaPD: Resolution-Agnostic Pixel Diffusion via Semantics-Enriched Implicit Representations
von: Ge, Yanhao, et al.
Veröffentlicht: (2026)
von: Ge, Yanhao, et al.
Veröffentlicht: (2026)
Bridging Pixels and Words: Mask-Aware Local Semantic Fusion for Multimodal Media Verification
von: Chen, Zizhao, et al.
Veröffentlicht: (2026)
von: Chen, Zizhao, et al.
Veröffentlicht: (2026)
UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding
von: Jiao, Yang, et al.
Veröffentlicht: (2025)
von: Jiao, Yang, et al.
Veröffentlicht: (2025)
UniCMs: A Unified Consistency Model For Efficient Multimodal Generation and Understanding
von: Xu, Chenkai, et al.
Veröffentlicht: (2025)
von: Xu, Chenkai, et al.
Veröffentlicht: (2025)
Ming-Flash-Omni: A Sparse, Unified Architecture for Multimodal Perception and Generation
von: AI, Inclusion, et al.
Veröffentlicht: (2025)
von: AI, Inclusion, et al.
Veröffentlicht: (2025)
PixelGen: Improving Pixel Diffusion with Perceptual Supervision
von: Ma, Zehong, et al.
Veröffentlicht: (2026)
von: Ma, Zehong, et al.
Veröffentlicht: (2026)
CAFe: Unifying Representation and Generation with Contrastive-Autoregressive Finetuning
von: Yu, Hao, et al.
Veröffentlicht: (2025)
von: Yu, Hao, et al.
Veröffentlicht: (2025)
From Skeletons to Pixels: Few-Shot Precise Event Spotting via Representation and Prediction Distillation
von: Yeoh, Zhong Han Ervin, et al.
Veröffentlicht: (2026)
von: Yeoh, Zhong Han Ervin, et al.
Veröffentlicht: (2026)
UniPose: A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing
von: Li, Yiheng, et al.
Veröffentlicht: (2024)
von: Li, Yiheng, et al.
Veröffentlicht: (2024)
UniFashion: A Unified Vision-Language Model for Multimodal Fashion Retrieval and Generation
von: Zhao, Xiangyu, et al.
Veröffentlicht: (2024)
von: Zhao, Xiangyu, et al.
Veröffentlicht: (2024)
Lumina-OmniLV: A Unified Multimodal Framework for General Low-Level Vision
von: Pu, Yuandong, et al.
Veröffentlicht: (2025)
von: Pu, Yuandong, et al.
Veröffentlicht: (2025)
PixelArena: A benchmark for Pixel-Precision Visual Intelligence
von: Liang, Feng, et al.
Veröffentlicht: (2025)
von: Liang, Feng, et al.
Veröffentlicht: (2025)
Pixel-Aligned Multi-View Generation with Depth Guided Decoder
von: Tang, Zhenggang, et al.
Veröffentlicht: (2024)
von: Tang, Zhenggang, et al.
Veröffentlicht: (2024)
L2P: Unlocking Latent Potential for Pixel Generation
von: Chen, Zhennan, et al.
Veröffentlicht: (2026)
von: Chen, Zhennan, et al.
Veröffentlicht: (2026)
PixelMan: Consistent Object Editing with Diffusion Models via Pixel Manipulation and Generation
von: Jiang, Liyao, et al.
Veröffentlicht: (2024)
von: Jiang, Liyao, et al.
Veröffentlicht: (2024)
SafePLUG: Empowering Multimodal LLMs with Pixel-Level Insight and Temporal Grounding for Traffic Accident Understanding
von: Sheng, Zihao, et al.
Veröffentlicht: (2025)
von: Sheng, Zihao, et al.
Veröffentlicht: (2025)
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
von: Wu, Chengyue, et al.
Veröffentlicht: (2024)
von: Wu, Chengyue, et al.
Veröffentlicht: (2024)
MergeVQ: A Unified Framework for Visual Generation and Representation with Disentangled Token Merging and Quantization
von: Li, Siyuan, et al.
Veröffentlicht: (2025)
von: Li, Siyuan, et al.
Veröffentlicht: (2025)
Humor in Pixels: Benchmarking Large Multimodal Models Understanding of Online Comics
von: Ryan, Yuriel, et al.
Veröffentlicht: (2025)
von: Ryan, Yuriel, et al.
Veröffentlicht: (2025)
Understanding and Harnessing Sparsity in Unified Multimodal Models
von: He, Shwai, et al.
Veröffentlicht: (2025)
von: He, Shwai, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
PixelBytes: Catching Unified Embedding for Multimodal Generation
von: Furfaro, Fabien
Veröffentlicht: (2024) -
Unified Multimodal Understanding via Byte-Pair Visual Encoding
von: Zhang, Wanpeng, et al.
Veröffentlicht: (2025) -
Enhancing Multimodal Unified Representations for Cross Modal Generalization
von: Huang, Hai, et al.
Veröffentlicht: (2024) -
From Pixels to Tokens: Byte-Pair Encoding on Quantized Visual Modalities
von: Zhang, Wanpeng, et al.
Veröffentlicht: (2024) -
UniPixel: Unified Object Referring and Segmentation for Pixel-Level Visual Reasoning
von: Liu, Ye, et al.
Veröffentlicht: (2025)