ReDDiT: Rehashing Noise for Discrete Visual Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Ma, Tianren, Zhang, Xiaosong, Yang, Boyu, Feng, Junlan, Ye, Qixiang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ClawMachine: Learning to Fetch Visual Tokens for Referential Comprehension
di: Ma, Tianren, et al.
Pubblicazione: (2024)
di: Ma, Tianren, et al.
Pubblicazione: (2024)
AceTone: Bridging Words and Colors for Conditional Image Grading
di: Ma, Tianren, et al.
Pubblicazione: (2026)
di: Ma, Tianren, et al.
Pubblicazione: (2026)
DDiT: Dynamic Patch Scheduling for Efficient Diffusion Transformers
di: Kim, Dahye, et al.
Pubblicazione: (2026)
di: Kim, Dahye, et al.
Pubblicazione: (2026)
ChatterBox: Multi-round Multimodal Referring and Grounding
di: Tian, Yunjie, et al.
Pubblicazione: (2024)
di: Tian, Yunjie, et al.
Pubblicazione: (2024)
Preserving Silent Features for Domain Generalization
di: Zhao, Chujie, et al.
Pubblicazione: (2024)
di: Zhao, Chujie, et al.
Pubblicazione: (2024)
VideoAnchor: Reinforcing Subspace-Structured Visual Cues for Coherent Visual-Spatial Reasoning
di: Wang, Zhaozhi, et al.
Pubblicazione: (2025)
di: Wang, Zhaozhi, et al.
Pubblicazione: (2025)
GeoTikzBridge: Advancing Multimodal Code Generation for Geometric Perception and Reasoning
di: Sun, Jiayin, et al.
Pubblicazione: (2026)
di: Sun, Jiayin, et al.
Pubblicazione: (2026)
Cubic Discrete Diffusion: Discrete Visual Generation on High-Dimensional Representation Tokens
di: Wang, Yuqing, et al.
Pubblicazione: (2026)
di: Wang, Yuqing, et al.
Pubblicazione: (2026)
Artemis: Towards Referential Understanding in Complex Videos
di: Qiu, Jihao, et al.
Pubblicazione: (2024)
di: Qiu, Jihao, et al.
Pubblicazione: (2024)
Rethinking Sampling Strategies for Unsupervised Person Re-identification
di: Han, Xumeng, et al.
Pubblicazione: (2021)
di: Han, Xumeng, et al.
Pubblicazione: (2021)
Bridging Continuous and Discrete Tokens for Autoregressive Visual Generation
di: Wang, Yuqing, et al.
Pubblicazione: (2025)
di: Wang, Yuqing, et al.
Pubblicazione: (2025)
Correspondence-Guided SfM-Free 3D Gaussian Splatting for NVS
di: Sun, Wei, et al.
Pubblicazione: (2024)
di: Sun, Wei, et al.
Pubblicazione: (2024)
Toward High-Fidelity Visual Reconstruction: From EEG-Based Conditioned Generation to Joint-Modal Guided Rebuilding
di: Gong, Zhijian, et al.
Pubblicazione: (2026)
di: Gong, Zhijian, et al.
Pubblicazione: (2026)
Thinking with Images via Self-Calling Agent
di: Yang, Wenxi, et al.
Pubblicazione: (2025)
di: Yang, Wenxi, et al.
Pubblicazione: (2025)
S&D Messenger: Exchanging Semantic and Domain Knowledge for Generic Semi-Supervised Medical Image Segmentation
di: Zhang, Qixiang, et al.
Pubblicazione: (2024)
di: Zhang, Qixiang, et al.
Pubblicazione: (2024)
Delving Deep into Semantic Relation Distillation
di: Yan, Zhaoyi, et al.
Pubblicazione: (2025)
di: Yan, Zhaoyi, et al.
Pubblicazione: (2025)
Mixture of Physical Priors Adapter for Parameter-Efficient Fine-Tuning
di: Wang, Zhaozhi, et al.
Pubblicazione: (2024)
di: Wang, Zhaozhi, et al.
Pubblicazione: (2024)
Reconstruction as a Bridge for Event-Based Visual Question Answering
di: Lou, Hanyue, et al.
Pubblicazione: (2025)
di: Lou, Hanyue, et al.
Pubblicazione: (2025)
UniEval: Unified Holistic Evaluation for Unified Multimodal Understanding and Generation
di: Li, Yi, et al.
Pubblicazione: (2025)
di: Li, Yi, et al.
Pubblicazione: (2025)
Instance-level Visual Active Tracking with Occlusion-Aware Planning
di: Sun, Haowei, et al.
Pubblicazione: (2026)
di: Sun, Haowei, et al.
Pubblicazione: (2026)
Tuning-free Instruction-based Video Editing Via Structural Noise Initialization and Guidance
di: Wu, Song, et al.
Pubblicazione: (2026)
di: Wu, Song, et al.
Pubblicazione: (2026)
Semantic-Enriched Latent Visual Reasoning
di: Xu, Tianrun, et al.
Pubblicazione: (2026)
di: Xu, Tianrun, et al.
Pubblicazione: (2026)
X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again
di: Geng, Zigang, et al.
Pubblicazione: (2025)
di: Geng, Zigang, et al.
Pubblicazione: (2025)
FAIL: Flow Matching Adversarial Imitation Learning for Image Generation
di: Ma, Yeyao, et al.
Pubblicazione: (2026)
di: Ma, Yeyao, et al.
Pubblicazione: (2026)
Re-Aligning Language to Visual Objects with an Agentic Workflow
di: Chen, Yuming, et al.
Pubblicazione: (2025)
di: Chen, Yuming, et al.
Pubblicazione: (2025)
SGG-R$^{\rm 3}$: From Next-Token Prediction to End-to-End Unbiased Scene Graph Generation
di: Feng, Jiaye, et al.
Pubblicazione: (2026)
di: Feng, Jiaye, et al.
Pubblicazione: (2026)
Towards All-in-One Medical Image Re-Identification
di: Tian, Yuan, et al.
Pubblicazione: (2025)
di: Tian, Yuan, et al.
Pubblicazione: (2025)
VMamba: Visual State Space Model
di: Liu, Yue, et al.
Pubblicazione: (2024)
di: Liu, Yue, et al.
Pubblicazione: (2024)
Expandable Residual Approximation for Knowledge Distillation
di: Yan, Zhaoyi, et al.
Pubblicazione: (2025)
di: Yan, Zhaoyi, et al.
Pubblicazione: (2025)
Spatial Transform Decoupling for Oriented Object Detection
di: Yu, Hongtian, et al.
Pubblicazione: (2023)
di: Yu, Hongtian, et al.
Pubblicazione: (2023)
Self-supervised Feature-Gate Coupling for Dynamic Network Pruning
di: Shi, Mengnan, et al.
Pubblicazione: (2021)
di: Shi, Mengnan, et al.
Pubblicazione: (2021)
CC-Diff: Enhancing Contextual Coherence in Remote Sensing Image Synthesis
di: Zhang, Mu, et al.
Pubblicazione: (2024)
di: Zhang, Mu, et al.
Pubblicazione: (2024)
Multi-Cue Adaptive Visual Token Pruning for Large Vision-Language Models
di: Luan, Bozhi, et al.
Pubblicazione: (2025)
di: Luan, Bozhi, et al.
Pubblicazione: (2025)
A Cognitive Process-Inspired Architecture for Subject-Agnostic Brain Visual Decoding
di: Lu, Jingyu, et al.
Pubblicazione: (2025)
di: Lu, Jingyu, et al.
Pubblicazione: (2025)
D$^{2}$-VPR: A Parameter-efficient Visual-foundation-model-based Visual Place Recognition Method via Knowledge Distillation and Deformable Aggregation
di: Zhang, Zheyuan, et al.
Pubblicazione: (2025)
di: Zhang, Zheyuan, et al.
Pubblicazione: (2025)
PAROAttention: Pattern-Aware ReOrdering for Efficient Sparse and Quantized Attention in Visual Generation Models
di: Zhao, Tianchen, et al.
Pubblicazione: (2025)
di: Zhao, Tianchen, et al.
Pubblicazione: (2025)
YOLOv12: Attention-Centric Real-Time Object Detectors
di: Tian, Yunjie, et al.
Pubblicazione: (2025)
di: Tian, Yunjie, et al.
Pubblicazione: (2025)
OpenView: Empowering MLLMs with Out-of-view VQA
di: Chen, Qixiang, et al.
Pubblicazione: (2025)
di: Chen, Qixiang, et al.
Pubblicazione: (2025)
Ray Denoising: Depth-aware Hard Negative Sampling for Multi-view 3D Object Detection
di: Liu, Feng, et al.
Pubblicazione: (2024)
di: Liu, Feng, et al.
Pubblicazione: (2024)
VAEVQ: Enhancing Discrete Visual Tokenization through Variational Modeling
di: Yang, Sicheng, et al.
Pubblicazione: (2025)
di: Yang, Sicheng, et al.
Pubblicazione: (2025)
Documenti analoghi
-
ClawMachine: Learning to Fetch Visual Tokens for Referential Comprehension
di: Ma, Tianren, et al.
Pubblicazione: (2024) -
AceTone: Bridging Words and Colors for Conditional Image Grading
di: Ma, Tianren, et al.
Pubblicazione: (2026) -
DDiT: Dynamic Patch Scheduling for Efficient Diffusion Transformers
di: Kim, Dahye, et al.
Pubblicazione: (2026) -
ChatterBox: Multi-round Multimodal Referring and Grounding
di: Tian, Yunjie, et al.
Pubblicazione: (2024) -
Preserving Silent Features for Domain Generalization
di: Zhao, Chujie, et al.
Pubblicazione: (2024)