RefDecoder: Enhancing Visual Generation with Conditional Video Decoding
Fuente:
arXiv
Saved in:
| Main Authors: | Fan, Xiang, Wang, Yuheng, Fang, Bohan, Ren, Zhongzheng, Krishna, Ranjay |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
RefTok: Reference-Based Tokenization for Video Generation
by: Fan, Xiang, et al.
Published: (2025)
by: Fan, Xiang, et al.
Published: (2025)
Videoshop: Localized Semantic Video Editing with Noise-Extrapolated Diffusion Inversion
by: Fan, Xiang, et al.
Published: (2024)
by: Fan, Xiang, et al.
Published: (2024)
PhysGen: Rigid-Body Physics-Grounded Image-to-Video Generation
by: Liu, Shaowei, et al.
Published: (2024)
by: Liu, Shaowei, et al.
Published: (2024)
VMDT: Decoding the Trustworthiness of Video Foundation Models
by: Potter, Yujin, et al.
Published: (2025)
by: Potter, Yujin, et al.
Published: (2025)
MultiRef: Controllable Image Generation with Multiple Visual References
by: Chen, Ruoxi, et al.
Published: (2025)
by: Chen, Ruoxi, et al.
Published: (2025)
VGS-Decoding: Visual Grounding Score Guided Decoding for Hallucination Mitigation in Medical VLMs
by: Kolli, Govinda, et al.
Published: (2026)
by: Kolli, Govinda, et al.
Published: (2026)
Generate Any Scene: Scene Graph Driven Data Synthesis for Visual Generation Training
by: Gao, Ziqi, et al.
Published: (2024)
by: Gao, Ziqi, et al.
Published: (2024)
Unified Text-Image-to-Video Generation: A Training-Free Approach to Flexible Visual Conditioning
by: Lai, Bolin, et al.
Published: (2025)
by: Lai, Bolin, et al.
Published: (2025)
REGEN: Learning Compact Video Embedding with (Re-)Generative Decoder
by: Zhang, Yitian, et al.
Published: (2025)
by: Zhang, Yitian, et al.
Published: (2025)
Perception Tokens Enhance Visual Reasoning in Multimodal Language Models
by: Bigverdi, Mahtab, et al.
Published: (2024)
by: Bigverdi, Mahtab, et al.
Published: (2024)
Human-Aligned Image Models Improve Visual Decoding from the Brain
by: Rajabi, Nona, et al.
Published: (2025)
by: Rajabi, Nona, et al.
Published: (2025)
SEED: Towards More Accurate Semantic Evaluation for Visual Brain Decoding
by: Park, Juhyeon, et al.
Published: (2025)
by: Park, Juhyeon, et al.
Published: (2025)
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning
by: Piergiovanni, AJ, et al.
Published: (2024)
by: Piergiovanni, AJ, et al.
Published: (2024)
Enhancing Vision-Language Model Reliability with Uncertainty-Guided Dropout Decoding
by: Fang, Yixiong, et al.
Published: (2024)
by: Fang, Yixiong, et al.
Published: (2024)
High-Fidelity Text-to-Image Generation from Pre-Trained Vision-Language Models via Distribution-Conditioned Diffusion Decoding
by: Hong, Ji Woo, et al.
Published: (2026)
by: Hong, Ji Woo, et al.
Published: (2026)
Multimodal ELBO with Diffusion Decoders
by: Wesego, Daniel, et al.
Published: (2024)
by: Wesego, Daniel, et al.
Published: (2024)
MEIcoder: Decoding Visual Stimuli from Neural Activity by Leveraging Most Exciting Inputs
by: Sobotka, Jan, et al.
Published: (2025)
by: Sobotka, Jan, et al.
Published: (2025)
LD-DETR: Loop Decoder DEtection TRansformer for Video Moment Retrieval and Highlight Detection
by: Zhao, Pengcheng, et al.
Published: (2025)
by: Zhao, Pengcheng, et al.
Published: (2025)
ResNet-50 with Class Reweighting and Anatomy-Guided Temporal Decoding for Gastrointestinal Video Analysis
by: Imtiaz, Romil, et al.
Published: (2026)
by: Imtiaz, Romil, et al.
Published: (2026)
VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognition
by: Yadav, Tanush, et al.
Published: (2026)
by: Yadav, Tanush, et al.
Published: (2026)
LILAC: Long-sequence Incremental Low-latency Arbitrary Motion Stylization via Streaming VAE-Diffusion with Causal Decoding
by: Ren, Peng, et al.
Published: (2025)
by: Ren, Peng, et al.
Published: (2025)
Toward Lightweight and Fast Decoders for Diffusion Models in Image and Video Generation
by: Buzovkin, Alexey, et al.
Published: (2025)
by: Buzovkin, Alexey, et al.
Published: (2025)
Ultra-Efficient Decoding for End-to-End Neural Compression and Reconstruction
by: Rogers, Ethan G., et al.
Published: (2025)
by: Rogers, Ethan G., et al.
Published: (2025)
AFN: Adaptive Fusion Normalization via an Encoder-Decoder Framework
by: Zhou, Zikai, et al.
Published: (2023)
by: Zhou, Zikai, et al.
Published: (2023)
FissionVAE: Federated Non-IID Image Generation with Latent Space and Decoder Decomposition
by: Hu, Chen, et al.
Published: (2024)
by: Hu, Chen, et al.
Published: (2024)
Finding Shared Decodable Concepts and their Negations in the Brain
by: Efird, Cory, et al.
Published: (2024)
by: Efird, Cory, et al.
Published: (2024)
Rationale-Enhanced Decoding for Multi-modal Chain-of-Thought
by: Yamaguchi, Shin'ya, et al.
Published: (2025)
by: Yamaguchi, Shin'ya, et al.
Published: (2025)
Pre-trained Visual Dynamics Representations for Efficient Policy Learning
by: Luo, Hao, et al.
Published: (2024)
by: Luo, Hao, et al.
Published: (2024)
DeCLIP: Decoding CLIP representations for deepfake localization
by: Smeu, Stefan, et al.
Published: (2024)
by: Smeu, Stefan, et al.
Published: (2024)
Decoding Federated Learning: The FedNAM+ Conformal Revolution
by: Balija, Sree Bhargavi, et al.
Published: (2025)
by: Balija, Sree Bhargavi, et al.
Published: (2025)
Gradient-free Decoder Inversion in Latent Diffusion Models
by: Hong, Seongmin, et al.
Published: (2024)
by: Hong, Seongmin, et al.
Published: (2024)
Domain Independent SVM for Transfer Learning in Brain Decoding
by: Zhou, Shuo, et al.
Published: (2019)
by: Zhou, Shuo, et al.
Published: (2019)
Rethinking Encoder-Decoder Flow Through Shared Structures
by: Laboyrie, Frederik, et al.
Published: (2025)
by: Laboyrie, Frederik, et al.
Published: (2025)
FocalLens: Instruction Tuning Enables Zero-Shot Conditional Image Representations
by: Hsieh, Cheng-Yu, et al.
Published: (2025)
by: Hsieh, Cheng-Yu, et al.
Published: (2025)
Variational Flow Maps: Make Some Noise for One-Step Conditional Generation
by: Mammadov, Abbas, et al.
Published: (2026)
by: Mammadov, Abbas, et al.
Published: (2026)
Masked Generative Nested Transformers with Decode Time Scaling
by: Goyal, Sahil, et al.
Published: (2025)
by: Goyal, Sahil, et al.
Published: (2025)
When Rubrics Fail: Error Enumeration as Reward in Reference-Free RL Post-Training for Virtual Try-On
by: Ikezogwo, Wisdom, et al.
Published: (2026)
by: Ikezogwo, Wisdom, et al.
Published: (2026)
Graph-Based Captioning: Enhancing Visual Descriptions by Interconnecting Region Captions
by: Hsieh, Yu-Guan, et al.
Published: (2024)
by: Hsieh, Yu-Guan, et al.
Published: (2024)
Unified Spatio-Temporal Token Scoring for Efficient Video VLMs
by: Zhang, Jianrui, et al.
Published: (2026)
by: Zhang, Jianrui, et al.
Published: (2026)
PSA: Pyramid Sparse Attention for Efficient Video Understanding and Generation
by: Li, Xiaolong, et al.
Published: (2025)
by: Li, Xiaolong, et al.
Published: (2025)
Similar Items
-
RefTok: Reference-Based Tokenization for Video Generation
by: Fan, Xiang, et al.
Published: (2025) -
Videoshop: Localized Semantic Video Editing with Noise-Extrapolated Diffusion Inversion
by: Fan, Xiang, et al.
Published: (2024) -
PhysGen: Rigid-Body Physics-Grounded Image-to-Video Generation
by: Liu, Shaowei, et al.
Published: (2024) -
VMDT: Decoding the Trustworthiness of Video Foundation Models
by: Potter, Yujin, et al.
Published: (2025) -
MultiRef: Controllable Image Generation with Multiple Visual References
by: Chen, Ruoxi, et al.
Published: (2025)