REGEN: Learning Compact Video Embedding with (Re-)Generative Decoder
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Yitian, Mai, Long, Mahapatra, Aniruddha, Bourgin, David, Hong, Yicong, Casebeer, Jonah, Liu, Feng, Fu, Yun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Progressive Growing of Video Tokenizers for Temporally Compact Latent Spaces
by: Mahapatra, Aniruddha, et al.
Published: (2025)
by: Mahapatra, Aniruddha, et al.
Published: (2025)
V2M-Zero: Zero-Pair Time-Aligned Video-to-Music Generation
by: Lin, Yan-Bo, et al.
Published: (2026)
by: Lin, Yan-Bo, et al.
Published: (2026)
DreamLoop: Controllable Cinemagraph Generation from a Single Photograph
by: Mahapatra, Aniruddha, et al.
Published: (2026)
by: Mahapatra, Aniruddha, et al.
Published: (2026)
Pushing the Boundaries of State Space Models for Image and Video Generation
by: Hong, Yicong, et al.
Published: (2025)
by: Hong, Yicong, et al.
Published: (2025)
MotionCanvas: Cinematic Shot Design with Controllable Image-to-Video Generation
by: Xing, Jinbo, et al.
Published: (2025)
by: Xing, Jinbo, et al.
Published: (2025)
TRACE: Object Motion Editing in Videos with First-Frame Trajectory Guidance
by: Phung, Quynh, et al.
Published: (2026)
by: Phung, Quynh, et al.
Published: (2026)
Diffusion Transformer-to-Mamba Distillation for High-Resolution Image Generation
by: Yao, Yuan, et al.
Published: (2025)
by: Yao, Yuan, et al.
Published: (2025)
TokenDial: Continuous Attribute Control in Text-to-Video via Spatiotemporal Token Offsets
by: Liu, Zhixuan, et al.
Published: (2026)
by: Liu, Zhixuan, et al.
Published: (2026)
REGEN: Real-Time Photorealism Enhancement in Games via a Dual-Stage Generative Network Framework
by: Pasios, Stefanos, et al.
Published: (2025)
by: Pasios, Stefanos, et al.
Published: (2025)
Generating, Fast and Slow: Scalable Parallel Video Generation with Video Interface Networks
by: Dedhia, Bhishma, et al.
Published: (2025)
by: Dedhia, Bhishma, et al.
Published: (2025)
DynamicEval: Rethinking Evaluation for Dynamic Text-to-Video Synthesis
by: Babu, Nithin C., et al.
Published: (2025)
by: Babu, Nithin C., et al.
Published: (2025)
MTS-DMAE: Dual-Masked Autoencoder for Unsupervised Multivariate Time Series Representation Learning
by: Xu, Yi, et al.
Published: (2025)
by: Xu, Yi, et al.
Published: (2025)
Causality Model for Semantic Understanding on Videos
by: Yicong, Li
Published: (2025)
by: Yicong, Li
Published: (2025)
On the Content Bias in Fréchet Video Distance
by: Ge, Songwei, et al.
Published: (2024)
by: Ge, Songwei, et al.
Published: (2024)
Video-Guided Foley Sound Generation with Multimodal Controls
by: Chen, Ziyang, et al.
Published: (2024)
by: Chen, Ziyang, et al.
Published: (2024)
Fine-T2I: An Open, Large-Scale, and Diverse Dataset for High-Quality T2I Fine-Tuning
by: Ma, Xu, et al.
Published: (2026)
by: Ma, Xu, et al.
Published: (2026)
Outlier-Aware Post-Training Quantization for Image Super-Resolution
by: Wang, Hailing, et al.
Published: (2025)
by: Wang, Hailing, et al.
Published: (2025)
Don't Judge by the Look: Towards Motion Coherent Video Representation
by: Zhang, Yitian, et al.
Published: (2024)
by: Zhang, Yitian, et al.
Published: (2024)
CoRe: Context-Regularized Text Embedding Learning for Text-to-Image Personalization
by: Wu, Feize, et al.
Published: (2024)
by: Wu, Feize, et al.
Published: (2024)
Distorted or Fabricated? A Survey on Hallucination in Video LLMs
by: Huang, Yiyang, et al.
Published: (2026)
by: Huang, Yiyang, et al.
Published: (2026)
Clapper: Compact Learning and Video Representation in VLMs
by: Kong, Lingyu, et al.
Published: (2025)
by: Kong, Lingyu, et al.
Published: (2025)
ReSpec: Relevance and Specificity Grounded Online Filtering for Learning on Video-Text Data Streams
by: Kim, Chris Dongjoo, et al.
Published: (2025)
by: Kim, Chris Dongjoo, et al.
Published: (2025)
EA-Swin: An Embedding-Agnostic Swin Transformer for AI-Generated Video Detection
by: Mai, Hung, et al.
Published: (2026)
by: Mai, Hung, et al.
Published: (2026)
VPG: Visual Prefix Guidance for Autoregressive Image and Video Generation
by: Liao, Xinyao, et al.
Published: (2026)
by: Liao, Xinyao, et al.
Published: (2026)
Making Your Dreams A Reality: Decoding the Dreams into a Coherent Video Story from fMRI Signals
by: Fu, Yanwei, et al.
Published: (2025)
by: Fu, Yanwei, et al.
Published: (2025)
ReNeg: Learning Negative Embedding with Reward Guidance
by: Li, Xiaomin, et al.
Published: (2024)
by: Li, Xiaomin, et al.
Published: (2024)
Progressive Autoregressive Video Diffusion Models
by: Xie, Desai, et al.
Published: (2024)
by: Xie, Desai, et al.
Published: (2024)
LineCounter: Learning Handwritten Text Line Segmentation by Counting
by: Li, Deng, et al.
Published: (2021)
by: Li, Deng, et al.
Published: (2021)
Multi-sentence Video Grounding for Long Video Generation
by: Feng, Wei, et al.
Published: (2024)
by: Feng, Wei, et al.
Published: (2024)
The Indra Representation Hypothesis for Multimodal Alignment
by: Lu, Jianglin, et al.
Published: (2026)
by: Lu, Jianglin, et al.
Published: (2026)
DecoFuse: Decomposing and Fusing the "What", "Where", and "How" for Brain-Inspired fMRI-to-Video Decoding
by: Li, Chong, et al.
Published: (2025)
by: Li, Chong, et al.
Published: (2025)
Joint2Human: High-quality 3D Human Generation via Compact Spherical Embedding of 3D Joints
by: Zhang, Muxin, et al.
Published: (2023)
by: Zhang, Muxin, et al.
Published: (2023)
GaussianVideo: Efficient Video Representation via Hierarchical Gaussian Splatting
by: Bond, Andrew, et al.
Published: (2025)
by: Bond, Andrew, et al.
Published: (2025)
Detection of Customer Interested Garments in Surveillance Video using Computer Vision
by: Ijjina, Earnest Paul, et al.
Published: (2025)
by: Ijjina, Earnest Paul, et al.
Published: (2025)
JEPA-VLA: Video Predictive Embedding is Needed for VLA Models
by: Miao, Shangchen, et al.
Published: (2026)
by: Miao, Shangchen, et al.
Published: (2026)
VGS-Decoding: Visual Grounding Score Guided Decoding for Hallucination Mitigation in Medical VLMs
by: Kolli, Govinda, et al.
Published: (2026)
by: Kolli, Govinda, et al.
Published: (2026)
Synthetic-To-Real Video Person Re-ID
by: Zhang, Xiangqun, et al.
Published: (2024)
by: Zhang, Xiangqun, et al.
Published: (2024)
Vid-Freeze: Protecting Images from Malicious Image-to-Video Generation via Temporal Freezing
by: Chowdhury, Rohit, et al.
Published: (2025)
by: Chowdhury, Rohit, et al.
Published: (2025)
ReCamMaster: Camera-Controlled Generative Rendering from A Single Video
by: Bai, Jianhong, et al.
Published: (2025)
by: Bai, Jianhong, et al.
Published: (2025)
Deep Learning based approach to detect Customer Age, Gender and Expression in Surveillance Video
by: Ijjina, Earnest Paul, et al.
Published: (2025)
by: Ijjina, Earnest Paul, et al.
Published: (2025)
Similar Items
-
Progressive Growing of Video Tokenizers for Temporally Compact Latent Spaces
by: Mahapatra, Aniruddha, et al.
Published: (2025) -
V2M-Zero: Zero-Pair Time-Aligned Video-to-Music Generation
by: Lin, Yan-Bo, et al.
Published: (2026) -
DreamLoop: Controllable Cinemagraph Generation from a Single Photograph
by: Mahapatra, Aniruddha, et al.
Published: (2026) -
Pushing the Boundaries of State Space Models for Image and Video Generation
by: Hong, Yicong, et al.
Published: (2025) -
MotionCanvas: Cinematic Shot Design with Controllable Image-to-Video Generation
by: Xing, Jinbo, et al.
Published: (2025)