Pushing the Boundaries of State Space Models for Image and Video Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hong, Yicong, Mai, Long, Yao, Yuan, Liu, Feng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
REGEN: Learning Compact Video Embedding with (Re-)Generative Decoder
von: Zhang, Yitian, et al.
Veröffentlicht: (2025)
von: Zhang, Yitian, et al.
Veröffentlicht: (2025)
Diffusion Transformer-to-Mamba Distillation for High-Resolution Image Generation
von: Yao, Yuan, et al.
Veröffentlicht: (2025)
von: Yao, Yuan, et al.
Veröffentlicht: (2025)
Progressive Autoregressive Video Diffusion Models
von: Xie, Desai, et al.
Veröffentlicht: (2024)
von: Xie, Desai, et al.
Veröffentlicht: (2024)
Rethinking Training Dynamics in Scale-wise Autoregressive Generation
von: Zhou, Gengze, et al.
Veröffentlicht: (2025)
von: Zhou, Gengze, et al.
Veröffentlicht: (2025)
Pushing Trade-Off Boundaries: Compact yet Effective Remote Sensing Change Detection
von: Xu, Luosheng, et al.
Veröffentlicht: (2025)
von: Xu, Luosheng, et al.
Veröffentlicht: (2025)
Skews in the Phenomenon Space Hinder Generalization in Text-to-Image Generation
von: Chang, Yingshan, et al.
Veröffentlicht: (2024)
von: Chang, Yingshan, et al.
Veröffentlicht: (2024)
AnchorDS: Anchoring Dynamic Sources for Semantically Consistent Text-to-3D Generation
von: Zhu, Jiayin, et al.
Veröffentlicht: (2025)
von: Zhu, Jiayin, et al.
Veröffentlicht: (2025)
Edify Image: High-Quality Image Generation with Pixel Space Laplacian Diffusion Models
von: NVIDIA, et al.
Veröffentlicht: (2024)
von: NVIDIA, et al.
Veröffentlicht: (2024)
GALOT: Generative Active Learning via Optimizable Zero-shot Text-to-image Generation
von: Hong, Hanbin, et al.
Veröffentlicht: (2024)
von: Hong, Hanbin, et al.
Veröffentlicht: (2024)
LRM: Large Reconstruction Model for Single Image to 3D
von: Hong, Yicong, et al.
Veröffentlicht: (2023)
von: Hong, Yicong, et al.
Veröffentlicht: (2023)
Learning Enriched Features via Selective State Spaces Model for Efficient Image Deblurring
von: Gao, Hu, et al.
Veröffentlicht: (2024)
von: Gao, Hu, et al.
Veröffentlicht: (2024)
On Structured State-Space Duality
von: Hu, Jerry Yao-Chieh, et al.
Veröffentlicht: (2025)
von: Hu, Jerry Yao-Chieh, et al.
Veröffentlicht: (2025)
SoliReward: Mitigating Susceptibility to Reward Hacking and Annotation Noise in Video Generation Reward Models
von: Lian, Jiesong, et al.
Veröffentlicht: (2025)
von: Lian, Jiesong, et al.
Veröffentlicht: (2025)
Physics-Driven Spatiotemporal Modeling for AI-Generated Video Detection
von: Zhang, Shuhai, et al.
Veröffentlicht: (2025)
von: Zhang, Shuhai, et al.
Veröffentlicht: (2025)
State Space Models for Event Cameras
von: Zubić, Nikola, et al.
Veröffentlicht: (2024)
von: Zubić, Nikola, et al.
Veröffentlicht: (2024)
Wan-Image: Pushing the Boundaries of Generative Visual Intelligence
von: Mao, Chaojie, et al.
Veröffentlicht: (2026)
von: Mao, Chaojie, et al.
Veröffentlicht: (2026)
SAME: Learning Generic Language-Guided Visual Navigation with State-Adaptive Mixture of Experts
von: Zhou, Gengze, et al.
Veröffentlicht: (2024)
von: Zhou, Gengze, et al.
Veröffentlicht: (2024)
Semantic-Guided Generative Image Augmentation Method with Diffusion Models for Image Classification
von: Li, Bohan, et al.
Veröffentlicht: (2023)
von: Li, Bohan, et al.
Veröffentlicht: (2023)
Gradient-Guided Exploration of Generative Model's Latent Space for Controlled Iris Image Augmentations
von: Mitcheff, Mahsa, et al.
Veröffentlicht: (2025)
von: Mitcheff, Mahsa, et al.
Veröffentlicht: (2025)
Video Reconstruction using Diffusion-based Image-to-Video Generation with Trajectory Guidance
von: Bompai, Stelio, et al.
Veröffentlicht: (2026)
von: Bompai, Stelio, et al.
Veröffentlicht: (2026)
iVideoGPT: Interactive VideoGPTs are Scalable World Models
von: Wu, Jialong, et al.
Veröffentlicht: (2024)
von: Wu, Jialong, et al.
Veröffentlicht: (2024)
Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model
von: Zhu, Lianghui, et al.
Veröffentlicht: (2024)
von: Zhu, Lianghui, et al.
Veröffentlicht: (2024)
Boundary-Constrained Diffusion Models for Floorplan Generation: Balancing Realism and Diversity
von: Stoppani, Leonardo, et al.
Veröffentlicht: (2026)
von: Stoppani, Leonardo, et al.
Veröffentlicht: (2026)
Contextualized Diffusion Models for Text-Guided Image and Video Generation
von: Yang, Ling, et al.
Veröffentlicht: (2024)
von: Yang, Ling, et al.
Veröffentlicht: (2024)
Fleximo: Towards Flexible Text-to-Human Motion Video Generation
von: Zhang, Yuhang, et al.
Veröffentlicht: (2024)
von: Zhang, Yuhang, et al.
Veröffentlicht: (2024)
VINCIE: Unlocking In-context Image Editing from Video
von: Qu, Leigang, et al.
Veröffentlicht: (2025)
von: Qu, Leigang, et al.
Veröffentlicht: (2025)
MamBEV: Enabling State Space Models to Learn Birds-Eye-View Representations
von: Ke, Hongyu, et al.
Veröffentlicht: (2025)
von: Ke, Hongyu, et al.
Veröffentlicht: (2025)
Zero-Shot Video Restoration and Enhancement Using Pre-Trained Image Diffusion Model
von: Cao, Cong, et al.
Veröffentlicht: (2024)
von: Cao, Cong, et al.
Veröffentlicht: (2024)
Latent Forcing: Reordering the Diffusion Trajectory for Pixel-Space Image Generation
von: Baade, Alan, et al.
Veröffentlicht: (2026)
von: Baade, Alan, et al.
Veröffentlicht: (2026)
FREPix: Frequency-Heterogeneous Flow Matching for Pixel-Space Image Generation
von: Lin, Mingfeng, et al.
Veröffentlicht: (2026)
von: Lin, Mingfeng, et al.
Veröffentlicht: (2026)
Pretrained Image-Text Models are Secretly Video Captioners
von: Zhang, Chunhui, et al.
Veröffentlicht: (2025)
von: Zhang, Chunhui, et al.
Veröffentlicht: (2025)
Progressive Compositionality in Text-to-Image Generative Models
von: Han, Evans Xu, et al.
Veröffentlicht: (2024)
von: Han, Evans Xu, et al.
Veröffentlicht: (2024)
SoFlow: Solution Flow Models for One-Step Generative Modeling
von: Luo, Tianze, et al.
Veröffentlicht: (2025)
von: Luo, Tianze, et al.
Veröffentlicht: (2025)
Multi-State-Action Tokenisation in Decision Transformers for Multi-Discrete Action Spaces
von: Moodley, Perusha, et al.
Veröffentlicht: (2024)
von: Moodley, Perusha, et al.
Veröffentlicht: (2024)
Understanding Image2Video Domain Shift in Food Segmentation: An Instance-level Analysis on Apples
von: Park, Keonvin, et al.
Veröffentlicht: (2026)
von: Park, Keonvin, et al.
Veröffentlicht: (2026)
DSS-GAN: Directional State Space GAN with Mamba backbone for Class-Conditional Image Synthesis
von: Ogonowski, Aleksander, et al.
Veröffentlicht: (2026)
von: Ogonowski, Aleksander, et al.
Veröffentlicht: (2026)
Shedding the Bits: Pushing the Boundaries of Quantization with Minifloats on FPGAs
von: Aggarwal, Shivam, et al.
Veröffentlicht: (2023)
von: Aggarwal, Shivam, et al.
Veröffentlicht: (2023)
Boosting Text-to-Image Diffusion Models via Core Token Attention-Based Seed Selection
von: Zhang, Yunzhe, et al.
Veröffentlicht: (2026)
von: Zhang, Yunzhe, et al.
Veröffentlicht: (2026)
TABNet: A Triplet Augmentation Self-Recovery Framework with Boundary-Aware Pseudo-Labels for Medical Image Segmentation
von: Zhang, Peilin, et al.
Veröffentlicht: (2025)
von: Zhang, Peilin, et al.
Veröffentlicht: (2025)
Direct Motion Models for Assessing Generated Videos
von: Allen, Kelsey, et al.
Veröffentlicht: (2025)
von: Allen, Kelsey, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
REGEN: Learning Compact Video Embedding with (Re-)Generative Decoder
von: Zhang, Yitian, et al.
Veröffentlicht: (2025) -
Diffusion Transformer-to-Mamba Distillation for High-Resolution Image Generation
von: Yao, Yuan, et al.
Veröffentlicht: (2025) -
Progressive Autoregressive Video Diffusion Models
von: Xie, Desai, et al.
Veröffentlicht: (2024) -
Rethinking Training Dynamics in Scale-wise Autoregressive Generation
von: Zhou, Gengze, et al.
Veröffentlicht: (2025) -
Pushing Trade-Off Boundaries: Compact yet Effective Remote Sensing Change Detection
von: Xu, Luosheng, et al.
Veröffentlicht: (2025)