Pyramidal Patchification Flow for Visual Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Hui, Chen, Baoyou, Zhang, Liwei, Li, Jiaye, Wang, Jingdong, Zhu, Siyu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Head-wise Adaptive Rotary Positional Encoding for Fine-Grained Image Generation
von: Li, Jiaye, et al.
Veröffentlicht: (2025)
von: Li, Jiaye, et al.
Veröffentlicht: (2025)
DPAR: Dynamic Patchification for Efficient Autoregressive Visual Generation
von: Srivastava, Divyansh, et al.
Veröffentlicht: (2025)
von: Srivastava, Divyansh, et al.
Veröffentlicht: (2025)
Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation
von: Xu, Mingwang, et al.
Veröffentlicht: (2024)
von: Xu, Mingwang, et al.
Veröffentlicht: (2024)
Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference Distillation
von: Li, Chunyu, et al.
Veröffentlicht: (2026)
von: Li, Chunyu, et al.
Veröffentlicht: (2026)
OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation
von: Li, Hui, et al.
Veröffentlicht: (2024)
von: Li, Hui, et al.
Veröffentlicht: (2024)
BARD: Bridging AutoRegressive and Diffusion Vision-Language Models Via Highly Efficient Progressive Block Merging and Stage-Wise Distillation
von: Chen, Baoyou, et al.
Veröffentlicht: (2026)
von: Chen, Baoyou, et al.
Veröffentlicht: (2026)
SlotMemory: Object-Centric KV Memory for Streaming Long-Video Generation
von: Dou, Weijia, et al.
Veröffentlicht: (2026)
von: Dou, Weijia, et al.
Veröffentlicht: (2026)
MixFlow Training: Alleviating Exposure Bias with Slowed Interpolation Mixture
von: Li, Hui, et al.
Veröffentlicht: (2025)
von: Li, Hui, et al.
Veröffentlicht: (2025)
Scaling Laws in Patchification: An Image Is Worth 50,176 Tokens And More
von: Wang, Feng, et al.
Veröffentlicht: (2025)
von: Wang, Feng, et al.
Veröffentlicht: (2025)
Portrait3D: Text-Guided High-Quality 3D Portrait Generation Using Pyramid Representation and GANs Prior
von: Wu, Yiqian, et al.
Veröffentlicht: (2024)
von: Wu, Yiqian, et al.
Veröffentlicht: (2024)
Exploring Effective Factors for Improving Visual In-Context Learning
von: Sun, Yanpeng, et al.
Veröffentlicht: (2023)
von: Sun, Yanpeng, et al.
Veröffentlicht: (2023)
Pyramid Forcing: Head-Aware Pyramid KV Cache Policy for High-Quality Long Video Generation
von: Chen, Jiayu, et al.
Veröffentlicht: (2026)
von: Chen, Jiayu, et al.
Veröffentlicht: (2026)
Pyramidal Flow Matching for Efficient Video Generative Modeling
von: Jin, Yang, et al.
Veröffentlicht: (2024)
von: Jin, Yang, et al.
Veröffentlicht: (2024)
Prompt Reinjection: Alleviating Prompt Forgetting in Multimodal Diffusion Transformers
von: Yao, Yuxuan, et al.
Veröffentlicht: (2026)
von: Yao, Yuxuan, et al.
Veröffentlicht: (2026)
Hallo2: Long-Duration and High-Resolution Audio-Driven Portrait Image Animation
von: Cui, Jiahao, et al.
Veröffentlicht: (2024)
von: Cui, Jiahao, et al.
Veröffentlicht: (2024)
InjectFlow: Weak Guides Strong via Orthogonal Injection for Flow Matching
von: Wang, Dayu, et al.
Veröffentlicht: (2026)
von: Wang, Dayu, et al.
Veröffentlicht: (2026)
Panther: Illuminate the Sight of Multimodal LLMs with Instruction-Guided Visual Prompts
von: Li, Honglin, et al.
Veröffentlicht: (2024)
von: Li, Honglin, et al.
Veröffentlicht: (2024)
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding
von: Chen, Zhanpeng, et al.
Veröffentlicht: (2025)
von: Chen, Zhanpeng, et al.
Veröffentlicht: (2025)
Pyramid Diffusion for Fine 3D Large Scene Generation
von: Liu, Yuheng, et al.
Veröffentlicht: (2023)
von: Liu, Yuheng, et al.
Veröffentlicht: (2023)
VP-LLM: Text-Driven 3D Volume Completion with Large Language Models through Patchification
von: Liu, Jianmeng, et al.
Veröffentlicht: (2024)
von: Liu, Jianmeng, et al.
Veröffentlicht: (2024)
Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding
von: Wang, Zhaokai, et al.
Veröffentlicht: (2025)
von: Wang, Zhaokai, et al.
Veröffentlicht: (2025)
Beyond Thinking: Imagining in 360$^\circ$ for Humanoid Visual Search
von: Zhang, Jingdong, et al.
Veröffentlicht: (2026)
von: Zhang, Jingdong, et al.
Veröffentlicht: (2026)
Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation
von: Zhang, Jianyu, et al.
Veröffentlicht: (2024)
von: Zhang, Jianyu, et al.
Veröffentlicht: (2024)
VRP-SAM: SAM with Visual Reference Prompt
von: Sun, Yanpeng, et al.
Veröffentlicht: (2024)
von: Sun, Yanpeng, et al.
Veröffentlicht: (2024)
DPFlow: Adaptive Optical Flow Estimation with a Dual-Pyramid Framework
von: Morimitsu, Henrique, et al.
Veröffentlicht: (2025)
von: Morimitsu, Henrique, et al.
Veröffentlicht: (2025)
PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction
von: Xing, Long, et al.
Veröffentlicht: (2024)
von: Xing, Long, et al.
Veröffentlicht: (2024)
Parameter-Inverted Image Pyramid Networks
von: Zhu, Xizhou, et al.
Veröffentlicht: (2024)
von: Zhu, Xizhou, et al.
Veröffentlicht: (2024)
PyramidMamba: Rethinking Pyramid Feature Fusion with Selective Space State Model for Semantic Segmentation of Remote Sensing Imagery
von: Wang, Libo, et al.
Veröffentlicht: (2024)
von: Wang, Libo, et al.
Veröffentlicht: (2024)
Guiding Visual Autoregressive Models through Spectrum Weakening
von: Wang, Chaoyang, et al.
Veröffentlicht: (2025)
von: Wang, Chaoyang, et al.
Veröffentlicht: (2025)
UniFlow: A Unified Pixel Flow Tokenizer for Visual Understanding and Generation
von: Yue, Zhengrong, et al.
Veröffentlicht: (2025)
von: Yue, Zhengrong, et al.
Veröffentlicht: (2025)
Enhancing Descriptive Captions with Visual Attributes for Multimodal Perception
von: Sun, Yanpeng, et al.
Veröffentlicht: (2024)
von: Sun, Yanpeng, et al.
Veröffentlicht: (2024)
Defective Convolutional Networks
von: Luo, Tiange, et al.
Veröffentlicht: (2019)
von: Luo, Tiange, et al.
Veröffentlicht: (2019)
VoxelSplat: Dynamic Gaussian Splatting as an Effective Loss for Occupancy and Flow Prediction
von: Zhu, Ziyue, et al.
Veröffentlicht: (2025)
von: Zhu, Ziyue, et al.
Veröffentlicht: (2025)
Rethinking Features-Fused-Pyramid-Neck for Object Detection
von: Li, Hulin
Veröffentlicht: (2025)
von: Li, Hulin
Veröffentlicht: (2025)
Deforming Videos to Masks: Flow Matching for Referring Video Segmentation
von: Wang, Zanyi, et al.
Veröffentlicht: (2025)
von: Wang, Zanyi, et al.
Veröffentlicht: (2025)
AuralSAM2: Enabling SAM2 Hear Through Pyramid Audio-Visual Feature Prompting
von: Liu, Yuyuan, et al.
Veröffentlicht: (2025)
von: Liu, Yuyuan, et al.
Veröffentlicht: (2025)
Learning Inverse Laplacian Pyramid for Progressive Depth Completion
von: Wang, Kun, et al.
Veröffentlicht: (2025)
von: Wang, Kun, et al.
Veröffentlicht: (2025)
Perception Before Reasoning: Two-Stage Reinforcement Learning for Visual Reasoning in Vision-Language Models
von: Chen, Yan, et al.
Veröffentlicht: (2025)
von: Chen, Yan, et al.
Veröffentlicht: (2025)
Improving Generalized Visual Grounding with Instance-aware Joint Learning
von: Dai, Ming, et al.
Veröffentlicht: (2025)
von: Dai, Ming, et al.
Veröffentlicht: (2025)
Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling
von: Wu, Keming, et al.
Veröffentlicht: (2026)
von: Wu, Keming, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Head-wise Adaptive Rotary Positional Encoding for Fine-Grained Image Generation
von: Li, Jiaye, et al.
Veröffentlicht: (2025) -
DPAR: Dynamic Patchification for Efficient Autoregressive Visual Generation
von: Srivastava, Divyansh, et al.
Veröffentlicht: (2025) -
Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation
von: Xu, Mingwang, et al.
Veröffentlicht: (2024) -
Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference Distillation
von: Li, Chunyu, et al.
Veröffentlicht: (2026) -
OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation
von: Li, Hui, et al.
Veröffentlicht: (2024)