One Attention, One Scale: Phase-Aligned Rotary Positional Embeddings for Mixed-Resolution Diffusion Transformer
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Haoyu, Xu, Jingyi, Miao, Qiaomu, Samaras, Dimitris, Le, Hieu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Importance-Based Token Merging for Efficient Image and Video Generation
by: Wu, Haoyu, et al.
Published: (2024)
by: Wu, Haoyu, et al.
Published: (2024)
Embedding Physical Reasoning into Diffusion-Based Shadow Generation
by: Hu, Shilin, et al.
Published: (2025)
by: Hu, Shilin, et al.
Published: (2025)
Assessing Sample Quality via the Latent Space of Generative Models
by: Xu, Jingyi, et al.
Published: (2024)
by: Xu, Jingyi, et al.
Published: (2024)
Talking Head Generation via AU-Guided Landmark Prediction
by: Chang, Shao-Yu, et al.
Published: (2025)
by: Chang, Shao-Yu, et al.
Published: (2025)
Diffusion-Refined VQA Annotations for Semi-Supervised Gaze Following
by: Miao, Qiaomu, et al.
Published: (2024)
by: Miao, Qiaomu, et al.
Published: (2024)
Multi-view Gaze Target Estimation
by: Miao, Qiaomu, et al.
Published: (2025)
by: Miao, Qiaomu, et al.
Published: (2025)
Cast and Attached Shadow Detection via Iterative Light and Geometry Reasoning
by: Hu, Shilin, et al.
Published: (2025)
by: Hu, Shilin, et al.
Published: (2025)
Weighting Pseudo-Labels via High-Activation Feature Index Similarity and Object Detection for Semi-Supervised Segmentation
by: Howlader, Prantik, et al.
Published: (2024)
by: Howlader, Prantik, et al.
Published: (2024)
Personalized Image Descriptions from Attention Sequences
by: Xue, Ruoyu, et al.
Published: (2025)
by: Xue, Ruoyu, et al.
Published: (2025)
Improving Contrastive Learning for Referring Expression Counting
by: Triaridis, Kostas, et al.
Published: (2025)
by: Triaridis, Kostas, et al.
Published: (2025)
CORA: Consistency-Guided Semi-Supervised Framework for Reasoning Segmentation
by: Howlader, Prantik, et al.
Published: (2025)
by: Howlader, Prantik, et al.
Published: (2025)
Beyond Pixels: Semi-Supervised Semantic Segmentation with a Multi-scale Patch-based Multi-Label Classifier
by: Howlader, Prantik, et al.
Published: (2024)
by: Howlader, Prantik, et al.
Published: (2024)
Rotary Position Embedding for Vision Transformer
by: Heo, Byeongho, et al.
Published: (2024)
by: Heo, Byeongho, et al.
Published: (2024)
Few-shot Personalized Scanpath Prediction
by: Xue, Ruoyu, et al.
Published: (2025)
by: Xue, Ruoyu, et al.
Published: (2025)
Diffusion MRI Transformer with a Diffusion Space Rotary Positional Embedding (D-RoPE)
by: Kung, Gustavo Chau Loo, et al.
Published: (2026)
by: Kung, Gustavo Chau Loo, et al.
Published: (2026)
Efficient Matrix Implementation for Rotary Position Embedding
by: Minqi, Chen, et al.
Published: (2026)
by: Minqi, Chen, et al.
Published: (2026)
Learning 3D Reconstruction with Priors in Test Time
by: Zhou, Lei, et al.
Published: (2026)
by: Zhou, Lei, et al.
Published: (2026)
TopoDiffusionNet: A Topology-aware Diffusion Model
by: Gupta, Saumya, et al.
Published: (2024)
by: Gupta, Saumya, et al.
Published: (2024)
Shadow Removal Refinement via Material-Consistent Shadow Edges
by: Hu, Shilin, et al.
Published: (2024)
by: Hu, Shilin, et al.
Published: (2024)
Cross-Axis Transformer with 3D Rotary Positional Embeddings
by: Erickson, Lily
Published: (2023)
by: Erickson, Lily
Published: (2023)
Spiral RoPE: Rotate Your Rotary Positional Embeddings in the 2D Plane
by: Liu, Haoyu, et al.
Published: (2026)
by: Liu, Haoyu, et al.
Published: (2026)
Phrase-Instance Alignment for Generalized Referring Segmentation
by: Nguyen, E-Ro, et al.
Published: (2024)
by: Nguyen, E-Ro, et al.
Published: (2024)
BinaryAttention: One-Bit QK-Attention for Vision and Diffusion Transformers
by: Xiao, Chaodong, et al.
Published: (2026)
by: Xiao, Chaodong, et al.
Published: (2026)
Self-supervised co-salient object detection via feature correspondence at multiple scales
by: Chakraborty, Souradeep, et al.
Published: (2024)
by: Chakraborty, Souradeep, et al.
Published: (2024)
PISCES: Annotation-free Text-to-Video Post-Training via Optimal Transport-Aligned Rewards
by: Le, Minh-Quan, et al.
Published: (2026)
by: Le, Minh-Quan, et al.
Published: (2026)
RomanTex: Decoupling 3D-aware Rotary Positional Embedded Multi-Attention Network for Texture Synthesis
by: Feng, Yifei, et al.
Published: (2025)
by: Feng, Yifei, et al.
Published: (2025)
Geometry-Aware Rotary Position Embedding for Consistent Video World Model
by: Xiang, Chendong, et al.
Published: (2026)
by: Xiang, Chendong, et al.
Published: (2026)
One-Step Diffusion Transformer for Controllable Real-World Image Super-Resolution
by: Fang, Yushun, et al.
Published: (2025)
by: Fang, Yushun, et al.
Published: (2025)
Learning Relighting and Intrinsic Decomposition in Neural Radiance Fields
by: Yang, Yixiong, et al.
Published: (2024)
by: Yang, Yixiong, et al.
Published: (2024)
MLI-NeRF: Multi-Light Intrinsic-Aware Neural Radiance Fields
by: Yang, Yixiong, et al.
Published: (2024)
by: Yang, Yixiong, et al.
Published: (2024)
VideoRoPE: What Makes for Good Video Rotary Position Embedding?
by: Wei, Xilin, et al.
Published: (2025)
by: Wei, Xilin, et al.
Published: (2025)
PathSegDiff: Pathology Segmentation using Diffusion model representations
by: Danisetty, Sachin Kumar, et al.
Published: (2025)
by: Danisetty, Sachin Kumar, et al.
Published: (2025)
Rethinking Diffusion Model-Based Video Super-Resolution: Leveraging Dense Guidance from Aligned Features
by: Xu, Jingyi, et al.
Published: (2025)
by: Xu, Jingyi, et al.
Published: (2025)
$\infty$-Brush: Controllable Large Image Synthesis with Diffusion Models in Infinite Dimensions
by: Le, Minh-Quan, et al.
Published: (2024)
by: Le, Minh-Quan, et al.
Published: (2024)
What about gravity in video generation? Post-Training Newton's Laws with Verifiable Rewards
by: Le, Minh-Quan, et al.
Published: (2025)
by: Le, Minh-Quan, et al.
Published: (2025)
Efficient Burst Super-Resolution with One-step Diffusion
by: Kawai, Kento, et al.
Published: (2025)
by: Kawai, Kento, et al.
Published: (2025)
DRoPE: Directional Rotary Position Embedding for Efficient Agent Interaction Modeling
by: Zhao, Jianbo, et al.
Published: (2025)
by: Zhao, Jianbo, et al.
Published: (2025)
Learning to Estimate Critical Gait Parameters from Single-View RGB Videos with Transformer-Based Attention Network
by: Le, Quoc Hung T., et al.
Published: (2023)
by: Le, Quoc Hung T., et al.
Published: (2023)
Multi-Level Feature Fusion Network for Lightweight Stereo Image Super-Resolution
by: Li, Yunxiang, et al.
Published: (2024)
by: Li, Yunxiang, et al.
Published: (2024)
Unleashing the Power of One-Step Diffusion based Image Super-Resolution via a Large-Scale Diffusion Discriminator
by: Li, Jianze, et al.
Published: (2024)
by: Li, Jianze, et al.
Published: (2024)
Similar Items
-
Importance-Based Token Merging for Efficient Image and Video Generation
by: Wu, Haoyu, et al.
Published: (2024) -
Embedding Physical Reasoning into Diffusion-Based Shadow Generation
by: Hu, Shilin, et al.
Published: (2025) -
Assessing Sample Quality via the Latent Space of Generative Models
by: Xu, Jingyi, et al.
Published: (2024) -
Talking Head Generation via AU-Guided Landmark Prediction
by: Chang, Shao-Yu, et al.
Published: (2025) -
Diffusion-Refined VQA Annotations for Semi-Supervised Gaze Following
by: Miao, Qiaomu, et al.
Published: (2024)