SEGA: Spectral-Energy Guided Attention for Resolution Extrapolation in Diffusion Transformers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Rajabi, Javad, Shaban, Kimia, Roohi, Koorosh, Lindell, David B., Taati, Babak |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
STARS: Self-supervised Tuning for 3D Action Recognition in Skeleton Sequences
von: Mehraban, Soroush, et al.
Veröffentlicht: (2024)
von: Mehraban, Soroush, et al.
Veröffentlicht: (2024)
When Does RL Help Medical VLMs? Disentangling Vision, SFT, and RL Gains
von: Jeddi, Ahmadreza, et al.
Veröffentlicht: (2026)
von: Jeddi, Ahmadreza, et al.
Veröffentlicht: (2026)
FastHMR: Accelerating Human Mesh Recovery via Token and Layer Merging with Diffusion Decoding
von: Mehraban, Soroush, et al.
Veröffentlicht: (2025)
von: Mehraban, Soroush, et al.
Veröffentlicht: (2025)
UltraImage: Rethinking Resolution Extrapolation in Image Diffusion Transformers
von: Zhao, Min, et al.
Veröffentlicht: (2025)
von: Zhao, Min, et al.
Veröffentlicht: (2025)
SUM: Saliency Unification through Mamba for Visual Attention Modeling
von: Hosseini, Alireza, et al.
Veröffentlicht: (2024)
von: Hosseini, Alireza, et al.
Veröffentlicht: (2024)
STAF: Sinusoidal Trainable Activation Functions for Implicit Neural Representation
von: Morsali, Alireza, et al.
Veröffentlicht: (2025)
von: Morsali, Alireza, et al.
Veröffentlicht: (2025)
Similarity-Aware Token Pruning: Your VLM but Faster
von: Jeddi, Ahmadreza, et al.
Veröffentlicht: (2025)
von: Jeddi, Ahmadreza, et al.
Veröffentlicht: (2025)
PuzzleCraft: Exploration-Aware Curriculum Learning for Puzzle-Based RLVR in VLMs
von: Jeddi, Ahmadreza, et al.
Veröffentlicht: (2025)
von: Jeddi, Ahmadreza, et al.
Veröffentlicht: (2025)
Mitigating 3D Prostate Biparametric MRI Data Scarcity through Domain Adaptation using Locally-Trained Latent Diffusion Models for Prostate Cancer Detection
von: Grabke, Emerson P., et al.
Veröffentlicht: (2025)
von: Grabke, Emerson P., et al.
Veröffentlicht: (2025)
Spectral and Trajectory Regularization for Diffusion Transformer Super-Resolution
von: Wang, Jingkai, et al.
Veröffentlicht: (2026)
von: Wang, Jingkai, et al.
Veröffentlicht: (2026)
LookHere: Vision Transformers with Directed Attention Generalize and Extrapolate
von: Fuller, Anthony, et al.
Veröffentlicht: (2024)
von: Fuller, Anthony, et al.
Veröffentlicht: (2024)
RAGSR: Regional Attention Guided Diffusion for Image Super-Resolution
von: He, Haodong, et al.
Veröffentlicht: (2025)
von: He, Haodong, et al.
Veröffentlicht: (2025)
DyPE: Dynamic Position Extrapolation for Ultra High Resolution Diffusion
von: Issachar, Noam, et al.
Veröffentlicht: (2025)
von: Issachar, Noam, et al.
Veröffentlicht: (2025)
PickStyle: Video-to-Video Style Transfer with Context-Style Adapters
von: Mehraban, Soroush, et al.
Veröffentlicht: (2025)
von: Mehraban, Soroush, et al.
Veröffentlicht: (2025)
LIFT: Latent Implicit Functions for Task- and Data-Agnostic Encoding
von: Kazerouni, Amirhossein, et al.
Veröffentlicht: (2025)
von: Kazerouni, Amirhossein, et al.
Veröffentlicht: (2025)
CAP4D: Creating Animatable 4D Portrait Avatars with Morphable Multi-View Diffusion Models
von: Taubner, Felix, et al.
Veröffentlicht: (2024)
von: Taubner, Felix, et al.
Veröffentlicht: (2024)
Leveraging Clinical Text and Class Conditioning for 3D Prostate MRI Generation
von: Grabke, Emerson P., et al.
Veröffentlicht: (2025)
von: Grabke, Emerson P., et al.
Veröffentlicht: (2025)
Pain in 3D: Generating Controllable Synthetic Faces for Automated Pain Assessment
von: Lin, Xin Lei, et al.
Veröffentlicht: (2025)
von: Lin, Xin Lei, et al.
Veröffentlicht: (2025)
AC3D: Analyzing and Improving 3D Camera Control in Video Diffusion Transformers
von: Bahmani, Sherwin, et al.
Veröffentlicht: (2024)
von: Bahmani, Sherwin, et al.
Veröffentlicht: (2024)
UltraViCo: Breaking Extrapolation Limits in Video Diffusion Transformers
von: Zhao, Min, et al.
Veröffentlicht: (2025)
von: Zhao, Min, et al.
Veröffentlicht: (2025)
RIFLEx: A Free Lunch for Length Extrapolation in Video Diffusion Transformers
von: Zhao, Min, et al.
Veröffentlicht: (2025)
von: Zhao, Min, et al.
Veröffentlicht: (2025)
SEGA: A Stepwise Evolution Paradigm for Content-Aware Layout Generation with Design Prior
von: Wang, Haoran, et al.
Veröffentlicht: (2025)
von: Wang, Haoran, et al.
Veröffentlicht: (2025)
Skewness-Guided Pruning of Multimodal Swin Transformers for Federated Skin Lesion Classification on Edge Devices
von: Paxton, Kuniko, et al.
Veröffentlicht: (2025)
von: Paxton, Kuniko, et al.
Veröffentlicht: (2025)
Dynamic Attention-Guided Diffusion for Image Super-Resolution
von: Moser, Brian B., et al.
Veröffentlicht: (2023)
von: Moser, Brian B., et al.
Veröffentlicht: (2023)
Transientangelo: Few-Viewpoint Surface Reconstruction Using Single-Photon Lidar
von: Luo, Weihan, et al.
Veröffentlicht: (2024)
von: Luo, Weihan, et al.
Veröffentlicht: (2024)
AccDiffusion v2: Towards More Accurate Higher-Resolution Diffusion Extrapolation
von: Lin, Zhihang, et al.
Veröffentlicht: (2024)
von: Lin, Zhihang, et al.
Veröffentlicht: (2024)
Distilling Latent Manifolds: Resolution Extrapolation by Variational Autoencoders
von: Chu, Jiaming, et al.
Veröffentlicht: (2026)
von: Chu, Jiaming, et al.
Veröffentlicht: (2026)
Face2Scene: Using Facial Degradation as an Oracle for Diffusion-Based Scene Restoration
von: Kazerouni, Amirhossein, et al.
Veröffentlicht: (2026)
von: Kazerouni, Amirhossein, et al.
Veröffentlicht: (2026)
SEGA: Drivable 3D Gaussian Head Avatar from a Single Image
von: Guo, Chen, et al.
Veröffentlicht: (2025)
von: Guo, Chen, et al.
Veröffentlicht: (2025)
TIDE: Text-Informed Dynamic Extrapolation with Step-Aware Temperature Control for Diffusion Transformers
von: Liu, Yihua, et al.
Veröffentlicht: (2026)
von: Liu, Yihua, et al.
Veröffentlicht: (2026)
Benchmarking Skeleton-based Motion Encoder Models for Clinical Applications: Estimating Parkinson's Disease Severity in Walking Sequences
von: Adeli, Vida, et al.
Veröffentlicht: (2024)
von: Adeli, Vida, et al.
Veröffentlicht: (2024)
MangaDiT: Reference-Guided Line Art Colorization with Hierarchical Attention in Diffusion Transformers
von: Qiu, Qianru, et al.
Veröffentlicht: (2025)
von: Qiu, Qianru, et al.
Veröffentlicht: (2025)
One Attention, One Scale: Phase-Aligned Rotary Positional Embeddings for Mixed-Resolution Diffusion Transformer
von: Wu, Haoyu, et al.
Veröffentlicht: (2025)
von: Wu, Haoyu, et al.
Veröffentlicht: (2025)
Novel View Extrapolation with Video Diffusion Priors
von: Liu, Kunhao, et al.
Veröffentlicht: (2024)
von: Liu, Kunhao, et al.
Veröffentlicht: (2024)
Automatic Aorta Segmentation with Heavily Augmented, High-Resolution 3-D ResUNet: Contribution to the SEG.A Challenge
von: Wodzinski, Marek, et al.
Veröffentlicht: (2023)
von: Wodzinski, Marek, et al.
Veröffentlicht: (2023)
SEGA: A Transferable Signed Ensemble Gaussian Black-Box Attack against No-Reference Image Quality Assessment Models
von: Liu, Yujia, et al.
Veröffentlicht: (2025)
von: Liu, Yujia, et al.
Veröffentlicht: (2025)
IGAF: Incremental Guided Attention Fusion for Depth Super-Resolution
von: Tragakis, Athanasios, et al.
Veröffentlicht: (2025)
von: Tragakis, Athanasios, et al.
Veröffentlicht: (2025)
FreeText: Training-Free Text Rendering in Diffusion Transformers via Attention Localization and Spectral Glyph Injection
von: Zhang, Ruiqiang, et al.
Veröffentlicht: (2026)
von: Zhang, Ruiqiang, et al.
Veröffentlicht: (2026)
VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control
von: Bahmani, Sherwin, et al.
Veröffentlicht: (2024)
von: Bahmani, Sherwin, et al.
Veröffentlicht: (2024)
Analysis of Attention in Video Diffusion Transformers
von: Wen, Yuxin, et al.
Veröffentlicht: (2025)
von: Wen, Yuxin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
STARS: Self-supervised Tuning for 3D Action Recognition in Skeleton Sequences
von: Mehraban, Soroush, et al.
Veröffentlicht: (2024) -
When Does RL Help Medical VLMs? Disentangling Vision, SFT, and RL Gains
von: Jeddi, Ahmadreza, et al.
Veröffentlicht: (2026) -
FastHMR: Accelerating Human Mesh Recovery via Token and Layer Merging with Diffusion Decoding
von: Mehraban, Soroush, et al.
Veröffentlicht: (2025) -
UltraImage: Rethinking Resolution Extrapolation in Image Diffusion Transformers
von: Zhao, Min, et al.
Veröffentlicht: (2025) -
SUM: Saliency Unification through Mamba for Visual Attention Modeling
von: Hosseini, Alireza, et al.
Veröffentlicht: (2024)