FEAT: Full-Dimensional Efficient Attention Transformer for Medical Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Huihan, Yang, Zhiwen, Zhang, Hui, Zhao, Dan, Wei, Bingzheng, Xu, Yan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Region Attention Transformer for Medical Image Restoration
by: Yang, Zhiwen, et al.
Published: (2024)
by: Yang, Zhiwen, et al.
Published: (2024)
Restore-RWKV: Efficient and Effective Medical Image Restoration with RWKV
by: Yang, Zhiwen, et al.
Published: (2024)
by: Yang, Zhiwen, et al.
Published: (2024)
TAT: Task-Adaptive Transformer for All-in-One Medical Image Restoration
by: Yang, Zhiwen, et al.
Published: (2025)
by: Yang, Zhiwen, et al.
Published: (2025)
All-In-One Medical Image Restoration via Task-Adaptive Routing
by: Yang, Zhiwen, et al.
Published: (2024)
by: Yang, Zhiwen, et al.
Published: (2024)
All-in-One Medical Image Restoration with Latent Diffusion-Enhanced Vector-Quantized Codebook Prior
by: Chen, Haowei, et al.
Published: (2025)
by: Chen, Haowei, et al.
Published: (2025)
FullDiT: Multi-Task Video Generative Foundation Model with Full Attention
by: Ju, Xuan, et al.
Published: (2025)
by: Ju, Xuan, et al.
Published: (2025)
iFlame: Interleaving Full and Linear Attention for Efficient Mesh Generation
by: Wang, Hanxiao, et al.
Published: (2025)
by: Wang, Hanxiao, et al.
Published: (2025)
FlashVGGT: Efficient and Scalable Visual Geometry Transformers with Compressed Descriptor Attention
by: Wang, Zipeng, et al.
Published: (2025)
by: Wang, Zipeng, et al.
Published: (2025)
FullTransNet: Full Transformer with Local-Global Attention for Video Summarization
by: Lan, Libin, et al.
Published: (2025)
by: Lan, Libin, et al.
Published: (2025)
NEWTON: Agentic Planning for Physically Grounded Video Generation
by: Feng, Yuxiang, et al.
Published: (2026)
by: Feng, Yuxiang, et al.
Published: (2026)
FullDiT2: Efficient In-Context Conditioning for Video Diffusion Transformers
by: He, Xuanhua, et al.
Published: (2025)
by: He, Xuanhua, et al.
Published: (2025)
Visual Textualization for Image Prompted Object Detection
by: Wu, Yongjian, et al.
Published: (2025)
by: Wu, Yongjian, et al.
Published: (2025)
Beyond Rigid: Benchmarking Non-Rigid Video Editing
by: Qu, Bingzheng, et al.
Published: (2026)
by: Qu, Bingzheng, et al.
Published: (2026)
Efficient-vDiT: Efficient Video Diffusion Transformers With Attention Tile
by: Ding, Hangliang, et al.
Published: (2025)
by: Ding, Hangliang, et al.
Published: (2025)
FrameDiT: Diffusion Transformer with Matrix Attention for Efficient Video Generation
by: Le, Minh Khoa, et al.
Published: (2026)
by: Le, Minh Khoa, et al.
Published: (2026)
Interspatial Attention for Efficient 4D Human Video Generation
by: Shao, Ruizhi, et al.
Published: (2025)
by: Shao, Ruizhi, et al.
Published: (2025)
SALAD: Achieve High-Sparsity Attention via Efficient Linear Attention Tuning for Video Diffusion Transformer
by: Fang, Tongcheng, et al.
Published: (2026)
by: Fang, Tongcheng, et al.
Published: (2026)
Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey
by: Qu, Bingzheng, et al.
Published: (2026)
by: Qu, Bingzheng, et al.
Published: (2026)
CTIS-QA: Clinical Template-Informed Slide-level Question Answering for Pathology
by: Lu, Hao, et al.
Published: (2026)
by: Lu, Hao, et al.
Published: (2026)
FEAT: Fashion Editing and Try-On from Any Design
by: Kwon, Soye, et al.
Published: (2026)
by: Kwon, Soye, et al.
Published: (2026)
MedVol-R1: Reward-Driven Evidence Grounding for Volumetric Reasoning Segmentation
by: Wang, Zichun, et al.
Published: (2026)
by: Wang, Zichun, et al.
Published: (2026)
Uncertainty-Aware Prototype Semantic Decoupling for Text-Based Person Search in Full Images
by: Luo, Zengli, et al.
Published: (2025)
by: Luo, Zengli, et al.
Published: (2025)
High-Speed FHD Full-Color Video Computer-Generated Holography
by: Zhang, Haomiao, et al.
Published: (2025)
by: Zhang, Haomiao, et al.
Published: (2025)
Video Super-Resolution Transformer with Masked Inter&Intra-Frame Attention
by: Zhou, Xingyu, et al.
Published: (2024)
by: Zhou, Xingyu, et al.
Published: (2024)
Video Compression Meets Video Generation: Latent Inter-Frame Pruning with Attention Recovery
by: Menn, Dennis, et al.
Published: (2026)
by: Menn, Dennis, et al.
Published: (2026)
ViDiT-Q: Efficient and Accurate Quantization of Diffusion Transformers for Image and Video Generation
by: Zhao, Tianchen, et al.
Published: (2024)
by: Zhao, Tianchen, et al.
Published: (2024)
Generative Model-Based Feature Attention Module for Video Action Analysis
by: Wang, Guiqin, et al.
Published: (2025)
by: Wang, Guiqin, et al.
Published: (2025)
SDPT: Synchronous Dual Prompt Tuning for Fusion-based Visual-Language Pre-trained Models
by: Zhou, Yang, et al.
Published: (2024)
by: Zhou, Yang, et al.
Published: (2024)
AttriPrompter: Auto-Prompting with Attribute Semantics for Zero-shot Nuclei Detection via Visual-Language Pre-trained Models
by: Wu, Yongjian, et al.
Published: (2024)
by: Wu, Yongjian, et al.
Published: (2024)
Generalizable and Animatable 3D Full-Head Gaussian Avatar from a Single Image
by: Zhao, Shuling, et al.
Published: (2026)
by: Zhao, Shuling, et al.
Published: (2026)
Attention Surgery: An Efficient Recipe to Linearize Your Video Diffusion Transformer
by: Ghafoorian, Mohsen, et al.
Published: (2025)
by: Ghafoorian, Mohsen, et al.
Published: (2025)
VMonarch: Efficient Video Diffusion Transformers with Structured Attention
by: Liang, Cheng, et al.
Published: (2026)
by: Liang, Cheng, et al.
Published: (2026)
Trainable Log-linear Sparse Attention for Efficient Diffusion Transformers
by: Zhou, Yifan, et al.
Published: (2025)
by: Zhou, Yifan, et al.
Published: (2025)
HAM: A Training-Free Style Transfer Approach via Heterogeneous Attention Modulation for Diffusion Models
by: He, Yeqi, et al.
Published: (2026)
by: He, Yeqi, et al.
Published: (2026)
Learning Online Scale Transformation for Talking Head Video Generation
by: Hong, Fa-Ting, et al.
Published: (2024)
by: Hong, Fa-Ting, et al.
Published: (2024)
Training-free and Adaptive Sparse Attention for Efficient Long Video Generation
by: Xia, Yifei, et al.
Published: (2025)
by: Xia, Yifei, et al.
Published: (2025)
TCSAFormer: Efficient Vision Transformer with Token Compression and Sparse Attention for Medical Image Segmentation
by: Xia, Zunhui, et al.
Published: (2025)
by: Xia, Zunhui, et al.
Published: (2025)
SegStitch: Multidimensional Transformer for Robust and Efficient Medical Imaging Segmentation
by: Tan, Shengbo, et al.
Published: (2024)
by: Tan, Shengbo, et al.
Published: (2024)
A Density-Guided Temporal Attention Transformer for Indiscernible Object Counting in Underwater Video
by: Yang, Cheng-Yen, et al.
Published: (2024)
by: Yang, Cheng-Yen, et al.
Published: (2024)
Polyline Path Masked Attention for Vision Transformer
by: Zhao, Zhongchen, et al.
Published: (2025)
by: Zhao, Zhongchen, et al.
Published: (2025)
Similar Items
-
Region Attention Transformer for Medical Image Restoration
by: Yang, Zhiwen, et al.
Published: (2024) -
Restore-RWKV: Efficient and Effective Medical Image Restoration with RWKV
by: Yang, Zhiwen, et al.
Published: (2024) -
TAT: Task-Adaptive Transformer for All-in-One Medical Image Restoration
by: Yang, Zhiwen, et al.
Published: (2025) -
All-In-One Medical Image Restoration via Task-Adaptive Routing
by: Yang, Zhiwen, et al.
Published: (2024) -
All-in-One Medical Image Restoration with Latent Diffusion-Enhanced Vector-Quantized Codebook Prior
by: Chen, Haowei, et al.
Published: (2025)