FrameDiT: Diffusion Transformer with Matrix Attention for Efficient Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Le, Minh Khoa, Do, Kien, Nguyen, Duc Thanh, Tran, Truyen |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Finding the Trigger: Causal Abductive Reasoning on Video Events
by: Le, Thao Minh, et al.
Published: (2025)
by: Le, Thao Minh, et al.
Published: (2025)
Towards Agentic AI for Multimodal-Guided Video Object Segmentation
by: Tran, Tuyen, et al.
Published: (2025)
by: Tran, Tuyen, et al.
Published: (2025)
h-Edit: Effective and Flexible Diffusion-Based Editing via Doob's h-Transform
by: Nguyen, Toan, et al.
Published: (2025)
by: Nguyen, Toan, et al.
Published: (2025)
Robust Deepfake Detection: Mitigating Spatial Attention Drift via Calibrated Complementary Ensembles
by: Le-Phan, Minh-Khoa, et al.
Published: (2026)
by: Le-Phan, Minh-Khoa, et al.
Published: (2026)
Planner-Refiner: Dynamic Space-Time Refinement for Vision-Language Alignment in Videos
by: Tran, Tuyen, et al.
Published: (2025)
by: Tran, Tuyen, et al.
Published: (2025)
Amodal Instance Segmentation with Diffusion Shape Prior Estimation
by: Tran, Minh, et al.
Published: (2024)
by: Tran, Minh, et al.
Published: (2024)
MaskDiff: Modeling Mask Distribution with Diffusion Probabilistic Model for Few-Shot Instance Segmentation
by: Le, Minh-Quan, et al.
Published: (2023)
by: Le, Minh-Quan, et al.
Published: (2023)
Efficient-vDiT: Efficient Video Diffusion Transformers With Attention Tile
by: Ding, Hangliang, et al.
Published: (2025)
by: Ding, Hangliang, et al.
Published: (2025)
Unified Framework with Consistency across Modalities for Human Activity Recognition
by: Tran, Tuyen, et al.
Published: (2024)
by: Tran, Tuyen, et al.
Published: (2024)
EDGER: EDge-Guided with HEatmap Refinement for Generalizable Image Forgery Localization
by: Le-Phan, Minh-Khoa, et al.
Published: (2026)
by: Le-Phan, Minh-Khoa, et al.
Published: (2026)
The Art of Camouflage: Few-Shot Learning for Animal Detection and Segmentation
by: Nguyen, Thanh-Danh, et al.
Published: (2023)
by: Nguyen, Thanh-Danh, et al.
Published: (2023)
Bidirectional Diffusion Bridge Models
by: Kieu, Duc, et al.
Published: (2025)
by: Kieu, Duc, et al.
Published: (2025)
ShapeFormer: Shape Prior Visible-to-Amodal Transformer-based Amodal Instance Segmentation
by: Tran, Minh, et al.
Published: (2024)
by: Tran, Minh, et al.
Published: (2024)
Universal Multi-Domain Translation via Diffusion Routers
by: Kieu, Duc, et al.
Published: (2025)
by: Kieu, Duc, et al.
Published: (2025)
PADM: A Physics-aware Diffusion Model for Attenuation Correction
by: Pham, Trung Kien, et al.
Published: (2025)
by: Pham, Trung Kien, et al.
Published: (2025)
Mono3DV: Monocular 3D Object Detection with 3D-Aware Bipartite Matching and Variational Query DeNoising
by: Vu, Kiet Dang, et al.
Published: (2026)
by: Vu, Kiet Dang, et al.
Published: (2026)
CamoFA: A Learnable Fourier-based Augmentation for Camouflage Segmentation
by: Le, Minh-Quan, et al.
Published: (2023)
by: Le, Minh-Quan, et al.
Published: (2023)
SADL: An Effective In-Context Learning Method for Compositional Visual QA
by: Dang, Long Hoang, et al.
Published: (2024)
by: Dang, Long Hoang, et al.
Published: (2024)
DiTPainter: Efficient Video Inpainting with Diffusion Transformers
by: Wu, Xian, et al.
Published: (2025)
by: Wu, Xian, et al.
Published: (2025)
PINNs for Medical Image Analysis: A Survey
by: Banerjee, Chayan, et al.
Published: (2024)
by: Banerjee, Chayan, et al.
Published: (2024)
HENASY: Learning to Assemble Scene-Entities for Egocentric Video-Language Model
by: Vo, Khoa, et al.
Published: (2024)
by: Vo, Khoa, et al.
Published: (2024)
Anatomical Attention Alignment representation for Radiology Report Generation
by: Nguyen, Quang Vinh, et al.
Published: (2025)
by: Nguyen, Quang Vinh, et al.
Published: (2025)
SuMa: A Subspace Mapping Approach for Robust and Effective Concept Erasure in Text-to-Image Diffusion Models
by: Nguyen, Kien, et al.
Published: (2025)
by: Nguyen, Kien, et al.
Published: (2025)
Multi-Perspective Data Augmentation for Few-shot Object Detection
by: Vu, Anh-Khoa Nguyen, et al.
Published: (2025)
by: Vu, Anh-Khoa Nguyen, et al.
Published: (2025)
ViDiT-Q: Efficient and Accurate Quantization of Diffusion Transformers for Image and Video Generation
by: Zhao, Tianchen, et al.
Published: (2024)
by: Zhao, Tianchen, et al.
Published: (2024)
AV-DiT: Efficient Audio-Visual Diffusion Transformer for Joint Audio and Video Generation
by: Wang, Kai, et al.
Published: (2024)
by: Wang, Kai, et al.
Published: (2024)
AISFormer: Amodal Instance Segmentation with Transformer
by: Tran, Minh, et al.
Published: (2022)
by: Tran, Minh, et al.
Published: (2022)
HAtt-Flow: Hierarchical Attention-Flow Mechanism for Group Activity Scene Graph Generation in Videos
by: Chappa, Naga VS Raviteja, et al.
Published: (2023)
by: Chappa, Naga VS Raviteja, et al.
Published: (2023)
ShowFlow: From Robust Single Concept to Condition-Free Multi-Concept Generation
by: Hoang, Trong-Vu, et al.
Published: (2025)
by: Hoang, Trong-Vu, et al.
Published: (2025)
EA-Swin: An Embedding-Agnostic Swin Transformer for AI-Generated Video Detection
by: Mai, Hung, et al.
Published: (2026)
by: Mai, Hung, et al.
Published: (2026)
FullDiT2: Efficient In-Context Conditioning for Video Diffusion Transformers
by: He, Xuanhua, et al.
Published: (2025)
by: He, Xuanhua, et al.
Published: (2025)
GenFlow: Interactive Modular System for Image Generation
by: Nguyen, Duc-Hung, et al.
Published: (2025)
by: Nguyen, Duc-Hung, et al.
Published: (2025)
DiVE: Efficient Multi-View Driving Scenes Generation Based on Video Diffusion Transformer
by: Jiang, Junpeng, et al.
Published: (2025)
by: Jiang, Junpeng, et al.
Published: (2025)
Online 3D Multi-Camera Perception through Robust 2D Tracking and Depth-based Late Aggregation
by: Le, Vu-Minh, et al.
Published: (2025)
by: Le, Vu-Minh, et al.
Published: (2025)
THYME: Temporal Hierarchical-Cyclic Interactivity Modeling for Video Scene Graphs in Aerial Footage
by: Nguyen, Trong-Thuan, et al.
Published: (2025)
by: Nguyen, Trong-Thuan, et al.
Published: (2025)
Self-supervised Video Object Segmentation with Distillation Learning of Deformable Attention
by: Truong, Quang-Trung, et al.
Published: (2024)
by: Truong, Quang-Trung, et al.
Published: (2024)
Driver Attention Tracking and Analysis
by: Nguyen, Dat Viet Thanh, et al.
Published: (2024)
by: Nguyen, Dat Viet Thanh, et al.
Published: (2024)
S2DiT: Sandwich Diffusion Transformer for Mobile Streaming Video Generation
by: Zhao, Lin, et al.
Published: (2026)
by: Zhao, Lin, et al.
Published: (2026)
Hierarchical Neural Collapse Detection Transformer for Class Incremental Object Detection
by: Pham, Duc Thanh, et al.
Published: (2025)
by: Pham, Duc Thanh, et al.
Published: (2025)
Enhancing Dataset Distillation via Non-Critical Region Refinement
by: Tran, Minh-Tuan, et al.
Published: (2025)
by: Tran, Minh-Tuan, et al.
Published: (2025)
Similar Items
-
Finding the Trigger: Causal Abductive Reasoning on Video Events
by: Le, Thao Minh, et al.
Published: (2025) -
Towards Agentic AI for Multimodal-Guided Video Object Segmentation
by: Tran, Tuyen, et al.
Published: (2025) -
h-Edit: Effective and Flexible Diffusion-Based Editing via Doob's h-Transform
by: Nguyen, Toan, et al.
Published: (2025) -
Robust Deepfake Detection: Mitigating Spatial Attention Drift via Calibrated Complementary Ensembles
by: Le-Phan, Minh-Khoa, et al.
Published: (2026) -
Planner-Refiner: Dynamic Space-Time Refinement for Vision-Language Alignment in Videos
by: Tran, Tuyen, et al.
Published: (2025)