Perceiving Longer Sequences With Bi-Directional Cross-Attention Transformers
Fuente:
arXiv
Saved in:
| Main Authors: | Hiller, Markus, Ehinger, Krista A., Drummond, Tom |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Improving Denoising Diffusion Models via Simultaneous Estimation of Image and Noise
by: Zhang, Zhenkai, et al.
Published: (2023)
by: Zhang, Zhenkai, et al.
Published: (2023)
Sequential Amodal Segmentation via Cumulative Occlusion Learning
by: Ao, Jiayang, et al.
Published: (2024)
by: Ao, Jiayang, et al.
Published: (2024)
Open-World Amodal Appearance Completion
by: Ao, Jiayang, et al.
Published: (2024)
by: Ao, Jiayang, et al.
Published: (2024)
KALE: An Artwork Image Captioning System Augmented with Heterogeneous Graph
by: Jiang, Yanbei, et al.
Published: (2024)
by: Jiang, Yanbei, et al.
Published: (2024)
PROPA: Toward Process-level Optimization in Visual Reasoning via Reinforcement Learning
by: Jiang, Yanbei, et al.
Published: (2025)
by: Jiang, Yanbei, et al.
Published: (2025)
Beyond Augmentation: Cross-Modal Transformer Fusion with Bi-directional Attention for Low-Data Aneurysm Screening
by: Titikhsha, Antara, et al.
Published: (2025)
by: Titikhsha, Antara, et al.
Published: (2025)
DeBiFormer: Vision Transformer with Deformable Agent Bi-level Routing Attention
by: Long, Nguyen Huu Bao, et al.
Published: (2024)
by: Long, Nguyen Huu Bao, et al.
Published: (2024)
AdaReTaKe: Adaptive Redundancy Reduction to Perceive Longer for Video-language Understanding
by: Wang, Xiao, et al.
Published: (2025)
by: Wang, Xiao, et al.
Published: (2025)
Beyond Perception: Evaluating Abstract Visual Reasoning through Multi-Stage Task
by: Jiang, Yanbei, et al.
Published: (2025)
by: Jiang, Yanbei, et al.
Published: (2025)
EViT: An Eagle Vision Transformer with Bi-Fovea Self-Attention
by: Shi, Yulong, et al.
Published: (2023)
by: Shi, Yulong, et al.
Published: (2023)
Reasoning Like Experts: Leveraging Multimodal Large Language Models for Drawing-based Psychoanalysis
by: Ma, Xueqi, et al.
Published: (2025)
by: Ma, Xueqi, et al.
Published: (2025)
Analysis of Attention in Video Diffusion Transformers
by: Wen, Yuxin, et al.
Published: (2025)
by: Wen, Yuxin, et al.
Published: (2025)
Cross-modulated Attention Transformer for RGBT Tracking
by: Xiao, Yun, et al.
Published: (2024)
by: Xiao, Yun, et al.
Published: (2024)
CrossLMM: Decoupling Long Video Sequences from LMMs via Dual Cross-Attention Mechanisms
by: Yan, Shilin, et al.
Published: (2025)
by: Yan, Shilin, et al.
Published: (2025)
LookHere: Vision Transformers with Directed Attention Generalize and Extrapolate
by: Fuller, Anthony, et al.
Published: (2024)
by: Fuller, Anthony, et al.
Published: (2024)
LoL: Longer than Longer, Scaling Video Generation to Hour
by: Cui, Justin, et al.
Published: (2026)
by: Cui, Justin, et al.
Published: (2026)
Controllable Longer Image Animation with Diffusion Models
by: Wang, Qiang, et al.
Published: (2024)
by: Wang, Qiang, et al.
Published: (2024)
ACIT: Attention-Guided Cross-Modal Interaction Transformer for Pedestrian Crossing Intention Prediction
by: Li, Yuanzhe, et al.
Published: (2025)
by: Li, Yuanzhe, et al.
Published: (2025)
StreetForward: Perceiving Dynamic Street with Feedforward Causal Attention
by: Yu, Zhongrui, et al.
Published: (2026)
by: Yu, Zhongrui, et al.
Published: (2026)
DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video Generation
by: Cai, Minghong, et al.
Published: (2024)
by: Cai, Minghong, et al.
Published: (2024)
RGB-Sonar Tracking Benchmark and Spatial Cross-Attention Transformer Tracker
by: Li, Yunfeng, et al.
Published: (2024)
by: Li, Yunfeng, et al.
Published: (2024)
Estimating Extreme 3D Image Rotation with Transformer Cross-Attention
by: Dekel, Shay, et al.
Published: (2023)
by: Dekel, Shay, et al.
Published: (2023)
Event-Based Motion Segmentation by Motion Compensation
by: Stoffregen, Timo, et al.
Published: (2019)
by: Stoffregen, Timo, et al.
Published: (2019)
Admitting Ignorance Helps the Video Question Answering Models to Answer
by: Li, Haopeng, et al.
Published: (2025)
by: Li, Haopeng, et al.
Published: (2025)
ROI-Aware Multiscale Cross-Attention Vision Transformer for Pest Image Identification
by: Kim, Ga-Eun, et al.
Published: (2023)
by: Kim, Ga-Eun, et al.
Published: (2023)
DBAT: Dynamic Backward Attention Transformer for Material Segmentation with Cross-Resolution Patches
by: Heng, Yuwen, et al.
Published: (2023)
by: Heng, Yuwen, et al.
Published: (2023)
Embedded Heterogeneous Attention Transformer for Cross-lingual Image Captioning
by: Song, Zijie, et al.
Published: (2023)
by: Song, Zijie, et al.
Published: (2023)
Multimodal Emotion Recognition via Bi-directional Cross-Attention and Temporal Modeling
by: Byeon, Junhyeong, et al.
Published: (2026)
by: Byeon, Junhyeong, et al.
Published: (2026)
Bi-Directional Deep Contextual Video Compression
by: Sheng, Xihua, et al.
Published: (2024)
by: Sheng, Xihua, et al.
Published: (2024)
Personalized Image Descriptions from Attention Sequences
by: Xue, Ruoyu, et al.
Published: (2025)
by: Xue, Ruoyu, et al.
Published: (2025)
Direct Visual Grounding by Directing Attention of Visual Tokens
by: Esmaeilkhani, Parsa, et al.
Published: (2025)
by: Esmaeilkhani, Parsa, et al.
Published: (2025)
Attention Normalization Impacts Cardinality Generalization in Slot Attention
by: Krimmel, Markus, et al.
Published: (2024)
by: Krimmel, Markus, et al.
Published: (2024)
FreeNoise: Tuning-Free Longer Video Diffusion via Noise Rescheduling
by: Qiu, Haonan, et al.
Published: (2023)
by: Qiu, Haonan, et al.
Published: (2023)
Free Geometry: Refining 3D Reconstruction from Longer Versions of Itself
by: Dai, Yuhang, et al.
Published: (2026)
by: Dai, Yuhang, et al.
Published: (2026)
DeforHMR: Vision Transformer with Deformable Cross-Attention for 3D Human Mesh Recovery
by: Heo, Jaewoo, et al.
Published: (2024)
by: Heo, Jaewoo, et al.
Published: (2024)
CATFace: Cross-Attribute-Guided Transformer with Self-Attention Distillation for Low-Quality Face Recognition
by: Talemi, Niloufar Alipour, et al.
Published: (2024)
by: Talemi, Niloufar Alipour, et al.
Published: (2024)
Lung Infection Severity Prediction Using Transformers with Conditional TransMix Augmentation and Cross-Attention
by: Slika, Bouthaina, et al.
Published: (2025)
by: Slika, Bouthaina, et al.
Published: (2025)
Dual Cross-Attention Siamese Transformer for Rectal Tumor Regrowth Assessment in Watch-and-Wait Endoscopy
by: Gomez, Jorge Tapias, et al.
Published: (2025)
by: Gomez, Jorge Tapias, et al.
Published: (2025)
Bi-directional Contextual Attention for 3D Dense Captioning
by: Kim, Minjung, et al.
Published: (2024)
by: Kim, Minjung, et al.
Published: (2024)
Cross-Attention is Not Always Needed: Dynamic Cross-Attention for Audio-Visual Dimensional Emotion Recognition
by: Praveen, R. Gnana, et al.
Published: (2024)
by: Praveen, R. Gnana, et al.
Published: (2024)
Similar Items
-
Improving Denoising Diffusion Models via Simultaneous Estimation of Image and Noise
by: Zhang, Zhenkai, et al.
Published: (2023) -
Sequential Amodal Segmentation via Cumulative Occlusion Learning
by: Ao, Jiayang, et al.
Published: (2024) -
Open-World Amodal Appearance Completion
by: Ao, Jiayang, et al.
Published: (2024) -
KALE: An Artwork Image Captioning System Augmented with Heterogeneous Graph
by: Jiang, Yanbei, et al.
Published: (2024) -
PROPA: Toward Process-level Optimization in Visual Reasoning via Reinforcement Learning
by: Jiang, Yanbei, et al.
Published: (2025)