DANTE-AD: Dual-Vision Attention Network for Long-Term Audio Description
Fuente:
arXiv
Guardado en:
| Autores principales: | Deganutti, Adrienne, Hadfield, Simon, Gilbert, Andrew |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
DistinctAD: Distinctive Audio Description Generation in Contexts
por: Fang, Bo, et al.
Publicado: (2024)
por: Fang, Bo, et al.
Publicado: (2024)
AutoAD-Zero: A Training-Free Framework for Zero-Shot Audio Description
por: Xie, Junyu, et al.
Publicado: (2024)
por: Xie, Junyu, et al.
Publicado: (2024)
LLM-AD: Large Language Model based Audio Description System
por: Chu, Peng, et al.
Publicado: (2024)
por: Chu, Peng, et al.
Publicado: (2024)
FocusedAD: Character-centric Movie Audio Description
por: Ye, Xiaojun, et al.
Publicado: (2025)
por: Ye, Xiaojun, et al.
Publicado: (2025)
Evaluating Design Video Generation: Metrics for Compositional Fidelity
por: Deganutti, Adrienne, et al.
Publicado: (2026)
por: Deganutti, Adrienne, et al.
Publicado: (2026)
InfinityHuman: Towards Long-Term Audio-Driven Human
por: Li, Xiaodi, et al.
Publicado: (2025)
por: Li, Xiaodi, et al.
Publicado: (2025)
Graphic-Design-Bench: A Comprehensive Benchmark for Evaluating AI on Graphic Design Tasks
por: Deganutti, Adrienne, et al.
Publicado: (2026)
por: Deganutti, Adrienne, et al.
Publicado: (2026)
Mamba2D: A Natively Multi-Dimensional State-Space Model for Vision Tasks
por: Baty, Enis, et al.
Publicado: (2024)
por: Baty, Enis, et al.
Publicado: (2024)
SpaGBOL: Spatial-Graph-Based Orientated Localisation
por: Shore, Tavis, et al.
Publicado: (2024)
por: Shore, Tavis, et al.
Publicado: (2024)
Deep Leakage with Generative Flow Matching Denoiser
por: Baglin, Isaac, et al.
Publicado: (2026)
por: Baglin, Isaac, et al.
Publicado: (2026)
DEL: Dense Event Localization for Multi-modal Audio-Visual Understanding
por: Ahmadian, Mona, et al.
Publicado: (2025)
por: Ahmadian, Mona, et al.
Publicado: (2025)
Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency
por: Jiang, Jianwen, et al.
Publicado: (2024)
por: Jiang, Jianwen, et al.
Publicado: (2024)
Tri-FusionNet: Enhancing Image Description Generation with Transformer-based Fusion Network and Dual Attention Mechanism
por: Agarwal, Lakshita, et al.
Publicado: (2025)
por: Agarwal, Lakshita, et al.
Publicado: (2025)
DualAD: Disentangling the Dynamic and Static World for End-to-End Driving
por: Doll, Simon, et al.
Publicado: (2024)
por: Doll, Simon, et al.
Publicado: (2024)
FlowDet: Unifying Object Detection and Generative Transport Flows
por: Baty, Enis, et al.
Publicado: (2025)
por: Baty, Enis, et al.
Publicado: (2025)
MemLens: Benchmarking Multimodal Long-Term Memory in Large Vision-Language Models
por: Ren, Xiyu, et al.
Publicado: (2026)
por: Ren, Xiyu, et al.
Publicado: (2026)
Structured Initialization for Attention in Vision Transformers
por: Zheng, Jianqiao, et al.
Publicado: (2024)
por: Zheng, Jianqiao, et al.
Publicado: (2024)
TACO: Trajectory Aligning Cross-view Optimisation
por: Shore, Tavis, et al.
Publicado: (2026)
por: Shore, Tavis, et al.
Publicado: (2026)
Cross-Modal Dual-Causal Learning for Long-Term Action Recognition
por: Shaowu, Xu, et al.
Publicado: (2025)
por: Shaowu, Xu, et al.
Publicado: (2025)
xLSTM-FER: Enhancing Student Expression Recognition with Extended Vision Long Short-Term Memory Network
por: Huang, Qionghao, et al.
Publicado: (2024)
por: Huang, Qionghao, et al.
Publicado: (2024)
GATE-AD: Graph Attention Network Encoding For Few-Shot Industrial Visual Anomaly Detection
por: Psiris, Aggelos, et al.
Publicado: (2026)
por: Psiris, Aggelos, et al.
Publicado: (2026)
Graph Convolutional Long Short-Term Memory Attention Network for Post-Stroke Compensatory Movement Detection Based on Skeleton Data
por: Fan, Jiaxing, et al.
Publicado: (2025)
por: Fan, Jiaxing, et al.
Publicado: (2025)
Shot-by-Shot: Film-Grammar-Aware Training-Free Audio Description Generation
por: Xie, Junyu, et al.
Publicado: (2025)
por: Xie, Junyu, et al.
Publicado: (2025)
YOLO-FireAD: Efficient Fire Detection via Attention-Guided Inverted Residual Learning and Dual-Pooling Feature Preservation
por: Pan, Weichao, et al.
Publicado: (2025)
por: Pan, Weichao, et al.
Publicado: (2025)
Vision and Intention Boost Large Language Model in Long-Term Action Anticipation
por: Cao, Congqi, et al.
Publicado: (2025)
por: Cao, Congqi, et al.
Publicado: (2025)
Dual-Hybrid Attention Network for Specular Highlight Removal
por: Guo, Xiaojiao, et al.
Publicado: (2024)
por: Guo, Xiaojiao, et al.
Publicado: (2024)
BEV-CV: Birds-Eye-View Transform for Cross-View Geo-Localisation
por: Shore, Tavis, et al.
Publicado: (2023)
por: Shore, Tavis, et al.
Publicado: (2023)
Single-image coherent reconstruction of objects and humans
por: Batra, Sarthak, et al.
Publicado: (2024)
por: Batra, Sarthak, et al.
Publicado: (2024)
More than a Moment: Towards Coherent Sequences of Audio Descriptions
por: Khandelwal, Eshika, et al.
Publicado: (2025)
por: Khandelwal, Eshika, et al.
Publicado: (2025)
Vision KAN: Towards an Attention-Free Backbone for Vision with Kolmogorov-Arnold Networks
por: Yang, Zhuoqin, et al.
Publicado: (2026)
por: Yang, Zhuoqin, et al.
Publicado: (2026)
Interpretable Long-term Action Quality Assessment
por: Dong, Xu, et al.
Publicado: (2024)
por: Dong, Xu, et al.
Publicado: (2024)
Generation Of Colors using Bidirectional Long Short Term Memory Networks
por: Sinha, A.
Publicado: (2023)
por: Sinha, A.
Publicado: (2023)
HYDRA: HYbrid knowledge Distillation and spectral Reconstruction Algorithm for high channel hyperspectral camera applications
por: Thirgood, Christopher, et al.
Publicado: (2025)
por: Thirgood, Christopher, et al.
Publicado: (2025)
FeatureSLAM: Feature-enriched 3D gaussian splatting SLAM in real time
por: Thirgood, Christopher, et al.
Publicado: (2026)
por: Thirgood, Christopher, et al.
Publicado: (2026)
AD-DINO: Attention-Dynamic DINO for Distance-Aware Embodied Reference Understanding
por: Guo, Hao, et al.
Publicado: (2024)
por: Guo, Hao, et al.
Publicado: (2024)
Personalized Image Descriptions from Attention Sequences
por: Xue, Ruoyu, et al.
Publicado: (2025)
por: Xue, Ruoyu, et al.
Publicado: (2025)
Artifacts and Attention Sinks: Structured Approximations for Efficient Vision Transformers
por: Lu, Andrew, et al.
Publicado: (2025)
por: Lu, Andrew, et al.
Publicado: (2025)
Dual Interaction Network with Cross-Image Attention for Medical Image Segmentation
por: Noh, Jeonghyun, et al.
Publicado: (2025)
por: Noh, Jeonghyun, et al.
Publicado: (2025)
CS-VLM: Compressed Sensing Attention for Efficient Vision-Language Representation Learning
por: Kiruluta, Andrew, et al.
Publicado: (2025)
por: Kiruluta, Andrew, et al.
Publicado: (2025)
MedFormer: Hierarchical Medical Vision Transformer with Content-Aware Dual Sparse Selection Attention
por: Xia, Zunhui, et al.
Publicado: (2025)
por: Xia, Zunhui, et al.
Publicado: (2025)
Ejemplares similares
-
DistinctAD: Distinctive Audio Description Generation in Contexts
por: Fang, Bo, et al.
Publicado: (2024) -
AutoAD-Zero: A Training-Free Framework for Zero-Shot Audio Description
por: Xie, Junyu, et al.
Publicado: (2024) -
LLM-AD: Large Language Model based Audio Description System
por: Chu, Peng, et al.
Publicado: (2024) -
FocusedAD: Character-centric Movie Audio Description
por: Ye, Xiaojun, et al.
Publicado: (2025) -
Evaluating Design Video Generation: Metrics for Compositional Fidelity
por: Deganutti, Adrienne, et al.
Publicado: (2026)