Transformer with Controlled Attention for Synchronous Motion Captioning
Fuente:
arXiv
Saved in:
| Main Authors: | Radouane, Karim, Ranwez, Sylvie, Lagarde, Julien, Tchechmedjiev, Andon |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Guided Attention for Interpretable Motion Captioning
by: Radouane, Karim, et al.
Published: (2023)
by: Radouane, Karim, et al.
Published: (2023)
MB-ORES: A Multi-Branch Object Reasoner for Visual Grounding in Remote Sensing
by: Radouane, Karim, et al.
Published: (2025)
by: Radouane, Karim, et al.
Published: (2025)
Words in Motion: Extracting Interpretable Control Vectors for Motion Transformers
by: Tas, Omer Sahin, et al.
Published: (2024)
by: Tas, Omer Sahin, et al.
Published: (2024)
Brazilian Portuguese Image Captioning with Transformers: A Study on Cross-Native-Translated Dataset
by: Bromonschenkel, Gabriel, et al.
Published: (2026)
by: Bromonschenkel, Gabriel, et al.
Published: (2026)
Head Pursuit: Probing Attention Specialization in Multimodal Transformers
by: Basile, Lorenzo, et al.
Published: (2025)
by: Basile, Lorenzo, et al.
Published: (2025)
CAT: Circular-Convolutional Attention for Sub-Quadratic Transformers
by: Yamada, Yoshihiro
Published: (2025)
by: Yamada, Yoshihiro
Published: (2025)
Image-Caption Encoding for Improving Zero-Shot Generalization
by: Yu, Eric Yang, et al.
Published: (2024)
by: Yu, Eric Yang, et al.
Published: (2024)
Linear Alignment of Vision-language Models for Image Captioning
by: Paischer, Fabian, et al.
Published: (2023)
by: Paischer, Fabian, et al.
Published: (2023)
Multimodal Arabic Captioning with Interpretable Visual Concept Integration
by: Elchafei, Passant, et al.
Published: (2025)
by: Elchafei, Passant, et al.
Published: (2025)
Wolf: Dense Video Captioning with a World Summarization Framework
by: Li, Boyi, et al.
Published: (2024)
by: Li, Boyi, et al.
Published: (2024)
From Pixels to Prose: A Large Dataset of Dense Image Captions
by: Singla, Vasu, et al.
Published: (2024)
by: Singla, Vasu, et al.
Published: (2024)
CIC-BART-SSA: Controllable Image Captioning with Structured Semantic Augmentation
by: Basioti, Kalliopi, et al.
Published: (2024)
by: Basioti, Kalliopi, et al.
Published: (2024)
Whats in a Video: Factorized Autoregressive Decoding for Online Dense Video Captioning
by: Piergiovanni, AJ, et al.
Published: (2024)
by: Piergiovanni, AJ, et al.
Published: (2024)
Bridging the Visual Gap: Fine-Tuning Multimodal Models with Knowledge-Adapted Captions
by: Yanuka, Moran, et al.
Published: (2024)
by: Yanuka, Moran, et al.
Published: (2024)
BioCAP: Exploiting Synthetic Captions Beyond Labels in Biological Foundation Models
by: Zhang, Ziheng, et al.
Published: (2025)
by: Zhang, Ziheng, et al.
Published: (2025)
CAPability: A Comprehensive Visual Caption Benchmark for Evaluating Both Correctness and Thoroughness
by: Liu, Zhihang, et al.
Published: (2025)
by: Liu, Zhihang, et al.
Published: (2025)
FigCaps-HF: A Figure-to-Caption Generative Framework and Benchmark with Human Feedback
by: Singh, Ashish, et al.
Published: (2023)
by: Singh, Ashish, et al.
Published: (2023)
TrafficVLM: A Controllable Visual Language Model for Traffic Video Captioning
by: Dinh, Quang Minh, et al.
Published: (2024)
by: Dinh, Quang Minh, et al.
Published: (2024)
Shotluck Holmes: A Family of Efficient Small-Scale Large Language Vision Models For Video Captioning and Summarization
by: Luo, Richard, et al.
Published: (2024)
by: Luo, Richard, et al.
Published: (2024)
Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M)
by: Merchant, Nicholas, et al.
Published: (2025)
by: Merchant, Nicholas, et al.
Published: (2025)
Unleashing Text-to-Image Diffusion Prior for Zero-Shot Image Captioning
by: Luo, Jianjie, et al.
Published: (2024)
by: Luo, Jianjie, et al.
Published: (2024)
SynC: Synthetic Image Caption Dataset Refinement with One-to-many Mapping for Zero-shot Image Captioning
by: Kim, Si-Woo, et al.
Published: (2025)
by: Kim, Si-Woo, et al.
Published: (2025)
LLM as a Neural Architect: Controlled Generation of Image Captioning Models Under Strict API Contracts
by: Jesani, Krunal, et al.
Published: (2025)
by: Jesani, Krunal, et al.
Published: (2025)
Reinforced Attention Learning
by: Li, Bangzheng, et al.
Published: (2026)
by: Li, Bangzheng, et al.
Published: (2026)
Modeling Caption Diversity in Contrastive Vision-Language Pretraining
by: Lavoie, Samuel, et al.
Published: (2024)
by: Lavoie, Samuel, et al.
Published: (2024)
ALOHa: A New Measure for Hallucination in Captioning Models
by: Petryk, Suzanne, et al.
Published: (2024)
by: Petryk, Suzanne, et al.
Published: (2024)
MTA: Multimodal Task Alignment for BEV Perception and Captioning
by: Ma, Yunsheng, et al.
Published: (2024)
by: Ma, Yunsheng, et al.
Published: (2024)
BAD: Bidirectional Auto-regressive Diffusion for Text-to-Motion Generation
by: Hosseyni, S. Rohollah, et al.
Published: (2024)
by: Hosseyni, S. Rohollah, et al.
Published: (2024)
Regional Attention-Enhanced Swin Transformer for Clinically Relevant Medical Image Captioning
by: Naz, Zubia, et al.
Published: (2025)
by: Naz, Zubia, et al.
Published: (2025)
Listen Then See: Video Alignment with Speaker Attention
by: Agrawal, Aviral, et al.
Published: (2024)
by: Agrawal, Aviral, et al.
Published: (2024)
OSCaR: Object State Captioning and State Change Representation
by: Nguyen, Nguyen, et al.
Published: (2024)
by: Nguyen, Nguyen, et al.
Published: (2024)
An Examination of the Robustness of Reference-Free Image Captioning Evaluation Metrics
by: Ahmadi, Saba, et al.
Published: (2023)
by: Ahmadi, Saba, et al.
Published: (2023)
AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers
by: Achtibat, Reduan, et al.
Published: (2024)
by: Achtibat, Reduan, et al.
Published: (2024)
Transformer-VQ: Linear-Time Transformers via Vector Quantization
by: Lingle, Lucas D.
Published: (2023)
by: Lingle, Lucas D.
Published: (2023)
General Transform: A Unified Framework for Adaptive Transform to Enhance Representations
by: Budiutama, Gekko, et al.
Published: (2025)
by: Budiutama, Gekko, et al.
Published: (2025)
MoTe: Learning Motion-Text Diffusion Model for Multiple Generation Tasks
by: Wu, Yiming, et al.
Published: (2024)
by: Wu, Yiming, et al.
Published: (2024)
The ART of Composition: Attention-Regularized Training for Compositional Visual Grounding
by: Luo, Jiayun, et al.
Published: (2024)
by: Luo, Jiayun, et al.
Published: (2024)
A Survey on Transformer Compression
by: Tang, Yehui, et al.
Published: (2024)
by: Tang, Yehui, et al.
Published: (2024)
ViPCap: Retrieval Text-Based Visual Prompts for Lightweight Image Captioning
by: Kim, Taewhan, et al.
Published: (2024)
by: Kim, Taewhan, et al.
Published: (2024)
Towards Better Multi-head Attention via Channel-wise Sample Permutation
by: Yuan, Shen, et al.
Published: (2024)
by: Yuan, Shen, et al.
Published: (2024)
Similar Items
-
Guided Attention for Interpretable Motion Captioning
by: Radouane, Karim, et al.
Published: (2023) -
MB-ORES: A Multi-Branch Object Reasoner for Visual Grounding in Remote Sensing
by: Radouane, Karim, et al.
Published: (2025) -
Words in Motion: Extracting Interpretable Control Vectors for Motion Transformers
by: Tas, Omer Sahin, et al.
Published: (2024) -
Brazilian Portuguese Image Captioning with Transformers: A Study on Cross-Native-Translated Dataset
by: Bromonschenkel, Gabriel, et al.
Published: (2026) -
Head Pursuit: Probing Attention Specialization in Multimodal Transformers
by: Basile, Lorenzo, et al.
Published: (2025)