Q-Adapter: Visual Query Adapter for Extracting Textually-related Features in Video Captioning
Fuente:
arXiv
Saved in:
| Main Authors: | Chen, Junan, Nguyen, Trung Thanh, Komamizu, Takahiro, Ide, Ichiro |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
One-Stage Open-Vocabulary Temporal Action Detection Leveraging Temporal Multi-scale and Action Label Features
by: Nguyen, Trung Thanh, et al.
Published: (2024)
by: Nguyen, Trung Thanh, et al.
Published: (2024)
Action Selection Learning for Multi-label Multi-view Action Recognition
by: Nguyen, Trung Thanh, et al.
Published: (2024)
by: Nguyen, Trung Thanh, et al.
Published: (2024)
MultiTSF: Transformer-based Sensor Fusion for Human-Centric Multi-view and Multi-modal Action Recognition
by: Nguyen, Trung Thanh, et al.
Published: (2025)
by: Nguyen, Trung Thanh, et al.
Published: (2025)
View-aware Cross-modal Distillation for Multi-view Action Recognition
by: Nguyen, Trung Thanh, et al.
Published: (2025)
by: Nguyen, Trung Thanh, et al.
Published: (2025)
MultiSensor-Home: A Wide-area Multi-modal Multi-view Dataset for Action Recognition and Transformer-based Sensor Fusion
by: Nguyen, Trung Thanh, et al.
Published: (2025)
by: Nguyen, Trung Thanh, et al.
Published: (2025)
ForestMamba: Sparse Mamba with Geometry-guided Queries for 3D Forest Point Cloud Segmentation
by: Nguyen, Trung Thanh, et al.
Published: (2026)
by: Nguyen, Trung Thanh, et al.
Published: (2026)
Tracking Small Birds by Detection Candidate Region Filtering and Detection History-aware Association
by: Liu, Tingwei, et al.
Published: (2024)
by: Liu, Tingwei, et al.
Published: (2024)
Multi-proposal Collaboration and Multi-task Training for Weakly-supervised Video Moment Retrieval
by: Zhang, Bolin, et al.
Published: (2026)
by: Zhang, Bolin, et al.
Published: (2026)
CapS-Adapter: Caption-based MultiModal Adapter in Zero-Shot Classification
by: Wang, Qijie, et al.
Published: (2024)
by: Wang, Qijie, et al.
Published: (2024)
Static and Dynamic Graph Alignment Network for Temporal Video Grounding
by: Hu, Zhanjie, et al.
Published: (2026)
by: Hu, Zhanjie, et al.
Published: (2026)
CLIP-Adapter: Better Vision-Language Models with Feature Adapters
by: Gao, Peng, et al.
Published: (2021)
by: Gao, Peng, et al.
Published: (2021)
LQ-Adapter: ViT-Adapter with Learnable Queries for Gallbladder Cancer Detection from Ultrasound Image
by: Madan, Chetan, et al.
Published: (2024)
by: Madan, Chetan, et al.
Published: (2024)
Small Object Detection for Birds with Swin Transformer
by: Huo, Da, et al.
Published: (2025)
by: Huo, Da, et al.
Published: (2025)
LLMVA-GEBC: Large Language Model with Video Adapter for Generic Event Boundary Captioning
by: Tang, Yolo Yunlong, et al.
Published: (2023)
by: Tang, Yolo Yunlong, et al.
Published: (2023)
I2V-Adapter: A General Image-to-Video Adapter for Diffusion Models
by: Guo, Xun, et al.
Published: (2023)
by: Guo, Xun, et al.
Published: (2023)
Hyperspectral Adapter for Object Tracking based on Hyperspectral Video
by: Gao, Long, et al.
Published: (2025)
by: Gao, Long, et al.
Published: (2025)
EVCtrl: Efficient Control Adapter for Visual Generation
by: Yang, Zixiang, et al.
Published: (2025)
by: Yang, Zixiang, et al.
Published: (2025)
ResAdapter: Domain Consistent Resolution Adapter for Diffusion Models
by: Cheng, Jiaxiang, et al.
Published: (2024)
by: Cheng, Jiaxiang, et al.
Published: (2024)
BT-Adapter: Video Conversation is Feasible Without Video Instruction Tuning
by: Liu, Ruyang, et al.
Published: (2023)
by: Liu, Ruyang, et al.
Published: (2023)
LLMs Meet Long Video: Advancing Long Video Question Answering with An Interactive Visual Adapter in LLMs
by: Li, Yunxin, et al.
Published: (2024)
by: Li, Yunxin, et al.
Published: (2024)
Dyn-Adapter: Towards Disentangled Representation for Efficient Visual Recognition
by: Zhang, Yurong, et al.
Published: (2024)
by: Zhang, Yurong, et al.
Published: (2024)
DeepFake-Adapter: Dual-Level Adapter for DeepFake Detection
by: Shao, Rui, et al.
Published: (2023)
by: Shao, Rui, et al.
Published: (2023)
Frequency Adapter with SAM for Generalized Medical Image Segmentation
by: Bui, Phuoc-Nguyen, et al.
Published: (2026)
by: Bui, Phuoc-Nguyen, et al.
Published: (2026)
RAP: Efficient Text-Video Retrieval with Sparse-and-Correlated Adapter
by: Cao, Meng, et al.
Published: (2024)
by: Cao, Meng, et al.
Published: (2024)
QueryAdapter: Rapid Adaptation of Vision-Language Models in Response to Natural Language Queries
by: Chapman, Nicolas Harvey, et al.
Published: (2025)
by: Chapman, Nicolas Harvey, et al.
Published: (2025)
PickStyle: Video-to-Video Style Transfer with Context-Style Adapters
by: Mehraban, Soroush, et al.
Published: (2025)
by: Mehraban, Soroush, et al.
Published: (2025)
MV-Adapter: Multimodal Video Transfer Learning for Video Text Retrieval
by: Jin, Xiaojie, et al.
Published: (2023)
by: Jin, Xiaojie, et al.
Published: (2023)
Inv-Adapter: ID Customization Generation via Image Inversion and Lightweight Adapter
by: Xing, Peng, et al.
Published: (2024)
by: Xing, Peng, et al.
Published: (2024)
T-Gated Adapter: A Lightweight Temporal Adapter for Vision-Language Medical Segmentation
by: Khadka, Pranjal
Published: (2026)
by: Khadka, Pranjal
Published: (2026)
Attention Lattice Adapter: Visual Explanation Generation for Visual Foundation Model
by: Hirano, Shinnosuke, et al.
Published: (2025)
by: Hirano, Shinnosuke, et al.
Published: (2025)
Hierarchical Awareness Adapters with Hybrid Pyramid Feature Fusion for Dense Depth Prediction
by: Su, Wuqi, et al.
Published: (2026)
by: Su, Wuqi, et al.
Published: (2026)
REEF: Relevance-Aware and Efficient LLM Adapter for Video Understanding
by: Reza, Sakib, et al.
Published: (2025)
by: Reza, Sakib, et al.
Published: (2025)
FE-Adapter: Adapting Image-based Emotion Classifiers to Videos
by: Gowda, Shreyank N, et al.
Published: (2024)
by: Gowda, Shreyank N, et al.
Published: (2024)
FontAdapter: Instant Font Adaptation in Visual Text Generation
by: Koo, Myungkyu, et al.
Published: (2025)
by: Koo, Myungkyu, et al.
Published: (2025)
RelationAdapter: Learning and Transferring Visual Relation with Diffusion Transformers
by: Gong, Yan, et al.
Published: (2025)
by: Gong, Yan, et al.
Published: (2025)
MoSA: Mixture of Sparse Adapters for Visual Efficient Tuning
by: Zhang, Qizhe, et al.
Published: (2023)
by: Zhang, Qizhe, et al.
Published: (2023)
Mod-Adapter: Tuning-Free and Versatile Multi-concept Personalization via Modulation Adapter
by: Zhong, Weizhi, et al.
Published: (2025)
by: Zhong, Weizhi, et al.
Published: (2025)
Motion-Adapter: A Diffusion Model Adapter for Text-to-Motion Generation of Compound Actions
by: Jiang, Yue, et al.
Published: (2026)
by: Jiang, Yue, et al.
Published: (2026)
Decoupled Prompt-Adapter Tuning for Continual Activity Recognition
by: Fu, Di, et al.
Published: (2024)
by: Fu, Di, et al.
Published: (2024)
TransAdapter: Vision Transformer for Feature-Centric Unsupervised Domain Adaptation
by: Doruk, A. Enes, et al.
Published: (2024)
by: Doruk, A. Enes, et al.
Published: (2024)
Similar Items
-
One-Stage Open-Vocabulary Temporal Action Detection Leveraging Temporal Multi-scale and Action Label Features
by: Nguyen, Trung Thanh, et al.
Published: (2024) -
Action Selection Learning for Multi-label Multi-view Action Recognition
by: Nguyen, Trung Thanh, et al.
Published: (2024) -
MultiTSF: Transformer-based Sensor Fusion for Human-Centric Multi-view and Multi-modal Action Recognition
by: Nguyen, Trung Thanh, et al.
Published: (2025) -
View-aware Cross-modal Distillation for Multi-view Action Recognition
by: Nguyen, Trung Thanh, et al.
Published: (2025) -
MultiSensor-Home: A Wide-area Multi-modal Multi-view Dataset for Action Recognition and Transformer-based Sensor Fusion
by: Nguyen, Trung Thanh, et al.
Published: (2025)