DistinctAD: Distinctive Audio Description Generation in Contexts
Fuente:
arXiv
Saved in:
| Main Authors: | Fang, Bo, Wu, Wenhao, Wu, Qiangqiang, Song, Yuxin, Chan, Antoni B. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders
by: Fang, Bo, et al.
Published: (2025)
by: Fang, Bo, et al.
Published: (2025)
ViSS-R1: Self-Supervised Reinforcement Video Reasoning
by: Fang, Bo, et al.
Published: (2025)
by: Fang, Bo, et al.
Published: (2025)
Learning Tracking Representations from Single Point Annotations
by: Wu, Qiangqiang, et al.
Published: (2024)
by: Wu, Qiangqiang, et al.
Published: (2024)
Group-based Distinctive Image Captioning with Memory Difference Encoding and Attention
by: Wang, Jiuniu, et al.
Published: (2025)
by: Wang, Jiuniu, et al.
Published: (2025)
Exploring the Distinctiveness and Fidelity of the Descriptions Generated by Large Vision-Language Models
by: Huang, Yuhang, et al.
Published: (2024)
by: Huang, Yuhang, et al.
Published: (2024)
Robust Zero-Shot Crowd Counting and Localization With Adaptive Resolution SAM
by: Wan, Jia, et al.
Published: (2024)
by: Wan, Jia, et al.
Published: (2024)
M2P: Improving Visual Foundation Models with Mask-to-Point Weakly-Supervised Learning for Dense Point Tracking
by: Wu, Qiangqiang, et al.
Published: (2026)
by: Wu, Qiangqiang, et al.
Published: (2026)
DropMAE: Learning Representations via Masked Autoencoders with Spatial-Attention Dropout for Temporal Matching Tasks
by: Wu, Qiangqiang, et al.
Published: (2023)
by: Wu, Qiangqiang, et al.
Published: (2023)
FocusedAD: Character-centric Movie Audio Description
by: Ye, Xiaojun, et al.
Published: (2025)
by: Ye, Xiaojun, et al.
Published: (2025)
LLM-AD: Large Language Model based Audio Description System
by: Chu, Peng, et al.
Published: (2024)
by: Chu, Peng, et al.
Published: (2024)
Guided Self-attention: Find the Generalized Necessarily Distinct Vectors for Grain Size Grading
by: Gao, Fang, et al.
Published: (2024)
by: Gao, Fang, et al.
Published: (2024)
DANTE-AD: Dual-Vision Attention Network for Long-Term Audio Description
by: Deganutti, Adrienne, et al.
Published: (2025)
by: Deganutti, Adrienne, et al.
Published: (2025)
This Looks Distinctly Like That: Grounding Interpretable Recognition in Stiefel Geometry against Neural Collapse
by: Jia, Junhao, et al.
Published: (2026)
by: Jia, Junhao, et al.
Published: (2026)
AutoAD-Zero: A Training-Free Framework for Zero-Shot Audio Description
by: Xie, Junyu, et al.
Published: (2024)
by: Xie, Junyu, et al.
Published: (2024)
NowYouSee Me: Context-Aware Automatic Audio Description
by: Lee, Seon-Ho, et al.
Published: (2024)
by: Lee, Seon-Ho, et al.
Published: (2024)
A Fixed-Point Approach to Unified Prompt-Based Counting
by: Lin, Wei, et al.
Published: (2024)
by: Lin, Wei, et al.
Published: (2024)
3D Crowd Counting via Geometric Attention-guided Multi-View Fusion
by: Zhang, Qi, et al.
Published: (2020)
by: Zhang, Qi, et al.
Published: (2020)
A Multimodal In-Context Tuning Approach for E-Commerce Product Description Generation
by: Li, Yunxin, et al.
Published: (2024)
by: Li, Yunxin, et al.
Published: (2024)
Pseudo-RIS: Distinctive Pseudo-supervision Generation for Referring Image Segmentation
by: Yu, Seonghoon, et al.
Published: (2024)
by: Yu, Seonghoon, et al.
Published: (2024)
Temporal Unlearnable Examples: Preventing Personal Video Data from Unauthorized Exploitation by Object Tracking
by: Wu, Qiangqiang, et al.
Published: (2025)
by: Wu, Qiangqiang, et al.
Published: (2025)
Style Composition within Distinct LoRA modules for Traditional Art
by: Lee, Jaehyun, et al.
Published: (2025)
by: Lee, Jaehyun, et al.
Published: (2025)
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs
by: Zhao, Jiahe, et al.
Published: (2025)
by: Zhao, Jiahe, et al.
Published: (2025)
GPT4Vis: What Can GPT-4 Do for Zero-shot Visual Recognition?
by: Wu, Wenhao, et al.
Published: (2023)
by: Wu, Wenhao, et al.
Published: (2023)
Improving Joint Audio-Video Generation with Cross-Modal Context Learning
by: Ma, Bingqi, et al.
Published: (2026)
by: Ma, Bingqi, et al.
Published: (2026)
GenAD: Generalized Predictive Model for Autonomous Driving
by: Yang, Jiazhi, et al.
Published: (2024)
by: Yang, Jiazhi, et al.
Published: (2024)
AnyAD: Unified Any-Modality Anomaly Detection in Incomplete Multi-Sequence MRI
by: Wu, Changwei, et al.
Published: (2025)
by: Wu, Changwei, et al.
Published: (2025)
Point-to-Region Loss for Semi-Supervised Point-Based Crowd Counting
by: Lin, Wei, et al.
Published: (2025)
by: Lin, Wei, et al.
Published: (2025)
Density-based Object Detection in Crowded Scenes
by: Zhao, Chenyang, et al.
Published: (2025)
by: Zhao, Chenyang, et al.
Published: (2025)
GeneQuery: A General QA-based Framework for Spatial Gene Expression Predictions from Histology Images
by: Xiong, Ying, et al.
Published: (2024)
by: Xiong, Ying, et al.
Published: (2024)
Open-Attribute Recognition for Person Retrieval: Finding People Through Distinctive and Novel Attributes
by: Park, Minjeong, et al.
Published: (2025)
by: Park, Minjeong, et al.
Published: (2025)
Scone: Bridging Composition and Distinction in Subject-Driven Image Generation via Unified Understanding-Generation Modeling
by: Wang, Yuran, et al.
Published: (2025)
by: Wang, Yuran, et al.
Published: (2025)
FreeDiff: Progressive Frequency Truncation for Image Editing with Diffusion Models
by: Wu, Wei, et al.
Published: (2024)
by: Wu, Wei, et al.
Published: (2024)
Baton: Explicit Semantic Blueprints for Joint Video-Audio Generation
by: Tu, Shuyuan, et al.
Published: (2026)
by: Tu, Shuyuan, et al.
Published: (2026)
EIANet: A Novel Domain Adaptation Approach to Maximize Class Distinction with Neural Collapse Principles
by: Pan, Zicheng, et al.
Published: (2024)
by: Pan, Zicheng, et al.
Published: (2024)
Landmarks Are Alike Yet Distinct: Harnessing Similarity and Individuality for One-Shot Medical Landmark Detection
by: He, Xu, et al.
Published: (2025)
by: He, Xu, et al.
Published: (2025)
Shot-by-Shot: Film-Grammar-Aware Training-Free Audio Description Generation
by: Xie, Junyu, et al.
Published: (2025)
by: Xie, Junyu, et al.
Published: (2025)
GenAD: Generative End-to-End Autonomous Driving
by: Zheng, Wenzhao, et al.
Published: (2024)
by: Zheng, Wenzhao, et al.
Published: (2024)
Beyond Boundary Frames: Context-Centric Video Interpolation with Audio-Visual Semantics
by: Deng, Yuchen, et al.
Published: (2025)
by: Deng, Yuchen, et al.
Published: (2025)
Distinctive Image Captioning: Leveraging Ground Truth Captions in CLIP Guided Reinforcement Learning
by: Chaffin, Antoine, et al.
Published: (2024)
by: Chaffin, Antoine, et al.
Published: (2024)
CLIP3D-AD: Extending CLIP for 3D Few-Shot Anomaly Detection with Multi-View Images Generation
by: Zuo, Zuo, et al.
Published: (2024)
by: Zuo, Zuo, et al.
Published: (2024)
Similar Items
-
Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders
by: Fang, Bo, et al.
Published: (2025) -
ViSS-R1: Self-Supervised Reinforcement Video Reasoning
by: Fang, Bo, et al.
Published: (2025) -
Learning Tracking Representations from Single Point Annotations
by: Wu, Qiangqiang, et al.
Published: (2024) -
Group-based Distinctive Image Captioning with Memory Difference Encoding and Attention
by: Wang, Jiuniu, et al.
Published: (2025) -
Exploring the Distinctiveness and Fidelity of the Descriptions Generated by Large Vision-Language Models
by: Huang, Yuhang, et al.
Published: (2024)