FocusedAD: Character-centric Movie Audio Description
Fuente:
arXiv
Saved in:
| Main Authors: | Ye, Xiaojun, Wang, Chun, Song, Yiren, Zhou, Sheng, Li, Liangcheng, Bu, Jiajun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FoR-Net: Learning to Focus on Hard Regions for Efficient Semantic Segmentation
by: Chan, Sheng-Wei, et al.
Published: (2026)
by: Chan, Sheng-Wei, et al.
Published: (2026)
Cinéaste: A Fine-grained Contextual Movie Question Answering Benchmark
by: Shah, Nisarg A., et al.
Published: (2025)
by: Shah, Nisarg A., et al.
Published: (2025)
Look, Listen, and Answer: Overcoming Biases for Audio-Visual Question Answering
by: Ma, Jie, et al.
Published: (2024)
by: Ma, Jie, et al.
Published: (2024)
Co-Speech Gesture and Facial Expression Generation for Non-Photorealistic 3D Characters
by: Omine, Taisei, et al.
Published: (2025)
by: Omine, Taisei, et al.
Published: (2025)
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding
by: Han, Yudong, et al.
Published: (2024)
by: Han, Yudong, et al.
Published: (2024)
Category-Agnostic Neural Object Rigging
by: He, Guangzhao, et al.
Published: (2025)
by: He, Guangzhao, et al.
Published: (2025)
Breaking the Resource Wall: Geometry-Guided Sequence Modeling for Efficient Semantic Segmentation
by: Chan, Sheng-Wei, et al.
Published: (2026)
by: Chan, Sheng-Wei, et al.
Published: (2026)
U-Net-Like Spiking Neural Networks for Single Image Dehazing
by: Li, Huibin, et al.
Published: (2025)
by: Li, Huibin, et al.
Published: (2025)
StoryMovie: A Dataset for Semantic Alignment of Visual Stories with Movie Scripts and Subtitles
by: Oliveira, Daniel, et al.
Published: (2026)
by: Oliveira, Daniel, et al.
Published: (2026)
RGB-Only Gaussian Splatting SLAM for Unbounded Outdoor Scenes
by: Yu, Sicheng, et al.
Published: (2025)
by: Yu, Sicheng, et al.
Published: (2025)
EmoVerse: A MLLMs-Driven Emotion Representation Dataset for Interpretable Visual Emotion Analysis
by: Guo, Yijie, et al.
Published: (2025)
by: Guo, Yijie, et al.
Published: (2025)
Birth and Death of a Rose
by: Geng, Chen, et al.
Published: (2024)
by: Geng, Chen, et al.
Published: (2024)
Web-Scale Collection of Video Data for 4D Animal Reconstruction
by: Zhao, Brian Nlong, et al.
Published: (2025)
by: Zhao, Brian Nlong, et al.
Published: (2025)
MINT: Mitigating Hallucinations in Large Vision-Language Models via Token Reduction
by: Wang, Chao, et al.
Published: (2025)
by: Wang, Chao, et al.
Published: (2025)
ERNet: Efficient Non-Rigid Registration Network for Point Sequences
by: He, Guangzhao, et al.
Published: (2025)
by: He, Guangzhao, et al.
Published: (2025)
SIFThinker: Spatially-Aware Image Focus for Visual Reasoning
by: Chen, Zhangquan, et al.
Published: (2025)
by: Chen, Zhangquan, et al.
Published: (2025)
LEGO: LoRA-Enabled Generator-Oriented Framework for Synthetic Image Detection
by: Xiao, Yutong, et al.
Published: (2026)
by: Xiao, Yutong, et al.
Published: (2026)
See-through: Single-image Layer Decomposition for Anime Characters
by: Lin, Jian, et al.
Published: (2026)
by: Lin, Jian, et al.
Published: (2026)
SimWorld: A Unified Benchmark for Simulator-Conditioned Scene Generation via World Model
by: Li, Xinqing, et al.
Published: (2025)
by: Li, Xinqing, et al.
Published: (2025)
Conscious Gaze: Adaptive Attention Mechanisms for Hallucination Mitigation in Vision-Language Models
by: Bu, Weijue, et al.
Published: (2025)
by: Bu, Weijue, et al.
Published: (2025)
Single-Shot Metric Depth from Focused Plenoptic Cameras
by: Lasheras-Hernandez, Blanca, et al.
Published: (2024)
by: Lasheras-Hernandez, Blanca, et al.
Published: (2024)
AnySurf: Any Surface Generation with Directed Edge
by: Shi, Wenda, et al.
Published: (2026)
by: Shi, Wenda, et al.
Published: (2026)
CoT4AD: A Vision-Language-Action Model with Explicit Chain-of-Thought Reasoning for Autonomous Driving
by: Wang, Zhaohui, et al.
Published: (2025)
by: Wang, Zhaohui, et al.
Published: (2025)
ReLKD: Inter-Class Relation Learning with Knowledge Distillation for Generalized Category Discovery
by: Zhou, Fang, et al.
Published: (2025)
by: Zhou, Fang, et al.
Published: (2025)
EventFormer: A Node-graph Hierarchical Attention Transformer for Action-centric Video Event Prediction
by: Su, Qile, et al.
Published: (2025)
by: Su, Qile, et al.
Published: (2025)
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
by: Raoufi, Behnam, et al.
Published: (2025)
by: Raoufi, Behnam, et al.
Published: (2025)
OmniVideo-R1: Reinforcing Audio-visual Reasoning with Query Intention and Modality Attention
by: Chen, Zhangquan, et al.
Published: (2026)
by: Chen, Zhangquan, et al.
Published: (2026)
Car Object Counting and Position Estimation via Extension of the CLIP-EBC Framework
by: Jung, Seoik, et al.
Published: (2025)
by: Jung, Seoik, et al.
Published: (2025)
An Evaluation of a Visual Question Answering Strategy for Zero-shot Facial Expression Recognition in Still Images
by: Castrillón-Santana, Modesto, et al.
Published: (2025)
by: Castrillón-Santana, Modesto, et al.
Published: (2025)
A Challenging Benchmark of Anime Style Recognition
by: Li, Haotang, et al.
Published: (2022)
by: Li, Haotang, et al.
Published: (2022)
Textual and Visual Guided Task Adaptation for Source-Free Cross-Domain Few-Shot Segmentation
by: Liu, Jianming, et al.
Published: (2025)
by: Liu, Jianming, et al.
Published: (2025)
FUSE-Flow: Scalable Real-Time Multi-View Point Cloud Reconstruction Using Confidence
by: Sun, Chentian
Published: (2026)
by: Sun, Chentian
Published: (2026)
YotoR-You Only Transform One Representation
by: Villa, José Ignacio Díaz, et al.
Published: (2024)
by: Villa, José Ignacio Díaz, et al.
Published: (2024)
MSTA3D: Multi-scale Twin-attention for 3D Instance Segmentation
by: Tran, Duc Dang Trung, et al.
Published: (2024)
by: Tran, Duc Dang Trung, et al.
Published: (2024)
GMAC: Global Multi-View Constraint for Automatic Multi-Camera Extrinsic Calibration
by: Sun, Chentian
Published: (2026)
by: Sun, Chentian
Published: (2026)
Escaping The Big Data Paradigm in Self-Supervised Representation Learning
by: García, Carlos Vélez, et al.
Published: (2025)
by: García, Carlos Vélez, et al.
Published: (2025)
A self-supervised cyclic neural-analytic approach for novel view synthesis and 3D reconstruction
by: Costea, Dragos, et al.
Published: (2025)
by: Costea, Dragos, et al.
Published: (2025)
NAC-TCN: Temporal Convolutional Networks with Causal Dilated Neighborhood Attention for Emotion Understanding
by: Mehta, Alexander, et al.
Published: (2023)
by: Mehta, Alexander, et al.
Published: (2023)
SemanticHuman-HD: High-Resolution Semantic Disentangled 3D Human Generation
by: Zheng, Peng, et al.
Published: (2024)
by: Zheng, Peng, et al.
Published: (2024)
Visual Enhanced Depth Scaling for Multimodal Latent Reasoning
by: Han, Yudong, et al.
Published: (2026)
by: Han, Yudong, et al.
Published: (2026)
Similar Items
-
FoR-Net: Learning to Focus on Hard Regions for Efficient Semantic Segmentation
by: Chan, Sheng-Wei, et al.
Published: (2026) -
Cinéaste: A Fine-grained Contextual Movie Question Answering Benchmark
by: Shah, Nisarg A., et al.
Published: (2025) -
Look, Listen, and Answer: Overcoming Biases for Audio-Visual Question Answering
by: Ma, Jie, et al.
Published: (2024) -
Co-Speech Gesture and Facial Expression Generation for Non-Photorealistic 3D Characters
by: Omine, Taisei, et al.
Published: (2025) -
DynFocus: Dynamic Cooperative Network Empowers LLMs with Video Understanding
by: Han, Yudong, et al.
Published: (2024)