Hallucination Localization in Video Captioning
Fuente:
arXiv
Guardado en:
| Autores principales: | Nakada, Shota, Saito, Kazuhiro, Ishikawa, Yuchi, Munakata, Hokuto, Komatsu, Tatsuya, Kondo, Masayoshi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
DETECLAP: Enhancing Audio-Visual Representation Learning with Object Information
por: Nakada, Shota, et al.
Publicado: (2024)
por: Nakada, Shota, et al.
Publicado: (2024)
Lighthouse: A User-Friendly Library for Reproducible Video Moment Retrieval and Highlight Detection
por: Nishimura, Taichi, et al.
Publicado: (2024)
por: Nishimura, Taichi, et al.
Publicado: (2024)
Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos
por: Ishikawa, Yuchi, et al.
Publicado: (2025)
por: Ishikawa, Yuchi, et al.
Publicado: (2025)
On the Audio Hallucinations in Large Audio-Video Language Models
por: Nishimura, Taichi, et al.
Publicado: (2024)
por: Nishimura, Taichi, et al.
Publicado: (2024)
Vision-Language Models Learn Super Images for Efficient Partially Relevant Video Retrieval
por: Nishimura, Taichi, et al.
Publicado: (2023)
por: Nishimura, Taichi, et al.
Publicado: (2023)
Language-based Audio Moment Retrieval
por: Munakata, Hokuto, et al.
Publicado: (2024)
por: Munakata, Hokuto, et al.
Publicado: (2024)
ProLAP: Probabilistic Language-Audio Pre-Training
por: Manabe, Toranosuke, et al.
Publicado: (2025)
por: Manabe, Toranosuke, et al.
Publicado: (2025)
CASTELLA: Long Audio Dataset with Captions and Temporal Boundaries
por: Munakata, Hokuto, et al.
Publicado: (2025)
por: Munakata, Hokuto, et al.
Publicado: (2025)
Cap2Sum: Learning to Summarize Videos by Generating Captions
por: Zhao, Cairong, et al.
Publicado: (2024)
por: Zhao, Cairong, et al.
Publicado: (2024)
Mitigating Image Captioning Hallucinations in Vision-Language Models
por: Zhao, Fei, et al.
Publicado: (2025)
por: Zhao, Fei, et al.
Publicado: (2025)
DetailVerifyBench: A Benchmark for Dense Hallucination Localization in Long Image Captions
por: Wang, Xinran, et al.
Publicado: (2026)
por: Wang, Xinran, et al.
Publicado: (2026)
Listening without Looking: Modality Bias in Audio-Visual Captioning
por: Ishikawa, Yuchi, et al.
Publicado: (2025)
por: Ishikawa, Yuchi, et al.
Publicado: (2025)
Data Collection-free Masked Video Modeling
por: Ishikawa, Yuchi, et al.
Publicado: (2024)
por: Ishikawa, Yuchi, et al.
Publicado: (2024)
PolySmart @ TRECVid 2024 Video Captioning (VTT)
por: Wu, Jiaxin, et al.
Publicado: (2024)
por: Wu, Jiaxin, et al.
Publicado: (2024)
Looking Backward: Streaming Video-to-Video Translation with Feature Banks
por: Liang, Feng, et al.
Publicado: (2024)
por: Liang, Feng, et al.
Publicado: (2024)
Multi Agents Semantic Emotion Aligned Music to Image Generation with Music Derived Captions
por: Shi, Junchang, et al.
Publicado: (2025)
por: Shi, Junchang, et al.
Publicado: (2025)
Voices, Faces, and Feelings: Multi-modal Emotion-Cognition Captioning for Mental Health Understanding
por: Zhou, Zhiyuan, et al.
Publicado: (2026)
por: Zhou, Zhiyuan, et al.
Publicado: (2026)
NewsCaption: Named-Entity aware Captioning for Out-of-Context Media
por: Singh, Anurag, et al.
Publicado: (2024)
por: Singh, Anurag, et al.
Publicado: (2024)
Unsupervised Ego- and Exo-centric Dense Procedural Activity Captioning via Gaze Consensus Adaptation
por: Shi, Zhaofeng, et al.
Publicado: (2025)
por: Shi, Zhaofeng, et al.
Publicado: (2025)
Edit As You Wish: Video Caption Editing with Multi-grained User Control
por: Yao, Linli, et al.
Publicado: (2023)
por: Yao, Linli, et al.
Publicado: (2023)
Divide and Conquer: Multimodal Video Deepfake Detection via Cross-Modal Fusion and Localization
por: Li, Qingcao, et al.
Publicado: (2026)
por: Li, Qingcao, et al.
Publicado: (2026)
Neural Compression of 360-Degree Equirectangular Videos using Quality Parameter Adaptation
por: Arai, Daichi, et al.
Publicado: (2025)
por: Arai, Daichi, et al.
Publicado: (2025)
Towards Alleviating Text-to-Image Retrieval Hallucination for CLIP in Zero-shot Learning
por: Wang, Hanyao, et al.
Publicado: (2024)
por: Wang, Hanyao, et al.
Publicado: (2024)
Video Summarization: Towards Entity-Aware Captions
por: Ayyubi, Hammad A., et al.
Publicado: (2023)
por: Ayyubi, Hammad A., et al.
Publicado: (2023)
DIBS: Enhancing Dense Video Captioning with Unlabeled Videos via Pseudo Boundary Enrichment and Online Refinement
por: Wu, Hao, et al.
Publicado: (2024)
por: Wu, Hao, et al.
Publicado: (2024)
STEAR: Layer-Aware Spatiotemporal Evidence Intervention for Hallucination Mitigation in Video Large Language Models
por: Fan, Linfeng, et al.
Publicado: (2026)
por: Fan, Linfeng, et al.
Publicado: (2026)
Implicit Location-Caption Alignment via Complementary Masking for Weakly-Supervised Dense Video Captioning
por: Ge, Shiping, et al.
Publicado: (2024)
por: Ge, Shiping, et al.
Publicado: (2024)
Synchronized Video Storytelling: Generating Video Narrations with Structured Storyline
por: Yang, Dingyi, et al.
Publicado: (2024)
por: Yang, Dingyi, et al.
Publicado: (2024)
FinCap: Topic-Aligned Captions for Short-Form Financial YouTube Videos
por: Sukhani, Siddhant, et al.
Publicado: (2025)
por: Sukhani, Siddhant, et al.
Publicado: (2025)
MSC: A Marine Wildlife Video Dataset with Grounded Segmentation and Clip-Level Captioning
por: Truong, Quang-Trung, et al.
Publicado: (2025)
por: Truong, Quang-Trung, et al.
Publicado: (2025)
SVD: Spatial Video Dataset
por: Izadimehr, M. H., et al.
Publicado: (2025)
por: Izadimehr, M. H., et al.
Publicado: (2025)
Music Grounding by Short Video
por: Xin, Zijie, et al.
Publicado: (2024)
por: Xin, Zijie, et al.
Publicado: (2024)
Resource-Efficient Reference-Free Evaluation of Audio Captions
por: Mahfuz, Rehana, et al.
Publicado: (2024)
por: Mahfuz, Rehana, et al.
Publicado: (2024)
Less for More: Enhanced Feedback-aligned Mixed LLMs for Molecule Caption Generation and Fine-Grained NLI Evaluation
por: Gkoumas, Dimitris, et al.
Publicado: (2024)
por: Gkoumas, Dimitris, et al.
Publicado: (2024)
diveXplore at the Video Browser Showdown 2024
por: Schoeffmann, Klaus, et al.
Publicado: (2025)
por: Schoeffmann, Klaus, et al.
Publicado: (2025)
OmniDPO: A Preference Optimization Framework to Address Omni-Modal Hallucination
por: Chen, Junzhe, et al.
Publicado: (2025)
por: Chen, Junzhe, et al.
Publicado: (2025)
CAMP-VQA: Caption-Embedded Multimodal Perception for No-Reference Quality Assessment of Compressed Video
por: Wang, Xinyi, et al.
Publicado: (2025)
por: Wang, Xinyi, et al.
Publicado: (2025)
Pre-training with Synthetic Patterns for Audio
por: Ishikawa, Yuchi, et al.
Publicado: (2024)
por: Ishikawa, Yuchi, et al.
Publicado: (2024)
Integrated Semantic and Temporal Alignment for Interactive Video Retrieval
por: Luu, Thanh-Danh, et al.
Publicado: (2025)
por: Luu, Thanh-Danh, et al.
Publicado: (2025)
Feedback-Driven Rate Control for Learned Video Compression
por: Xu, Zhiheng, et al.
Publicado: (2026)
por: Xu, Zhiheng, et al.
Publicado: (2026)
Ejemplares similares
-
DETECLAP: Enhancing Audio-Visual Representation Learning with Object Information
por: Nakada, Shota, et al.
Publicado: (2024) -
Lighthouse: A User-Friendly Library for Reproducible Video Moment Retrieval and Highlight Detection
por: Nishimura, Taichi, et al.
Publicado: (2024) -
Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos
por: Ishikawa, Yuchi, et al.
Publicado: (2025) -
On the Audio Hallucinations in Large Audio-Video Language Models
por: Nishimura, Taichi, et al.
Publicado: (2024) -
Vision-Language Models Learn Super Images for Efficient Partially Relevant Video Retrieval
por: Nishimura, Taichi, et al.
Publicado: (2023)