Failures to Surface Harmful Contents in Video Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Cao, Yuxin, Song, Wei, Wang, Derui, Xue, Jingling, Dong, Jin Song |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VideoSTF: Stress-Testing Output Repetition in Video Large Language Models
by: Cao, Yuxin, et al.
Published: (2026)
by: Cao, Yuxin, et al.
Published: (2026)
Poisoning Prompt-Guided Sampling in Video Large Language Models
by: Cao, Yuxin, et al.
Published: (2025)
by: Cao, Yuxin, et al.
Published: (2025)
CPSL: Representing Volumetric Video via Content-Promoted Scene Layers
by: Hu, Kaiyuan, et al.
Published: (2025)
by: Hu, Kaiyuan, et al.
Published: (2025)
Segmentation-Based Attention Entropy: Detecting and Mitigating Object Hallucinations in Large Vision-Language Models
by: Song, Jiale, et al.
Published: (2026)
by: Song, Jiale, et al.
Published: (2026)
Context-Enhanced Video Moment Retrieval with Large Language Models
by: Liu, Weijia, et al.
Published: (2024)
by: Liu, Weijia, et al.
Published: (2024)
When Video Coding Meets Multimodal Large Language Models: A Unified Paradigm for Video Coding
by: Zhang, Pingping, et al.
Published: (2024)
by: Zhang, Pingping, et al.
Published: (2024)
Grounded Chain-of-Thought for Multimodal Large Language Models
by: Wu, Qiong, et al.
Published: (2025)
by: Wu, Qiong, et al.
Published: (2025)
Scene Graph Generation with Role-Playing Large Language Models
by: Chen, Guikun, et al.
Published: (2024)
by: Chen, Guikun, et al.
Published: (2024)
AudioStory: Generating Long-Form Narrative Audio with Large Language Models
by: Guo, Yuxin, et al.
Published: (2025)
by: Guo, Yuxin, et al.
Published: (2025)
TextRefiner: Internal Visual Feature as Efficient Refiner for Vision-Language Models Prompt Tuning
by: Xie, Jingjing, et al.
Published: (2024)
by: Xie, Jingjing, et al.
Published: (2024)
Enhancing Partially Relevant Video Retrieval with Robust Alignment Learning
by: Zhang, Long, et al.
Published: (2025)
by: Zhang, Long, et al.
Published: (2025)
ChatVTG: Video Temporal Grounding via Chat with Video Dialogue Large Language Models
by: Qu, Mengxue, et al.
Published: (2024)
by: Qu, Mengxue, et al.
Published: (2024)
Context Guided Transformer Entropy Modeling for Video Compression
by: Tong, Junlong, et al.
Published: (2025)
by: Tong, Junlong, et al.
Published: (2025)
Break-for-Make: Modular Low-Rank Adaptations for Composable Content-Style Customization
by: Xu, Yu, et al.
Published: (2024)
by: Xu, Yu, et al.
Published: (2024)
VidCompress: Memory-Enhanced Temporal Compression for Video Understanding in Large Language Models
by: Lan, Xiaohan, et al.
Published: (2024)
by: Lan, Xiaohan, et al.
Published: (2024)
Learning Segment Similarity and Alignment in Large-Scale Content Based Video Retrieval
by: Jiang, Chen, et al.
Published: (2023)
by: Jiang, Chen, et al.
Published: (2023)
OmniScript: Towards Audio-Visual Script Generation for Long-Form Cinematic Video
by: Pu, Junfu, et al.
Published: (2026)
by: Pu, Junfu, et al.
Published: (2026)
Generalized Video Anomaly Event Detection: Systematic Taxonomy and Comparison of Deep Models
by: Liu, Yang, et al.
Published: (2023)
by: Liu, Yang, et al.
Published: (2023)
Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models
by: Xu, Yifang, et al.
Published: (2025)
by: Xu, Yifang, et al.
Published: (2025)
STEAR: Layer-Aware Spatiotemporal Evidence Intervention for Hallucination Mitigation in Video Large Language Models
by: Fan, Linfeng, et al.
Published: (2026)
by: Fan, Linfeng, et al.
Published: (2026)
AIS 2024 Challenge on Video Quality Assessment of User-Generated Content: Methods and Results
by: Conde, Marcos V., et al.
Published: (2024)
by: Conde, Marcos V., et al.
Published: (2024)
ChronusOmni: Improving Time Awareness of Omni Large Language Models
by: Chen, Yijing, et al.
Published: (2025)
by: Chen, Yijing, et al.
Published: (2025)
MotionScape: A Large-Scale Real-World Highly Dynamic UAV Video Dataset for World Models
by: Guo, Zile, et al.
Published: (2026)
by: Guo, Zile, et al.
Published: (2026)
NeR-SC: Adapting Neural Video Representation to Screen Content
by: Shi, Ruohan, et al.
Published: (2026)
by: Shi, Ruohan, et al.
Published: (2026)
D-FCGS: Feedforward Compression of Dynamic Gaussian Splatting for Free-Viewpoint Videos
by: Zhang, Wenkang, et al.
Published: (2025)
by: Zhang, Wenkang, et al.
Published: (2025)
Incorporating Visual Experts to Resolve the Information Loss in Multimodal Large Language Models
by: He, Xin, et al.
Published: (2024)
by: He, Xin, et al.
Published: (2024)
Multimodal Large Language Model is a Human-Aligned Annotator for Text-to-Image Generation
by: Wu, Xun, et al.
Published: (2024)
by: Wu, Xun, et al.
Published: (2024)
Towards Emotion Analysis in Short-form Videos: A Large-Scale Dataset and Baseline
by: Wu, Xuecheng, et al.
Published: (2023)
by: Wu, Xuecheng, et al.
Published: (2023)
Hierarchical Action Recognition: A Contrastive Video-Language Approach with Hierarchical Interactions
by: Zhang, Rui, et al.
Published: (2024)
by: Zhang, Rui, et al.
Published: (2024)
Enrich and Detect: Video Temporal Grounding with Multimodal LLMs
by: Pramanick, Shraman, et al.
Published: (2025)
by: Pramanick, Shraman, et al.
Published: (2025)
CBVS: A Large-Scale Chinese Image-Text Benchmark for Real-World Short Video Search Scenarios
by: Qiao, Xiangshuo, et al.
Published: (2024)
by: Qiao, Xiangshuo, et al.
Published: (2024)
Linguistics-Vision Monotonic Consistent Network for Sign Language Production
by: Wang, Xu, et al.
Published: (2024)
by: Wang, Xu, et al.
Published: (2024)
Unified Generative and Discriminative Training for Multi-modal Large Language Models
by: Chow, Wei, et al.
Published: (2024)
by: Chow, Wei, et al.
Published: (2024)
Hallu-PI: Evaluating Hallucination in Multi-modal Large Language Models within Perturbed Inputs
by: Ding, Peng, et al.
Published: (2024)
by: Ding, Peng, et al.
Published: (2024)
Rate-aware Compression for NeRF-based Volumetric Video
by: Zhang, Zhiyu, et al.
Published: (2024)
by: Zhang, Zhiyu, et al.
Published: (2024)
Scene-Text Grounding for Text-Based Video Question Answering
by: Zhou, Sheng, et al.
Published: (2024)
by: Zhou, Sheng, et al.
Published: (2024)
Exposure Completing for Temporally Consistent Neural High Dynamic Range Video Rendering
by: Cui, Jiahao, et al.
Published: (2024)
by: Cui, Jiahao, et al.
Published: (2024)
Adaptive Low Light Enhancement via Joint Global-Local Illumination Adjustment
by: Wang, Haodian, et al.
Published: (2025)
by: Wang, Haodian, et al.
Published: (2025)
QGFace: Quality-Guided Joint Training For Mixed-Quality Face Recognition
by: Song, Youzhe, et al.
Published: (2023)
by: Song, Youzhe, et al.
Published: (2023)
Multi-task Prompt Words Learning for Social Media Content Generation
by: Xue, Haochen, et al.
Published: (2024)
by: Xue, Haochen, et al.
Published: (2024)
Similar Items
-
VideoSTF: Stress-Testing Output Repetition in Video Large Language Models
by: Cao, Yuxin, et al.
Published: (2026) -
Poisoning Prompt-Guided Sampling in Video Large Language Models
by: Cao, Yuxin, et al.
Published: (2025) -
CPSL: Representing Volumetric Video via Content-Promoted Scene Layers
by: Hu, Kaiyuan, et al.
Published: (2025) -
Segmentation-Based Attention Entropy: Detecting and Mitigating Object Hallucinations in Large Vision-Language Models
by: Song, Jiale, et al.
Published: (2026) -
Context-Enhanced Video Moment Retrieval with Large Language Models
by: Liu, Weijia, et al.
Published: (2024)