Video-LevelGauge: Investigating Contextual Positional Bias in Large Video Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xia, Hou, Fu, Zheren, Ling, Fangcan, Li, Jiajun, Tu, Yi, Mao, Zhendong, Zhang, Yongdong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Uncertainty-Aware Exploratory Direct Preference Optimization for Multimodal Large Language Models
von: Zhang, Huatian, et al.
Veröffentlicht: (2026)
von: Zhang, Huatian, et al.
Veröffentlicht: (2026)
Stream-R1: Reliability-Perplexity Aware Reward Distillation for Streaming Video Generation
von: Wu, Bin, et al.
Veröffentlicht: (2026)
von: Wu, Bin, et al.
Veröffentlicht: (2026)
SparseRM: A Lightweight Preference Modeling with Sparse Autoencoder
von: Liu, Dengcan, et al.
Veröffentlicht: (2025)
von: Liu, Dengcan, et al.
Veröffentlicht: (2025)
DualReal: Adaptive Joint Training for Lossless Identity-Motion Fusion in Video Customization
von: Wang, Wenchuan, et al.
Veröffentlicht: (2025)
von: Wang, Wenchuan, et al.
Veröffentlicht: (2025)
FACE-net: Factual Calibration and Emotion Augmentation for Retrieval-enhanced Emotional Video Captioning
von: Chen, Weidong, et al.
Veröffentlicht: (2026)
von: Chen, Weidong, et al.
Veröffentlicht: (2026)
SVLTA: Benchmarking Vision-Language Temporal Alignment via Synthetic Video Situation
von: Du, Hao, et al.
Veröffentlicht: (2025)
von: Du, Hao, et al.
Veröffentlicht: (2025)
Exploring Enhanced Contextual Information for Video-Level Object Tracking
von: Kang, Ben, et al.
Veröffentlicht: (2024)
von: Kang, Ben, et al.
Veröffentlicht: (2024)
Lance: Unified Multimodal Modeling by Multi-Task Synergy
von: Fu, Fengyi, et al.
Veröffentlicht: (2026)
von: Fu, Fengyi, et al.
Veröffentlicht: (2026)
Sentiment-oriented Transformer-based Variational Autoencoder Network for Live Video Commenting
von: Fu, Fengyi, et al.
Veröffentlicht: (2024)
von: Fu, Fengyi, et al.
Veröffentlicht: (2024)
SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding Capability
von: Wang, Jiankang, et al.
Veröffentlicht: (2025)
von: Wang, Jiankang, et al.
Veröffentlicht: (2025)
Stream-T1: Test-Time Scaling for Streaming Video Generation
von: Tu, Yijing, et al.
Veröffentlicht: (2026)
von: Tu, Yijing, et al.
Veröffentlicht: (2026)
VideoLLM-online: Online Video Large Language Model for Streaming Video
von: Chen, Joya, et al.
Veröffentlicht: (2024)
von: Chen, Joya, et al.
Veröffentlicht: (2024)
RealCustom: Narrowing Real Text Word for Real-Time Open-Domain Text-to-Image Customization
von: Huang, Mengqi, et al.
Veröffentlicht: (2024)
von: Huang, Mengqi, et al.
Veröffentlicht: (2024)
DSE-GAN: Dynamic Semantic Evolution Generative Adversarial Network for Text-to-Image Generation
von: Huang, Mengqi, et al.
Veröffentlicht: (2022)
von: Huang, Mengqi, et al.
Veröffentlicht: (2022)
RealGeneral: Unifying Visual Generation via Temporal In-Context Learning with Video Models
von: Lin, Yijing, et al.
Veröffentlicht: (2025)
von: Lin, Yijing, et al.
Veröffentlicht: (2025)
Investigating Video Reasoning Capability of Large Language Models with Tropes in Movies
von: Su, Hung-Ting, et al.
Veröffentlicht: (2024)
von: Su, Hung-Ting, et al.
Veröffentlicht: (2024)
Dual-path Collaborative Generation Network for Emotional Video Captioning
von: Ye, Cheng, et al.
Veröffentlicht: (2024)
von: Ye, Cheng, et al.
Veröffentlicht: (2024)
Leveraging Robust Optimization for LLM Alignment under Distribution Shifts
von: Zhu, Mingye, et al.
Veröffentlicht: (2025)
von: Zhu, Mingye, et al.
Veröffentlicht: (2025)
Geometry-Aware Rotary Position Embedding for Consistent Video World Model
von: Xiang, Chendong, et al.
Veröffentlicht: (2026)
von: Xiang, Chendong, et al.
Veröffentlicht: (2026)
RealCustom++: Representing Images as Real Textual Word for Real-Time Customization
von: Mao, Zhendong, et al.
Veröffentlicht: (2024)
von: Mao, Zhendong, et al.
Veröffentlicht: (2024)
T-SVG: Text-Driven Stereoscopic Video Generation
von: Jin, Qiao, et al.
Veröffentlicht: (2024)
von: Jin, Qiao, et al.
Veröffentlicht: (2024)
Video Summarization with Large Language Models
von: Lee, Min Jung, et al.
Veröffentlicht: (2025)
von: Lee, Min Jung, et al.
Veröffentlicht: (2025)
ManiVideo: Generating Hand-Object Manipulation Video with Dexterous and Generalizable Grasping
von: Pang, Youxin, et al.
Veröffentlicht: (2024)
von: Pang, Youxin, et al.
Veröffentlicht: (2024)
Vision-Language Models Assisted Unsupervised Video Anomaly Detection
von: Jiang, Yalong, et al.
Veröffentlicht: (2024)
von: Jiang, Yalong, et al.
Veröffentlicht: (2024)
VideoGLaMM: A Large Multimodal Model for Pixel-Level Visual Grounding in Videos
von: Munasinghe, Shehan, et al.
Veröffentlicht: (2024)
von: Munasinghe, Shehan, et al.
Veröffentlicht: (2024)
ViLLa: Video Reasoning Segmentation with Large Language Model
von: Zheng, Rongkun, et al.
Veröffentlicht: (2024)
von: Zheng, Rongkun, et al.
Veröffentlicht: (2024)
Tango: Taming Visual Signals for Efficient Video Large Language Models
von: Yin, Shukang, et al.
Veröffentlicht: (2026)
von: Yin, Shukang, et al.
Veröffentlicht: (2026)
Contextualized Diffusion Models for Text-Guided Image and Video Generation
von: Yang, Ling, et al.
Veröffentlicht: (2024)
von: Yang, Ling, et al.
Veröffentlicht: (2024)
Linear Scaling Video VLMs for Long Video Understanding
von: Eyzaguirre, Cristobal, et al.
Veröffentlicht: (2026)
von: Eyzaguirre, Cristobal, et al.
Veröffentlicht: (2026)
Video SimpleQA: Towards Factuality Evaluation in Large Video Language Models
von: Cao, Meng, et al.
Veröffentlicht: (2025)
von: Cao, Meng, et al.
Veröffentlicht: (2025)
VideoHallucer: Evaluating Intrinsic and Extrinsic Hallucinations in Large Video-Language Models
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
von: Wang, Yuxuan, et al.
Veröffentlicht: (2024)
Enhancing Video Large Language Models with Structured Multi-Video Collaborative Reasoning
von: He, Zhihao, et al.
Veröffentlicht: (2025)
von: He, Zhihao, et al.
Veröffentlicht: (2025)
DynaTok: Temporally Adaptive and Positional Bias-Aware Token Compression for Video-LLMs
von: Park, Minyoung, et al.
Veröffentlicht: (2026)
von: Park, Minyoung, et al.
Veröffentlicht: (2026)
Large Language Models for Video Surveillance Applications
von: De Silva, Ulindu, et al.
Veröffentlicht: (2025)
von: De Silva, Ulindu, et al.
Veröffentlicht: (2025)
Prompts to Summaries: Zero-Shot Language-Guided Video Summarization with Large Language and Video Models
von: Barbara, Mario, et al.
Veröffentlicht: (2025)
von: Barbara, Mario, et al.
Veröffentlicht: (2025)
Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos
von: Fei, Jiajun, et al.
Veröffentlicht: (2024)
von: Fei, Jiajun, et al.
Veröffentlicht: (2024)
Video-RAG: Visually-aligned Retrieval-Augmented Long Video Comprehension
von: Luo, Yongdong, et al.
Veröffentlicht: (2024)
von: Luo, Yongdong, et al.
Veröffentlicht: (2024)
Generative Ghost: Investigating Ranking Bias Hidden in AI-Generated Videos
von: Gao, Haowen, et al.
Veröffentlicht: (2025)
von: Gao, Haowen, et al.
Veröffentlicht: (2025)
A Large-Scale Analysis on Contextual Self-Supervised Video Representation Learning
von: Kumar, Akash, et al.
Veröffentlicht: (2025)
von: Kumar, Akash, et al.
Veröffentlicht: (2025)
Video Compression Commander: Plug-and-Play Inference Acceleration for Video Large Language Models
von: Liu, Xuyang, et al.
Veröffentlicht: (2025)
von: Liu, Xuyang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Uncertainty-Aware Exploratory Direct Preference Optimization for Multimodal Large Language Models
von: Zhang, Huatian, et al.
Veröffentlicht: (2026) -
Stream-R1: Reliability-Perplexity Aware Reward Distillation for Streaming Video Generation
von: Wu, Bin, et al.
Veröffentlicht: (2026) -
SparseRM: A Lightweight Preference Modeling with Sparse Autoencoder
von: Liu, Dengcan, et al.
Veröffentlicht: (2025) -
DualReal: Adaptive Joint Training for Lossless Identity-Motion Fusion in Video Customization
von: Wang, Wenchuan, et al.
Veröffentlicht: (2025) -
FACE-net: Factual Calibration and Emotion Augmentation for Retrieval-enhanced Emotional Video Captioning
von: Chen, Weidong, et al.
Veröffentlicht: (2026)