PaMi-VDPO: Mitigating Video Hallucinations by Prompt-Aware Multi-Instance Video Preference Learning
Fuente:
arXiv
Guardado en:
| Autores principales: | Ding, Xinpeng, Zhang, Kui, Han, Jianhua, Hong, Lanqing, Xu, Hang, Li, Xiaomeng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
HiLM-D: Enhancing MLLMs with Multi-Scale High-Resolution Details for Autonomous Driving
por: Ding, Xinpeng, et al.
Publicado: (2023)
por: Ding, Xinpeng, et al.
Publicado: (2023)
Hallucination Mitigation Prompts Long-term Video Understanding
por: Sun, Yiwei, et al.
Publicado: (2024)
por: Sun, Yiwei, et al.
Publicado: (2024)
Mitigating Hallucination in VideoLLMs via Temporal-Aware Activation Engineering
por: Cai, Jianfeng, et al.
Publicado: (2025)
por: Cai, Jianfeng, et al.
Publicado: (2025)
Holistic Autonomous Driving Understanding by Bird's-Eye-View Injected Multi-Modal Large Models
por: Ding, Xinpeng, et al.
Publicado: (2024)
por: Ding, Xinpeng, et al.
Publicado: (2024)
Hierarchical Visual Prompt Learning for Continual Video Instance Segmentation
por: Dong, Jiahua, et al.
Publicado: (2025)
por: Dong, Jiahua, et al.
Publicado: (2025)
CAVIS: Context-Aware Video Instance Segmentation
por: Lee, Seunghun, et al.
Publicado: (2024)
por: Lee, Seunghun, et al.
Publicado: (2024)
VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations on Synthetic Video Understanding
por: Li, Zongxia, et al.
Publicado: (2025)
por: Li, Zongxia, et al.
Publicado: (2025)
SmartSight: Mitigating Hallucination in Video-LLMs Without Compromising Video Understanding via Temporal Attention Collapse
por: Sun, Yiming, et al.
Publicado: (2025)
por: Sun, Yiming, et al.
Publicado: (2025)
SlowFocus: Enhancing Fine-grained Temporal Understanding in Video LLM
por: Nie, Ming, et al.
Publicado: (2026)
por: Nie, Ming, et al.
Publicado: (2026)
iMOVE: Instance-Motion-Aware Video Understanding
por: Li, Jiaze, et al.
Publicado: (2025)
por: Li, Jiaze, et al.
Publicado: (2025)
InstanceAnimator: Multi-Instance Sketch Video Colorization
por: Zhang, Yinhan, et al.
Publicado: (2026)
por: Zhang, Yinhan, et al.
Publicado: (2026)
Physics-Aware Video Instance Removal Benchmark
por: Li, Zirui, et al.
Publicado: (2026)
por: Li, Zirui, et al.
Publicado: (2026)
CRISP: Contrastive Residual Injection and Semantic Prompting for Continual Video Instance Segmentation
por: Liu, Baichen, et al.
Publicado: (2025)
por: Liu, Baichen, et al.
Publicado: (2025)
An Instance-Aware Prompting Framework for Training-free Camouflaged Object Segmentation
por: Yin, Chao, et al.
Publicado: (2025)
por: Yin, Chao, et al.
Publicado: (2025)
STEAR: Layer-Aware Spatiotemporal Evidence Intervention for Hallucination Mitigation in Video Large Language Models
por: Fan, Linfeng, et al.
Publicado: (2026)
por: Fan, Linfeng, et al.
Publicado: (2026)
Mitigating Hallucinations in Video Large Language Models via Spatiotemporal-Semantic Contrastive Decoding
por: Gao, Yuansheng, et al.
Publicado: (2026)
por: Gao, Yuansheng, et al.
Publicado: (2026)
FramePainter: Endowing Interactive Image Editing with Video Diffusion Priors
por: Zhang, Yabo, et al.
Publicado: (2025)
por: Zhang, Yabo, et al.
Publicado: (2025)
Task-customized Masked AutoEncoder via Mixture of Cluster-conditional Experts
por: Liu, Zhili, et al.
Publicado: (2024)
por: Liu, Zhili, et al.
Publicado: (2024)
InstaDrive: Instance-Aware Driving World Models for Realistic and Consistent Video Generation
por: Yang, Zhuoran, et al.
Publicado: (2026)
por: Yang, Zhuoran, et al.
Publicado: (2026)
A Temporal Modeling Framework for Video Pre-Training on Video Instance Segmentation
por: Zhong, Qing, et al.
Publicado: (2025)
por: Zhong, Qing, et al.
Publicado: (2025)
IAP: Improving Continual Learning of Vision-Language Models via Instance-Aware Prompting
por: Fu, Hao, et al.
Publicado: (2025)
por: Fu, Hao, et al.
Publicado: (2025)
Taming Video Diffusion Prior with Scene-Grounding Guidance for 3D Gaussian Splatting from Sparse Inputs
por: Zhong, Yingji, et al.
Publicado: (2025)
por: Zhong, Yingji, et al.
Publicado: (2025)
Make-A-Protagonist: Generic Video Editing with An Ensemble of Experts
por: Zhao, Yuyang, et al.
Publicado: (2023)
por: Zhao, Yuyang, et al.
Publicado: (2023)
InstanceV: Instance-Level Video Generation
por: Chen, Yuheng, et al.
Publicado: (2025)
por: Chen, Yuheng, et al.
Publicado: (2025)
MiCo: Multiple Instance Learning with Context-Aware Clustering for Whole Slide Image Analysis
por: Li, Junjian, et al.
Publicado: (2025)
por: Li, Junjian, et al.
Publicado: (2025)
Exploring Hallucination of Large Multimodal Models in Video Understanding: Benchmark, Analysis and Mitigation
por: Gao, Hongcheng, et al.
Publicado: (2025)
por: Gao, Hongcheng, et al.
Publicado: (2025)
Relaxing Anchor-Frame Dominance for Mitigating Hallucinations in Video Large Language Models
por: Liu, Zijian, et al.
Publicado: (2026)
por: Liu, Zijian, et al.
Publicado: (2026)
VIRES: Video Instance Repainting via Sketch and Text Guided Generation
por: Weng, Shuchen, et al.
Publicado: (2024)
por: Weng, Shuchen, et al.
Publicado: (2024)
Prompt-A-Video: Prompt Your Video Diffusion Model via Preference-Aligned LLM
por: Ji, Yatai, et al.
Publicado: (2024)
por: Ji, Yatai, et al.
Publicado: (2024)
SAPNet++: Evolving Point-Prompted Instance Segmentation with Semantic and Spatial Awareness
por: Wei, Zhaoyang, et al.
Publicado: (2026)
por: Wei, Zhaoyang, et al.
Publicado: (2026)
MagicDrive-V2: High-Resolution Long Video Generation for Autonomous Driving with Adaptive Control
por: Gao, Ruiyuan, et al.
Publicado: (2024)
por: Gao, Ruiyuan, et al.
Publicado: (2024)
UnityVideo: Unified Multi-Modal Multi-Task Learning for Enhancing World-Aware Video Generation
por: Huang, Jiehui, et al.
Publicado: (2025)
por: Huang, Jiehui, et al.
Publicado: (2025)
X-Prompt: Multi-modal Visual Prompt for Video Object Segmentation
por: Guo, Pinxue, et al.
Publicado: (2024)
por: Guo, Pinxue, et al.
Publicado: (2024)
A2VIS: Amodal-Aware Approach to Video Instance Segmentation
por: Tran, Minh, et al.
Publicado: (2024)
por: Tran, Minh, et al.
Publicado: (2024)
MedHorizon: Towards Long-context Medical Video Understanding in the Wild
por: Du, Bodong, et al.
Publicado: (2026)
por: Du, Bodong, et al.
Publicado: (2026)
Mitigating Hallucinations in Multimodal Spatial Relations through Constraint-Aware Prompting
por: Wu, Jiarui, et al.
Publicado: (2025)
por: Wu, Jiarui, et al.
Publicado: (2025)
Learning Instance-Aware Correspondences for Robust Multi-Instance Point Cloud Registration in Cluttered Scenes
por: Yu, Zhiyuan, et al.
Publicado: (2024)
por: Yu, Zhiyuan, et al.
Publicado: (2024)
Collaboratively Self-supervised Video Representation Learning for Action Recognition
por: Zhang, Jie, et al.
Publicado: (2024)
por: Zhang, Jie, et al.
Publicado: (2024)
VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation
por: Xu, Jiazheng, et al.
Publicado: (2024)
por: Xu, Jiazheng, et al.
Publicado: (2024)
MJ-VIDEO: Fine-Grained Benchmarking and Rewarding Video Preferences in Video Generation
por: Tong, Haibo, et al.
Publicado: (2025)
por: Tong, Haibo, et al.
Publicado: (2025)
Ejemplares similares
-
HiLM-D: Enhancing MLLMs with Multi-Scale High-Resolution Details for Autonomous Driving
por: Ding, Xinpeng, et al.
Publicado: (2023) -
Hallucination Mitigation Prompts Long-term Video Understanding
por: Sun, Yiwei, et al.
Publicado: (2024) -
Mitigating Hallucination in VideoLLMs via Temporal-Aware Activation Engineering
por: Cai, Jianfeng, et al.
Publicado: (2025) -
Holistic Autonomous Driving Understanding by Bird's-Eye-View Injected Multi-Modal Large Models
por: Ding, Xinpeng, et al.
Publicado: (2024) -
Hierarchical Visual Prompt Learning for Continual Video Instance Segmentation
por: Dong, Jiahua, et al.
Publicado: (2025)