Gespeichert in:
| Hauptverfasser: | Liu, Ming, Zhang, Wensheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2503.05977 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Natural Reflection Backdoor Attack on Vision Language Model for Autonomous Driving
von: Liu, Ming, et al.
Veröffentlicht: (2025)
von: Liu, Ming, et al.
Veröffentlicht: (2025)
TRAVL: A Recipe for Making Video-Language Models Better Judges of Physics Implausibility
von: Motamed, Saman, et al.
Veröffentlicht: (2025)
von: Motamed, Saman, et al.
Veröffentlicht: (2025)
EvoStreaming: Your Offline Video Model Is a Natively Streaming Assistant
von: Wen, Zichen, et al.
Veröffentlicht: (2026)
von: Wen, Zichen, et al.
Veröffentlicht: (2026)
MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark
von: Chen, Dongping, et al.
Veröffentlicht: (2024)
von: Chen, Dongping, et al.
Veröffentlicht: (2024)
WorldModelBench: Judging Video Generation Models As World Models
von: Li, Dacheng, et al.
Veröffentlicht: (2025)
von: Li, Dacheng, et al.
Veröffentlicht: (2025)
PRIME: Protect Your Videos From Malicious Editing
von: Li, Guanlin, et al.
Veröffentlicht: (2024)
von: Li, Guanlin, et al.
Veröffentlicht: (2024)
Fuse Your Latents: Video Editing with Multi-source Latent Diffusion Models
von: Lu, Tianyi, et al.
Veröffentlicht: (2023)
von: Lu, Tianyi, et al.
Veröffentlicht: (2023)
StreamBridge: Turning Your Offline Video Large Language Model into a Proactive Streaming Assistant
von: Wang, Haibo, et al.
Veröffentlicht: (2025)
von: Wang, Haibo, et al.
Veröffentlicht: (2025)
Temporal Regularization Makes Your Video Generator Stronger
von: Chen, Harold Haodong, et al.
Veröffentlicht: (2025)
von: Chen, Harold Haodong, et al.
Veröffentlicht: (2025)
AdaptGCD: Multi-Expert Adapter Tuning for Generalized Category Discovery
von: Qu, Yuxun, et al.
Veröffentlicht: (2024)
von: Qu, Yuxun, et al.
Veröffentlicht: (2024)
DeVAn: Dense Video Annotation for Video-Language Models
von: Liu, Tingkai, et al.
Veröffentlicht: (2023)
von: Liu, Tingkai, et al.
Veröffentlicht: (2023)
Pack and Force Your Memory: Long-form and Consistent Video Generation
von: Wu, Xiaofei, et al.
Veröffentlicht: (2025)
von: Wu, Xiaofei, et al.
Veröffentlicht: (2025)
LLM-as-Judge Framework for Evaluating Tone-Induced Hallucination in Vision-Language Models
von: Jiang, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Jiang, Zhiyuan, et al.
Veröffentlicht: (2026)
VideoPoet: A Large Language Model for Zero-Shot Video Generation
von: Kondratyuk, Dan, et al.
Veröffentlicht: (2023)
von: Kondratyuk, Dan, et al.
Veröffentlicht: (2023)
Learning to Decode Against Compositional Hallucination in Video Multimodal Large Language Models
von: Xing, Wenbin, et al.
Veröffentlicht: (2026)
von: Xing, Wenbin, et al.
Veröffentlicht: (2026)
OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?
von: Li, Yifei, et al.
Veröffentlicht: (2025)
von: Li, Yifei, et al.
Veröffentlicht: (2025)
Your One-Stop Solution for AI-Generated Video Detection
von: Ma, Long, et al.
Veröffentlicht: (2026)
von: Ma, Long, et al.
Veröffentlicht: (2026)
Med-RewardBench: Benchmarking Reward Models and Judges for Medical Multimodal Large Language Models
von: Ding, Meidan, et al.
Veröffentlicht: (2025)
von: Ding, Meidan, et al.
Veröffentlicht: (2025)
GameVerse: Can Vision-Language Models Learn from Video-based Reflection?
von: Zhang, Kuan, et al.
Veröffentlicht: (2026)
von: Zhang, Kuan, et al.
Veröffentlicht: (2026)
To Trust Or Not To Trust Your Vision-Language Model's Prediction
von: Dong, Hao, et al.
Veröffentlicht: (2025)
von: Dong, Hao, et al.
Veröffentlicht: (2025)
Multimodal Video Emotion Recognition with Reliable Reasoning Priors
von: Wang, Zhepeng, et al.
Veröffentlicht: (2025)
von: Wang, Zhepeng, et al.
Veröffentlicht: (2025)
Can Vision Language Models Judge Action Quality? An Empirical Evaluation
von: Freitas, Miguel Monte e, et al.
Veröffentlicht: (2026)
von: Freitas, Miguel Monte e, et al.
Veröffentlicht: (2026)
Don't Judge by the Look: Towards Motion Coherent Video Representation
von: Zhang, Yitian, et al.
Veröffentlicht: (2024)
von: Zhang, Yitian, et al.
Veröffentlicht: (2024)
Your Vision-Language Model Can't Even Count to 20: Exposing the Failures of VLMs in Compositional Counting
von: Guo, Xuyang, et al.
Veröffentlicht: (2025)
von: Guo, Xuyang, et al.
Veröffentlicht: (2025)
When Models Judge Themselves: Unsupervised Self-Evolution for Multimodal Reasoning
von: Wu, Zhengxian, et al.
Veröffentlicht: (2026)
von: Wu, Zhengxian, et al.
Veröffentlicht: (2026)
You Only Judge Once: Multi-response Reward Modeling in a Single Forward Pass
von: Yang, Yinuo, et al.
Veröffentlicht: (2026)
von: Yang, Yinuo, et al.
Veröffentlicht: (2026)
SketchJudge: A Diagnostic Benchmark for Grading Hand-drawn Diagrams with Multimodal Large Language Models
von: Su, Yuhang, et al.
Veröffentlicht: (2026)
von: Su, Yuhang, et al.
Veröffentlicht: (2026)
Explicit Abstention Knobs for Predictable Reliability in Video Question Answering
von: Ortiz, Jorge
Veröffentlicht: (2025)
von: Ortiz, Jorge
Veröffentlicht: (2025)
Knowing Your Target: Target-Aware Transformer Makes Better Spatio-Temporal Video Grounding
von: Gu, Xin, et al.
Veröffentlicht: (2025)
von: Gu, Xin, et al.
Veröffentlicht: (2025)
Bootstrap Your Own Views: Masked Ego-Exo Modeling for Fine-grained View-invariant Video Representations
von: Park, Jungin, et al.
Veröffentlicht: (2025)
von: Park, Jungin, et al.
Veröffentlicht: (2025)
SurgVidLM: Towards Multi-grained Surgical Video Understanding with Large Language Model
von: Wang, Guankun, et al.
Veröffentlicht: (2025)
von: Wang, Guankun, et al.
Veröffentlicht: (2025)
VideoCogQA: A Controllable Benchmark for Evaluating Cognitive Abilities in Video-Language Models
von: Li, Chenglin, et al.
Veröffentlicht: (2024)
von: Li, Chenglin, et al.
Veröffentlicht: (2024)
CopyJudge: Automated Copyright Infringement Identification and Mitigation in Text-to-Image Diffusion Models
von: Liu, Shunchang, et al.
Veröffentlicht: (2025)
von: Liu, Shunchang, et al.
Veröffentlicht: (2025)
TEMPLE: Incentivizing Temporal Understanding of Video Large Language Models via Progressive Pre-SFT Alignment
von: Li, Shicheng, et al.
Veröffentlicht: (2025)
von: Li, Shicheng, et al.
Veröffentlicht: (2025)
MedProbCLIP: Probabilistic Adaptation of Vision-Language Foundation Model for Reliable Radiograph-Report Retrieval
von: Elallaf, Ahmad, et al.
Veröffentlicht: (2026)
von: Elallaf, Ahmad, et al.
Veröffentlicht: (2026)
TRCE: Towards Reliable Malicious Concept Erasure in Text-to-Image Diffusion Models
von: Chen, Ruidong, et al.
Veröffentlicht: (2025)
von: Chen, Ruidong, et al.
Veröffentlicht: (2025)
COLT: Enhancing Video Large Language Models with Continual Tool Usage
von: Liu, Yuyang, et al.
Veröffentlicht: (2025)
von: Liu, Yuyang, et al.
Veröffentlicht: (2025)
AD-EE: Early Exiting for Fast and Reliable Vision-Language Models in Autonomous Driving
von: Huang, Lianming, et al.
Veröffentlicht: (2025)
von: Huang, Lianming, et al.
Veröffentlicht: (2025)
From Evaluation to Defense: Advancing Safety in Video Large Language Models
von: Sun, Yiwei, et al.
Veröffentlicht: (2025)
von: Sun, Yiwei, et al.
Veröffentlicht: (2025)
ProJudge: A Multi-Modal Multi-Discipline Benchmark and Instruction-Tuning Dataset for MLLM-based Process Judges
von: Ai, Jiaxin, et al.
Veröffentlicht: (2025)
von: Ai, Jiaxin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Natural Reflection Backdoor Attack on Vision Language Model for Autonomous Driving
von: Liu, Ming, et al.
Veröffentlicht: (2025) -
TRAVL: A Recipe for Making Video-Language Models Better Judges of Physics Implausibility
von: Motamed, Saman, et al.
Veröffentlicht: (2025) -
EvoStreaming: Your Offline Video Model Is a Natively Streaming Assistant
von: Wen, Zichen, et al.
Veröffentlicht: (2026) -
MLLM-as-a-Judge: Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark
von: Chen, Dongping, et al.
Veröffentlicht: (2024) -
WorldModelBench: Judging Video Generation Models As World Models
von: Li, Dacheng, et al.
Veröffentlicht: (2025)