Salvato in:
| Autori principali: | Ma, Yan, Chern, Steffi, Shen, Xuyang, Zhong, Yiran, Liu, Pengfei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2504.02587 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Can Large Language Models be Trusted for Evaluation? Scalable Meta-Evaluation of LLMs as Evaluators via Agent Debate
di: Chern, Steffi, et al.
Pubblicazione: (2024)
di: Chern, Steffi, et al.
Pubblicazione: (2024)
Thinking with Generated Images
di: Chern, Ethan, et al.
Pubblicazione: (2025)
di: Chern, Ethan, et al.
Pubblicazione: (2025)
One RL to See Them All: Visual Triple Unified Reinforcement Learning
di: Ma, Yan, et al.
Pubblicazione: (2025)
di: Ma, Yan, et al.
Pubblicazione: (2025)
Halu-J: Critique-Based Hallucination Judge
di: Wang, Binjie, et al.
Pubblicazione: (2024)
di: Wang, Binjie, et al.
Pubblicazione: (2024)
ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation
di: Chern, Ethan, et al.
Pubblicazione: (2024)
di: Chern, Ethan, et al.
Pubblicazione: (2024)
Scaling Laws for Linear Complexity Language Models
di: Shen, Xuyang, et al.
Pubblicazione: (2024)
di: Shen, Xuyang, et al.
Pubblicazione: (2024)
You Only Scan Once: Efficient Multi-dimension Sequential Modeling with LightNet
di: Qin, Zhen, et al.
Pubblicazione: (2024)
di: Qin, Zhen, et al.
Pubblicazione: (2024)
BeHonest: Benchmarking Honesty in Large Language Models
di: Chern, Steffi, et al.
Pubblicazione: (2024)
di: Chern, Steffi, et al.
Pubblicazione: (2024)
Evaluating Fairness in Large Vision-Language Models Across Diverse Demographic Attributes and Prompts
di: Wu, Xuyang, et al.
Pubblicazione: (2024)
di: Wu, Xuyang, et al.
Pubblicazione: (2024)
MarsRetrieval: Benchmarking Vision-Language Models for Planetary-Scale Geospatial Retrieval on Mars
di: Wang, Shuoyuan, et al.
Pubblicazione: (2026)
di: Wang, Shuoyuan, et al.
Pubblicazione: (2026)
Rethinking Visual Dependency in Long-Context Reasoning for Large Vision-Language Models
di: Zhou, Yucheng, et al.
Pubblicazione: (2024)
di: Zhou, Yucheng, et al.
Pubblicazione: (2024)
Elucidating the Design Space of Decay in Linear Attention
di: Qin, Zhen, et al.
Pubblicazione: (2025)
di: Qin, Zhen, et al.
Pubblicazione: (2025)
Rethinking Multilingual Vision-Language Translation: Dataset, Evaluation, and Adaptation
di: Wang, Xintong, et al.
Pubblicazione: (2025)
di: Wang, Xintong, et al.
Pubblicazione: (2025)
VScan: Rethinking Visual Token Reduction for Efficient Large Vision-Language Models
di: Zhang, Ce, et al.
Pubblicazione: (2025)
di: Zhang, Ce, et al.
Pubblicazione: (2025)
Enhancing Large Vision Language Models with Self-Training on Image Comprehension
di: Deng, Yihe, et al.
Pubblicazione: (2024)
di: Deng, Yihe, et al.
Pubblicazione: (2024)
Combating Adversarial Attacks with Multi-Agent Debate
di: Chern, Steffi, et al.
Pubblicazione: (2024)
di: Chern, Steffi, et al.
Pubblicazione: (2024)
Various Lengths, Constant Speed: Efficient Language Modeling with Lightning Attention
di: Qin, Zhen, et al.
Pubblicazione: (2024)
di: Qin, Zhen, et al.
Pubblicazione: (2024)
LiveTalk: Real-Time Multimodal Interactive Video Diffusion via Improved On-Policy Distillation
di: Chern, Ethan, et al.
Pubblicazione: (2025)
di: Chern, Ethan, et al.
Pubblicazione: (2025)
Judging the Judges: Can Large Vision-Language Models Fairly Evaluate Chart Comprehension and Reasoning?
di: Laskar, Md Tahmid Rahman, et al.
Pubblicazione: (2025)
di: Laskar, Md Tahmid Rahman, et al.
Pubblicazione: (2025)
EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models
di: Wang, Zekun, et al.
Pubblicazione: (2025)
di: Wang, Zekun, et al.
Pubblicazione: (2025)
Rethinking Bottlenecks in Safety Fine-Tuning of Vision Language Models
di: Ding, Yi, et al.
Pubblicazione: (2025)
di: Ding, Yi, et al.
Pubblicazione: (2025)
S-GRPO: Unified Post-Training for Large Vision-Language Models
di: Yan, Yuming, et al.
Pubblicazione: (2026)
di: Yan, Yuming, et al.
Pubblicazione: (2026)
HPE-CogVLM: Advancing Vision Language Models with a Head Pose Grounding Task
di: Tian, Yu, et al.
Pubblicazione: (2024)
di: Tian, Yu, et al.
Pubblicazione: (2024)
VEGA: Learning Interleaved Image-Text Comprehension in Vision-Language Large Models
di: Zhou, Chenyu, et al.
Pubblicazione: (2024)
di: Zhou, Chenyu, et al.
Pubblicazione: (2024)
Guiding Medical Vision-Language Models with Explicit Visual Prompts: Framework Design and Comprehensive Exploration of Prompt Variations
di: Zhu, Kangyu, et al.
Pubblicazione: (2025)
di: Zhu, Kangyu, et al.
Pubblicazione: (2025)
Evaluating Vision-Language Models as Evaluators in Path Planning
di: Aghzal, Mohamed, et al.
Pubblicazione: (2024)
di: Aghzal, Mohamed, et al.
Pubblicazione: (2024)
ScratchEval: Are GPT-4o Smarter than My Child? Evaluating Large Multimodal Models with Visual Programming Challenges
di: Fu, Rao, et al.
Pubblicazione: (2024)
di: Fu, Rao, et al.
Pubblicazione: (2024)
NPHardEval4V: Dynamic Evaluation of Large Vision-Language Models with Effects of Vision
di: Li, Xiang, et al.
Pubblicazione: (2024)
di: Li, Xiang, et al.
Pubblicazione: (2024)
Coordinated Robustness Evaluation Framework for Vision-Language Models
di: Babu, Ashwin Ramesh, et al.
Pubblicazione: (2025)
di: Babu, Ashwin Ramesh, et al.
Pubblicazione: (2025)
Lightning Attention-2: A Free Lunch for Handling Unlimited Sequence Lengths in Large Language Models
di: Qin, Zhen, et al.
Pubblicazione: (2024)
di: Qin, Zhen, et al.
Pubblicazione: (2024)
Evaluating Vision-Language Models for Emotion Recognition
di: Bhattacharyya, Sree, et al.
Pubblicazione: (2025)
di: Bhattacharyya, Sree, et al.
Pubblicazione: (2025)
Evaluation of Cultural Competence of Vision-Language Models
di: Yadav, Srishti, et al.
Pubblicazione: (2025)
di: Yadav, Srishti, et al.
Pubblicazione: (2025)
NaturalBench: Evaluating Vision-Language Models on Natural Adversarial Samples
di: Li, Baiqi, et al.
Pubblicazione: (2024)
di: Li, Baiqi, et al.
Pubblicazione: (2024)
AlignMMBench: Evaluating Chinese Multimodal Alignment in Large Vision-Language Models
di: Wu, Yuhang, et al.
Pubblicazione: (2024)
di: Wu, Yuhang, et al.
Pubblicazione: (2024)
Rethinking Misalignment in Vision-Language Model Adaptation from a Causal Perspective
di: Zhang, Yanan, et al.
Pubblicazione: (2024)
di: Zhang, Yanan, et al.
Pubblicazione: (2024)
A Comprehensive Analysis for Visual Object Hallucination in Large Vision-Language Models
di: Jing, Liqiang, et al.
Pubblicazione: (2025)
di: Jing, Liqiang, et al.
Pubblicazione: (2025)
OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles
di: Deng, Yihe, et al.
Pubblicazione: (2025)
di: Deng, Yihe, et al.
Pubblicazione: (2025)
From Local Concepts to Universals: Evaluating the Multicultural Understanding of Vision-Language Models
di: Bhatia, Mehar, et al.
Pubblicazione: (2024)
di: Bhatia, Mehar, et al.
Pubblicazione: (2024)
OpenClaw-RL: Train Any Agent Simply by Talking
di: Wang, Yinjie, et al.
Pubblicazione: (2026)
di: Wang, Yinjie, et al.
Pubblicazione: (2026)
IRR: Image Review Ranking Framework for Evaluating Vision-Language Models
di: Hayashi, Kazuki, et al.
Pubblicazione: (2024)
di: Hayashi, Kazuki, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Can Large Language Models be Trusted for Evaluation? Scalable Meta-Evaluation of LLMs as Evaluators via Agent Debate
di: Chern, Steffi, et al.
Pubblicazione: (2024) -
Thinking with Generated Images
di: Chern, Ethan, et al.
Pubblicazione: (2025) -
One RL to See Them All: Visual Triple Unified Reinforcement Learning
di: Ma, Yan, et al.
Pubblicazione: (2025) -
Halu-J: Critique-Based Hallucination Judge
di: Wang, Binjie, et al.
Pubblicazione: (2024) -
ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation
di: Chern, Ethan, et al.
Pubblicazione: (2024)