HRScene: How Far Are VLMs from Effective High-Resolution Image Understanding?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Yusen, Zheng, Wenliang, Madasu, Aashrith, Shi, Peng, Kamoi, Ryo, Zhou, Hao, Zou, Zhuoyang, Zhao, Shu, Das, Sarkar Snigdha Sarathi, Gupta, Vipul, Lu, Xiaoxin, Zhang, Nan, Zhang, Ranran Haoran, Iyer, Avitej, Lou, Renze, Yin, Wenpeng, Zhang, Rui |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Efficient PRM Training Data Synthesis via Formal Verification
von: Kamoi, Ryo, et al.
Veröffentlicht: (2025)
von: Kamoi, Ryo, et al.
Veröffentlicht: (2025)
VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information
von: Kamoi, Ryo, et al.
Veröffentlicht: (2024)
von: Kamoi, Ryo, et al.
Veröffentlicht: (2024)
Evaluating LLMs at Detecting Errors in LLM Responses
von: Kamoi, Ryo, et al.
Veröffentlicht: (2024)
von: Kamoi, Ryo, et al.
Veröffentlicht: (2024)
GReaTer: Gradients over Reasoning Makes Smaller Language Models Strong Prompt Optimizers
von: Das, Sarkar Snigdha Sarathi, et al.
Veröffentlicht: (2024)
von: Das, Sarkar Snigdha Sarathi, et al.
Veröffentlicht: (2024)
GREATERPROMPT: A Unified, Customizable, and High-Performing Open-Source Toolkit for Prompt Optimization
von: Zheng, Wenliang, et al.
Veröffentlicht: (2025)
von: Zheng, Wenliang, et al.
Veröffentlicht: (2025)
Verbosity $\neq$ Veracity: Demystify Verbosity Compensation Behavior of Large Language Models
von: Zhang, Yusen, et al.
Veröffentlicht: (2024)
von: Zhang, Yusen, et al.
Veröffentlicht: (2024)
Enhance Multimodal Consistency and Coherence for Text-Image Plan Generation
von: Lu, Xiaoxin, et al.
Veröffentlicht: (2025)
von: Lu, Xiaoxin, et al.
Veröffentlicht: (2025)
When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs
von: Kamoi, Ryo, et al.
Veröffentlicht: (2024)
von: Kamoi, Ryo, et al.
Veröffentlicht: (2024)
Large Language Model Instruction Following: A Survey of Progresses and Challenges
von: Lou, Renze, et al.
Veröffentlicht: (2023)
von: Lou, Renze, et al.
Veröffentlicht: (2023)
Toward Zero-Shot Instruction Following
von: Lou, Renze, et al.
Veröffentlicht: (2023)
von: Lou, Renze, et al.
Veröffentlicht: (2023)
AAAR-1.0: Assessing AI's Potential to Assist Research
von: Lou, Renze, et al.
Veröffentlicht: (2024)
von: Lou, Renze, et al.
Veröffentlicht: (2024)
Can LLMs Rank the Harmfulness of Smaller LLMs? We are Not There Yet
von: Atil, Berk, et al.
Veröffentlicht: (2025)
von: Atil, Berk, et al.
Veröffentlicht: (2025)
Direct-Inverse Prompting: Analyzing LLMs' Discriminative Capacity in Self-Improving Generation
von: Ahn, Jihyun Janice, et al.
Veröffentlicht: (2024)
von: Ahn, Jihyun Janice, et al.
Veröffentlicht: (2024)
Fair Abstractive Summarization of Diverse Perspectives
von: Zhang, Yusen, et al.
Veröffentlicht: (2023)
von: Zhang, Yusen, et al.
Veröffentlicht: (2023)
Large Language Models for Mathematical Reasoning: Progresses and Challenges
von: Ahn, Janice, et al.
Veröffentlicht: (2024)
von: Ahn, Janice, et al.
Veröffentlicht: (2024)
Chain-of-Scrutiny: Detecting Backdoor Attacks for Large Language Models
von: Li, Xi, et al.
Veröffentlicht: (2024)
von: Li, Xi, et al.
Veröffentlicht: (2024)
Coverage-based Fairness in Multi-document Summarization
von: Li, Haoyuan, et al.
Veröffentlicht: (2024)
von: Li, Haoyuan, et al.
Veröffentlicht: (2024)
UMIE: Unified Multimodal Information Extraction with Instruction Tuning
von: Sun, Lin, et al.
Veröffentlicht: (2024)
von: Sun, Lin, et al.
Veröffentlicht: (2024)
MUFFIN: Curating Multi-Faceted Instructions for Improving Instruction-Following
von: Lou, Renze, et al.
Veröffentlicht: (2023)
von: Lou, Renze, et al.
Veröffentlicht: (2023)
M2-Verify: A Large-Scale Multidomain Benchmark for Checking Multimodal Claim Consistency
von: Ansari, Abolfazl, et al.
Veröffentlicht: (2026)
von: Ansari, Abolfazl, et al.
Veröffentlicht: (2026)
DIAGPaper: Diagnosing Valid and Specific Weaknesses in Scientific Papers via Multi-Agent Reasoning
von: Zou, Zhuoyang, et al.
Veröffentlicht: (2026)
von: Zou, Zhuoyang, et al.
Veröffentlicht: (2026)
Some interesting number theory problems
von: Zhang, Wenpeng
Veröffentlicht: (2025)
von: Zhang, Wenpeng
Veröffentlicht: (2025)
A century problem related to the Legendre symbol modulo p
von: Zhang, Wenpeng
Veröffentlicht: (2025)
von: Zhang, Wenpeng
Veröffentlicht: (2025)
Bridging the Know-Act Gap via Task-Level Autoregressive Reasoning
von: Ahn, Jihyun Janice, et al.
Veröffentlicht: (2026)
von: Ahn, Jihyun Janice, et al.
Veröffentlicht: (2026)
Adaptive Chameleon or Stubborn Sloth: Revealing the Behavior of Large Language Models in Knowledge Conflicts
von: Xie, Jian, et al.
Veröffentlicht: (2023)
von: Xie, Jian, et al.
Veröffentlicht: (2023)
Evaluating LLM-Simulated Conversations in Modeling Inconsistent and Uncollaborative Behaviors in Human Social Interaction
von: Kamoi, Ryo, et al.
Veröffentlicht: (2026)
von: Kamoi, Ryo, et al.
Veröffentlicht: (2026)
A General Benchmark Framework is Dynamic Graph Neural Network Need
von: Zhang, Yusen
Veröffentlicht: (2024)
von: Zhang, Yusen
Veröffentlicht: (2024)
How Do LLMs and VLMs Understand Viewpoint Rotation Without Vision? An Interpretability Study
von: Yang, Zhen, et al.
Veröffentlicht: (2026)
von: Yang, Zhen, et al.
Veröffentlicht: (2026)
How Far are VLMs from Visual Spatial Intelligence? A Benchmark-Driven Perspective
von: Yu, Songsong, et al.
Veröffentlicht: (2025)
von: Yu, Songsong, et al.
Veröffentlicht: (2025)
The Second Vanishing Theorem for Local Cohomology Modules
von: Zhang, Wenliang
Veröffentlicht: (2021)
von: Zhang, Wenliang
Veröffentlicht: (2021)
UHR-Micro: Diagnosing and Mitigating the Resolution Illusion in Earth Observation VLMs
von: Ni, Shuo, et al.
Veröffentlicht: (2026)
von: Ni, Shuo, et al.
Veröffentlicht: (2026)
Space Syntax-guided Post-training for Residential Floor Plan Generation
von: Jiang, Zhuoyang, et al.
Veröffentlicht: (2026)
von: Jiang, Zhuoyang, et al.
Veröffentlicht: (2026)
The Tool Illusion: Rethinking Tool Use in Web Agents
von: Lou, Renze, et al.
Veröffentlicht: (2026)
von: Lou, Renze, et al.
Veröffentlicht: (2026)
Wall effects of an eccentric fluid sphere: Happel's and Kuwabara's models
von: Krishna Prasad Madasu
Veröffentlicht: (2024)
von: Krishna Prasad Madasu
Veröffentlicht: (2024)
Precision and Variability: Exploring Osteotomy Cuts in Sagittal Ramus Osteotomy—A Review
von: Prajesh Dubey, et al.
Veröffentlicht: (2025)
von: Prajesh Dubey, et al.
Veröffentlicht: (2025)
The Model Agreed, But Didn't Learn: Diagnosing Surface Compliance in Large Language Models
von: Gu, Xiaojie, et al.
Veröffentlicht: (2026)
von: Gu, Xiaojie, et al.
Veröffentlicht: (2026)
Nexus : An Agentic Framework for Time Series Forecasting
von: Das, Sarkar Snigdha Sarathi, et al.
Veröffentlicht: (2026)
von: Das, Sarkar Snigdha Sarathi, et al.
Veröffentlicht: (2026)
DocMath-Eval: Evaluating Math Reasoning Capabilities of LLMs in Understanding Long and Specialized Documents
von: Zhao, Yilun, et al.
Veröffentlicht: (2023)
von: Zhao, Yilun, et al.
Veröffentlicht: (2023)
SPACENUM: Revisiting Spatial Numerical Understanding in VLMs
von: Zhang, Jianshu, et al.
Veröffentlicht: (2026)
von: Zhang, Jianshu, et al.
Veröffentlicht: (2026)
LLMs' Classification Performance is Overclaimed
von: Xu, Hanzi, et al.
Veröffentlicht: (2024)
von: Xu, Hanzi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Efficient PRM Training Data Synthesis via Formal Verification
von: Kamoi, Ryo, et al.
Veröffentlicht: (2025) -
VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information
von: Kamoi, Ryo, et al.
Veröffentlicht: (2024) -
Evaluating LLMs at Detecting Errors in LLM Responses
von: Kamoi, Ryo, et al.
Veröffentlicht: (2024) -
GReaTer: Gradients over Reasoning Makes Smaller Language Models Strong Prompt Optimizers
von: Das, Sarkar Snigdha Sarathi, et al.
Veröffentlicht: (2024) -
GREATERPROMPT: A Unified, Customizable, and High-Performing Open-Source Toolkit for Prompt Optimization
von: Zheng, Wenliang, et al.
Veröffentlicht: (2025)