Evaluating Cross-Modal Reasoning Ability and Problem Characteristics with Multimodal Item Response Theory
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Uebayashi, Shunki, Masui, Kento, Atarashi, Kyohei, Bao, Han, Kashima, Hisashi, Inoue, Naoto, Otani, Mayu, Takeuchi, Koh |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Emulating Retrieval Augmented Generation via Prompt Engineering for Enhanced Long Context Comprehension in LLMs
von: Park, Joon, et al.
Veröffentlicht: (2025)
von: Park, Joon, et al.
Veröffentlicht: (2025)
Exploring Causes and Mitigation of Hallucinations in Large Vision Language Models
von: Sun, Yaqi, et al.
Veröffentlicht: (2025)
von: Sun, Yaqi, et al.
Veröffentlicht: (2025)
Counterfactual Evaluation for Blind Attack Detection in LLM-based Evaluation Systems
von: Liu, Lijia, et al.
Veröffentlicht: (2025)
von: Liu, Lijia, et al.
Veröffentlicht: (2025)
LayoutFlow: Flow Matching for Layout Generation
von: Guerreiro, Julian Jorge Andrade, et al.
Veröffentlicht: (2024)
von: Guerreiro, Julian Jorge Andrade, et al.
Veröffentlicht: (2024)
Robust Anomaly Detection Under Normality Distribution Shift in Dynamic Graphs
von: Xu, Xiaoyang, et al.
Veröffentlicht: (2025)
von: Xu, Xiaoyang, et al.
Veröffentlicht: (2025)
AHP-Powered LLM Reasoning for Multi-Criteria Evaluation of Open-Ended Responses
von: Lu, Xiaotian, et al.
Veröffentlicht: (2024)
von: Lu, Xiaotian, et al.
Veröffentlicht: (2024)
Unpacking the Implicit Norm Dynamics of Sharpness-Aware Minimization in Tensorized Models
von: Cao, Tianxiao, et al.
Veröffentlicht: (2025)
von: Cao, Tianxiao, et al.
Veröffentlicht: (2025)
Harnessing the Latent Diffusion Model for Training-Free Image Style Transfer
von: Masui, Kento, et al.
Veröffentlicht: (2024)
von: Masui, Kento, et al.
Veröffentlicht: (2024)
Cognitive Biases in Large Language Models: A Survey and Mitigation Experiments
von: Sumita, Yasuaki, et al.
Veröffentlicht: (2024)
von: Sumita, Yasuaki, et al.
Veröffentlicht: (2024)
Probability Bounding: Post-Hoc Calibration via Box-Constrained Softmax
von: Atarashi, Kyohei, et al.
Veröffentlicht: (2025)
von: Atarashi, Kyohei, et al.
Veröffentlicht: (2025)
LTSim: Layout Transportation-based Similarity Measure for Evaluating Layout Generation
von: Otani, Mayu, et al.
Veröffentlicht: (2024)
von: Otani, Mayu, et al.
Veröffentlicht: (2024)
OpenCOLE: Towards Reproducible Automatic Graphic Design Generation
von: Inoue, Naoto, et al.
Veröffentlicht: (2024)
von: Inoue, Naoto, et al.
Veröffentlicht: (2024)
Multimodal Markup Document Models for Graphic Design Completion
von: Kikuchi, Kotaro, et al.
Veröffentlicht: (2024)
von: Kikuchi, Kotaro, et al.
Veröffentlicht: (2024)
Neural Double Auction Mechanism
von: Suehara, Tsuyoshi, et al.
Veröffentlicht: (2024)
von: Suehara, Tsuyoshi, et al.
Veröffentlicht: (2024)
Evaluating Saliency Explanations in NLP by Crowdsourcing
von: Lu, Xiaotian, et al.
Veröffentlicht: (2024)
von: Lu, Xiaotian, et al.
Veröffentlicht: (2024)
Learning Neural Strategy-Proof Matching Mechanism from Examples
von: Maruo, Ryota, et al.
Veröffentlicht: (2024)
von: Maruo, Ryota, et al.
Veröffentlicht: (2024)
Dynamic Feature Selection from Variable Feature Sets Using Features of Features
von: Takahashi, Katsumi, et al.
Veröffentlicht: (2025)
von: Takahashi, Katsumi, et al.
Veröffentlicht: (2025)
Learning Fair and Preferable Allocations through Neural Network
von: Maruo, Ryota, et al.
Veröffentlicht: (2024)
von: Maruo, Ryota, et al.
Veröffentlicht: (2024)
Hierarchical Text Classification Using Black Box Large Language Models
von: Yoshimura, Kosuke, et al.
Veröffentlicht: (2025)
von: Yoshimura, Kosuke, et al.
Veröffentlicht: (2025)
Efficient Preference Elicitation in Iterative Combinatorial Auctions with Many Participants
von: Maruo, Ryota, et al.
Veröffentlicht: (2024)
von: Maruo, Ryota, et al.
Veröffentlicht: (2024)
Mitigating Cognitive Biases in Multi-Criteria Crowd Assessment
von: Ito, Shun, et al.
Veröffentlicht: (2024)
von: Ito, Shun, et al.
Veröffentlicht: (2024)
Estimating LLM Grading Ability and Response Difficulty in Automatic Short Answer Grading via Item Response Theory
von: Cong, Longwei, et al.
Veröffentlicht: (2026)
von: Cong, Longwei, et al.
Veröffentlicht: (2026)
Measure Twice, Cut Once: A Semantic-Oriented Approach to Video Temporal Localization with Video LLMs
von: Pang, Zongshang, et al.
Veröffentlicht: (2025)
von: Pang, Zongshang, et al.
Veröffentlicht: (2025)
RIDE: Difficulty Evolving Perturbation with Item Response Theory for Mathematical Reasoning
von: Li, Xinyuan, et al.
Veröffentlicht: (2025)
von: Li, Xinyuan, et al.
Veröffentlicht: (2025)
DiverXplorer: Stock Image Exploration via Diversity Adjustment for Graphic Design
von: Tejero-de-Pablos, Antonio, et al.
Veröffentlicht: (2026)
von: Tejero-de-Pablos, Antonio, et al.
Veröffentlicht: (2026)
The Impact of Item-Writing Flaws on Difficulty and Discrimination in Item Response Theory
von: Schmucker, Robin, et al.
Veröffentlicht: (2025)
von: Schmucker, Robin, et al.
Veröffentlicht: (2025)
JE-IRT: A Geometric Lens on LLM Abilities through Joint Embedding Item Response Theory
von: Yao, Louie Hong, et al.
Veröffentlicht: (2025)
von: Yao, Louie Hong, et al.
Veröffentlicht: (2025)
Auditing LLM Benchmarks with Item Response Theory
von: Land, Sander, et al.
Veröffentlicht: (2026)
von: Land, Sander, et al.
Veröffentlicht: (2026)
Fairness Evaluation with Item Response Theory
von: Xu, Ziqi, et al.
Veröffentlicht: (2024)
von: Xu, Ziqi, et al.
Veröffentlicht: (2024)
A User Study on the Suitability of Teleoperation Interfaces for Primitive Manipulation Tasks
von: Aoki, Jun, et al.
Veröffentlicht: (2026)
von: Aoki, Jun, et al.
Veröffentlicht: (2026)
A Generalized Model for Multidimensional Intransitivity
von: Duan, Jiuding, et al.
Veröffentlicht: (2024)
von: Duan, Jiuding, et al.
Veröffentlicht: (2024)
Lost in Benchmarks? Rethinking Large Language Model Benchmarking with Item Response Theory
von: Zhou, Hongli, et al.
Veröffentlicht: (2025)
von: Zhou, Hongli, et al.
Veröffentlicht: (2025)
ColorGPT: Leveraging Large Language Models for Multimodal Color Recommendation
von: Xia, Ding, et al.
Veröffentlicht: (2025)
von: Xia, Ding, et al.
Veröffentlicht: (2025)
MangaDiT: Reference-Guided Line Art Colorization with Hierarchical Attention in Diffusion Transformers
von: Qiu, Qianru, et al.
Veröffentlicht: (2025)
von: Qiu, Qianru, et al.
Veröffentlicht: (2025)
Actial: Activate Spatial Reasoning Ability of Multimodal Large Language Models
von: Zhan, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Zhan, Xiaoyu, et al.
Veröffentlicht: (2025)
Would Deep Generative Models Amplify Bias in Future Models?
von: Chen, Tianwei, et al.
Veröffentlicht: (2024)
von: Chen, Tianwei, et al.
Veröffentlicht: (2024)
SMART: Simulated Students Aligned with Item Response Theory for Question Difficulty Prediction
von: Scarlatos, Alexander, et al.
Veröffentlicht: (2025)
von: Scarlatos, Alexander, et al.
Veröffentlicht: (2025)
Formally Explaining Decision Tree Models with Answer Set Programming
von: Takemura, Akihiro, et al.
Veröffentlicht: (2026)
von: Takemura, Akihiro, et al.
Veröffentlicht: (2026)
Reliability-Targeted Simulation of Item Response Data: Solving the Inverse Design Problem
von: Lee, JoonHo
Veröffentlicht: (2025)
von: Lee, JoonHo
Veröffentlicht: (2025)
Adaptive Discovery of Interpretable Audio Attributes with Multimodal LLMs for Low-Resource Classification
von: Yoshimura, Kosuke, et al.
Veröffentlicht: (2026)
von: Yoshimura, Kosuke, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Emulating Retrieval Augmented Generation via Prompt Engineering for Enhanced Long Context Comprehension in LLMs
von: Park, Joon, et al.
Veröffentlicht: (2025) -
Exploring Causes and Mitigation of Hallucinations in Large Vision Language Models
von: Sun, Yaqi, et al.
Veröffentlicht: (2025) -
Counterfactual Evaluation for Blind Attack Detection in LLM-based Evaluation Systems
von: Liu, Lijia, et al.
Veröffentlicht: (2025) -
LayoutFlow: Flow Matching for Layout Generation
von: Guerreiro, Julian Jorge Andrade, et al.
Veröffentlicht: (2024) -
Robust Anomaly Detection Under Normality Distribution Shift in Dynamic Graphs
von: Xu, Xiaoyang, et al.
Veröffentlicht: (2025)