Saved in:
| Main Authors: | de Curtò, J., de Zarzà, I., García, Pablo, Cabot, Jordi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2510.26732 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Semantic Invariance in Agentic AI
by: de Zarzà, I., et al.
Published: (2026)
by: de Zarzà, I., et al.
Published: (2026)
LLM Constitutional Multi-Agent Governance
by: de Curtò, J., et al.
Published: (2026)
by: de Curtò, J., et al.
Published: (2026)
Language-Conditioned Visual Grounding with CLIP Multilingual
by: de Curtò, J., et al.
Published: (2026)
by: de Curtò, J., et al.
Published: (2026)
Evaluating Consistency and Reasoning Capabilities of Large Language Models
by: Saxena, Yash, et al.
Published: (2024)
by: Saxena, Yash, et al.
Published: (2024)
Trade-offs in Large Reasoning Models: An Empirical Analysis of Deliberative and Adaptive Reasoning over Foundational Capabilities
by: Zhao, Weixiang, et al.
Published: (2025)
by: Zhao, Weixiang, et al.
Published: (2025)
CHANCERY: Evaluating Corporate Governance Reasoning Capabilities in Language Models
by: Irwin, Lucas, et al.
Published: (2025)
by: Irwin, Lucas, et al.
Published: (2025)
Using Large Language Models to Enrich the Documentation of Datasets for Machine Learning
by: Giner-Miguelez, Joan, et al.
Published: (2024)
by: Giner-Miguelez, Joan, et al.
Published: (2024)
Mind the Language Gap: Automated and Augmented Evaluation of Bias in LLMs for High- and Low-Resource Languages
by: Buscemi, Alessio, et al.
Published: (2025)
by: Buscemi, Alessio, et al.
Published: (2025)
TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models
by: Shangguan, Ziyao, et al.
Published: (2024)
by: Shangguan, Ziyao, et al.
Published: (2024)
Assessing Logical Reasoning Capabilities of Encoder-Only Transformer Models
by: Pirozelli, Paulo, et al.
Published: (2023)
by: Pirozelli, Paulo, et al.
Published: (2023)
Oolong: Evaluating Long Context Reasoning and Aggregation Capabilities
by: Bertsch, Amanda, et al.
Published: (2025)
by: Bertsch, Amanda, et al.
Published: (2025)
Analysis of instruction-based LLMs' capabilities to score and judge text-input problems in an academic setting
by: Ramirez-Garcia, Valeria, et al.
Published: (2025)
by: Ramirez-Garcia, Valeria, et al.
Published: (2025)
Lost in the Logic: An Evaluation of Large Language Models' Reasoning Capabilities on LSAT Logic Games
by: Malik, Saumya
Published: (2024)
by: Malik, Saumya
Published: (2024)
CrossWordBench: Evaluating the Reasoning Capabilities of LLMs and LVLMs with Controllable Puzzle Generation
by: Leng, Jixuan, et al.
Published: (2025)
by: Leng, Jixuan, et al.
Published: (2025)
Evaluating Interventional Reasoning Capabilities of Large Language Models
by: Kasetty, Tejas, et al.
Published: (2024)
by: Kasetty, Tejas, et al.
Published: (2024)
SeaEval for Multilingual Foundation Models: From Cross-Lingual Alignment to Cultural Reasoning
by: Wang, Bin, et al.
Published: (2023)
by: Wang, Bin, et al.
Published: (2023)
ACE-$M^3$: Automatic Capability Evaluator for Multimodal Medical Models
by: Zhang, Xiechi, et al.
Published: (2024)
by: Zhang, Xiechi, et al.
Published: (2024)
TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models
by: Li, Ce, et al.
Published: (2025)
by: Li, Ce, et al.
Published: (2025)
Evaluation of Multilingual LLMs Personalized Text Generation Capabilities Targeting Groups and Social-Media Platforms
by: Macko, Dominik
Published: (2026)
by: Macko, Dominik
Published: (2026)
Cross-lingual Collapse: How Language-Centric Foundation Models Shape Reasoning in Large Language Models
by: Park, Cheonbok, et al.
Published: (2025)
by: Park, Cheonbok, et al.
Published: (2025)
Reasoning Capabilities of Large Language Models on Dynamic Tasks
by: Wong, Annie, et al.
Published: (2025)
by: Wong, Annie, et al.
Published: (2025)
Distilling Mathematical Reasoning Capabilities into Small Language Models
by: Zhu, Xunyu, et al.
Published: (2024)
by: Zhu, Xunyu, et al.
Published: (2024)
Adapting Large Language Models for Education: Foundational Capabilities, Potentials, and Challenges
by: Li, Qingyao, et al.
Published: (2023)
by: Li, Qingyao, et al.
Published: (2023)
Foundation Models for Geospatial Reasoning: Assessing Capabilities of Large Language Models in Understanding Geometries and Topological Spatial Relations
by: Ji, Yuhan, et al.
Published: (2025)
by: Ji, Yuhan, et al.
Published: (2025)
OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models
by: Xu, Hainiu, et al.
Published: (2024)
by: Xu, Hainiu, et al.
Published: (2024)
Steamroller Problems: An Evaluation of LLM Reasoning Capability with Automated Theorem Prover Strategies
by: McGinness, Lachlan, et al.
Published: (2024)
by: McGinness, Lachlan, et al.
Published: (2024)
Inclusion Arena: An Open Platform for Evaluating Large Foundation Models with Real-World Apps
by: Wang, Kangyu, et al.
Published: (2025)
by: Wang, Kangyu, et al.
Published: (2025)
ReasonAny: Incorporating Reasoning Capability to Any Model via Simple and Effective Model Merging
by: Yang, Junyao, et al.
Published: (2026)
by: Yang, Junyao, et al.
Published: (2026)
Exploring the System 1 Thinking Capability of Large Reasoning Models
by: Zhang, Wenyuan, et al.
Published: (2025)
by: Zhang, Wenyuan, et al.
Published: (2025)
Unlocking Reasoning Capability on Machine Translation in Large Language Models
by: Rajaee, Sara, et al.
Published: (2026)
by: Rajaee, Sara, et al.
Published: (2026)
M3Kang: Evaluating Multilingual Multimodal Mathematical Reasoning in Vision-Language Models
by: Torres-Camps, Aleix, et al.
Published: (2026)
by: Torres-Camps, Aleix, et al.
Published: (2026)
Are Your LLMs Capable of Stable Reasoning?
by: Liu, Junnan, et al.
Published: (2024)
by: Liu, Junnan, et al.
Published: (2024)
Word Sense Disambiguation in Native Spanish: A Comprehensive Lexical Evaluation Resource
by: Ortega, Pablo, et al.
Published: (2024)
by: Ortega, Pablo, et al.
Published: (2024)
The Science of Evaluating Foundation Models
by: Yuan, Jiayi, et al.
Published: (2025)
by: Yuan, Jiayi, et al.
Published: (2025)
Towards Foundation Models for Knowledge Graph Reasoning
by: Galkin, Mikhail, et al.
Published: (2023)
by: Galkin, Mikhail, et al.
Published: (2023)
Beyond Correlation: Interpretable Evaluation of Machine Translation Metrics
by: Perrella, Stefano, et al.
Published: (2024)
by: Perrella, Stefano, et al.
Published: (2024)
TS-Reasoner: Aligning Time Series Foundation Models with LLM Reasoning
by: Yu, Fangxu, et al.
Published: (2025)
by: Yu, Fangxu, et al.
Published: (2025)
Cooperative Strategic Planning Enhances Reasoning Capabilities in Large Language Models
by: Wang, Danqing, et al.
Published: (2024)
by: Wang, Danqing, et al.
Published: (2024)
ECG-Reasoning-Benchmark: A Benchmark for Evaluating Clinical Reasoning Capabilities in ECG Interpretation
by: Oh, Jungwoo, et al.
Published: (2026)
by: Oh, Jungwoo, et al.
Published: (2026)
Cross-Platform Evaluation of Large Language Model Safety in Pediatric Consultations: Evolution of Adversarial Robustness and the Scale Paradox
by: Zolfaghari, Vahideh
Published: (2025)
by: Zolfaghari, Vahideh
Published: (2025)
Similar Items
-
Semantic Invariance in Agentic AI
by: de Zarzà, I., et al.
Published: (2026) -
LLM Constitutional Multi-Agent Governance
by: de Curtò, J., et al.
Published: (2026) -
Language-Conditioned Visual Grounding with CLIP Multilingual
by: de Curtò, J., et al.
Published: (2026) -
Evaluating Consistency and Reasoning Capabilities of Large Language Models
by: Saxena, Yash, et al.
Published: (2024) -
Trade-offs in Large Reasoning Models: An Empirical Analysis of Deliberative and Adaptive Reasoning over Foundational Capabilities
by: Zhao, Weixiang, et al.
Published: (2025)