Cross-Platform Evaluation of Reasoning Capabilities in Foundation Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | de Curtò, J., de Zarzà, I., García, Pablo, Cabot, Jordi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Semantic Invariance in Agentic AI
von: de Zarzà, I., et al.
Veröffentlicht: (2026)
von: de Zarzà, I., et al.
Veröffentlicht: (2026)
Evaluating Consistency and Reasoning Capabilities of Large Language Models
von: Saxena, Yash, et al.
Veröffentlicht: (2024)
von: Saxena, Yash, et al.
Veröffentlicht: (2024)
Trade-offs in Large Reasoning Models: An Empirical Analysis of Deliberative and Adaptive Reasoning over Foundational Capabilities
von: Zhao, Weixiang, et al.
Veröffentlicht: (2025)
von: Zhao, Weixiang, et al.
Veröffentlicht: (2025)
LLM Constitutional Multi-Agent Governance
von: de Curtò, J., et al.
Veröffentlicht: (2026)
von: de Curtò, J., et al.
Veröffentlicht: (2026)
CHANCERY: Evaluating Corporate Governance Reasoning Capabilities in Language Models
von: Irwin, Lucas, et al.
Veröffentlicht: (2025)
von: Irwin, Lucas, et al.
Veröffentlicht: (2025)
Oolong: Evaluating Long Context Reasoning and Aggregation Capabilities
von: Bertsch, Amanda, et al.
Veröffentlicht: (2025)
von: Bertsch, Amanda, et al.
Veröffentlicht: (2025)
Assessing Logical Reasoning Capabilities of Encoder-Only Transformer Models
von: Pirozelli, Paulo, et al.
Veröffentlicht: (2023)
von: Pirozelli, Paulo, et al.
Veröffentlicht: (2023)
TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation Models
von: Shangguan, Ziyao, et al.
Veröffentlicht: (2024)
von: Shangguan, Ziyao, et al.
Veröffentlicht: (2024)
Lost in the Logic: An Evaluation of Large Language Models' Reasoning Capabilities on LSAT Logic Games
von: Malik, Saumya
Veröffentlicht: (2024)
von: Malik, Saumya
Veröffentlicht: (2024)
Mind the Language Gap: Automated and Augmented Evaluation of Bias in LLMs for High- and Low-Resource Languages
von: Buscemi, Alessio, et al.
Veröffentlicht: (2025)
von: Buscemi, Alessio, et al.
Veröffentlicht: (2025)
SeaEval for Multilingual Foundation Models: From Cross-Lingual Alignment to Cultural Reasoning
von: Wang, Bin, et al.
Veröffentlicht: (2023)
von: Wang, Bin, et al.
Veröffentlicht: (2023)
TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models
von: Li, Ce, et al.
Veröffentlicht: (2025)
von: Li, Ce, et al.
Veröffentlicht: (2025)
Cross-lingual Collapse: How Language-Centric Foundation Models Shape Reasoning in Large Language Models
von: Park, Cheonbok, et al.
Veröffentlicht: (2025)
von: Park, Cheonbok, et al.
Veröffentlicht: (2025)
Evaluation of Multilingual LLMs Personalized Text Generation Capabilities Targeting Groups and Social-Media Platforms
von: Macko, Dominik
Veröffentlicht: (2026)
von: Macko, Dominik
Veröffentlicht: (2026)
Language-Conditioned Visual Grounding with CLIP Multilingual
von: de Curtò, J., et al.
Veröffentlicht: (2026)
von: de Curtò, J., et al.
Veröffentlicht: (2026)
CrossWordBench: Evaluating the Reasoning Capabilities of LLMs and LVLMs with Controllable Puzzle Generation
von: Leng, Jixuan, et al.
Veröffentlicht: (2025)
von: Leng, Jixuan, et al.
Veröffentlicht: (2025)
Evaluating Interventional Reasoning Capabilities of Large Language Models
von: Kasetty, Tejas, et al.
Veröffentlicht: (2024)
von: Kasetty, Tejas, et al.
Veröffentlicht: (2024)
Reasoning Capabilities of Large Language Models on Dynamic Tasks
von: Wong, Annie, et al.
Veröffentlicht: (2025)
von: Wong, Annie, et al.
Veröffentlicht: (2025)
Distilling Mathematical Reasoning Capabilities into Small Language Models
von: Zhu, Xunyu, et al.
Veröffentlicht: (2024)
von: Zhu, Xunyu, et al.
Veröffentlicht: (2024)
OpenToM: A Comprehensive Benchmark for Evaluating Theory-of-Mind Reasoning Capabilities of Large Language Models
von: Xu, Hainiu, et al.
Veröffentlicht: (2024)
von: Xu, Hainiu, et al.
Veröffentlicht: (2024)
Adapting Large Language Models for Education: Foundational Capabilities, Potentials, and Challenges
von: Li, Qingyao, et al.
Veröffentlicht: (2023)
von: Li, Qingyao, et al.
Veröffentlicht: (2023)
ACE-$M^3$: Automatic Capability Evaluator for Multimodal Medical Models
von: Zhang, Xiechi, et al.
Veröffentlicht: (2024)
von: Zhang, Xiechi, et al.
Veröffentlicht: (2024)
Steamroller Problems: An Evaluation of LLM Reasoning Capability with Automated Theorem Prover Strategies
von: McGinness, Lachlan, et al.
Veröffentlicht: (2024)
von: McGinness, Lachlan, et al.
Veröffentlicht: (2024)
Using Large Language Models to Enrich the Documentation of Datasets for Machine Learning
von: Giner-Miguelez, Joan, et al.
Veröffentlicht: (2024)
von: Giner-Miguelez, Joan, et al.
Veröffentlicht: (2024)
ReasonAny: Incorporating Reasoning Capability to Any Model via Simple and Effective Model Merging
von: Yang, Junyao, et al.
Veröffentlicht: (2026)
von: Yang, Junyao, et al.
Veröffentlicht: (2026)
Foundation Models for Geospatial Reasoning: Assessing Capabilities of Large Language Models in Understanding Geometries and Topological Spatial Relations
von: Ji, Yuhan, et al.
Veröffentlicht: (2025)
von: Ji, Yuhan, et al.
Veröffentlicht: (2025)
Exploring the System 1 Thinking Capability of Large Reasoning Models
von: Zhang, Wenyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Wenyuan, et al.
Veröffentlicht: (2025)
Unlocking Reasoning Capability on Machine Translation in Large Language Models
von: Rajaee, Sara, et al.
Veröffentlicht: (2026)
von: Rajaee, Sara, et al.
Veröffentlicht: (2026)
Analysis of instruction-based LLMs' capabilities to score and judge text-input problems in an academic setting
von: Ramirez-Garcia, Valeria, et al.
Veröffentlicht: (2025)
von: Ramirez-Garcia, Valeria, et al.
Veröffentlicht: (2025)
Inclusion Arena: An Open Platform for Evaluating Large Foundation Models with Real-World Apps
von: Wang, Kangyu, et al.
Veröffentlicht: (2025)
von: Wang, Kangyu, et al.
Veröffentlicht: (2025)
Towards Foundation Models for Knowledge Graph Reasoning
von: Galkin, Mikhail, et al.
Veröffentlicht: (2023)
von: Galkin, Mikhail, et al.
Veröffentlicht: (2023)
Are Your LLMs Capable of Stable Reasoning?
von: Liu, Junnan, et al.
Veröffentlicht: (2024)
von: Liu, Junnan, et al.
Veröffentlicht: (2024)
The Science of Evaluating Foundation Models
von: Yuan, Jiayi, et al.
Veröffentlicht: (2025)
von: Yuan, Jiayi, et al.
Veröffentlicht: (2025)
TS-Reasoner: Aligning Time Series Foundation Models with LLM Reasoning
von: Yu, Fangxu, et al.
Veröffentlicht: (2025)
von: Yu, Fangxu, et al.
Veröffentlicht: (2025)
Cooperative Strategic Planning Enhances Reasoning Capabilities in Large Language Models
von: Wang, Danqing, et al.
Veröffentlicht: (2024)
von: Wang, Danqing, et al.
Veröffentlicht: (2024)
Cross-Platform Evaluation of Large Language Model Safety in Pediatric Consultations: Evolution of Adversarial Robustness and the Scale Paradox
von: Zolfaghari, Vahideh
Veröffentlicht: (2025)
von: Zolfaghari, Vahideh
Veröffentlicht: (2025)
M3Kang: Evaluating Multilingual Multimodal Mathematical Reasoning in Vision-Language Models
von: Torres-Camps, Aleix, et al.
Veröffentlicht: (2026)
von: Torres-Camps, Aleix, et al.
Veröffentlicht: (2026)
Stop Before You Fail: Operational Capability Boundaries for Mitigating Unproductive Reasoning in Large Reasoning Models
von: Zhang, Qingjie, et al.
Veröffentlicht: (2025)
von: Zhang, Qingjie, et al.
Veröffentlicht: (2025)
Capabilities of GPT-5 on Multimodal Medical Reasoning
von: Wang, Shansong, et al.
Veröffentlicht: (2025)
von: Wang, Shansong, et al.
Veröffentlicht: (2025)
Explore the Reasoning Capability of LLMs in the Chess Testbed
von: Wang, Shu, et al.
Veröffentlicht: (2024)
von: Wang, Shu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Semantic Invariance in Agentic AI
von: de Zarzà, I., et al.
Veröffentlicht: (2026) -
Evaluating Consistency and Reasoning Capabilities of Large Language Models
von: Saxena, Yash, et al.
Veröffentlicht: (2024) -
Trade-offs in Large Reasoning Models: An Empirical Analysis of Deliberative and Adaptive Reasoning over Foundational Capabilities
von: Zhao, Weixiang, et al.
Veröffentlicht: (2025) -
LLM Constitutional Multi-Agent Governance
von: de Curtò, J., et al.
Veröffentlicht: (2026) -
CHANCERY: Evaluating Corporate Governance Reasoning Capabilities in Language Models
von: Irwin, Lucas, et al.
Veröffentlicht: (2025)