Beyond Accuracy: Evaluating Strategy Diversity in LLM Mathematical Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yang, Xia, Zhang, Xuanyi, Hu, Hao, Ji, Feng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Impact of LLM Self-Consistency and Reasoning Effort on Automated Scoring Accuracy and Cost
von: Frohn, Scott
Veröffentlicht: (2026)
von: Frohn, Scott
Veröffentlicht: (2026)
Exploring Communication Strategies for Collaborative LLM Agents in Mathematical Problem-Solving
von: Zhang, Liang, et al.
Veröffentlicht: (2025)
von: Zhang, Liang, et al.
Veröffentlicht: (2025)
MedThink: Enhancing Diagnostic Accuracy in Small Models via Teacher-Guided Reasoning Correction
von: Su, Xinchun, et al.
Veröffentlicht: (2026)
von: Su, Xinchun, et al.
Veröffentlicht: (2026)
Human-in-the-Loop LLM Grading for Handwritten Mathematics Assessments
von: Vanhoyweghen, Arne, et al.
Veröffentlicht: (2026)
von: Vanhoyweghen, Arne, et al.
Veröffentlicht: (2026)
SHIELD: Evaluation and Defense Strategies for Copyright Compliance in LLM Text Generation
von: Liu, Xiaoze, et al.
Veröffentlicht: (2024)
von: Liu, Xiaoze, et al.
Veröffentlicht: (2024)
Evaluating Implicit Biases in LLM Reasoning through Logic Grid Puzzles
von: Jahara, Fatima, et al.
Veröffentlicht: (2025)
von: Jahara, Fatima, et al.
Veröffentlicht: (2025)
SafeMCP: Proactive Power Regulation for LLM Agent Defense via Environment-Grounded Look-Ahead Reasoning
von: Wang, Lichao, et al.
Veröffentlicht: (2026)
von: Wang, Lichao, et al.
Veröffentlicht: (2026)
LLM Biases
von: Han, Jinhui, et al.
Veröffentlicht: (2026)
von: Han, Jinhui, et al.
Veröffentlicht: (2026)
(A)I Am Not a Lawyer, But...: Engaging Legal Experts towards Responsible LLM Policies for Legal Advice
von: Cheong, Inyoung, et al.
Veröffentlicht: (2024)
von: Cheong, Inyoung, et al.
Veröffentlicht: (2024)
A More Advanced Group Polarization Measurement Approach Based on LLM-Based Agents and Graphs
von: Liu, Zixin, et al.
Veröffentlicht: (2024)
von: Liu, Zixin, et al.
Veröffentlicht: (2024)
Mathematics Teachers Interactions with a Multi-Agent System for Personalized Problem Generation
von: Walkington, Candace, et al.
Veröffentlicht: (2026)
von: Walkington, Candace, et al.
Veröffentlicht: (2026)
XCR-Bench: A Multi-Task Benchmark for Evaluating Cultural Reasoning in LLMs
von: Kabir, Mohsinul, et al.
Veröffentlicht: (2026)
von: Kabir, Mohsinul, et al.
Veröffentlicht: (2026)
Beyond Access: Guided LLM Scaffolding for Independent Learning in Undergraduate Statistics
von: Amanlou, Mohammad, et al.
Veröffentlicht: (2026)
von: Amanlou, Mohammad, et al.
Veröffentlicht: (2026)
An Evaluation of Cultural Value Alignment in LLM
von: Sukiennik, Nicholas, et al.
Veröffentlicht: (2025)
von: Sukiennik, Nicholas, et al.
Veröffentlicht: (2025)
Sustainable LLM Inference for Edge AI: Evaluating Quantized LLMs for Energy Efficiency, Output Accuracy, and Inference Latency
von: Husom, Erik Johannes, et al.
Veröffentlicht: (2025)
von: Husom, Erik Johannes, et al.
Veröffentlicht: (2025)
LLM-Driven Personalized Answer Generation and Evaluation
von: Molavi, Mohammadreza, et al.
Veröffentlicht: (2025)
von: Molavi, Mohammadreza, et al.
Veröffentlicht: (2025)
Perturbation Effects on Accuracy and Fairness among Similar Individuals
von: Li, Xuran, et al.
Veröffentlicht: (2024)
von: Li, Xuran, et al.
Veröffentlicht: (2024)
ALSO: Adversarial Online Strategy Optimization for Social Agents
von: Li, Xiang, et al.
Veröffentlicht: (2026)
von: Li, Xiang, et al.
Veröffentlicht: (2026)
AI Detectors Fail Diverse Student Populations: A Mathematical Framing of Structural Detection Limits
von: Garland, Nathan
Veröffentlicht: (2026)
von: Garland, Nathan
Veröffentlicht: (2026)
Examining and Addressing Barriers to Diversity in LLM-Generated Ideas
von: Deng, Yuting, et al.
Veröffentlicht: (2026)
von: Deng, Yuting, et al.
Veröffentlicht: (2026)
Beyond Translation: LLM-Based Data Generation for Multilingual Fact-Checking
von: Chung, Yi-Ling, et al.
Veröffentlicht: (2025)
von: Chung, Yi-Ling, et al.
Veröffentlicht: (2025)
Truth Knows No Language: Evaluating Truthfulness Beyond English
von: Figueras, Blanca Calvo, et al.
Veröffentlicht: (2025)
von: Figueras, Blanca Calvo, et al.
Veröffentlicht: (2025)
AI-Augmented Predictions: LLM Assistants Improve Human Forecasting Accuracy
von: Schoenegger, Philipp, et al.
Veröffentlicht: (2024)
von: Schoenegger, Philipp, et al.
Veröffentlicht: (2024)
Same Verdict, Different Reasons: LLM-as-a-Judge and Clinician Disagreement on Medical Chatbot Completeness
von: DeLucia, Alexandra, et al.
Veröffentlicht: (2026)
von: DeLucia, Alexandra, et al.
Veröffentlicht: (2026)
Beyond Static Question Banks: Dynamic Knowledge Expansion via LLM-Automated Graph Construction and Adaptive Generation
von: Wang, Yingquan, et al.
Veröffentlicht: (2026)
von: Wang, Yingquan, et al.
Veröffentlicht: (2026)
MEDEQUALQA: Evaluating Biases in LLMs with Counterfactual Reasoning
von: Ghosh, Rajarshi, et al.
Veröffentlicht: (2025)
von: Ghosh, Rajarshi, et al.
Veröffentlicht: (2025)
Magic, Madness, Heaven, Sin: LLM Output Diversity is Everything, Everywhere, All at Once
von: Dhingra, Harnoor
Veröffentlicht: (2026)
von: Dhingra, Harnoor
Veröffentlicht: (2026)
In-Situ Behavioral Evaluation for LLM Fairness, Not Standardized-Test Scores
von: Tang, Zeyu, et al.
Veröffentlicht: (2026)
von: Tang, Zeyu, et al.
Veröffentlicht: (2026)
MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents
von: Zhu, Kunlun, et al.
Veröffentlicht: (2025)
von: Zhu, Kunlun, et al.
Veröffentlicht: (2025)
Beyond Accuracy: Rethinking Hallucination and Regulatory Response in Generative AI
von: Li, Zihao, et al.
Veröffentlicht: (2025)
von: Li, Zihao, et al.
Veröffentlicht: (2025)
Fair Machine Learning in Healthcare: A Review
von: Feng, Qizhang, et al.
Veröffentlicht: (2022)
von: Feng, Qizhang, et al.
Veröffentlicht: (2022)
Evaluating Prompt Engineering Techniques for Accuracy and Confidence Elicitation in Medical LLMs
von: Naderi, Nariman, et al.
Veröffentlicht: (2025)
von: Naderi, Nariman, et al.
Veröffentlicht: (2025)
A Large-Scale Real-World Evaluation of LLM-Based Virtual Teaching Assistant
von: Kweon, Sunjun, et al.
Veröffentlicht: (2025)
von: Kweon, Sunjun, et al.
Veröffentlicht: (2025)
Wisdom of the Silicon Crowd: LLM Ensemble Prediction Capabilities Rival Human Crowd Accuracy
von: Schoenegger, Philipp, et al.
Veröffentlicht: (2024)
von: Schoenegger, Philipp, et al.
Veröffentlicht: (2024)
LLM Agents Predict Social Media Reactions but Do Not Outperform Text Classifiers: Benchmarking Simulation Accuracy Using 120K+ Personas of 1511 Humans
von: Bojic, Ljubisa, et al.
Veröffentlicht: (2026)
von: Bojic, Ljubisa, et al.
Veröffentlicht: (2026)
A Study on Individual Spatiotemporal Activity Generation Method Using MCP-Enhanced Chain-of-Thought Large Language Models
von: Zhang, Yu, et al.
Veröffentlicht: (2025)
von: Zhang, Yu, et al.
Veröffentlicht: (2025)
PopResume: Causal Fairness Evaluation of LLM/VLM Resume Screeners with Population-Representative Dataset
von: Yu, Sumin, et al.
Veröffentlicht: (2026)
von: Yu, Sumin, et al.
Veröffentlicht: (2026)
Can You Share Your Story? Modeling Clients' Metacognition and Openness for LLM Therapist Evaluation
von: Kim, Minju, et al.
Veröffentlicht: (2025)
von: Kim, Minju, et al.
Veröffentlicht: (2025)
Baseline Performance of AI Tools in Classifying Cognitive Demand of Mathematical Tasks
von: Fox, Danielle S., et al.
Veröffentlicht: (2026)
von: Fox, Danielle S., et al.
Veröffentlicht: (2026)
Are LLMs Court-Ready? Evaluating Frontier Models on Indian Legal Reasoning
von: Juvekar, Kush, et al.
Veröffentlicht: (2025)
von: Juvekar, Kush, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
The Impact of LLM Self-Consistency and Reasoning Effort on Automated Scoring Accuracy and Cost
von: Frohn, Scott
Veröffentlicht: (2026) -
Exploring Communication Strategies for Collaborative LLM Agents in Mathematical Problem-Solving
von: Zhang, Liang, et al.
Veröffentlicht: (2025) -
MedThink: Enhancing Diagnostic Accuracy in Small Models via Teacher-Guided Reasoning Correction
von: Su, Xinchun, et al.
Veröffentlicht: (2026) -
Human-in-the-Loop LLM Grading for Handwritten Mathematics Assessments
von: Vanhoyweghen, Arne, et al.
Veröffentlicht: (2026) -
SHIELD: Evaluation and Defense Strategies for Copyright Compliance in LLM Text Generation
von: Liu, Xiaoze, et al.
Veröffentlicht: (2024)