Gespeichert in:
| Hauptverfasser: | Chernyshev, Konstantin, Polshkov, Vitaliy, Artemova, Ekaterina, Myasnikov, Alex, Stepanov, Vlad, Miasnikov, Alexei, Tilga, Sergei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2412.03205 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Hands-On Tutorial: Labeling with LLM and Human-in-the-Loop
von: Artemova, Ekaterina, et al.
Veröffentlicht: (2024)
von: Artemova, Ekaterina, et al.
Veröffentlicht: (2024)
Voices of Freelance Professional Writers on AI: Limitations, Expectations, and Fears
von: Ivanova, Anastasiia, et al.
Veröffentlicht: (2025)
von: Ivanova, Anastasiia, et al.
Veröffentlicht: (2025)
Beemo: Benchmark of Expert-edited Machine-generated Outputs
von: Artemova, Ekaterina, et al.
Veröffentlicht: (2024)
von: Artemova, Ekaterina, et al.
Veröffentlicht: (2024)
Tendem: A Hybrid AI+Human Platform
von: Chernyshev, Konstantin, et al.
Veröffentlicht: (2026)
von: Chernyshev, Konstantin, et al.
Veröffentlicht: (2026)
JEEM: Vision-Language Understanding in Four Arabic Dialects
von: Kadaoui, Karima, et al.
Veröffentlicht: (2025)
von: Kadaoui, Karima, et al.
Veröffentlicht: (2025)
Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models
von: Gao, Bofei, et al.
Veröffentlicht: (2024)
von: Gao, Bofei, et al.
Veröffentlicht: (2024)
KisMATH: Do LLMs Have Knowledge of Implicit Structures in Mathematical Reasoning?
von: Saha, Soumadeep, et al.
Veröffentlicht: (2025)
von: Saha, Soumadeep, et al.
Veröffentlicht: (2025)
MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations
von: Huang, Kaixuan, et al.
Veröffentlicht: (2025)
von: Huang, Kaixuan, et al.
Veröffentlicht: (2025)
Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset
von: Wang, Ke, et al.
Veröffentlicht: (2024)
von: Wang, Ke, et al.
Veröffentlicht: (2024)
A Compute-Matched Re-Evaluation of TroVE on MATH
von: Sesterhenn, Tobias, et al.
Veröffentlicht: (2025)
von: Sesterhenn, Tobias, et al.
Veröffentlicht: (2025)
GAUSS: Benchmarking Structured Mathematical Skills for Large Language Models
von: Zhang, Yue, et al.
Veröffentlicht: (2025)
von: Zhang, Yue, et al.
Veröffentlicht: (2025)
Low-Resource, High-Impact: Building Corpora for Inclusive Language Technologies
von: Artemova, Ekaterina, et al.
Veröffentlicht: (2025)
von: Artemova, Ekaterina, et al.
Veröffentlicht: (2025)
Can LLMs Reliably Simulate Real Students' Abilities in Mathematics and Reading Comprehension?
von: Srivatsa, KV Aditya, et al.
Veröffentlicht: (2025)
von: Srivatsa, KV Aditya, et al.
Veröffentlicht: (2025)
RuBLiMP: Russian Benchmark of Linguistic Minimal Pairs
von: Taktasheva, Ekaterina, et al.
Veröffentlicht: (2024)
von: Taktasheva, Ekaterina, et al.
Veröffentlicht: (2024)
LLMs Are Not Intelligent Thinkers: Introducing Mathematical Topic Tree Benchmark for Comprehensive Evaluation of LLMs
von: Davoodi, Arash Gholami, et al.
Veröffentlicht: (2024)
von: Davoodi, Arash Gholami, et al.
Veröffentlicht: (2024)
REPA: Russian Error Types Annotation for Evaluating Text Generation and Judgment Capabilities
von: Pugachev, Alexander, et al.
Veröffentlicht: (2025)
von: Pugachev, Alexander, et al.
Veröffentlicht: (2025)
Exploring the Robustness of Task-oriented Dialogue Systems for Colloquial German Varieties
von: Artemova, Ekaterina, et al.
Veröffentlicht: (2024)
von: Artemova, Ekaterina, et al.
Veröffentlicht: (2024)
Unmasking Reasoning Processes: A Process-aware Benchmark for Evaluating Structural Mathematical Reasoning in LLMs
von: Zheng, Xiang, et al.
Veröffentlicht: (2026)
von: Zheng, Xiang, et al.
Veröffentlicht: (2026)
Facilitating large language model Russian adaptation with Learned Embedding Propagation
von: Tikhomirov, Mikhail, et al.
Veröffentlicht: (2024)
von: Tikhomirov, Mikhail, et al.
Veröffentlicht: (2024)
FormalMATH: Benchmarking Formal Mathematical Reasoning of Large Language Models
von: Yu, Zhouliang, et al.
Veröffentlicht: (2025)
von: Yu, Zhouliang, et al.
Veröffentlicht: (2025)
An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
von: Hao, Yuren, et al.
Veröffentlicht: (2025)
von: Hao, Yuren, et al.
Veröffentlicht: (2025)
Benchmarking LLMs' Mathematical Reasoning with Unseen Random Variables Questions
von: Hong, Zijin, et al.
Veröffentlicht: (2025)
von: Hong, Zijin, et al.
Veröffentlicht: (2025)
TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering
von: Han, Liying, et al.
Veröffentlicht: (2026)
von: Han, Liying, et al.
Veröffentlicht: (2026)
LUNA: A Framework for Language Understanding and Naturalness Assessment
von: Saidov, Marat, et al.
Veröffentlicht: (2024)
von: Saidov, Marat, et al.
Veröffentlicht: (2024)
TeXpert: A Multi-Level Benchmark for Evaluating LaTeX Code Generation by LLMs
von: Kale, Sahil, et al.
Veröffentlicht: (2025)
von: Kale, Sahil, et al.
Veröffentlicht: (2025)
Identifying and Manipulating Personality Traits in LLMs Through Activation Engineering
von: Allbert, Rumi, et al.
Veröffentlicht: (2024)
von: Allbert, Rumi, et al.
Veröffentlicht: (2024)
MathOdyssey: Benchmarking Mathematical Problem-Solving Skills in Large Language Models Using Odyssey Math Data
von: Fang, Meng, et al.
Veröffentlicht: (2024)
von: Fang, Meng, et al.
Veröffentlicht: (2024)
RuBia: A Russian Language Bias Detection Dataset
von: Grigoreva, Veronika, et al.
Veröffentlicht: (2024)
von: Grigoreva, Veronika, et al.
Veröffentlicht: (2024)
AIpom at SemEval-2024 Task 8: Detecting AI-produced Outputs in M4
von: Shirnin, Alexander, et al.
Veröffentlicht: (2024)
von: Shirnin, Alexander, et al.
Veröffentlicht: (2024)
Papilusion at DAGPap24: Paper or Illusion? Detecting AI-generated Scientific Papers
von: Andreev, Nikita, et al.
Veröffentlicht: (2024)
von: Andreev, Nikita, et al.
Veröffentlicht: (2024)
Best-of-L: Cross-Lingual Reward Modeling for Mathematical Reasoning
von: Rajaee, Sara, et al.
Veröffentlicht: (2025)
von: Rajaee, Sara, et al.
Veröffentlicht: (2025)
Fairness Evaluation and Inference Level Mitigation in LLMs
von: Nadeem, Afrozah, et al.
Veröffentlicht: (2025)
von: Nadeem, Afrozah, et al.
Veröffentlicht: (2025)
Understanding the Role of LLMs in Multimodal Evaluation Benchmarks
von: Jiang, Botian, et al.
Veröffentlicht: (2024)
von: Jiang, Botian, et al.
Veröffentlicht: (2024)
On Robustness and Reliability of Benchmark-Based Evaluation of LLMs
von: Lunardi, Riccardo, et al.
Veröffentlicht: (2025)
von: Lunardi, Riccardo, et al.
Veröffentlicht: (2025)
DHP Benchmark: Are LLMs Good NLG Evaluators?
von: Wang, Yicheng, et al.
Veröffentlicht: (2024)
von: Wang, Yicheng, et al.
Veröffentlicht: (2024)
Mathematical Reasoning in Large Language Models: Benchmarks, Architectures, Evaluation, and Open Challenges
von: Amjad, Husnain, et al.
Veröffentlicht: (2026)
von: Amjad, Husnain, et al.
Veröffentlicht: (2026)
EduAdapt: A Question Answer Benchmark Dataset for Evaluating Grade-Level Adaptability in LLMs
von: Naeem, Numaan, et al.
Veröffentlicht: (2025)
von: Naeem, Numaan, et al.
Veröffentlicht: (2025)
UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models
von: Xu, Xin, et al.
Veröffentlicht: (2025)
von: Xu, Xin, et al.
Veröffentlicht: (2025)
VAR-MATH: Probing True Mathematical Reasoning in LLMS via Symbolic Multi-Instance Benchmarks
von: Yao, Jian, et al.
Veröffentlicht: (2025)
von: Yao, Jian, et al.
Veröffentlicht: (2025)
What Makes Cryptic Crosswords Challenging for LLMs?
von: Sadallah, Abdelrahman, et al.
Veröffentlicht: (2024)
von: Sadallah, Abdelrahman, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Hands-On Tutorial: Labeling with LLM and Human-in-the-Loop
von: Artemova, Ekaterina, et al.
Veröffentlicht: (2024) -
Voices of Freelance Professional Writers on AI: Limitations, Expectations, and Fears
von: Ivanova, Anastasiia, et al.
Veröffentlicht: (2025) -
Beemo: Benchmark of Expert-edited Machine-generated Outputs
von: Artemova, Ekaterina, et al.
Veröffentlicht: (2024) -
Tendem: A Hybrid AI+Human Platform
von: Chernyshev, Konstantin, et al.
Veröffentlicht: (2026) -
JEEM: Vision-Language Understanding in Four Arabic Dialects
von: Kadaoui, Karima, et al.
Veröffentlicht: (2025)