Judging What We Cannot Solve: A Consequence-Based Approach for Oracle-Free Evaluation of Research-Level Math
Fuente:
arXiv
Saved in:
| Main Authors: | Son, Guijin, Yang, Donghun, Patel, Hitesh Laxmichand, Ko, Hyunwoo, Agarwal, Amit, Ahn, Sunghee, Lee, Kyong-Ha, Yu, Youngjae |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Pushing on Multilingual Reasoning Models with Language-Mixed Chain-of-Thought
by: Son, Guijin, et al.
Published: (2025)
by: Son, Guijin, et al.
Published: (2025)
LLM-as-a-Judge & Reward Model: What They Can and Cannot Do
by: Son, Guijin, et al.
Published: (2024)
by: Son, Guijin, et al.
Published: (2024)
ResearchMath-14K: Scaling Research-Level Mathematics via Agents
by: Son, Guijin, et al.
Published: (2026)
by: Son, Guijin, et al.
Published: (2026)
Self-Improving CAD Generation Agents with Finite Element Analysis as Feedback
by: Son, Guijin, et al.
Published: (2026)
by: Son, Guijin, et al.
Published: (2026)
Understand, Solve and Translate: Bridging the Multilingual Mathematical Reasoning Gap
by: Ko, Hyunwoo, et al.
Published: (2025)
by: Ko, Hyunwoo, et al.
Published: (2025)
AccessEval: Benchmarking Disability Bias in Large Language Models
by: Panda, Srikant, et al.
Published: (2025)
by: Panda, Srikant, et al.
Published: (2025)
Tokenization Matters: Improving Zero-Shot NER for Indic Languages
by: Pattnayak, Priyaranjan, et al.
Published: (2025)
by: Pattnayak, Priyaranjan, et al.
Published: (2025)
LLM-Guided Lifecycle-Aware Clustering of Multi-Turn Customer Support Conversations
by: Pattnayak, Priyaranjan, et al.
Published: (2026)
by: Pattnayak, Priyaranjan, et al.
Published: (2026)
Multi-Step Reasoning in Korean and the Emergent Mirage
by: Son, Guijin, et al.
Published: (2025)
by: Son, Guijin, et al.
Published: (2025)
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation
by: Kim, Eunsu, et al.
Published: (2025)
by: Kim, Eunsu, et al.
Published: (2025)
Hybrid AI for Responsive Multi-Turn Online Conversations with Novel Dynamic Routing and Feedback Adaptation
by: Pattnayak, Priyaranjan, et al.
Published: (2025)
by: Pattnayak, Priyaranjan, et al.
Published: (2025)
Hard Negative Mining for Domain-Specific Retrieval in Enterprise Systems
by: Meghwani, Hansa, et al.
Published: (2025)
by: Meghwani, Hansa, et al.
Published: (2025)
Who's Asking? Investigating Bias Through the Lens of Disability Framed Queries in LLMs
by: Hari, Vishnu, et al.
Published: (2025)
by: Hari, Vishnu, et al.
Published: (2025)
LLM for Barcodes: Generating Diverse Synthetic Data for Identity Documents
by: Patel, Hitesh Laxmichand, et al.
Published: (2024)
by: Patel, Hitesh Laxmichand, et al.
Published: (2024)
RECOR: Reasoning-focused Multi-turn Conversational Retrieval Benchmark
by: Ali, Mohammed, et al.
Published: (2026)
by: Ali, Mohammed, et al.
Published: (2026)
Won: Establishing Best Practices for Korean Financial NLP
by: Son, Guijin, et al.
Published: (2025)
by: Son, Guijin, et al.
Published: (2025)
KAIO: A Collection of More Challenging Korean Questions
by: Lee, Nahyun, et al.
Published: (2025)
by: Lee, Nahyun, et al.
Published: (2025)
Linguistic Generalizability of Test-Time Scaling in Mathematical Reasoning
by: Son, Guijin, et al.
Published: (2025)
by: Son, Guijin, et al.
Published: (2025)
Controlling Language Confusion in Multilingual LLMs
by: Lee, Nahyun, et al.
Published: (2025)
by: Lee, Nahyun, et al.
Published: (2025)
Aligning LLMs for Multilingual Consistency in Enterprise Applications
by: Agarwal, Amit, et al.
Published: (2025)
by: Agarwal, Amit, et al.
Published: (2025)
Clinical QA 2.0: Multi-Task Learning for Answer Extraction and Categorization
by: Pattnayak, Priyaranjan, et al.
Published: (2025)
by: Pattnayak, Priyaranjan, et al.
Published: (2025)
Survey of Large Multimodal Model Datasets, Application Categories and Taxonomy
by: Pattnayak, Priyaranjan, et al.
Published: (2024)
by: Pattnayak, Priyaranjan, et al.
Published: (2024)
What Users Leave Unsaid: Under-Specified Queries Limit Vision-Language Models
by: Choi, Dasol, et al.
Published: (2026)
by: Choi, Dasol, et al.
Published: (2026)
DAIQ: Auditing Demographic Attribute Inference from Question in LLMs
by: Panda, Srikant, et al.
Published: (2025)
by: Panda, Srikant, et al.
Published: (2025)
SpeechWeave: Diverse Multilingual Synthetic Text & Audio Data Generation Pipeline for Training Text to Speech Models
by: Dua, Karan, et al.
Published: (2025)
by: Dua, Karan, et al.
Published: (2025)
When AI Co-Scientists Fail: SPOT-a Benchmark for Automated Verification of Scientific Research
by: Son, Guijin, et al.
Published: (2025)
by: Son, Guijin, et al.
Published: (2025)
PCRI: Measuring Context Robustness in Multimodal Models for Enterprise Applications
by: Patel, Hitesh Laxmichand, et al.
Published: (2025)
by: Patel, Hitesh Laxmichand, et al.
Published: (2025)
KMMMU: Evaluation of Massive Multi-discipline Multimodal Understanding in Korean Language and Context
by: Lee, Nahyun, et al.
Published: (2026)
by: Lee, Nahyun, et al.
Published: (2026)
SPENCE: A Syntactic Probe for Detecting Contamination in NL2SQL Benchmarks
by: Safarzadeh, Mohammadtaher, et al.
Published: (2026)
by: Safarzadeh, Mohammadtaher, et al.
Published: (2026)
Seeing as We Cannot
by: Carlos A. Postlethwaite
Published: (2021)
by: Carlos A. Postlethwaite
Published: (2021)
Pushing the Boundaries of Multiple Choice Evaluation to One Hundred Options
by: Lee, Nahyun, et al.
Published: (2026)
by: Lee, Nahyun, et al.
Published: (2026)
RCI: A Score for Evaluating Global and Local Reasoning in Multimodal Benchmarks
by: Agarwal, Amit, et al.
Published: (2025)
by: Agarwal, Amit, et al.
Published: (2025)
SweEval: Do LLMs Really Swear? A Safety Benchmark for Testing Limits for Enterprise Use
by: Patel, Hitesh Laxmichand, et al.
Published: (2025)
by: Patel, Hitesh Laxmichand, et al.
Published: (2025)
FlexDoc: Parameterized Sampling for Diverse Multilingual Synthetic Documents for Training Document Understanding Models
by: Dua, Karan, et al.
Published: (2025)
by: Dua, Karan, et al.
Published: (2025)
We Need Knowledge Distillation for Solving Math Word Problems
by: Shen, Zhenquan, et al.
Published: (2025)
by: Shen, Zhenquan, et al.
Published: (2025)
VLM Judges Can Rank but Cannot Score: Task-Dependent Uncertainty in Multimodal Evaluation
by: Kumar, Divake, et al.
Published: (2026)
by: Kumar, Divake, et al.
Published: (2026)
Beyond Interactions: Node-Level Graph Generation for Knowledge-Free Augmentation in Recommender Systems
by: Wang, Zhaoyan, et al.
Published: (2025)
by: Wang, Zhaoyan, et al.
Published: (2025)
E-Book Usability in Educational Technology Classes: Teachers and Teacher Candidates' Perception toward E-Book for Teaching and Learning
by: Shin, Sunghee
Published: (2014)
by: Shin, Sunghee
Published: (2014)
Keep What Audio Cannot Say: Context-Preserving Token Pruning for Omni-LLMs
by: Jung, Chaeyoung, et al.
Published: (2026)
by: Jung, Chaeyoung, et al.
Published: (2026)
Robust Audio-Text Retrieval via Cross-Modal Attention and Hybrid Loss
by: Liu, Meizhu, et al.
Published: (2026)
by: Liu, Meizhu, et al.
Published: (2026)
Similar Items
-
Pushing on Multilingual Reasoning Models with Language-Mixed Chain-of-Thought
by: Son, Guijin, et al.
Published: (2025) -
LLM-as-a-Judge & Reward Model: What They Can and Cannot Do
by: Son, Guijin, et al.
Published: (2024) -
ResearchMath-14K: Scaling Research-Level Mathematics via Agents
by: Son, Guijin, et al.
Published: (2026) -
Self-Improving CAD Generation Agents with Finite Element Analysis as Feedback
by: Son, Guijin, et al.
Published: (2026) -
Understand, Solve and Translate: Bridging the Multilingual Mathematical Reasoning Gap
by: Ko, Hyunwoo, et al.
Published: (2025)