HellaSwag-Pro: A Large-Scale Bilingual Benchmark for Evaluating the Robustness of LLMs in Commonsense Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Xiaoyuan, Li, Moxin, Men, Rui, Zhang, Yichang, Bao, Keqin, Wang, Wenjie, Feng, Fuli, Liu, Dayiheng, Lin, Junyang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
What the HellaSwag? On the Validity of Common-Sense Reasoning Benchmarks
by: Chizhov, Pavel, et al.
Published: (2025)
by: Chizhov, Pavel, et al.
Published: (2025)
MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning Evaluation
by: Li, Xiaoyuan, et al.
Published: (2025)
by: Li, Xiaoyuan, et al.
Published: (2025)
MathOPEval: A Fine-grained Evaluation Benchmark for Visual Operations of MLLMs in Mathematical Reasoning
by: Li, Xiaoyuan, et al.
Published: (2025)
by: Li, Xiaoyuan, et al.
Published: (2025)
SAGE: Scalable Automated Robustness Augmentation for LLM Knowledge Evaluation
by: Li, Xiaoyuan, et al.
Published: (2026)
by: Li, Xiaoyuan, et al.
Published: (2026)
ARES: Automated Rubric Synthesis for Scalable LLM Reinforcement Learning
by: Li, Xiaoyuan, et al.
Published: (2026)
by: Li, Xiaoyuan, et al.
Published: (2026)
SkillGraph: Skill-Augmented Reinforcement Learning for Agents via Evolving Skill Graphs
by: Li, Xiaoyuan, et al.
Published: (2026)
by: Li, Xiaoyuan, et al.
Published: (2026)
On Predicting the Post-training Potential of Pre-trained LLMs
by: Li, Xiaoyuan, et al.
Published: (2026)
by: Li, Xiaoyuan, et al.
Published: (2026)
Evaluating Mathematical Reasoning of Large Language Models: A Focus on Error Identification and Correction
by: Li, Xiaoyuan, et al.
Published: (2024)
by: Li, Xiaoyuan, et al.
Published: (2024)
Unified Data Selection for LLM Reasoning
by: Li, Xiaoyuan, et al.
Published: (2026)
by: Li, Xiaoyuan, et al.
Published: (2026)
Teaching LLM to Reason: Reinforcement Learning from Algorithmic Problems without Code
by: Bao, Keqin, et al.
Published: (2025)
by: Bao, Keqin, et al.
Published: (2025)
Chain of Execution Supervision Promotes General Reasoning in Large Language Models
by: Chen, Nuo, et al.
Published: (2025)
by: Chen, Nuo, et al.
Published: (2025)
Counterfactual Debating with Preset Stances for Hallucination Elimination of LLMs
by: Fang, Yi, et al.
Published: (2024)
by: Fang, Yi, et al.
Published: (2024)
Robust Prompt Optimization for Large Language Models Against Distribution Shifts
by: Li, Moxin, et al.
Published: (2023)
by: Li, Moxin, et al.
Published: (2023)
Teaching Language Models to Reason with Tools
by: Li, Chengpeng, et al.
Published: (2025)
by: Li, Chengpeng, et al.
Published: (2025)
Controllable LLM Reasoning via Sparse Autoencoder-Based Steering
by: Fang, Yi, et al.
Published: (2026)
by: Fang, Yi, et al.
Published: (2026)
CoRT: Code-integrated Reasoning within Thinking
by: Li, Chengpeng, et al.
Published: (2025)
by: Li, Chengpeng, et al.
Published: (2025)
Think Twice Before Trusting: Self-Detection for Large Language Models through Comprehensive Answer Reflection
by: Li, Moxin, et al.
Published: (2024)
by: Li, Moxin, et al.
Published: (2024)
Text-like Encoding of Collaborative Information in Large Language Models for Recommendation
by: Zhang, Yang, et al.
Published: (2024)
by: Zhang, Yang, et al.
Published: (2024)
Plant, Persist, Trigger: Sleeper Attack on Large Language Model Agents
by: Li, Yongxiang, et al.
Published: (2026)
by: Li, Yongxiang, et al.
Published: (2026)
Byzantion Nea Hellás
Published: (2011)
Published: (2011)
TAT-LLM: A Specialized Language Model for Discrete Reasoning over Tabular and Textual Data
by: Zhu, Fengbin, et al.
Published: (2024)
by: Zhu, Fengbin, et al.
Published: (2024)
Benchmarking Chinese Commonsense Reasoning of LLMs: From Chinese-Specifics to Reasoning-Memorization Correlations
by: Sun, Jiaxing, et al.
Published: (2024)
by: Sun, Jiaxing, et al.
Published: (2024)
Assistant-Guided Mitigation of Teacher Preference Bias in LLM-as-a-Judge
by: Liu, Zhuo, et al.
Published: (2025)
by: Liu, Zhuo, et al.
Published: (2025)
Doc2SoarGraph: Discrete Reasoning over Visually-Rich Table-Text Documents via Semantic-Oriented Hierarchical Graphs
by: Zhu, Fengbin, et al.
Published: (2023)
by: Zhu, Fengbin, et al.
Published: (2023)
Benchmarking Chinese Commonsense Reasoning with a Multi-hop Reasoning Perspective
by: You, Wangjie, et al.
Published: (2025)
by: You, Wangjie, et al.
Published: (2025)
Self-Improvement Towards Pareto Optimality: Mitigating Preference Conflicts in Multi-Objective Alignment
by: Li, Moxin, et al.
Published: (2025)
by: Li, Moxin, et al.
Published: (2025)
Real-Time Personalization for LLM-based Recommendation with Customized In-Context Learning
by: Bao, Keqin, et al.
Published: (2024)
by: Bao, Keqin, et al.
Published: (2024)
Item-side Fairness of Large Language Model-based Recommendation System
by: Jiang, Meng, et al.
Published: (2024)
by: Jiang, Meng, et al.
Published: (2024)
SCoRE: Benchmarking Long-Chain Reasoning in Commonsense Scenarios
by: Zhan, Weidong, et al.
Published: (2025)
by: Zhan, Weidong, et al.
Published: (2025)
GRASP: A Grid-Based Benchmark for Evaluating Commonsense Spatial Reasoning
by: Tang, Zhisheng, et al.
Published: (2024)
by: Tang, Zhisheng, et al.
Published: (2024)
Causal Debiasing for Visual Commonsense Reasoning
by: Zou, Jiayi, et al.
Published: (2025)
by: Zou, Jiayi, et al.
Published: (2025)
CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
by: Quan, Shanghaoran, et al.
Published: (2025)
by: Quan, Shanghaoran, et al.
Published: (2025)
Prospect Personalized Recommendation on Large Language Model-based Agent Platform
by: Zhang, Jizhi, et al.
Published: (2024)
by: Zhang, Jizhi, et al.
Published: (2024)
Boosting Parameter Efficiency in LLM-Based Recommendation through Sophisticated Pruning
by: Zheng, Shanle, et al.
Published: (2025)
by: Zheng, Shanle, et al.
Published: (2025)
CoLLM: Integrating Collaborative Embeddings into Large Language Models for Recommendation
by: Zhang, Yang, et al.
Published: (2023)
by: Zhang, Yang, et al.
Published: (2023)
Decoding Matters: Addressing Amplification Bias and Homogeneity Issue for LLM-based Recommendation
by: Bao, Keqin, et al.
Published: (2024)
by: Bao, Keqin, et al.
Published: (2024)
MSQA: Benchmarking LLMs on Graduate-Level Materials Science Reasoning and Knowledge
by: Cheung, Jerry Junyang, et al.
Published: (2025)
by: Cheung, Jerry Junyang, et al.
Published: (2025)
Leveraging LLMs for Influence Path Planning in Proactive Recommendation
by: Wang, Mingze, et al.
Published: (2024)
by: Wang, Mingze, et al.
Published: (2024)
CAPSUL: A Comprehensive Human Protein Benchmark for Subcellular Localization
by: Hu, Yicheng, et al.
Published: (2026)
by: Hu, Yicheng, et al.
Published: (2026)
Disentangling Reasoning Tokens and Boilerplate Tokens For Language Model Fine-tuning
by: Ye, Ziang, et al.
Published: (2024)
by: Ye, Ziang, et al.
Published: (2024)
Similar Items
-
What the HellaSwag? On the Validity of Common-Sense Reasoning Benchmarks
by: Chizhov, Pavel, et al.
Published: (2025) -
MTR-Bench: A Comprehensive Benchmark for Multi-Turn Reasoning Evaluation
by: Li, Xiaoyuan, et al.
Published: (2025) -
MathOPEval: A Fine-grained Evaluation Benchmark for Visual Operations of MLLMs in Mathematical Reasoning
by: Li, Xiaoyuan, et al.
Published: (2025) -
SAGE: Scalable Automated Robustness Augmentation for LLM Knowledge Evaluation
by: Li, Xiaoyuan, et al.
Published: (2026) -
ARES: Automated Rubric Synthesis for Scalable LLM Reinforcement Learning
by: Li, Xiaoyuan, et al.
Published: (2026)