Instance-level Randomization: Toward More Stable LLM Evaluations
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Yiyang, Wu, Yonghuang, Luo, Ying, Sun, Liangtai, Qin, Zishu, Qiu, Lin, Cao, Xuezhi, Cai, Xunliang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AgentEscapeBench: Evaluating Out-of-Domain Tool-Grounded Reasoning in LLM Agents
by: Guo, Zhengkang, et al.
Published: (2026)
by: Guo, Zhengkang, et al.
Published: (2026)
AMemGym: Interactive Memory Benchmarking for Assistants in Long-Horizon Conversations
by: Jiayang, Cheng, et al.
Published: (2026)
by: Jiayang, Cheng, et al.
Published: (2026)
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics
by: Zhu, Yaoming, et al.
Published: (2025)
by: Zhu, Yaoming, et al.
Published: (2025)
CATArena: Evaluating Evolutionary Capabilities of Code Agents via Iterative Tournaments
by: Fu, Lingyue, et al.
Published: (2025)
by: Fu, Lingyue, et al.
Published: (2025)
META-GUI: Towards Multi-modal Conversational Agents on Mobile GUI
by: Sun, Liangtai, et al.
Published: (2022)
by: Sun, Liangtai, et al.
Published: (2022)
MUSE: MCTS-Driven Red Teaming Framework for Enhanced Multi-Turn Dialogue Safety in Large Language Models
by: Yan, Siyu, et al.
Published: (2025)
by: Yan, Siyu, et al.
Published: (2025)
MaskPrune: Mask-based LLM Pruning for Layer-wise Uniform Structures
by: Qin, Jiayu, et al.
Published: (2025)
by: Qin, Jiayu, et al.
Published: (2025)
Making Mathematical Reasoning Adaptive
by: Lai, Zhejian, et al.
Published: (2025)
by: Lai, Zhejian, et al.
Published: (2025)
General365: Benchmarking General Reasoning in Large Language Models Across Diverse and Challenging Tasks
by: Liu, Junlin, et al.
Published: (2026)
by: Liu, Junlin, et al.
Published: (2026)
Seeing the Needle in the Haystack: Towards Weakly-Supervised Log Instance Anomaly Localization via Counterfactual Perturbation
by: Wong, Yutszyuk, et al.
Published: (2026)
by: Wong, Yutszyuk, et al.
Published: (2026)
AMO-Bench: Large Language Models Still Struggle in High School Math Competitions
by: An, Shengnan, et al.
Published: (2025)
by: An, Shengnan, et al.
Published: (2025)
A Simple Ensemble Strategy for LLM Inference: Towards More Stable Text Classification
by: Niimi, Junichiro
Published: (2025)
by: Niimi, Junichiro
Published: (2025)
R-Horizon: How Far Can Your Large Reasoning Model Really Go in Breadth and Depth?
by: Lu, Yi, et al.
Published: (2025)
by: Lu, Yi, et al.
Published: (2025)
MemBench: Towards More Comprehensive Evaluation on the Memory of LLM-based Agents
by: Tan, Haoran, et al.
Published: (2025)
by: Tan, Haoran, et al.
Published: (2025)
Towards Evaluation for Real-World LLM Unlearning
by: Miao, Ke, et al.
Published: (2025)
by: Miao, Ke, et al.
Published: (2025)
Model Tampering Attacks Enable More Rigorous Evaluations of LLM Capabilities
by: Che, Zora, et al.
Published: (2025)
by: Che, Zora, et al.
Published: (2025)
Toward a More Complete OMR Solution
by: Yang, Guang, et al.
Published: (2024)
by: Yang, Guang, et al.
Published: (2024)
MASPO: Unifying Gradient Utilization, Probability Mass, and Signal Reliability for Robust and Sample-Efficient LLM Reasoning
by: Fu, Xiaoliang, et al.
Published: (2026)
by: Fu, Xiaoliang, et al.
Published: (2026)
Cross-Patient Pseudo Bags Generation and Curriculum Contrastive Learning for Imbalanced Multiclassification of Whole Slide Image
by: Wu, Yonghuang, et al.
Published: (2024)
by: Wu, Yonghuang, et al.
Published: (2024)
Prompt Group-Aware Training for Robust Text-Guided Nuclei Segmentation
by: Wu, Yonghuang, et al.
Published: (2026)
by: Wu, Yonghuang, et al.
Published: (2026)
UNO-Bench: A Unified Benchmark for Exploring the Compositional Law Between Uni-modal and Omni-modal in Omni Models
by: Chen, Chen, et al.
Published: (2025)
by: Chen, Chen, et al.
Published: (2025)
Leveraging Dual Process Theory in Language Agent Framework for Real-time Simultaneous Human-AI Collaboration
by: Zhang, Shao, et al.
Published: (2025)
by: Zhang, Shao, et al.
Published: (2025)
AutoContext: Instance-Level Context Learning for LLM Agents
by: Cai, Kuntai, et al.
Published: (2025)
by: Cai, Kuntai, et al.
Published: (2025)
Towards Reliable LLM Evaluation: Correcting the Winner's Curse in Adaptive Benchmarking
by: Xu, Yang, et al.
Published: (2026)
by: Xu, Yang, et al.
Published: (2026)
C2T: A Classifier-Based Tree Construction Method in Speculative Decoding
by: Huo, Feiye, et al.
Published: (2025)
by: Huo, Feiye, et al.
Published: (2025)
Neural Network Optimal Power Flow via Energy Gradient Flow and Unified Dynamics
by: Liu, Xuezhi
Published: (2025)
by: Liu, Xuezhi
Published: (2025)
Physics-Constrained Neural Dynamics: A Unified Manifold Framework for Large-Scale Power Flow Computation
by: Liu, Xuezhi
Published: (2025)
by: Liu, Xuezhi
Published: (2025)
Privacy in Action: Towards Realistic Privacy Mitigation and Evaluation for LLM-Powered Agents
by: Wang, Shouju, et al.
Published: (2025)
by: Wang, Shouju, et al.
Published: (2025)
Evaluative Fingerprints: Stable and Systematic Differences in LLM Evaluator Behavior
by: Nasser, Wajid
Published: (2026)
by: Nasser, Wajid
Published: (2026)
AgentNoiseBench: Benchmarking Robustness of Tool-Using LLM Agents Under Noisy Condition
by: Wang, Ruipeng, et al.
Published: (2026)
by: Wang, Ruipeng, et al.
Published: (2026)
S^3cMath: Spontaneous Step-level Self-correction Makes Large Language Models Better Mathematical Reasoners
by: Yan, Yuchen, et al.
Published: (2024)
by: Yan, Yuchen, et al.
Published: (2024)
FeatBench: Towards More Realistic Evaluation of Feature-level Code Generation
by: Chen, Haorui, et al.
Published: (2025)
by: Chen, Haorui, et al.
Published: (2025)
Look Before You Leap: Autonomous Exploration for LLM Agents
by: Ye, Ziang, et al.
Published: (2026)
by: Ye, Ziang, et al.
Published: (2026)
Interactive Learning for LLM Reasoning
by: Lin, Hehai, et al.
Published: (2025)
by: Lin, Hehai, et al.
Published: (2025)
Towards Self-Robust LLMs: Intrinsic Prompt Noise Resistance via CoIPO
by: Yang, Xin, et al.
Published: (2026)
by: Yang, Xin, et al.
Published: (2026)
Does Theory of Mind Improvement Really Benefit Human-AI Interactions? Empirical Findings from Interactive Evaluations
by: Gong, Nanxu, et al.
Published: (2026)
by: Gong, Nanxu, et al.
Published: (2026)
If an LLM Were a Character, Would It Know Its Own Story? Evaluating Lifelong Learning in LLMs
by: Fan, Siqi, et al.
Published: (2025)
by: Fan, Siqi, et al.
Published: (2025)
Not All Instances Are Equally Valuable: Towards Influence-Weighted Dataset Distillation
by: Deng, Qiyan, et al.
Published: (2025)
by: Deng, Qiyan, et al.
Published: (2025)
Bridging Synthetic and Real Routing Problems via LLM-Guided Instance Generation and Progressive Adaptation
by: Zhu, Jianghan, et al.
Published: (2025)
by: Zhu, Jianghan, et al.
Published: (2025)
Understanding and Optimizing Agentic Workflows via Shapley value
by: Yang, Yingxuan, et al.
Published: (2025)
by: Yang, Yingxuan, et al.
Published: (2025)
Similar Items
-
AgentEscapeBench: Evaluating Out-of-Domain Tool-Grounded Reasoning in LLM Agents
by: Guo, Zhengkang, et al.
Published: (2026) -
AMemGym: Interactive Memory Benchmarking for Assistants in Long-Horizon Conversations
by: Jiayang, Cheng, et al.
Published: (2026) -
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics
by: Zhu, Yaoming, et al.
Published: (2025) -
CATArena: Evaluating Evolutionary Capabilities of Code Agents via Iterative Tournaments
by: Fu, Lingyue, et al.
Published: (2025) -
META-GUI: Towards Multi-modal Conversational Agents on Mobile GUI
by: Sun, Liangtai, et al.
Published: (2022)