How Far Can LLMs Improve from Experience? Measuring Test-Time Learning Ability in LLMs with Human Comparison
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Jiayin, Guo, Zhiquang, Ma, Weizhi, Zhang, Min |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
How Far Are We on the Decision-Making of LLMs? Evaluating LLMs' Gaming Ability in Multi-Agent Environments
por: Huang, Jen-tse, et al.
Publicado: (2024)
por: Huang, Jen-tse, et al.
Publicado: (2024)
Can LLMs Reason with Rules? Logic Scaffolding for Stress-Testing and Improving LLMs
por: Wang, Siyuan, et al.
Publicado: (2024)
por: Wang, Siyuan, et al.
Publicado: (2024)
StepTool: Enhancing Multi-Step Tool Usage in LLMs via Step-Grained Reinforcement Learning
por: Yu, Yuanqing, et al.
Publicado: (2024)
por: Yu, Yuanqing, et al.
Publicado: (2024)
LLMs Improving LLMs: Agentic Discovery for Test-Time Scaling
por: Zheng, Tong, et al.
Publicado: (2026)
por: Zheng, Tong, et al.
Publicado: (2026)
Model Editing for LLMs4Code: How Far are We?
por: Li, Xiaopeng, et al.
Publicado: (2024)
por: Li, Xiaopeng, et al.
Publicado: (2024)
LingGym: How Far Are LLMs from Thinking Like Field Linguists?
por: Yang, Changbing, et al.
Publicado: (2025)
por: Yang, Changbing, et al.
Publicado: (2025)
Teaching Human Behavior Improves Content Understanding Abilities Of LLMs
por: Singh, Somesh, et al.
Publicado: (2024)
por: Singh, Somesh, et al.
Publicado: (2024)
How Far Are LLMs from Believable AI? A Benchmark for Evaluating the Believability of Human Behavior Simulation
por: Xiao, Yang, et al.
Publicado: (2023)
por: Xiao, Yang, et al.
Publicado: (2023)
How Well Can LLMs Echo Us? Evaluating AI Chatbots' Role-Play Ability with ECHO
por: Ng, Man Tik, et al.
Publicado: (2024)
por: Ng, Man Tik, et al.
Publicado: (2024)
LLMs for Relational Reasoning: How Far are We?
por: Li, Zhiming, et al.
Publicado: (2024)
por: Li, Zhiming, et al.
Publicado: (2024)
How Far Are We? Systematic Evaluation of LLMs vs. Human Experts in Mathematical Contest in Modeling
por: Liu, Yuhang, et al.
Publicado: (2026)
por: Liu, Yuhang, et al.
Publicado: (2026)
Diversity of Thought Improves Reasoning Abilities of LLMs
por: Naik, Ranjita, et al.
Publicado: (2023)
por: Naik, Ranjita, et al.
Publicado: (2023)
Learning Beyond the Surface: How Far Can Continual Pre-Training with LoRA Enhance LLMs' Domain-Specific Insight Learning?
por: Pezeshkpour, Pouya, et al.
Publicado: (2025)
por: Pezeshkpour, Pouya, et al.
Publicado: (2025)
How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs
por: Zeng, Yi, et al.
Publicado: (2024)
por: Zeng, Yi, et al.
Publicado: (2024)
A User-Centric Multi-Intent Benchmark for Evaluating Large Language Models
por: Wang, Jiayin, et al.
Publicado: (2024)
por: Wang, Jiayin, et al.
Publicado: (2024)
Measuring Bargaining Abilities of LLMs: A Benchmark and A Buyer-Enhancement Method
por: Xia, Tian, et al.
Publicado: (2024)
por: Xia, Tian, et al.
Publicado: (2024)
Improving Fairness in LLMs Through Testing-Time Adversaries
por: Gregio, Isabela Pereira, et al.
Publicado: (2025)
por: Gregio, Isabela Pereira, et al.
Publicado: (2025)
Adaptive Layer-skipping in Pre-trained LLMs
por: Luo, Xuan, et al.
Publicado: (2025)
por: Luo, Xuan, et al.
Publicado: (2025)
Can LLMs Learn from Previous Mistakes? Investigating LLMs' Errors to Boost for Reasoning
por: Tong, Yongqi, et al.
Publicado: (2024)
por: Tong, Yongqi, et al.
Publicado: (2024)
Can We Trust LLMs on Memristors? Diving into Reasoning Ability under Non-Ideality
por: Wu, Taiqiang, et al.
Publicado: (2026)
por: Wu, Taiqiang, et al.
Publicado: (2026)
DORA Explorer: Improving the Exploration Ability of LLMs Without Training
por: Gurjar, Priya, et al.
Publicado: (2026)
por: Gurjar, Priya, et al.
Publicado: (2026)
Can LLMs Learn to Map the World from Local Descriptions?
por: Xia, Sirui, et al.
Publicado: (2025)
por: Xia, Sirui, et al.
Publicado: (2025)
Can LLMs Capture Human Preferences?
por: Goli, Ali, et al.
Publicado: (2023)
por: Goli, Ali, et al.
Publicado: (2023)
Can LLMs Reliably Simulate Real Students' Abilities in Mathematics and Reading Comprehension?
por: Srivatsa, KV Aditya, et al.
Publicado: (2025)
por: Srivatsa, KV Aditya, et al.
Publicado: (2025)
How Far Are We From AGI: Are LLMs All We Need?
por: Feng, Tao, et al.
Publicado: (2024)
por: Feng, Tao, et al.
Publicado: (2024)
Large Language Models as Evaluators for Recommendation Explanations
por: Zhang, Xiaoyu, et al.
Publicado: (2024)
por: Zhang, Xiaoyu, et al.
Publicado: (2024)
Distill Visual Chart Reasoning Ability from LLMs to MLLMs
por: He, Wei, et al.
Publicado: (2024)
por: He, Wei, et al.
Publicado: (2024)
Can LLMs Translate Human Instructions into a Reinforcement Learning Agent's Internal Emergent Symbolic Representation?
por: Ma, Ziqi, et al.
Publicado: (2025)
por: Ma, Ziqi, et al.
Publicado: (2025)
Understanding the Ability of LLMs to Handle Character-Level Perturbation
por: Zhuo, Anyuan, et al.
Publicado: (2025)
por: Zhuo, Anyuan, et al.
Publicado: (2025)
Do LLMs Have the Generalization Ability in Conducting Causal Inference?
por: Wang, Chen, et al.
Publicado: (2024)
por: Wang, Chen, et al.
Publicado: (2024)
Quantifying the Reasoning Abilities of LLMs on Real-world Clinical Cases
por: Qiu, Pengcheng, et al.
Publicado: (2025)
por: Qiu, Pengcheng, et al.
Publicado: (2025)
RareBench: Can LLMs Serve as Rare Diseases Specialists?
por: Chen, Xuanzhong, et al.
Publicado: (2024)
por: Chen, Xuanzhong, et al.
Publicado: (2024)
How Good Are LLMs for Literary Translation, Really? Literary Translation Evaluation with Humans and LLMs
por: Zhang, Ran, et al.
Publicado: (2024)
por: Zhang, Ran, et al.
Publicado: (2024)
ScholarSearch: Benchmarking Scholar Searching Ability of LLMs
por: Zhou, Junting, et al.
Publicado: (2025)
por: Zhou, Junting, et al.
Publicado: (2025)
Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning
por: Fei, Zhaoye, et al.
Publicado: (2025)
por: Fei, Zhaoye, et al.
Publicado: (2025)
How Deep is Love in LLMs' Hearts? Exploring Semantic Size in Human-like Cognition
por: Yao, Yao, et al.
Publicado: (2025)
por: Yao, Yao, et al.
Publicado: (2025)
Can LLMs Beat Humans in Debating? A Dynamic Multi-agent Framework for Competitive Debate
por: Zhang, Yiqun, et al.
Publicado: (2024)
por: Zhang, Yiqun, et al.
Publicado: (2024)
Using RL to Identify Divisive Perspectives Improves LLMs Abilities to Identify Communities on Social Media
por: Mehta, Nikhil, et al.
Publicado: (2024)
por: Mehta, Nikhil, et al.
Publicado: (2024)
Effective Distillation of Table-based Reasoning Ability from LLMs
por: Yang, Bohao, et al.
Publicado: (2023)
por: Yang, Bohao, et al.
Publicado: (2023)
Beyond English-Centric Training: How Reinforcement Learning Improves Cross-Lingual Reasoning in LLMs
por: Huang, Shulin, et al.
Publicado: (2025)
por: Huang, Shulin, et al.
Publicado: (2025)
Ejemplares similares
-
How Far Are We on the Decision-Making of LLMs? Evaluating LLMs' Gaming Ability in Multi-Agent Environments
por: Huang, Jen-tse, et al.
Publicado: (2024) -
Can LLMs Reason with Rules? Logic Scaffolding for Stress-Testing and Improving LLMs
por: Wang, Siyuan, et al.
Publicado: (2024) -
StepTool: Enhancing Multi-Step Tool Usage in LLMs via Step-Grained Reinforcement Learning
por: Yu, Yuanqing, et al.
Publicado: (2024) -
LLMs Improving LLMs: Agentic Discovery for Test-Time Scaling
por: Zheng, Tong, et al.
Publicado: (2026) -
Model Editing for LLMs4Code: How Far are We?
por: Li, Xiaopeng, et al.
Publicado: (2024)