EVOLvE: Evaluating and Optimizing LLMs For In-Context Exploration
Fuente:
arXiv
Salvato in:
| Autori principali: | Nie, Allen, Su, Yi, Chang, Bo, Lee, Jonathan N., Chi, Ed H., Le, Quoc V., Chen, Minmin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
NATURAL PLAN: Benchmarking LLMs on Natural Language Planning
di: Zheng, Huaixiu Steven, et al.
Pubblicazione: (2024)
di: Zheng, Huaixiu Steven, et al.
Pubblicazione: (2024)
In-Context Learning with Long-Context Models: An In-Depth Exploration
di: Bertsch, Amanda, et al.
Pubblicazione: (2024)
di: Bertsch, Amanda, et al.
Pubblicazione: (2024)
SEE: Strategic Exploration and Exploitation for Cohesive In-Context Prompt Optimization
di: Cui, Wendi, et al.
Pubblicazione: (2024)
di: Cui, Wendi, et al.
Pubblicazione: (2024)
UIO-LLMs: Unbiased Incremental Optimization for Long-Context LLMs
di: Li, Wenhao, et al.
Pubblicazione: (2024)
di: Li, Wenhao, et al.
Pubblicazione: (2024)
Take a Step Back: Evoking Reasoning via Abstraction in Large Language Models
di: Zheng, Huaixiu Steven, et al.
Pubblicazione: (2023)
di: Zheng, Huaixiu Steven, et al.
Pubblicazione: (2023)
HELMET: How to Evaluate Long-Context Language Models Effectively and Thoroughly
di: Yen, Howard, et al.
Pubblicazione: (2024)
di: Yen, Howard, et al.
Pubblicazione: (2024)
QwenLong-CPRS: Towards $\infty$-LLMs with Dynamic Context Optimization
di: Shen, Weizhou, et al.
Pubblicazione: (2025)
di: Shen, Weizhou, et al.
Pubblicazione: (2025)
Evaluating the Sensitivity of LLMs to Prior Context
di: Hankache, Robert, et al.
Pubblicazione: (2025)
di: Hankache, Robert, et al.
Pubblicazione: (2025)
Long Context vs. RAG for LLMs: An Evaluation and Revisits
di: Li, Xinze, et al.
Pubblicazione: (2024)
di: Li, Xinze, et al.
Pubblicazione: (2024)
Michelangelo: Long Context Evaluations Beyond Haystacks via Latent Structure Queries
di: Vodrahalli, Kiran, et al.
Pubblicazione: (2024)
di: Vodrahalli, Kiran, et al.
Pubblicazione: (2024)
How LLMs Comprehend Temporal Meaning in Narratives: A Case Study in Cognitive Evaluation of LLMs
di: de Langis, Karin, et al.
Pubblicazione: (2025)
di: de Langis, Karin, et al.
Pubblicazione: (2025)
Conversations Gone Awry, But Then? Evaluating Conversational Forecasting Models
di: Tran, Son Quoc, et al.
Pubblicazione: (2025)
di: Tran, Son Quoc, et al.
Pubblicazione: (2025)
Benchmarking Cognitive Domains for LLMs: Insights from Taiwanese Hakka Culture
di: Chang, Chen-Chi, et al.
Pubblicazione: (2024)
di: Chang, Chen-Chi, et al.
Pubblicazione: (2024)
DoubleDipper: Improving Long-Context LLMs via Context Recycling
di: Cattan, Arie, et al.
Pubblicazione: (2024)
di: Cattan, Arie, et al.
Pubblicazione: (2024)
PICACO: Pluralistic In-Context Value Alignment of LLMs via Total Correlation Optimization
di: Jiang, Han, et al.
Pubblicazione: (2025)
di: Jiang, Han, et al.
Pubblicazione: (2025)
Agents that Matter: Optimizing Multi-Agent LLMs via Removal-Based Attribution
di: Lu, Mingyu, et al.
Pubblicazione: (2026)
di: Lu, Mingyu, et al.
Pubblicazione: (2026)
CURIE: Evaluating LLMs On Multitask Scientific Long Context Understanding and Reasoning
di: Cui, Hao, et al.
Pubblicazione: (2025)
di: Cui, Hao, et al.
Pubblicazione: (2025)
When Names Disappear: Revealing What LLMs Actually Understand About Code
di: Le, Cuong Chi, et al.
Pubblicazione: (2025)
di: Le, Cuong Chi, et al.
Pubblicazione: (2025)
Systematic Evaluation of Long-Context LLMs on Financial Concepts
di: Gupta, Lavanya, et al.
Pubblicazione: (2024)
di: Gupta, Lavanya, et al.
Pubblicazione: (2024)
DYCP: Dynamic Context Pruning for Long-Form Dialogue with LLMs
di: Choi, Nayoung, et al.
Pubblicazione: (2026)
di: Choi, Nayoung, et al.
Pubblicazione: (2026)
Enhancing Low-Resource Minority Language Translation with LLMs and Retrieval-Augmented Generation for Cultural Nuances
di: Chang, Chen-Chi, et al.
Pubblicazione: (2025)
di: Chang, Chen-Chi, et al.
Pubblicazione: (2025)
Self-Discover: Large Language Models Self-Compose Reasoning Structures
di: Zhou, Pei, et al.
Pubblicazione: (2024)
di: Zhou, Pei, et al.
Pubblicazione: (2024)
SAGE: Shaping Anchors for Guided Exploration in RLVR of LLMs
di: Lee, Chanuk, et al.
Pubblicazione: (2026)
di: Lee, Chanuk, et al.
Pubblicazione: (2026)
Stochastic Chameleons: Irrelevant Context Hallucinations Reveal Class-Based (Mis)Generalization in LLMs
di: Cheng, Ziling, et al.
Pubblicazione: (2025)
di: Cheng, Ziling, et al.
Pubblicazione: (2025)
An Exploration of Higher Education Course Evaluation by Large Language Models
di: Yuan, Bo, et al.
Pubblicazione: (2024)
di: Yuan, Bo, et al.
Pubblicazione: (2024)
Large Language Models as Optimizers
di: Yang, Chengrun, et al.
Pubblicazione: (2023)
di: Yang, Chengrun, et al.
Pubblicazione: (2023)
EarthSE: A Benchmark for Evaluating Earth Scientific Exploration Capability of LLMs
di: Xu, Wanghan, et al.
Pubblicazione: (2025)
di: Xu, Wanghan, et al.
Pubblicazione: (2025)
FineSurE: Fine-grained Summarization Evaluation using LLMs
di: Song, Hwanjun, et al.
Pubblicazione: (2024)
di: Song, Hwanjun, et al.
Pubblicazione: (2024)
EEPO: Exploration-Enhanced Policy Optimization via Sample-Then-Forget
di: Chen, Liang, et al.
Pubblicazione: (2025)
di: Chen, Liang, et al.
Pubblicazione: (2025)
Knapsack RL: Unlocking Exploration of LLMs via Optimizing Budget Allocation
di: Li, Ziniu, et al.
Pubblicazione: (2025)
di: Li, Ziniu, et al.
Pubblicazione: (2025)
Not All Contexts Are Equal: Teaching LLMs Credibility-aware Generation
di: Pan, Ruotong, et al.
Pubblicazione: (2024)
di: Pan, Ruotong, et al.
Pubblicazione: (2024)
InfiniPot: Infinite Context Processing on Memory-Constrained LLMs
di: Kim, Minsoo, et al.
Pubblicazione: (2024)
di: Kim, Minsoo, et al.
Pubblicazione: (2024)
Dissecting Multiplication in Transformers: Insights into LLMs
di: Qiu, Luyu, et al.
Pubblicazione: (2024)
di: Qiu, Luyu, et al.
Pubblicazione: (2024)
Training-Inference Consistent Segmented Execution for Long-Context LLMs
di: Shang, Xianpeng, et al.
Pubblicazione: (2026)
di: Shang, Xianpeng, et al.
Pubblicazione: (2026)
Multi-Modal Data Exploration via Language Agents
di: Nooralahzadeh, Farhad, et al.
Pubblicazione: (2024)
di: Nooralahzadeh, Farhad, et al.
Pubblicazione: (2024)
Gavel: Agent Meets Checklist for Evaluating LLMs on Long-Context Legal Summarization
di: Dou, Yao, et al.
Pubblicazione: (2026)
di: Dou, Yao, et al.
Pubblicazione: (2026)
Evaluating LLMs in the Context of a Functional Programming Course: A Comprehensive Study
di: Zhang, Yihan, et al.
Pubblicazione: (2026)
di: Zhang, Yihan, et al.
Pubblicazione: (2026)
ArithmAttack: Evaluating Robustness of LLMs to Noisy Context in Math Problem Solving
di: Abedin, Zain Ul, et al.
Pubblicazione: (2025)
di: Abedin, Zain Ul, et al.
Pubblicazione: (2025)
Data-efficient Meta-models for Evaluation of Context-based Questions and Answers in LLMs
di: Belikova, Julia, et al.
Pubblicazione: (2025)
di: Belikova, Julia, et al.
Pubblicazione: (2025)
Evaluating Gender Bias of LLMs in Making Morality Judgements
di: Bajaj, Divij, et al.
Pubblicazione: (2024)
di: Bajaj, Divij, et al.
Pubblicazione: (2024)
Documenti analoghi
-
NATURAL PLAN: Benchmarking LLMs on Natural Language Planning
di: Zheng, Huaixiu Steven, et al.
Pubblicazione: (2024) -
In-Context Learning with Long-Context Models: An In-Depth Exploration
di: Bertsch, Amanda, et al.
Pubblicazione: (2024) -
SEE: Strategic Exploration and Exploitation for Cohesive In-Context Prompt Optimization
di: Cui, Wendi, et al.
Pubblicazione: (2024) -
UIO-LLMs: Unbiased Incremental Optimization for Long-Context LLMs
di: Li, Wenhao, et al.
Pubblicazione: (2024) -
Take a Step Back: Evoking Reasoning via Abstraction in Large Language Models
di: Zheng, Huaixiu Steven, et al.
Pubblicazione: (2023)