Efficient Exploration for LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Dwaracherla, Vikranth, Asghari, Seyed Mohammad, Hao, Botao, Van Roy, Benjamin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Efficient Exploration at Scale
por: Asghari, Seyed Mohammad, et al.
Publicado: (2026)
por: Asghari, Seyed Mohammad, et al.
Publicado: (2026)
Removing Spurious Correlation from Neural Network Interpretations
por: Fotouhi, Milad, et al.
Publicado: (2024)
por: Fotouhi, Milad, et al.
Publicado: (2024)
Evaluating Interventional Reasoning Capabilities of Large Language Models
por: Kasetty, Tejas, et al.
Publicado: (2024)
por: Kasetty, Tejas, et al.
Publicado: (2024)
AutoEval Done Right: Using Synthetic Data for Model Evaluation
por: Boyeau, Pierre, et al.
Publicado: (2024)
por: Boyeau, Pierre, et al.
Publicado: (2024)
ALCM: Autonomous LLM-Augmented Causal Discovery Framework
por: Khatibi, Elahe, et al.
Publicado: (2024)
por: Khatibi, Elahe, et al.
Publicado: (2024)
Adaptive Uncertainty Quantification for Generative AI
por: Kim, Jungeum, et al.
Publicado: (2024)
por: Kim, Jungeum, et al.
Publicado: (2024)
Industrial-Grade Smart Troubleshooting through Causal Technical Language Processing: a Proof of Concept
por: Trilla, Alexandre, et al.
Publicado: (2024)
por: Trilla, Alexandre, et al.
Publicado: (2024)
Propagation and Pitfalls: Reasoning-based Assessment of Knowledge Editing through Counterfactual Tasks
por: Hua, Wenyue, et al.
Publicado: (2024)
por: Hua, Wenyue, et al.
Publicado: (2024)
CLEAR: Can Language Models Really Understand Causal Graphs?
por: Chen, Sirui, et al.
Publicado: (2024)
por: Chen, Sirui, et al.
Publicado: (2024)
Text Rationalization for Robust Causal Effect Estimation
por: Zhang, Lijinghua, et al.
Publicado: (2025)
por: Zhang, Lijinghua, et al.
Publicado: (2025)
A Causal Lens for Evaluating Faithfulness Metrics
por: Zaman, Kerem, et al.
Publicado: (2025)
por: Zaman, Kerem, et al.
Publicado: (2025)
RCT Rejection Sampling for Causal Estimation Evaluation
por: Keith, Katherine A., et al.
Publicado: (2023)
por: Keith, Katherine A., et al.
Publicado: (2023)
From Ground Truth to Measurement: A Statistical Framework for Human Labeling
por: Chew, Robert, et al.
Publicado: (2026)
por: Chew, Robert, et al.
Publicado: (2026)
The Leaderboard Illusion
por: Singh, Shivalika, et al.
Publicado: (2025)
por: Singh, Shivalika, et al.
Publicado: (2025)
Majority of the Bests: Improving Best-of-N via Bootstrapping
por: Rakhsha, Amin, et al.
Publicado: (2025)
por: Rakhsha, Amin, et al.
Publicado: (2025)
Embedding Trust: Semantic Isotropy Predicts Nonfactuality in Long-Form Text Generation
por: Bhardwaj, Dhrupad, et al.
Publicado: (2025)
por: Bhardwaj, Dhrupad, et al.
Publicado: (2025)
(Mis)Fitting: A Survey of Scaling Laws
por: Li, Margaret, et al.
Publicado: (2025)
por: Li, Margaret, et al.
Publicado: (2025)
Enhancing Causal Reasoning in Large Language Models: A Causal Attribution Model for Precision Fine-Tuning
por: Cai, Hengrui, et al.
Publicado: (2023)
por: Cai, Hengrui, et al.
Publicado: (2023)
Language Models as Causal Effect Generators
por: Bynum, Lucius E. J., et al.
Publicado: (2024)
por: Bynum, Lucius E. J., et al.
Publicado: (2024)
LLMs in the Imaginarium: Tool Learning through Simulated Trial and Error
por: Wang, Boshi, et al.
Publicado: (2024)
por: Wang, Boshi, et al.
Publicado: (2024)
A Comparative Study of DSPy Teleprompter Algorithms for Aligning Large Language Models Evaluation Metrics to Human Evaluation
por: Sarmah, Bhaskarjit, et al.
Publicado: (2024)
por: Sarmah, Bhaskarjit, et al.
Publicado: (2024)
Mind Your Tone: Investigating How Prompt Politeness Affects LLM Accuracy (short paper)
por: Dobariya, Om, et al.
Publicado: (2025)
por: Dobariya, Om, et al.
Publicado: (2025)
SEQR: Secure and Efficient QR-based LoRA Routing
por: Fleshman, William, et al.
Publicado: (2025)
por: Fleshman, William, et al.
Publicado: (2025)
LLMs Are Not Intelligent Thinkers: Introducing Mathematical Topic Tree Benchmark for Comprehensive Evaluation of LLMs
por: Davoodi, Arash Gholami, et al.
Publicado: (2024)
por: Davoodi, Arash Gholami, et al.
Publicado: (2024)
EVOLvE: Evaluating and Optimizing LLMs For In-Context Exploration
por: Nie, Allen, et al.
Publicado: (2024)
por: Nie, Allen, et al.
Publicado: (2024)
SAGE: Shaping Anchors for Guided Exploration in RLVR of LLMs
por: Lee, Chanuk, et al.
Publicado: (2026)
por: Lee, Chanuk, et al.
Publicado: (2026)
AdapterSwap: Continuous Training of LLMs with Data Removal and Access-Control Guarantees
por: Fleshman, William, et al.
Publicado: (2024)
por: Fleshman, William, et al.
Publicado: (2024)
How to Train Data-Efficient LLMs
por: Sachdeva, Noveen, et al.
Publicado: (2024)
por: Sachdeva, Noveen, et al.
Publicado: (2024)
Unlocking Reasoning Capabilities in LLMs via Reinforcement Learning Exploration
por: Deng, Wenhao, et al.
Publicado: (2025)
por: Deng, Wenhao, et al.
Publicado: (2025)
Diffusion Generative Flow Samplers: Improving learning signals through partial trajectory optimization
por: Zhang, Dinghuai, et al.
Publicado: (2023)
por: Zhang, Dinghuai, et al.
Publicado: (2023)
LogicTree: Structured Proof Exploration for Coherent and Rigorous Logical Reasoning with Large Language Models
por: He, Kang, et al.
Publicado: (2025)
por: He, Kang, et al.
Publicado: (2025)
Sanity Checks Revisited: An Exploration to Repair the Model Parameter Randomisation Test
por: Hedström, Anna, et al.
Publicado: (2024)
por: Hedström, Anna, et al.
Publicado: (2024)
LETS-C: Leveraging Text Embedding for Time Series Classification
por: Kaur, Rachneet, et al.
Publicado: (2024)
por: Kaur, Rachneet, et al.
Publicado: (2024)
Knapsack RL: Unlocking Exploration of LLMs via Optimizing Budget Allocation
por: Li, Ziniu, et al.
Publicado: (2025)
por: Li, Ziniu, et al.
Publicado: (2025)
SELF-[IN]CORRECT: LLMs Struggle with Discriminating Self-Generated Responses
por: Jiang, Dongwei, et al.
Publicado: (2024)
por: Jiang, Dongwei, et al.
Publicado: (2024)
Sample-Efficient Online Learning in LM Agents via Hindsight Trajectory Rewriting
por: Hu, Michael Y., et al.
Publicado: (2025)
por: Hu, Michael Y., et al.
Publicado: (2025)
Causal Reasoning and Large Language Models: Opening a New Frontier for Causality
por: Kıcıman, Emre, et al.
Publicado: (2023)
por: Kıcıman, Emre, et al.
Publicado: (2023)
Sample-Efficient Alignment for LLMs
por: Liu, Zichen, et al.
Publicado: (2024)
por: Liu, Zichen, et al.
Publicado: (2024)
Less is More: Extreme Gradient Boost Rank-1 Adaption for Efficient Finetuning of LLMs
por: Zhang, Yifei, et al.
Publicado: (2024)
por: Zhang, Yifei, et al.
Publicado: (2024)
Efficient multi-prompt evaluation of LLMs
por: Polo, Felipe Maia, et al.
Publicado: (2024)
por: Polo, Felipe Maia, et al.
Publicado: (2024)
Ejemplares similares
-
Efficient Exploration at Scale
por: Asghari, Seyed Mohammad, et al.
Publicado: (2026) -
Removing Spurious Correlation from Neural Network Interpretations
por: Fotouhi, Milad, et al.
Publicado: (2024) -
Evaluating Interventional Reasoning Capabilities of Large Language Models
por: Kasetty, Tejas, et al.
Publicado: (2024) -
AutoEval Done Right: Using Synthetic Data for Model Evaluation
por: Boyeau, Pierre, et al.
Publicado: (2024) -
ALCM: Autonomous LLM-Augmented Causal Discovery Framework
por: Khatibi, Elahe, et al.
Publicado: (2024)