Disentangling Exploration of Large Language Models by Optimal Exploitation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Grams, Tim, Betz, Patrick, Marton, Sascha, Lüdtke, Stefan, Bartelt, Christian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
GradTree: Learning Axis-Aligned Decision Trees with Gradient Descent
von: Marton, Sascha, et al.
Veröffentlicht: (2023)
von: Marton, Sascha, et al.
Veröffentlicht: (2023)
A Data-Centric Perspective on Evaluating Machine Learning Models for Tabular Data
von: Tschalzev, Andrej, et al.
Veröffentlicht: (2024)
von: Tschalzev, Andrej, et al.
Veröffentlicht: (2024)
Explaining Neural Networks without Access to Training Data
von: Marton, Sascha, et al.
Veröffentlicht: (2022)
von: Marton, Sascha, et al.
Veröffentlicht: (2022)
Mitigating Information Loss in Tree-Based Reinforcement Learning via Direct Optimization
von: Marton, Sascha, et al.
Veröffentlicht: (2024)
von: Marton, Sascha, et al.
Veröffentlicht: (2024)
Which LIME should I trust? Concepts, Challenges, and Solutions
von: Knab, Patrick, et al.
Veröffentlicht: (2025)
von: Knab, Patrick, et al.
Veröffentlicht: (2025)
Interpreting Outliers in Time Series Data through Decoding Autoencoder
von: Knab, Patrick, et al.
Veröffentlicht: (2024)
von: Knab, Patrick, et al.
Veröffentlicht: (2024)
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
von: Chen, Zhipeng, et al.
Veröffentlicht: (2025)
von: Chen, Zhipeng, et al.
Veröffentlicht: (2025)
GRANDE: Gradient-Based Decision Tree Ensembles for Tabular Data
von: Marton, Sascha, et al.
Veröffentlicht: (2023)
von: Marton, Sascha, et al.
Veröffentlicht: (2023)
Beyond Pixels: Enhancing LIME with Hierarchical Features and Segmentation Foundation Models
von: Knab, Patrick, et al.
Veröffentlicht: (2024)
von: Knab, Patrick, et al.
Veröffentlicht: (2024)
Exploration vs Exploitation: Rethinking RLVR through Clipping, Entropy, and Spurious Reward
von: Chen, Peter, et al.
Veröffentlicht: (2025)
von: Chen, Peter, et al.
Veröffentlicht: (2025)
$ϕ$-Decoding: Adaptive Foresight Sampling for Balanced Inference-Time Exploration and Exploitation
von: Xu, Fangzhi, et al.
Veröffentlicht: (2025)
von: Xu, Fangzhi, et al.
Veröffentlicht: (2025)
B-STaR: Monitoring and Balancing Exploration and Exploitation in Self-Taught Reasoners
von: Zeng, Weihao, et al.
Veröffentlicht: (2024)
von: Zeng, Weihao, et al.
Veröffentlicht: (2024)
Fine-tuning Large Language Models with Limited Data: A Survey and Practical Guide
von: Szep, Marton, et al.
Veröffentlicht: (2024)
von: Szep, Marton, et al.
Veröffentlicht: (2024)
Disentangling Logic: The Role of Context in Large Language Model Reasoning Capabilities
von: Hua, Wenyue, et al.
Veröffentlicht: (2024)
von: Hua, Wenyue, et al.
Veröffentlicht: (2024)
Understanding and Mitigating Tokenization Bias in Language Models
von: Phan, Buu, et al.
Veröffentlicht: (2024)
von: Phan, Buu, et al.
Veröffentlicht: (2024)
Training Optimal Large Diffusion Language Models
von: Ni, Jinjie, et al.
Veröffentlicht: (2025)
von: Ni, Jinjie, et al.
Veröffentlicht: (2025)
A Scalable Multi-LLM Collaboration System with Retrieval-based Selection and Exploration-Exploitation-Driven Enhancement
von: Tang, Shengji, et al.
Veröffentlicht: (2025)
von: Tang, Shengji, et al.
Veröffentlicht: (2025)
HTPO: Towards Exploration-Exploitation Balanced Policy Optimization via Hierarchical Token-level Objective Control
von: Yao, Xincheng, et al.
Veröffentlicht: (2026)
von: Yao, Xincheng, et al.
Veröffentlicht: (2026)
Large Language Models Are Overparameterized Text Encoders
von: K, Thennal D, et al.
Veröffentlicht: (2024)
von: K, Thennal D, et al.
Veröffentlicht: (2024)
CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models
von: Dai, Runpeng, et al.
Veröffentlicht: (2025)
von: Dai, Runpeng, et al.
Veröffentlicht: (2025)
Latent Concept Disentanglement in Transformer-based Language Models
von: Hong, Guan Zhe, et al.
Veröffentlicht: (2025)
von: Hong, Guan Zhe, et al.
Veröffentlicht: (2025)
Explorations of Self-Repair in Language Models
von: Rushing, Cody, et al.
Veröffentlicht: (2024)
von: Rushing, Cody, et al.
Veröffentlicht: (2024)
LogicTree: Structured Proof Exploration for Coherent and Rigorous Logical Reasoning with Large Language Models
von: He, Kang, et al.
Veröffentlicht: (2025)
von: He, Kang, et al.
Veröffentlicht: (2025)
Do Large Language Models Truly Grasp Mathematics? An Empirical Exploration From Cognitive Psychology
von: Xie, Wei, et al.
Veröffentlicht: (2024)
von: Xie, Wei, et al.
Veröffentlicht: (2024)
Auto-Intent: Automated Intent Discovery and Self-Exploration for Large Language Model Web Agents
von: Kim, Jaekyeom, et al.
Veröffentlicht: (2024)
von: Kim, Jaekyeom, et al.
Veröffentlicht: (2024)
Beyond Either-Or Reasoning: Transduction and Induction as Cooperative Problem-Solving Paradigms
von: Zenkner, Janis, et al.
Veröffentlicht: (2025)
von: Zenkner, Janis, et al.
Veröffentlicht: (2025)
Improving Reasoning Performance in Large Language Models via Representation Engineering
von: Højer, Bertram, et al.
Veröffentlicht: (2025)
von: Højer, Bertram, et al.
Veröffentlicht: (2025)
Semi-Supervised Learning for Large Language Models Safety and Content Moderation
von: Dinuta, Eduard Stefan, et al.
Veröffentlicht: (2025)
von: Dinuta, Eduard Stefan, et al.
Veröffentlicht: (2025)
Alzheimer's Dementia Detection Using Perplexity from Paired Large Language Models
von: Xiao, Yao, et al.
Veröffentlicht: (2025)
von: Xiao, Yao, et al.
Veröffentlicht: (2025)
Towards Automated Regulatory Compliance Verification in Financial Auditing with Large Language Models
von: Berger, Armin, et al.
Veröffentlicht: (2025)
von: Berger, Armin, et al.
Veröffentlicht: (2025)
Your Finetuned Large Language Model is Already a Powerful Out-of-distribution Detector
von: Zhang, Andi, et al.
Veröffentlicht: (2024)
von: Zhang, Andi, et al.
Veröffentlicht: (2024)
Comparing Exploration-Exploitation Strategies of LLMs and Humans: Insights from Standard Multi-armed Bandit Experiments
von: Zhang, Ziyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Ziyuan, et al.
Veröffentlicht: (2025)
Unintended Memorization of Sensitive Information in Fine-Tuned Language Models
von: Szep, Marton, et al.
Veröffentlicht: (2026)
von: Szep, Marton, et al.
Veröffentlicht: (2026)
Large Language Models are Powerful Electronic Health Record Encoders
von: Hegselmann, Stefan, et al.
Veröffentlicht: (2025)
von: Hegselmann, Stefan, et al.
Veröffentlicht: (2025)
The Human Factor in Detecting Errors of Large Language Models: A Systematic Literature Review and Future Research Directions
von: Schiller, Christian A.
Veröffentlicht: (2024)
von: Schiller, Christian A.
Veröffentlicht: (2024)
DaMoC: Efficiently Selecting the Optimal Large Language Model for Fine-tuning Domain Tasks Based on Data and Model Compression
von: Huang, Wei, et al.
Veröffentlicht: (2025)
von: Huang, Wei, et al.
Veröffentlicht: (2025)
On the Optimal Reasoning Length for RL-Trained Language Models
von: Nohara, Daisuke, et al.
Veröffentlicht: (2026)
von: Nohara, Daisuke, et al.
Veröffentlicht: (2026)
Better Estimation of the Kullback--Leibler Divergence Between Language Models
von: Amini, Afra, et al.
Veröffentlicht: (2025)
von: Amini, Afra, et al.
Veröffentlicht: (2025)
Tracing Uncertainty in Language Model "Reasoning"
von: Grünefeld, Nils, et al.
Veröffentlicht: (2026)
von: Grünefeld, Nils, et al.
Veröffentlicht: (2026)
Syntactic Control of Language Models by Posterior Inference
von: Xefteri, Vicky, et al.
Veröffentlicht: (2025)
von: Xefteri, Vicky, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
GradTree: Learning Axis-Aligned Decision Trees with Gradient Descent
von: Marton, Sascha, et al.
Veröffentlicht: (2023) -
A Data-Centric Perspective on Evaluating Machine Learning Models for Tabular Data
von: Tschalzev, Andrej, et al.
Veröffentlicht: (2024) -
Explaining Neural Networks without Access to Training Data
von: Marton, Sascha, et al.
Veröffentlicht: (2022) -
Mitigating Information Loss in Tree-Based Reinforcement Learning via Direct Optimization
von: Marton, Sascha, et al.
Veröffentlicht: (2024) -
Which LIME should I trust? Concepts, Challenges, and Solutions
von: Knab, Patrick, et al.
Veröffentlicht: (2025)