LLM Library Learning Fails: A LEGO-Prover Case Study
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Berlot-Attwell, Ian, Rudzicz, Frank, Si, Xujie |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Library Learning Doesn't: The Curious Case of the Single-Use "Library"
von: Berlot-Attwell, Ian, et al.
Veröffentlicht: (2024)
von: Berlot-Attwell, Ian, et al.
Veröffentlicht: (2024)
Attribute Diversity Determines the Systematicity Gap in VQA
von: Berlot-Attwell, Ian, et al.
Veröffentlicht: (2023)
von: Berlot-Attwell, Ian, et al.
Veröffentlicht: (2023)
Chronosymbolic Learning: Efficient CHC Solving with Symbolic Reasoning and Inductive Learning
von: Luo, Ziyan, et al.
Veröffentlicht: (2023)
von: Luo, Ziyan, et al.
Veröffentlicht: (2023)
LEGO: Language Model Building Blocks
von: Bhansali, Shrenik, et al.
Veröffentlicht: (2024)
von: Bhansali, Shrenik, et al.
Veröffentlicht: (2024)
CASSANDRA: Programmatic and Probabilistic Learning and Inference for Stochastic World Modeling
von: Lymperopoulos, Panagiotis, et al.
Veröffentlicht: (2026)
von: Lymperopoulos, Panagiotis, et al.
Veröffentlicht: (2026)
REAL-Prover: Retrieval Augmented Lean Prover for Mathematical Reasoning
von: Shen, Ziju, et al.
Veröffentlicht: (2025)
von: Shen, Ziju, et al.
Veröffentlicht: (2025)
Plug and Play with Prompts: A Prompt Tuning Approach for Controlling Text Generation
von: Ajwani, Rohan Deepak, et al.
Veröffentlicht: (2024)
von: Ajwani, Rohan Deepak, et al.
Veröffentlicht: (2024)
Detection Is Cheap, Routing Is Learned: Why Refusal-Based Alignment Evaluation Fails
von: Frank, Gregory N.
Veröffentlicht: (2026)
von: Frank, Gregory N.
Veröffentlicht: (2026)
A Compute-Matched Re-Evaluation of TroVE on MATH
von: Sesterhenn, Tobias, et al.
Veröffentlicht: (2025)
von: Sesterhenn, Tobias, et al.
Veröffentlicht: (2025)
Learning Minimal Neural Specifications
von: Geng, Chuqin, et al.
Veröffentlicht: (2024)
von: Geng, Chuqin, et al.
Veröffentlicht: (2024)
Evaluating Defences against Unsafe Feedback in RLHF
von: Rosati, Domenic, et al.
Veröffentlicht: (2024)
von: Rosati, Domenic, et al.
Veröffentlicht: (2024)
Representation Noising: A Defence Mechanism Against Harmful Finetuning
von: Rosati, Domenic, et al.
Veröffentlicht: (2024)
von: Rosati, Domenic, et al.
Veröffentlicht: (2024)
PutnamBench: Evaluating Neural Theorem-Provers on the Putnam Mathematical Competition
von: Tsoukalas, George, et al.
Veröffentlicht: (2024)
von: Tsoukalas, George, et al.
Veröffentlicht: (2024)
Single-Position Intervention Fails: Distributed Output Templates Drive In-Context Learning
von: Cheng, Bryan, et al.
Veröffentlicht: (2026)
von: Cheng, Bryan, et al.
Veröffentlicht: (2026)
ScenicProver: A Framework for Compositional Probabilistic Verification of Learning-Enabled Systems
von: Vin, Eric, et al.
Veröffentlicht: (2025)
von: Vin, Eric, et al.
Veröffentlicht: (2025)
When Refusals Fail: Unstable Safety Mechanisms in Long-Context LLM Agents
von: Hadeliya, Tsimur, et al.
Veröffentlicht: (2025)
von: Hadeliya, Tsimur, et al.
Veröffentlicht: (2025)
When LLM Judge Scores Look Good but Best-of-N Decisions Fail
von: Landesberg, Eddie
Veröffentlicht: (2026)
von: Landesberg, Eddie
Veröffentlicht: (2026)
APPL: A Prompt Programming Language for Harmonious Integration of Programs and Large Language Model Prompts
von: Dong, Honghua, et al.
Veröffentlicht: (2024)
von: Dong, Honghua, et al.
Veröffentlicht: (2024)
The Illusion of Certainty: Uncertainty Quantification for LLMs Fails under Ambiguity
von: Tomov, Tim, et al.
Veröffentlicht: (2025)
von: Tomov, Tim, et al.
Veröffentlicht: (2025)
Failing to Falsify: Evaluating and Mitigating Confirmation Bias in Language Models
von: Jhaveri, Ayush Rajesh, et al.
Veröffentlicht: (2026)
von: Jhaveri, Ayush Rajesh, et al.
Veröffentlicht: (2026)
Why Reasoning Fails to Plan: A Planning-Centric Analysis of Long-Horizon Decision Making in LLM Agents
von: Wang, Zehong, et al.
Veröffentlicht: (2026)
von: Wang, Zehong, et al.
Veröffentlicht: (2026)
ImProver 2: Iteratively Self-Improving LMs for Neurosymbolic Proof Optimization
von: Ahuja, Riyaz, et al.
Veröffentlicht: (2026)
von: Ahuja, Riyaz, et al.
Veröffentlicht: (2026)
Duo-LLM: A Framework for Studying Adaptive Computation in Large Language Models
von: Alizadeh, Keivan, et al.
Veröffentlicht: (2024)
von: Alizadeh, Keivan, et al.
Veröffentlicht: (2024)
Improving Diversity in Language Models: When Temperature Fails, Change the Loss
von: Verine, Alexandre, et al.
Veröffentlicht: (2025)
von: Verine, Alexandre, et al.
Veröffentlicht: (2025)
Two-stage LLM Fine-tuning with Less Specialization and More Generalization
von: Wang, Yihan, et al.
Veröffentlicht: (2022)
von: Wang, Yihan, et al.
Veröffentlicht: (2022)
Compile to Compress: Boosting Formal Theorem Provers by Compiler Outputs
von: Li, Guchan, et al.
Veröffentlicht: (2026)
von: Li, Guchan, et al.
Veröffentlicht: (2026)
ACCORD: Closing the Commonsense Measurability Gap
von: Roewer-Després, François, et al.
Veröffentlicht: (2024)
von: Roewer-Després, François, et al.
Veröffentlicht: (2024)
Dense Retrievers Can Fail on Simple Queries: Revealing The Granularity Dilemma of Embeddings
von: Xu, Liyan, et al.
Veröffentlicht: (2025)
von: Xu, Liyan, et al.
Veröffentlicht: (2025)
Why Semantic Entropy Fails: Geometry-Aware and Calibrated Uncertainty for Policy Optimization
von: Zhang, Zheyuan, et al.
Veröffentlicht: (2026)
von: Zhang, Zheyuan, et al.
Veröffentlicht: (2026)
Merging LoRAs like Playing LEGO: Pushing the Modularity of LoRA to Extremes Through Rank-Wise Clustering
von: Zhao, Ziyu, et al.
Veröffentlicht: (2024)
von: Zhao, Ziyu, et al.
Veröffentlicht: (2024)
Transfer Learning via Lexical Relatedness: A Sarcasm and Hate Speech Case Study
von: Cabrera, Angelly, et al.
Veröffentlicht: (2025)
von: Cabrera, Angelly, et al.
Veröffentlicht: (2025)
Enhancing LLM Factual Accuracy with RAG to Counter Hallucinations: A Case Study on Domain-Specific Queries in Private Knowledge-Bases
von: Li, Jiarui, et al.
Veröffentlicht: (2024)
von: Li, Jiarui, et al.
Veröffentlicht: (2024)
DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Search
von: Xin, Huajian, et al.
Veröffentlicht: (2024)
von: Xin, Huajian, et al.
Veröffentlicht: (2024)
Initialisation Determines the Basin: Efficient Codebook Optimisation for Extreme LLM Quantization
von: Kennedy, Ian W., et al.
Veröffentlicht: (2026)
von: Kennedy, Ian W., et al.
Veröffentlicht: (2026)
Predicting Rental Price of Lane Houses in Shanghai with Machine Learning Methods and Large Language Models
von: Chen, Tingting, et al.
Veröffentlicht: (2024)
von: Chen, Tingting, et al.
Veröffentlicht: (2024)
CORAL: Learning Consistent Representations across Multi-step Training with Lighter Speculative Drafter
von: Weng, Yepeng, et al.
Veröffentlicht: (2025)
von: Weng, Yepeng, et al.
Veröffentlicht: (2025)
torchdistill Meets Hugging Face Libraries for Reproducible, Coding-Free Deep Learning Studies: A Case Study on NLP
von: Matsubara, Yoshitomo
Veröffentlicht: (2023)
von: Matsubara, Yoshitomo
Veröffentlicht: (2023)
When In-Distribution Gains Fail: Evaluating Weak-to-Strong Reward Models under Preference Shift
von: Le, Khoi, et al.
Veröffentlicht: (2026)
von: Le, Khoi, et al.
Veröffentlicht: (2026)
Selective Fine-tuning on LLM-labeled Data May Reduce Reliance on Human Annotation: A Case Study Using Schedule-of-Event Table Detection
von: Kumar, Bhawesh, et al.
Veröffentlicht: (2024)
von: Kumar, Bhawesh, et al.
Veröffentlicht: (2024)
Why Bonds Fail Differently? Explainable Multimodal Learning for Multi-Class Default Prediction
von: Lu, Yi, et al.
Veröffentlicht: (2025)
von: Lu, Yi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Library Learning Doesn't: The Curious Case of the Single-Use "Library"
von: Berlot-Attwell, Ian, et al.
Veröffentlicht: (2024) -
Attribute Diversity Determines the Systematicity Gap in VQA
von: Berlot-Attwell, Ian, et al.
Veröffentlicht: (2023) -
Chronosymbolic Learning: Efficient CHC Solving with Symbolic Reasoning and Inductive Learning
von: Luo, Ziyan, et al.
Veröffentlicht: (2023) -
LEGO: Language Model Building Blocks
von: Bhansali, Shrenik, et al.
Veröffentlicht: (2024) -
CASSANDRA: Programmatic and Probabilistic Learning and Inference for Stochastic World Modeling
von: Lymperopoulos, Panagiotis, et al.
Veröffentlicht: (2026)