Increasing LLM Coding Capabilities through Diverse Synthetic Coding Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Abed, Amal, Lukic, Ivan, Franke, Jörg K. H., Hutter, Frank |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Balancing Synthetic Data and Replay for Enhancing Task-Specific Capabilities
by: Spiegelhalter, Urs, et al.
Published: (2025)
by: Spiegelhalter, Urs, et al.
Published: (2025)
Fast Optimizer Benchmark
by: Blauth, Simon, et al.
Published: (2024)
by: Blauth, Simon, et al.
Published: (2024)
KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding
by: Xu, Zhangchen, et al.
Published: (2025)
by: Xu, Zhangchen, et al.
Published: (2025)
Applied Machine Learning Methods with Long-Short Term Memory Based Recurrent Neural Networks for Multivariate Temperature Prediction
by: Lukić, Bojan
Published: (2025)
by: Lukić, Bojan
Published: (2025)
HW-GPT-Bench: Hardware-Aware Architecture Benchmark for Language Models
by: Sukthanker, Rhea Sanjay, et al.
Published: (2024)
by: Sukthanker, Rhea Sanjay, et al.
Published: (2024)
One-shot World Models Using a Transformer Trained on a Synthetic Prior
by: Ferreira, Fabio, et al.
Published: (2024)
by: Ferreira, Fabio, et al.
Published: (2024)
TempoPFN: Synthetic Pre-training of Linear RNNs for Zero-shot Time Series Forecasting
by: Moroshan, Vladyslav, et al.
Published: (2025)
by: Moroshan, Vladyslav, et al.
Published: (2025)
Speeding Up Multi-Objective Hyperparameter Optimization by Task Similarity-Based Meta-Learning for the Tree-Structured Parzen Estimator
by: Watanabe, Shuhei, et al.
Published: (2022)
by: Watanabe, Shuhei, et al.
Published: (2022)
c-TPE: Tree-structured Parzen Estimator with Inequality Constraints for Expensive Hyperparameter Optimization
by: Watanabe, Shuhei, et al.
Published: (2022)
by: Watanabe, Shuhei, et al.
Published: (2022)
Tuning LLM Judge Design Decisions for 1/1000 of the Cost
by: Salinas, David, et al.
Published: (2025)
by: Salinas, David, et al.
Published: (2025)
CodeScaler: Scaling Code LLM Training and Test-Time Inference via Reward Models
by: Zhu, Xiao, et al.
Published: (2026)
by: Zhu, Xiao, et al.
Published: (2026)
Beyond Synthetic Benchmarks: Evaluating LLM Performance on Real-World Class-Level Code Generation
by: Rahman, Musfiqur, et al.
Published: (2025)
by: Rahman, Musfiqur, et al.
Published: (2025)
Enhancing LLM Planning Capabilities through Intrinsic Self-Critique
by: Bohnet, Bernd, et al.
Published: (2025)
by: Bohnet, Bernd, et al.
Published: (2025)
CodeARC: Benchmarking Reasoning Capabilities of LLM Agents for Inductive Program Synthesis
by: Wei, Anjiang, et al.
Published: (2025)
by: Wei, Anjiang, et al.
Published: (2025)
[Re] Benchmarking LLM Capabilities in Negotiation through Scoreable Games
by: Pollo, Jorge Carrasco, et al.
Published: (2026)
by: Pollo, Jorge Carrasco, et al.
Published: (2026)
A Deep Dive into Scaling RL for Code Generation with Synthetic Data and Curricula
by: Sancaktar, Cansu, et al.
Published: (2026)
by: Sancaktar, Cansu, et al.
Published: (2026)
CodeEditorBench: Evaluating Code Editing Capability of Large Language Models
by: Guo, Jiawei, et al.
Published: (2024)
by: Guo, Jiawei, et al.
Published: (2024)
Improving Deep Learning Optimization through Constrained Parameter Regularization
by: Franke, Jörg K. H., et al.
Published: (2023)
by: Franke, Jörg K. H., et al.
Published: (2023)
One Model, Many Skills: Parameter-Efficient Fine-Tuning for Multitask Code Analysis
by: Akli, Amal, et al.
Published: (2026)
by: Akli, Amal, et al.
Published: (2026)
ICQuant: Index Coding enables Low-bit LLM Quantization
by: Li, Xinlin, et al.
Published: (2025)
by: Li, Xinlin, et al.
Published: (2025)
Code Simulation as a Proxy for High-order Tasks in Large Language Models
by: La Malfa, Emanuele, et al.
Published: (2025)
by: La Malfa, Emanuele, et al.
Published: (2025)
Unveiling the Magic of Code Reasoning through Hypothesis Decomposition and Amendment
by: Zhao, Yuze, et al.
Published: (2025)
by: Zhao, Yuze, et al.
Published: (2025)
Relative Positioning Based Code Chunking Method For Rich Context Retrieval In Repository Level Code Completion Task With Code Language Model
by: Rahman, Imranur, et al.
Published: (2025)
by: Rahman, Imranur, et al.
Published: (2025)
EquiTabPFN: A Target-Permutation Equivariant Prior Fitted Networks
by: Arbel, Michael, et al.
Published: (2025)
by: Arbel, Michael, et al.
Published: (2025)
Agentic NL2SQL to Reduce Computational Costs
by: Jehle, Dominik, et al.
Published: (2025)
by: Jehle, Dominik, et al.
Published: (2025)
Efficient Search for Customized Activation Functions with Gradient Descent
by: Strack, Lukas, et al.
Published: (2024)
by: Strack, Lukas, et al.
Published: (2024)
Large Language Models Engineer Too Many Simple Features For Tabular Data
by: Küken, Jaris, et al.
Published: (2024)
by: Küken, Jaris, et al.
Published: (2024)
Don't Waste Your Time: Early Stopping Cross-Validation
by: Bergman, Edward, et al.
Published: (2024)
by: Bergman, Edward, et al.
Published: (2024)
Rewriting Pre-Training Data Boosts LLM Performance in Math and Code
by: Fujii, Kazuki, et al.
Published: (2025)
by: Fujii, Kazuki, et al.
Published: (2025)
CodeRefine: A Pipeline for Enhancing LLM-Generated Code Implementations of Research Papers
by: Trofimova, Ekaterina, et al.
Published: (2024)
by: Trofimova, Ekaterina, et al.
Published: (2024)
Improving LLM-based Global Optimization with Search Space Partitioning
by: Schwanke, Andrej, et al.
Published: (2025)
by: Schwanke, Andrej, et al.
Published: (2025)
ComplexityNet: Increasing LLM Inference Efficiency by Learning Task Complexity
by: Bae, Henry, et al.
Published: (2023)
by: Bae, Henry, et al.
Published: (2023)
I Know Which LLM Wrote Your Code Last Summer: LLM generated Code Stylometry for Authorship Attribution
by: Bisztray, Tamas, et al.
Published: (2025)
by: Bisztray, Tamas, et al.
Published: (2025)
Generalizable End-to-End Tool-Use RL with Synthetic CodeGym
by: Du, Weihua, et al.
Published: (2025)
by: Du, Weihua, et al.
Published: (2025)
Classical Planning with LLM-Generated Heuristics: Challenging the State of the Art with Python Code
by: Corrêa, Augusto B., et al.
Published: (2025)
by: Corrêa, Augusto B., et al.
Published: (2025)
VERIRL: Boosting the LLM-based Verilog Code Generation via Reinforcement Learning
by: Teng, Fu, et al.
Published: (2025)
by: Teng, Fu, et al.
Published: (2025)
Taxon: Hierarchical Tax Code Prediction with Semantically Aligned LLM Expert Guidance
by: Li, Jihang, et al.
Published: (2026)
by: Li, Jihang, et al.
Published: (2026)
ChatGPT Code Detection: Techniques for Uncovering the Source of Code
by: Oedingen, Marc, et al.
Published: (2024)
by: Oedingen, Marc, et al.
Published: (2024)
DreamPRM-Code: Function-as-Step Process Reward Model with Label Correction for LLM Coding
by: Zhang, Ruiyi, et al.
Published: (2025)
by: Zhang, Ruiyi, et al.
Published: (2025)
On Mitigating Code LLM Hallucinations with API Documentation
by: Jain, Nihal, et al.
Published: (2024)
by: Jain, Nihal, et al.
Published: (2024)
Similar Items
-
Balancing Synthetic Data and Replay for Enhancing Task-Specific Capabilities
by: Spiegelhalter, Urs, et al.
Published: (2025) -
Fast Optimizer Benchmark
by: Blauth, Simon, et al.
Published: (2024) -
KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding
by: Xu, Zhangchen, et al.
Published: (2025) -
Applied Machine Learning Methods with Long-Short Term Memory Based Recurrent Neural Networks for Multivariate Temperature Prediction
by: Lukić, Bojan
Published: (2025) -
HW-GPT-Bench: Hardware-Aware Architecture Benchmark for Language Models
by: Sukthanker, Rhea Sanjay, et al.
Published: (2024)