Smaller Models, Smarter Rewards: A Two-Sided Approach to Process and Outcome Rewards
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Groeneveld, Jan Niklas, Qin, Xi, Schaefer, Alexander, Oren, Yaad |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Learning Software Bug Reports: A Systematic Literature Review
von: Long, Guoming, et al.
Veröffentlicht: (2025)
von: Long, Guoming, et al.
Veröffentlicht: (2025)
A Framework for Testing and Adapting REST APIs as LLM Tools
von: Bandlamudi, Jayachandu, et al.
Veröffentlicht: (2025)
von: Bandlamudi, Jayachandu, et al.
Veröffentlicht: (2025)
Finetuning LLMs for Automatic Form Interaction on Web-Browser in Selenium Testing Framework
von: Le, Nguyen-Khang, et al.
Veröffentlicht: (2025)
von: Le, Nguyen-Khang, et al.
Veröffentlicht: (2025)
Mind the GAP: Text Safety Does Not Transfer to Tool-Call Safety in LLM Agents
von: Cartagena, Arnold, et al.
Veröffentlicht: (2026)
von: Cartagena, Arnold, et al.
Veröffentlicht: (2026)
Comprehensive Evaluation and Insights into the Use of Large Language Models in the Automation of Behavior-Driven Development Acceptance Test Formulation
von: Karpurapu, Shanthi, et al.
Veröffentlicht: (2024)
von: Karpurapu, Shanthi, et al.
Veröffentlicht: (2024)
A Systematic Approach for Assessing Large Language Models' Test Case Generation Capability
von: Chang, Hung-Fu, et al.
Veröffentlicht: (2025)
von: Chang, Hung-Fu, et al.
Veröffentlicht: (2025)
Mechanistic Understanding of Language Models in Syntactic Code Completion
von: Miller, Samuel, et al.
Veröffentlicht: (2025)
von: Miller, Samuel, et al.
Veröffentlicht: (2025)
Benchmarking Energy Efficiency of Large Language Models Using vLLM
von: Pronk, K., et al.
Veröffentlicht: (2025)
von: Pronk, K., et al.
Veröffentlicht: (2025)
Vibe Code Bench: Evaluating AI Models on End-to-End Web Application Development
von: Tran, Hung, et al.
Veröffentlicht: (2026)
von: Tran, Hung, et al.
Veröffentlicht: (2026)
CIDR: A Large-Scale Industrial Source Code Dataset for Software Engineering Research
von: Savenkov, Vladislav
Veröffentlicht: (2026)
von: Savenkov, Vladislav
Veröffentlicht: (2026)
Generative AI Toolkit -- a framework for increasing the quality of LLM-based applications over their whole life cycle
von: Kohl, Jens, et al.
Veröffentlicht: (2024)
von: Kohl, Jens, et al.
Veröffentlicht: (2024)
ContractBench: Can LLM Agents Preserve Observation Contracts?
von: Wang, Jicheng, et al.
Veröffentlicht: (2026)
von: Wang, Jicheng, et al.
Veröffentlicht: (2026)
FREYR: A Framework for Recognizing and Executing Your Requests
von: Gallotta, Roberto, et al.
Veröffentlicht: (2025)
von: Gallotta, Roberto, et al.
Veröffentlicht: (2025)
Failure by Interference: Language Models Make Balanced Parentheses Errors When Faulty Mechanisms Overshadow Sound Ones
von: Rai, Daking, et al.
Veröffentlicht: (2025)
von: Rai, Daking, et al.
Veröffentlicht: (2025)
Automated Web Application Testing: End-to-End Test Case Generation with Large Language Models and Screen Transition Graphs
von: Le, Nguyen-Khang, et al.
Veröffentlicht: (2025)
von: Le, Nguyen-Khang, et al.
Veröffentlicht: (2025)
A Serverless Architecture for Real-Time Stock Analysis using Large Language Models: An Iterative Development and Debugging Case Study
von: Ashraf, Taniv
Veröffentlicht: (2025)
von: Ashraf, Taniv
Veröffentlicht: (2025)
The Impact of Large Language Models on Open-source Innovation: Evidence from GitHub Copilot
von: Yeverechyahu, Doron, et al.
Veröffentlicht: (2024)
von: Yeverechyahu, Doron, et al.
Veröffentlicht: (2024)
PRPO: Aligning Process Reward with Outcome Reward in Policy Optimization
von: Ding, Ruiyi, et al.
Veröffentlicht: (2026)
von: Ding, Ruiyi, et al.
Veröffentlicht: (2026)
TCProF: Time-Complexity Prediction SSL Framework
von: Hahn, Joonghyuk, et al.
Veröffentlicht: (2025)
von: Hahn, Joonghyuk, et al.
Veröffentlicht: (2025)
MEC$^3$O: Multi-Expert Consensus for Code Time Complexity Prediction
von: Hahn, Joonghyuk, et al.
Veröffentlicht: (2025)
von: Hahn, Joonghyuk, et al.
Veröffentlicht: (2025)
ContractEval: A Benchmark for Evaluating Contract-Satisfying Assertions in Code Generation
von: Lim, Soohan, et al.
Veröffentlicht: (2025)
von: Lim, Soohan, et al.
Veröffentlicht: (2025)
Tool-Schema Compression Enables Agentic RAG Under Constrained Context Budgets
von: Sakizli, Furkan
Veröffentlicht: (2026)
von: Sakizli, Furkan
Veröffentlicht: (2026)
VeriGuard: Enhancing LLM Agent Safety via Verified Code Generation
von: Miculicich, Lesly, et al.
Veröffentlicht: (2025)
von: Miculicich, Lesly, et al.
Veröffentlicht: (2025)
Narrow Transformer: StarCoder-Based Java-LM For Desktop
von: Rathinasamy, Kamalkumar, et al.
Veröffentlicht: (2024)
von: Rathinasamy, Kamalkumar, et al.
Veröffentlicht: (2024)
Fine-Tuning LLMs to Analyze Multiple Dimensions of Code Review: A Maximum Entropy Regulated Long Chain-of-Thought Approach
von: Yu, Yongda, et al.
Veröffentlicht: (2025)
von: Yu, Yongda, et al.
Veröffentlicht: (2025)
Beyond Greenfield: The D3 Framework for AI-Driven Productivity in Brownfield Engineering
von: Sharma, Krishna Kumaar
Veröffentlicht: (2025)
von: Sharma, Krishna Kumaar
Veröffentlicht: (2025)
Leveraging Large Language Models for Use Case Model Generation from Software Requirements
von: Eisenreich, Tobias, et al.
Veröffentlicht: (2025)
von: Eisenreich, Tobias, et al.
Veröffentlicht: (2025)
Distilling Desired Comments for Enhanced Code Review with Large Language Models
von: Yu, Yongda, et al.
Veröffentlicht: (2024)
von: Yu, Yongda, et al.
Veröffentlicht: (2024)
Assessing Data Augmentation-Induced Bias in Training and Testing of Machine Learning Models
von: More, Riddhi, et al.
Veröffentlicht: (2025)
von: More, Riddhi, et al.
Veröffentlicht: (2025)
Software Defined Vehicle Code Generation: A Few-Shot Prompting Approach
von: Nguyen, Quang-Dung, et al.
Veröffentlicht: (2025)
von: Nguyen, Quang-Dung, et al.
Veröffentlicht: (2025)
CWM: An Open-Weights LLM for Research on Code Generation with World Models
von: FAIR CodeGen team, et al.
Veröffentlicht: (2025)
von: FAIR CodeGen team, et al.
Veröffentlicht: (2025)
From Untestable to Testable: Metamorphic Testing in the Age of LLMs
von: Terragni, Valerio
Veröffentlicht: (2026)
von: Terragni, Valerio
Veröffentlicht: (2026)
Automating Domain-Driven Design: Experience with a Prompting Framework
von: Eisenreich, Tobias, et al.
Veröffentlicht: (2026)
von: Eisenreich, Tobias, et al.
Veröffentlicht: (2026)
An Analysis of LLM Fine-Tuning and Few-Shot Learning for Flaky Test Detection and Classification
von: More, Riddhi, et al.
Veröffentlicht: (2025)
von: More, Riddhi, et al.
Veröffentlicht: (2025)
Achieving Tool Calling Functionality in LLMs Using Only Prompt Engineering Without Fine-Tuning
von: He, Shengtao
Veröffentlicht: (2024)
von: He, Shengtao
Veröffentlicht: (2024)
Prompt Engineering Strategies for LLM-based Qualitative Coding of Psychological Safety in Software Engineering Communities: A Controlled Empirical Study
von: Alshaikh, Moaath, et al.
Veröffentlicht: (2026)
von: Alshaikh, Moaath, et al.
Veröffentlicht: (2026)
REPOT: Recoverable Program-of-Thought via Checkpoint Repair
von: Mazaheri, Parsa
Veröffentlicht: (2026)
von: Mazaheri, Parsa
Veröffentlicht: (2026)
Plan with Code: Comparing approaches for robust NL to DSL generation
von: Bassamzadeh, Nastaran, et al.
Veröffentlicht: (2024)
von: Bassamzadeh, Nastaran, et al.
Veröffentlicht: (2024)
A Comparative Study of DSL Code Generation: Fine-Tuning vs. Optimized Retrieval Augmentation
von: Bassamzadeh, Nastaran, et al.
Veröffentlicht: (2024)
von: Bassamzadeh, Nastaran, et al.
Veröffentlicht: (2024)
Automated Bug Triaging using Instruction-Tuned Large Language Models
von: Kiashemshaki, Kiana, et al.
Veröffentlicht: (2025)
von: Kiashemshaki, Kiana, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Learning Software Bug Reports: A Systematic Literature Review
von: Long, Guoming, et al.
Veröffentlicht: (2025) -
A Framework for Testing and Adapting REST APIs as LLM Tools
von: Bandlamudi, Jayachandu, et al.
Veröffentlicht: (2025) -
Finetuning LLMs for Automatic Form Interaction on Web-Browser in Selenium Testing Framework
von: Le, Nguyen-Khang, et al.
Veröffentlicht: (2025) -
Mind the GAP: Text Safety Does Not Transfer to Tool-Call Safety in LLM Agents
von: Cartagena, Arnold, et al.
Veröffentlicht: (2026) -
Comprehensive Evaluation and Insights into the Use of Large Language Models in the Automation of Behavior-Driven Development Acceptance Test Formulation
von: Karpurapu, Shanthi, et al.
Veröffentlicht: (2024)