Themis: Training Robust Multilingual Code Reward Models for Flexible Multi-Criteria Scoring
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Paul, Indraneil, Glavaš, Goran, Gurevych, Iryna |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ObscuraCoder: Powering Efficient Code LM Pre-Training Via Obfuscation Grounding
von: Paul, Indraneil, et al.
Veröffentlicht: (2025)
von: Paul, Indraneil, et al.
Veröffentlicht: (2025)
AICD Bench: A Challenging Benchmark for AI-Generated Code Detection
von: Orel, Daniil, et al.
Veröffentlicht: (2026)
von: Orel, Daniil, et al.
Veröffentlicht: (2026)
IRCoder: Intermediate Representations Make Language Models Robust Multilingual Code Generators
von: Paul, Indraneil, et al.
Veröffentlicht: (2024)
von: Paul, Indraneil, et al.
Veröffentlicht: (2024)
Aletheia: What Makes RLVR For Code Verifiers Tick?
von: Venkatkrishna, Vatsal, et al.
Veröffentlicht: (2026)
von: Venkatkrishna, Vatsal, et al.
Veröffentlicht: (2026)
$\texttt{Droid}$: A Resource Suite for AI-Generated Code Detection
von: Orel, Daniil, et al.
Veröffentlicht: (2025)
von: Orel, Daniil, et al.
Veröffentlicht: (2025)
FunPRM: Function-as-Step Process Reward Model with Meta Reward Correction for Code Generation
von: Zhang, Ruiyi, et al.
Veröffentlicht: (2026)
von: Zhang, Ruiyi, et al.
Veröffentlicht: (2026)
You Only Train Once: A Flexible Training Framework for Code Vulnerability Detection Driven by Vul-Vector
von: Tian, Bowen, et al.
Veröffentlicht: (2025)
von: Tian, Bowen, et al.
Veröffentlicht: (2025)
Leveraging Reward Models for Guiding Code Review Comment Generation
von: Sghaier, Oussama Ben, et al.
Veröffentlicht: (2025)
von: Sghaier, Oussama Ben, et al.
Veröffentlicht: (2025)
Trained Without My Consent: Detecting Code Inclusion In Language Models Trained on Code
von: Majdinasab, Vahid, et al.
Veröffentlicht: (2024)
von: Majdinasab, Vahid, et al.
Veröffentlicht: (2024)
CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-X
von: Zheng, Qinkai, et al.
Veröffentlicht: (2023)
von: Zheng, Qinkai, et al.
Veröffentlicht: (2023)
Understanding Robustness of Model Editing in Code LLMs
von: Chhetri, Vinaik, et al.
Veröffentlicht: (2025)
von: Chhetri, Vinaik, et al.
Veröffentlicht: (2025)
Domain-Adaptable Reinforcement Learning for Code Generation with Dense Rewards
von: Jolfaei, Erfan Aghadavoodi, et al.
Veröffentlicht: (2026)
von: Jolfaei, Erfan Aghadavoodi, et al.
Veröffentlicht: (2026)
Large Language Models for Multilingual Code Intelligence: A Survey
von: Jiang, Chao, et al.
Veröffentlicht: (2026)
von: Jiang, Chao, et al.
Veröffentlicht: (2026)
QiMeng-PRepair: Precise Code Repair via Edit-Aware Reward Optimization
von: Ke, Changxin, et al.
Veröffentlicht: (2026)
von: Ke, Changxin, et al.
Veröffentlicht: (2026)
An Exploratory Investigation into Code License Infringements in Large Language Model Training Datasets
von: Katzy, Jonathan, et al.
Veröffentlicht: (2024)
von: Katzy, Jonathan, et al.
Veröffentlicht: (2024)
Robust Learning of Diverse Code Edits
von: Aggarwal, Tushar, et al.
Veröffentlicht: (2025)
von: Aggarwal, Tushar, et al.
Veröffentlicht: (2025)
Operational Robustness of LLMs on Code Generation
von: Paul, Debalina Ghosh, et al.
Veröffentlicht: (2026)
von: Paul, Debalina Ghosh, et al.
Veröffentlicht: (2026)
AgentForge: A Flexible Low-Code Platform for Reinforcement Learning Agent Design
von: Junior, Francisco Erivaldo Fernandes, et al.
Veröffentlicht: (2024)
von: Junior, Francisco Erivaldo Fernandes, et al.
Veröffentlicht: (2024)
LeetCodeDataset: A Temporal Dataset for Robust Evaluation and Efficient Training of Code LLMs
von: Xia, Yunhui, et al.
Veröffentlicht: (2025)
von: Xia, Yunhui, et al.
Veröffentlicht: (2025)
Model Cascading for Code: A Cascaded Black-Box Multi-Model Framework for Cost-Efficient Code Completion with Self-Testing
von: Chen, Boyuan, et al.
Veröffentlicht: (2024)
von: Chen, Boyuan, et al.
Veröffentlicht: (2024)
Pre-Training Representations of Binary Code Using Contrastive Learning
von: Zhang, Yifan, et al.
Veröffentlicht: (2022)
von: Zhang, Yifan, et al.
Veröffentlicht: (2022)
CodeSAM: Source Code Representation Learning by Infusing Self-Attention with Multi-Code-View Graphs
von: Mathai, Alex, et al.
Veröffentlicht: (2024)
von: Mathai, Alex, et al.
Veröffentlicht: (2024)
TritonRL: Training LLMs to Think and Code Triton Without Cheating
von: Woo, Jiin, et al.
Veröffentlicht: (2025)
von: Woo, Jiin, et al.
Veröffentlicht: (2025)
RM -RF: Reward Model for Run-Free Unit Test Evaluation
von: Bruches, Elena, et al.
Veröffentlicht: (2026)
von: Bruches, Elena, et al.
Veröffentlicht: (2026)
Exploring Pass-Rate Reward in Reinforcement Learning for Code Generation
von: Li, Xin-Ye, et al.
Veröffentlicht: (2026)
von: Li, Xin-Ye, et al.
Veröffentlicht: (2026)
LoRA-MME: Multi-Model Ensemble of LoRA-Tuned Encoders for Code Comment Classification
von: Haider, Md Akib, et al.
Veröffentlicht: (2026)
von: Haider, Md Akib, et al.
Veröffentlicht: (2026)
CodeIF: Benchmarking the Instruction-Following Capabilities of Large Language Models for Code Generation
von: Yan, Kaiwen, et al.
Veröffentlicht: (2025)
von: Yan, Kaiwen, et al.
Veröffentlicht: (2025)
Benchmarking Reward Hack Detection in Code Environments via Contrastive Analysis
von: Deshpande, Darshan, et al.
Veröffentlicht: (2026)
von: Deshpande, Darshan, et al.
Veröffentlicht: (2026)
Calibration and Correctness of Language Models for Code
von: Spiess, Claudio, et al.
Veröffentlicht: (2024)
von: Spiess, Claudio, et al.
Veröffentlicht: (2024)
Automatic Semantic Augmentation of Language Model Prompts (for Code Summarization)
von: Ahmed, Toufique, et al.
Veröffentlicht: (2023)
von: Ahmed, Toufique, et al.
Veröffentlicht: (2023)
Can Code Language Models Learn Clarification-Seeking Behaviors?
von: Wu, Jie JW, et al.
Veröffentlicht: (2025)
von: Wu, Jie JW, et al.
Veröffentlicht: (2025)
Ensuring Functional Correctness of Large Code Models with Selective Generation
von: Jeong, Jaewoo, et al.
Veröffentlicht: (2025)
von: Jeong, Jaewoo, et al.
Veröffentlicht: (2025)
ReCode: Reinforcing Code Generation with Reasoning-Process Rewards
von: Fan, Lishui, et al.
Veröffentlicht: (2025)
von: Fan, Lishui, et al.
Veröffentlicht: (2025)
LEANCODE: Understanding Models Better for Code Simplification of Pre-trained Large Language Models
von: Wang, Yan, et al.
Veröffentlicht: (2025)
von: Wang, Yan, et al.
Veröffentlicht: (2025)
CodeFuse-13B: A Pretrained Multi-lingual Code Large Language Model
von: Di, Peng, et al.
Veröffentlicht: (2023)
von: Di, Peng, et al.
Veröffentlicht: (2023)
Enhancing Large Language Models with Faster Code Preprocessing for Vulnerability Detection
von: Gonçalves, José, et al.
Veröffentlicht: (2025)
von: Gonçalves, José, et al.
Veröffentlicht: (2025)
GitChameleon: Unmasking the Version-Switching Capabilities of Code Generation Models
von: Islah, Nizar, et al.
Veröffentlicht: (2024)
von: Islah, Nizar, et al.
Veröffentlicht: (2024)
Language Models are Better Bug Detector Through Code-Pair Classification
von: Alrashedy, Kamel, et al.
Veröffentlicht: (2023)
von: Alrashedy, Kamel, et al.
Veröffentlicht: (2023)
SemRep: Generative Code Representation Learning with Code Transformations
von: Li, Weichen, et al.
Veröffentlicht: (2026)
von: Li, Weichen, et al.
Veröffentlicht: (2026)
ReCodeAgent: A Multi-Agent Workflow for Language-agnostic Translation and Validation of Large-scale Repositories
von: Ibrahimzada, Ali Reza, et al.
Veröffentlicht: (2026)
von: Ibrahimzada, Ali Reza, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
ObscuraCoder: Powering Efficient Code LM Pre-Training Via Obfuscation Grounding
von: Paul, Indraneil, et al.
Veröffentlicht: (2025) -
AICD Bench: A Challenging Benchmark for AI-Generated Code Detection
von: Orel, Daniil, et al.
Veröffentlicht: (2026) -
IRCoder: Intermediate Representations Make Language Models Robust Multilingual Code Generators
von: Paul, Indraneil, et al.
Veröffentlicht: (2024) -
Aletheia: What Makes RLVR For Code Verifiers Tick?
von: Venkatkrishna, Vatsal, et al.
Veröffentlicht: (2026) -
$\texttt{Droid}$: A Resource Suite for AI-Generated Code Detection
von: Orel, Daniil, et al.
Veröffentlicht: (2025)