KodCode: A Diverse, Challenging, and Verifiable Synthetic Dataset for Coding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xu, Zhangchen, Liu, Yang, Yin, Yueqin, Zhou, Mingyuan, Poovendran, Radha |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning
von: Xu, Zhangchen, et al.
Veröffentlicht: (2025)
von: Xu, Zhangchen, et al.
Veröffentlicht: (2025)
Relative Preference Optimization: Enhancing LLM Alignment through Contrasting Responses across Identical and Diverse Prompts
von: Yin, Yueqin, et al.
Veröffentlicht: (2024)
von: Yin, Yueqin, et al.
Veröffentlicht: (2024)
TOUCAN: Synthesizing 1.5M Tool-Agentic Data from Real-World MCP Environments
von: Xu, Zhangchen, et al.
Veröffentlicht: (2025)
von: Xu, Zhangchen, et al.
Veröffentlicht: (2025)
Scoring Verifiers: Evaluating Synthetic Verification for Code and Reasoning
von: Ficek, Aleksander, et al.
Veröffentlicht: (2025)
von: Ficek, Aleksander, et al.
Veröffentlicht: (2025)
Stronger Models are NOT Stronger Teachers for Instruction Tuning
von: Xu, Zhangchen, et al.
Veröffentlicht: (2024)
von: Xu, Zhangchen, et al.
Veröffentlicht: (2024)
VisualSphinx: Large-Scale Synthetic Vision Logic Puzzles for RL
von: Feng, Yichen, et al.
Veröffentlicht: (2025)
von: Feng, Yichen, et al.
Veröffentlicht: (2025)
ChatBug: A Common Vulnerability of Aligned LLMs Induced by Chat Templates
von: Jiang, Fengqing, et al.
Veröffentlicht: (2024)
von: Jiang, Fengqing, et al.
Veröffentlicht: (2024)
Generalizable End-to-End Tool-Use RL with Synthetic CodeGym
von: Du, Weihua, et al.
Veröffentlicht: (2025)
von: Du, Weihua, et al.
Veröffentlicht: (2025)
Steering Multimodal Large Language Models Decoding for Context-Aware Safety
von: Liu, Zheyuan, et al.
Veröffentlicht: (2025)
von: Liu, Zheyuan, et al.
Veröffentlicht: (2025)
ACE: A Model Poisoning Attack on Contribution Evaluation Methods in Federated Learning
von: Xu, Zhangchen, et al.
Veröffentlicht: (2024)
von: Xu, Zhangchen, et al.
Veröffentlicht: (2024)
ArtPrompt: ASCII Art-based Jailbreak Attacks against Aligned LLMs
von: Jiang, Fengqing, et al.
Veröffentlicht: (2024)
von: Jiang, Fengqing, et al.
Veröffentlicht: (2024)
APIGen: Automated Pipeline for Generating Verifiable and Diverse Function-Calling Datasets
von: Liu, Zuxin, et al.
Veröffentlicht: (2024)
von: Liu, Zuxin, et al.
Veröffentlicht: (2024)
ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning
von: Lin, Bill Yuchen, et al.
Veröffentlicht: (2025)
von: Lin, Bill Yuchen, et al.
Veröffentlicht: (2025)
SafeDecoding: Defending against Jailbreak Attacks via Safety-Aware Decoding
von: Xu, Zhangchen, et al.
Veröffentlicht: (2024)
von: Xu, Zhangchen, et al.
Veröffentlicht: (2024)
Magpie: Alignment Data Synthesis from Scratch by Prompting Aligned LLMs with Nothing
von: Xu, Zhangchen, et al.
Veröffentlicht: (2024)
von: Xu, Zhangchen, et al.
Veröffentlicht: (2024)
Code Simulation Challenges for Large Language Models
von: La Malfa, Emanuele, et al.
Veröffentlicht: (2024)
von: La Malfa, Emanuele, et al.
Veröffentlicht: (2024)
VERINA: Benchmarking Verifiable Code Generation
von: Ye, Zhe, et al.
Veröffentlicht: (2025)
von: Ye, Zhe, et al.
Veröffentlicht: (2025)
Code to Think, Think to Code: A Survey on Code-Enhanced Reasoning and Reasoning-Driven Code Intelligence in LLMs
von: Yang, Dayu, et al.
Veröffentlicht: (2025)
von: Yang, Dayu, et al.
Veröffentlicht: (2025)
Building a Foundational Guardrail for General Agentic Systems via Synthetic Data
von: Huang, Yue, et al.
Veröffentlicht: (2025)
von: Huang, Yue, et al.
Veröffentlicht: (2025)
SafeChain: Safety of Language Models with Long Chain-of-Thought Reasoning Capabilities
von: Jiang, Fengqing, et al.
Veröffentlicht: (2025)
von: Jiang, Fengqing, et al.
Veröffentlicht: (2025)
To Code or not to Code? Adaptive Tool Integration for Math Language Models via Expectation-Maximization
von: Wang, Haozhe, et al.
Veröffentlicht: (2025)
von: Wang, Haozhe, et al.
Veröffentlicht: (2025)
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
von: Chen, Yang, et al.
Veröffentlicht: (2025)
von: Chen, Yang, et al.
Veröffentlicht: (2025)
A Synthetic Dataset for Personal Attribute Inference
von: Yukhymenko, Hanna, et al.
Veröffentlicht: (2024)
von: Yukhymenko, Hanna, et al.
Veröffentlicht: (2024)
Enhancing Rare Codes via Probability-Biased Directed Graph Attention for Long-Tail ICD Coding
von: Chen, Tianlei, et al.
Veröffentlicht: (2025)
von: Chen, Tianlei, et al.
Veröffentlicht: (2025)
Temporal Sampling for Forgotten Reasoning in LLMs
von: Li, Yuetai, et al.
Veröffentlicht: (2025)
von: Li, Yuetai, et al.
Veröffentlicht: (2025)
SecureCode: A Production-Grade Multi-Turn Dataset for Training Security-Aware Code Generation Models
von: Thornton, Scott
Veröffentlicht: (2025)
von: Thornton, Scott
Veröffentlicht: (2025)
CodeRefine: A Pipeline for Enhancing LLM-Generated Code Implementations of Research Papers
von: Trofimova, Ekaterina, et al.
Veröffentlicht: (2024)
von: Trofimova, Ekaterina, et al.
Veröffentlicht: (2024)
The Strongest Teacher Is Not Always the Best Teacher: Student-Centric Answer Selection
von: Hu, Zhengyu, et al.
Veröffentlicht: (2026)
von: Hu, Zhengyu, et al.
Veröffentlicht: (2026)
Verifying the Verifiers: Unveiling Pitfalls and Potentials in Fact Verifiers
von: Seo, Wooseok, et al.
Veröffentlicht: (2025)
von: Seo, Wooseok, et al.
Veröffentlicht: (2025)
Towards Effective Code-Integrated Reasoning
von: Bai, Fei, et al.
Veröffentlicht: (2025)
von: Bai, Fei, et al.
Veröffentlicht: (2025)
CoRelation: Boosting Automatic ICD Coding Through Contextualized Code Relation Learning
von: Luo, Junyu, et al.
Veröffentlicht: (2024)
von: Luo, Junyu, et al.
Veröffentlicht: (2024)
MedAgentGym: A Scalable Agentic Training Environment for Code-Centric Reasoning in Biomedical Data Science
von: Xu, Ran, et al.
Veröffentlicht: (2025)
von: Xu, Ran, et al.
Veröffentlicht: (2025)
Improving Multilingual Instruction Finetuning via Linguistically Natural and Diverse Datasets
von: Indurthi, Sathish Reddy, et al.
Veröffentlicht: (2024)
von: Indurthi, Sathish Reddy, et al.
Veröffentlicht: (2024)
CLEVER: A Curated Benchmark for Formally Verified Code Generation
von: Thakur, Amitayush, et al.
Veröffentlicht: (2025)
von: Thakur, Amitayush, et al.
Veröffentlicht: (2025)
Multimodal Medical Code Tokenizer
von: Su, Xiaorui, et al.
Veröffentlicht: (2025)
von: Su, Xiaorui, et al.
Veröffentlicht: (2025)
DreamPRM-Code: Function-as-Step Process Reward Model with Label Correction for LLM Coding
von: Zhang, Ruiyi, et al.
Veröffentlicht: (2025)
von: Zhang, Ruiyi, et al.
Veröffentlicht: (2025)
CoSPlay: Cooperative Self-Play at Test-Time with Self-Generated Code and Unit Test
von: Hu, Zhangyi, et al.
Veröffentlicht: (2026)
von: Hu, Zhangyi, et al.
Veröffentlicht: (2026)
SCI-Verifier: Scientific Verifier with Thinking
von: Zheng, Shenghe, et al.
Veröffentlicht: (2025)
von: Zheng, Shenghe, et al.
Veröffentlicht: (2025)
Investigating the Transferability of Code Repair for Low-Resource Programming Languages
von: Wong, Kyle, et al.
Veröffentlicht: (2024)
von: Wong, Kyle, et al.
Veröffentlicht: (2024)
StRuCom: A Novel Dataset of Structured Code Comments in Russian
von: Dziuba, Maria, et al.
Veröffentlicht: (2025)
von: Dziuba, Maria, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
TinyV: Reducing False Negatives in Verification Improves RL for LLM Reasoning
von: Xu, Zhangchen, et al.
Veröffentlicht: (2025) -
Relative Preference Optimization: Enhancing LLM Alignment through Contrasting Responses across Identical and Diverse Prompts
von: Yin, Yueqin, et al.
Veröffentlicht: (2024) -
TOUCAN: Synthesizing 1.5M Tool-Agentic Data from Real-World MCP Environments
von: Xu, Zhangchen, et al.
Veröffentlicht: (2025) -
Scoring Verifiers: Evaluating Synthetic Verification for Code and Reasoning
von: Ficek, Aleksander, et al.
Veröffentlicht: (2025) -
Stronger Models are NOT Stronger Teachers for Instruction Tuning
von: Xu, Zhangchen, et al.
Veröffentlicht: (2024)