Code Representation Learning At Scale
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Dejiao, Ahmad, Wasi, Tan, Ming, Ding, Hantian, Nallapati, Ramesh, Roth, Dan, Ma, Xiaofei, Xiang, Bing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Repoformer: Selective Retrieval for Repository-Level Code Completion
von: Wu, Di, et al.
Veröffentlicht: (2024)
von: Wu, Di, et al.
Veröffentlicht: (2024)
From Output to Evaluation: Does Raw Instruction-Tuned Code LLMs Output Suffice for Fill-in-the-Middle Code Generation?
von: Ahmad, Wasi Uddin, et al.
Veröffentlicht: (2025)
von: Ahmad, Wasi Uddin, et al.
Veröffentlicht: (2025)
Fewer Truncations Improve Language Modeling
von: Ding, Hantian, et al.
Veröffentlicht: (2024)
von: Ding, Hantian, et al.
Veröffentlicht: (2024)
BASS: Batched Attention-optimized Speculative Sampling
von: Qian, Haifeng, et al.
Veröffentlicht: (2024)
von: Qian, Haifeng, et al.
Veröffentlicht: (2024)
Planning-Aware Code Infilling via Horizon-Length Prediction
von: Ding, Yifeng, et al.
Veröffentlicht: (2024)
von: Ding, Yifeng, et al.
Veröffentlicht: (2024)
ShieldedCode: Learning Robust Representations for Virtual Machine Protected Code
von: Mo, Mingqiao, et al.
Veröffentlicht: (2026)
von: Mo, Mingqiao, et al.
Veröffentlicht: (2026)
Learning Code Preference via Synthetic Evolution
von: Liu, Jiawei, et al.
Veröffentlicht: (2024)
von: Liu, Jiawei, et al.
Veröffentlicht: (2024)
OpenCodeReasoning-II: A Simple Test Time Scaling Approach via Self-Critique
von: Ahmad, Wasi Uddin, et al.
Veröffentlicht: (2025)
von: Ahmad, Wasi Uddin, et al.
Veröffentlicht: (2025)
OpenCodeReasoning: Advancing Data Distillation for Competitive Coding
von: Ahmad, Wasi Uddin, et al.
Veröffentlicht: (2025)
von: Ahmad, Wasi Uddin, et al.
Veröffentlicht: (2025)
Token Alignment via Character Matching for Subword Completion
von: Athiwaratkun, Ben, et al.
Veröffentlicht: (2024)
von: Athiwaratkun, Ben, et al.
Veröffentlicht: (2024)
Genetic Instruct: Scaling up Synthetic Generation of Coding Instructions for Large Language Models
von: Majumdar, Somshubra, et al.
Veröffentlicht: (2024)
von: Majumdar, Somshubra, et al.
Veröffentlicht: (2024)
OpenCodeInstruct: A Large-scale Instruction Tuning Dataset for Code LLMs
von: Ahmad, Wasi Uddin, et al.
Veröffentlicht: (2025)
von: Ahmad, Wasi Uddin, et al.
Veröffentlicht: (2025)
On Leveraging Encoder-only Pre-trained Language Models for Effective Keyphrase Generation
von: Wu, Di, et al.
Veröffentlicht: (2024)
von: Wu, Di, et al.
Veröffentlicht: (2024)
Pre-trained Language Models for Keyphrase Generation: A Thorough Empirical Study
von: Wu, Di, et al.
Veröffentlicht: (2022)
von: Wu, Di, et al.
Veröffentlicht: (2022)
Lightweight reranking for language model generations
von: Jain, Siddhartha, et al.
Veröffentlicht: (2023)
von: Jain, Siddhartha, et al.
Veröffentlicht: (2023)
Reasoning in Token Economies: Budget-Aware Evaluation of LLM Reasoning Strategies
von: Wang, Junlin, et al.
Veröffentlicht: (2024)
von: Wang, Junlin, et al.
Veröffentlicht: (2024)
Active Use of Latent Constituency Representation in both Humans and Large Language Models
von: Liu, Wei, et al.
Veröffentlicht: (2024)
von: Liu, Wei, et al.
Veröffentlicht: (2024)
From SWE-ZERO to SWE-HERO: Execution-free to Execution-based Fine-tuning for Software Engineering Agents
von: Ludwig, Nikolai, et al.
Veröffentlicht: (2026)
von: Ludwig, Nikolai, et al.
Veröffentlicht: (2026)
Reasoning is about giving reasons
von: Shah, Krunal, et al.
Veröffentlicht: (2025)
von: Shah, Krunal, et al.
Veröffentlicht: (2025)
Conflicts in Texts: Data, Implications and Challenges
von: Liu, Siyi, et al.
Veröffentlicht: (2025)
von: Liu, Siyi, et al.
Veröffentlicht: (2025)
Scaling Agentic Verifier for Competitive Coding
von: Ma, Zeyao, et al.
Veröffentlicht: (2026)
von: Ma, Zeyao, et al.
Veröffentlicht: (2026)
Cyber-Zero: Training Cybersecurity Agents without Runtime
von: Zhuo, Terry Yue, et al.
Veröffentlicht: (2025)
von: Zhuo, Terry Yue, et al.
Veröffentlicht: (2025)
UniCoder: Scaling Code Large Language Model via Universal Code
von: Sun, Tao, et al.
Veröffentlicht: (2024)
von: Sun, Tao, et al.
Veröffentlicht: (2024)
IITR-CIOL@NLU of Devanagari Script Languages 2025: Multilingual Hate Speech Detection and Target Identification in Devanagari-Scripted Languages
von: Gupta, Siddhant, et al.
Veröffentlicht: (2024)
von: Gupta, Siddhant, et al.
Veröffentlicht: (2024)
SRLCG: Self-Rectified Large-Scale Code Generation with Multidimensional Chain-of-Thought and Dynamic Backtracking
von: Ma, Hongru, et al.
Veröffentlicht: (2025)
von: Ma, Hongru, et al.
Veröffentlicht: (2025)
HRGraph: Leveraging LLMs for HR Data Knowledge Graphs with Information Propagation-based Job Recommendation
von: Wasi, Azmine Toushik
Veröffentlicht: (2024)
von: Wasi, Azmine Toushik
Veröffentlicht: (2024)
Benchmarking LLM Guardrails in Handling Multilingual Toxicity
von: Yang, Yahan, et al.
Veröffentlicht: (2024)
von: Yang, Yahan, et al.
Veröffentlicht: (2024)
Scaling Test-Time Compute to Achieve IOI Gold Medal with Open-Weight Models
von: Samadi, Mehrzad, et al.
Veröffentlicht: (2025)
von: Samadi, Mehrzad, et al.
Veröffentlicht: (2025)
SocREval: Large Language Models with the Socratic Method for Reference-Free Reasoning Evaluation
von: He, Hangfeng, et al.
Veröffentlicht: (2023)
von: He, Hangfeng, et al.
Veröffentlicht: (2023)
Training Language Model Agents to Find Vulnerabilities with CTF-Dojo
von: Zhuo, Terry Yue, et al.
Veröffentlicht: (2025)
von: Zhuo, Terry Yue, et al.
Veröffentlicht: (2025)
Unified Scaling Laws for Compressed Representations
von: Panferov, Andrei, et al.
Veröffentlicht: (2025)
von: Panferov, Andrei, et al.
Veröffentlicht: (2025)
Dhoroni: Exploring Bengali Climate Change and Environmental Views with a Multi-Perspective News Dataset and Natural Language Processing
von: Wasi, Azmine Toushik, et al.
Veröffentlicht: (2024)
von: Wasi, Azmine Toushik, et al.
Veröffentlicht: (2024)
Dynamic Scaling of Unit Tests for Code Reward Modeling
von: Ma, Zeyao, et al.
Veröffentlicht: (2025)
von: Ma, Zeyao, et al.
Veröffentlicht: (2025)
Generating Grounded Responses to Counter Misinformation via Learning Efficient Fine-Grained Critiques
von: Xu, Xiaofei, et al.
Veröffentlicht: (2025)
von: Xu, Xiaofei, et al.
Veröffentlicht: (2025)
Explainable Identification of Hate Speech towards Islam using Graph Neural Networks
von: Wasi, Azmine Toushik
Veröffentlicht: (2023)
von: Wasi, Azmine Toushik
Veröffentlicht: (2023)
Self-supervised Analogical Learning using Language Models
von: Zhou, Ben, et al.
Veröffentlicht: (2025)
von: Zhou, Ben, et al.
Veröffentlicht: (2025)
Enhancing Source Code Classification Effectiveness via Prompt Learning Incorporating Knowledge Features
von: Ma, Yong, et al.
Veröffentlicht: (2024)
von: Ma, Yong, et al.
Veröffentlicht: (2024)
CodeSimpleQA: Scaling Factuality in Code Large Language Models
von: Yang, Jian, et al.
Veröffentlicht: (2025)
von: Yang, Jian, et al.
Veröffentlicht: (2025)
MrGuard: A Multilingual Reasoning Guardrail for Universal LLM Safety
von: Yang, Yahan, et al.
Veröffentlicht: (2025)
von: Yang, Yahan, et al.
Veröffentlicht: (2025)
Infinite-Instruct: Synthesizing Scaling Code instruction Data with Bidirectional Synthesis and Static Verification
von: Xing, Wenjing, et al.
Veröffentlicht: (2025)
von: Xing, Wenjing, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Repoformer: Selective Retrieval for Repository-Level Code Completion
von: Wu, Di, et al.
Veröffentlicht: (2024) -
From Output to Evaluation: Does Raw Instruction-Tuned Code LLMs Output Suffice for Fill-in-the-Middle Code Generation?
von: Ahmad, Wasi Uddin, et al.
Veröffentlicht: (2025) -
Fewer Truncations Improve Language Modeling
von: Ding, Hantian, et al.
Veröffentlicht: (2024) -
BASS: Batched Attention-optimized Speculative Sampling
von: Qian, Haifeng, et al.
Veröffentlicht: (2024) -
Planning-Aware Code Infilling via Horizon-Length Prediction
von: Ding, Yifeng, et al.
Veröffentlicht: (2024)