Skill over Scale: The Case for Medium, Domain-Specific Models for SE
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Mukherjee, Manisha, Hellendoorn, Vincent J. |
|---|---|
| Format: | Preprint |
| Publié: |
2023
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
SOSecure: Safer Code Generation with RAG and StackOverflow Discussions
par: Mukherjee, Manisha, et autres
Publié: (2025)
par: Mukherjee, Manisha, et autres
Publié: (2025)
Inference-Time Safety For Code LLMs Via Retrieval-Augmented Revision
par: Mukherjee, Manisha, et autres
Publié: (2026)
par: Mukherjee, Manisha, et autres
Publié: (2026)
EvoCodeBench: An Evolving Code Generation Benchmark with Domain-Specific Evaluations
par: Li, Jia, et autres
Publié: (2024)
par: Li, Jia, et autres
Publié: (2024)
SkillCraft: Can LLM Agents Learn to Use Tools Skillfully?
par: Chen, Shiqi, et autres
Publié: (2026)
par: Chen, Shiqi, et autres
Publié: (2026)
EffiSkill: Agent Skill Based Automated Code Efficiency Optimization
par: Wang, Zimu, et autres
Publié: (2026)
par: Wang, Zimu, et autres
Publié: (2026)
ConCodeEval: Evaluating Large Language Models for Code Constraints in Domain-Specific Languages
par: Kammakomati, Mehant, et autres
Publié: (2024)
par: Kammakomati, Mehant, et autres
Publié: (2024)
Towards AI Evaluation in Domain-Specific RAG Systems: The AgriHubi Case Study
par: Hasan, Md. Toufique, et autres
Publié: (2026)
par: Hasan, Md. Toufique, et autres
Publié: (2026)
Adapting Large Language Models to Log Analysis with Interpretable Domain Knowledge
par: Ji, Yuhe, et autres
Publié: (2024)
par: Ji, Yuhe, et autres
Publié: (2024)
Dynamic Scaling of Unit Tests for Code Reward Modeling
par: Ma, Zeyao, et autres
Publié: (2025)
par: Ma, Zeyao, et autres
Publié: (2025)
FASTRIC: Prompt Specification Language for Verifiable LLM Interactions
par: Jin, Wen-Long
Publié: (2025)
par: Jin, Wen-Long
Publié: (2025)
From Procedural Skills to Strategy Genes: Towards Experience-Driven Test-Time Evolution
par: Wang, Junjie, et autres
Publié: (2026)
par: Wang, Junjie, et autres
Publié: (2026)
CodeSpecBench: Benchmarking LLMs for Executable Behavioral Specification Generation
par: Chen, Zaoyu, et autres
Publié: (2026)
par: Chen, Zaoyu, et autres
Publié: (2026)
Enhancing Project-Specific Code Completion by Inferring Internal API Information
par: Deng, Le, et autres
Publié: (2025)
par: Deng, Le, et autres
Publié: (2025)
Evaluation of the Programming Skills of Large Language Models
par: Heitz, Luc Bryan, et autres
Publié: (2024)
par: Heitz, Luc Bryan, et autres
Publié: (2024)
DI-BENCH: Benchmarking Large Language Models on Dependency Inference with Testable Repositories at Scale
par: Zhang, Linghao, et autres
Publié: (2025)
par: Zhang, Linghao, et autres
Publié: (2025)
ScaleBox: Enabling High-Fidelity and Scalable Code Verification for Large Language Models
par: Zheng, Jiasheng, et autres
Publié: (2026)
par: Zheng, Jiasheng, et autres
Publié: (2026)
On the Potential and Limitations of Few-Shot In-Context Learning to Generate Metamorphic Specifications for Tax Preparation Software
par: Srinivas, Dananjay, et autres
Publié: (2023)
par: Srinivas, Dananjay, et autres
Publié: (2023)
Prompting Large Language Models to Tackle the Full Software Development Lifecycle: A Case Study
par: Li, Bowen, et autres
Publié: (2024)
par: Li, Bowen, et autres
Publié: (2024)
An Exploratory Study of ML Sketches and Visual Code Assistants
par: Gomes, Luís F., et autres
Publié: (2024)
par: Gomes, Luís F., et autres
Publié: (2024)
Using Large Language Models for Student-Code Guided Test Case Generation in Computer Science Education
par: Kumar, Nischal Ashok, et autres
Publié: (2024)
par: Kumar, Nischal Ashok, et autres
Publié: (2024)
Zero-Shot Cross-Domain Code Search without Fine-Tuning
par: Liang, Keyu, et autres
Publié: (2025)
par: Liang, Keyu, et autres
Publié: (2025)
Seed-Guided Fine-Grained Entity Typing in Science and Engineering Domains
par: Zhang, Yu, et autres
Publié: (2024)
par: Zhang, Yu, et autres
Publié: (2024)
SV-TrustEval-C: Evaluating Structure and Semantic Reasoning in Large Language Models for Source Code Vulnerability Analysis
par: Li, Yansong, et autres
Publié: (2025)
par: Li, Yansong, et autres
Publié: (2025)
From Prediction to Application: Language Model-based Code Knowledge Tracing with Domain Adaptive Pre-Training and Automatic Feedback System with Pedagogical Prompting for Comprehensive Programming Education
par: Lee, Unggi, et autres
Publié: (2024)
par: Lee, Unggi, et autres
Publié: (2024)
ClassEval-Pro: A Cross-Domain Benchmark for Class-Level Code Generation
par: Chen, Yeheng, et autres
Publié: (2026)
par: Chen, Yeheng, et autres
Publié: (2026)
Revisiting Unnaturalness for Automated Program Repair in the Era of Large Language Models
par: Yang, Aidan Z. H., et autres
Publié: (2024)
par: Yang, Aidan Z. H., et autres
Publié: (2024)
Can LLMs Generate High-Quality Test Cases for Algorithm Problems? TestCase-Eval: A Systematic Evaluation of Fault Coverage and Exposure
par: Yang, Zheyuan, et autres
Publié: (2025)
par: Yang, Zheyuan, et autres
Publié: (2025)
Agent Skills in the Wild: An Empirical Study of Security Vulnerabilities at Scale
par: Liu, Yi, et autres
Publié: (2026)
par: Liu, Yi, et autres
Publié: (2026)
Combinatorial Synthesis: Scaling Code RLVR via Atomic Decomposition and Recombination
par: Zheng, Jiasheng, et autres
Publié: (2026)
par: Zheng, Jiasheng, et autres
Publié: (2026)
TeamUp: Semantic Project Matching and Team Formation for Learning at Scale
par: Gulwani, Dhruv, et autres
Publié: (2026)
par: Gulwani, Dhruv, et autres
Publié: (2026)
SWE-rebench V2: Language-Agnostic SWE Task Collection at Scale
par: Badertdinov, Ibragim, et autres
Publié: (2026)
par: Badertdinov, Ibragim, et autres
Publié: (2026)
CodeContests+: High-Quality Test Case Generation for Competitive Programming
par: Wang, Zihan, et autres
Publié: (2025)
par: Wang, Zihan, et autres
Publié: (2025)
Using an LLM to Help With Code Understanding
par: Nam, Daye, et autres
Publié: (2023)
par: Nam, Daye, et autres
Publié: (2023)
Firefly: Illuminating Large-Scale Verified Tool-Call Data Generation from Real APIs
par: Lu, Yuxuan, et autres
Publié: (2026)
par: Lu, Yuxuan, et autres
Publié: (2026)
SRLCG: Self-Rectified Large-Scale Code Generation with Multidimensional Chain-of-Thought and Dynamic Backtracking
par: Ma, Hongru, et autres
Publié: (2025)
par: Ma, Hongru, et autres
Publié: (2025)
Explainable Compliance Detection with Multi-Hop Natural Language Inference on Assurance Case Structure
par: Ikhwantri, Fariz, et autres
Publié: (2025)
par: Ikhwantri, Fariz, et autres
Publié: (2025)
An LLM-as-Judge Metric for Bridging the Gap with Human Evaluation in SE Tasks
par: Zhou, Xin, et autres
Publié: (2025)
par: Zhou, Xin, et autres
Publié: (2025)
CPSLint: A Domain-Specific Language Providing Data Validation and Sanitisation for Industrial Cyber-Physical Systems
par: Odyurt, Uraz, et autres
Publié: (2025)
par: Odyurt, Uraz, et autres
Publié: (2025)
Kimi-Dev: Agentless Training as Skill Prior for SWE-Agents
par: Yang, Zonghan, et autres
Publié: (2025)
par: Yang, Zonghan, et autres
Publié: (2025)
MoSE: Hierarchical Self-Distillation Enhances Early Layer Embeddings
par: Gurioli, Andrea, et autres
Publié: (2025)
par: Gurioli, Andrea, et autres
Publié: (2025)
Documents similaires
-
SOSecure: Safer Code Generation with RAG and StackOverflow Discussions
par: Mukherjee, Manisha, et autres
Publié: (2025) -
Inference-Time Safety For Code LLMs Via Retrieval-Augmented Revision
par: Mukherjee, Manisha, et autres
Publié: (2026) -
EvoCodeBench: An Evolving Code Generation Benchmark with Domain-Specific Evaluations
par: Li, Jia, et autres
Publié: (2024) -
SkillCraft: Can LLM Agents Learn to Use Tools Skillfully?
par: Chen, Shiqi, et autres
Publié: (2026) -
EffiSkill: Agent Skill Based Automated Code Efficiency Optimization
par: Wang, Zimu, et autres
Publié: (2026)