Gespeichert in:
| Hauptverfasser: | Mu, Wenchuan, Lim, Kwan Hui |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2404.16457 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Towards Reliable Evaluation of Neural Program Repair with Natural Robustness Testing
von: Le-Cong, Thanh, et al.
Veröffentlicht: (2024)
von: Le-Cong, Thanh, et al.
Veröffentlicht: (2024)
Towards Better Correctness and Efficiency in Code Generation
von: Feng, Yunlong, et al.
Veröffentlicht: (2025)
von: Feng, Yunlong, et al.
Veröffentlicht: (2025)
Label-Free Topic-Focused Summarization Using Query Augmentation
von: Mu, Wenchuan, et al.
Veröffentlicht: (2024)
von: Mu, Wenchuan, et al.
Veröffentlicht: (2024)
OLAF: Towards Robust LLM-Based Annotation Framework in Empirical Software Engineering
von: Imran, Mia Mohammad, et al.
Veröffentlicht: (2025)
von: Imran, Mia Mohammad, et al.
Veröffentlicht: (2025)
Benchmarking Harmonized Tariff Schedule Classification Models
von: Judy, Bryce
Veröffentlicht: (2024)
von: Judy, Bryce
Veröffentlicht: (2024)
Model Provenance via Model DNA
von: Mu, Xin, et al.
Veröffentlicht: (2023)
von: Mu, Xin, et al.
Veröffentlicht: (2023)
Towards Robust LLM Post-Training: Automatic Failure Management for Reinforcement Fine-Tuning
von: Zhang, Lingzhe, et al.
Veröffentlicht: (2026)
von: Zhang, Lingzhe, et al.
Veröffentlicht: (2026)
Human-Machine Co-Boosted Bug Report Identification with Mutualistic Neural Active Learning
von: Long, Guoming, et al.
Veröffentlicht: (2026)
von: Long, Guoming, et al.
Veröffentlicht: (2026)
Towards a Classification of Open-Source ML Models and Datasets for Software Engineering
von: González, Alexandra, et al.
Veröffentlicht: (2024)
von: González, Alexandra, et al.
Veröffentlicht: (2024)
Beyond Retrieval: A Multitask Benchmark and Model for Code Search
von: Xue, Siqiao, et al.
Veröffentlicht: (2026)
von: Xue, Siqiao, et al.
Veröffentlicht: (2026)
Conventional Commit Classification using Large Language Models and Prompt Engineering
von: Quadir, H. M. Sazzad, et al.
Veröffentlicht: (2026)
von: Quadir, H. M. Sazzad, et al.
Veröffentlicht: (2026)
CodeFort: Robust Training for Code Generation Models
von: Zhang, Yuhao, et al.
Veröffentlicht: (2024)
von: Zhang, Yuhao, et al.
Veröffentlicht: (2024)
Exploring the Potential of Large Language Models in Fine-Grained Review Comment Classification
von: Nguyen, Linh, et al.
Veröffentlicht: (2025)
von: Nguyen, Linh, et al.
Veröffentlicht: (2025)
DRAGON: Robust Classification for Very Large Collections of Software Repositories
von: Balla, Stefano, et al.
Veröffentlicht: (2026)
von: Balla, Stefano, et al.
Veröffentlicht: (2026)
Towards a General Framework for HTN Modeling with LLMs
von: Puerta-Merino, Israel, et al.
Veröffentlicht: (2025)
von: Puerta-Merino, Israel, et al.
Veröffentlicht: (2025)
Towards Leveraging Large Language Model Summaries for Topic Modeling in Source Code
von: Carissimi, Michele, et al.
Veröffentlicht: (2025)
von: Carissimi, Michele, et al.
Veröffentlicht: (2025)
Unveiling Project-Specific Bias in Neural Code Models
von: Li, Zhiming, et al.
Veröffentlicht: (2022)
von: Li, Zhiming, et al.
Veröffentlicht: (2022)
Precision in Practice: Knowledge Guided Code Summarizing Grounded in Industrial Expectations
von: Li, Jintai, et al.
Veröffentlicht: (2026)
von: Li, Jintai, et al.
Veröffentlicht: (2026)
Towards a Neural Debugger for Python
von: Beck, Maximilian, et al.
Veröffentlicht: (2026)
von: Beck, Maximilian, et al.
Veröffentlicht: (2026)
Enhancing Deployment-Time Predictive Model Robustness for Code Analysis and Optimization
von: Wang, Huanting, et al.
Veröffentlicht: (2024)
von: Wang, Huanting, et al.
Veröffentlicht: (2024)
Post-Incorporating Code Structural Knowledge into Pretrained Models via ICL for Code Translation
von: Du, Yali, et al.
Veröffentlicht: (2025)
von: Du, Yali, et al.
Veröffentlicht: (2025)
Towards a Domain-Specific Modelling Environment for Reinforcement Learning
von: Sinani, Natalie, et al.
Veröffentlicht: (2024)
von: Sinani, Natalie, et al.
Veröffentlicht: (2024)
CodeSSM: Towards State Space Models for Code Understanding
von: Verma, Shweta, et al.
Veröffentlicht: (2025)
von: Verma, Shweta, et al.
Veröffentlicht: (2025)
Are Large Language Models Robust in Understanding Code Against Semantics-Preserving Mutations?
von: Orvalho, Pedro, et al.
Veröffentlicht: (2025)
von: Orvalho, Pedro, et al.
Veröffentlicht: (2025)
Towards Better Code Understanding in Decoder-Only Models with Contrastive Learning
von: Lin, Jiayi, et al.
Veröffentlicht: (2024)
von: Lin, Jiayi, et al.
Veröffentlicht: (2024)
Towards a Digital Twin Modeling Method for Container Terminal Port
von: Hakimi, Faouzi, et al.
Veröffentlicht: (2025)
von: Hakimi, Faouzi, et al.
Veröffentlicht: (2025)
An Experience Report on Regression-Free Repair of Deep Neural Network Model
von: Nakagawa, Takao, et al.
Veröffentlicht: (2025)
von: Nakagawa, Takao, et al.
Veröffentlicht: (2025)
Robustness and Reasoning Fidelity of Large Language Models in Long-Context Code Question Answering
von: Maharaj, Kishan, et al.
Veröffentlicht: (2026)
von: Maharaj, Kishan, et al.
Veröffentlicht: (2026)
A-ProS: Towards Reliable Autonomous Programming Through Multi-Model Feedback
von: Tabassum, Anika, et al.
Veröffentlicht: (2026)
von: Tabassum, Anika, et al.
Veröffentlicht: (2026)
Rethinking Scientific Modeling: Toward Physically Consistent and Simulation-Executable Programmatic Generation
von: Jiang, Yongqing, et al.
Veröffentlicht: (2026)
von: Jiang, Yongqing, et al.
Veröffentlicht: (2026)
WebCompass: Towards Multimodal Web Coding Evaluation for Code Language Models
von: Lei, Xinping, et al.
Veröffentlicht: (2026)
von: Lei, Xinping, et al.
Veröffentlicht: (2026)
Towards Advancing Code Generation with Large Language Models: A Research Roadmap
von: Jin, Haolin, et al.
Veröffentlicht: (2025)
von: Jin, Haolin, et al.
Veröffentlicht: (2025)
Toward a Theory of Causation for Interpreting Neural Code Models
von: Palacio, David N., et al.
Veröffentlicht: (2023)
von: Palacio, David N., et al.
Veröffentlicht: (2023)
PACIFIC: a framework for generating benchmarks to check Precise Automatically Checked Instruction Following In Code
von: Dreyfuss, Itay, et al.
Veröffentlicht: (2025)
von: Dreyfuss, Itay, et al.
Veröffentlicht: (2025)
Towards Robust Agentic CUDA Kernel Benchmarking, Verification, and Optimization
von: Lange, Robert Tjarko, et al.
Veröffentlicht: (2025)
von: Lange, Robert Tjarko, et al.
Veröffentlicht: (2025)
Toward Automated Validation of Language Model Synthesized Test Cases using Semantic Entropy
von: Taherkhani, Hamed, et al.
Veröffentlicht: (2024)
von: Taherkhani, Hamed, et al.
Veröffentlicht: (2024)
SWE-Compass: Towards Unified Evaluation of Agentic Coding Abilities for Large Language Models
von: Xu, Jingxuan, et al.
Veröffentlicht: (2025)
von: Xu, Jingxuan, et al.
Veröffentlicht: (2025)
RobuNFR: Evaluating the Robustness of Large Language Models on Non-Functional Requirements Aware Code Generation
von: Lin, Feng, et al.
Veröffentlicht: (2025)
von: Lin, Feng, et al.
Veröffentlicht: (2025)
When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions
von: Larbi, Maya, et al.
Veröffentlicht: (2025)
von: Larbi, Maya, et al.
Veröffentlicht: (2025)
The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management
von: Lindenbauer, Tobias, et al.
Veröffentlicht: (2025)
von: Lindenbauer, Tobias, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Towards Reliable Evaluation of Neural Program Repair with Natural Robustness Testing
von: Le-Cong, Thanh, et al.
Veröffentlicht: (2024) -
Towards Better Correctness and Efficiency in Code Generation
von: Feng, Yunlong, et al.
Veröffentlicht: (2025) -
Label-Free Topic-Focused Summarization Using Query Augmentation
von: Mu, Wenchuan, et al.
Veröffentlicht: (2024) -
OLAF: Towards Robust LLM-Based Annotation Framework in Empirical Software Engineering
von: Imran, Mia Mohammad, et al.
Veröffentlicht: (2025) -
Benchmarking Harmonized Tariff Schedule Classification Models
von: Judy, Bryce
Veröffentlicht: (2024)