Salvato in:
| Autori principali: | Mu, Wenchuan, Lim, Kwan Hui |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2404.16457 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Towards Reliable Evaluation of Neural Program Repair with Natural Robustness Testing
di: Le-Cong, Thanh, et al.
Pubblicazione: (2024)
di: Le-Cong, Thanh, et al.
Pubblicazione: (2024)
Towards Better Correctness and Efficiency in Code Generation
di: Feng, Yunlong, et al.
Pubblicazione: (2025)
di: Feng, Yunlong, et al.
Pubblicazione: (2025)
Label-Free Topic-Focused Summarization Using Query Augmentation
di: Mu, Wenchuan, et al.
Pubblicazione: (2024)
di: Mu, Wenchuan, et al.
Pubblicazione: (2024)
OLAF: Towards Robust LLM-Based Annotation Framework in Empirical Software Engineering
di: Imran, Mia Mohammad, et al.
Pubblicazione: (2025)
di: Imran, Mia Mohammad, et al.
Pubblicazione: (2025)
Benchmarking Harmonized Tariff Schedule Classification Models
di: Judy, Bryce
Pubblicazione: (2024)
di: Judy, Bryce
Pubblicazione: (2024)
Model Provenance via Model DNA
di: Mu, Xin, et al.
Pubblicazione: (2023)
di: Mu, Xin, et al.
Pubblicazione: (2023)
Towards Robust LLM Post-Training: Automatic Failure Management for Reinforcement Fine-Tuning
di: Zhang, Lingzhe, et al.
Pubblicazione: (2026)
di: Zhang, Lingzhe, et al.
Pubblicazione: (2026)
Human-Machine Co-Boosted Bug Report Identification with Mutualistic Neural Active Learning
di: Long, Guoming, et al.
Pubblicazione: (2026)
di: Long, Guoming, et al.
Pubblicazione: (2026)
Towards a Classification of Open-Source ML Models and Datasets for Software Engineering
di: González, Alexandra, et al.
Pubblicazione: (2024)
di: González, Alexandra, et al.
Pubblicazione: (2024)
Beyond Retrieval: A Multitask Benchmark and Model for Code Search
di: Xue, Siqiao, et al.
Pubblicazione: (2026)
di: Xue, Siqiao, et al.
Pubblicazione: (2026)
Conventional Commit Classification using Large Language Models and Prompt Engineering
di: Quadir, H. M. Sazzad, et al.
Pubblicazione: (2026)
di: Quadir, H. M. Sazzad, et al.
Pubblicazione: (2026)
CodeFort: Robust Training for Code Generation Models
di: Zhang, Yuhao, et al.
Pubblicazione: (2024)
di: Zhang, Yuhao, et al.
Pubblicazione: (2024)
Exploring the Potential of Large Language Models in Fine-Grained Review Comment Classification
di: Nguyen, Linh, et al.
Pubblicazione: (2025)
di: Nguyen, Linh, et al.
Pubblicazione: (2025)
DRAGON: Robust Classification for Very Large Collections of Software Repositories
di: Balla, Stefano, et al.
Pubblicazione: (2026)
di: Balla, Stefano, et al.
Pubblicazione: (2026)
Towards a General Framework for HTN Modeling with LLMs
di: Puerta-Merino, Israel, et al.
Pubblicazione: (2025)
di: Puerta-Merino, Israel, et al.
Pubblicazione: (2025)
Towards Leveraging Large Language Model Summaries for Topic Modeling in Source Code
di: Carissimi, Michele, et al.
Pubblicazione: (2025)
di: Carissimi, Michele, et al.
Pubblicazione: (2025)
Unveiling Project-Specific Bias in Neural Code Models
di: Li, Zhiming, et al.
Pubblicazione: (2022)
di: Li, Zhiming, et al.
Pubblicazione: (2022)
Precision in Practice: Knowledge Guided Code Summarizing Grounded in Industrial Expectations
di: Li, Jintai, et al.
Pubblicazione: (2026)
di: Li, Jintai, et al.
Pubblicazione: (2026)
Towards a Neural Debugger for Python
di: Beck, Maximilian, et al.
Pubblicazione: (2026)
di: Beck, Maximilian, et al.
Pubblicazione: (2026)
Enhancing Deployment-Time Predictive Model Robustness for Code Analysis and Optimization
di: Wang, Huanting, et al.
Pubblicazione: (2024)
di: Wang, Huanting, et al.
Pubblicazione: (2024)
Post-Incorporating Code Structural Knowledge into Pretrained Models via ICL for Code Translation
di: Du, Yali, et al.
Pubblicazione: (2025)
di: Du, Yali, et al.
Pubblicazione: (2025)
Towards a Domain-Specific Modelling Environment for Reinforcement Learning
di: Sinani, Natalie, et al.
Pubblicazione: (2024)
di: Sinani, Natalie, et al.
Pubblicazione: (2024)
CodeSSM: Towards State Space Models for Code Understanding
di: Verma, Shweta, et al.
Pubblicazione: (2025)
di: Verma, Shweta, et al.
Pubblicazione: (2025)
Are Large Language Models Robust in Understanding Code Against Semantics-Preserving Mutations?
di: Orvalho, Pedro, et al.
Pubblicazione: (2025)
di: Orvalho, Pedro, et al.
Pubblicazione: (2025)
Towards Better Code Understanding in Decoder-Only Models with Contrastive Learning
di: Lin, Jiayi, et al.
Pubblicazione: (2024)
di: Lin, Jiayi, et al.
Pubblicazione: (2024)
Towards a Digital Twin Modeling Method for Container Terminal Port
di: Hakimi, Faouzi, et al.
Pubblicazione: (2025)
di: Hakimi, Faouzi, et al.
Pubblicazione: (2025)
An Experience Report on Regression-Free Repair of Deep Neural Network Model
di: Nakagawa, Takao, et al.
Pubblicazione: (2025)
di: Nakagawa, Takao, et al.
Pubblicazione: (2025)
Robustness and Reasoning Fidelity of Large Language Models in Long-Context Code Question Answering
di: Maharaj, Kishan, et al.
Pubblicazione: (2026)
di: Maharaj, Kishan, et al.
Pubblicazione: (2026)
A-ProS: Towards Reliable Autonomous Programming Through Multi-Model Feedback
di: Tabassum, Anika, et al.
Pubblicazione: (2026)
di: Tabassum, Anika, et al.
Pubblicazione: (2026)
Rethinking Scientific Modeling: Toward Physically Consistent and Simulation-Executable Programmatic Generation
di: Jiang, Yongqing, et al.
Pubblicazione: (2026)
di: Jiang, Yongqing, et al.
Pubblicazione: (2026)
WebCompass: Towards Multimodal Web Coding Evaluation for Code Language Models
di: Lei, Xinping, et al.
Pubblicazione: (2026)
di: Lei, Xinping, et al.
Pubblicazione: (2026)
Towards Advancing Code Generation with Large Language Models: A Research Roadmap
di: Jin, Haolin, et al.
Pubblicazione: (2025)
di: Jin, Haolin, et al.
Pubblicazione: (2025)
Toward a Theory of Causation for Interpreting Neural Code Models
di: Palacio, David N., et al.
Pubblicazione: (2023)
di: Palacio, David N., et al.
Pubblicazione: (2023)
PACIFIC: a framework for generating benchmarks to check Precise Automatically Checked Instruction Following In Code
di: Dreyfuss, Itay, et al.
Pubblicazione: (2025)
di: Dreyfuss, Itay, et al.
Pubblicazione: (2025)
Towards Robust Agentic CUDA Kernel Benchmarking, Verification, and Optimization
di: Lange, Robert Tjarko, et al.
Pubblicazione: (2025)
di: Lange, Robert Tjarko, et al.
Pubblicazione: (2025)
Toward Automated Validation of Language Model Synthesized Test Cases using Semantic Entropy
di: Taherkhani, Hamed, et al.
Pubblicazione: (2024)
di: Taherkhani, Hamed, et al.
Pubblicazione: (2024)
SWE-Compass: Towards Unified Evaluation of Agentic Coding Abilities for Large Language Models
di: Xu, Jingxuan, et al.
Pubblicazione: (2025)
di: Xu, Jingxuan, et al.
Pubblicazione: (2025)
RobuNFR: Evaluating the Robustness of Large Language Models on Non-Functional Requirements Aware Code Generation
di: Lin, Feng, et al.
Pubblicazione: (2025)
di: Lin, Feng, et al.
Pubblicazione: (2025)
When Prompts Go Wrong: Evaluating Code Model Robustness to Ambiguous, Contradictory, and Incomplete Task Descriptions
di: Larbi, Maya, et al.
Pubblicazione: (2025)
di: Larbi, Maya, et al.
Pubblicazione: (2025)
The Complexity Trap: Simple Observation Masking Is as Efficient as LLM Summarization for Agent Context Management
di: Lindenbauer, Tobias, et al.
Pubblicazione: (2025)
di: Lindenbauer, Tobias, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Towards Reliable Evaluation of Neural Program Repair with Natural Robustness Testing
di: Le-Cong, Thanh, et al.
Pubblicazione: (2024) -
Towards Better Correctness and Efficiency in Code Generation
di: Feng, Yunlong, et al.
Pubblicazione: (2025) -
Label-Free Topic-Focused Summarization Using Query Augmentation
di: Mu, Wenchuan, et al.
Pubblicazione: (2024) -
OLAF: Towards Robust LLM-Based Annotation Framework in Empirical Software Engineering
di: Imran, Mia Mohammad, et al.
Pubblicazione: (2025) -
Benchmarking Harmonized Tariff Schedule Classification Models
di: Judy, Bryce
Pubblicazione: (2024)