LLMs taking shortcuts in test generation: A study with SAP HANA and LevelDB
Fuente:
arXiv
Saved in:
| Main Authors: | Bekmyradov, Vekil, Pütz, Noah C., Bartz-Beielstein, Thomas |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
The Transformative Influence of LLMs on Software Development & Developer Productivity
by: Jalil, Sajed
Published: (2023)
by: Jalil, Sajed
Published: (2023)
MedMemoryBench: Benchmarking Agent Memory in Personalized Healthcare
by: Wang, Yihao, et al.
Published: (2026)
by: Wang, Yihao, et al.
Published: (2026)
Collaborative LLM Agents for C4 Software Architecture Design Automation
by: Szczepanik, Kamil, et al.
Published: (2025)
by: Szczepanik, Kamil, et al.
Published: (2025)
Large Language Models in Software Documentation and Modeling: A Literature Review and Findings
by: Radosky, Lukas, et al.
Published: (2026)
by: Radosky, Lukas, et al.
Published: (2026)
DRS-OSS: Practical Diff Risk Scoring with LLMs
by: Sayedsalehi, Ali, et al.
Published: (2025)
by: Sayedsalehi, Ali, et al.
Published: (2025)
Thinking Machines: Mathematical Reasoning in the Age of LLMs
by: Asperti, Andrea, et al.
Published: (2025)
by: Asperti, Andrea, et al.
Published: (2025)
Navigating WebAI: Training Agents to Complete Web Tasks with Large Language Models and Reinforcement Learning
by: Thil, Lucas-Andreï, et al.
Published: (2024)
by: Thil, Lucas-Andreï, et al.
Published: (2024)
SLEAN: Simple Lightweight Ensemble Analysis Network for Multi-Provider LLM Coordination: Design, Implementation, and Vibe Coding Bug Investigation Case Study
by: Vargas, Matheus J. T.
Published: (2025)
by: Vargas, Matheus J. T.
Published: (2025)
The Curious Case of In-Training Compression of State Space Models
by: Chahine, Makram, et al.
Published: (2025)
by: Chahine, Makram, et al.
Published: (2025)
NeuronSpark: A Spiking Neural Network Language Model with Selective State Space Dynamics
by: Tang, Zhengzheng
Published: (2026)
by: Tang, Zhengzheng
Published: (2026)
PerkwE_COQA: Enhanced Persian Conversational Question Answering by combining contextual keyword extraction with Large Language Models
by: Moradbeiki, Pardis, et al.
Published: (2024)
by: Moradbeiki, Pardis, et al.
Published: (2024)
AI Agents-as-Judge: Automated Assessment of Accuracy, Consistency, Completeness and Clarity for Enterprise Documents
by: Dasgupta, Sudip, et al.
Published: (2025)
by: Dasgupta, Sudip, et al.
Published: (2025)
elsciRL: Integrating Language Solutions into Reinforcement Learning Problem Settings
by: Osborne, Philip, et al.
Published: (2025)
by: Osborne, Philip, et al.
Published: (2025)
Reference-Guided Verdict: LLMs-as-Judges in Automatic Evaluation of Free-Form QA
by: Badshah, Sher, et al.
Published: (2024)
by: Badshah, Sher, et al.
Published: (2024)
VeriPlan: Integrating Formal Verification and LLMs into End-User Planning
by: Lee, Christine, et al.
Published: (2025)
by: Lee, Christine, et al.
Published: (2025)
When Does Content-Based Routing Work? Representation Requirements for Selective Attention in Hybrid Sequence Models
by: Basu, Abhinaba
Published: (2026)
by: Basu, Abhinaba
Published: (2026)
MAcPNN: Mutual Assisted Learning on Data Streams with Temporal Dependence
by: Giannini, Federico, et al.
Published: (2026)
by: Giannini, Federico, et al.
Published: (2026)
NOTAI.AI: Explainable Detection of Machine-Generated Text via Curvature and Feature Attribution
by: Breneur, Oleksandr Marchenko, et al.
Published: (2026)
by: Breneur, Oleksandr Marchenko, et al.
Published: (2026)
Rethinking the Multilingual Reasoning Gap with Layer Swap
by: Lasbordes, Maxence, et al.
Published: (2026)
by: Lasbordes, Maxence, et al.
Published: (2026)
Natural Language Summarization Enables Multi-Repository Bug Localization by LLMs in Microservice Architectures
by: Oskooei, Amirkia Rafiei, et al.
Published: (2025)
by: Oskooei, Amirkia Rafiei, et al.
Published: (2025)
RelRepair: Enhancing Automated Program Repair by Retrieving Relevant Code
by: Liu, Shunyu, et al.
Published: (2025)
by: Liu, Shunyu, et al.
Published: (2025)
Constitution or Collapse? Exploring Constitutional AI with Llama 3-8B
by: Zhang, Xue
Published: (2025)
by: Zhang, Xue
Published: (2025)
FREYR: A Framework for Recognizing and Executing Your Requests
by: Gallotta, Roberto, et al.
Published: (2025)
by: Gallotta, Roberto, et al.
Published: (2025)
Comprehensive Evaluation and Insights into the Use of Large Language Models in the Automation of Behavior-Driven Development Acceptance Test Formulation
by: Karpurapu, Shanthi, et al.
Published: (2024)
by: Karpurapu, Shanthi, et al.
Published: (2024)
Towards Human-AI Synergy in Requirements Engineering: A Framework and Preliminary Study
by: Abbasi, Mateen Ahmed, et al.
Published: (2025)
by: Abbasi, Mateen Ahmed, et al.
Published: (2025)
DeepSeek-V3, GPT-4, Phi-4, and LLaMA-3.3 generate correct code for LoRaWAN-related engineering tasks
by: Fernandes, Daniel, et al.
Published: (2025)
by: Fernandes, Daniel, et al.
Published: (2025)
An Epidemiological Knowledge Graph extracted from the World Health Organization's Disease Outbreak News
by: Consoli, Sergio, et al.
Published: (2025)
by: Consoli, Sergio, et al.
Published: (2025)
Semantic Modeling for World-Centered Architectures
by: Mantsivoda, Andrei, et al.
Published: (2026)
by: Mantsivoda, Andrei, et al.
Published: (2026)
Bug In the Code Stack: Can LLMs Find Bugs in Large Python Code Stacks
by: Lee, Hokyung, et al.
Published: (2024)
by: Lee, Hokyung, et al.
Published: (2024)
BudgetMLAgent: A Cost-Effective LLM Multi-Agent system for Automating Machine Learning Tasks
by: Gandhi, Shubham, et al.
Published: (2024)
by: Gandhi, Shubham, et al.
Published: (2024)
Harnessing non-adversarial robustness in large language models
by: Zhou, Qinghua, et al.
Published: (2026)
by: Zhou, Qinghua, et al.
Published: (2026)
Causal Dimensionality of Transformer Representations: Measurement, Scaling, and Layer Structure
by: Sarkar, Nilesh, et al.
Published: (2026)
by: Sarkar, Nilesh, et al.
Published: (2026)
DeepPersona: A Generative Engine for Scaling Deep Synthetic Personas
by: Wang, Zhen, et al.
Published: (2025)
by: Wang, Zhen, et al.
Published: (2025)
Anonymization-Enhanced Privacy Protection for Mobile GUI Agents: Available but Invisible
by: Zhao, Lepeng, et al.
Published: (2026)
by: Zhao, Lepeng, et al.
Published: (2026)
TRIZ Agents: A Multi-Agent LLM Approach for TRIZ-Based Innovation
by: Szczepanik, Kamil, et al.
Published: (2025)
by: Szczepanik, Kamil, et al.
Published: (2025)
CLMN: Concept based Language Models via Neural Symbolic Reasoning
by: Yang, Yibo
Published: (2025)
by: Yang, Yibo
Published: (2025)
QHackBench: Benchmarking Large Language Models for Quantum Code Generation Using PennyLane Hackathon Challenges
by: Basit, Abdul, et al.
Published: (2025)
by: Basit, Abdul, et al.
Published: (2025)
Inference acceleration for large language models using "stairs" assisted greedy generation
by: Grigaliūnas, Domas, et al.
Published: (2024)
by: Grigaliūnas, Domas, et al.
Published: (2024)
PennyCoder: Efficient Domain-Specific LLMs for PennyLane-Based Quantum Code Generation
by: Basit, Abdul, et al.
Published: (2025)
by: Basit, Abdul, et al.
Published: (2025)
Optimizing Large Language Models for OpenAPI Code Completion
by: Petryshyn, Bohdan, et al.
Published: (2024)
by: Petryshyn, Bohdan, et al.
Published: (2024)
Similar Items
-
The Transformative Influence of LLMs on Software Development & Developer Productivity
by: Jalil, Sajed
Published: (2023) -
MedMemoryBench: Benchmarking Agent Memory in Personalized Healthcare
by: Wang, Yihao, et al.
Published: (2026) -
Collaborative LLM Agents for C4 Software Architecture Design Automation
by: Szczepanik, Kamil, et al.
Published: (2025) -
Large Language Models in Software Documentation and Modeling: A Literature Review and Findings
by: Radosky, Lukas, et al.
Published: (2026) -
DRS-OSS: Practical Diff Risk Scoring with LLMs
by: Sayedsalehi, Ali, et al.
Published: (2025)