Interactive Evaluation of Large Language Models for Multi-Requirement Software Engineering Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Rontogiannis, Dimitrios, Peyrard, Maxime, Baldwin, Nicolas, Josifoski, Martin, West, Robert, Gunopulos, Dimitrios |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Grammar-Constrained Decoding for Structured NLP Tasks without Finetuning
by: Geng, Saibo, et al.
Published: (2023)
by: Geng, Saibo, et al.
Published: (2023)
Agentic AI: The Era of Semantic Decoding
by: Peyrard, Maxime, et al.
Published: (2024)
by: Peyrard, Maxime, et al.
Published: (2024)
Symbolic Autoencoding for Self-Supervised Sequence Learning
by: Amani, Mohammad Hossein, et al.
Published: (2024)
by: Amani, Mohammad Hossein, et al.
Published: (2024)
Evaluating Language Model Agency through Negotiations
by: Davidson, Tim R., et al.
Published: (2024)
by: Davidson, Tim R., et al.
Published: (2024)
A Glitch in the Matrix? Locating and Detecting Language Model Grounding with Fakepedia
by: Monea, Giovanni, et al.
Published: (2023)
by: Monea, Giovanni, et al.
Published: (2023)
A Framework for Feasible Counterfactual Exploration incorporating Causality, Sparsity and Density
by: Markou, Kleopatra, et al.
Published: (2024)
by: Markou, Kleopatra, et al.
Published: (2024)
Flows: Building Blocks of Reasoning and Collaborating AI
by: Josifoski, Martin, et al.
Published: (2023)
by: Josifoski, Martin, et al.
Published: (2023)
What Really Controls Temporal Reasoning in Large Language Models: Tokenisation or Representation of Time?
by: Bhatia, Gagan, et al.
Published: (2026)
by: Bhatia, Gagan, et al.
Published: (2026)
Interpretability-by-Design with Accurate Locally Additive Models and Conditional Feature Effects
by: Gkolemis, Vasilis, et al.
Published: (2026)
by: Gkolemis, Vasilis, et al.
Published: (2026)
Meta-Statistical Learning: Supervised Learning of Statistical Estimators
by: Peyrard, Maxime, et al.
Published: (2025)
by: Peyrard, Maxime, et al.
Published: (2025)
The Dead Salmons of AI Interpretability
by: Méloux, Maxime, et al.
Published: (2025)
by: Méloux, Maxime, et al.
Published: (2025)
Mechanistic Interpretability as Statistical Estimation: A Variance Analysis
by: Méloux, Maxime, et al.
Published: (2025)
by: Méloux, Maxime, et al.
Published: (2025)
Test-Time Compute Games
by: Velasco, Ander Artola, et al.
Published: (2026)
by: Velasco, Ander Artola, et al.
Published: (2026)
Date Fragments: A Hidden Bottleneck of Tokenization for Temporal Reasoning
by: Bhatia, Gagan, et al.
Published: (2025)
by: Bhatia, Gagan, et al.
Published: (2025)
Toward Explaining Large Language Models in Software Engineering Tasks
by: Vitale, Antonio, et al.
Published: (2025)
by: Vitale, Antonio, et al.
Published: (2025)
Everything, Everywhere, All at Once: Is Mechanistic Interpretability Identifiable?
by: Méloux, Maxime, et al.
Published: (2025)
by: Méloux, Maxime, et al.
Published: (2025)
Automated Non-Functional Requirements Generation in Software Engineering with Large Language Models: A Comparative Study
by: Almonte, Jomar Thomas, et al.
Published: (2025)
by: Almonte, Jomar Thomas, et al.
Published: (2025)
Engineering Safety Requirements for Autonomous Driving with Large Language Models
by: Nouri, Ali, et al.
Published: (2024)
by: Nouri, Ali, et al.
Published: (2024)
PickLLM: Context-Aware RL-Assisted Large Language Model Routing
by: Sikeridis, Dimitrios, et al.
Published: (2024)
by: Sikeridis, Dimitrios, et al.
Published: (2024)
SWE-MERA: A Dynamic Benchmark for Agenticly Evaluating Large Language Models on Software Engineering Tasks
by: Adamenko, Pavel, et al.
Published: (2025)
by: Adamenko, Pavel, et al.
Published: (2025)
Evaluating Large Language Models for Real-World Engineering Tasks
by: Heesch, Rene, et al.
Published: (2025)
by: Heesch, Rene, et al.
Published: (2025)
Hands-on Evaluation of Visual Transformers for Object Recognition and Detection
by: Vlachogiannis, Dimitrios N., et al.
Published: (2025)
by: Vlachogiannis, Dimitrios N., et al.
Published: (2025)
Large Language Model Partitioning for Low-Latency Inference at the Edge
by: Kafetzis, Dimitrios, et al.
Published: (2025)
by: Kafetzis, Dimitrios, et al.
Published: (2025)
Trustworthy Scheduling for Big Data Applications
by: Tomaras, Dimitrios, et al.
Published: (2026)
by: Tomaras, Dimitrios, et al.
Published: (2026)
SEER: Sustainability Enhanced Engineering of Software Requirements
by: Roy, Mandira, et al.
Published: (2025)
by: Roy, Mandira, et al.
Published: (2025)
Geo-OLM: Enabling Sustainable Earth Observation Studies with Cost-Efficient Open Language Models & State-Driven Workflows
by: Stamoulis, Dimitrios, et al.
Published: (2025)
by: Stamoulis, Dimitrios, et al.
Published: (2025)
CarbonCall: Sustainability-Aware Function Calling for Large Language Models on Edge Devices
by: Paramanayakam, Varatheepan, et al.
Published: (2025)
by: Paramanayakam, Varatheepan, et al.
Published: (2025)
Towards Intent-Based Network Management: Large Language Models for Intent Extraction in 5G Core Networks
by: Manias, Dimitrios Michael, et al.
Published: (2024)
by: Manias, Dimitrios Michael, et al.
Published: (2024)
An Evaluation of Large Language Models on Text Summarization Tasks Using Prompt Engineering Techniques
by: Aly, Walid Mohamed, et al.
Published: (2025)
by: Aly, Walid Mohamed, et al.
Published: (2025)
ArchCode: Incorporating Software Requirements in Code Generation with Large Language Models
by: Han, Hojae, et al.
Published: (2024)
by: Han, Hojae, et al.
Published: (2024)
At First Sight: Zero-Shot Classification of Astronomical Images with Large Multimodal Models
by: Tanoglidis, Dimitrios, et al.
Published: (2024)
by: Tanoglidis, Dimitrios, et al.
Published: (2024)
GRAD: Generative Retrieval-Aligned Demonstration Sampler for Efficient Few-Shot Reasoning
by: Gabouj, Oussama, et al.
Published: (2025)
by: Gabouj, Oussama, et al.
Published: (2025)
From What to How: Bridging User Requirements with Software Development Using Large Language Models
by: He, Xiao, et al.
Published: (2026)
by: He, Xiao, et al.
Published: (2026)
A Scalable Multi-GPU Framework for Encrypted Large-Model Inference
by: Jayashankar, Siddharth, et al.
Published: (2025)
by: Jayashankar, Siddharth, et al.
Published: (2025)
CORE: Full-Path Evaluation of LLM Agents Beyond Final State
by: Michelakis, Panagiotis, et al.
Published: (2025)
by: Michelakis, Panagiotis, et al.
Published: (2025)
Intent-Driven Smart Manufacturing Integrating Knowledge Graphs and Large Language Models
by: Jradi, Takoua, et al.
Published: (2026)
by: Jradi, Takoua, et al.
Published: (2026)
RepoDebug: Repository-Level Multi-Task and Multi-Language Debugging Evaluation of Large Language Models
by: Liu, Jingjing, et al.
Published: (2025)
by: Liu, Jingjing, et al.
Published: (2025)
Psychometric Predictive Power of Large Language Models
by: Kuribayashi, Tatsuki, et al.
Published: (2023)
by: Kuribayashi, Tatsuki, et al.
Published: (2023)
Finding the Subjective Truth: Collecting 2 Million Votes for Comprehensive Gen-AI Model Evaluation
by: Christodoulou, Dimitrios, et al.
Published: (2024)
by: Christodoulou, Dimitrios, et al.
Published: (2024)
Large Language Models for Software Engineering: A Systematic Literature Review
by: Hou, Xinyi, et al.
Published: (2023)
by: Hou, Xinyi, et al.
Published: (2023)
Similar Items
-
Grammar-Constrained Decoding for Structured NLP Tasks without Finetuning
by: Geng, Saibo, et al.
Published: (2023) -
Agentic AI: The Era of Semantic Decoding
by: Peyrard, Maxime, et al.
Published: (2024) -
Symbolic Autoencoding for Self-Supervised Sequence Learning
by: Amani, Mohammad Hossein, et al.
Published: (2024) -
Evaluating Language Model Agency through Negotiations
by: Davidson, Tim R., et al.
Published: (2024) -
A Glitch in the Matrix? Locating and Detecting Language Model Grounding with Fakepedia
by: Monea, Giovanni, et al.
Published: (2023)