Benchmarking Requirement-to-Architecture Generation with Hybrid Evaluation
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Minxiao, Yan, Shuying, Zhang, Li, Liu, Yang, Liu, Fang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
A System Model Generation Benchmark from Natural Language Requirements
di: Jin, Dongming, et al.
Pubblicazione: (2025)
di: Jin, Dongming, et al.
Pubblicazione: (2025)
CodeIF-Bench: Evaluating Instruction-Following Capabilities of Large Language Models in Interactive Code Generation
di: Wang, Peiding, et al.
Pubblicazione: (2025)
di: Wang, Peiding, et al.
Pubblicazione: (2025)
Incorporating Verification Standards for Security Requirements Generation from Functional Specifications
di: Lian, Xiaoli, et al.
Pubblicazione: (2025)
di: Lian, Xiaoli, et al.
Pubblicazione: (2025)
A Scalable Benchmark for Repository-Oriented Long-Horizon Conversational Context Management
di: Liu, Yang, et al.
Pubblicazione: (2026)
di: Liu, Yang, et al.
Pubblicazione: (2026)
MRG-Bench: Evaluating and Exploring the Requirements of Context for Repository-Level Code Generation
di: Li, Haiyang
Pubblicazione: (2025)
di: Li, Haiyang
Pubblicazione: (2025)
EfficientEdit: Accelerating Code Editing via Edit-Oriented Speculative Decoding
di: Wang, Peiding, et al.
Pubblicazione: (2025)
di: Wang, Peiding, et al.
Pubblicazione: (2025)
ReqElicitGym: An Evaluation Environment for Interview Competence in Conversational Requirements Elicitation
di: Jin, Dongming, et al.
Pubblicazione: (2026)
di: Jin, Dongming, et al.
Pubblicazione: (2026)
Bridging Requirements and Architecture: Multi-Agent Orchestration with External Knowledge and Hierarchical Memory
di: Li, Ruiyin, et al.
Pubblicazione: (2026)
di: Li, Ruiyin, et al.
Pubblicazione: (2026)
Benchmarking and Evaluating VLMs for Software Architecture Diagram Understanding
di: Ouyang, Shuyin, et al.
Pubblicazione: (2026)
di: Ouyang, Shuyin, et al.
Pubblicazione: (2026)
Requirements Development and Formalization for Reliable Code Generation: A Multi-Agent Vision
di: Lu, Xu, et al.
Pubblicazione: (2025)
di: Lu, Xu, et al.
Pubblicazione: (2025)
UserTrace: User-Level Requirements Generation and Traceability Recovery from Software Project Repositories
di: Jin, Dongming, et al.
Pubblicazione: (2025)
di: Jin, Dongming, et al.
Pubblicazione: (2025)
Are They All Good? Evaluating the Quality of CoTs in LLM-based Code Generation
di: Zhang, Binquan, et al.
Pubblicazione: (2025)
di: Zhang, Binquan, et al.
Pubblicazione: (2025)
RealBench: A Repo-Level Code Generation Benchmark Aligned with Real-World Software Development Practices
di: Li, Jia, et al.
Pubblicazione: (2026)
di: Li, Jia, et al.
Pubblicazione: (2026)
An Evaluation of Requirements Modeling for Cyber-Physical Systems via LLMs
di: Jin, Dongming, et al.
Pubblicazione: (2024)
di: Jin, Dongming, et al.
Pubblicazione: (2024)
REprompt: Prompt Generation for Intelligent Software Development Guided by Requirements Engineering
di: Shi, Junjie, et al.
Pubblicazione: (2026)
di: Shi, Junjie, et al.
Pubblicazione: (2026)
SolContractEval: A Benchmark for Evaluating Contract-Level Solidity Code Generation
di: Ye, Zhifan, et al.
Pubblicazione: (2025)
di: Ye, Zhifan, et al.
Pubblicazione: (2025)
Requirements for Active Assistance of Natural Questions in Software Architecture
di: Lemos, Diogo, et al.
Pubblicazione: (2025)
di: Lemos, Diogo, et al.
Pubblicazione: (2025)
RepoScope: Leveraging Call Chain-Aware Multi-View Context for Repository-Level Code Generation
di: Liu, Yang, et al.
Pubblicazione: (2025)
di: Liu, Yang, et al.
Pubblicazione: (2025)
Towards Requirements Engineering for GenAI-Enabled Software: Bridging Responsibility Gaps through Human Oversight Requirements
di: Mao, Zhenyu, et al.
Pubblicazione: (2025)
di: Mao, Zhenyu, et al.
Pubblicazione: (2025)
Bridging the Gap between User Intent and LLM: A Requirement Alignment Approach for Code Generation
di: Li, Jia, et al.
Pubblicazione: (2026)
di: Li, Jia, et al.
Pubblicazione: (2026)
From Chat to Interview: Agentic Requirements Elicitation with an Experience Ontology
di: Jin, Dongming, et al.
Pubblicazione: (2026)
di: Jin, Dongming, et al.
Pubblicazione: (2026)
Knowledge-Based Multi-Agent Framework for Automated Software Architecture Design
di: Zhang, Yiran, et al.
Pubblicazione: (2025)
di: Zhang, Yiran, et al.
Pubblicazione: (2025)
Towards Realistic Project-Level Code Generation via Multi-Agent Collaboration and Semantic Architecture Modeling
di: Zhao, Qianhui, et al.
Pubblicazione: (2025)
di: Zhao, Qianhui, et al.
Pubblicazione: (2025)
Requirements Volatility in Software Architecture Design: An Exploratory Case Study
di: Aaramaa, Sanja, et al.
Pubblicazione: (2026)
di: Aaramaa, Sanja, et al.
Pubblicazione: (2026)
Enhancing User-Feedback Driven Requirements Prioritization
di: Chattopadhyay, Aurek, et al.
Pubblicazione: (2026)
di: Chattopadhyay, Aurek, et al.
Pubblicazione: (2026)
AdaptiveLLM: A Framework for Selecting Optimal Cost-Efficient LLM for Code-Generation Based on CoT Length
di: Cheng, Junhang, et al.
Pubblicazione: (2025)
di: Cheng, Junhang, et al.
Pubblicazione: (2025)
ArchBench: Benchmarking Generative-AI for Software Architecture Tasks
di: Adnan, Bassam, et al.
Pubblicazione: (2026)
di: Adnan, Bassam, et al.
Pubblicazione: (2026)
SR-Eval: Evaluating LLMs on Code Generation under Stepwise Requirement Refinement
di: Zhan, Zexun, et al.
Pubblicazione: (2025)
di: Zhan, Zexun, et al.
Pubblicazione: (2025)
Usability as a Weapon: Attacking the Safety of LLM-Based Code Generation via Usability Requirements
di: Li, Yue, et al.
Pubblicazione: (2026)
di: Li, Yue, et al.
Pubblicazione: (2026)
A Study to Evaluate the Impact of LoRA Fine-tuning on the Performance of Non-functional Requirements Classification
di: Li, Xia, et al.
Pubblicazione: (2025)
di: Li, Xia, et al.
Pubblicazione: (2025)
LONGCODEU: Benchmarking Long-Context Language Models on Long Code Understanding
di: Li, Jia, et al.
Pubblicazione: (2025)
di: Li, Jia, et al.
Pubblicazione: (2025)
WebCoderBench: Benchmarking Web Application Generation with Comprehensive and Interpretable Evaluation Metrics
di: Liu, Chenxu, et al.
Pubblicazione: (2026)
di: Liu, Chenxu, et al.
Pubblicazione: (2026)
Generating Project-Specific Test Cases with Requirement Validation Intention
di: Qi, Binhang, et al.
Pubblicazione: (2025)
di: Qi, Binhang, et al.
Pubblicazione: (2025)
Requirements-Based Test Generation: A Comprehensive Survey
di: Yang, Zhenzhen, et al.
Pubblicazione: (2025)
di: Yang, Zhenzhen, et al.
Pubblicazione: (2025)
QUPER-MAn: Benchmark-Guided Target Setting for Maintainability Requirements
di: Borg, Markus, et al.
Pubblicazione: (2025)
di: Borg, Markus, et al.
Pubblicazione: (2025)
RobuNFR: Evaluating the Robustness of Large Language Models on Non-Functional Requirements Aware Code Generation
di: Lin, Feng, et al.
Pubblicazione: (2025)
di: Lin, Feng, et al.
Pubblicazione: (2025)
UniCoR: Modality Collaboration for Robust Cross-Language Hybrid Code Retrieval
di: Yang, Yang, et al.
Pubblicazione: (2025)
di: Yang, Yang, et al.
Pubblicazione: (2025)
ProjectEval: A Benchmark for Programming Agents Automated Evaluation on Project-Level Code Generation
di: Liu, Kaiyuan, et al.
Pubblicazione: (2025)
di: Liu, Kaiyuan, et al.
Pubblicazione: (2025)
A Benchmark for Evaluating Repository-Level Code Agents with Intermediate Reasoning on Feature Addition Task
di: Liu, Shuhan, et al.
Pubblicazione: (2026)
di: Liu, Shuhan, et al.
Pubblicazione: (2026)
Assessing the Impact of Requirement Ambiguity on LLM-based Function-Level Code Generation
di: Yang, Di, et al.
Pubblicazione: (2026)
di: Yang, Di, et al.
Pubblicazione: (2026)
Documenti analoghi
-
A System Model Generation Benchmark from Natural Language Requirements
di: Jin, Dongming, et al.
Pubblicazione: (2025) -
CodeIF-Bench: Evaluating Instruction-Following Capabilities of Large Language Models in Interactive Code Generation
di: Wang, Peiding, et al.
Pubblicazione: (2025) -
Incorporating Verification Standards for Security Requirements Generation from Functional Specifications
di: Lian, Xiaoli, et al.
Pubblicazione: (2025) -
A Scalable Benchmark for Repository-Oriented Long-Horizon Conversational Context Management
di: Liu, Yang, et al.
Pubblicazione: (2026) -
MRG-Bench: Evaluating and Exploring the Requirements of Context for Repository-Level Code Generation
di: Li, Haiyang
Pubblicazione: (2025)