LaQual: A Novel Framework for Automated Evaluation of LLM App Quality
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Yan, Hou, Xinyi, Zhao, Yanjie, Lin, Weiguo, Wang, Haoyu, Si, Junjun |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLM Applications: Current Paradigms and the Next Frontier
by: Hou, Xinyi, et al.
Published: (2025)
by: Hou, Xinyi, et al.
Published: (2025)
Unveiling the Landscape of LLM Deployment in the Wild: An Empirical Study
by: Hou, Xinyi, et al.
Published: (2025)
by: Hou, Xinyi, et al.
Published: (2025)
LLM App Store Analysis: A Vision and Roadmap
by: Zhao, Yanjie, et al.
Published: (2024)
by: Zhao, Yanjie, et al.
Published: (2024)
Software Engineering for Large Language Models: Research Status, Challenges and the Road Ahead
by: Rao, Hongzhou, et al.
Published: (2025)
by: Rao, Hongzhou, et al.
Published: (2025)
Gradient-Based Model Fingerprinting for LLM Similarity Detection and Family Classification
by: Wu, Zehao, et al.
Published: (2025)
by: Wu, Zehao, et al.
Published: (2025)
Large Language Models for Software Engineering: A Systematic Literature Review
by: Hou, Xinyi, et al.
Published: (2023)
by: Hou, Xinyi, et al.
Published: (2023)
Unsafe by Flow: Uncovering Bidirectional Data-Flow Risks in MCP Ecosystem
by: Hou, Xinyi, et al.
Published: (2026)
by: Hou, Xinyi, et al.
Published: (2026)
Voices from the Frontier: A Comprehensive Analysis of the OpenAI Developer Forum
by: Hou, Xinyi, et al.
Published: (2024)
by: Hou, Xinyi, et al.
Published: (2024)
An LLM-based Quantitative Framework for Evaluating High-Stealthy Backdoor Risks in OSS Supply Chains
by: Yan, Zihe, et al.
Published: (2025)
by: Yan, Zihe, et al.
Published: (2025)
GPTZoo: A Large-scale Dataset of GPTs for the Research Community
by: Hou, Xinyi, et al.
Published: (2024)
by: Hou, Xinyi, et al.
Published: (2024)
Large Language Model Supply Chain: A Research Agenda
by: Wang, Shenao, et al.
Published: (2024)
by: Wang, Shenao, et al.
Published: (2024)
Z-Space: A Multi-Agent Tool Orchestration Framework for Enterprise-Grade LLM Automation
by: He, Qingsong, et al.
Published: (2025)
by: He, Qingsong, et al.
Published: (2025)
On the (In)Security of LLM App Stores
by: Hou, Xinyi, et al.
Published: (2024)
by: Hou, Xinyi, et al.
Published: (2024)
Bridging Design and Development with Automated Declarative UI Code Generation
by: Zhou, Ting, et al.
Published: (2024)
by: Zhou, Ting, et al.
Published: (2024)
Skeleton-Guided-Translation: A Benchmarking Framework for Code Repository Translation with Fine-Grained Quality Evaluation
by: Zhang, Xing, et al.
Published: (2025)
by: Zhang, Xing, et al.
Published: (2025)
SWE-Factory: Your Automated Factory for Issue Resolution Training Data and Evaluation Benchmarks
by: Guo, Lianghong, et al.
Published: (2025)
by: Guo, Lianghong, et al.
Published: (2025)
Automated Validation of LLM-based Evaluators for Software Engineering Artifacts
by: Fandina, Ora Nova, et al.
Published: (2025)
by: Fandina, Ora Nova, et al.
Published: (2025)
Fine-grained Approaches for Confidence Calibration of LLMs in Automated Code Revision
by: Lin, Hong Yi, et al.
Published: (2026)
by: Lin, Hong Yi, et al.
Published: (2026)
Getting Inspiration for Feature Elicitation: App Store- vs. LLM-based Approach
by: Wei, Jialiang, et al.
Published: (2024)
by: Wei, Jialiang, et al.
Published: (2024)
ProbGuard: Probabilistic Runtime Monitoring for LLM Agent Safety
by: Wang, Haoyu, et al.
Published: (2025)
by: Wang, Haoyu, et al.
Published: (2025)
AgenticTCAD: A LLM-based Multi-Agent Framework for Automated TCAD Code Generation and Device Optimization
by: Fan, Guangxi, et al.
Published: (2025)
by: Fan, Guangxi, et al.
Published: (2025)
Automated Self-Testing as a Quality Gate: Evidence-Driven Release Management for LLM Applications
by: Maiorano, Alexandre Cristovão
Published: (2026)
by: Maiorano, Alexandre Cristovão
Published: (2026)
Advancing Automated Ethical Profiling in SE: a Zero-Shot Evaluation of LLM Reasoning
by: Migliarini, Patrizio, et al.
Published: (2025)
by: Migliarini, Patrizio, et al.
Published: (2025)
WebApp1K: A Practical Code-Generation Benchmark for Web App Development
by: Cui, Yi
Published: (2024)
by: Cui, Yi
Published: (2024)
A Dual-Loop Agent Framework for Automated Vulnerability Reproduction
by: Liu, Bin, et al.
Published: (2026)
by: Liu, Bin, et al.
Published: (2026)
BiasScope: Towards Automated Detection of Bias in LLM-as-a-Judge Evaluation
by: Lai, Peng, et al.
Published: (2026)
by: Lai, Peng, et al.
Published: (2026)
Vul-R2: A Reasoning LLM for Automated Vulnerability Repair
by: Wen, Xin-Cheng, et al.
Published: (2025)
by: Wen, Xin-Cheng, et al.
Published: (2025)
Skeet: Towards a Lightweight Serverless Framework Supporting Modern AI-Driven App Development
by: Fumitake, Kawasaki, et al.
Published: (2024)
by: Fumitake, Kawasaki, et al.
Published: (2024)
app.build: A Production Framework for Scaling Agentic Prompt-to-App Generation with Environment Scaffolding
by: Kniazev, Evgenii, et al.
Published: (2025)
by: Kniazev, Evgenii, et al.
Published: (2025)
LLM Agents for Automated Web Vulnerability Reproduction: Are We There Yet?
by: Liu, Bin, et al.
Published: (2025)
by: Liu, Bin, et al.
Published: (2025)
AutoDroid: LLM-powered Task Automation in Android
by: Wen, Hao, et al.
Published: (2023)
by: Wen, Hao, et al.
Published: (2023)
AppBench: Planning of Multiple APIs from Various APPs for Complex User Instruction
by: Wang, Hongru, et al.
Published: (2024)
by: Wang, Hongru, et al.
Published: (2024)
EvolveTool-Bench: Evaluating the Quality of LLM-Generated Tool Libraries as Software Artifacts
by: Kaliyev, Alibek T., et al.
Published: (2026)
by: Kaliyev, Alibek T., et al.
Published: (2026)
Evolving Excellence: Automated Optimization of LLM-based Agents
by: Brookes, Paul, et al.
Published: (2025)
by: Brookes, Paul, et al.
Published: (2025)
An Empirical Study of Agent Developer Practices in AI Agent Frameworks
by: Wang, Yanlin, et al.
Published: (2025)
by: Wang, Yanlin, et al.
Published: (2025)
AXIOM: Benchmarking LLM-as-a-Judge for Code via Rule-Based Perturbation and Multisource Quality Calibration
by: Wang, Ruiqi, et al.
Published: (2025)
by: Wang, Ruiqi, et al.
Published: (2025)
Toward Understanding Bugs in Vector Database Management Systems
by: Xie, Yinglin, et al.
Published: (2025)
by: Xie, Yinglin, et al.
Published: (2025)
ToolFactory: Automating Tool Generation by Leveraging LLM to Understand REST API Documentations
by: Ni, Xinyi, et al.
Published: (2025)
by: Ni, Xinyi, et al.
Published: (2025)
OrcaLoca: An LLM Agent Framework for Software Issue Localization
by: Yu, Zhongming, et al.
Published: (2025)
by: Yu, Zhongming, et al.
Published: (2025)
GPT Store Mining and Analysis
by: Su, Dongxun, et al.
Published: (2024)
by: Su, Dongxun, et al.
Published: (2024)
Similar Items
-
LLM Applications: Current Paradigms and the Next Frontier
by: Hou, Xinyi, et al.
Published: (2025) -
Unveiling the Landscape of LLM Deployment in the Wild: An Empirical Study
by: Hou, Xinyi, et al.
Published: (2025) -
LLM App Store Analysis: A Vision and Roadmap
by: Zhao, Yanjie, et al.
Published: (2024) -
Software Engineering for Large Language Models: Research Status, Challenges and the Road Ahead
by: Rao, Hongzhou, et al.
Published: (2025) -
Gradient-Based Model Fingerprinting for LLM Similarity Detection and Family Classification
by: Wu, Zehao, et al.
Published: (2025)