DI-BENCH: Benchmarking Large Language Models on Dependency Inference with Testable Repositories at Scale
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Linghao, Wang, Junhao, He, Shilin, Zhang, Chaoyun, Kang, Yu, Li, Bowen, Wen, Jiaheng, Xie, Chengxing, Wang, Maoquan, Huang, Yufan, Nallipogu, Elsie, Lin, Qingwei, Dang, Yingnong, Rajmohan, Saravan, Zhang, Dongmei, Zhang, Qi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SWE-bench Goes Live!
von: Zhang, Linghao, et al.
Veröffentlicht: (2025)
von: Zhang, Linghao, et al.
Veröffentlicht: (2025)
Skeleton-Guided-Translation: A Benchmarking Framework for Code Repository Translation with Fine-Grained Quality Evaluation
von: Zhang, Xing, et al.
Veröffentlicht: (2025)
von: Zhang, Xing, et al.
Veröffentlicht: (2025)
RepoLaunch: Automating Build&Test Pipeline of Code Repositories on ANY Language and ANY Platform
von: Li, Kenan, et al.
Veröffentlicht: (2026)
von: Li, Kenan, et al.
Veröffentlicht: (2026)
API Agents vs. GUI Agents: Divergence and Convergence
von: Zhang, Chaoyun, et al.
Veröffentlicht: (2025)
von: Zhang, Chaoyun, et al.
Veröffentlicht: (2025)
From Reasoning to Answer: Empirical, Attention-Based and Mechanistic Insights into Distilled DeepSeek R1 Models
von: Zhang, Jue, et al.
Veröffentlicht: (2025)
von: Zhang, Jue, et al.
Veröffentlicht: (2025)
Text2Grad: Reinforcement Learning from Natural Language Feedback
von: Wang, Hanyang, et al.
Veröffentlicht: (2025)
von: Wang, Hanyang, et al.
Veröffentlicht: (2025)
ORACLE-SWE: Quantifying the Contribution of Oracle Information Signals on SWE Agents
von: Li, Kenan, et al.
Veröffentlicht: (2026)
von: Li, Kenan, et al.
Veröffentlicht: (2026)
Enabling Autonomic Microservice Management through Self-Learning Agents
von: Yu, Fenglin, et al.
Veröffentlicht: (2025)
von: Yu, Fenglin, et al.
Veröffentlicht: (2025)
Zipage: Maintain High Request Concurrency for LLM Reasoning through Compressed PagedAttention
von: Liao, Mengqi, et al.
Veröffentlicht: (2026)
von: Liao, Mengqi, et al.
Veröffentlicht: (2026)
Everything of Thoughts: Defying the Law of Penrose Triangle for Thought Generation
von: Ding, Ruomeng, et al.
Veröffentlicht: (2023)
von: Ding, Ruomeng, et al.
Veröffentlicht: (2023)
VEM: Environment-Free Exploration for Training GUI Agent with Value Environment Model
von: Zheng, Jiani, et al.
Veröffentlicht: (2025)
von: Zheng, Jiani, et al.
Veröffentlicht: (2025)
UFO: A UI-Focused Agent for Windows OS Interaction
von: Zhang, Chaoyun, et al.
Veröffentlicht: (2024)
von: Zhang, Chaoyun, et al.
Veröffentlicht: (2024)
Large Language Model-Brained GUI Agents: A Survey
von: Zhang, Chaoyun, et al.
Veröffentlicht: (2024)
von: Zhang, Chaoyun, et al.
Veröffentlicht: (2024)
Large Language Models can Deliver Accurate and Interpretable Time Series Anomaly Detection
von: Liu, Jun, et al.
Veröffentlicht: (2024)
von: Liu, Jun, et al.
Veröffentlicht: (2024)
Why does Prediction Accuracy Decrease over Time? Uncertain Positive Learning for Cloud Failure Prediction
von: Li, Haozhe, et al.
Veröffentlicht: (2024)
von: Li, Haozhe, et al.
Veröffentlicht: (2024)
Contrastive Attribution in the Wild: An Interpretability Analysis of LLM Failures on Realistic Benchmarks
von: Tan, Rongyuan, et al.
Veröffentlicht: (2026)
von: Tan, Rongyuan, et al.
Veröffentlicht: (2026)
G-KV: Decoding-Time KV Cache Eviction with Global Attention
von: Liao, Mengqi, et al.
Veröffentlicht: (2025)
von: Liao, Mengqi, et al.
Veröffentlicht: (2025)
Lita: Light Agent Uncovers the Agentic Coding Capabilities of LLMs
von: Dai, Hankun, et al.
Veröffentlicht: (2025)
von: Dai, Hankun, et al.
Veröffentlicht: (2025)
RepoGenesis: Benchmarking End-to-End Microservice Generation from Readme to Repository
von: Peng, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Peng, Zhiyuan, et al.
Veröffentlicht: (2026)
StepFly: Agentic Troubleshooting Guide Automation for Incident Diagnosis
von: Mao, Jiayi, et al.
Veröffentlicht: (2025)
von: Mao, Jiayi, et al.
Veröffentlicht: (2025)
LLM Reasoning as Trajectories: Step-Specific Representation Geometry and Correctness Signals
von: Sun, Lihao, et al.
Veröffentlicht: (2026)
von: Sun, Lihao, et al.
Veröffentlicht: (2026)
DoVer: Intervention-Driven Auto Debugging for LLM Multi-Agent Systems
von: Ma, Ming, et al.
Veröffentlicht: (2025)
von: Ma, Ming, et al.
Veröffentlicht: (2025)
Privacy in Action: Towards Realistic Privacy Mitigation and Evaluation for LLM-Powered Agents
von: Wang, Shouju, et al.
Veröffentlicht: (2025)
von: Wang, Shouju, et al.
Veröffentlicht: (2025)
UFO3: Weaving the Digital Agent Galaxy
von: Zhang, Chaoyun, et al.
Veröffentlicht: (2025)
von: Zhang, Chaoyun, et al.
Veröffentlicht: (2025)
MEETING DELEGATE: Benchmarking LLMs on Attending Meetings on Our Behalf
von: Hu, Lingxiang, et al.
Veröffentlicht: (2025)
von: Hu, Lingxiang, et al.
Veröffentlicht: (2025)
From Task to Tutorial: An Automated GUI Framework for Excel Tutorial Document and Video Creation
von: Xie, Yuhang, et al.
Veröffentlicht: (2025)
von: Xie, Yuhang, et al.
Veröffentlicht: (2025)
Call Me When Necessary: LLMs can Efficiently and Faithfully Reason over Structured Environments
von: Cheng, Sitao, et al.
Veröffentlicht: (2024)
von: Cheng, Sitao, et al.
Veröffentlicht: (2024)
An Empirical Study of Production Incidents in Generative AI Cloud Services
von: Yan, Haoran, et al.
Veröffentlicht: (2025)
von: Yan, Haoran, et al.
Veröffentlicht: (2025)
CI-Work: Benchmarking Contextual Integrity in Enterprise LLM Agents
von: Fu, Wenjie, et al.
Veröffentlicht: (2026)
von: Fu, Wenjie, et al.
Veröffentlicht: (2026)
Navigating the Unknown: A Chat-Based Collaborative Interface for Personalized Exploratory Tasks
von: Peng, Yingzhe, et al.
Veröffentlicht: (2024)
von: Peng, Yingzhe, et al.
Veröffentlicht: (2024)
The Vision of Autonomic Computing: Can LLMs Make It a Reality?
von: Zhang, Zhiyang, et al.
Veröffentlicht: (2024)
von: Zhang, Zhiyang, et al.
Veröffentlicht: (2024)
AllHands: Ask Me Anything on Large-scale Verbatim Feedback via Large Language Models
von: Zhang, Chaoyun, et al.
Veröffentlicht: (2024)
von: Zhang, Chaoyun, et al.
Veröffentlicht: (2024)
AXIS: Efficient Human-Agent-Computer Interaction with API-First LLM-Based Agents
von: Lu, Junting, et al.
Veröffentlicht: (2024)
von: Lu, Junting, et al.
Veröffentlicht: (2024)
Pretrain Value, Not Reward: Decoupled Value Policy Optimization
von: Huang, Chenghua, et al.
Veröffentlicht: (2025)
von: Huang, Chenghua, et al.
Veröffentlicht: (2025)
RuAG: Learned-rule-augmented Generation for Large Language Models
von: Zhang, Yudi, et al.
Veröffentlicht: (2024)
von: Zhang, Yudi, et al.
Veröffentlicht: (2024)
AutoRAG-HP: Automatic Online Hyper-Parameter Tuning for Retrieval-Augmented Generation
von: Fu, Jia, et al.
Veröffentlicht: (2024)
von: Fu, Jia, et al.
Veröffentlicht: (2024)
Learning to Refine: Self-Refinement of Parallel Reasoning in LLMs
von: Wang, Qibin, et al.
Veröffentlicht: (2025)
von: Wang, Qibin, et al.
Veröffentlicht: (2025)
Distill Not Only Data but Also Rewards: Can Smaller Language Models Surpass Larger Ones?
von: Zhang, Yudi, et al.
Veröffentlicht: (2025)
von: Zhang, Yudi, et al.
Veröffentlicht: (2025)
Self-Evolved Reward Learning for LLMs
von: Huang, Chenghua, et al.
Veröffentlicht: (2024)
von: Huang, Chenghua, et al.
Veröffentlicht: (2024)
Sharingan: Extract User Action Sequence from Desktop Recordings
von: Chen, Yanting, et al.
Veröffentlicht: (2024)
von: Chen, Yanting, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
SWE-bench Goes Live!
von: Zhang, Linghao, et al.
Veröffentlicht: (2025) -
Skeleton-Guided-Translation: A Benchmarking Framework for Code Repository Translation with Fine-Grained Quality Evaluation
von: Zhang, Xing, et al.
Veröffentlicht: (2025) -
RepoLaunch: Automating Build&Test Pipeline of Code Repositories on ANY Language and ANY Platform
von: Li, Kenan, et al.
Veröffentlicht: (2026) -
API Agents vs. GUI Agents: Divergence and Convergence
von: Zhang, Chaoyun, et al.
Veröffentlicht: (2025) -
From Reasoning to Answer: Empirical, Attention-Based and Mechanistic Insights into Distilled DeepSeek R1 Models
von: Zhang, Jue, et al.
Veröffentlicht: (2025)