Salvato in:
| Autori principali: | Weidener, Lukas, Brkić, Marko, Jovanović, Mihailo, Ulgac, Emre, Meduri, Aakaash |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2605.21545 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Rethinking the AI Scientist: Interactive Multi-Agent Workflows for Scientific Discovery
di: Weidener, Lukas, et al.
Pubblicazione: (2026)
di: Weidener, Lukas, et al.
Pubblicazione: (2026)
From Task Executors to Research Partners: Evaluating AI Co-Pilots Through Workflow Integration in Biomedical Research
di: Weidener, Lukas, et al.
Pubblicazione: (2025)
di: Weidener, Lukas, et al.
Pubblicazione: (2025)
RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models
di: Muhamed, Aashiq, et al.
Pubblicazione: (2025)
di: Muhamed, Aashiq, et al.
Pubblicazione: (2025)
ORFuzz: Fuzzing the "Other Side" of LLM Safety -- Testing Over-Refusal
di: Zhang, Haonan, et al.
Pubblicazione: (2025)
di: Zhang, Haonan, et al.
Pubblicazione: (2025)
Enhancing Educational Efficiency: Generative AI Chatbots and DevOps in Education 4.0
di: Mekić, Edis, et al.
Pubblicazione: (2024)
di: Mekić, Edis, et al.
Pubblicazione: (2024)
ImproBR: Bug Report Improver Using LLMs
di: Akyol, Emre Furkan, et al.
Pubblicazione: (2026)
di: Akyol, Emre Furkan, et al.
Pubblicazione: (2026)
FullStack Bench: Evaluating LLMs as Full Stack Coders
di: Bytedance-Seed-Foundation-Code-Team, et al.
Pubblicazione: (2024)
di: Bytedance-Seed-Foundation-Code-Team, et al.
Pubblicazione: (2024)
Neuron-Guided Interpretation of Code LLMs: Where, Why, and How?
di: Yin, Zhe, et al.
Pubblicazione: (2025)
di: Yin, Zhe, et al.
Pubblicazione: (2025)
The SWE-Bench Illusion: When State-of-the-Art LLMs Remember Instead of Reason
di: Liang, Shanchao, et al.
Pubblicazione: (2025)
di: Liang, Shanchao, et al.
Pubblicazione: (2025)
Codev-Bench: How Do LLMs Understand Developer-Centric Code Completion?
di: Pan, Zhenyu, et al.
Pubblicazione: (2024)
di: Pan, Zhenyu, et al.
Pubblicazione: (2024)
When the Code Autopilot Breaks: Why LLMs Falter in Embedded Machine Learning
di: Morabito, Roberto, et al.
Pubblicazione: (2025)
di: Morabito, Roberto, et al.
Pubblicazione: (2025)
Terminus-4B: Can a Smaller Model Replace Frontier LLMs at Agentic Execution Tasks?
di: Garg, Spandan, et al.
Pubblicazione: (2026)
di: Garg, Spandan, et al.
Pubblicazione: (2026)
ResearchEnvBench: Benchmarking Agents on Environment Synthesis for Research Code Execution
di: Wang, Yubang, et al.
Pubblicazione: (2026)
di: Wang, Yubang, et al.
Pubblicazione: (2026)
FrontendBench: A Benchmark for Evaluating LLMs on Front-End Development via Automatic Evaluation
di: Zhu, Hongda, et al.
Pubblicazione: (2025)
di: Zhu, Hongda, et al.
Pubblicazione: (2025)
Prompting for Performance: Exploring LLMs for Configuring Software
di: Spieker, Helge, et al.
Pubblicazione: (2025)
di: Spieker, Helge, et al.
Pubblicazione: (2025)
Mind the Prompt: Self-adaptive Generation of Task Plan Explanations via LLMs
di: Vázquez, Gricel, et al.
Pubblicazione: (2026)
di: Vázquez, Gricel, et al.
Pubblicazione: (2026)
Prompt engineering and framework: implementation to increase code reliability based guideline for LLMs
di: Cruz, Rogelio, et al.
Pubblicazione: (2025)
di: Cruz, Rogelio, et al.
Pubblicazione: (2025)
Prompt Stability in Code LLMs: Measuring Sensitivity across Emotion- and Personality-Driven Variations
di: Ma, Wei, et al.
Pubblicazione: (2025)
di: Ma, Wei, et al.
Pubblicazione: (2025)
Are We SOLID Yet? An Empirical Study on Prompting LLMs to Detect Design Principle Violations
di: Pehlivan, Fatih, et al.
Pubblicazione: (2025)
di: Pehlivan, Fatih, et al.
Pubblicazione: (2025)
From Agent-Only Social Networks to Autonomous Scientific Research: Lessons from OpenClaw and Moltbook, and the Architecture of ClawdLab and Beach.Science
di: Weidener, Lukas, et al.
Pubblicazione: (2026)
di: Weidener, Lukas, et al.
Pubblicazione: (2026)
BenchEvolver: Frontier Task Synthesis via Solution-Centric Evolution
di: Wu, Yangzhen, et al.
Pubblicazione: (2026)
di: Wu, Yangzhen, et al.
Pubblicazione: (2026)
LLM Applications: Current Paradigms and the Next Frontier
di: Hou, Xinyi, et al.
Pubblicazione: (2025)
di: Hou, Xinyi, et al.
Pubblicazione: (2025)
A Study of LLMs' Preferences for Libraries and Programming Languages
di: Twist, Lukas, et al.
Pubblicazione: (2025)
di: Twist, Lukas, et al.
Pubblicazione: (2025)
LLMs: A Game-Changer for Software Engineers?
di: Haque, Md Asraful
Pubblicazione: (2024)
di: Haque, Md Asraful
Pubblicazione: (2024)
Insights from Benchmarking Frontier Language Models on Web App Code Generation
di: Cui, Yi
Pubblicazione: (2024)
di: Cui, Yi
Pubblicazione: (2024)
Get on the Train or be Left on the Station: Using LLMs for Software Engineering Research
di: Trinkenreich, Bianca, et al.
Pubblicazione: (2025)
di: Trinkenreich, Bianca, et al.
Pubblicazione: (2025)
PromptPex: Automatic Test Generation for Language Model Prompts
di: Sharma, Reshabh K, et al.
Pubblicazione: (2025)
di: Sharma, Reshabh K, et al.
Pubblicazione: (2025)
CangjieBench: Benchmarking LLMs on a Low-Resource General-Purpose Programming Language
di: Cheng, Junhang, et al.
Pubblicazione: (2026)
di: Cheng, Junhang, et al.
Pubblicazione: (2026)
KernelBench: Can LLMs Write Efficient GPU Kernels?
di: Ouyang, Anne, et al.
Pubblicazione: (2025)
di: Ouyang, Anne, et al.
Pubblicazione: (2025)
CrackMeBench: Binary Reverse Engineering for Agents
di: David, Isaac, et al.
Pubblicazione: (2026)
di: David, Isaac, et al.
Pubblicazione: (2026)
Automated Root-Cause Subclassification and No-Code Fix Generation for Invalid Bug Reports
di: Gon, Mahmut Furkan, et al.
Pubblicazione: (2026)
di: Gon, Mahmut Furkan, et al.
Pubblicazione: (2026)
SciReplicate-Bench: Benchmarking LLMs in Agent-driven Algorithmic Reproduction from Research Papers
di: Xiang, Yanzheng, et al.
Pubblicazione: (2025)
di: Xiang, Yanzheng, et al.
Pubblicazione: (2025)
Prompt Less, Smile More: MTP with Semantic Engineering in Lieu of Prompt Engineering
di: Dantanarayana, Jayanaka L., et al.
Pubblicazione: (2025)
di: Dantanarayana, Jayanaka L., et al.
Pubblicazione: (2025)
SWE Context Bench: A Benchmark for Context Learning in Coding
di: Zhu, Jiayuan, et al.
Pubblicazione: (2026)
di: Zhu, Jiayuan, et al.
Pubblicazione: (2026)
PBT-Bench: Benchmarking AI Agents on Property-Based Testing
di: Jing, Lucas, et al.
Pubblicazione: (2026)
di: Jing, Lucas, et al.
Pubblicazione: (2026)
FeatureBench: Benchmarking Agentic Coding for Complex Feature Development
di: Zhou, Qixing, et al.
Pubblicazione: (2026)
di: Zhou, Qixing, et al.
Pubblicazione: (2026)
Multi-Sample Prompting and Actor-Critic Prompt Optimization for Diverse Synthetic Data Generation
di: El-Hajjami, Abdelkarim, et al.
Pubblicazione: (2025)
di: El-Hajjami, Abdelkarim, et al.
Pubblicazione: (2025)
WebSuite: Systematically Evaluating Why Web Agents Fail
di: Li, Eric, et al.
Pubblicazione: (2024)
di: Li, Eric, et al.
Pubblicazione: (2024)
Code2Bench: Scaling Source and Rigor for Dynamic Benchmark Construction
di: Zhang, Zhe, et al.
Pubblicazione: (2025)
di: Zhang, Zhe, et al.
Pubblicazione: (2025)
ProgramBench: Can Language Models Rebuild Programs From Scratch?
di: Yang, John, et al.
Pubblicazione: (2026)
di: Yang, John, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Rethinking the AI Scientist: Interactive Multi-Agent Workflows for Scientific Discovery
di: Weidener, Lukas, et al.
Pubblicazione: (2026) -
From Task Executors to Research Partners: Evaluating AI Co-Pilots Through Workflow Integration in Biomedical Research
di: Weidener, Lukas, et al.
Pubblicazione: (2025) -
RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models
di: Muhamed, Aashiq, et al.
Pubblicazione: (2025) -
ORFuzz: Fuzzing the "Other Side" of LLM Safety -- Testing Over-Refusal
di: Zhang, Haonan, et al.
Pubblicazione: (2025) -
Enhancing Educational Efficiency: Generative AI Chatbots and DevOps in Education 4.0
di: Mekić, Edis, et al.
Pubblicazione: (2024)