ScribeAgent: Towards Specialized Web Agents Using Production-Scale Workflow Data
Fuente:
arXiv
Salvato in:
| Autori principali: | Shen, Junhong, Jain, Atishay, Xiao, Zedian, Amlekar, Ishan, Hadji, Mouad, Podolny, Aaron, Talwalkar, Ameet |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
UPS: Efficiently Building Foundation Models for PDE Solving via Cross-Modal Adaptation
di: Shen, Junhong, et al.
Pubblicazione: (2024)
di: Shen, Junhong, et al.
Pubblicazione: (2024)
The Impact of Element Ordering on LM Agent Performance
di: Chi, Wayne, et al.
Pubblicazione: (2024)
di: Chi, Wayne, et al.
Pubblicazione: (2024)
Specialized Foundation Models Struggle to Beat Supervised Baselines
di: Xu, Zongzhe, et al.
Pubblicazione: (2024)
di: Xu, Zongzhe, et al.
Pubblicazione: (2024)
CoMind: Towards Community-Driven Agents for Machine Learning Engineering
di: Li, Sijie, et al.
Pubblicazione: (2025)
di: Li, Sijie, et al.
Pubblicazione: (2025)
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction
di: Shen, Junhong, et al.
Pubblicazione: (2025)
di: Shen, Junhong, et al.
Pubblicazione: (2025)
Code with Me or for Me? How Increasing AI Automation Transforms Developer Workflows
di: Chen, Valerie, et al.
Pubblicazione: (2025)
di: Chen, Valerie, et al.
Pubblicazione: (2025)
WebInject: Prompt Injection Attack to Web Agents
di: Wang, Xilong, et al.
Pubblicazione: (2025)
di: Wang, Xilong, et al.
Pubblicazione: (2025)
CodePDE: An Inference Framework for LLM-driven PDE Solver Generation
di: Li, Shanda, et al.
Pubblicazione: (2025)
di: Li, Shanda, et al.
Pubblicazione: (2025)
RECODE: Reasoning Through Code Generation for Visual Question Answering
di: Shen, Junhong, et al.
Pubblicazione: (2025)
di: Shen, Junhong, et al.
Pubblicazione: (2025)
Beyond the Commit: Developer Perspectives on Productivity with AI Coding Assistants
di: Chen, Valerie, et al.
Pubblicazione: (2026)
di: Chen, Valerie, et al.
Pubblicazione: (2026)
Completion $\neq$ Collaboration: Scaling Collaborative Effort with Agents
di: Shen, Shannon Zejiang, et al.
Pubblicazione: (2025)
di: Shen, Shannon Zejiang, et al.
Pubblicazione: (2025)
FrontierCO: Real-World and Large-Scale Evaluation of Machine Learning Solvers for Combinatorial Optimization
di: Feng, Shengyu, et al.
Pubblicazione: (2025)
di: Feng, Shengyu, et al.
Pubblicazione: (2025)
Centre driven Controlled Evolution of Wireless Virtual Networks based on Broadcast Tokens
di: Babu, Vignesh, et al.
Pubblicazione: (2025)
di: Babu, Vignesh, et al.
Pubblicazione: (2025)
Why Do Decision Makers (Not) Use AI? A Cross-Domain Analysis of Factors Impacting AI Adoption
di: Yu, Rebecca, et al.
Pubblicazione: (2025)
di: Yu, Rebecca, et al.
Pubblicazione: (2025)
Agreement-Based Cascading for Efficient Inference
di: Kolawole, Steven, et al.
Pubblicazione: (2024)
di: Kolawole, Steven, et al.
Pubblicazione: (2024)
Spend Less, Fit Better: Budget-Efficient Scaling Law Fitting via Active Experiment Selection
di: Li, Sijie, et al.
Pubblicazione: (2026)
di: Li, Sijie, et al.
Pubblicazione: (2026)
Provably tuning the ElasticNet across instances
di: Balcan, Maria-Florina, et al.
Pubblicazione: (2022)
di: Balcan, Maria-Florina, et al.
Pubblicazione: (2022)
Learning to Relax: Setting Solver Parameters Across a Sequence of Linear System Instances
di: Khodak, Mikhail, et al.
Pubblicazione: (2023)
di: Khodak, Mikhail, et al.
Pubblicazione: (2023)
Where Does My Model Underperform? A Human Evaluation of Slice Discovery Algorithms
di: Johnson, Nari, et al.
Pubblicazione: (2023)
di: Johnson, Nari, et al.
Pubblicazione: (2023)
AgentCompass: Towards Reliable Evaluation of Agentic Workflows in Production
di: Kartik, NVJK, et al.
Pubblicazione: (2025)
di: Kartik, NVJK, et al.
Pubblicazione: (2025)
ARFBench: Benchmarking Time Series Question Answering Ability for Software Incident Response
di: Xie, Stephan, et al.
Pubblicazione: (2026)
di: Xie, Stephan, et al.
Pubblicazione: (2026)
Pre-Generating Multi-Difficulty PDE Data for Few-Shot Neural PDE Solvers
di: Choudhary, Naman, et al.
Pubblicazione: (2025)
di: Choudhary, Naman, et al.
Pubblicazione: (2025)
MedScribe: Clinically Grounded CT Reporting through Agentic Workflows
di: Orlando, Giuseppe A., et al.
Pubblicazione: (2026)
di: Orlando, Giuseppe A., et al.
Pubblicazione: (2026)
Beyond the All-in-One Agent: Benchmarking Role-Specialized Multi-Agent Collaboration in Enterprise Workflows
di: Yu, Tao, et al.
Pubblicazione: (2026)
di: Yu, Tao, et al.
Pubblicazione: (2026)
PolyKV: A Shared Asymmetrically-Compressed KV Cache Pool for Multi-Agent LLM Inference
di: Patel, Ishan, et al.
Pubblicazione: (2026)
di: Patel, Ishan, et al.
Pubblicazione: (2026)
Do LLMs exhibit human-like response biases? A case study in survey design
di: Tjuatja, Lindia, et al.
Pubblicazione: (2023)
di: Tjuatja, Lindia, et al.
Pubblicazione: (2023)
AgentWebBench: Benchmarking Multi-Agent Coordination in Agentic Web
di: Zhong, Shanshan, et al.
Pubblicazione: (2026)
di: Zhong, Shanshan, et al.
Pubblicazione: (2026)
TAMAS: Benchmarking Adversarial Risks in Multi-Agent LLM Systems
di: Kavathekar, Ishan, et al.
Pubblicazione: (2025)
di: Kavathekar, Ishan, et al.
Pubblicazione: (2025)
ReUseIt: Synthesizing Reusable AI Agent Workflows for Web Automation
di: Liu, Yimeng, et al.
Pubblicazione: (2025)
di: Liu, Yimeng, et al.
Pubblicazione: (2025)
Learning Causality for Longitudinal Data
di: Bouchattaoui, Mouad EL
Pubblicazione: (2025)
di: Bouchattaoui, Mouad EL
Pubblicazione: (2025)
Sample Complexity and Representation Ability of Test-time Scaling Paradigms
di: Huang, Baihe, et al.
Pubblicazione: (2025)
di: Huang, Baihe, et al.
Pubblicazione: (2025)
Large Language Models as Pokémon Battle Agents: Strategic Play and Content Generation
di: Jain, Daksh, et al.
Pubblicazione: (2025)
di: Jain, Daksh, et al.
Pubblicazione: (2025)
Foam-Agent: Towards Automated Intelligent CFD Workflows
di: Yue, Ling, et al.
Pubblicazione: (2025)
di: Yue, Ling, et al.
Pubblicazione: (2025)
Need Help? Designing Proactive AI Assistants for Programming
di: Chen, Valerie, et al.
Pubblicazione: (2024)
di: Chen, Valerie, et al.
Pubblicazione: (2024)
CodingGenie: A Proactive LLM-Powered Programming Assistant
di: Zhao, Sebastian, et al.
Pubblicazione: (2025)
di: Zhao, Sebastian, et al.
Pubblicazione: (2025)
When Benchmarks Talk: Re-Evaluating Code LLMs with Interactive Feedback
di: Pan, Jane, et al.
Pubblicazione: (2025)
di: Pan, Jane, et al.
Pubblicazione: (2025)
Multitask Learning Can Improve Worst-Group Outcomes
di: Kulkarni, Atharva, et al.
Pubblicazione: (2023)
di: Kulkarni, Atharva, et al.
Pubblicazione: (2023)
Outside Knowledge Conversational Video (OKCV) Dataset -- Dialoguing over Videos
di: Reichman, Benjamin, et al.
Pubblicazione: (2025)
di: Reichman, Benjamin, et al.
Pubblicazione: (2025)
A Survey of WebAgents: Towards Next-Generation AI Agents for Web Automation with Large Foundation Models
di: Ning, Liangbo, et al.
Pubblicazione: (2025)
di: Ning, Liangbo, et al.
Pubblicazione: (2025)
WebWorld: A Large-Scale World Model for Web Agent Training
di: Xiao, Zikai, et al.
Pubblicazione: (2026)
di: Xiao, Zikai, et al.
Pubblicazione: (2026)
Documenti analoghi
-
UPS: Efficiently Building Foundation Models for PDE Solving via Cross-Modal Adaptation
di: Shen, Junhong, et al.
Pubblicazione: (2024) -
The Impact of Element Ordering on LM Agent Performance
di: Chi, Wayne, et al.
Pubblicazione: (2024) -
Specialized Foundation Models Struggle to Beat Supervised Baselines
di: Xu, Zongzhe, et al.
Pubblicazione: (2024) -
CoMind: Towards Community-Driven Agents for Machine Learning Engineering
di: Li, Sijie, et al.
Pubblicazione: (2025) -
Thinking vs. Doing: Agents that Reason by Scaling Test-Time Interaction
di: Shen, Junhong, et al.
Pubblicazione: (2025)