Cybernaut: Towards Reliable Web Automation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tomar, Ankur, Liang, Hengyue, Bhattacharya, Indranil, Larios, Natalia, Carbone, Francesco |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
A Dual-Helix Governance Approach Towards Reliable Agentic AI for WebGIS Development
von: Boyuan, et al.
Veröffentlicht: (2026)
von: Boyuan, et al.
Veröffentlicht: (2026)
WebCompass: Towards Multimodal Web Coding Evaluation for Code Language Models
von: Lei, Xinping, et al.
Veröffentlicht: (2026)
von: Lei, Xinping, et al.
Veröffentlicht: (2026)
Toward Reliable Design of LLM-Enabled Agentic Workflows: Optimizing Latency-Reliability-Cost Tradeoffs
von: Yang, Ya-Ting, et al.
Veröffentlicht: (2026)
von: Yang, Ya-Ting, et al.
Veröffentlicht: (2026)
Evaluating the Use of LLMs for Automated DOM-Level Resolution of Web Performance Issues
von: Peters, Gideon, et al.
Veröffentlicht: (2026)
von: Peters, Gideon, et al.
Veröffentlicht: (2026)
Explore-Construct-Filter: An Automated Framework for Rich and Reliable API Knowledge Graph Construction
von: Sun, Yanbang, et al.
Veröffentlicht: (2025)
von: Sun, Yanbang, et al.
Veröffentlicht: (2025)
WALL: A Web Application for Automated Quality Assurance using Large Language Models
von: Abtahi, Seyed Moein, et al.
Veröffentlicht: (2025)
von: Abtahi, Seyed Moein, et al.
Veröffentlicht: (2025)
Towards Reliable LLM-Driven Fuzz Testing: Vision and Road Ahead
von: Cheng, Yiran, et al.
Veröffentlicht: (2025)
von: Cheng, Yiran, et al.
Veröffentlicht: (2025)
Towards Reliable Evaluation of Neural Program Repair with Natural Robustness Testing
von: Le-Cong, Thanh, et al.
Veröffentlicht: (2024)
von: Le-Cong, Thanh, et al.
Veröffentlicht: (2024)
Learning From Developers: Towards Reliable Patch Validation at Scale for Linux
von: Lin, Chih-En, et al.
Veröffentlicht: (2026)
von: Lin, Chih-En, et al.
Veröffentlicht: (2026)
Towards Automated Formal Verification of Backend Systems with LLMs
von: Xu, Kangping, et al.
Veröffentlicht: (2025)
von: Xu, Kangping, et al.
Veröffentlicht: (2025)
A-ProS: Towards Reliable Autonomous Programming Through Multi-Model Feedback
von: Tabassum, Anika, et al.
Veröffentlicht: (2026)
von: Tabassum, Anika, et al.
Veröffentlicht: (2026)
From Mirage to Grounding: Towards Reliable Multimodal Circuit-to-Verilog Code Generation
von: Yang, Guang, et al.
Veröffentlicht: (2026)
von: Yang, Guang, et al.
Veröffentlicht: (2026)
Toward Automated and Trustworthy Scientific Analysis and Visualization with LLM-Generated Code
von: Chakroborti, Apu Kumar, et al.
Veröffentlicht: (2025)
von: Chakroborti, Apu Kumar, et al.
Veröffentlicht: (2025)
Toward Scalable Automated Repository-Level Datasets for Software Vulnerability Detection
von: Lbath, Amine
Veröffentlicht: (2026)
von: Lbath, Amine
Veröffentlicht: (2026)
AutoWebWorld: Synthesizing Infinite Verifiable Web Environments via Finite State Machines
von: Wu, Yifan, et al.
Veröffentlicht: (2026)
von: Wu, Yifan, et al.
Veröffentlicht: (2026)
Towards an Ontology for Scenario Definition for the Assessment of Automated Vehicles: An Object-Oriented Framework
von: de Gelder, E., et al.
Veröffentlicht: (2020)
von: de Gelder, E., et al.
Veröffentlicht: (2020)
AI-Driven Self-Evolving Software: A Promising Path Toward Software Automation
von: Cai, Liyi, et al.
Veröffentlicht: (2025)
von: Cai, Liyi, et al.
Veröffentlicht: (2025)
Toward Automated Validation of Language Model Synthesized Test Cases using Semantic Entropy
von: Taherkhani, Hamed, et al.
Veröffentlicht: (2024)
von: Taherkhani, Hamed, et al.
Veröffentlicht: (2024)
Generating Reliable Adverse event Profiles for Health through Automated Integrated Data (GRAPH-AID): A Semi-Automated Ontology Building Approach
von: Gadusu, Srikar Reddy, et al.
Veröffentlicht: (2025)
von: Gadusu, Srikar Reddy, et al.
Veröffentlicht: (2025)
WebSuite: Systematically Evaluating Why Web Agents Fail
von: Li, Eric, et al.
Veröffentlicht: (2024)
von: Li, Eric, et al.
Veröffentlicht: (2024)
EmbeWebAgent: Embedding Web Agents into Any Customized UI
von: Ma, Chenyang, et al.
Veröffentlicht: (2026)
von: Ma, Chenyang, et al.
Veröffentlicht: (2026)
WebDevJudge: Evaluating (M)LLMs as Critiques for Web Development Quality
von: Li, Chunyang, et al.
Veröffentlicht: (2025)
von: Li, Chunyang, et al.
Veröffentlicht: (2025)
QSpark: Towards Reliable Qiskit Code Generation
von: Kheiri, Kiana, et al.
Veröffentlicht: (2025)
von: Kheiri, Kiana, et al.
Veröffentlicht: (2025)
Automatically Generating Web Applications from Requirements Via Multi-Agent Test-Driven Development
von: Wan, Yuxuan, et al.
Veröffentlicht: (2025)
von: Wan, Yuxuan, et al.
Veröffentlicht: (2025)
WebCoderBench: Benchmarking Web Application Generation with Comprehensive and Interpretable Evaluation Metrics
von: Liu, Chenxu, et al.
Veröffentlicht: (2026)
von: Liu, Chenxu, et al.
Veröffentlicht: (2026)
WebApp1K: A Practical Code-Generation Benchmark for Web App Development
von: Cui, Yi
Veröffentlicht: (2024)
von: Cui, Yi
Veröffentlicht: (2024)
StressWeb: A Diagnostic Benchmark for Web Agent Robustness under Realistic Interaction Variability
von: Bai, Haoyue, et al.
Veröffentlicht: (2026)
von: Bai, Haoyue, et al.
Veröffentlicht: (2026)
WebVIA: A Web-based Vision-Language Agentic Framework for Interactive and Verifiable UI-to-Code Generation
von: Xu, Mingde, et al.
Veröffentlicht: (2025)
von: Xu, Mingde, et al.
Veröffentlicht: (2025)
MAAD: Automate Software Architecture Design through Knowledge-Driven Multi-Agent Collaboration
von: Li, Ruiyin, et al.
Veröffentlicht: (2025)
von: Li, Ruiyin, et al.
Veröffentlicht: (2025)
LLMs and Stack Overflow Discussions: Reliability, Impact, and Challenges
von: Da Silva, Leuson, et al.
Veröffentlicht: (2024)
von: Da Silva, Leuson, et al.
Veröffentlicht: (2024)
Multimodal Auto Validation For Self-Refinement in Web Agents
von: Azam, Ruhana, et al.
Veröffentlicht: (2024)
von: Azam, Ruhana, et al.
Veröffentlicht: (2024)
Previously on... Automating Code Review
von: Heumüller, Robert, et al.
Veröffentlicht: (2025)
von: Heumüller, Robert, et al.
Veröffentlicht: (2025)
MacroBench: A Novel Testbed for Web Automation Scripts via Large Language Models
von: Kim, Hyunjun, et al.
Veröffentlicht: (2025)
von: Kim, Hyunjun, et al.
Veröffentlicht: (2025)
ACCESS: Prompt Engineering for Automated Web Accessibility Violation Corrections
von: Huang, Calista, et al.
Veröffentlicht: (2024)
von: Huang, Calista, et al.
Veröffentlicht: (2024)
ReXCL: A Tool for Requirement Document Extraction and Classification
von: Bhattacharya, Paheli, et al.
Veröffentlicht: (2025)
von: Bhattacharya, Paheli, et al.
Veröffentlicht: (2025)
Application Modernization with LLMs: Addressing Core Challenges in Reliability, Security, and Quality
von: Ponnusamy, Ahilan Ayyachamy Nadar
Veröffentlicht: (2025)
von: Ponnusamy, Ahilan Ayyachamy Nadar
Veröffentlicht: (2025)
TREAT: A Code LLMs Trustworthiness / Reliability Evaluation and Testing Framework
von: Gao, Shuzheng, et al.
Veröffentlicht: (2025)
von: Gao, Shuzheng, et al.
Veröffentlicht: (2025)
A Self-Healing Framework for Reliable LLM-Based Autonomous Agents
von: Jeong, Cheonsu, et al.
Veröffentlicht: (2026)
von: Jeong, Cheonsu, et al.
Veröffentlicht: (2026)
Chain of Targeted Verification Questions to Improve the Reliability of Code Generated by LLMs
von: Ngassom, Sylvain Kouemo, et al.
Veröffentlicht: (2024)
von: Ngassom, Sylvain Kouemo, et al.
Veröffentlicht: (2024)
Are LLMs Reliable Code Reviewers? Systematic Overcorrection in Requirement Conformance Judgement
von: Jin, Haolin, et al.
Veröffentlicht: (2026)
von: Jin, Haolin, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
A Dual-Helix Governance Approach Towards Reliable Agentic AI for WebGIS Development
von: Boyuan, et al.
Veröffentlicht: (2026) -
WebCompass: Towards Multimodal Web Coding Evaluation for Code Language Models
von: Lei, Xinping, et al.
Veröffentlicht: (2026) -
Toward Reliable Design of LLM-Enabled Agentic Workflows: Optimizing Latency-Reliability-Cost Tradeoffs
von: Yang, Ya-Ting, et al.
Veröffentlicht: (2026) -
Evaluating the Use of LLMs for Automated DOM-Level Resolution of Web Performance Issues
von: Peters, Gideon, et al.
Veröffentlicht: (2026) -
Explore-Construct-Filter: An Automated Framework for Rich and Reliable API Knowledge Graph Construction
von: Sun, Yanbang, et al.
Veröffentlicht: (2025)