Large Language Models Can Self-Improve At Web Agent Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Patel, Ajay, Hofmarcher, Markus, Leoveanu-Condrei, Claudiu, Dinu, Marius-Constantin, Callison-Burch, Chris, Hochreiter, Sepp |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SymbolicAI: A framework for logic-based approaches combining generative models and solvers
by: Dinu, Marius-Constantin, et al.
Published: (2024)
by: Dinu, Marius-Constantin, et al.
Published: (2024)
HyDRA: A Hybrid-Driven Reasoning Architecture for Verifiable Knowledge Graphs
by: Kaiser, Adrian, et al.
Published: (2025)
by: Kaiser, Adrian, et al.
Published: (2025)
A DbC Inspired Neurosymbolic Layer for Trustworthy Agent Design
by: Leoveanu-Condrei, Claudiu
Published: (2025)
by: Leoveanu-Condrei, Claudiu
Published: (2025)
Low-Resource Authorship Style Transfer: Can Non-Famous Authors Be Imitated?
by: Patel, Ajay, et al.
Published: (2022)
by: Patel, Ajay, et al.
Published: (2022)
Linear Alignment of Vision-language Models for Image Captioning
by: Paischer, Fabian, et al.
Published: (2023)
by: Paischer, Fabian, et al.
Published: (2023)
FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale
by: Patel, Ajay, et al.
Published: (2026)
by: Patel, Ajay, et al.
Published: (2026)
DataDreamer: A Tool for Synthetic Data Generation and Reproducible LLM Workflows
by: Patel, Ajay, et al.
Published: (2024)
by: Patel, Ajay, et al.
Published: (2024)
FPC-Net: Revisiting SuperPoint with Descriptor-Free Keypoint Detection via Feature Pyramids and Consistency-Based Implicit Matching
by: Grigore, Ionuţ, et al.
Published: (2025)
by: Grigore, Ionuţ, et al.
Published: (2025)
Unlocking the Working Memory of Large Language Models for Latent Reasoning
by: Aichberger, Lukas, et al.
Published: (2026)
by: Aichberger, Lukas, et al.
Published: (2026)
BibTeX Citation Hallucinations in Scientific Publishing Agents: Evaluation and Mitigation
by: Rao, Delip, et al.
Published: (2026)
by: Rao, Delip, et al.
Published: (2026)
Overhearing LLM Agents: A Survey, Taxonomy, and Roadmap
by: Zhu, Andrew, et al.
Published: (2025)
by: Zhu, Andrew, et al.
Published: (2025)
mStyleDistance: Multilingual Style Embeddings and their Evaluation
by: Qiu, Justin, et al.
Published: (2025)
by: Qiu, Justin, et al.
Published: (2025)
WHAT-IF: Exploring Branching Narratives by Meta-Prompting Large Language Models
by: Huang, Runsheng "Anson", et al.
Published: (2024)
by: Huang, Runsheng "Anson", et al.
Published: (2024)
This Land is {Your, My} Land: Evaluating Geopolitical Biases in Language Models
by: Li, Bryan, et al.
Published: (2023)
by: Li, Bryan, et al.
Published: (2023)
Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents
by: Rao, Delip, et al.
Published: (2026)
by: Rao, Delip, et al.
Published: (2026)
Contrastive Abstraction for Reinforcement Learning
by: Patil, Vihang, et al.
Published: (2024)
by: Patil, Vihang, et al.
Published: (2024)
What Do Claim Verification Datasets Actually Test? A Reasoning Trace Analysis
by: Rao, Delip, et al.
Published: (2026)
by: Rao, Delip, et al.
Published: (2026)
ParaGuide: Guided Diffusion Paraphrasers for Plug-and-Play Textual Style Transfer
by: Horvitz, Zachary, et al.
Published: (2023)
by: Horvitz, Zachary, et al.
Published: (2023)
Uncovering Differences in Persuasive Language in Russian versus English Wikipedia
by: Li, Bryan, et al.
Published: (2024)
by: Li, Bryan, et al.
Published: (2024)
FanOutQA: A Multi-Hop, Multi-Document Question Answering Benchmark for Large Language Models
by: Zhu, Andrew, et al.
Published: (2024)
by: Zhu, Andrew, et al.
Published: (2024)
Autorubric: Unifying Rubric-based LLM Evaluation
by: Rao, Delip, et al.
Published: (2026)
by: Rao, Delip, et al.
Published: (2026)
Towards Faithful Model Explanation in NLP: A Survey
by: Lyu, Qing, et al.
Published: (2022)
by: Lyu, Qing, et al.
Published: (2022)
First Steps Towards Overhearing LLM Agents: A Case Study With Dungeons & Dragons Gameplay
by: Zhu, Andrew, et al.
Published: (2025)
by: Zhu, Andrew, et al.
Published: (2025)
TinyStyler: Efficient Few-Shot Text Style Transfer with Authorship Embeddings
by: Horvitz, Zachary, et al.
Published: (2024)
by: Horvitz, Zachary, et al.
Published: (2024)
ReDel: A Toolkit for LLM-Powered Recursive Multi-Agent Systems
by: Zhu, Andrew, et al.
Published: (2024)
by: Zhu, Andrew, et al.
Published: (2024)
Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why
by: Rao, Delip, et al.
Published: (2026)
by: Rao, Delip, et al.
Published: (2026)
ThinknCheck: Grounded Claim Verification with Compact, Reasoning-Driven, and Interpretable Models
by: Rao, Delip, et al.
Published: (2026)
by: Rao, Delip, et al.
Published: (2026)
Toward Beginner-Friendly LLMs for Language Learning: Controlling Difficulty in Conversation
by: Jin, Meiqing, et al.
Published: (2025)
by: Jin, Meiqing, et al.
Published: (2025)
Designing NLP Systems That Adapt to Diverse Worldviews
by: Creanga, Claudiu, et al.
Published: (2024)
by: Creanga, Claudiu, et al.
Published: (2024)
Transformer based neural networks for emotion recognition in conversations
by: Creanga, Claudiu, et al.
Published: (2024)
by: Creanga, Claudiu, et al.
Published: (2024)
Automated Text Identification Using CNN and Training Dynamics
by: Creanga, Claudiu, et al.
Published: (2024)
by: Creanga, Claudiu, et al.
Published: (2024)
You Have Thirteen Hours in Which to Solve the Labyrinth: Enhancing AI Game Masters with Function Calling
by: Song, Jaewoo, et al.
Published: (2024)
by: Song, Jaewoo, et al.
Published: (2024)
OpenPI2.0: An Improved Dataset for Entity Tracking in Texts
by: Zhang, Li, et al.
Published: (2023)
by: Zhang, Li, et al.
Published: (2023)
WithdrarXiv: A Large-Scale Dataset for Retraction Study
by: Rao, Delip, et al.
Published: (2024)
by: Rao, Delip, et al.
Published: (2024)
StyleDistance: Stronger Content-Independent Style Embeddings with Synthetic Parallel Examples
by: Patel, Ajay, et al.
Published: (2024)
by: Patel, Ajay, et al.
Published: (2024)
Calibrating Large Language Models with Sample Consistency
by: Lyu, Qing, et al.
Published: (2024)
by: Lyu, Qing, et al.
Published: (2024)
Transformer and Hybrid Deep Learning Based Models for Machine-Generated Text Detection
by: Marchitan, Teodor-George, et al.
Published: (2024)
by: Marchitan, Teodor-George, et al.
Published: (2024)
NSF-SciFy: Mining the NSF Awards Database for Scientific Claims
by: Rao, Delip, et al.
Published: (2025)
by: Rao, Delip, et al.
Published: (2025)
When Verification Fails: How Compositionally Infeasible Claims Escape Rejection
by: Liu, Muxin, et al.
Published: (2026)
by: Liu, Muxin, et al.
Published: (2026)
Primality Testing via Circulant Matrix Eigenvalue Structure: A Novel Approach Using Cyclotomic Field Theory
by: Dinu, Marius-Constantin
Published: (2025)
by: Dinu, Marius-Constantin
Published: (2025)
Similar Items
-
SymbolicAI: A framework for logic-based approaches combining generative models and solvers
by: Dinu, Marius-Constantin, et al.
Published: (2024) -
HyDRA: A Hybrid-Driven Reasoning Architecture for Verifiable Knowledge Graphs
by: Kaiser, Adrian, et al.
Published: (2025) -
A DbC Inspired Neurosymbolic Layer for Trustworthy Agent Design
by: Leoveanu-Condrei, Claudiu
Published: (2025) -
Low-Resource Authorship Style Transfer: Can Non-Famous Authors Be Imitated?
by: Patel, Ajay, et al.
Published: (2022) -
Linear Alignment of Vision-language Models for Image Captioning
by: Paischer, Fabian, et al.
Published: (2023)