WebCanvas: Benchmarking Web Agents in Online Environments
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pan, Yichen, Kong, Dehan, Zhou, Sida, Cui, Cheng, Leng, Yifei, Jiang, Bing, Liu, Hangyu, Shang, Yanyi, Zhou, Shuyan, Wu, Tongshuang, Wu, Zhengyang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
WebChain: A Large-Scale Human-Annotated Dataset of Real-World Web Interaction Traces
von: Fan, Sicheng, et al.
Veröffentlicht: (2026)
von: Fan, Sicheng, et al.
Veröffentlicht: (2026)
Mapping the Web of Science, a large-scale graph and text-based dataset with LLM embeddings
von: Kunt, Tim, et al.
Veröffentlicht: (2026)
von: Kunt, Tim, et al.
Veröffentlicht: (2026)
WeDLM: Reconciling Diffusion Language Models with Standard Causal Attention for Fast Inference
von: Liu, Aiwei, et al.
Veröffentlicht: (2025)
von: Liu, Aiwei, et al.
Veröffentlicht: (2025)
XPath Agent: An Efficient XPath Programming Agent Based on LLM for Web Crawler
von: Li, Yu, et al.
Veröffentlicht: (2024)
von: Li, Yu, et al.
Veröffentlicht: (2024)
WaterSeeker: Pioneering Efficient Detection of Watermarked Segments in Large Documents
von: Pan, Leyi, et al.
Veröffentlicht: (2024)
von: Pan, Leyi, et al.
Veröffentlicht: (2024)
Omni-SafetyBench: A Benchmark for Safety Evaluation of Audio-Visual Large Language Models
von: Pan, Leyi, et al.
Veröffentlicht: (2025)
von: Pan, Leyi, et al.
Veröffentlicht: (2025)
MicroRemed: Benchmarking LLMs in Microservices Remediation
von: Zhang, Lingzhe, et al.
Veröffentlicht: (2025)
von: Zhang, Lingzhe, et al.
Veröffentlicht: (2025)
A Knowledge Enhanced Learning and Semantic Composition Model for Multi-Claim Fact Checking
von: Wang, Shuai, et al.
Veröffentlicht: (2021)
von: Wang, Shuai, et al.
Veröffentlicht: (2021)
AI-assisted German Employment Contract Review: A Benchmark Dataset
von: Wardas, Oliver, et al.
Veröffentlicht: (2025)
von: Wardas, Oliver, et al.
Veröffentlicht: (2025)
Compression Method Matters: Benchmark-Dependent Output Dynamics in LLM Prompt Compression
von: Johnson, Warren
Veröffentlicht: (2026)
von: Johnson, Warren
Veröffentlicht: (2026)
Parametric Social Identity Injection and Diversification in Public Opinion Simulation
von: Wang, Hexi, et al.
Veröffentlicht: (2026)
von: Wang, Hexi, et al.
Veröffentlicht: (2026)
Can Watermarked LLMs be Identified by Users via Crafted Prompts?
von: Liu, Aiwei, et al.
Veröffentlicht: (2024)
von: Liu, Aiwei, et al.
Veröffentlicht: (2024)
Distractor Injection Attacks on Large Reasoning Models: Characterization and Defense
von: Zhang, Zhehao, et al.
Veröffentlicht: (2025)
von: Zhang, Zhehao, et al.
Veröffentlicht: (2025)
SAGED: A Holistic Bias-Benchmarking Pipeline for Language Models with Customisable Fairness Calibration
von: Guan, Xin, et al.
Veröffentlicht: (2024)
von: Guan, Xin, et al.
Veröffentlicht: (2024)
The Superalignment of Superhuman Intelligence with Large Language Models
von: Huang, Minlie, et al.
Veröffentlicht: (2024)
von: Huang, Minlie, et al.
Veröffentlicht: (2024)
Direct Large Language Model Alignment Through Self-Rewarding Contrastive Prompt Distillation
von: Liu, Aiwei, et al.
Veröffentlicht: (2024)
von: Liu, Aiwei, et al.
Veröffentlicht: (2024)
WebMap -- Large Language Model-assisted Semantic Link Induction in the Web
von: Pokharel, Shiraj, et al.
Veröffentlicht: (2025)
von: Pokharel, Shiraj, et al.
Veröffentlicht: (2025)
MELoRA: Mini-Ensemble Low-Rank Adapters for Parameter-Efficient Fine-Tuning
von: Ren, Pengjie, et al.
Veröffentlicht: (2024)
von: Ren, Pengjie, et al.
Veröffentlicht: (2024)
Exploiting Web Search Tools of AI Agents for Data Exfiltration
von: Rall, Dennis, et al.
Veröffentlicht: (2025)
von: Rall, Dennis, et al.
Veröffentlicht: (2025)
An Unforgeable Publicly Verifiable Watermark for Large Language Models
von: Liu, Aiwei, et al.
Veröffentlicht: (2023)
von: Liu, Aiwei, et al.
Veröffentlicht: (2023)
TIS-DPO: Token-level Importance Sampling for Direct Preference Optimization With Estimated Weights
von: Liu, Aiwei, et al.
Veröffentlicht: (2024)
von: Liu, Aiwei, et al.
Veröffentlicht: (2024)
Can LLM Watermarks Robustly Prevent Unauthorized Knowledge Distillation?
von: Pan, Leyi, et al.
Veröffentlicht: (2025)
von: Pan, Leyi, et al.
Veröffentlicht: (2025)
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark
von: Cheng, Ziming, et al.
Veröffentlicht: (2025)
von: Cheng, Ziming, et al.
Veröffentlicht: (2025)
A Survey of Text Watermarking in the Era of Large Language Models
von: Liu, Aiwei, et al.
Veröffentlicht: (2023)
von: Liu, Aiwei, et al.
Veröffentlicht: (2023)
MORABLES: A Benchmark for Assessing Abstract Moral Reasoning in LLMs with Fables
von: Marcuzzo, Matteo, et al.
Veröffentlicht: (2025)
von: Marcuzzo, Matteo, et al.
Veröffentlicht: (2025)
d-TreeRPO: Towards More Reliable Policy Optimization for Diffusion Language Models
von: Pan, Leyi, et al.
Veröffentlicht: (2025)
von: Pan, Leyi, et al.
Veröffentlicht: (2025)
MarkLLM: An Open-Source Toolkit for LLM Watermarking
von: Pan, Leyi, et al.
Veröffentlicht: (2024)
von: Pan, Leyi, et al.
Veröffentlicht: (2024)
LayerTracer: A Joint Task-Particle and Vulnerable-Layer Analysis framework for Arbitrary Large Language Model Architectures
von: Wu, Yuhang, et al.
Veröffentlicht: (2026)
von: Wu, Yuhang, et al.
Veröffentlicht: (2026)
CIKT: A Collaborative and Iterative Knowledge Tracing Framework with Large Language Models
von: Li, Runze, et al.
Veröffentlicht: (2025)
von: Li, Runze, et al.
Veröffentlicht: (2025)
Towards Effective and Efficient Continual Pre-training of Large Language Models
von: Chen, Jie, et al.
Veröffentlicht: (2024)
von: Chen, Jie, et al.
Veröffentlicht: (2024)
PaperAudit-Bench: Benchmarking Error Detection in Research Papers for Critical Automated Peer Review
von: Tu, Songjun, et al.
Veröffentlicht: (2026)
von: Tu, Songjun, et al.
Veröffentlicht: (2026)
Sure! Here's a short and concise title for your paper: "Contamination in Generated Text Detection Benchmarks"
von: Dingfelder, Philipp, et al.
Veröffentlicht: (2025)
von: Dingfelder, Philipp, et al.
Veröffentlicht: (2025)
IFMTBench: A Comprehensive Benchmark for Multilingual Translation Instruction Following
von: Sun, Mingrui, et al.
Veröffentlicht: (2026)
von: Sun, Mingrui, et al.
Veröffentlicht: (2026)
Prompt Engineering and the Effectiveness of Large Language Models in Enhancing Human Productivity
von: Anam, Rizal Khoirul
Veröffentlicht: (2025)
von: Anam, Rizal Khoirul
Veröffentlicht: (2025)
Softmax Linear Attention: Reclaiming Global Competition
von: Xu, Mingwei, et al.
Veröffentlicht: (2026)
von: Xu, Mingwei, et al.
Veröffentlicht: (2026)
WebFactory: Automated Compression of Foundational Language Intelligence into Grounded Web Agents
von: Fan, Sicheng, et al.
Veröffentlicht: (2026)
von: Fan, Sicheng, et al.
Veröffentlicht: (2026)
A Semantic Invariant Robust Watermark for Large Language Models
von: Liu, Aiwei, et al.
Veröffentlicht: (2023)
von: Liu, Aiwei, et al.
Veröffentlicht: (2023)
Homogenization of Non-homogeneous Incompressible Navier-Stokes System in Critically Perforated Domains
von: Pan, Jiaojiao
Veröffentlicht: (2024)
von: Pan, Jiaojiao
Veröffentlicht: (2024)
Raw Text is All you Need: Knowledge-intensive Multi-turn Instruction Tuning for Large Language Model
von: Hou, Xia, et al.
Veröffentlicht: (2024)
von: Hou, Xia, et al.
Veröffentlicht: (2024)
Bug In the Code Stack: Can LLMs Find Bugs in Large Python Code Stacks
von: Lee, Hokyung, et al.
Veröffentlicht: (2024)
von: Lee, Hokyung, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
WebChain: A Large-Scale Human-Annotated Dataset of Real-World Web Interaction Traces
von: Fan, Sicheng, et al.
Veröffentlicht: (2026) -
Mapping the Web of Science, a large-scale graph and text-based dataset with LLM embeddings
von: Kunt, Tim, et al.
Veröffentlicht: (2026) -
WeDLM: Reconciling Diffusion Language Models with Standard Causal Attention for Fast Inference
von: Liu, Aiwei, et al.
Veröffentlicht: (2025) -
XPath Agent: An Efficient XPath Programming Agent Based on LLM for Web Crawler
von: Li, Yu, et al.
Veröffentlicht: (2024) -
WaterSeeker: Pioneering Efficient Detection of Watermarked Segments in Large Documents
von: Pan, Leyi, et al.
Veröffentlicht: (2024)