Tur[k]ingBench: A Challenge Benchmark for Web Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Kevin, Kordi, Yeganeh, Nayak, Tanay, Asija, Adi, Wang, Yizhong, Sanders, Kate, Byerly, Adam, Zhang, Jingyu, Van Durme, Benjamin, Khashabi, Daniel |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
GOLD PANNING: Strategic Context Shuffling for Needle-in-Haystack Reasoning
by: Byerly, Adam, et al.
Published: (2025)
by: Byerly, Adam, et al.
Published: (2025)
Self-Consistency Falls Short! The Adverse Effects of Positional Bias on Long-Context Problems
by: Byerly, Adam, et al.
Published: (2024)
by: Byerly, Adam, et al.
Published: (2024)
Are Finer Citations Always Better? Rethinking Granularity for Attributed Generation
by: Wang, Hexuan, et al.
Published: (2026)
by: Wang, Hexuan, et al.
Published: (2026)
Hell or High Water: Evaluating Agentic Recovery from External Failures
by: Wang, Andrew, et al.
Published: (2025)
by: Wang, Andrew, et al.
Published: (2025)
Bonsai: Interpretable Tree-Adaptive Grounded Reasoning
by: Sanders, Kate, et al.
Published: (2025)
by: Sanders, Kate, et al.
Published: (2025)
A Survey of Video Datasets for Grounded Event Understanding
by: Sanders, Kate, et al.
Published: (2024)
by: Sanders, Kate, et al.
Published: (2024)
Core: Robust Factual Precision with Informative Sub-Claim Identification
by: Jiang, Zhengping, et al.
Published: (2024)
by: Jiang, Zhengping, et al.
Published: (2024)
Verifiable by Design: Aligning Language Models to Quote from Pre-Training Data
by: Zhang, Jingyu, et al.
Published: (2024)
by: Zhang, Jingyu, et al.
Published: (2024)
Certified Mitigation of Worst-Case LLM Copyright Infringement
by: Zhang, Jingyu, et al.
Published: (2025)
by: Zhang, Jingyu, et al.
Published: (2025)
SocialNLI: A Dialogue-Centric Social Inference Dataset
by: Deo, Akhil, et al.
Published: (2025)
by: Deo, Akhil, et al.
Published: (2025)
Crystal: Characterizing Relative Impact of Scholarly Publications
by: Collison, Hannah, et al.
Published: (2026)
by: Collison, Hannah, et al.
Published: (2026)
Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirements
by: Zhang, Jingyu, et al.
Published: (2024)
by: Zhang, Jingyu, et al.
Published: (2024)
TV-TREES: Multimodal Entailment Trees for Neuro-Symbolic Video Reasoning
by: Sanders, Kate, et al.
Published: (2024)
by: Sanders, Kate, et al.
Published: (2024)
DeonticBench: A Benchmark for Reasoning over Rules
by: Dou, Guangyao, et al.
Published: (2026)
by: Dou, Guangyao, et al.
Published: (2026)
arXiv2Table: Toward Realistic Benchmarking and Evaluation for LLM-Based Literature-Review Table Generation
by: Wang, Weiqi, et al.
Published: (2025)
by: Wang, Weiqi, et al.
Published: (2025)
Hidden in the Haystack: Smaller Needles are More Difficult for LLMs to Find
by: Bianchi, Owen, et al.
Published: (2025)
by: Bianchi, Owen, et al.
Published: (2025)
Insights into LLM Long-Context Failures: When Transformers Know but Don't Tell
by: Lu, Taiming, et al.
Published: (2024)
by: Lu, Taiming, et al.
Published: (2024)
WorldAPIs: The World Is Worth How Many APIs? A Thought Experiment
by: Ou, Jiefu, et al.
Published: (2024)
by: Ou, Jiefu, et al.
Published: (2024)
Many-Tier Instruction Hierarchy in LLM Agents
by: Zhang, Jingyu, et al.
Published: (2026)
by: Zhang, Jingyu, et al.
Published: (2026)
Revisiting Generalization Across Difficulty Levels: It's Not So Easy
by: Kordi, Yeganeh, et al.
Published: (2025)
by: Kordi, Yeganeh, et al.
Published: (2025)
Jailbreak Distillation: Renewable Safety Benchmarking
by: Zhang, Jingyu, et al.
Published: (2025)
by: Zhang, Jingyu, et al.
Published: (2025)
SELF-[IN]CORRECT: LLMs Struggle with Discriminating Self-Generated Responses
by: Jiang, Dongwei, et al.
Published: (2024)
by: Jiang, Dongwei, et al.
Published: (2024)
BenchCLAMP: A Benchmark for Evaluating Language Models on Syntactic and Semantic Parsing
by: Roy, Subhro, et al.
Published: (2022)
by: Roy, Subhro, et al.
Published: (2022)
k-SemStamp: A Clustering-Based Semantic Watermark for Detection of Machine-Generated Text
by: Hou, Abe Bohan, et al.
Published: (2024)
by: Hou, Abe Bohan, et al.
Published: (2024)
RATIONALYST: Mining Implicit Rationales for Process Supervision of Reasoning
by: Jiang, Dongwei, et al.
Published: (2024)
by: Jiang, Dongwei, et al.
Published: (2024)
RORA: Robust Free-Text Rationale Evaluation
by: Jiang, Zhengping, et al.
Published: (2024)
by: Jiang, Zhengping, et al.
Published: (2024)
Dated Data: Tracing Knowledge Cutoffs in Large Language Models
by: Cheng, Jeffrey, et al.
Published: (2024)
by: Cheng, Jeffrey, et al.
Published: (2024)
TurBLiMP: A Turkish Benchmark of Linguistic Minimal Pairs
by: Başar, Ezgi, et al.
Published: (2025)
by: Başar, Ezgi, et al.
Published: (2025)
Challenging the Evaluator: LLM Sycophancy Under User Rebuttal
by: Kim, Sungwon, et al.
Published: (2025)
by: Kim, Sungwon, et al.
Published: (2025)
Grounding Partially-Defined Events in Multimodal Data
by: Sanders, Kate, et al.
Published: (2024)
by: Sanders, Kate, et al.
Published: (2024)
"According to ...": Prompting Language Models Improves Quoting from Pre-Training Data
by: Weller, Orion, et al.
Published: (2023)
by: Weller, Orion, et al.
Published: (2023)
DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models
by: Tiwari, Utkarsh, et al.
Published: (2025)
by: Tiwari, Utkarsh, et al.
Published: (2025)
The Alignment Waltz: Jointly Training Agents to Collaborate for Safety
by: Zhang, Jingyu, et al.
Published: (2025)
by: Zhang, Jingyu, et al.
Published: (2025)
MultiVENT 2.0: A Massive Multilingual Benchmark for Event-Centric Video Retrieval
by: Kriz, Reno, et al.
Published: (2024)
by: Kriz, Reno, et al.
Published: (2024)
WikiVideo: Article Generation from Multiple Videos
by: Martin, Alexander, et al.
Published: (2025)
by: Martin, Alexander, et al.
Published: (2025)
Seeing Through the MiRAGE: Evaluating Multimodal Retrieval Augmented Generation
by: Martin, Alexander, et al.
Published: (2025)
by: Martin, Alexander, et al.
Published: (2025)
Computer Cache. Online Recess--Web Games for Play and Fun
by: Byerly, Greg, et al.
Published: (2005)
by: Byerly, Greg, et al.
Published: (2005)
LLMs Provide Unstable Answers to Legal Questions
by: Blair-Stanek, Andrew, et al.
Published: (2025)
by: Blair-Stanek, Andrew, et al.
Published: (2025)
SemStamp: A Semantic Watermark with Paraphrastic Robustness for Text Generation
by: Hou, Abe Bohan, et al.
Published: (2023)
by: Hou, Abe Bohan, et al.
Published: (2023)
TurQUaz at CheckThat! 2025: Debating Large Language Models for Scientific Web Discourse Detection
by: Saraç, Tarık, et al.
Published: (2025)
by: Saraç, Tarık, et al.
Published: (2025)
Similar Items
-
GOLD PANNING: Strategic Context Shuffling for Needle-in-Haystack Reasoning
by: Byerly, Adam, et al.
Published: (2025) -
Self-Consistency Falls Short! The Adverse Effects of Positional Bias on Long-Context Problems
by: Byerly, Adam, et al.
Published: (2024) -
Are Finer Citations Always Better? Rethinking Granularity for Attributed Generation
by: Wang, Hexuan, et al.
Published: (2026) -
Hell or High Water: Evaluating Agentic Recovery from External Failures
by: Wang, Andrew, et al.
Published: (2025) -
Bonsai: Interpretable Tree-Adaptive Grounded Reasoning
by: Sanders, Kate, et al.
Published: (2025)