Cookie-Bench: Continuous On-screen Key Interaction Evaluation for Web Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Haoyue, Shen, Zhangxiao, Ding, Fan, Lou, Hangting, Kou, Yifeng, Yu, Haoqing, Li, Jingyao, Wu, Zhengfan, Bao, Siqi, Liu, Jing, Wu, Hua |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
I-WebGenBench : Evaluating Interactivity in LLM-Generated Scientific Web Applications
by: Dai, Dasen, et al.
Published: (2026)
by: Dai, Dasen, et al.
Published: (2026)
Chain-of-Programming (CoP) : Empowering Large Language Models for Geospatial Code Generation
by: Hou, Shuyang, et al.
Published: (2024)
by: Hou, Shuyang, et al.
Published: (2024)
Intractable Cookie Crumbs: Unveiling the Nexus of Stateful Banner Interaction and Tracking Cookies
by: Rasaii, Ali, et al.
Published: (2025)
by: Rasaii, Ali, et al.
Published: (2025)
Geo-FuB: A Method for Constructing an Operator-Function Knowledge Base for Geospatial Code Generation Tasks Using Large Language Models
by: Hou, Shuyang, et al.
Published: (2024)
by: Hou, Shuyang, et al.
Published: (2024)
WAKE: Watermarking Audio with Key Enrichment
by: Xu, Yaoxun, et al.
Published: (2025)
by: Xu, Yaoxun, et al.
Published: (2025)
Invisible, Unreadable, and Inaudible Cookie Notices: An Evaluation of Cookie Notices for Users with Visual Impairments
by: Clarke, James M., et al.
Published: (2023)
by: Clarke, James M., et al.
Published: (2023)
AutoGEEval++: A Multi-Level and Multi-Geospatial-Modality Automated Evaluation Framework for Large Language Models in Geospatial Code Generation on Google Earth Engine
by: Hou, Shuyang, et al.
Published: (2025)
by: Hou, Shuyang, et al.
Published: (2025)
AutoGEEval: A Multimodal and Automated Framework for Geospatial Code Generation on GEE with Large Language Models
by: Hou, Shuyang, et al.
Published: (2025)
by: Hou, Shuyang, et al.
Published: (2025)
X-WebAgentBench: A Multilingual Interactive Web Benchmark for Evaluating Global Agentic System
by: Wang, Peng, et al.
Published: (2025)
by: Wang, Peng, et al.
Published: (2025)
Not All Tokens Learn Alike: Attention Entropy Reveals Heterogeneous Signals in RL Reasoning
by: Li, Gengyang, et al.
Published: (2026)
by: Li, Gengyang, et al.
Published: (2026)
PQDT: Pseudo-Query Dual Transformer for Robust Point Cloud Restoration
by: Wu, Haoqing, et al.
Published: (2026)
by: Wu, Haoqing, et al.
Published: (2026)
Design and Reachable Domain Analysis of On‐orbit Electromagnetic Launcher for CubeSats
by: Haoqing Feng, et al.
Published: (2025)
by: Haoqing Feng, et al.
Published: (2025)
New Advances in Periodontal Functional Materials Based on Antibacterial, Anti‐Inflammatory, and Tissue Regeneration Strategies
by: Haoyue Wu, et al.
Published: (2025)
by: Haoyue Wu, et al.
Published: (2025)
SiMing-Bench: Evaluating Procedural Correctness from Continuous Interactions in Clinical Skill Videos
by: Huang, Xiyang, et al.
Published: (2026)
by: Huang, Xiyang, et al.
Published: (2026)
Demystifying Cookie Sharing Risks in WebView-based Mobile App-in-app Ecosystems
by: Zhang, Miao, et al.
Published: (2025)
by: Zhang, Miao, et al.
Published: (2025)
Can Large Language Models Generate Geospatial Code?
by: Hou, Shuyang, et al.
Published: (2024)
by: Hou, Shuyang, et al.
Published: (2024)
WebGen-Bench: Evaluating LLMs on Generating Interactive and Functional Websites from Scratch
by: Lu, Zimu, et al.
Published: (2025)
by: Lu, Zimu, et al.
Published: (2025)
Cookies, Coleslaw, and Stoops
by: van der Sijs, Nicoline
Published: (2010)
by: van der Sijs, Nicoline
Published: (2010)
The Proof is in the Almond Cookies
by: van Trijp, Remi, et al.
Published: (2025)
by: van Trijp, Remi, et al.
Published: (2025)
Inflammasomes and their roles in autoimmune diseases
by: Minghui Pan, et al.
Published: (2024)
by: Minghui Pan, et al.
Published: (2024)
Continuous Target Speech Extraction: Enhancing Personalized Diarization and Extraction on Complex Recordings
by: Zhao, He, et al.
Published: (2024)
by: Zhao, He, et al.
Published: (2024)
AgentWebBench: Benchmarking Multi-Agent Coordination in Agentic Web
by: Zhong, Shanshan, et al.
Published: (2026)
by: Zhong, Shanshan, et al.
Published: (2026)
AudioComposer: Towards Fine-grained Audio Generation with Natural Language Descriptions
by: Wang, Yuanyuan, et al.
Published: (2024)
by: Wang, Yuanyuan, et al.
Published: (2024)
Dual-Constrained Dynamical Neural ODEs for Ambiguity-aware Continuous Emotion Prediction
by: Wu, Jingyao, et al.
Published: (2024)
by: Wu, Jingyao, et al.
Published: (2024)
Hybrid-Domain Synergistic Transformer for Hyperspectral Image Denoising
by: Li, Haoyue, et al.
Published: (2025)
by: Li, Haoyue, et al.
Published: (2025)
StressWeb: A Diagnostic Benchmark for Web Agent Robustness under Realistic Interaction Variability
by: Bai, Haoyue, et al.
Published: (2026)
by: Bai, Haoyue, et al.
Published: (2026)
Crumbled Cookie Exploring E-commerce Websites Cookie Policies with Data Protection Regulations
by: Singh, Nivedita, et al.
Published: (2024)
by: Singh, Nivedita, et al.
Published: (2024)
Browsing without Third-Party Cookies: What Do You See?
by: Lin, Maxwell, et al.
Published: (2024)
by: Lin, Maxwell, et al.
Published: (2024)
DeepWeb-Bench: A Deep Research Benchmark Demanding Massive Cross-Source Evidence and Long-Horizon Derivation
by: Xie, Sixiong, et al.
Published: (2026)
by: Xie, Sixiong, et al.
Published: (2026)
Pass the Cookies and Uphold the Privacy.
by: Guenther, Kim
Published: (2001)
by: Guenther, Kim
Published: (2001)
Omni-WorldBench: Towards a Comprehensive Interaction-Centric Evaluation for World Models
by: Wu, Meiqi, et al.
Published: (2026)
by: Wu, Meiqi, et al.
Published: (2026)
Securing Web Applications Against Vulnerabilities Using Quantum Cryptography and Reverse Proxy Authentication for Cookie Protection
by: Dr.G.Aravind Swaminathan, Ajitha Devadharshini B
Published: (2026)
by: Dr.G.Aravind Swaminathan, Ajitha Devadharshini B
Published: (2026)
GeoCode-GPT: A Large Language Model for Geospatial Code Generation Tasks
by: Hou, Shuyang, et al.
Published: (2024)
by: Hou, Shuyang, et al.
Published: (2024)
Rethinking Supervised Fine-Tuning: Emphasizing Key Answer Tokens for Improved LLM Accuracy
by: Shi, Xiaofeng, et al.
Published: (2025)
by: Shi, Xiaofeng, et al.
Published: (2025)
AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories
by: Lù, Xing Han, et al.
Published: (2025)
by: Lù, Xing Han, et al.
Published: (2025)
Dataset Cookie 20 Museum Indonesia
by: Maharani, Kayla Putri, et al.
Published: (2025)
by: Maharani, Kayla Putri, et al.
Published: (2025)
Cookie cutters: Bisections with fixed shapes
by: Schnider, Patrick, et al.
Published: (2025)
by: Schnider, Patrick, et al.
Published: (2025)
ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents
by: Levy, Ido, et al.
Published: (2024)
by: Levy, Ido, et al.
Published: (2024)
WebCoderBench: Benchmarking Web Application Generation with Comprehensive and Interpretable Evaluation Metrics
by: Liu, Chenxu, et al.
Published: (2026)
by: Liu, Chenxu, et al.
Published: (2026)
LiveClawBench: Benchmarking LLM Agents on Complex, Real-World Assistant Tasks
by: Long, Xiang, et al.
Published: (2026)
by: Long, Xiang, et al.
Published: (2026)
Similar Items
-
I-WebGenBench : Evaluating Interactivity in LLM-Generated Scientific Web Applications
by: Dai, Dasen, et al.
Published: (2026) -
Chain-of-Programming (CoP) : Empowering Large Language Models for Geospatial Code Generation
by: Hou, Shuyang, et al.
Published: (2024) -
Intractable Cookie Crumbs: Unveiling the Nexus of Stateful Banner Interaction and Tracking Cookies
by: Rasaii, Ali, et al.
Published: (2025) -
Geo-FuB: A Method for Constructing an Operator-Function Knowledge Base for Geospatial Code Generation Tasks Using Large Language Models
by: Hou, Shuyang, et al.
Published: (2024) -
WAKE: Watermarking Audio with Key Enrichment
by: Xu, Yaoxun, et al.
Published: (2025)