VideoWebArena: Evaluating Long Context Multimodal Agents with Video Understanding Web Tasks
Fuente:
arXiv
Saved in:
| Main Authors: | Jang, Lawrence, Li, Yinheng, Zhao, Dan, Ding, Charles, Lin, Justin, Liang, Paul Pu, Bonatti, Rogerio, Koishida, Kazuhito |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale
by: Bonatti, Rogerio, et al.
Published: (2024)
by: Bonatti, Rogerio, et al.
Published: (2024)
Data Generation Using Large Language Models for Text Classification: An Empirical Case Study
by: Li, Yinheng, et al.
Published: (2024)
by: Li, Yinheng, et al.
Published: (2024)
Instruction Agent: Enhancing Agent with Expert Demonstration
by: Li, Yinheng, et al.
Published: (2025)
by: Li, Yinheng, et al.
Published: (2025)
VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
by: Koh, Jing Yu, et al.
Published: (2024)
by: Koh, Jing Yu, et al.
Published: (2024)
WinClick: GUI Grounding with Multimodal Large Language Models
by: Hui, Zheng, et al.
Published: (2025)
by: Hui, Zheng, et al.
Published: (2025)
Hierarchical Self-Attention: Generalizing Neural Attention Mechanics to Multi-Scale Problems
by: Amizadeh, Saeed, et al.
Published: (2025)
by: Amizadeh, Saeed, et al.
Published: (2025)
WebChoreArena: Evaluating Web Browsing Agents on Realistic Tedious Web Tasks
by: Miyai, Atsuyuki, et al.
Published: (2025)
by: Miyai, Atsuyuki, et al.
Published: (2025)
Odysseys: Benchmarking Web Agents on Realistic Long Horizon Tasks
by: Jang, Lawrence Keunho, et al.
Published: (2026)
by: Jang, Lawrence Keunho, et al.
Published: (2026)
BrowserArena: Evaluating LLM Agents on Real-World Web Navigation Tasks
by: Anupam, Sagnik, et al.
Published: (2025)
by: Anupam, Sagnik, et al.
Published: (2025)
EconWebArena: Benchmarking Autonomous Agents on Economic Tasks in Realistic Web Environments
by: Liu, Zefang, et al.
Published: (2025)
by: Liu, Zefang, et al.
Published: (2025)
SafeArena: Evaluating the Safety of Autonomous Web Agents
by: Tur, Ada Defne, et al.
Published: (2025)
by: Tur, Ada Defne, et al.
Published: (2025)
Evaluating Long-Context Reasoning in LLM-Based WebAgents
by: Chung, Andy, et al.
Published: (2025)
by: Chung, Andy, et al.
Published: (2025)
SecureWebArena: A Holistic Security Evaluation Benchmark for LVLM-based Web Agents
by: Ying, Zonghao, et al.
Published: (2025)
by: Ying, Zonghao, et al.
Published: (2025)
WebArena: A Realistic Web Environment for Building Autonomous Agents
by: Zhou, Shuyan, et al.
Published: (2023)
by: Zhou, Shuyan, et al.
Published: (2023)
Single-channel speech enhancement using learnable loss mixup
by: Chang, Oscar, et al.
Published: (2023)
by: Chang, Oscar, et al.
Published: (2023)
Learned Image Compression with Text Quality Enhancement
by: Lai, Chih-Yu, et al.
Published: (2024)
by: Lai, Chih-Yu, et al.
Published: (2024)
WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?
by: Drouin, Alexandre, et al.
Published: (2024)
by: Drouin, Alexandre, et al.
Published: (2024)
A Real-World WebAgent with Planning, Long Context Understanding, and Program Synthesis
by: Gur, Izzeddin, et al.
Published: (2023)
by: Gur, Izzeddin, et al.
Published: (2023)
Ego2Web: A Web Agent Benchmark Grounded in Egocentric Videos
by: Yu, Shoubin, et al.
Published: (2026)
by: Yu, Shoubin, et al.
Published: (2026)
VideoAutoArena: An Automated Arena for Evaluating Large Multimodal Models in Video Analysis through User Simulation
by: Luo, Ziyang, et al.
Published: (2024)
by: Luo, Ziyang, et al.
Published: (2024)
WebSP-Eval: Evaluating Web Agents on Website Security and Privacy Tasks
by: Ramesh, Guruprasad Viswanathan, et al.
Published: (2026)
by: Ramesh, Guruprasad Viswanathan, et al.
Published: (2026)
Weakly-supervised Audio Separation via Bi-modal Semantic Similarity
by: Mahmud, Tanvir, et al.
Published: (2024)
by: Mahmud, Tanvir, et al.
Published: (2024)
LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes
by: Dang, Trung, et al.
Published: (2024)
by: Dang, Trung, et al.
Published: (2024)
AgentFold: Long-Horizon Web Agents with Proactive Context Management
by: Ye, Rui, et al.
Published: (2025)
by: Ye, Rui, et al.
Published: (2025)
WebArXiv: Evaluating Multimodal Agents on Time-Invariant arXiv Tasks
by: Sun, Zihao, et al.
Published: (2025)
by: Sun, Zihao, et al.
Published: (2025)
GTA: Generating Long-Horizon Tasks for Web Agents at Scale
by: Huang, Tenghao, et al.
Published: (2026)
by: Huang, Tenghao, et al.
Published: (2026)
Data Processing Techniques for Modern Multimodal Models
by: Li, Yinheng, et al.
Published: (2024)
by: Li, Yinheng, et al.
Published: (2024)
Deep Research Bench: Evaluating AI Web Research Agents
by: FutureSearch, et al.
Published: (2025)
by: FutureSearch, et al.
Published: (2025)
VCA: Video Curious Agent for Long Video Understanding
by: Yang, Zeyuan, et al.
Published: (2024)
by: Yang, Zeyuan, et al.
Published: (2024)
CUA-Skill: Develop Skills for Computer Using Agent
by: Chen, Tianyi, et al.
Published: (2026)
by: Chen, Tianyi, et al.
Published: (2026)
Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models
by: Ma, Martin Q., et al.
Published: (2026)
by: Ma, Martin Q., et al.
Published: (2026)
AutoScraper: A Progressive Understanding Web Agent for Web Scraper Generation
by: Huang, Wenhao, et al.
Published: (2024)
by: Huang, Wenhao, et al.
Published: (2024)
DAVE: A VLM Vision Encoder for Document Understanding and Web Agents
by: Huang, Brandon, et al.
Published: (2025)
by: Huang, Brandon, et al.
Published: (2025)
VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding
by: Fan, Yue, et al.
Published: (2024)
by: Fan, Yue, et al.
Published: (2024)
Towards a Realistic Long-Term Benchmark for Open-Web Research Agents
by: Mühlbacher, Peter, et al.
Published: (2024)
by: Mühlbacher, Peter, et al.
Published: (2024)
Can Agent Conquer Web? Exploring the Frontiers of ChatGPT Atlas Agent in Web Games
by: Zhang, Jingran, et al.
Published: (2025)
by: Zhang, Jingran, et al.
Published: (2025)
Mixup Helps Understanding Multimodal Video Better
by: Ma, Xiaoyu, et al.
Published: (2025)
by: Ma, Xiaoyu, et al.
Published: (2025)
PATHWAYS: Evaluating Investigation and Context Discovery in AI Web Agents
by: Arman, Shifat E., et al.
Published: (2026)
by: Arman, Shifat E., et al.
Published: (2026)
MM-WebAgent: A Hierarchical Multimodal Web Agent for Webpage Generation
by: Li, Yan, et al.
Published: (2026)
by: Li, Yan, et al.
Published: (2026)
Understanding Long Videos with Multimodal Language Models
by: Ranasinghe, Kanchana, et al.
Published: (2024)
by: Ranasinghe, Kanchana, et al.
Published: (2024)
Similar Items
-
Windows Agent Arena: Evaluating Multi-Modal OS Agents at Scale
by: Bonatti, Rogerio, et al.
Published: (2024) -
Data Generation Using Large Language Models for Text Classification: An Empirical Case Study
by: Li, Yinheng, et al.
Published: (2024) -
Instruction Agent: Enhancing Agent with Expert Demonstration
by: Li, Yinheng, et al.
Published: (2025) -
VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
by: Koh, Jing Yu, et al.
Published: (2024) -
WinClick: GUI Grounding with Multimodal Large Language Models
by: Hui, Zheng, et al.
Published: (2025)