Saved in:
| Main Authors: | Yuan, Peng, Yin, Yuyang, Cai, Yuxuan, Wei, Zheng |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2604.10988 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Weblica: Scalable and Reproducible Training Environments for Visual Web Agents
by: Kar, Oğuzhan Fatih, et al.
Published: (2026)
by: Kar, Oğuzhan Fatih, et al.
Published: (2026)
Realism Control One-step Diffusion for Real-World Image Super-Resolution
by: Wu, Zongliang, et al.
Published: (2025)
by: Wu, Zongliang, et al.
Published: (2025)
WEBEYETRACK: Scalable Eye-Tracking for the Browser via On-Device Few-Shot Personalization
by: Davalos, Eduardo, et al.
Published: (2025)
by: Davalos, Eduardo, et al.
Published: (2025)
InfiniteWeb: Scalable Web Environment Synthesis for GUI Agent Training
by: Zhang, Ziyun, et al.
Published: (2026)
by: Zhang, Ziyun, et al.
Published: (2026)
SpatialForge: Bootstrapping 3D-Aware Spatial Reasoning from Open-World 2D Images
by: Liu, Zishan, et al.
Published: (2026)
by: Liu, Zishan, et al.
Published: (2026)
ProductWebGen: Benchmarking Multimodal Product Webpage Generation
by: Liu, Zhihong, et al.
Published: (2026)
by: Liu, Zhihong, et al.
Published: (2026)
PyVision-RL: Forging Open Agentic Vision Models via RL
by: Zhao, Shitian, et al.
Published: (2026)
by: Zhao, Shitian, et al.
Published: (2026)
Ego2Web: A Web Agent Benchmark Grounded in Egocentric Videos
by: Yu, Shoubin, et al.
Published: (2026)
by: Yu, Shoubin, et al.
Published: (2026)
ChronoForge-RL: Chronological Forging through Reinforcement Learning for Enhanced Video Understanding
by: Chen, Kehua
Published: (2025)
by: Chen, Kehua
Published: (2025)
WebGuard: Building a Generalizable Guardrail for Web Agents
by: Zheng, Boyuan, et al.
Published: (2025)
by: Zheng, Boyuan, et al.
Published: (2025)
Benchmarking Interaction, Beyond Policy: a Reproducible Benchmark for Collaborative Instance Object Navigation
by: Zorzi, Edoardo, et al.
Published: (2026)
by: Zorzi, Edoardo, et al.
Published: (2026)
On the Diversity and Realism of Distilled Dataset: An Efficient Dataset Distillation Paradigm
by: Sun, Peng, et al.
Published: (2023)
by: Sun, Peng, et al.
Published: (2023)
Tri-Modal Motion Retrieval by Learning a Joint Embedding Space
by: Yin, Kangning, et al.
Published: (2024)
by: Yin, Kangning, et al.
Published: (2024)
Video2Game: Real-time, Interactive, Realistic and Browser-Compatible Environment from a Single Video
by: Xia, Hongchi, et al.
Published: (2024)
by: Xia, Hongchi, et al.
Published: (2024)
ImAgent: A Unified Multimodal Agent Framework for Test-Time Scalable Image Generation
by: Wang, Kaishen, et al.
Published: (2025)
by: Wang, Kaishen, et al.
Published: (2025)
Beyond the Generative Learning Trilemma: Generative Model Assessment in Data Scarcity Domains
by: Salmè, Marco, et al.
Published: (2025)
by: Salmè, Marco, et al.
Published: (2025)
Devling into Adversarial Transferability on Image Classification: Review, Benchmark, and Evaluation
by: Wang, Xiaosen, et al.
Published: (2026)
by: Wang, Xiaosen, et al.
Published: (2026)
WebSight: A Vision-First Architecture for Robust Web Agents
by: Bhathal, Tanvir, et al.
Published: (2025)
by: Bhathal, Tanvir, et al.
Published: (2025)
BodyMetric: Evaluating the Realism of Human Bodies in Text-to-Image Generation
by: Andreou, Nefeli, et al.
Published: (2024)
by: Andreou, Nefeli, et al.
Published: (2024)
Disentangling Foreground and Background Motion for Enhanced Realism in Human Video Generation
by: Liu, Jinlin, et al.
Published: (2024)
by: Liu, Jinlin, et al.
Published: (2024)
MM-WebAgent: A Hierarchical Multimodal Web Agent for Webpage Generation
by: Li, Yan, et al.
Published: (2026)
by: Li, Yan, et al.
Published: (2026)
Breaking Self-Attention Failure: Rethinking Query Initialization for Infrared Small Target Detection
by: Liu, Yuteng, et al.
Published: (2026)
by: Liu, Yuteng, et al.
Published: (2026)
SAFIRE: Segment Any Forged Image Region
by: Kwon, Myung-Joon, et al.
Published: (2024)
by: Kwon, Myung-Joon, et al.
Published: (2024)
DeepForgeSeal: Latent Space-Driven Semi-Fragile Watermarking for Deepfake Detection Using Multi-Agent Adversarial Reinforcement Learning
by: Fernando, Tharindu, et al.
Published: (2025)
by: Fernando, Tharindu, et al.
Published: (2025)
Zero-shot Synthetic Video Realism Enhancement via Structure-aware Denoising
by: Wang, Yifan, et al.
Published: (2025)
by: Wang, Yifan, et al.
Published: (2025)
VISTA: An End-to-End Benchmark for Visual Spec-to-Web-App Coding Agents
by: Guo, JunJia, et al.
Published: (2026)
by: Guo, JunJia, et al.
Published: (2026)
Watching, Reasoning, and Searching: A Video Deep Research Benchmark on Open Web for Agentic Video Reasoning
by: Liu, Chengwen, et al.
Published: (2026)
by: Liu, Chengwen, et al.
Published: (2026)
VideoWebArena: Evaluating Long Context Multimodal Agents with Video Understanding Web Tasks
by: Jang, Lawrence, et al.
Published: (2024)
by: Jang, Lawrence, et al.
Published: (2024)
ForgeVLA: Federated Vision-Language-Action Learning without Language Annotations
by: Zhou, Yuhao, et al.
Published: (2026)
by: Zhou, Yuhao, et al.
Published: (2026)
"PhyWorldBench": A Comprehensive Evaluation of Physical Realism in Text-to-Video Models
by: Gu, Jing, et al.
Published: (2025)
by: Gu, Jing, et al.
Published: (2025)
Pareto-Enhanced Portrait Generation: Vision-Aligned Text Supervision for Alignment, Realism, and Aesthetics
by: Wang, Yunlong, et al.
Published: (2026)
by: Wang, Yunlong, et al.
Published: (2026)
Is Visual Realism Enough? Evaluating Gait Biometric Fidelity in Generative AI Human Animation
by: DeAndres-Tame, Ivan, et al.
Published: (2025)
by: DeAndres-Tame, Ivan, et al.
Published: (2025)
Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction
by: Tian, Keyu, et al.
Published: (2024)
by: Tian, Keyu, et al.
Published: (2024)
RealCompo: Balancing Realism and Compositionality Improves Text-to-Image Diffusion Models
by: Zhang, Xinchen, et al.
Published: (2024)
by: Zhang, Xinchen, et al.
Published: (2024)
Embodied Web Agents: Bridging Physical-Digital Realms for Integrated Agent Intelligence
by: Hong, Yining, et al.
Published: (2025)
by: Hong, Yining, et al.
Published: (2025)
FedOnco-Bench: A Reproducible Benchmark for Privacy-Aware Federated Tumor Segmentation with Synthetic CT Data
by: Marella, Viswa Chaitanya, et al.
Published: (2025)
by: Marella, Viswa Chaitanya, et al.
Published: (2025)
ReLE: A Scalable System and Structured Benchmark for Diagnosing Capability Anisotropy in Chinese LLMs
by: Fang, Rui, et al.
Published: (2026)
by: Fang, Rui, et al.
Published: (2026)
StyleForge: Enhancing Text-to-Image Synthesis for Any Artistic Styles with Dual Binding
by: Park, Junseo, et al.
Published: (2024)
by: Park, Junseo, et al.
Published: (2024)
Towards Scalable Foundation Model for Multi-modal and Hyperspectral Geospatial Data
by: Si, Haozhe, et al.
Published: (2025)
by: Si, Haozhe, et al.
Published: (2025)
v-HUB: A Benchmark for Video Humor Understanding from Vision and Sound
by: Shi, Zhengpeng, et al.
Published: (2025)
by: Shi, Zhengpeng, et al.
Published: (2025)
Similar Items
-
Weblica: Scalable and Reproducible Training Environments for Visual Web Agents
by: Kar, Oğuzhan Fatih, et al.
Published: (2026) -
Realism Control One-step Diffusion for Real-World Image Super-Resolution
by: Wu, Zongliang, et al.
Published: (2025) -
WEBEYETRACK: Scalable Eye-Tracking for the Browser via On-Device Few-Shot Personalization
by: Davalos, Eduardo, et al.
Published: (2025) -
InfiniteWeb: Scalable Web Environment Synthesis for GUI Agent Training
by: Zhang, Ziyun, et al.
Published: (2026) -
SpatialForge: Bootstrapping 3D-Aware Spatial Reasoning from Open-World 2D Images
by: Liu, Zishan, et al.
Published: (2026)