WebRPG: Automatic Web Rendering Parameters Generation for Visual Presentation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shao, Zirui, Gao, Feiyu, Xing, Hangdi, Zhu, Zepeng, Yu, Zhi, Bu, Jiajun, Zheng, Qi, Yao, Cong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Is Cognition Consistent with Perception? Assessing and Mitigating Multimodal Knowledge Conflicts in Document Understanding
von: Shao, Zirui, et al.
Veröffentlicht: (2024)
von: Shao, Zirui, et al.
Veröffentlicht: (2024)
BrailleLLM: Braille Instruction Tuning with Large Language Models for Braille Domain Tasks
von: Huang, Tianyuan, et al.
Veröffentlicht: (2025)
von: Huang, Tianyuan, et al.
Veröffentlicht: (2025)
A Simple yet Effective Layout Token in Large Language Models for Document Understanding
von: Zhu, Zhaoqing, et al.
Veröffentlicht: (2025)
von: Zhu, Zhaoqing, et al.
Veröffentlicht: (2025)
Towards Scalable Web Accessibility Audit with MLLMs as Copilots
von: Gu, Ming, et al.
Veröffentlicht: (2025)
von: Gu, Ming, et al.
Veröffentlicht: (2025)
LORE++: Logical Location Regression Network for Table Structure Recognition with Pre-training
von: Long, Rujiao, et al.
Veröffentlicht: (2024)
von: Long, Rujiao, et al.
Veröffentlicht: (2024)
Doc-CoB: Enhancing Document Understanding with Visual Chain-of-Boxes Reasoning
von: Mo, Ye, et al.
Veröffentlicht: (2025)
von: Mo, Ye, et al.
Veröffentlicht: (2025)
Learning Only with Images: Visual Reinforcement Learning with Reasoning, Rendering, and Visual Feedback
von: Chen, Yang, et al.
Veröffentlicht: (2025)
von: Chen, Yang, et al.
Veröffentlicht: (2025)
ProcTag: Process Tagging for Assessing the Efficacy of Document Instruction Data
von: Shen, Yufan, et al.
Veröffentlicht: (2024)
von: Shen, Yufan, et al.
Veröffentlicht: (2024)
WebRenderBench: Enhancing Web Interface Generation through Layout-Style Consistency and Reinforcement Learning
von: Lai, Peichao, et al.
Veröffentlicht: (2025)
von: Lai, Peichao, et al.
Veröffentlicht: (2025)
City-on-Web: Real-time Neural Rendering of Large-scale Scenes on the Web
von: Song, Kaiwen, et al.
Veröffentlicht: (2023)
von: Song, Kaiwen, et al.
Veröffentlicht: (2023)
Automatic Generation of Web Censorship Probe Lists
von: Tang, Jenny, et al.
Veröffentlicht: (2024)
von: Tang, Jenny, et al.
Veröffentlicht: (2024)
Automatic Welding of Corrugated Steel Webs on Composite Box Girder with Corrugated Steel Webs
von: Yunfei Wu, et al.
Veröffentlicht: (2024)
von: Yunfei Wu, et al.
Veröffentlicht: (2024)
TongUI: Internet-Scale Trajectories from Multimodal Web Tutorials for Generalized GUI Agents
von: Zhang, Bofei, et al.
Veröffentlicht: (2025)
von: Zhang, Bofei, et al.
Veröffentlicht: (2025)
Visual Text Generation in the Wild
von: Zhu, Yuanzhi, et al.
Veröffentlicht: (2024)
von: Zhu, Yuanzhi, et al.
Veröffentlicht: (2024)
Web-Based Slide Presentations.
von: Just, Melissa L.
Veröffentlicht: (1997)
von: Just, Melissa L.
Veröffentlicht: (1997)
MolmoWeb: Open Visual Web Agent and Open Data for the Open Web
von: Gupta, Tanmay, et al.
Veröffentlicht: (2026)
von: Gupta, Tanmay, et al.
Veröffentlicht: (2026)
VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
von: Koh, Jing Yu, et al.
Veröffentlicht: (2024)
von: Koh, Jing Yu, et al.
Veröffentlicht: (2024)
GuideWeb: A Benchmark for Automatic In-App Guide Generation on Real-World Web UIs
von: Gan, Chengguang, et al.
Veröffentlicht: (2026)
von: Gan, Chengguang, et al.
Veröffentlicht: (2026)
WebGen-V Bench: Structured Representation for Enhancing Visual Design in LLM-based Web Generation and Evaluation
von: Wang, Kuang-Da, et al.
Veröffentlicht: (2025)
von: Wang, Kuang-Da, et al.
Veröffentlicht: (2025)
OpenWebRL: Demystifying Online Multi-turn Reinforcement Learning for Visual Web Agents
von: Yang, Rui, et al.
Veröffentlicht: (2026)
von: Yang, Rui, et al.
Veröffentlicht: (2026)
MPDocBench-Parse: Benchmarking Practical Multi-page Document Parsing
von: Zhou, Bangbang, et al.
Veröffentlicht: (2026)
von: Zhou, Bangbang, et al.
Veröffentlicht: (2026)
GUIDE: Resolving Domain Bias in GUI Agents through Real-Time Web Video Retrieval and Plug-and-Play Annotation
von: Xie, Rui, et al.
Veröffentlicht: (2026)
von: Xie, Rui, et al.
Veröffentlicht: (2026)
The Turán number of the triangular pyramid of 4-layers
von: Chen, Hangdi, et al.
Veröffentlicht: (2026)
von: Chen, Hangdi, et al.
Veröffentlicht: (2026)
Dual-View Visual Contextualization for Web Navigation
von: Kil, Jihyung, et al.
Veröffentlicht: (2024)
von: Kil, Jihyung, et al.
Veröffentlicht: (2024)
Coverage-Aware Web Crawling for Domain-Specific Supplier Discovery via a Web--Knowledge--Web Pipeline
von: Qi, Yijiashun, et al.
Veröffentlicht: (2026)
von: Qi, Yijiashun, et al.
Veröffentlicht: (2026)
Enhancing Web Agents with a Hierarchical Memory Tree
von: Tan, Yunteng, et al.
Veröffentlicht: (2026)
von: Tan, Yunteng, et al.
Veröffentlicht: (2026)
Structured Distillation of Web Agent Capabilities Enables Generalization
von: Lù, Xing Han, et al.
Veröffentlicht: (2026)
von: Lù, Xing Han, et al.
Veröffentlicht: (2026)
WebXSkill: Skill Learning for Autonomous Web Agents
von: Wang, Zhaoyang, et al.
Veröffentlicht: (2026)
von: Wang, Zhaoyang, et al.
Veröffentlicht: (2026)
AgentRewardBench: Evaluating Automatic Evaluations of Web Agent Trajectories
von: Lù, Xing Han, et al.
Veröffentlicht: (2025)
von: Lù, Xing Han, et al.
Veröffentlicht: (2025)
Web Diagrams of Cluster Variables for Grassmannian Gr(4,8)
von: Zhang, Wen Ting, et al.
Veröffentlicht: (2025)
von: Zhang, Wen Ting, et al.
Veröffentlicht: (2025)
WebInject: Prompt Injection Attack to Web Agents
von: Wang, Xilong, et al.
Veröffentlicht: (2025)
von: Wang, Xilong, et al.
Veröffentlicht: (2025)
REAL: Resolving Knowledge Conflicts in Knowledge-Intensive Visual Question Answering via Reasoning-Pivot Alignment
von: Ye, Kai, et al.
Veröffentlicht: (2026)
von: Ye, Kai, et al.
Veröffentlicht: (2026)
OphIn-500K: Curating Web-Scale Visual Instructions for Scaling Ophthalmic Multimodal Large Language Models
von: Dong, Xuanzhao, et al.
Veröffentlicht: (2026)
von: Dong, Xuanzhao, et al.
Veröffentlicht: (2026)
Constructing a Spider‐Web Polymer Blocking Layer on Separator for the High‐Loading Li‐S Battery
von: Qian Zhang, et al.
Veröffentlicht: (2024)
von: Qian Zhang, et al.
Veröffentlicht: (2024)
WebGuard: Building a Generalizable Guardrail for Web Agents
von: Zheng, Boyuan, et al.
Veröffentlicht: (2025)
von: Zheng, Boyuan, et al.
Veröffentlicht: (2025)
WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models
von: He, Hongliang, et al.
Veröffentlicht: (2024)
von: He, Hongliang, et al.
Veröffentlicht: (2024)
Selecting a Web 2.0 Presentation Tool
von: Hodges, Charles B., et al.
Veröffentlicht: (2011)
von: Hodges, Charles B., et al.
Veröffentlicht: (2011)
RPG: A Repository Planning Graph for Unified and Scalable Codebase Generation
von: Luo, Jane, et al.
Veröffentlicht: (2025)
von: Luo, Jane, et al.
Veröffentlicht: (2025)
AI-Assisted Adaptive Rendering for High-Frequency Security Telemetry in Web Interfaces
von: Rajhans, Mona
Veröffentlicht: (2026)
von: Rajhans, Mona
Veröffentlicht: (2026)
WebGym: Scaling Training Environments for Visual Web Agents with Realistic Tasks
von: Bai, Hao, et al.
Veröffentlicht: (2026)
von: Bai, Hao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Is Cognition Consistent with Perception? Assessing and Mitigating Multimodal Knowledge Conflicts in Document Understanding
von: Shao, Zirui, et al.
Veröffentlicht: (2024) -
BrailleLLM: Braille Instruction Tuning with Large Language Models for Braille Domain Tasks
von: Huang, Tianyuan, et al.
Veröffentlicht: (2025) -
A Simple yet Effective Layout Token in Large Language Models for Document Understanding
von: Zhu, Zhaoqing, et al.
Veröffentlicht: (2025) -
Towards Scalable Web Accessibility Audit with MLLMs as Copilots
von: Gu, Ming, et al.
Veröffentlicht: (2025) -
LORE++: Logical Location Regression Network for Table Structure Recognition with Pre-training
von: Long, Rujiao, et al.
Veröffentlicht: (2024)