VisRefiner: Learning from Visual Differences for Screenshot-to-Code Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Deng, Jie, Yao, Kaichun, Zhang, Libo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Rethinking Token Pruning for Historical Screenshots in GUI Visual Agents: Semantic, Spatial, and Temporal Perspectives
von: Li, Daiqiang, et al.
Veröffentlicht: (2026)
von: Li, Daiqiang, et al.
Veröffentlicht: (2026)
VisKnow: Constructing Visual Knowledge Base for Object Understanding
von: Yao, Ziwei, et al.
Veröffentlicht: (2025)
von: Yao, Ziwei, et al.
Veröffentlicht: (2025)
Improving Language Understanding from Screenshots
von: Gao, Tianyu, et al.
Veröffentlicht: (2024)
von: Gao, Tianyu, et al.
Veröffentlicht: (2024)
Unlocking the conversion of Web Screenshots into HTML Code with the WebSight Dataset
von: Laurençon, Hugo, et al.
Veröffentlicht: (2024)
von: Laurençon, Hugo, et al.
Veröffentlicht: (2024)
GPT4Vis: What Can GPT-4 Do for Zero-shot Visual Recognition?
von: Wu, Wenhao, et al.
Veröffentlicht: (2023)
von: Wu, Wenhao, et al.
Veröffentlicht: (2023)
UAV-VisLoc: A Large-scale Dataset for UAV Visual Localization
von: Xu, Wenjia, et al.
Veröffentlicht: (2024)
von: Xu, Wenjia, et al.
Veröffentlicht: (2024)
VisualToolAgent (VisTA): A Reinforcement Learning Framework for Visual Tool Selection
von: Huang, Zeyi, et al.
Veröffentlicht: (2025)
von: Huang, Zeyi, et al.
Veröffentlicht: (2025)
UISearch: Graph-Based Embeddings for Multimodal Enterprise UI Screenshots Retrieval
von: Ayli, Maroun, et al.
Veröffentlicht: (2025)
von: Ayli, Maroun, et al.
Veröffentlicht: (2025)
CoVis: A Collaborative Framework for Fine-grained Graphic Visual Understanding
von: Deng, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Deng, Xiaoyu, et al.
Veröffentlicht: (2024)
VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models
von: Jiang, Lingjie, et al.
Veröffentlicht: (2025)
von: Jiang, Lingjie, et al.
Veröffentlicht: (2025)
VisGuard: Securing Visualization Dissemination through Tamper-Resistant Data Retrieval
von: Ye, Huayuan, et al.
Veröffentlicht: (2025)
von: Ye, Huayuan, et al.
Veröffentlicht: (2025)
VisR-Bench: An Empirical Study on Visual Retrieval-Augmented Generation for Multilingual Long Document Understanding
von: Chen, Jian, et al.
Veröffentlicht: (2025)
von: Chen, Jian, et al.
Veröffentlicht: (2025)
Generative Refinement Networks for Visual Synthesis
von: Han, Jian, et al.
Veröffentlicht: (2026)
von: Han, Jian, et al.
Veröffentlicht: (2026)
LiVisSfM: Accurate and Robust Structure-from-Motion with LiDAR and Visual Cues
von: Jiang, Hanqing, et al.
Veröffentlicht: (2024)
von: Jiang, Hanqing, et al.
Veröffentlicht: (2024)
VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMs
von: Törtei, Brigitta Malagurski, et al.
Veröffentlicht: (2025)
von: Törtei, Brigitta Malagurski, et al.
Veröffentlicht: (2025)
Text2Vis: A Challenging and Diverse Benchmark for Generating Multimodal Visualizations from Text
von: Rahman, Mizanur, et al.
Veröffentlicht: (2025)
von: Rahman, Mizanur, et al.
Veröffentlicht: (2025)
VisMin: Visual Minimal-Change Understanding
von: Awal, Rabiul, et al.
Veröffentlicht: (2024)
von: Awal, Rabiul, et al.
Veröffentlicht: (2024)
RobustVisRAG: Causality-Aware Vision-Based Retrieval-Augmented Generation under Visual Degradations
von: Chen, I-Hsiang, et al.
Veröffentlicht: (2026)
von: Chen, I-Hsiang, et al.
Veröffentlicht: (2026)
VisNec: Measuring and Leveraging Visual Necessity for Multimodal Instruction Tuning
von: Dong, Mingkang, et al.
Veröffentlicht: (2026)
von: Dong, Mingkang, et al.
Veröffentlicht: (2026)
UI2Code^N: UI-to-Code Generation as Interactive Visual Optimization
von: Yang, Zhen, et al.
Veröffentlicht: (2025)
von: Yang, Zhen, et al.
Veröffentlicht: (2025)
VisTa: Visual-contextual and Text-augmented Zero-shot Object-level OOD Detection
von: Zhang, Bin, et al.
Veröffentlicht: (2025)
von: Zhang, Bin, et al.
Veröffentlicht: (2025)
Focus-Scan-Refine: From Human Visual Perception to Efficient Visual Token Pruning
von: Tong, Enwei, et al.
Veröffentlicht: (2026)
von: Tong, Enwei, et al.
Veröffentlicht: (2026)
PrecisionCUA: Iterative Visual Refinement for Pixel-Precise Cursor Grounding in Code Editors
von: Mittal, Himangi, et al.
Veröffentlicht: (2026)
von: Mittal, Himangi, et al.
Veröffentlicht: (2026)
VisAnalog: A Diagnostic Suite for Visual Concept Transfer on Natural Images
von: Li, Zhaonan, et al.
Veröffentlicht: (2026)
von: Li, Zhaonan, et al.
Veröffentlicht: (2026)
DynamicVis: Dynamic Visual Perception for Efficient Remote Sensing Foundation Models
von: Chen, Keyan, et al.
Veröffentlicht: (2025)
von: Chen, Keyan, et al.
Veröffentlicht: (2025)
VisJudge-Bench: Aesthetics and Quality Assessment of Visualizations
von: Xie, Yupeng, et al.
Veröffentlicht: (2025)
von: Xie, Yupeng, et al.
Veröffentlicht: (2025)
VisRAG 2.0: Evidence-Guided Multi-Image Reasoning in Visual Retrieval-Augmented Generation
von: Sun, Yubo, et al.
Veröffentlicht: (2025)
von: Sun, Yubo, et al.
Veröffentlicht: (2025)
SimVecVis: A Dataset for Enhancing MLLMs in Visualization Understanding
von: Liu, Can, et al.
Veröffentlicht: (2025)
von: Liu, Can, et al.
Veröffentlicht: (2025)
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models
von: Reilly, Dominick, et al.
Veröffentlicht: (2025)
von: Reilly, Dominick, et al.
Veröffentlicht: (2025)
VisPCO: Visual Token Pruning Configuration Optimization via Budget-Aware Pareto-Frontier Learning for Vision-Language Models
von: Ji, Huawei, et al.
Veröffentlicht: (2026)
von: Ji, Huawei, et al.
Veröffentlicht: (2026)
VisRet: Visualization Improves Knowledge-Intensive Text-to-Image Retrieval
von: Wu, Di, et al.
Veröffentlicht: (2025)
von: Wu, Di, et al.
Veröffentlicht: (2025)
G3CN: Gaussian Topology Refinement Gated Graph Convolutional Network for Skeleton-Based Action Recognition
von: Ren, Haiqing, et al.
Veröffentlicht: (2025)
von: Ren, Haiqing, et al.
Veröffentlicht: (2025)
VisAgent: Narrative-Preserving Story Visualization Framework
von: Kim, Seungkwon, et al.
Veröffentlicht: (2025)
von: Kim, Seungkwon, et al.
Veröffentlicht: (2025)
VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information
von: Kamoi, Ryo, et al.
Veröffentlicht: (2024)
von: Kamoi, Ryo, et al.
Veröffentlicht: (2024)
From EduVisBench to EduVisAgent: A Benchmark and Multi-Agent Framework for Reasoning-Driven Pedagogical Visualization
von: Ji, Haonian, et al.
Veröffentlicht: (2025)
von: Ji, Haonian, et al.
Veröffentlicht: (2025)
Insights from Visual Cognition: Understanding Human Action Dynamics with Overall Glance and Refined Gaze Transformer
von: Xing, Bohao, et al.
Veröffentlicht: (2026)
von: Xing, Bohao, et al.
Veröffentlicht: (2026)
GPT-5 Model Corrected GPT-4V's Chart Reading Errors, Not Prompting
von: Yang, Kaichun, et al.
Veröffentlicht: (2025)
von: Yang, Kaichun, et al.
Veröffentlicht: (2025)
Semantically Guided Dynamic Visual Prototype Refinement for Compositional Zero-Shot Learning
von: Peng, Zhong, et al.
Veröffentlicht: (2025)
von: Peng, Zhong, et al.
Veröffentlicht: (2025)
Draft and Refine with Visual Experts
von: Jeong, Sungheon, et al.
Veröffentlicht: (2025)
von: Jeong, Sungheon, et al.
Veröffentlicht: (2025)
Lightweight Pixel Difference Networks for Efficient Visual Representation Learning
von: Su, Zhuo, et al.
Veröffentlicht: (2024)
von: Su, Zhuo, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Rethinking Token Pruning for Historical Screenshots in GUI Visual Agents: Semantic, Spatial, and Temporal Perspectives
von: Li, Daiqiang, et al.
Veröffentlicht: (2026) -
VisKnow: Constructing Visual Knowledge Base for Object Understanding
von: Yao, Ziwei, et al.
Veröffentlicht: (2025) -
Improving Language Understanding from Screenshots
von: Gao, Tianyu, et al.
Veröffentlicht: (2024) -
Unlocking the conversion of Web Screenshots into HTML Code with the WebSight Dataset
von: Laurençon, Hugo, et al.
Veröffentlicht: (2024) -
GPT4Vis: What Can GPT-4 Do for Zero-shot Visual Recognition?
von: Wu, Wenhao, et al.
Veröffentlicht: (2023)