Unlocking the conversion of Web Screenshots into HTML Code with the WebSight Dataset
Fuente:
arXiv
Saved in:
| Main Authors: | Laurençon, Hugo, Tronchon, Léo, Sanh, Victor |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
What matters when building vision-language models?
by: Laurençon, Hugo, et al.
Published: (2024)
by: Laurençon, Hugo, et al.
Published: (2024)
Building and better understanding vision-language models: insights and future directions
by: Laurençon, Hugo, et al.
Published: (2024)
by: Laurençon, Hugo, et al.
Published: (2024)
WebSight: A Vision-First Architecture for Robust Web Agents
by: Bhathal, Tanvir, et al.
Published: (2025)
by: Bhathal, Tanvir, et al.
Published: (2025)
PixelWeb: The First Web GUI Dataset with Pixel-Wise Labels
by: Yang, Qi, et al.
Published: (2025)
by: Yang, Qi, et al.
Published: (2025)
WebAccessVL: Violation-Aware VLM for Web Accessibility
by: Zheng, Amber Yijia, et al.
Published: (2025)
by: Zheng, Amber Yijia, et al.
Published: (2025)
Sightation Counts: Leveraging Sighted User Feedback in Building a BLV-aligned Dataset of Diagram Descriptions
by: Kang, Wan Ju, et al.
Published: (2025)
by: Kang, Wan Ju, et al.
Published: (2025)
OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web
by: Kapoor, Raghav, et al.
Published: (2024)
by: Kapoor, Raghav, et al.
Published: (2024)
Tur[k]ingBench: A Challenge Benchmark for Web Agents
by: Xu, Kevin, et al.
Published: (2024)
by: Xu, Kevin, et al.
Published: (2024)
LIVS: A Pluralistic Alignment Dataset for Inclusive Public Spaces
by: Mushkani, Rashid, et al.
Published: (2025)
by: Mushkani, Rashid, et al.
Published: (2025)
Egocentric Co-Pilot: Web-Native Smart-Glasses Agents for Assistive Egocentric AI
by: Yang, Sicheng, et al.
Published: (2026)
by: Yang, Sicheng, et al.
Published: (2026)
Aria Everyday Activities Dataset
by: Lv, Zhaoyang, et al.
Published: (2024)
by: Lv, Zhaoyang, et al.
Published: (2024)
Real-Time Feedback and Benchmark Dataset for Isometric Pose Evaluation
by: Jaiswal, Abhishek, et al.
Published: (2025)
by: Jaiswal, Abhishek, et al.
Published: (2025)
SigmaCollab: An Application-Driven Dataset for Physically Situated Collaboration
by: Bohus, Dan, et al.
Published: (2025)
by: Bohus, Dan, et al.
Published: (2025)
Few-Shot VLM-Based G-Code and HMI Verification in CNC Machining
by: Pour, Yasaman Hashem, et al.
Published: (2025)
by: Pour, Yasaman Hashem, et al.
Published: (2025)
ChartGen: Scaling Chart Understanding Via Code-Guided Synthetic Chart Generation
by: Kondic, Jovana, et al.
Published: (2025)
by: Kondic, Jovana, et al.
Published: (2025)
ScreenQA: Large-Scale Question-Answer Pairs over Mobile App Screenshots
by: Hsiao, Yu-Chung, et al.
Published: (2022)
by: Hsiao, Yu-Chung, et al.
Published: (2022)
Can Vision-Language Models Understand and Interpret Dynamic Gestures from Pedestrians? Pilot Datasets and Exploration Towards Instructive Nonverbal Commands for Cooperative Autonomous Vehicles
by: Bossen, Tonko E. W., et al.
Published: (2025)
by: Bossen, Tonko E. W., et al.
Published: (2025)
How Good (Or Bad) Are LLMs at Detecting Misleading Visualizations?
by: Lo, Leo Yu-Ho, et al.
Published: (2024)
by: Lo, Leo Yu-Ho, et al.
Published: (2024)
Code2World: A GUI World Model via Renderable Code Generation
by: Zheng, Yuhao, et al.
Published: (2026)
by: Zheng, Yuhao, et al.
Published: (2026)
IndEgo: A Dataset of Industrial Scenarios and Collaborative Work for Egocentric Assistants
by: Chavan, Vivek, et al.
Published: (2025)
by: Chavan, Vivek, et al.
Published: (2025)
Handwritten Code Recognition for Pen-and-Paper CS Education
by: Islam, Md Sazzad, et al.
Published: (2024)
by: Islam, Md Sazzad, et al.
Published: (2024)
ASL STEM Wiki: Dataset and Benchmark for Interpreting STEM Articles
by: Yin, Kayo, et al.
Published: (2024)
by: Yin, Kayo, et al.
Published: (2024)
Code2Video: A Code-centric Paradigm for Educational Video Generation
by: Chen, Yanzhe, et al.
Published: (2025)
by: Chen, Yanzhe, et al.
Published: (2025)
OpenDriver: An Open-Road Driver State Detection Dataset
by: Liu, Delong, et al.
Published: (2023)
by: Liu, Delong, et al.
Published: (2023)
HandS3C: 3D Hand Mesh Reconstruction with State Space Spatial Channel Attention from RGB images
by: Jiao, Zixun, et al.
Published: (2024)
by: Jiao, Zixun, et al.
Published: (2024)
Benchmarking XAI Explanations with Human-Aligned Evaluations
by: Kazmierczak, Rémi, et al.
Published: (2024)
by: Kazmierczak, Rémi, et al.
Published: (2024)
Large Language Models estimate fine-grained human color-concept associations
by: Mukherjee, Kushin, et al.
Published: (2024)
by: Mukherjee, Kushin, et al.
Published: (2024)
It's a Feature, Not a Bug: Measuring Creative Fluidity in Image Generators
by: Ramaswamy, Aditi, et al.
Published: (2024)
by: Ramaswamy, Aditi, et al.
Published: (2024)
Posture-Informed Muscular Force Learning for Robust Hand Pressure Estimation
by: Seo, Kyungjin, et al.
Published: (2024)
by: Seo, Kyungjin, et al.
Published: (2024)
AV-EmoDialog: Chat with Audio-Visual Users Leveraging Emotional Cues
by: Park, Se Jin, et al.
Published: (2024)
by: Park, Se Jin, et al.
Published: (2024)
ASAP: Interpretable Analysis and Summarization of AI-generated Image Patterns at Scale
by: Huang, Jinbin, et al.
Published: (2024)
by: Huang, Jinbin, et al.
Published: (2024)
Learning User Embeddings from Human Gaze for Personalised Saliency Prediction
by: Strohm, Florian, et al.
Published: (2024)
by: Strohm, Florian, et al.
Published: (2024)
Dermatologist-like explainable AI enhances melanoma diagnosis accuracy: eye-tracking study
by: Chanda, Tirtha, et al.
Published: (2024)
by: Chanda, Tirtha, et al.
Published: (2024)
SelfReDepth: Self-Supervised Real-Time Depth Restoration for Consumer-Grade Sensors
by: Duarte, Alexandre, et al.
Published: (2024)
by: Duarte, Alexandre, et al.
Published: (2024)
RITA: A Real-time Interactive Talking Avatars Framework
by: Cheng, Wuxinlin, et al.
Published: (2024)
by: Cheng, Wuxinlin, et al.
Published: (2024)
Let's Go Real Talk: Spoken Dialogue Model for Face-to-Face Conversation
by: Park, Se Jin, et al.
Published: (2024)
by: Park, Se Jin, et al.
Published: (2024)
VerSe: Integrating Multiple Queries as Prompts for Versatile Cardiac MRI Segmentation
by: Guo, Bangwei, et al.
Published: (2024)
by: Guo, Bangwei, et al.
Published: (2024)
Sketch2Prototype: Rapid Conceptual Design Exploration and Prototyping with Generative AI
by: Edwards, Kristen M., et al.
Published: (2024)
by: Edwards, Kristen M., et al.
Published: (2024)
Classification Metrics for Image Explanations: Towards Building Reliable XAI-Evaluations
by: Fresz, Benjamin, et al.
Published: (2024)
by: Fresz, Benjamin, et al.
Published: (2024)
Embracing Diversity: Interpretable Zero-shot classification beyond one vector per class
by: Moayeri, Mazda, et al.
Published: (2024)
by: Moayeri, Mazda, et al.
Published: (2024)
Similar Items
-
What matters when building vision-language models?
by: Laurençon, Hugo, et al.
Published: (2024) -
Building and better understanding vision-language models: insights and future directions
by: Laurençon, Hugo, et al.
Published: (2024) -
WebSight: A Vision-First Architecture for Robust Web Agents
by: Bhathal, Tanvir, et al.
Published: (2025) -
PixelWeb: The First Web GUI Dataset with Pixel-Wise Labels
by: Yang, Qi, et al.
Published: (2025) -
WebAccessVL: Violation-Aware VLM for Web Accessibility
by: Zheng, Amber Yijia, et al.
Published: (2025)