VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search
Fuente:
arXiv
Saved in:
| Main Authors: | Jia, Yiming, Li, Jiachen, Yue, Xiang, Li, Bo, Nie, Ping, Zou, Kai, Chen, Wenhu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MAmmoTH2: Scaling Instructions from the Web
by: Yue, Xiang, et al.
Published: (2024)
by: Yue, Xiang, et al.
Published: (2024)
VisCoder: Fine-Tuning LLMs for Executable Python Visualization Code Generation
by: Ni, Yuansheng, et al.
Published: (2025)
by: Ni, Yuansheng, et al.
Published: (2025)
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
by: Guo, Jarvis, et al.
Published: (2024)
by: Guo, Jarvis, et al.
Published: (2024)
OphIn-500K: Curating Web-Scale Visual Instructions for Scaling Ophthalmic Multimodal Large Language Models
by: Dong, Xuanzhao, et al.
Published: (2026)
by: Dong, Xuanzhao, et al.
Published: (2026)
VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?
by: Liu, Junpeng, et al.
Published: (2024)
by: Liu, Junpeng, et al.
Published: (2024)
MM-Instruct: Generated Visual Instructions for Large Multimodal Model Alignment
by: Liu, Jihao, et al.
Published: (2024)
by: Liu, Jihao, et al.
Published: (2024)
Learning to Instruct for Visual Instruction Tuning
by: Zhou, Zhihan, et al.
Published: (2025)
by: Zhou, Zhihan, et al.
Published: (2025)
Scaling Towards the Information Boundary of Instruction Sets: The Infinity Instruct Subject Technical Report
by: Du, Li, et al.
Published: (2025)
by: Du, Li, et al.
Published: (2025)
BrowserAgent: Building Web Agents with Human-Inspired Web Browsing Actions
by: Yu, Tao, et al.
Published: (2025)
by: Yu, Tao, et al.
Published: (2025)
Instruct-Imagen: Image Generation with Multi-modal Instruction
by: Hu, Hexiang, et al.
Published: (2024)
by: Hu, Hexiang, et al.
Published: (2024)
Infinity Instruct: Scaling Instruction Selection and Synthesis to Enhance Language Models
by: Li, Jijie, et al.
Published: (2025)
by: Li, Jijie, et al.
Published: (2025)
Unleashing the Reasoning Potential of Pre-trained LLMs by Critique Fine-Tuning on One Problem
by: Wang, Yubo, et al.
Published: (2025)
by: Wang, Yubo, et al.
Published: (2025)
VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
by: Koh, Jing Yu, et al.
Published: (2024)
by: Koh, Jing Yu, et al.
Published: (2024)
Genetic Instruct: Scaling up Synthetic Generation of Coding Instructions for Large Language Models
by: Majumdar, Somshubra, et al.
Published: (2024)
by: Majumdar, Somshubra, et al.
Published: (2024)
WebExplorer: Explore and Evolve for Training Long-Horizon Web Agents
by: Liu, Junteng, et al.
Published: (2025)
by: Liu, Junteng, et al.
Published: (2025)
The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale
by: Penedo, Guilherme, et al.
Published: (2024)
by: Penedo, Guilherme, et al.
Published: (2024)
Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data
by: Gu, Shuhao, et al.
Published: (2024)
by: Gu, Shuhao, et al.
Published: (2024)
VisCoder2: Building Multi-Language Visualization Coding Agents
by: Ni, Yuansheng, et al.
Published: (2025)
by: Ni, Yuansheng, et al.
Published: (2025)
WebWeaver: Structuring Web-Scale Evidence with Dynamic Outlines for Open-Ended Deep Research
by: Li, Zijian, et al.
Published: (2025)
by: Li, Zijian, et al.
Published: (2025)
WebCiteS: Attributed Query-Focused Summarization on Chinese Web Search Results with Citations
by: Deng, Haolin, et al.
Published: (2024)
by: Deng, Haolin, et al.
Published: (2024)
SCRIBES: Web-Scale Script-Based Semi-Structured Data Extraction with Reinforcement Learning
by: Liu, Shicheng, et al.
Published: (2025)
by: Liu, Shicheng, et al.
Published: (2025)
WebSeer: Training Deeper Search Agents through Reinforcement Learning with Self-Reflection
by: He, Guanzhong, et al.
Published: (2025)
by: He, Guanzhong, et al.
Published: (2025)
DeepDiver: Adaptive Search Intensity Scaling via Open-Web Reinforcement Learning
by: Shi, Wenxuan, et al.
Published: (2025)
by: Shi, Wenxuan, et al.
Published: (2025)
Instruction-Tuning Data Synthesis from Scratch via Web Reconstruction
by: Jiang, Yuxin, et al.
Published: (2025)
by: Jiang, Yuxin, et al.
Published: (2025)
EditReward: A Human-Aligned Reward Model for Instruction-Guided Image Editing
by: Wu, Keming, et al.
Published: (2025)
by: Wu, Keming, et al.
Published: (2025)
WebUIBench: A Comprehensive Benchmark for Evaluating Multimodal Large Language Models in WebUI-to-Code
by: Lin, Zhiyu, et al.
Published: (2025)
by: Lin, Zhiyu, et al.
Published: (2025)
InstructRetro: Instruction Tuning post Retrieval-Augmented Pretraining
by: Wang, Boxin, et al.
Published: (2023)
by: Wang, Boxin, et al.
Published: (2023)
Critique Fine-Tuning: Learning to Critique is More Effective than Learning to Imitate
by: Wang, Yubo, et al.
Published: (2025)
by: Wang, Yubo, et al.
Published: (2025)
SearchInstruct: Enhancing Domain Adaptation via Retrieval-Based Instruction Dataset Creation
by: Barati, Iman, et al.
Published: (2025)
by: Barati, Iman, et al.
Published: (2025)
Avenir-Web: Human-Experience-Imitating Multimodal Web Agents with Mixture of Grounding Experts
by: Li, Aiden Yiliu, et al.
Published: (2026)
by: Li, Aiden Yiliu, et al.
Published: (2026)
MM-WebAgent: A Hierarchical Multimodal Web Agent for Webpage Generation
by: Li, Yan, et al.
Published: (2026)
by: Li, Yan, et al.
Published: (2026)
WEPO: Web Element Preference Optimization for LLM-based Web Navigation
by: Liu, Jiarun, et al.
Published: (2024)
by: Liu, Jiarun, et al.
Published: (2024)
Infinite-Instruct: Synthesizing Scaling Code instruction Data with Bidirectional Synthesis and Static Verification
by: Xing, Wenjing, et al.
Published: (2025)
by: Xing, Wenjing, et al.
Published: (2025)
DoG-Instruct: Towards Premium Instruction-Tuning Data via Text-Grounded Instruction Wrapping
by: Chen, Yongrui, et al.
Published: (2023)
by: Chen, Yongrui, et al.
Published: (2023)
Dual-View Visual Contextualization for Web Navigation
by: Kil, Jihyung, et al.
Published: (2024)
by: Kil, Jihyung, et al.
Published: (2024)
SearchAttack: Red-Teaming LLMs against Knowledge-to-Action Threats under Online Web Search
by: Yan, Yu, et al.
Published: (2026)
by: Yan, Yu, et al.
Published: (2026)
InstructDoc: A Dataset for Zero-Shot Generalization of Visual Document Understanding with Instructions
by: Tanaka, Ryota, et al.
Published: (2024)
by: Tanaka, Ryota, et al.
Published: (2024)
ExpSeek: Self-Triggered Experience Seeking for Web Agents
by: Zhang, Wenyuan, et al.
Published: (2026)
by: Zhang, Wenyuan, et al.
Published: (2026)
FineWeb-zhtw: Scalable Curation of Traditional Chinese Text Data from the Web
by: Lin, Cheng-Wei, et al.
Published: (2024)
by: Lin, Cheng-Wei, et al.
Published: (2024)
EvolveCoder: Evolving Test Cases via Adversarial Verification for Code Reinforcement Learning
by: Ruan, Chi, et al.
Published: (2026)
by: Ruan, Chi, et al.
Published: (2026)
Similar Items
-
MAmmoTH2: Scaling Instructions from the Web
by: Yue, Xiang, et al.
Published: (2024) -
VisCoder: Fine-Tuning LLMs for Executable Python Visualization Code Generation
by: Ni, Yuansheng, et al.
Published: (2025) -
MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
by: Guo, Jarvis, et al.
Published: (2024) -
OphIn-500K: Curating Web-Scale Visual Instructions for Scaling Ophthalmic Multimodal Large Language Models
by: Dong, Xuanzhao, et al.
Published: (2026) -
VisualWebBench: How Far Have Multimodal LLMs Evolved in Web Page Understanding and Grounding?
by: Liu, Junpeng, et al.
Published: (2024)