Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data Generation
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Yang, Yue, Patel, Ajay, Deitke, Matt, Gupta, Tanmay, Weihs, Luca, Head, Andrew, Yatskar, Mark, Callison-Burch, Chris, Krishna, Ranjay, Kembhavi, Aniruddha, Clark, Christopher |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
CodeNav: Beyond tool-use to using real-world codebases with LLM agents
par: Gupta, Tanmay, et autres
Publié: (2024)
par: Gupta, Tanmay, et autres
Publié: (2024)
Holodeck: Language Guided Generation of 3D Embodied AI Environments
par: Yang, Yue, et autres
Publié: (2023)
par: Yang, Yue, et autres
Publié: (2023)
FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale
par: Patel, Ajay, et autres
Publié: (2026)
par: Patel, Ajay, et autres
Publié: (2026)
DataDreamer: A Tool for Synthetic Data Generation and Reproducible LLM Workflows
par: Patel, Ajay, et autres
Publié: (2024)
par: Patel, Ajay, et autres
Publié: (2024)
MutaGReP: Execution-Free Repository-Grounded Plan Search for Code-Use
par: Khan, Zaid, et autres
Publié: (2025)
par: Khan, Zaid, et autres
Publié: (2025)
Interpretable-by-Design Text Understanding with Iteratively Generated Concept Bottleneck
par: Ludan, Josh Magnus, et autres
Publié: (2023)
par: Ludan, Josh Magnus, et autres
Publié: (2023)
Low-Resource Authorship Style Transfer: Can Non-Famous Authors Be Imitated?
par: Patel, Ajay, et autres
Publié: (2022)
par: Patel, Ajay, et autres
Publié: (2022)
Iterated Learning Improves Compositionality in Large Vision-Language Models
par: Zheng, Chenhao, et autres
Publié: (2024)
par: Zheng, Chenhao, et autres
Publié: (2024)
CoMo: Controllable Motion Generation through Language Guided Pose Code Editing
par: Huang, Yiming, et autres
Publié: (2024)
par: Huang, Yiming, et autres
Publié: (2024)
Generate Any Scene: Scene Graph Driven Data Synthesis for Visual Generation Training
par: Gao, Ziqi, et autres
Publié: (2024)
par: Gao, Ziqi, et autres
Publié: (2024)
TinyStyler: Efficient Few-Shot Text Style Transfer with Authorship Embeddings
par: Horvitz, Zachary, et autres
Publié: (2024)
par: Horvitz, Zachary, et autres
Publié: (2024)
mStyleDistance: Multilingual Style Embeddings and their Evaluation
par: Qiu, Justin, et autres
Publié: (2025)
par: Qiu, Justin, et autres
Publié: (2025)
CALYPSO: LLMs as Dungeon Masters' Assistants
par: Zhu, Andrew, et autres
Publié: (2023)
par: Zhu, Andrew, et autres
Publié: (2023)
Overhearing LLM Agents: A Survey, Taxonomy, and Roadmap
par: Zhu, Andrew, et autres
Publié: (2025)
par: Zhu, Andrew, et autres
Publié: (2025)
StyleDistance: Stronger Content-Independent Style Embeddings with Synthetic Parallel Examples
par: Patel, Ajay, et autres
Publié: (2024)
par: Patel, Ajay, et autres
Publié: (2024)
Agreement Metrics for LLM-as-Judge Evaluation: What to Report and Why
par: Rao, Delip, et autres
Publié: (2026)
par: Rao, Delip, et autres
Publié: (2026)
Seeing the Unseen: Visual Common Sense for Semantic Placement
par: Ramrakhya, Ram, et autres
Publié: (2024)
par: Ramrakhya, Ram, et autres
Publié: (2024)
What Do Claim Verification Datasets Actually Test? A Reasoning Trace Analysis
par: Rao, Delip, et autres
Publié: (2026)
par: Rao, Delip, et autres
Publié: (2026)
ParaGuide: Guided Diffusion Paraphrasers for Plug-and-Play Textual Style Transfer
par: Horvitz, Zachary, et autres
Publié: (2023)
par: Horvitz, Zachary, et autres
Publié: (2023)
MiRAGeNews: Multimodal Realistic AI-Generated News Detection
par: Huang, Runsheng, et autres
Publié: (2024)
par: Huang, Runsheng, et autres
Publié: (2024)
Autorubric: Unifying Rubric-based LLM Evaluation
par: Rao, Delip, et autres
Publié: (2026)
par: Rao, Delip, et autres
Publié: (2026)
BibTeX Citation Hallucinations in Scientific Publishing Agents: Evaluation and Mitigation
par: Rao, Delip, et autres
Publié: (2026)
par: Rao, Delip, et autres
Publié: (2026)
Task Me Anything
par: Zhang, Jieyu, et autres
Publié: (2024)
par: Zhang, Jieyu, et autres
Publié: (2024)
SPOC: Imitating Shortest Paths in Simulation Enables Effective Navigation and Manipulation in the Real World
par: Ehsani, Kiana, et autres
Publié: (2023)
par: Ehsani, Kiana, et autres
Publié: (2023)
One Diffusion to Generate Them All
par: Le, Duong H., et autres
Publié: (2024)
par: Le, Duong H., et autres
Publié: (2024)
First Steps Towards Overhearing LLM Agents: A Case Study With Dungeons & Dragons Gameplay
par: Zhu, Andrew, et autres
Publié: (2025)
par: Zhu, Andrew, et autres
Publié: (2025)
Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
par: Deitke, Matt, et autres
Publié: (2024)
par: Deitke, Matt, et autres
Publié: (2024)
m&m's: A Benchmark to Evaluate Tool-Use for multi-step multi-modal Tasks
par: Ma, Zixian, et autres
Publié: (2024)
par: Ma, Zixian, et autres
Publié: (2024)
ViUniT: Visual Unit Tests for More Robust Visual Programming
par: Panagopoulou, Artemis, et autres
Publié: (2024)
par: Panagopoulou, Artemis, et autres
Publié: (2024)
Toward Beginner-Friendly LLMs for Language Learning: Controlling Difficulty in Conversation
par: Jin, Meiqing, et autres
Publié: (2025)
par: Jin, Meiqing, et autres
Publié: (2025)
A Textbook Remedy for Domain Shifts: Knowledge Priors for Medical Image Analysis
par: Yang, Yue, et autres
Publié: (2024)
par: Yang, Yue, et autres
Publié: (2024)
Machine Text Detectors are Membership Inference Attacks
par: Koike, Ryuto, et autres
Publié: (2025)
par: Koike, Ryuto, et autres
Publié: (2025)
Towards Faithful Model Explanation in NLP: A Survey
par: Lyu, Qing, et autres
Publié: (2022)
par: Lyu, Qing, et autres
Publié: (2022)
Uncovering Differences in Persuasive Language in Russian versus English Wikipedia
par: Li, Bryan, et autres
Publié: (2024)
par: Li, Bryan, et autres
Publié: (2024)
Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents
par: Rao, Delip, et autres
Publié: (2026)
par: Rao, Delip, et autres
Publié: (2026)
This Land is {Your, My} Land: Evaluating Geopolitical Biases in Language Models
par: Li, Bryan, et autres
Publié: (2023)
par: Li, Bryan, et autres
Publié: (2023)
From an Image to a Scene: Learning to Imagine the World from a Million 360 Videos
par: Wallingford, Matthew, et autres
Publié: (2024)
par: Wallingford, Matthew, et autres
Publié: (2024)
OpenPI2.0: An Improved Dataset for Entity Tracking in Texts
par: Zhang, Li, et autres
Publié: (2023)
par: Zhang, Li, et autres
Publié: (2023)
ThinknCheck: Grounded Claim Verification with Compact, Reasoning-Driven, and Interpretable Models
par: Rao, Delip, et autres
Publié: (2026)
par: Rao, Delip, et autres
Publié: (2026)
You Have Thirteen Hours in Which to Solve the Labyrinth: Enhancing AI Game Masters with Function Calling
par: Song, Jaewoo, et autres
Publié: (2024)
par: Song, Jaewoo, et autres
Publié: (2024)
Documents similaires
-
CodeNav: Beyond tool-use to using real-world codebases with LLM agents
par: Gupta, Tanmay, et autres
Publié: (2024) -
Holodeck: Language Guided Generation of 3D Embodied AI Environments
par: Yang, Yue, et autres
Publié: (2023) -
FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale
par: Patel, Ajay, et autres
Publié: (2026) -
DataDreamer: A Tool for Synthetic Data Generation and Reproducible LLM Workflows
par: Patel, Ajay, et autres
Publié: (2024) -
MutaGReP: Execution-Free Repository-Grounded Plan Search for Code-Use
par: Khan, Zaid, et autres
Publié: (2025)