Visual Program Distillation with Template-Based Augmentation
Fuente:
arXiv
Saved in:
| Main Authors: | Shlapentokh-Rothman, Michal, Wang, Yu-Xiong, Hoiem, Derek |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Decomposing Queries into Tool Calls for Long-Video Keyframe Retrieval
by: Shlapentokh-Rothman, Michal, et al.
Published: (2026)
by: Shlapentokh-Rothman, Michal, et al.
Published: (2026)
Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models
by: Zhou, Andy, et al.
Published: (2023)
by: Zhou, Andy, et al.
Published: (2023)
Region-Based Representations Revisited
by: Shlapentokh-Rothman, Michal, et al.
Published: (2024)
by: Shlapentokh-Rothman, Michal, et al.
Published: (2024)
RELOCATE: A Simple Training-Free Baseline for Visual Query Localization Using Region-Based Representations
by: Khosla, Savya, et al.
Published: (2024)
by: Khosla, Savya, et al.
Published: (2024)
Visual Program Distillation: Distilling Tools and Programmatic Reasoning into Vision-Language Models
by: Hu, Yushi, et al.
Published: (2023)
by: Hu, Yushi, et al.
Published: (2023)
Beyond Meme Templates: Limitations of Visual Similarity Measures in Meme Matching
by: Hazman, Muzhaffar, et al.
Published: (2025)
by: Hazman, Muzhaffar, et al.
Published: (2025)
Overcoming Language Priors for Visual Question Answering Based on Knowledge Distillation
by: Peng, Daowan, et al.
Published: (2025)
by: Peng, Daowan, et al.
Published: (2025)
Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries
by: Wu, Yin, et al.
Published: (2025)
by: Wu, Yin, et al.
Published: (2025)
Efficient End-to-End Visual Document Understanding with Rationale Distillation
by: Zhu, Wang, et al.
Published: (2023)
by: Zhu, Wang, et al.
Published: (2023)
Recursive Visual Programming
by: Ge, Jiaxin, et al.
Published: (2023)
by: Ge, Jiaxin, et al.
Published: (2023)
REN: Fast and Efficient Region Encodings from Patch-Based Image Encoders
by: Khosla, Savya, et al.
Published: (2025)
by: Khosla, Savya, et al.
Published: (2025)
Continual Learning in Open-vocabulary Classification with Complementary Memory Systems
by: Zhu, Zhen, et al.
Published: (2023)
by: Zhu, Zhen, et al.
Published: (2023)
Unveil: Unified Visual-Textual Integration and Distillation for Multi-modal Document Retrieval
by: Sun, Hao, et al.
Published: (2026)
by: Sun, Hao, et al.
Published: (2026)
Anytime Continual Learning for Open Vocabulary Classification
by: Zhu, Zhen, et al.
Published: (2024)
by: Zhu, Zhen, et al.
Published: (2024)
Retrieval-Augmented Perception: High-Resolution Image Perception Meets Visual RAG
by: Wang, Wenbin, et al.
Published: (2025)
by: Wang, Wenbin, et al.
Published: (2025)
Synthesize Step-by-Step: Tools, Templates and LLMs as Data Generators for Reasoning-Based Chart VQA
by: Li, Zhuowan, et al.
Published: (2024)
by: Li, Zhuowan, et al.
Published: (2024)
VimRAG: Navigating Massive Visual Context in Retrieval-Augmented Generation via Multimodal Memory Graph
by: Wang, Qiuchen, et al.
Published: (2026)
by: Wang, Qiuchen, et al.
Published: (2026)
Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding
by: Yoon, Hee Suk, et al.
Published: (2026)
by: Yoon, Hee Suk, et al.
Published: (2026)
VDebugger: Harnessing Execution Feedback for Debugging Visual Programs
by: Wu, Xueqing, et al.
Published: (2024)
by: Wu, Xueqing, et al.
Published: (2024)
VisRAG 2.0: Evidence-Guided Multi-Image Reasoning in Visual Retrieval-Augmented Generation
by: Sun, Yubo, et al.
Published: (2025)
by: Sun, Yubo, et al.
Published: (2025)
RAVENEA: A Benchmark for Multimodal Retrieval-Augmented Visual Culture Understanding
by: Li, Jiaang, et al.
Published: (2025)
by: Li, Jiaang, et al.
Published: (2025)
MonoPatchNeRF: Improving Neural Radiance Fields with Patch-based Monocular Guidance
by: Wu, Yuqun, et al.
Published: (2024)
by: Wu, Yuqun, et al.
Published: (2024)
Plenoptic PNG: Real-Time Neural Radiance Fields in 150 KB
by: Lee, Jae Yong, et al.
Published: (2024)
by: Lee, Jae Yong, et al.
Published: (2024)
T-REN: Learning Text-Aligned Region Tokens Improves Dense Vision-Language Alignment and Scalability
by: Khosla, Savya, et al.
Published: (2026)
by: Khosla, Savya, et al.
Published: (2026)
Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation
by: Yu, Seonghoon, et al.
Published: (2026)
by: Yu, Seonghoon, et al.
Published: (2026)
DocLens : A Tool-Augmented Multi-Agent Framework for Long Visual Document Understanding
by: Zhu, Dawei, et al.
Published: (2025)
by: Zhu, Dawei, et al.
Published: (2025)
DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies
by: Song, Wei, et al.
Published: (2025)
by: Song, Wei, et al.
Published: (2025)
Growing a Multi-head Twig via Distillation and Reinforcement Learning to Accelerate Large Vision-Language Models
by: Shao, Zhenwei, et al.
Published: (2025)
by: Shao, Zhenwei, et al.
Published: (2025)
Generating Reports or Repeating Templates? Measuring and Mitigating Template Collapse in 3D CT Report Generation
by: Maye-Lasserre, Tom, et al.
Published: (2026)
by: Maye-Lasserre, Tom, et al.
Published: (2026)
$A^2R^2$: Advancing Img2LaTeX Conversion via Visual Reasoning with Attention-Guided Refinement
by: Li, Zhecheng, et al.
Published: (2025)
by: Li, Zhecheng, et al.
Published: (2025)
OphIn-500K: Curating Web-Scale Visual Instructions for Scaling Ophthalmic Multimodal Large Language Models
by: Dong, Xuanzhao, et al.
Published: (2026)
by: Dong, Xuanzhao, et al.
Published: (2026)
Distill CLIP (DCLIP): Enhancing Image-Text Retrieval via Cross-Modal Transformer Distillation
by: Csizmadia, Daniel, et al.
Published: (2025)
by: Csizmadia, Daniel, et al.
Published: (2025)
Latent Visual Reasoning
by: Li, Bangzheng, et al.
Published: (2025)
by: Li, Bangzheng, et al.
Published: (2025)
The Devil is in the Prompts: Retrieval-Augmented Prompt Optimization for Text-to-Video Generation
by: Gao, Bingjie, et al.
Published: (2025)
by: Gao, Bingjie, et al.
Published: (2025)
LinguDistill: Recovering Linguistic Ability in Vision-Language Models via Selective Cross-Modal Distillation
by: Irawan, Patrick Amadeus, et al.
Published: (2026)
by: Irawan, Patrick Amadeus, et al.
Published: (2026)
Harnessing Webpage UIs for Text-Rich Visual Understanding
by: Liu, Junpeng, et al.
Published: (2024)
by: Liu, Junpeng, et al.
Published: (2024)
TextRegion: Text-Aligned Region Tokens from Frozen Image-Text Models
by: Xiao, Yao, et al.
Published: (2025)
by: Xiao, Yao, et al.
Published: (2025)
What Makes for Good Visual Instructions? Synthesizing Complex Visual Reasoning Instructions for Visual Instruction Tuning
by: Du, Yifan, et al.
Published: (2023)
by: Du, Yifan, et al.
Published: (2023)
Subject or Style: Adaptive and Training-Free Mixture of LoRAs
by: Zhang, Jia-Chen, et al.
Published: (2025)
by: Zhang, Jia-Chen, et al.
Published: (2025)
MLLMs-Augmented Visual-Language Representation Learning
by: Liu, Yanqing, et al.
Published: (2023)
by: Liu, Yanqing, et al.
Published: (2023)
Similar Items
-
Decomposing Queries into Tool Calls for Long-Video Keyframe Retrieval
by: Shlapentokh-Rothman, Michal, et al.
Published: (2026) -
Language Agent Tree Search Unifies Reasoning Acting and Planning in Language Models
by: Zhou, Andy, et al.
Published: (2023) -
Region-Based Representations Revisited
by: Shlapentokh-Rothman, Michal, et al.
Published: (2024) -
RELOCATE: A Simple Training-Free Baseline for Visual Query Localization Using Region-Based Representations
by: Khosla, Savya, et al.
Published: (2024) -
Visual Program Distillation: Distilling Tools and Programmatic Reasoning into Vision-Language Models
by: Hu, Yushi, et al.
Published: (2023)