Visual Program Distillation: Distilling Tools and Programmatic Reasoning into Vision-Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Hu, Yushi, Stretcu, Otilia, Lu, Chun-Ta, Viswanathan, Krishnamurthy, Hata, Kenji, Luo, Enming, Krishna, Ranjay, Fuxman, Ariel |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Agile Deliberation: Concept Deliberation for Subjective Visual Classification
di: Wang, Leijie, et al.
Pubblicazione: (2025)
di: Wang, Leijie, et al.
Pubblicazione: (2025)
Modeling Collaborator: Enabling Subjective Vision Classification With Minimal Human Effort via LLM Tool-Use
di: Toubal, Imad Eddine, et al.
Pubblicazione: (2024)
di: Toubal, Imad Eddine, et al.
Pubblicazione: (2024)
Scaling Up LLM Reviews for Google Ads Content Moderation
di: Qiao, Wei, et al.
Pubblicazione: (2024)
di: Qiao, Wei, et al.
Pubblicazione: (2024)
Why Fine-grained Labels in Pretraining Benefit Generalization?
di: Hong, Guan Zhe, et al.
Pubblicazione: (2024)
di: Hong, Guan Zhe, et al.
Pubblicazione: (2024)
Zero-Shot Image Moderation in Google Ads with LLM-Assisted Textual Descriptions and Cross-modal Co-embeddings
di: Luo, Enming, et al.
Pubblicazione: (2024)
di: Luo, Enming, et al.
Pubblicazione: (2024)
AdaReasoner: Dynamic Tool Orchestration for Iterative Visual Reasoning
di: Song, Mingyang, et al.
Pubblicazione: (2026)
di: Song, Mingyang, et al.
Pubblicazione: (2026)
Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding
di: Yoon, Hee Suk, et al.
Pubblicazione: (2026)
di: Yoon, Hee Suk, et al.
Pubblicazione: (2026)
Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models
di: Hu, Yushi, et al.
Pubblicazione: (2024)
di: Hu, Yushi, et al.
Pubblicazione: (2024)
Reinforced Visual Perception with Tools
di: Zhou, Zetong, et al.
Pubblicazione: (2025)
di: Zhou, Zetong, et al.
Pubblicazione: (2025)
Visual Program Distillation with Template-Based Augmentation
di: Shlapentokh-Rothman, Michal, et al.
Pubblicazione: (2024)
di: Shlapentokh-Rothman, Michal, et al.
Pubblicazione: (2024)
Training Task Experts through Retrieval Based Distillation
di: Ge, Jiaxin, et al.
Pubblicazione: (2024)
di: Ge, Jiaxin, et al.
Pubblicazione: (2024)
ProVision: Programmatically Scaling Vision-centric Instruction Data for Multimodal Language Models
di: Zhang, Jieyu, et al.
Pubblicazione: (2024)
di: Zhang, Jieyu, et al.
Pubblicazione: (2024)
Structure From Tracking: Distilling Structure-Preserving Motion for Video Generation
di: Fei, Yang, et al.
Pubblicazione: (2025)
di: Fei, Yang, et al.
Pubblicazione: (2025)
SKDF: A Simple Knowledge Distillation Framework for Distilling Open-Vocabulary Knowledge to Open-world Object Detector
di: Ma, Shuailei, et al.
Pubblicazione: (2023)
di: Ma, Shuailei, et al.
Pubblicazione: (2023)
TableMind++: An Uncertainty-Aware Programmatic Agent for Tool-Augmented Table Reasoning
di: Cheng, Mingyue, et al.
Pubblicazione: (2026)
di: Cheng, Mingyue, et al.
Pubblicazione: (2026)
Ablate-to-Validate: Are Vision-Language Models Really Using Continuous Thought Tokens?
di: Zhang, Tianyi, et al.
Pubblicazione: (2026)
di: Zhang, Tianyi, et al.
Pubblicazione: (2026)
Scale Can't Overcome Pragmatics: The Impact of Reporting Bias on Vision-Language Reasoning
di: Kamath, Amita, et al.
Pubblicazione: (2026)
di: Kamath, Amita, et al.
Pubblicazione: (2026)
Skill-Conditioned Gated Self-Distillation for LLM Reasoning
di: Huang, Jiazhen, et al.
Pubblicazione: (2026)
di: Huang, Jiazhen, et al.
Pubblicazione: (2026)
Visual-Advantage On-Policy Distillation for Vision-Language Models
di: Liu, Ruiqi, et al.
Pubblicazione: (2026)
di: Liu, Ruiqi, et al.
Pubblicazione: (2026)
Distill Visual Chart Reasoning Ability from LLMs to MLLMs
di: He, Wei, et al.
Pubblicazione: (2024)
di: He, Wei, et al.
Pubblicazione: (2024)
A Stepwise Distillation Learning Strategy for Non-differentiable Visual Programming Frameworks on Visual Reasoning Tasks
di: Wan, Wentao, et al.
Pubblicazione: (2023)
di: Wan, Wentao, et al.
Pubblicazione: (2023)
Vision Transformers with Self-Distilled Registers
di: Chen, Yinjie, et al.
Pubblicazione: (2025)
di: Chen, Yinjie, et al.
Pubblicazione: (2025)
CodeTool: Enhancing Programmatic Tool Invocation of LLMs via Process Supervision
di: Lu, Yifei, et al.
Pubblicazione: (2025)
di: Lu, Yifei, et al.
Pubblicazione: (2025)
Skill-Aware Data Selection and Fine-Tuning for Data-Efficient Reasoning Distillation
di: Zhang, Lechen, et al.
Pubblicazione: (2026)
di: Zhang, Lechen, et al.
Pubblicazione: (2026)
m&m's: A Benchmark to Evaluate Tool-Use for multi-step multi-modal Tasks
di: Ma, Zixian, et al.
Pubblicazione: (2024)
di: Ma, Zixian, et al.
Pubblicazione: (2024)
Distilling Algorithmic Reasoning from LLMs via Explaining Solution Programs
di: Li, Jierui, et al.
Pubblicazione: (2024)
di: Li, Jierui, et al.
Pubblicazione: (2024)
GenEval 2: Addressing Benchmark Drift in Text-to-Image Evaluation
di: Kamath, Amita, et al.
Pubblicazione: (2025)
di: Kamath, Amita, et al.
Pubblicazione: (2025)
The Hard Positive Truth about Vision-Language Compositionality
di: Kamath, Amita, et al.
Pubblicazione: (2024)
di: Kamath, Amita, et al.
Pubblicazione: (2024)
Distilling Feedback into Memory-as-a-Tool
di: Gallego, Víctor
Pubblicazione: (2026)
di: Gallego, Víctor
Pubblicazione: (2026)
Counting Circuits: Mechanistic Interpretability of Visual Reasoning in Large Vision-Language Models
di: Che, Liwei, et al.
Pubblicazione: (2026)
di: Che, Liwei, et al.
Pubblicazione: (2026)
SGMD: Score Gradient Matching Distillation for Few-Step Video Diffusion Distillation
di: Wu, Zhuguanyu, et al.
Pubblicazione: (2026)
di: Wu, Zhuguanyu, et al.
Pubblicazione: (2026)
Iterated Learning Improves Compositionality in Large Vision-Language Models
di: Zheng, Chenhao, et al.
Pubblicazione: (2024)
di: Zheng, Chenhao, et al.
Pubblicazione: (2024)
MM-CondChain: A Programmatically Verified Benchmark for Visually Grounded Deep Compositional Reasoning
di: Shen, Haozhan, et al.
Pubblicazione: (2026)
di: Shen, Haozhan, et al.
Pubblicazione: (2026)
Perception Tokens Enhance Visual Reasoning in Multimodal Language Models
di: Bigverdi, Mahtab, et al.
Pubblicazione: (2024)
di: Bigverdi, Mahtab, et al.
Pubblicazione: (2024)
Hide to See: Reasoning-prefix Masking for Visual-anchored Thinking in VLM Distillation
di: Yu, Seonghoon, et al.
Pubblicazione: (2026)
di: Yu, Seonghoon, et al.
Pubblicazione: (2026)
Distilling the Essence: Efficient Reasoning Distillation via Sequence Truncation
di: Chen, Wei-Rui, et al.
Pubblicazione: (2025)
di: Chen, Wei-Rui, et al.
Pubblicazione: (2025)
Training with Harnesses: On-Policy Harness Self-Distillation for Complex Reasoning
di: Zhao, Zhengyang, et al.
Pubblicazione: (2026)
di: Zhao, Zhengyang, et al.
Pubblicazione: (2026)
Native Parallel Reasoner: Reasoning in Parallelism via Self-Distilled Reinforcement Learning
di: Wu, Tong, et al.
Pubblicazione: (2025)
di: Wu, Tong, et al.
Pubblicazione: (2025)
The Cartesian Shortcut: Re-evaluate Vision Reasoning in Polar Coordinate Space
di: Hu, Xia, et al.
Pubblicazione: (2026)
di: Hu, Xia, et al.
Pubblicazione: (2026)
Distilling Vision-Language Models on Millions of Videos
di: Zhao, Yue, et al.
Pubblicazione: (2024)
di: Zhao, Yue, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Agile Deliberation: Concept Deliberation for Subjective Visual Classification
di: Wang, Leijie, et al.
Pubblicazione: (2025) -
Modeling Collaborator: Enabling Subjective Vision Classification With Minimal Human Effort via LLM Tool-Use
di: Toubal, Imad Eddine, et al.
Pubblicazione: (2024) -
Scaling Up LLM Reviews for Google Ads Content Moderation
di: Qiao, Wei, et al.
Pubblicazione: (2024) -
Why Fine-grained Labels in Pretraining Benefit Generalization?
di: Hong, Guan Zhe, et al.
Pubblicazione: (2024) -
Zero-Shot Image Moderation in Google Ads with LLM-Assisted Textual Descriptions and Cross-modal Co-embeddings
di: Luo, Enming, et al.
Pubblicazione: (2024)