Cache-of-Thought: Master-Apprentice Framework for Cost-Effective Vision Language Model Reasoning
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Mingyuan, Jiang, Jize, Zheng, Haozhen, Li, Meitang, Li, Zhaoheng, Tian, Beitong, Chen, Bo, Park, Yongjoo, Zhang, Minjia, Zhai, Chengxiang, Nahrstedt, Klara |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UOUO: Uncontextualized Uncommon Objects for Measuring Knowledge Horizons of Vision Language Models
by: Pi, Xinyu, et al.
Published: (2024)
by: Pi, Xinyu, et al.
Published: (2024)
VTool-R1: VLMs Learn to Think with Images via Reinforcement Learning on Multimodal Tool Use
by: Wu, Mingyuan, et al.
Published: (2025)
by: Wu, Mingyuan, et al.
Published: (2025)
Aha Moment Revisited: Are VLMs Truly Capable of Self Verification in Inference-time Scaling?
by: Wu, Mingyuan, et al.
Published: (2025)
by: Wu, Mingyuan, et al.
Published: (2025)
Spatio-Temporal LLM: Reasoning about Environments and Actions
by: Zheng, Haozhen, et al.
Published: (2025)
by: Zheng, Haozhen, et al.
Published: (2025)
AquaVLM: Improving Underwater Situation Awareness with Mobile Vision Language Models
by: Tian, Beitong, et al.
Published: (2025)
by: Tian, Beitong, et al.
Published: (2025)
AquaScope: Reliable Underwater Image Transmission on Mobile Devices
by: Tian, Beitong, et al.
Published: (2025)
by: Tian, Beitong, et al.
Published: (2025)
Spreadsheet-RL: Advancing Large Language Model Agents on Realistic Spreadsheet Tasks via Reinforcement Learning
by: Chi, Banghao, et al.
Published: (2026)
by: Chi, Banghao, et al.
Published: (2026)
QStore: Quantization-Aware Compressed Model Storage
by: Shah, Raunak, et al.
Published: (2025)
by: Shah, Raunak, et al.
Published: (2025)
TraceNet: Segment one thing efficiently
by: Wu, Mingyuan, et al.
Published: (2024)
by: Wu, Mingyuan, et al.
Published: (2024)
SIEVE: Effective Filtered Vector Search with Collection of Indexes
by: Li, Zhaoheng, et al.
Published: (2025)
by: Li, Zhaoheng, et al.
Published: (2025)
MojoFrame: Dataframe Library in Mojo Language
by: Huang, Shengya, et al.
Published: (2025)
by: Huang, Shengya, et al.
Published: (2025)
Circuit Tracing in Vision-Language Models: Understanding the Internal Mechanisms of Multimodal Thinking
by: Yang, Jingcheng, et al.
Published: (2026)
by: Yang, Jingcheng, et al.
Published: (2026)
ACT360: An Efficient 360-Degree Action Detection and Summarization Framework for Mission-Critical Training and Debriefing
by: Tiwari, Aditi, et al.
Published: (2025)
by: Tiwari, Aditi, et al.
Published: (2025)
FDM-Bench: A Comprehensive Benchmark for Evaluating Large Language Models in Additive Manufacturing Tasks
by: Eslaminia, Ahmadreza, et al.
Published: (2024)
by: Eslaminia, Ahmadreza, et al.
Published: (2024)
Chipmink: Efficient Delta Identification for Massive Object Graph
by: Chockchowwat, Supawit, et al.
Published: (2025)
by: Chockchowwat, Supawit, et al.
Published: (2025)
Kishu: Time-Traveling for Computational Notebooks
by: Li, Zhaoheng, et al.
Published: (2024)
by: Li, Zhaoheng, et al.
Published: (2024)
Performance Characterization of Containers in Edge Computing
by: Gupta, Ragini, et al.
Published: (2025)
by: Gupta, Ragini, et al.
Published: (2025)
ElasticNotebook: Enabling Live Migration for Computational Notebooks
by: Li, Zhaoheng, et al.
Published: (2023)
by: Li, Zhaoheng, et al.
Published: (2023)
EcoLens: Leveraging Multi-Objective Bayesian Optimization for Energy-Efficient Video Processing on Edge Devices
by: Civjan, Benjamin, et al.
Published: (2025)
by: Civjan, Benjamin, et al.
Published: (2025)
Viewport-based Neural 360° Image Compression
by: Liao, Jingwei, et al.
Published: (2026)
by: Liao, Jingwei, et al.
Published: (2026)
QoS-QoE Translation with Large Language Model
by: Yu, Yingjie, et al.
Published: (2026)
by: Yu, Yingjie, et al.
Published: (2026)
Revisiting Disaggregated Large Language Model Serving for Performance and Energy Implications
by: Li, Jiaxi, et al.
Published: (2025)
by: Li, Jiaxi, et al.
Published: (2025)
Report on the NSF Workshop on Sustainable Computing for Sustainability (NSF WSCS 2024)
by: Guérin, Roch, et al.
Published: (2024)
by: Guérin, Roch, et al.
Published: (2024)
Pseudo Dataset Generation for Out-of-Domain Multi-Camera View Recommendation
by: Lee, Kuan-Ying, et al.
Published: (2024)
by: Lee, Kuan-Ying, et al.
Published: (2024)
ORBIT: Cost-Effective Dataset Curation for Large Language Model Domain Adaptation with an Astronomy Case Study
by: Modesitt, Eric, et al.
Published: (2024)
by: Modesitt, Eric, et al.
Published: (2024)
Hold Onto That Thought: Assessing KV Cache Compression On Reasoning
by: Liu, Minghui, et al.
Published: (2025)
by: Liu, Minghui, et al.
Published: (2025)
Federated Transfer Learning with Task Personalization for Condition Monitoring in Ultrasonic Metal Welding
by: Eslaminia, Ahmadreza, et al.
Published: (2024)
by: Eslaminia, Ahmadreza, et al.
Published: (2024)
Cloud-Native Vector Search: A Comprehensive Performance Analysis
by: Li, Zhaoheng, et al.
Published: (2025)
by: Li, Zhaoheng, et al.
Published: (2025)
PBE Meets LLM: When Few Examples Aren't Few-Shot Enough
by: Zhang, Shuning, et al.
Published: (2025)
by: Zhang, Shuning, et al.
Published: (2025)
FlexGaussian: Flexible and Cost-Effective Training-Free Compression for 3D Gaussian Splatting
by: Tian, Boyuan, et al.
Published: (2025)
by: Tian, Boyuan, et al.
Published: (2025)
What Makes In-context Learning Effective for Mathematical Reasoning: A Theoretical Analysis
by: Liu, Jiayu, et al.
Published: (2024)
by: Liu, Jiayu, et al.
Published: (2024)
Adaptive Unknown Fault Detection and Few-Shot Continual Learning for Condition Monitoring in Ultrasonic Metal Welding
by: Eslaminia, Ahmadreza, et al.
Published: (2026)
by: Eslaminia, Ahmadreza, et al.
Published: (2026)
Beyond Chain-of-Thought, Effective Graph-of-Thought Reasoning in Language Models
by: Yao, Yao, et al.
Published: (2023)
by: Yao, Yao, et al.
Published: (2023)
MegaCacheX: Towards Cost-Effective Hierarchical Collaborative Content Caching in Emerging Mega-Constellations
by: Shi, Haoyang, et al.
Published: (2025)
by: Shi, Haoyang, et al.
Published: (2025)
The Cost of Reasoning: Chain-of-Thought Induces Overconfidence in Vision-Language Models
by: Welch, Robert, et al.
Published: (2026)
by: Welch, Robert, et al.
Published: (2026)
Enhancing Neural Adaptive Wireless Video Streaming via Lower-Layer Information Exposure and Online Tuning
by: Zhao, Lingzhi, et al.
Published: (2025)
by: Zhao, Lingzhi, et al.
Published: (2025)
Apprentice training in South Australia
Published: (1925)
Published: (1925)
Apprentice of the Year award for SVN
Published: (2025)
Published: (2025)
Evaluating Spatio-Temporal Forecasting Trade-offs Between Graph Neural Networks and Foundation Models
by: Gupta, Ragini, et al.
Published: (2025)
by: Gupta, Ragini, et al.
Published: (2025)
LLM-ADAM: A Generalizable LLM Agent Framework for Pre-Print Anomaly Detection in Additive Manufacturing
by: Eslaminia, Ahmadreza, et al.
Published: (2026)
by: Eslaminia, Ahmadreza, et al.
Published: (2026)
Similar Items
-
UOUO: Uncontextualized Uncommon Objects for Measuring Knowledge Horizons of Vision Language Models
by: Pi, Xinyu, et al.
Published: (2024) -
VTool-R1: VLMs Learn to Think with Images via Reinforcement Learning on Multimodal Tool Use
by: Wu, Mingyuan, et al.
Published: (2025) -
Aha Moment Revisited: Are VLMs Truly Capable of Self Verification in Inference-time Scaling?
by: Wu, Mingyuan, et al.
Published: (2025) -
Spatio-Temporal LLM: Reasoning about Environments and Actions
by: Zheng, Haozhen, et al.
Published: (2025) -
AquaVLM: Improving Underwater Situation Awareness with Mobile Vision Language Models
by: Tian, Beitong, et al.
Published: (2025)