Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Janjua, Muhammad Kamran, Silva, Hugo, Niu, Di, Rashidi, Bahador |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Panoptic Pairwise Distortion Graph
von: Janjua, Muhammad Kamran, et al.
Veröffentlicht: (2026)
von: Janjua, Muhammad Kamran, et al.
Veröffentlicht: (2026)
Show, Don't Tell: Morphing Latent Reasoning into Image Generation
von: Chen, Harold Haodong, et al.
Veröffentlicht: (2026)
von: Chen, Harold Haodong, et al.
Veröffentlicht: (2026)
Show, Don't Tell: Detecting Novel Objects by Watching Human Videos
von: Akl, James, et al.
Veröffentlicht: (2026)
von: Akl, James, et al.
Veröffentlicht: (2026)
Learning Truncated Causal History Model for Video Restoration
von: Ghasemabadi, Amirhosein, et al.
Veröffentlicht: (2024)
von: Ghasemabadi, Amirhosein, et al.
Veröffentlicht: (2024)
Tell, Don't Show!: Language Guidance Eases Transfer Across Domains in Images and Videos
von: Kalluri, Tarun, et al.
Veröffentlicht: (2024)
von: Kalluri, Tarun, et al.
Veröffentlicht: (2024)
ShowTable: Unlocking Creative Table Visualization with Collaborative Reflection and Refinement
von: Liu, Zhihang, et al.
Veröffentlicht: (2025)
von: Liu, Zhihang, et al.
Veröffentlicht: (2025)
Pixels Don't Lie (But Your Detector Might): Bootstrapping MLLM-as-a-Judge for Trustworthy Deepfake Detection and Reasoning Supervision
von: Kuckreja, Kartik, et al.
Veröffentlicht: (2026)
von: Kuckreja, Kartik, et al.
Veröffentlicht: (2026)
Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation
von: Zhang, David Junhao, et al.
Veröffentlicht: (2023)
von: Zhang, David Junhao, et al.
Veröffentlicht: (2023)
Show, Tell and Summarize: Dense Video Captioning Using Visual Cue Aided Sentence Summarization
von: Zhang, Zhiwang, et al.
Veröffentlicht: (2025)
von: Zhang, Zhiwang, et al.
Veröffentlicht: (2025)
Grounding Degradations in Natural Language for All-In-One Video Restoration
von: Janjua, Muhammad Kamran, et al.
Veröffentlicht: (2025)
von: Janjua, Muhammad Kamran, et al.
Veröffentlicht: (2025)
Know-Show: Benchmarking Video-Language Models on Spatio-Temporal Grounded Reasoning
von: Sugandhika, Chinthani, et al.
Veröffentlicht: (2025)
von: Sugandhika, Chinthani, et al.
Veröffentlicht: (2025)
Show and Guide: Instructional-Plan Grounded Vision and Language Model
von: Glória-Silva, Diogo, et al.
Veröffentlicht: (2024)
von: Glória-Silva, Diogo, et al.
Veröffentlicht: (2024)
CascadedGaze: Efficiency in Global Context Extraction for Image Restoration
von: Ghasemabadi, Amirhosein, et al.
Veröffentlicht: (2024)
von: Ghasemabadi, Amirhosein, et al.
Veröffentlicht: (2024)
PatchCue: Enhancing Vision-Language Model Reasoning with Patch-Based Visual Cues
von: Qi, Yukun, et al.
Veröffentlicht: (2026)
von: Qi, Yukun, et al.
Veröffentlicht: (2026)
Show and Segment: Universal Medical Image Segmentation via In-Context Learning
von: Gao, Yunhe, et al.
Veröffentlicht: (2025)
von: Gao, Yunhe, et al.
Veröffentlicht: (2025)
Show and Tell: Visually Explainable Deep Neural Nets via Spatially-Aware Concept Bottleneck Models
von: Benou, Itay, et al.
Veröffentlicht: (2025)
von: Benou, Itay, et al.
Veröffentlicht: (2025)
Don't Let the Video Speak: Audio-Contrastive Preference Optimization for Audio-Visual Language Models
von: Baid, Ami, et al.
Veröffentlicht: (2026)
von: Baid, Ami, et al.
Veröffentlicht: (2026)
ShowHowTo: Generating Scene-Conditioned Step-by-Step Visual Instructions
von: Souček, Tomáš, et al.
Veröffentlicht: (2024)
von: Souček, Tomáš, et al.
Veröffentlicht: (2024)
Vision Transformers Exhibit Human-Like Biases: Evidence of Orientation and Color Selectivity, Categorical Perception, and Phase Transitions
von: Bahador, Nooshin
Veröffentlicht: (2025)
von: Bahador, Nooshin
Veröffentlicht: (2025)
ShowUI: One Vision-Language-Action Model for GUI Visual Agent
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
Don't Just Chase "Highlighted Tokens" in MLLMs: Revisiting Visual Holistic Context Retention
von: Zou, Xin, et al.
Veröffentlicht: (2025)
von: Zou, Xin, et al.
Veröffentlicht: (2025)
Show or Tell? Effectively prompting Vision-Language Models for semantic segmentation
von: Avogaro, Niccolo, et al.
Veröffentlicht: (2025)
von: Avogaro, Niccolo, et al.
Veröffentlicht: (2025)
Show or Tell? A Benchmark To Evaluate Visual and Textual Prompts in Semantic Segmentation
von: Rosi, Gabriele, et al.
Veröffentlicht: (2025)
von: Rosi, Gabriele, et al.
Veröffentlicht: (2025)
Surely Large Multimodal Models (Don't) Excel in Visual Species Recognition?
von: Liu, Tian, et al.
Veröffentlicht: (2025)
von: Liu, Tian, et al.
Veröffentlicht: (2025)
CLIPSwarm: Generating Drone Shows from Text Prompts with Vision-Language Models
von: Pueyo, Pablo, et al.
Veröffentlicht: (2024)
von: Pueyo, Pablo, et al.
Veröffentlicht: (2024)
Multimodal Latent Reasoning via Hierarchical Visual Cues Injection
von: Zhang, Yiming, et al.
Veröffentlicht: (2026)
von: Zhang, Yiming, et al.
Veröffentlicht: (2026)
Show, don't tell -- Providing Visual Error Feedback for Handwritten Documents
von: Yasin, Said, et al.
Veröffentlicht: (2026)
von: Yasin, Said, et al.
Veröffentlicht: (2026)
Show-o2: Improved Native Unified Multimodal Models
von: Xie, Jinheng, et al.
Veröffentlicht: (2025)
von: Xie, Jinheng, et al.
Veröffentlicht: (2025)
Visual Program Distillation: Distilling Tools and Programmatic Reasoning into Vision-Language Models
von: Hu, Yushi, et al.
Veröffentlicht: (2023)
von: Hu, Yushi, et al.
Veröffentlicht: (2023)
Visually Dehallucinative Instruction Generation: Know What You Don't Know
von: Cha, Sungguk, et al.
Veröffentlicht: (2024)
von: Cha, Sungguk, et al.
Veröffentlicht: (2024)
Show Me: Unifying Instructional Image and Video Generation with Diffusion Models
von: Pu, Yujiang, et al.
Veröffentlicht: (2025)
von: Pu, Yujiang, et al.
Veröffentlicht: (2025)
ShowUI-Aloha: Human-Taught GUI Agent
von: Zhang, Yichun, et al.
Veröffentlicht: (2026)
von: Zhang, Yichun, et al.
Veröffentlicht: (2026)
KAN or MLP? Point Cloud Shows the Way Forward
von: Shi, Yan, et al.
Veröffentlicht: (2025)
von: Shi, Yan, et al.
Veröffentlicht: (2025)
Describe, Don't Dictate: Semantic Image Editing with Natural Language Intent
von: Ci, En, et al.
Veröffentlicht: (2025)
von: Ci, En, et al.
Veröffentlicht: (2025)
Computer Vision Models Show Human-Like Sensitivity to Geometric and Topological Concepts
von: Wang, Zekun, et al.
Veröffentlicht: (2025)
von: Wang, Zekun, et al.
Veröffentlicht: (2025)
Diffusion Is Your Friend in Show, Suggest and Tell
von: Hu, Jia Cheng, et al.
Veröffentlicht: (2025)
von: Hu, Jia Cheng, et al.
Veröffentlicht: (2025)
Mutual Information Analysis in Multimodal Learning Systems
von: Hadizadeh, Hadi, et al.
Veröffentlicht: (2024)
von: Hadizadeh, Hadi, et al.
Veröffentlicht: (2024)
SMOL-MapSeg: Show Me One Label as prompt
von: Yuan, Yunshuang, et al.
Veröffentlicht: (2025)
von: Yuan, Yunshuang, et al.
Veröffentlicht: (2025)
NeIn: Telling What You Don't Want
von: Bui, Nhat-Tan, et al.
Veröffentlicht: (2024)
von: Bui, Nhat-Tan, et al.
Veröffentlicht: (2024)
PixelMan: Consistent Object Editing with Diffusion Models via Pixel Manipulation and Generation
von: Jiang, Liyao, et al.
Veröffentlicht: (2024)
von: Jiang, Liyao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Panoptic Pairwise Distortion Graph
von: Janjua, Muhammad Kamran, et al.
Veröffentlicht: (2026) -
Show, Don't Tell: Morphing Latent Reasoning into Image Generation
von: Chen, Harold Haodong, et al.
Veröffentlicht: (2026) -
Show, Don't Tell: Detecting Novel Objects by Watching Human Videos
von: Akl, James, et al.
Veröffentlicht: (2026) -
Learning Truncated Causal History Model for Video Restoration
von: Ghasemabadi, Amirhosein, et al.
Veröffentlicht: (2024) -
Tell, Don't Show!: Language Guidance Eases Transfer Across Domains in Images and Videos
von: Kalluri, Tarun, et al.
Veröffentlicht: (2024)