See&Trek: Training-Free Spatial Prompting for Multimodal Large Language Model
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Pengteng, Song, Pinhao, Li, Wuyang, Guo, Weiyu, Yao, Huizai, Xu, Yijie, Liu, Dugang, Xiong, Hui |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
EventVL: Understand Event Streams via Multimodal Large Language Model
von: Li, Pengteng, et al.
Veröffentlicht: (2025)
von: Li, Pengteng, et al.
Veröffentlicht: (2025)
DeblurSplat: SfM-free 3D Gaussian Splatting with Event Camera for Robust Deblurring
von: Li, Pengteng, et al.
Veröffentlicht: (2025)
von: Li, Pengteng, et al.
Veröffentlicht: (2025)
Beyond Boundaries: Leveraging Vision Foundation Models for Source-Free Object Detection
von: Yao, Huizai, et al.
Veröffentlicht: (2025)
von: Yao, Huizai, et al.
Veröffentlicht: (2025)
SEE: See Everything Every Time -- Adaptive Brightness Adjustment for Broad Light Range Images via Events
von: Lu, Yunfan, et al.
Veröffentlicht: (2025)
von: Lu, Yunfan, et al.
Veröffentlicht: (2025)
From Events to Clarity: The Event-Guided Diffusion Framework for Dehazing
von: Wang, Ling, et al.
Veröffentlicht: (2025)
von: Wang, Ling, et al.
Veröffentlicht: (2025)
From Events to Enhancement: A Survey on Event-Based Imaging Technologies
von: Lu, Yunfan, et al.
Veröffentlicht: (2025)
von: Lu, Yunfan, et al.
Veröffentlicht: (2025)
You only need 4 extra tokens: Synergistic Test-time Adaptation for LLMs
von: Xu, Yijie, et al.
Veröffentlicht: (2025)
von: Xu, Yijie, et al.
Veröffentlicht: (2025)
Robust Object Detection of Underwater Robot based on Domain Generalization
von: Song, Pinhao
Veröffentlicht: (2025)
von: Song, Pinhao
Veröffentlicht: (2025)
Less is More: Token-Efficient Video-QA via Adaptive Frame-Pruning and Semantic Graph Integration
von: Wang, Shaoguang, et al.
Veröffentlicht: (2025)
von: Wang, Shaoguang, et al.
Veröffentlicht: (2025)
Event Camera Demosaicing via Swin Transformer and Pixel-focus Loss
von: Lu, Yunfan, et al.
Veröffentlicht: (2024)
von: Lu, Yunfan, et al.
Veröffentlicht: (2024)
ControlMLLM: Training-Free Visual Prompt Learning for Multimodal Large Language Models
von: Wu, Mingrui, et al.
Veröffentlicht: (2024)
von: Wu, Mingrui, et al.
Veröffentlicht: (2024)
Where to Focus: Query-Modulated Multimodal Keyframe Selection for Long Video Understanding
von: Wang, Shaoguang, et al.
Veröffentlicht: (2026)
von: Wang, Shaoguang, et al.
Veröffentlicht: (2026)
Seeing Together: Multi-Robot Cooperative Egocentric Spatial Reasoning with Multimodal Large Language Models
von: Peng, Kunyu, et al.
Veröffentlicht: (2026)
von: Peng, Kunyu, et al.
Veröffentlicht: (2026)
Seeing is Believing? Mitigating OCR Hallucinations in Multimodal Large Language Models
von: He, Zhentao, et al.
Veröffentlicht: (2025)
von: He, Zhentao, et al.
Veröffentlicht: (2025)
Prompt Reinjection: Alleviating Prompt Forgetting in Multimodal Diffusion Transformers
von: Yao, Yuxuan, et al.
Veröffentlicht: (2026)
von: Yao, Yuxuan, et al.
Veröffentlicht: (2026)
HR-INR: Continuous Space-Time Video Super-Resolution via Event Camera
von: Lu, Yunfan, et al.
Veröffentlicht: (2024)
von: Lu, Yunfan, et al.
Veröffentlicht: (2024)
Source-Free Object Detection with Detection Transformer
von: Yao, Huizai, et al.
Veröffentlicht: (2025)
von: Yao, Huizai, et al.
Veröffentlicht: (2025)
Learning to See What You Need: Gaze Attention for Multimodal Large Language Models
von: Song, Junha, et al.
Veröffentlicht: (2026)
von: Song, Junha, et al.
Veröffentlicht: (2026)
Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces
von: Yang, Jihan, et al.
Veröffentlicht: (2024)
von: Yang, Jihan, et al.
Veröffentlicht: (2024)
See It, Say It, Sorted: An Iterative Training-Free Framework for Visually-Grounded Multimodal Reasoning in LVLMs
von: Zhang, Yongchang, et al.
Veröffentlicht: (2026)
von: Zhang, Yongchang, et al.
Veröffentlicht: (2026)
FinePOSE: Fine-Grained Prompt-Driven 3D Human Pose Estimation via Diffusion Models
von: Xu, Jinglin, et al.
Veröffentlicht: (2024)
von: Xu, Jinglin, et al.
Veröffentlicht: (2024)
Bridge the Gap Between Visual and Linguistic Comprehension for Generalized Zero-shot Semantic Segmentation
von: Guo, Xiaoqing, et al.
Veröffentlicht: (2025)
von: Guo, Xiaoqing, et al.
Veröffentlicht: (2025)
AmPLe: Supporting Vision-Language Models via Adaptive-Debiased Ensemble Multi-Prompt Learning
von: Song, Fei, et al.
Veröffentlicht: (2025)
von: Song, Fei, et al.
Veröffentlicht: (2025)
FreeMotion: MoCap-Free Human Motion Synthesis with Multimodal Large Language Models
von: Zhang, Zhikai, et al.
Veröffentlicht: (2024)
von: Zhang, Zhikai, et al.
Veröffentlicht: (2024)
Seeing Clearly without Training: Mitigating Hallucinations in Multimodal LLMs for Remote Sensing
von: Liu, Yi, et al.
Veröffentlicht: (2026)
von: Liu, Yi, et al.
Veröffentlicht: (2026)
VP-Bench: A Comprehensive Benchmark for Visual Prompting in Multimodal Large Language Models
von: Xu, Mingjie, et al.
Veröffentlicht: (2025)
von: Xu, Mingjie, et al.
Veröffentlicht: (2025)
Domain Similarity-Perceived Label Assignment for Domain Generalized Underwater Object Detection
von: Li, Xisheng, et al.
Veröffentlicht: (2023)
von: Li, Xisheng, et al.
Veröffentlicht: (2023)
Spatial Memory for Out-of-Vision Manipulation in Vision-Language-Action
von: Li, Pengteng, et al.
Veröffentlicht: (2026)
von: Li, Pengteng, et al.
Veröffentlicht: (2026)
Adversarial Prompt Injection Attack on Multimodal Large Language Models
von: Ding, Meiwen, et al.
Veröffentlicht: (2026)
von: Ding, Meiwen, et al.
Veröffentlicht: (2026)
BLINK: Multimodal Large Language Models Can See but Not Perceive
von: Fu, Xingyu, et al.
Veröffentlicht: (2024)
von: Fu, Xingyu, et al.
Veröffentlicht: (2024)
Seeing to Act, Prompting to Specify: A Bayesian Factorization of Vision Language Action Policy
von: Xu, Kechun, et al.
Veröffentlicht: (2025)
von: Xu, Kechun, et al.
Veröffentlicht: (2025)
Unleashing the Temporal-Spatial Reasoning Capacity of GPT for Training-Free Audio and Language Referenced Video Object Segmentation
von: Huang, Shaofei, et al.
Veröffentlicht: (2024)
von: Huang, Shaofei, et al.
Veröffentlicht: (2024)
A Gated Cross-domain Collaborative Network for Underwater Object Detection
von: Dai, Linhui, et al.
Veröffentlicht: (2023)
von: Dai, Linhui, et al.
Veröffentlicht: (2023)
Visual Prompting in Multimodal Large Language Models: A Survey
von: Wu, Junda, et al.
Veröffentlicht: (2024)
von: Wu, Junda, et al.
Veröffentlicht: (2024)
Seeing the Arrow of Time in Large Multimodal Models
von: Xue, Zihui, et al.
Veröffentlicht: (2025)
von: Xue, Zihui, et al.
Veröffentlicht: (2025)
LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding
von: Li, Hongyu, et al.
Veröffentlicht: (2025)
von: Li, Hongyu, et al.
Veröffentlicht: (2025)
Divide, Conquer and Combine: A Training-Free Framework for High-Resolution Image Perception in Multimodal Large Language Models
von: Wang, Wenbin, et al.
Veröffentlicht: (2024)
von: Wang, Wenbin, et al.
Veröffentlicht: (2024)
FashionLOGO: Prompting Multimodal Large Language Models for Fashion Logo Embeddings
von: Wang, Zhen, et al.
Veröffentlicht: (2023)
von: Wang, Zhen, et al.
Veröffentlicht: (2023)
Actial: Activate Spatial Reasoning Ability of Multimodal Large Language Models
von: Zhan, Xiaoyu, et al.
Veröffentlicht: (2025)
von: Zhan, Xiaoyu, et al.
Veröffentlicht: (2025)
Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes
von: Feng, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Feng, Zhiyuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
EventVL: Understand Event Streams via Multimodal Large Language Model
von: Li, Pengteng, et al.
Veröffentlicht: (2025) -
DeblurSplat: SfM-free 3D Gaussian Splatting with Event Camera for Robust Deblurring
von: Li, Pengteng, et al.
Veröffentlicht: (2025) -
Beyond Boundaries: Leveraging Vision Foundation Models for Source-Free Object Detection
von: Yao, Huizai, et al.
Veröffentlicht: (2025) -
SEE: See Everything Every Time -- Adaptive Brightness Adjustment for Broad Light Range Images via Events
von: Lu, Yunfan, et al.
Veröffentlicht: (2025) -
From Events to Clarity: The Event-Guided Diffusion Framework for Dehazing
von: Wang, Ling, et al.
Veröffentlicht: (2025)