Chain-of-Caption: Training-free improvement of multimodal large language model on referring expression comprehension
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pang, Yik Lung, Oh, Changjae |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Sparse multi-view hand-object reconstruction for unseen environments
von: Pang, Yik Lung, et al.
Veröffentlicht: (2024)
von: Pang, Yik Lung, et al.
Veröffentlicht: (2024)
Learning human-to-robot handovers through 3D scene reconstruction
von: Wu, Yuekun, et al.
Veröffentlicht: (2025)
von: Wu, Yuekun, et al.
Veröffentlicht: (2025)
Stereo Hand-Object Reconstruction for Human-to-Robot Handover
von: Pang, Yik Lung, et al.
Veröffentlicht: (2024)
von: Pang, Yik Lung, et al.
Veröffentlicht: (2024)
Toward Human-Robot Teaming: Learning Handover Behaviors from 3D Scenes
von: Wu, Yuekun, et al.
Veröffentlicht: (2025)
von: Wu, Yuekun, et al.
Veröffentlicht: (2025)
The Detector Teaches Itself: Lightweight Self-Supervised Adaptation for Open-Vocabulary Object Detection
von: Wan, Yazhe, et al.
Veröffentlicht: (2026)
von: Wan, Yazhe, et al.
Veröffentlicht: (2026)
MIMO: A medical vision language model with visual referring multimodal input and pixel grounding multimodal output
von: Chen, Yanyuan, et al.
Veröffentlicht: (2025)
von: Chen, Yanyuan, et al.
Veröffentlicht: (2025)
FlowOVD: Learning Generative Latent Flows for Zero-shot Open-vocabulary Detection
von: Wei, Yao, et al.
Veröffentlicht: (2026)
von: Wei, Yao, et al.
Veröffentlicht: (2026)
Improving Image De-raining Using Reference-Guided Transformers
von: Ye, Zihao, et al.
Veröffentlicht: (2024)
von: Ye, Zihao, et al.
Veröffentlicht: (2024)
In-context learning enables multimodal large language models to classify cancer pathology images
von: Ferber, Dyke, et al.
Veröffentlicht: (2024)
von: Ferber, Dyke, et al.
Veröffentlicht: (2024)
Evaluating point-light biological motion in multimodal large language models
von: Kadambi, Akila, et al.
Veröffentlicht: (2025)
von: Kadambi, Akila, et al.
Veröffentlicht: (2025)
Learning by Erasing: Conditional Entropy based Transferable Out-Of-Distribution Detection
von: Xing, Meng, et al.
Veröffentlicht: (2022)
von: Xing, Meng, et al.
Veröffentlicht: (2022)
GPT-4o: Visual perception performance of multimodal large language models in piglet activity understanding
von: Wu, Yiqi, et al.
Veröffentlicht: (2024)
von: Wu, Yiqi, et al.
Veröffentlicht: (2024)
On the robustness of multimodal language model towards distractions
von: Liu, Ming, et al.
Veröffentlicht: (2025)
von: Liu, Ming, et al.
Veröffentlicht: (2025)
A benchmark multimodal oro-dental dataset for large vision-language models
von: Lv, Haoxin, et al.
Veröffentlicht: (2025)
von: Lv, Haoxin, et al.
Veröffentlicht: (2025)
Improving Generalization of Language-Conditioned Robot Manipulation
von: Cui, Chenglin, et al.
Veröffentlicht: (2025)
von: Cui, Chenglin, et al.
Veröffentlicht: (2025)
Open-vocabulary object 6D pose estimation
von: Corsetti, Jaime, et al.
Veröffentlicht: (2023)
von: Corsetti, Jaime, et al.
Veröffentlicht: (2023)
Diffusion-driven GAN Inversion for Multi-Modal Face Image Generation
von: Kim, Jihyun, et al.
Veröffentlicht: (2024)
von: Kim, Jihyun, et al.
Veröffentlicht: (2024)
MedViLaM: A multimodal large language model with advanced generalizability and explainability for medical data understanding and generation
von: Xu, Lijian, et al.
Veröffentlicht: (2024)
von: Xu, Lijian, et al.
Veröffentlicht: (2024)
When language and vision meet road safety: leveraging multimodal large language models for video-based traffic accident analysis
von: Zhang, Ruixuan, et al.
Veröffentlicht: (2025)
von: Zhang, Ruixuan, et al.
Veröffentlicht: (2025)
SalsaAgent: A multimodal embodied language model for interactive dance generation
von: Yazdian, Payam Jome, et al.
Veröffentlicht: (2026)
von: Yazdian, Payam Jome, et al.
Veröffentlicht: (2026)
VLA-Mark: A cross modal watermark for large vision-language alignment model
von: Liu, Shuliang, et al.
Veröffentlicht: (2025)
von: Liu, Shuliang, et al.
Veröffentlicht: (2025)
Adaptive Multi-Modal Control of Digital Human Hand Synthesis Using a Region-Aware Cycle Loss
von: Fu, Qifan, et al.
Veröffentlicht: (2024)
von: Fu, Qifan, et al.
Veröffentlicht: (2024)
HanDrawer: Leveraging Spatial Information to Render Realistic Hands Using a Conditional Diffusion Model in Single Stage
von: Fu, Qifan, et al.
Veröffentlicht: (2025)
von: Fu, Qifan, et al.
Veröffentlicht: (2025)
High-resolution open-vocabulary object 6D pose estimation
von: Corsetti, Jaime, et al.
Veröffentlicht: (2024)
von: Corsetti, Jaime, et al.
Veröffentlicht: (2024)
Patch Matters: Training-free Fine-grained Image Caption Enhancement via Local Perception
von: Peng, Ruotian, et al.
Veröffentlicht: (2025)
von: Peng, Ruotian, et al.
Veröffentlicht: (2025)
Expert-level vision-language foundation model for real-world radiology and comprehensive evaluation
von: Liu, Xiaohong, et al.
Veröffentlicht: (2024)
von: Liu, Xiaohong, et al.
Veröffentlicht: (2024)
ECCV Caption: Correcting False Negatives by Collecting Machine-and-Human-verified Image-Caption Associations for MS-COCO
von: Chun, Sanghyuk, et al.
Veröffentlicht: (2022)
von: Chun, Sanghyuk, et al.
Veröffentlicht: (2022)
ResCLIP: Residual Attention for Training-free Dense Vision-language Inference
von: Yang, Yuhang, et al.
Veröffentlicht: (2024)
von: Yang, Yuhang, et al.
Veröffentlicht: (2024)
Assessing the alignment between infants' visual and linguistic experience using multimodal language models
von: Tan, Alvin Wei Ming, et al.
Veröffentlicht: (2025)
von: Tan, Alvin Wei Ming, et al.
Veröffentlicht: (2025)
Decoupled Video Generation with Chain of Training-free Diffusion Model Experts
von: Li, Wenhao, et al.
Veröffentlicht: (2024)
von: Li, Wenhao, et al.
Veröffentlicht: (2024)
VidChain: Chain-of-Tasks with Metric-based Direct Preference Optimization for Dense Video Captioning
von: Lee, Ji Soo, et al.
Veröffentlicht: (2025)
von: Lee, Ji Soo, et al.
Veröffentlicht: (2025)
A Training-free Synthetic Data Selection Method for Semantic Segmentation
von: Tang, Hao, et al.
Veröffentlicht: (2025)
von: Tang, Hao, et al.
Veröffentlicht: (2025)
Attacks on multimodal models
von: Iablochnikov, Viacheslav, et al.
Veröffentlicht: (2024)
von: Iablochnikov, Viacheslav, et al.
Veröffentlicht: (2024)
Human-like object concept representations emerge naturally in multimodal large language models
von: Du, Changde, et al.
Veröffentlicht: (2024)
von: Du, Changde, et al.
Veröffentlicht: (2024)
Do large language vision models understand 3D shapes?
von: Eppel, Sagi
Veröffentlicht: (2024)
von: Eppel, Sagi
Veröffentlicht: (2024)
Generating crossmodal gene expression from cancer histopathology improves multimodal AI predictions
von: Dey, Samiran, et al.
Veröffentlicht: (2025)
von: Dey, Samiran, et al.
Veröffentlicht: (2025)
Multilingual Training-Free Remote Sensing Image Captioning
von: Rebelo, Carlos, et al.
Veröffentlicht: (2025)
von: Rebelo, Carlos, et al.
Veröffentlicht: (2025)
Generalizing vision-language models to novel domains: A comprehensive survey
von: Li, Xinyao, et al.
Veröffentlicht: (2025)
von: Li, Xinyao, et al.
Veröffentlicht: (2025)
iSeg: An Iterative Refinement-based Framework for Training-free Segmentation
von: Sun, Lin, et al.
Veröffentlicht: (2024)
von: Sun, Lin, et al.
Veröffentlicht: (2024)
What do vision-language models see in the context? Investigating multimodal in-context learning
von: Santos, Gabriel O. dos, et al.
Veröffentlicht: (2025)
von: Santos, Gabriel O. dos, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Sparse multi-view hand-object reconstruction for unseen environments
von: Pang, Yik Lung, et al.
Veröffentlicht: (2024) -
Learning human-to-robot handovers through 3D scene reconstruction
von: Wu, Yuekun, et al.
Veröffentlicht: (2025) -
Stereo Hand-Object Reconstruction for Human-to-Robot Handover
von: Pang, Yik Lung, et al.
Veröffentlicht: (2024) -
Toward Human-Robot Teaming: Learning Handover Behaviors from 3D Scenes
von: Wu, Yuekun, et al.
Veröffentlicht: (2025) -
The Detector Teaches Itself: Lightweight Self-Supervised Adaptation for Open-Vocabulary Object Detection
von: Wan, Yazhe, et al.
Veröffentlicht: (2026)