CALICO: Part-Focused Semantic Co-Segmentation with Large Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Nguyen, Kiet A., Juvekar, Adheesh, Yu, Tianjiao, Wahed, Muntasir, Lourentzou, Ismini |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PRIMA: Multi-Image Vision-Language Models for Reasoning Segmentation
by: Wahed, Muntasir, et al.
Published: (2024)
by: Wahed, Muntasir, et al.
Published: (2024)
Counterfactual Segmentation Reasoning: Diagnosing and Mitigating Pixel-Grounding Hallucination
by: Li, Xinzhuo, et al.
Published: (2025)
by: Li, Xinzhuo, et al.
Published: (2025)
Part$^{2}$GS: Part-aware Modeling of Articulated Objects using 3D Gaussian Splatting
by: Yu, Tianjiao, et al.
Published: (2025)
by: Yu, Tianjiao, et al.
Published: (2025)
DreamPartGen: Semantically Grounded Part-Level 3D Generation via Collaborative Latent Denoising
by: Yu, Tianjiao, et al.
Published: (2026)
by: Yu, Tianjiao, et al.
Published: (2026)
Uncertainty in Action: Confidence Elicitation in Embodied Agents
by: Yu, Tianjiao, et al.
Published: (2025)
by: Yu, Tianjiao, et al.
Published: (2025)
PyraTok: Language-Aligned Pyramidal Tokenizer for Video Understanding and Generation
by: Susladkar, Onkar, et al.
Published: (2026)
by: Susladkar, Onkar, et al.
Published: (2026)
RewardFlow: Generate Images by Optimizing What You Reward
by: Susladkar, Onkar, et al.
Published: (2026)
by: Susladkar, Onkar, et al.
Published: (2026)
Tiny but Trusted: Efficient Vision-Language Reasoning for Time-Series Anomaly Detection
by: Zhou, Xiaona, et al.
Published: (2026)
by: Zhou, Xiaona, et al.
Published: (2026)
CoRe3D: Collaborative Reasoning as a Foundation for 3D Intelligence
by: Yu, Tianjiao, et al.
Published: (2025)
by: Yu, Tianjiao, et al.
Published: (2025)
Commonsense for Zero-Shot Natural Language Video Localization
by: Holla, Meghana, et al.
Published: (2023)
by: Holla, Meghana, et al.
Published: (2023)
MOCHA: Are Code Language Models Robust Against Multi-Turn Malicious Coding Prompts?
by: Wahed, Muntasir, et al.
Published: (2025)
by: Wahed, Muntasir, et al.
Published: (2025)
Best of Both Worlds: Multimodal Reasoning and Generation via Unified Discrete Flow Matching
by: Susladkar, Onkar, et al.
Published: (2026)
by: Susladkar, Onkar, et al.
Published: (2026)
3D-VCD: Hallucination Mitigation in 3D-LLM Embodied Agents through Visual Contrastive Decoding
by: Ogunleye, Makanjuola, et al.
Published: (2026)
by: Ogunleye, Makanjuola, et al.
Published: (2026)
Phantom: Physics-Infused Video Generation via Joint Modeling of Visual and Latent Physical Dynamics
by: Shen, Ying, et al.
Published: (2026)
by: Shen, Ying, et al.
Published: (2026)
ELBA: Learning by Asking for Embodied Visual Navigation and Task Completion
by: Shen, Ying, et al.
Published: (2023)
by: Shen, Ying, et al.
Published: (2023)
Toward Cognitive Supersensing in Multimodal Large Language Model
by: Li, Boyi, et al.
Published: (2026)
by: Li, Boyi, et al.
Published: (2026)
CALICO: Confident Active Learning with Integrated Calibration
by: Querol, Lorenzo S., et al.
Published: (2024)
by: Querol, Lorenzo S., et al.
Published: (2024)
Weather-Robust Cross-View Geo-Localization via Prototype-Based Semantic Part Discovery
by: Tran, Chi-Nguyen, et al.
Published: (2026)
by: Tran, Chi-Nguyen, et al.
Published: (2026)
Context-Aware Semantic Segmentation: Enhancing Pixel-Level Understanding with Large Language Models for Advanced Vision Applications
by: Rahman, Ben
Published: (2025)
by: Rahman, Ben
Published: (2025)
Representation Separation for Semantic Segmentation with Vision Transformers
by: Hong, Yuanduo, et al.
Published: (2022)
by: Hong, Yuanduo, et al.
Published: (2022)
Vision-Language Model Purified Semi-Supervised Semantic Segmentation for Remote Sensing Images
by: Wang, Shanwen, et al.
Published: (2026)
by: Wang, Shanwen, et al.
Published: (2026)
Emergent Open-Vocabulary Semantic Segmentation from Off-the-shelf Vision-Language Models
by: Luo, Jiayun, et al.
Published: (2023)
by: Luo, Jiayun, et al.
Published: (2023)
A Study on Unsupervised Domain Adaptation for Semantic Segmentation in the Era of Vision-Language Models
by: Schwonberg, Manuel, et al.
Published: (2024)
by: Schwonberg, Manuel, et al.
Published: (2024)
Uncertainty-guided Compositional Alignment with Part-to-Whole Semantic Representativeness in Hyperbolic Vision-Language Models
by: Kim, Hayeon, et al.
Published: (2026)
by: Kim, Hayeon, et al.
Published: (2026)
Urban Socio-Semantic Segmentation with Vision-Language Reasoning
by: Wang, Yu, et al.
Published: (2026)
by: Wang, Yu, et al.
Published: (2026)
Cross-Layer Vision Smoothing: Enhancing Visual Understanding via Sustained Focus on Key Objects in Large Vision-Language Models
by: Zhao, Jianfei, et al.
Published: (2025)
by: Zhao, Jianfei, et al.
Published: (2025)
A Visual Semantic Adaptive Watermark grounded by Prefix-Tuning for Large Vision-Language Model
by: Zheng, Qi, et al.
Published: (2026)
by: Zheng, Qi, et al.
Published: (2026)
Attention Prompting on Image for Large Vision-Language Models
by: Yu, Runpeng, et al.
Published: (2024)
by: Yu, Runpeng, et al.
Published: (2024)
G4Seg: Generation for Inexact Segmentation Refinement with Diffusion Models
by: Zhang, Tianjiao, et al.
Published: (2025)
by: Zhang, Tianjiao, et al.
Published: (2025)
XAI-Enhanced Semantic Segmentation Models for Visual Quality Inspection
by: Clement, Tobias, et al.
Published: (2024)
by: Clement, Tobias, et al.
Published: (2024)
Compound Expression Recognition via Large Vision-Language Models
by: Yu, Jun, et al.
Published: (2025)
by: Yu, Jun, et al.
Published: (2025)
Continuous Vision-Language-Action Co-Learning with Semantic-Physical Alignment for Behavioral Cloning
by: Qi, Xiuxiu, et al.
Published: (2025)
by: Qi, Xiuxiu, et al.
Published: (2025)
Leveraging Chat-Based Large Vision Language Models for Multimodal Out-Of-Context Detection
by: Shalabi, Fatma, et al.
Published: (2024)
by: Shalabi, Fatma, et al.
Published: (2024)
Remodeling Semantic Relationships in Vision-Language Fine-Tuning
by: Wu, Xiangyang, et al.
Published: (2025)
by: Wu, Xiangyang, et al.
Published: (2025)
Power of Boundary and Reflection: Semantic Transparent Object Segmentation using Pyramid Vision Transformer with Transparent Cues
by: Vu, Tuan-Anh, et al.
Published: (2025)
by: Vu, Tuan-Anh, et al.
Published: (2025)
Vision Transformer-Conditioned UNet for Domain-Adaptive Semantic Segmentation
by: Ortega, Joel Valdivia, et al.
Published: (2026)
by: Ortega, Joel Valdivia, et al.
Published: (2026)
Sanitizing Manufacturing Dataset Labels Using Vision-Language Models
by: Mahjourian, Nazanin, et al.
Published: (2025)
by: Mahjourian, Nazanin, et al.
Published: (2025)
CoCo-SAM3: Harnessing Concept Conflict in Open-Vocabulary Semantic Segmentation
by: Chen, Yanhui, et al.
Published: (2026)
by: Chen, Yanhui, et al.
Published: (2026)
CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models
by: Cheng, Zihui, et al.
Published: (2024)
by: Cheng, Zihui, et al.
Published: (2024)
Global Semantic-Guided Sub-image Feature Weight Allocation in High-Resolution Large Vision-Language Models
by: Liang, Yuxuan, et al.
Published: (2025)
by: Liang, Yuxuan, et al.
Published: (2025)
Similar Items
-
PRIMA: Multi-Image Vision-Language Models for Reasoning Segmentation
by: Wahed, Muntasir, et al.
Published: (2024) -
Counterfactual Segmentation Reasoning: Diagnosing and Mitigating Pixel-Grounding Hallucination
by: Li, Xinzhuo, et al.
Published: (2025) -
Part$^{2}$GS: Part-aware Modeling of Articulated Objects using 3D Gaussian Splatting
by: Yu, Tianjiao, et al.
Published: (2025) -
DreamPartGen: Semantically Grounded Part-Level 3D Generation via Collaborative Latent Denoising
by: Yu, Tianjiao, et al.
Published: (2026) -
Uncertainty in Action: Confidence Elicitation in Embodied Agents
by: Yu, Tianjiao, et al.
Published: (2025)