Vision Harnessing Agent for Open Ad-hoc Segmentation
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Zilin, Yu, Stella X. |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Open Ad-hoc Categorization with Contextualized Feature Learning
por: Wang, Zilin, et al.
Publicado: (2025)
por: Wang, Zilin, et al.
Publicado: (2025)
Free-Grained Hierarchical Visual Recognition
por: Park, Seulki, et al.
Publicado: (2025)
por: Park, Seulki, et al.
Publicado: (2025)
Normalize Filters! Classical Wisdom for Deep Vision
por: Perez, Gustavo, et al.
Publicado: (2025)
por: Perez, Gustavo, et al.
Publicado: (2025)
Harnessing Vision Foundation Models for High-Performance, Training-Free Open Vocabulary Segmentation
por: Shi, Yuheng, et al.
Publicado: (2024)
por: Shi, Yuheng, et al.
Publicado: (2024)
Interpretable Embedding for Ad-hoc Video Search
por: Wu, Jiaxin, et al.
Publicado: (2024)
por: Wu, Jiaxin, et al.
Publicado: (2024)
SHED Light on Segmentation for Dense Prediction
por: Lee, Seung Hyun, et al.
Publicado: (2026)
por: Lee, Seung Hyun, et al.
Publicado: (2026)
Poster: Reliable 3D Reconstruction for Ad-hoc Edge Implementations
por: Absur, Md Nurul, et al.
Publicado: (2024)
por: Absur, Md Nurul, et al.
Publicado: (2024)
A Unified Framework for Event-based Frame Interpolation with Ad-hoc Deblurring in the Wild
por: Sun, Lei, et al.
Publicado: (2023)
por: Sun, Lei, et al.
Publicado: (2023)
Stronger, Fewer, & Superior: Harnessing Vision Foundation Models for Domain Generalized Semantic Segmentation
por: Wei, Zhixiang, et al.
Publicado: (2023)
por: Wei, Zhixiang, et al.
Publicado: (2023)
Harnessing Vision-Language Pretrained Models with Temporal-Aware Adaptation for Referring Video Object Segmentation
por: Zhou, Zikun, et al.
Publicado: (2024)
por: Zhou, Zikun, et al.
Publicado: (2024)
Aligning Forest and Trees in Images & Long Captions for Visually Grounded Understanding
por: Woo, Byeongju, et al.
Publicado: (2026)
por: Woo, Byeongju, et al.
Publicado: (2026)
Improving Interpretable Embeddings for Ad-hoc Video Search with Generative Captions and Multi-word Concept Bank
por: Wu, Jiaxin, et al.
Publicado: (2024)
por: Wu, Jiaxin, et al.
Publicado: (2024)
Structural Pruning via Spatial-aware Information Redundancy for Semantic Segmentation
por: Wu, Dongyue, et al.
Publicado: (2024)
por: Wu, Dongyue, et al.
Publicado: (2024)
Open-Vocabulary Camouflaged Object Segmentation with Cascaded Vision Language Models
por: Zhao, Kai, et al.
Publicado: (2025)
por: Zhao, Kai, et al.
Publicado: (2025)
Decomposed Vision-Language Alignment for Fine-Grained Open-Vocabulary Segmentation
por: Wang, Chenhao, et al.
Publicado: (2026)
por: Wang, Chenhao, et al.
Publicado: (2026)
CoCo-SAM3: Harnessing Concept Conflict in Open-Vocabulary Semantic Segmentation
por: Chen, Yanhui, et al.
Publicado: (2026)
por: Chen, Yanhui, et al.
Publicado: (2026)
FairCLIP: Harnessing Fairness in Vision-Language Learning
por: Luo, Yan, et al.
Publicado: (2024)
por: Luo, Yan, et al.
Publicado: (2024)
Selective, Regularized, and Calibrated: Harnessing Vision Foundation Models for Cross-Domain Few-Shot Semantic Segmentation
por: Ma, Junyuan, et al.
Publicado: (2026)
por: Ma, Junyuan, et al.
Publicado: (2026)
Collaborative Vision-Text Representation Optimizing for Open-Vocabulary Segmentation
por: Jiao, Siyu, et al.
Publicado: (2024)
por: Jiao, Siyu, et al.
Publicado: (2024)
Learning Partially-Decorrelated Common Spaces for Ad-hoc Video Search
por: Hu, Fan, et al.
Publicado: (2025)
por: Hu, Fan, et al.
Publicado: (2025)
TSegAgent: Zero-Shot Tooth Segmentation via Geometry-Aware Vision-Language Agents
por: Zhuang, Shaojie, et al.
Publicado: (2026)
por: Zhuang, Shaojie, et al.
Publicado: (2026)
Novel Category Discovery with X-Agent Attention for Open-Vocabulary Semantic Segmentation
por: Li, Jiahao, et al.
Publicado: (2025)
por: Li, Jiahao, et al.
Publicado: (2025)
Post-hoc Probabilistic Vision-Language Models
por: Baumann, Anton, et al.
Publicado: (2024)
por: Baumann, Anton, et al.
Publicado: (2024)
AerOSeg: Harnessing SAM for Open-Vocabulary Segmentation in Remote Sensing Images
por: Dutta, Saikat, et al.
Publicado: (2025)
por: Dutta, Saikat, et al.
Publicado: (2025)
Token Transformation Matters: Towards Faithful Post-hoc Explanation for Vision Transformer
por: Wu, Junyi, et al.
Publicado: (2024)
por: Wu, Junyi, et al.
Publicado: (2024)
Beyond-Labels: Advancing Open-Vocabulary Segmentation With Vision-Language Models
por: Rahman, Muhammad Atta ur, et al.
Publicado: (2025)
por: Rahman, Muhammad Atta ur, et al.
Publicado: (2025)
Novel View Synthesis from A Few Glimpses via Test-Time Natural Video Completion
por: Xu, Yan, et al.
Publicado: (2025)
por: Xu, Yan, et al.
Publicado: (2025)
Next-Embedding Prediction Makes Strong Vision Learners
por: Xu, Sihan, et al.
Publicado: (2025)
por: Xu, Sihan, et al.
Publicado: (2025)
Exploring Vision-Language Models for Open-Vocabulary Zero-Shot Action Segmentation
por: Unmesh, Asim, et al.
Publicado: (2026)
por: Unmesh, Asim, et al.
Publicado: (2026)
Test-Time Adaptation of Vision-Language Models for Open-Vocabulary Semantic Segmentation
por: Noori, Mehrdad, et al.
Publicado: (2025)
por: Noori, Mehrdad, et al.
Publicado: (2025)
Adapting Vision-Language Model with Fine-grained Semantics for Open-Vocabulary Segmentation
por: Chng, Yong Xien, et al.
Publicado: (2024)
por: Chng, Yong Xien, et al.
Publicado: (2024)
REALM: An MLLM-Agent Framework for Open World 3D Reasoning Segmentation and Editing on Gaussian Splatting
por: Shi, Changyue, et al.
Publicado: (2025)
por: Shi, Changyue, et al.
Publicado: (2025)
Test-time Contrastive Concepts for Open-world Semantic Segmentation with Vision-Language Models
por: Wysoczańska, Monika, et al.
Publicado: (2024)
por: Wysoczańska, Monika, et al.
Publicado: (2024)
Open-RGBT: Open-vocabulary RGB-T Zero-shot Semantic Segmentation in Open-world Environments
por: Yu, Meng, et al.
Publicado: (2024)
por: Yu, Meng, et al.
Publicado: (2024)
Segment then Splat: Unified 3D Open-Vocabulary Segmentation via Gaussian Splatting
por: Lu, Yiren, et al.
Publicado: (2025)
por: Lu, Yiren, et al.
Publicado: (2025)
Leveraging Vision-Language Models for Open-Vocabulary Instance Segmentation and Tracking
por: Pätzold, Bastian, et al.
Publicado: (2025)
por: Pätzold, Bastian, et al.
Publicado: (2025)
Unsupervised Domain Adaption Harnessing Vision-Language Pre-training
por: Zhou, Wenlve, et al.
Publicado: (2024)
por: Zhou, Wenlve, et al.
Publicado: (2024)
Open Panoramic Segmentation
por: Zheng, Junwei, et al.
Publicado: (2024)
por: Zheng, Junwei, et al.
Publicado: (2024)
Conformal Semantic Image Segmentation: Post-hoc Quantification of Predictive Uncertainty
por: Mossina, Luca, et al.
Publicado: (2024)
por: Mossina, Luca, et al.
Publicado: (2024)
Pose-Aware Self-Supervised Learning with Viewpoint Trajectory Regularization
por: Wang, Jiayun, et al.
Publicado: (2024)
por: Wang, Jiayun, et al.
Publicado: (2024)
Ejemplares similares
-
Open Ad-hoc Categorization with Contextualized Feature Learning
por: Wang, Zilin, et al.
Publicado: (2025) -
Free-Grained Hierarchical Visual Recognition
por: Park, Seulki, et al.
Publicado: (2025) -
Normalize Filters! Classical Wisdom for Deep Vision
por: Perez, Gustavo, et al.
Publicado: (2025) -
Harnessing Vision Foundation Models for High-Performance, Training-Free Open Vocabulary Segmentation
por: Shi, Yuheng, et al.
Publicado: (2024) -
Interpretable Embedding for Ad-hoc Video Search
por: Wu, Jiaxin, et al.
Publicado: (2024)