AlignZeg: Mitigating Objective Misalignment for Zero-shot Semantic Segmentation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ge, Jiannan, Xie, Lingxi, Xie, Hongtao, Li, Pandeng, Zhang, Xiaopeng, Zhang, Yongdong, Tian, Qi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding Capability
von: Wang, Jiankang, et al.
Veröffentlicht: (2025)
von: Wang, Jiankang, et al.
Veröffentlicht: (2025)
Segment Any 3D Gaussians
von: Cen, Jiazhong, et al.
Veröffentlicht: (2023)
von: Cen, Jiazhong, et al.
Veröffentlicht: (2023)
GaussianEditor: Editing 3D Gaussians Delicately with Text Instructions
von: Wang, Junjie, et al.
Veröffentlicht: (2023)
von: Wang, Junjie, et al.
Veröffentlicht: (2023)
Segment Anything in 3D with Radiance Fields
von: Cen, Jiazhong, et al.
Veröffentlicht: (2023)
von: Cen, Jiazhong, et al.
Veröffentlicht: (2023)
SAM-CP: Marrying SAM with Composable Prompts for Versatile Segmentation
von: Chen, Pengfei, et al.
Veröffentlicht: (2024)
von: Chen, Pengfei, et al.
Veröffentlicht: (2024)
Tackling View-Dependent Semantics in 3D Language Gaussian Splatting
von: Cen, Jiazhong, et al.
Veröffentlicht: (2025)
von: Cen, Jiazhong, et al.
Veröffentlicht: (2025)
Generalizable Semantic Vision Query Generation for Zero-shot Panoptic and Semantic Segmentation
von: Chen, Jialei, et al.
Veröffentlicht: (2024)
von: Chen, Jialei, et al.
Veröffentlicht: (2024)
Mixpert: Mitigating Multimodal Learning Conflicts with Efficient Mixture-of-Vision-Experts
von: He, Xin, et al.
Veröffentlicht: (2025)
von: He, Xin, et al.
Veröffentlicht: (2025)
Reasoning to Align: Implicit Reasoning in Diffusion Transformers for Video Editing
von: Li, Yan, et al.
Veröffentlicht: (2026)
von: Li, Yan, et al.
Veröffentlicht: (2026)
Split Matching for Inductive Zero-shot Semantic Segmentation
von: Chen, Jialei, et al.
Veröffentlicht: (2025)
von: Chen, Jialei, et al.
Veröffentlicht: (2025)
Cascade-Zero123: One Image to Highly Consistent 3D with Self-Prompted Nearby Views
von: Chen, Yabo, et al.
Veröffentlicht: (2023)
von: Chen, Yabo, et al.
Veröffentlicht: (2023)
DiffAM: Diffusion-based Adversarial Makeup Transfer for Facial Privacy Protection
von: Sun, Yuhao, et al.
Veröffentlicht: (2024)
von: Sun, Yuhao, et al.
Veröffentlicht: (2024)
Parameter Efficient Fine-tuning via Cross Block Orchestration for Segment Anything Model
von: Peng, Zelin, et al.
Veröffentlicht: (2023)
von: Peng, Zelin, et al.
Veröffentlicht: (2023)
GaussianObject: High-Quality 3D Object Reconstruction from Four Views with Gaussian Splatting
von: Yang, Chen, et al.
Veröffentlicht: (2024)
von: Yang, Chen, et al.
Veröffentlicht: (2024)
Discover, Segment, and Select: A Progressive Mechanism for Zero-shot Camouflaged Object Segmentation
von: Yang, Yilong, et al.
Veröffentlicht: (2026)
von: Yang, Yilong, et al.
Veröffentlicht: (2026)
Incorporating Visual Experts to Resolve the Information Loss in Multimodal Large Language Models
von: He, Xin, et al.
Veröffentlicht: (2024)
von: He, Xin, et al.
Veröffentlicht: (2024)
Boosting Semi-Supervised Scene Text Recognition via Viewing and Summarizing
von: Qu, Yadong, et al.
Veröffentlicht: (2024)
von: Qu, Yadong, et al.
Veröffentlicht: (2024)
Bridging Granularity Gaps: Hierarchical Semantic Learning for Cross-domain Few-shot Segmentation
von: Sun, Sujun, et al.
Veröffentlicht: (2025)
von: Sun, Sujun, et al.
Veröffentlicht: (2025)
Exploring Reliable Matching with Phase Enhancement for Night-time Semantic Segmentation
von: Pan, Yuwen, et al.
Veröffentlicht: (2024)
von: Pan, Yuwen, et al.
Veröffentlicht: (2024)
Make It Up: Fake Images, Real Gains in Generalized Few-shot Semantic Segmentation
von: Xie, Guohuan, et al.
Veröffentlicht: (2026)
von: Xie, Guohuan, et al.
Veröffentlicht: (2026)
Zero-shot Composed Text-Image Retrieval
von: Liu, Yikun, et al.
Veröffentlicht: (2023)
von: Liu, Yikun, et al.
Veröffentlicht: (2023)
Partial CLIP is Enough: Chimera-Seg for Zero-shot Semantic Segmentation
von: Chen, Jialei, et al.
Veröffentlicht: (2025)
von: Chen, Jialei, et al.
Veröffentlicht: (2025)
Spectral Prompt Tuning:Unveiling Unseen Classes for Zero-Shot Semantic Segmentation
von: Xu, Wenhao, et al.
Veröffentlicht: (2023)
von: Xu, Wenhao, et al.
Veröffentlicht: (2023)
Bridge the Gap Between Visual and Linguistic Comprehension for Generalized Zero-shot Semantic Segmentation
von: Guo, Xiaoqing, et al.
Veröffentlicht: (2025)
von: Guo, Xiaoqing, et al.
Veröffentlicht: (2025)
LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding
von: Qiu, Jihao, et al.
Veröffentlicht: (2026)
von: Qiu, Jihao, et al.
Veröffentlicht: (2026)
DEADiff: An Efficient Stylization Diffusion Model with Disentangled Representations
von: Qi, Tianhao, et al.
Veröffentlicht: (2024)
von: Qi, Tianhao, et al.
Veröffentlicht: (2024)
GaussianDreamer: Fast Generation from Text to 3D Gaussians by Bridging 2D and 3D Diffusion Models
von: Yi, Taoran, et al.
Veröffentlicht: (2023)
von: Yi, Taoran, et al.
Veröffentlicht: (2023)
4D Gaussian Splatting for Real-Time Dynamic Scene Rendering
von: Wu, Guanjun, et al.
Veröffentlicht: (2023)
von: Wu, Guanjun, et al.
Veröffentlicht: (2023)
LiftImage3D: Lifting Any Single Image to 3D Gaussians with Video Generation Priors
von: Chen, Yabo, et al.
Veröffentlicht: (2024)
von: Chen, Yabo, et al.
Veröffentlicht: (2024)
Semantic Segmentation of Transparent and Opaque Drinking Glasses with the Help of Zero-shot Learning
von: Blänsdorf, Annalena, et al.
Veröffentlicht: (2025)
von: Blänsdorf, Annalena, et al.
Veröffentlicht: (2025)
IGD: Instructional Graphic Design with Multimodal Layer Generation
von: Qu, Yadong, et al.
Veröffentlicht: (2025)
von: Qu, Yadong, et al.
Veröffentlicht: (2025)
CLIP Is Also a Good Teacher: A New Learning Framework for Inductive Zero-shot Semantic Segmentation
von: Chen, Jialei, et al.
Veröffentlicht: (2023)
von: Chen, Jialei, et al.
Veröffentlicht: (2023)
MeaCap: Memory-Augmented Zero-shot Image Captioning
von: Zeng, Zequn, et al.
Veröffentlicht: (2024)
von: Zeng, Zequn, et al.
Veröffentlicht: (2024)
SAQ-SAM: Semantically-Aligned Quantization for Segment Anything Model
von: Zhang, Jing, et al.
Veröffentlicht: (2025)
von: Zhang, Jing, et al.
Veröffentlicht: (2025)
Hallucination Mitigation Prompts Long-term Video Understanding
von: Sun, Yiwei, et al.
Veröffentlicht: (2024)
von: Sun, Yiwei, et al.
Veröffentlicht: (2024)
Exploring Semantic Consistency and Style Diversity for Domain Generalized Semantic Segmentation
von: Niu, Hongwei, et al.
Veröffentlicht: (2024)
von: Niu, Hongwei, et al.
Veröffentlicht: (2024)
ShowTable: Unlocking Creative Table Visualization with Collaborative Reflection and Refinement
von: Liu, Zhihang, et al.
Veröffentlicht: (2025)
von: Liu, Zhihang, et al.
Veröffentlicht: (2025)
EMMA: Efficient Multimodal Understanding, Generation, and Editing with a Unified Architecture
von: He, Xin, et al.
Veröffentlicht: (2025)
von: He, Xin, et al.
Veröffentlicht: (2025)
GaussianDreamerPro: Text to Manipulable 3D Gaussians with Highly Enhanced Quality
von: Yi, Taoran, et al.
Veröffentlicht: (2024)
von: Yi, Taoran, et al.
Veröffentlicht: (2024)
Open-RGBT: Open-vocabulary RGB-T Zero-shot Semantic Segmentation in Open-world Environments
von: Yu, Meng, et al.
Veröffentlicht: (2024)
von: Yu, Meng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
SpaceVLLM: Endowing Multimodal Large Language Model with Spatio-Temporal Video Grounding Capability
von: Wang, Jiankang, et al.
Veröffentlicht: (2025) -
Segment Any 3D Gaussians
von: Cen, Jiazhong, et al.
Veröffentlicht: (2023) -
GaussianEditor: Editing 3D Gaussians Delicately with Text Instructions
von: Wang, Junjie, et al.
Veröffentlicht: (2023) -
Segment Anything in 3D with Radiance Fields
von: Cen, Jiazhong, et al.
Veröffentlicht: (2023) -
SAM-CP: Marrying SAM with Composable Prompts for Versatile Segmentation
von: Chen, Pengfei, et al.
Veröffentlicht: (2024)