Leveraging MLLM Embeddings and Attribute Smoothing for Compositional Zero-Shot Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Yan, Xudong, Feng, Songhe, Zhang, Yang, Yang, Jian, Lin, Yueguan, Fei, Haojun |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
TOMCAT: Test-time Comprehensive Knowledge Accumulation for Compositional Zero-Shot Learning
by: Yan, Xudong, et al.
Published: (2025)
by: Yan, Xudong, et al.
Published: (2025)
Hybrid Discriminative Attribute-Object Embedding Network for Compositional Zero-Shot Learning
by: Liu, Yang, et al.
Published: (2024)
by: Liu, Yang, et al.
Published: (2024)
WARM-CAT: Warm-Started Test-Time Comprehensive Knowledge Accumulation for Compositional Zero-Shot Learning
by: Yan, Xudong, et al.
Published: (2026)
by: Yan, Xudong, et al.
Published: (2026)
No Need For Real Anomaly: MLLM Empowered Zero-Shot Video Anomaly Detection
by: Dai, Zunkai, et al.
Published: (2026)
by: Dai, Zunkai, et al.
Published: (2026)
UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning
by: Gu, Tiancheng, et al.
Published: (2025)
by: Gu, Tiancheng, et al.
Published: (2025)
Shot Sequence Ordering for Video Editing: Benchmarks, Metrics, and Cinematology-Inspired Computing Methods
by: Li, Yuzhi, et al.
Published: (2025)
by: Li, Yuzhi, et al.
Published: (2025)
Learning Primitive Relations for Compositional Zero-Shot Learning
by: Lee, Insu, et al.
Published: (2025)
by: Lee, Insu, et al.
Published: (2025)
Prompt-Based Continual Compositional Zero-Shot Learning
by: Maryam, Sauda, et al.
Published: (2025)
by: Maryam, Sauda, et al.
Published: (2025)
SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM
by: Yu, An, et al.
Published: (2025)
by: Yu, An, et al.
Published: (2025)
PostEdit: Posterior Sampling for Efficient Zero-Shot Image Editing
by: Tian, Feng, et al.
Published: (2024)
by: Tian, Feng, et al.
Published: (2024)
Decoupling Endpoint and Semantic Transition Learning for Zero-Shot Composed Image Retrieval
by: Liu, Mingyu, et al.
Published: (2026)
by: Liu, Mingyu, et al.
Published: (2026)
Self-Attention with State-Object Weighted Combination for Compositional Zero Shot Learning
by: Chang, Cheng-Hong, et al.
Published: (2025)
by: Chang, Cheng-Hong, et al.
Published: (2025)
LIME: Less Is More for MLLM Evaluation
by: Zhu, King, et al.
Published: (2024)
by: Zhu, King, et al.
Published: (2024)
Pseudo-label Based Domain Adaptation for Zero-Shot Text Steganalysis
by: Luo, Yufei, et al.
Published: (2024)
by: Luo, Yufei, et al.
Published: (2024)
Bootstrapping Diffusion: Diffusion Model Training Leveraging Partial and Corrupted Data
by: Ma, Xudong
Published: (2025)
by: Ma, Xudong
Published: (2025)
Unleashing the Potential of Vision-Language Pre-Training for 3D Zero-Shot Lesion Segmentation via Mask-Attribute Alignment
by: Jiang, Yankai, et al.
Published: (2024)
by: Jiang, Yankai, et al.
Published: (2024)
SafeEditor: Unified MLLM for Efficient Post-hoc T2I Safety Editing
by: Zhang, Ruiyang, et al.
Published: (2025)
by: Zhang, Ruiyang, et al.
Published: (2025)
SCOT: Self-Supervised Contrastive Pretraining For Zero-Shot Compositional Retrieval
by: Jawade, Bhavin, et al.
Published: (2025)
by: Jawade, Bhavin, et al.
Published: (2025)
SalientFusion: Context-Aware Compositional Zero-Shot Food Recognition
by: Song, Jiajun, et al.
Published: (2025)
by: Song, Jiajun, et al.
Published: (2025)
VidVec: Unlocking Video MLLM Embeddings for Video-Text Retrieval
by: Tzachor, Issar, et al.
Published: (2026)
by: Tzachor, Issar, et al.
Published: (2026)
SpatialNav: Leveraging Spatial Scene Graphs for Zero-Shot Vision-and-Language Navigation
by: Zhang, Jiwen, et al.
Published: (2026)
by: Zhang, Jiwen, et al.
Published: (2026)
Compositional Few-Shot Class-Incremental Learning
by: Zou, Yixiong, et al.
Published: (2024)
by: Zou, Yixiong, et al.
Published: (2024)
Dog-IQA: Standard-guided Zero-shot MLLM for Mix-grained Image Quality Assessment
by: Liu, Kai, et al.
Published: (2024)
by: Liu, Kai, et al.
Published: (2024)
MAC: A Benchmark for Multiple Attributes Compositional Zero-Shot Learning
by: Xu, Shuo, et al.
Published: (2024)
by: Xu, Shuo, et al.
Published: (2024)
Leveraging Unknown Objects to Construct Labeled-Unlabeled Meta-Relationships for Zero-Shot Object Navigation
by: Zheng, Yanwei, et al.
Published: (2024)
by: Zheng, Yanwei, et al.
Published: (2024)
Zero-Training Task-Specific Model Synthesis for Few-Shot Medical Image Classification
by: Qin, Yao, et al.
Published: (2025)
by: Qin, Yao, et al.
Published: (2025)
Transductive Zero-Shot and Few-Shot CLIP
by: Martin, Ségolène, et al.
Published: (2024)
by: Martin, Ségolène, et al.
Published: (2024)
CLIP Architecture for Abdominal CT Image-Text Alignment and Zero-Shot Learning: Investigating Batch Composition and Data Scaling
by: Shivika, et al.
Published: (2026)
by: Shivika, et al.
Published: (2026)
Sce2DriveX: A Generalized MLLM Framework for Scene-to-Drive Learning
by: Zhao, Rui, et al.
Published: (2025)
by: Zhao, Rui, et al.
Published: (2025)
Revisiting Cross-Attention Mechanisms: Leveraging Beneficial Noise for Domain-Adaptive Learning
by: Zang, Zelin, et al.
Published: (2026)
by: Zang, Zelin, et al.
Published: (2026)
Compositional Attribute Imbalance in Vision Datasets
by: Chen, Jiayi, et al.
Published: (2025)
by: Chen, Jiayi, et al.
Published: (2025)
TinyVLM: Zero-Shot Object Detection on Microcontrollers via Vision-Language Distillation with Matryoshka Embeddings
by: Wilson, Bibin
Published: (2026)
by: Wilson, Bibin
Published: (2026)
AttrSeg: Open-Vocabulary Semantic Segmentation via Attribute Decomposition-Aggregation
by: Ma, Chaofan, et al.
Published: (2023)
by: Ma, Chaofan, et al.
Published: (2023)
MLLM-CL: Continual Learning for Multimodal Large Language Models
by: Zhao, Hongbo, et al.
Published: (2025)
by: Zhao, Hongbo, et al.
Published: (2025)
CompDiff: Hierarchical Compositional Diffusion for Fair and Zero-Shot Intersectional Medical Image Generation
by: Ibrahim, Mahmoud, et al.
Published: (2026)
by: Ibrahim, Mahmoud, et al.
Published: (2026)
Semantic Surgery: Zero-Shot Concept Erasure in Diffusion Models
by: Xiong, Lexiang, et al.
Published: (2025)
by: Xiong, Lexiang, et al.
Published: (2025)
Place-it-R1: Unlocking Environment-aware Reasoning Potential of MLLM for Video Object Insertion
by: Gu, Bohai, et al.
Published: (2026)
by: Gu, Bohai, et al.
Published: (2026)
Nuanced Emotion Recognition Based on a Segment-based MLLM Framework Leveraging Qwen3-Omni for AH Detection
by: Tang, Liang, et al.
Published: (2026)
by: Tang, Liang, et al.
Published: (2026)
Re-purposing SAM into Efficient Visual Projectors for MLLM-Based Referring Image Segmentation
by: Yang, Xiaobo, et al.
Published: (2025)
by: Yang, Xiaobo, et al.
Published: (2025)
Tempo-R0: A Video-MLLM for Temporal Video Grounding through Efficient Temporal Sensing Reinforcement Learning
by: Yue, Feng, et al.
Published: (2025)
by: Yue, Feng, et al.
Published: (2025)
Similar Items
-
TOMCAT: Test-time Comprehensive Knowledge Accumulation for Compositional Zero-Shot Learning
by: Yan, Xudong, et al.
Published: (2025) -
Hybrid Discriminative Attribute-Object Embedding Network for Compositional Zero-Shot Learning
by: Liu, Yang, et al.
Published: (2024) -
WARM-CAT: Warm-Started Test-Time Comprehensive Knowledge Accumulation for Compositional Zero-Shot Learning
by: Yan, Xudong, et al.
Published: (2026) -
No Need For Real Anomaly: MLLM Empowered Zero-Shot Video Anomaly Detection
by: Dai, Zunkai, et al.
Published: (2026) -
UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning
by: Gu, Tiancheng, et al.
Published: (2025)