UniFGVC: Universal Training-Free Few-Shot Fine-Grained Vision Classification via Attribute-Aware Multimodal Retrieval
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Guo, Hongyu, Hao, Xiangzhao, Guo, Jiarui, Guo, Haiyun, Wang, Jinqiao, Chua, Tat-Seng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ReCALL: Recalibrating Capability Degradation for MLLM-based Composed Image Retrieval
von: Yang, Tianyu, et al.
Veröffentlicht: (2026)
von: Yang, Tianyu, et al.
Veröffentlicht: (2026)
WISER: Wider Search, Deeper Thinking, and Adaptive Fusion for Training-Free Zero-Shot Composed Image Retrieval
von: Wang, Tianyue, et al.
Veröffentlicht: (2026)
von: Wang, Tianyue, et al.
Veröffentlicht: (2026)
TRACE: Task-Adaptive Reasoning and Representation Learning for Universal Multimodal Retrieval
von: Hao, Xiangzhao, et al.
Veröffentlicht: (2026)
von: Hao, Xiangzhao, et al.
Veröffentlicht: (2026)
Referring Expression Instance Retrieval and A Strong End-to-End Baseline
von: Hao, Xiangzhao, et al.
Veröffentlicht: (2025)
von: Hao, Xiangzhao, et al.
Veröffentlicht: (2025)
Active Zero: Self-Evolving Vision-Language Models through Active Environment Exploration
von: He, Jinghan, et al.
Veröffentlicht: (2026)
von: He, Jinghan, et al.
Veröffentlicht: (2026)
PLUME: Latent Reasoning Based Universal Multimodal Embedding
von: He, Chenwei, et al.
Veröffentlicht: (2026)
von: He, Chenwei, et al.
Veröffentlicht: (2026)
Cracking the Code of Hallucination in LVLMs with Vision-aware Head Divergence
von: He, Jinghan, et al.
Veröffentlicht: (2024)
von: He, Jinghan, et al.
Veröffentlicht: (2024)
FOCUS: Fine-grained Optimization with Semantic Guided Understanding for Pedestrian Attributes Recognition
von: An, Hongyan, et al.
Veröffentlicht: (2025)
von: An, Hongyan, et al.
Veröffentlicht: (2025)
ST-Prune: Training-Free Spatio-Temporal Token Pruning for Vision-Language Models in Autonomous Driving
von: Sha, Lin, et al.
Veröffentlicht: (2026)
von: Sha, Lin, et al.
Veröffentlicht: (2026)
MLLM-CTBench: A Benchmark for Continual Instruction Tuning with Reasoning Process Diagnosis
von: Guo, Haiyun, et al.
Veröffentlicht: (2025)
von: Guo, Haiyun, et al.
Veröffentlicht: (2025)
Unifying Group-Relative and Self-Distillation Policy Optimization via Sample Routing
von: Li, Gengsheng, et al.
Veröffentlicht: (2026)
von: Li, Gengsheng, et al.
Veröffentlicht: (2026)
Fine-Grained Zero-Shot Learning with Attribute-Centric Representations
von: Chen, Zhi, et al.
Veröffentlicht: (2025)
von: Chen, Zhi, et al.
Veröffentlicht: (2025)
Text and Image Are Mutually Beneficial: Enhancing Training-Free Few-Shot Classification with CLIP
von: Li, Yayuan, et al.
Veröffentlicht: (2024)
von: Li, Yayuan, et al.
Veröffentlicht: (2024)
Fine-Grained Verifiers: Preference Modeling as Next-token Prediction in Vision-Language Alignment
von: Cui, Chenhang, et al.
Veröffentlicht: (2024)
von: Cui, Chenhang, et al.
Veröffentlicht: (2024)
Training-Free Multimodal Large Language Model Orchestration
von: Xie, Tianyu, et al.
Veröffentlicht: (2025)
von: Xie, Tianyu, et al.
Veröffentlicht: (2025)
A Survey on Neural Question Generation: Methods, Applications, and Prospects
von: Guo, Shasha, et al.
Veröffentlicht: (2024)
von: Guo, Shasha, et al.
Veröffentlicht: (2024)
Universal Scene Graph Generation
von: Wu, Shengqiong, et al.
Veröffentlicht: (2025)
von: Wu, Shengqiong, et al.
Veröffentlicht: (2025)
FiLo++: Zero-/Few-Shot Anomaly Detection by Fused Fine-Grained Descriptions and Deformable Localization
von: Gu, Zhaopeng, et al.
Veröffentlicht: (2025)
von: Gu, Zhaopeng, et al.
Veröffentlicht: (2025)
Generative Cross-Modal Retrieval: Memorizing Images in Multimodal Language Models for Retrieval and Beyond
von: Li, Yongqi, et al.
Veröffentlicht: (2024)
von: Li, Yongqi, et al.
Veröffentlicht: (2024)
Transfer Learning and Mixup for Fine-Grained Few-Shot Fungi Classification
von: Tam, Jason Kahei, et al.
Veröffentlicht: (2025)
von: Tam, Jason Kahei, et al.
Veröffentlicht: (2025)
UniDetect: LLM-Driven Universal Fraud Detection across Heterogeneous Blockchains
von: Miao, Shuyi, et al.
Veröffentlicht: (2026)
von: Miao, Shuyi, et al.
Veröffentlicht: (2026)
Compose Your Aesthetics: Empowering Text-to-Image Models with the Principles of Art
von: Jin, Zhe, et al.
Veröffentlicht: (2025)
von: Jin, Zhe, et al.
Veröffentlicht: (2025)
Simple but Effective Raw-Data Level Multimodal Fusion for Composed Image Retrieval
von: Wen, Haokun, et al.
Veröffentlicht: (2024)
von: Wen, Haokun, et al.
Veröffentlicht: (2024)
MMDocBench: Benchmarking Large Vision-Language Models for Fine-Grained Visual Document Understanding
von: Zhu, Fengbin, et al.
Veröffentlicht: (2024)
von: Zhu, Fengbin, et al.
Veröffentlicht: (2024)
Pre-Training and Prompting for Few-Shot Node Classification on Text-Attributed Graphs
von: Zhao, Huanjing, et al.
Veröffentlicht: (2024)
von: Zhao, Huanjing, et al.
Veröffentlicht: (2024)
Zero-1-to-A: Zero-Shot One Image to Animatable Head Avatars Using Video Diffusion
von: Zhou, Zhenglin, et al.
Veröffentlicht: (2025)
von: Zhou, Zhenglin, et al.
Veröffentlicht: (2025)
Exploring Training and Inference Scaling Laws in Generative Retrieval
von: Cai, Hongru, et al.
Veröffentlicht: (2025)
von: Cai, Hongru, et al.
Veröffentlicht: (2025)
Complementary Subspace Low-Rank Adaptation of Vision-Language Models for Few-Shot Classification
von: Wang, Zhongqi, et al.
Veröffentlicht: (2025)
von: Wang, Zhongqi, et al.
Veröffentlicht: (2025)
A Visual RAG Pipeline for Few-Shot Fine-Grained Product Classification
von: Lamm, Bianca, et al.
Veröffentlicht: (2025)
von: Lamm, Bianca, et al.
Veröffentlicht: (2025)
UniVAD: A Training-free Unified Model for Few-shot Visual Anomaly Detection
von: Gu, Zhaopeng, et al.
Veröffentlicht: (2024)
von: Gu, Zhaopeng, et al.
Veröffentlicht: (2024)
Towards Fine-Grained Vision-Language Alignment for Few-Shot Anomaly Detection
von: Fan, Yuanting, et al.
Veröffentlicht: (2025)
von: Fan, Yuanting, et al.
Veröffentlicht: (2025)
Principled Multimodal Representation Learning
von: Liu, Xiaohao, et al.
Veröffentlicht: (2025)
von: Liu, Xiaohao, et al.
Veröffentlicht: (2025)
Continual Multimodal Contrastive Learning
von: Liu, Xiaohao, et al.
Veröffentlicht: (2025)
von: Liu, Xiaohao, et al.
Veröffentlicht: (2025)
Turing Patterns for Multimedia: Reaction-Diffusion Multi-Modal Fusion for Language-Guided Video Moment Retrieval
von: Fang, Xiang, et al.
Veröffentlicht: (2026)
von: Fang, Xiang, et al.
Veröffentlicht: (2026)
Task Attribute Distance for Few-Shot Learning: Theoretical Analysis and Applications
von: Hu, Minyang, et al.
Veröffentlicht: (2024)
von: Hu, Minyang, et al.
Veröffentlicht: (2024)
XNLP: An Interactive Demonstration System for Universal Structured NLP
von: Fei, Hao, et al.
Veröffentlicht: (2023)
von: Fei, Hao, et al.
Veröffentlicht: (2023)
SWIFT: Sliding Window Reconstruction for Few-Shot Training-Free Generated Video Attribution
von: Wang, Chao, et al.
Veröffentlicht: (2026)
von: Wang, Chao, et al.
Veröffentlicht: (2026)
TIGeR: Unifying Text-to-Image Generation and Retrieval with Large Multimodal Models
von: Qu, Leigang, et al.
Veröffentlicht: (2024)
von: Qu, Leigang, et al.
Veröffentlicht: (2024)
CLEAR: Unlocking Generative Potential for Degraded Image Understanding in Unified Multimodal Models
von: Hao, Xiangzhao, et al.
Veröffentlicht: (2026)
von: Hao, Xiangzhao, et al.
Veröffentlicht: (2026)
3D-TAFS: A Training-free Framework for 3D Affordance Segmentation
von: Chu, Meng, et al.
Veröffentlicht: (2024)
von: Chu, Meng, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
ReCALL: Recalibrating Capability Degradation for MLLM-based Composed Image Retrieval
von: Yang, Tianyu, et al.
Veröffentlicht: (2026) -
WISER: Wider Search, Deeper Thinking, and Adaptive Fusion for Training-Free Zero-Shot Composed Image Retrieval
von: Wang, Tianyue, et al.
Veröffentlicht: (2026) -
TRACE: Task-Adaptive Reasoning and Representation Learning for Universal Multimodal Retrieval
von: Hao, Xiangzhao, et al.
Veröffentlicht: (2026) -
Referring Expression Instance Retrieval and A Strong End-to-End Baseline
von: Hao, Xiangzhao, et al.
Veröffentlicht: (2025) -
Active Zero: Self-Evolving Vision-Language Models through Active Environment Exploration
von: He, Jinghan, et al.
Veröffentlicht: (2026)