Gespeichert in:
| Hauptverfasser: | Chen, Cheng, Guo, Yuyu, Zeng, Pengpeng, Song, Jingkuan, Di, Peng, Yu, Hang, Gao, Lianli |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2601.10710 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Skip Tuning: Pre-trained Vision-Language Models are Effective and Efficient Adapters Themselves
von: Wu, Shihan, et al.
Veröffentlicht: (2024)
von: Wu, Shihan, et al.
Veröffentlicht: (2024)
CFReID: Continual Few-shot Person Re-Identification
von: Ni, Hao, et al.
Veröffentlicht: (2025)
von: Ni, Hao, et al.
Veröffentlicht: (2025)
ProS: Prompting-to-simulate Generalized knowledge for Universal Cross-Domain Retrieval
von: Fang, Kaipeng, et al.
Veröffentlicht: (2023)
von: Fang, Kaipeng, et al.
Veröffentlicht: (2023)
SeMv-3D: Towards Concurrency of Semantic and Multi-view Consistency in General Text-to-3D Generation
von: Cai, Xiao, et al.
Veröffentlicht: (2024)
von: Cai, Xiao, et al.
Veröffentlicht: (2024)
TIMI: Training-Free Image-to-3D Multi-Instance Generation with Spatial Fidelity
von: Cai, Xiao, et al.
Veröffentlicht: (2026)
von: Cai, Xiao, et al.
Veröffentlicht: (2026)
Towards Generalized and Training-Free Text-Guided Semantic Manipulation
von: Hong, Yu, et al.
Veröffentlicht: (2025)
von: Hong, Yu, et al.
Veröffentlicht: (2025)
Text-Video Retrieval with Global-Local Semantic Consistent Learning
von: Zhang, Haonan, et al.
Veröffentlicht: (2024)
von: Zhang, Haonan, et al.
Veröffentlicht: (2024)
A Survey on Efficient Vision-Language-Action Models
von: Yu, Zhaoshu, et al.
Veröffentlicht: (2025)
von: Yu, Zhaoshu, et al.
Veröffentlicht: (2025)
Alleviating Hallucinations in Large Vision-Language Models through Hallucination-Induced Optimization
von: Lyu, Xinyu, et al.
Veröffentlicht: (2024)
von: Lyu, Xinyu, et al.
Veröffentlicht: (2024)
Dynamic Pattern Alignment Learning for Pretraining Lightweight Human-Centric Vision Models
von: Wang, Xuanhan, et al.
Veröffentlicht: (2025)
von: Wang, Xuanhan, et al.
Veröffentlicht: (2025)
Structure-aware Prompt Adaptation from Seen to Unseen for Open-Vocabulary Compositional Zero-Shot Learning
von: Duan, Yihang, et al.
Veröffentlicht: (2026)
von: Duan, Yihang, et al.
Veröffentlicht: (2026)
A Closer Look at Conditional Prompt Tuning for Vision-Language Models
von: Zhang, Ji, et al.
Veröffentlicht: (2025)
von: Zhang, Ji, et al.
Veröffentlicht: (2025)
Prototype-based Aleatoric Uncertainty Quantification for Cross-modal Retrieval
von: Li, Hao, et al.
Veröffentlicht: (2023)
von: Li, Hao, et al.
Veröffentlicht: (2023)
Scale-Aware Pre-Training for Human-Centric Visual Perception: Enabling Lightweight and Generalizable Models
von: Wang, Xuanhan, et al.
Veröffentlicht: (2025)
von: Wang, Xuanhan, et al.
Veröffentlicht: (2025)
F3-Pruning: A Training-Free and Generalized Pruning Strategy towards Faster and Finer Text-to-Video Synthesis
von: Su, Sitong, et al.
Veröffentlicht: (2023)
von: Su, Sitong, et al.
Veröffentlicht: (2023)
Reversible Inversion for Training-Free Exemplar-guided Image Editing
von: Li, Yuke, et al.
Veröffentlicht: (2025)
von: Li, Yuke, et al.
Veröffentlicht: (2025)
Training-Free Semantic Video Composition via Pre-trained Diffusion Model
von: Guo, Jiaqi, et al.
Veröffentlicht: (2024)
von: Guo, Jiaqi, et al.
Veröffentlicht: (2024)
CoIN: A Benchmark of Continual Instruction tuNing for Multimodel Large Language Model
von: Chen, Cheng, et al.
Veröffentlicht: (2024)
von: Chen, Cheng, et al.
Veröffentlicht: (2024)
Informative Scene Graph Generation via Debiasing
von: Gao, Lianli, et al.
Veröffentlicht: (2023)
von: Gao, Lianli, et al.
Veröffentlicht: (2023)
GT23D-Bench: A Comprehensive General Text-to-3D Generation Benchmark
von: Cai, Xiao, et al.
Veröffentlicht: (2024)
von: Cai, Xiao, et al.
Veröffentlicht: (2024)
Debiased Orthogonal Boundary-Driven Efficient Noise Mitigation
von: Li, Hao, et al.
Veröffentlicht: (2024)
von: Li, Hao, et al.
Veröffentlicht: (2024)
FlexAC: Towards Flexible Control of Associative Reasoning in Multimodal Large Language Models
von: Yuan, Shengming, et al.
Veröffentlicht: (2025)
von: Yuan, Shengming, et al.
Veröffentlicht: (2025)
Unlocking Smarter Device Control: Foresighted Planning with a World Model-Driven Code Execution Approach
von: Yin, Xiaoran, et al.
Veröffentlicht: (2025)
von: Yin, Xiaoran, et al.
Veröffentlicht: (2025)
From Channel Bias to Feature Redundancy: Uncovering the "Less is More" Principle in Few-Shot Learning
von: Zhang, Ji, et al.
Veröffentlicht: (2023)
von: Zhang, Ji, et al.
Veröffentlicht: (2023)
Multimodal Adversarial Defense for Vision-Language Models by Leveraging One-To-Many Relationships
von: Waseda, Futa, et al.
Veröffentlicht: (2024)
von: Waseda, Futa, et al.
Veröffentlicht: (2024)
Learning Generalizable and Efficient Image Watermarking via Hierarchical Two-Stage Optimization
von: Liu, Ke, et al.
Veröffentlicht: (2025)
von: Liu, Ke, et al.
Veröffentlicht: (2025)
Any Target Can be Offense: Adversarial Example Generation via Generalized Latent Infection
von: Sun, Youheng, et al.
Veröffentlicht: (2024)
von: Sun, Youheng, et al.
Veröffentlicht: (2024)
Pairing Regularization for Mitigating Many-to-One Collapse in GANs
von: Lin, Kuan-Yu, et al.
Veröffentlicht: (2026)
von: Lin, Kuan-Yu, et al.
Veröffentlicht: (2026)
DePT: Decoupled Prompt Tuning
von: Zhang, Ji, et al.
Veröffentlicht: (2023)
von: Zhang, Ji, et al.
Veröffentlicht: (2023)
AICL: Action In-Context Learning for Video Diffusion Model
von: Liu, Jianzhi, et al.
Veröffentlicht: (2024)
von: Liu, Jianzhi, et al.
Veröffentlicht: (2024)
Attention Hijackers: Detect and Disentangle Attention Hijacking in LVLMs for Hallucination Mitigation
von: Chen, Beitao, et al.
Veröffentlicht: (2025)
von: Chen, Beitao, et al.
Veröffentlicht: (2025)
Memories are One-to-Many Mapping Alleviators in Talking Face Generation
von: Tang, Anni, et al.
Veröffentlicht: (2022)
von: Tang, Anni, et al.
Veröffentlicht: (2022)
Improving Deep Generative Models on Many-To-One Image-to-Image Translation
von: Saxena, Sagar, et al.
Veröffentlicht: (2024)
von: Saxena, Sagar, et al.
Veröffentlicht: (2024)
Reliable Few-shot Learning under Dual Noises
von: Zhang, Ji, et al.
Veröffentlicht: (2025)
von: Zhang, Ji, et al.
Veröffentlicht: (2025)
SafePTR: Token-Level Jailbreak Defense in Multimodal LLMs via Prune-then-Restore Mechanism
von: Chen, Beitao, et al.
Veröffentlicht: (2025)
von: Chen, Beitao, et al.
Veröffentlicht: (2025)
Benchmarking Few-shot Transferability of Pre-trained Models with Improved Evaluation Protocols
von: Luo, Xu, et al.
Veröffentlicht: (2026)
von: Luo, Xu, et al.
Veröffentlicht: (2026)
Video Individual Counting With Implicit One-to-Many Matching
von: Zhu, Xuhui, et al.
Veröffentlicht: (2025)
von: Zhu, Xuhui, et al.
Veröffentlicht: (2025)
Impact of Sunglasses on One-to-Many Facial Identification Accuracy
von: Tian, Sicong, et al.
Veröffentlicht: (2024)
von: Tian, Sicong, et al.
Veröffentlicht: (2024)
MotionDreamer: One-to-Many Motion Synthesis with Localized Generative Masked Transformer
von: Wang, Yilin, et al.
Veröffentlicht: (2025)
von: Wang, Yilin, et al.
Veröffentlicht: (2025)
SOLAR: Self-supervised Joint Learning for Symmetric Multimodal Retrieval
von: Yang, Wenjie, et al.
Veröffentlicht: (2026)
von: Yang, Wenjie, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Skip Tuning: Pre-trained Vision-Language Models are Effective and Efficient Adapters Themselves
von: Wu, Shihan, et al.
Veröffentlicht: (2024) -
CFReID: Continual Few-shot Person Re-Identification
von: Ni, Hao, et al.
Veröffentlicht: (2025) -
ProS: Prompting-to-simulate Generalized knowledge for Universal Cross-Domain Retrieval
von: Fang, Kaipeng, et al.
Veröffentlicht: (2023) -
SeMv-3D: Towards Concurrency of Semantic and Multi-view Consistency in General Text-to-3D Generation
von: Cai, Xiao, et al.
Veröffentlicht: (2024) -
TIMI: Training-Free Image-to-3D Multi-Instance Generation with Spatial Fidelity
von: Cai, Xiao, et al.
Veröffentlicht: (2026)