Salvato in:
| Autori principali: | Yin, Xiaoran, Luo, Xu, Wu, Hao, Gao, Lianli, Song, Jingkuan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2505.16422 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Debiased Orthogonal Boundary-Driven Efficient Noise Mitigation
di: Li, Hao, et al.
Pubblicazione: (2024)
di: Li, Hao, et al.
Pubblicazione: (2024)
Scale-Aware Pre-Training for Human-Centric Visual Perception: Enabling Lightweight and Generalizable Models
di: Wang, Xuanhan, et al.
Pubblicazione: (2025)
di: Wang, Xuanhan, et al.
Pubblicazione: (2025)
CoIN: A Benchmark of Continual Instruction tuNing for Multimodel Large Language Model
di: Chen, Cheng, et al.
Pubblicazione: (2024)
di: Chen, Cheng, et al.
Pubblicazione: (2024)
Benchmarking Few-shot Transferability of Pre-trained Models with Improved Evaluation Protocols
di: Luo, Xu, et al.
Pubblicazione: (2026)
di: Luo, Xu, et al.
Pubblicazione: (2026)
F3-Pruning: A Training-Free and Generalized Pruning Strategy towards Faster and Finer Text-to-Video Synthesis
di: Su, Sitong, et al.
Pubblicazione: (2023)
di: Su, Sitong, et al.
Pubblicazione: (2023)
FlexAC: Towards Flexible Control of Associative Reasoning in Multimodal Large Language Models
di: Yuan, Shengming, et al.
Pubblicazione: (2025)
di: Yuan, Shengming, et al.
Pubblicazione: (2025)
Training-Free Semantic Video Composition via Pre-trained Diffusion Model
di: Guo, Jiaqi, et al.
Pubblicazione: (2024)
di: Guo, Jiaqi, et al.
Pubblicazione: (2024)
Prototype-based Aleatoric Uncertainty Quantification for Cross-modal Retrieval
di: Li, Hao, et al.
Pubblicazione: (2023)
di: Li, Hao, et al.
Pubblicazione: (2023)
CFReID: Continual Few-shot Person Re-Identification
di: Ni, Hao, et al.
Pubblicazione: (2025)
di: Ni, Hao, et al.
Pubblicazione: (2025)
AICL: Action In-Context Learning for Video Diffusion Model
di: Liu, Jianzhi, et al.
Pubblicazione: (2024)
di: Liu, Jianzhi, et al.
Pubblicazione: (2024)
DePT: Decoupled Prompt Tuning
di: Zhang, Ji, et al.
Pubblicazione: (2023)
di: Zhang, Ji, et al.
Pubblicazione: (2023)
Alleviating Hallucinations in Large Vision-Language Models through Hallucination-Induced Optimization
di: Lyu, Xinyu, et al.
Pubblicazione: (2024)
di: Lyu, Xinyu, et al.
Pubblicazione: (2024)
From Channel Bias to Feature Redundancy: Uncovering the "Less is More" Principle in Few-Shot Learning
di: Zhang, Ji, et al.
Pubblicazione: (2023)
di: Zhang, Ji, et al.
Pubblicazione: (2023)
A Closer Look at Conditional Prompt Tuning for Vision-Language Models
di: Zhang, Ji, et al.
Pubblicazione: (2025)
di: Zhang, Ji, et al.
Pubblicazione: (2025)
Skip Tuning: Pre-trained Vision-Language Models are Effective and Efficient Adapters Themselves
di: Wu, Shihan, et al.
Pubblicazione: (2024)
di: Wu, Shihan, et al.
Pubblicazione: (2024)
Shortcut Learning in Generalist Robot Policies: The Role of Dataset Diversity and Fragmentation
di: Xing, Youguang, et al.
Pubblicazione: (2025)
di: Xing, Youguang, et al.
Pubblicazione: (2025)
Dynamic Pattern Alignment Learning for Pretraining Lightweight Human-Centric Vision Models
di: Wang, Xuanhan, et al.
Pubblicazione: (2025)
di: Wang, Xuanhan, et al.
Pubblicazione: (2025)
Learning Generalizable and Efficient Image Watermarking via Hierarchical Two-Stage Optimization
di: Liu, Ke, et al.
Pubblicazione: (2025)
di: Liu, Ke, et al.
Pubblicazione: (2025)
Attention Hijackers: Detect and Disentangle Attention Hijacking in LVLMs for Hallucination Mitigation
di: Chen, Beitao, et al.
Pubblicazione: (2025)
di: Chen, Beitao, et al.
Pubblicazione: (2025)
AstraNav-World: World Model for Foresight Control and Consistency
di: Chen, Jintao, et al.
Pubblicazione: (2025)
di: Chen, Jintao, et al.
Pubblicazione: (2025)
Any Target Can be Offense: Adversarial Example Generation via Generalized Latent Infection
di: Sun, Youheng, et al.
Pubblicazione: (2024)
di: Sun, Youheng, et al.
Pubblicazione: (2024)
Reliable Few-shot Learning under Dual Noises
di: Zhang, Ji, et al.
Pubblicazione: (2025)
di: Zhang, Ji, et al.
Pubblicazione: (2025)
SafePTR: Token-Level Jailbreak Defense in Multimodal LLMs via Prune-then-Restore Mechanism
di: Chen, Beitao, et al.
Pubblicazione: (2025)
di: Chen, Beitao, et al.
Pubblicazione: (2025)
SeMv-3D: Towards Concurrency of Semantic and Multi-view Consistency in General Text-to-3D Generation
di: Cai, Xiao, et al.
Pubblicazione: (2024)
di: Cai, Xiao, et al.
Pubblicazione: (2024)
TIMI: Training-Free Image-to-3D Multi-Instance Generation with Spatial Fidelity
di: Cai, Xiao, et al.
Pubblicazione: (2026)
di: Cai, Xiao, et al.
Pubblicazione: (2026)
Text-Video Retrieval with Global-Local Semantic Consistent Learning
di: Zhang, Haonan, et al.
Pubblicazione: (2024)
di: Zhang, Haonan, et al.
Pubblicazione: (2024)
From One-to-One to Many-to-Many: Dynamic Cross-Layer Injection for Deep Vision-Language Fusion
di: Chen, Cheng, et al.
Pubblicazione: (2026)
di: Chen, Cheng, et al.
Pubblicazione: (2026)
Practical No-box Adversarial Attacks with Training-free Hybrid Image Transformation
di: Zhang, Qilong, et al.
Pubblicazione: (2022)
di: Zhang, Qilong, et al.
Pubblicazione: (2022)
ALF: Adaptive Label Finetuning for Scene Graph Generation
di: Chen, Qishen, et al.
Pubblicazione: (2023)
di: Chen, Qishen, et al.
Pubblicazione: (2023)
Towards Generalized and Training-Free Text-Guided Semantic Manipulation
di: Hong, Yu, et al.
Pubblicazione: (2025)
di: Hong, Yu, et al.
Pubblicazione: (2025)
ProS: Prompting-to-simulate Generalized knowledge for Universal Cross-Domain Retrieval
di: Fang, Kaipeng, et al.
Pubblicazione: (2023)
di: Fang, Kaipeng, et al.
Pubblicazione: (2023)
Structure-aware Prompt Adaptation from Seen to Unseen for Open-Vocabulary Compositional Zero-Shot Learning
di: Duan, Yihang, et al.
Pubblicazione: (2026)
di: Duan, Yihang, et al.
Pubblicazione: (2026)
Informative Scene Graph Generation via Debiasing
di: Gao, Lianli, et al.
Pubblicazione: (2023)
di: Gao, Lianli, et al.
Pubblicazione: (2023)
MiVLA: Towards Generalizable Vision-Language-Action Model with Human-Robot Mutual Imitation Pre-training
di: Yin, Zhenhan, et al.
Pubblicazione: (2025)
di: Yin, Zhenhan, et al.
Pubblicazione: (2025)
Reversible Inversion for Training-Free Exemplar-guided Image Editing
di: Li, Yuke, et al.
Pubblicazione: (2025)
di: Li, Yuke, et al.
Pubblicazione: (2025)
Towards Unified World Models for Visual Navigation via Memory-Augmented Planning and Foresight
di: Dong, Yifei, et al.
Pubblicazione: (2025)
di: Dong, Yifei, et al.
Pubblicazione: (2025)
GT23D-Bench: A Comprehensive General Text-to-3D Generation Benchmark
di: Cai, Xiao, et al.
Pubblicazione: (2024)
di: Cai, Xiao, et al.
Pubblicazione: (2024)
Thinking Ahead: Foresight Intelligence in MLLMs and World Models
di: Gong, Zhantao, et al.
Pubblicazione: (2025)
di: Gong, Zhantao, et al.
Pubblicazione: (2025)
RoScenes: A Large-scale Multi-view 3D Dataset for Roadside Perception
di: Zhu, Xiaosu, et al.
Pubblicazione: (2024)
di: Zhu, Xiaosu, et al.
Pubblicazione: (2024)
See Tomorrow, Act Today: Foresight-Driven Autonomous Driving
di: Zhang, Bozhou, et al.
Pubblicazione: (2026)
di: Zhang, Bozhou, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Debiased Orthogonal Boundary-Driven Efficient Noise Mitigation
di: Li, Hao, et al.
Pubblicazione: (2024) -
Scale-Aware Pre-Training for Human-Centric Visual Perception: Enabling Lightweight and Generalizable Models
di: Wang, Xuanhan, et al.
Pubblicazione: (2025) -
CoIN: A Benchmark of Continual Instruction tuNing for Multimodel Large Language Model
di: Chen, Cheng, et al.
Pubblicazione: (2024) -
Benchmarking Few-shot Transferability of Pre-trained Models with Improved Evaluation Protocols
di: Luo, Xu, et al.
Pubblicazione: (2026) -
F3-Pruning: A Training-Free and Generalized Pruning Strategy towards Faster and Finer Text-to-Video Synthesis
di: Su, Sitong, et al.
Pubblicazione: (2023)