Saved in:
| Main Authors: | Han, Jiayi, Du, Liang, Wu, Yiwen, Zhou, Xiangguo, Du, Hongwei, Zheng, Weibo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2501.09532 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ESP-Zero: Unsupervised enhancement of zero-shot classification for Extremely Sparse Point cloud
by: Han, Jiayi, et al.
Published: (2024)
by: Han, Jiayi, et al.
Published: (2024)
SLIM: Let LLM Learn More and Forget Less with Soft LoRA and Identity Mixture
by: Han, Jiayi, et al.
Published: (2024)
by: Han, Jiayi, et al.
Published: (2024)
Rethinking Precision of Pseudo Label: Test-Time Adaptation via Complementary Learning
by: Han, Jiayi, et al.
Published: (2023)
by: Han, Jiayi, et al.
Published: (2023)
Efficient Multi-modal Large Language Models via Visual Token Grouping
by: Huang, Minbin, et al.
Published: (2024)
by: Huang, Minbin, et al.
Published: (2024)
AdaGen: Learning Adaptive Policy for Image Synthesis
by: Ni, Zanlin, et al.
Published: (2026)
by: Ni, Zanlin, et al.
Published: (2026)
PAFedFV: Personalized and Asynchronous Federated Learning for Finger Vein Recognition
by: Mu, Hengyu, et al.
Published: (2024)
by: Mu, Hengyu, et al.
Published: (2024)
Rethinking VLM Representation for VLA Initialization
by: Lin, Weifeng, et al.
Published: (2026)
by: Lin, Weifeng, et al.
Published: (2026)
CAD-Judge: Toward Efficient Morphological Grading and Verification for Text-to-CAD Generation
by: Zhou, Zheyuan, et al.
Published: (2025)
by: Zhou, Zheyuan, et al.
Published: (2025)
Rethinking Visual Token Reduction in LVLMs Under Cross-Modal Misalignment
by: Xu, Rui, et al.
Published: (2025)
by: Xu, Rui, et al.
Published: (2025)
LaxMotion: Rethinking Supervision Granularity for 3D Human Motion Generation
by: Liu, Sheng, et al.
Published: (2025)
by: Liu, Sheng, et al.
Published: (2025)
A VLM-based Method for Visual Anomaly Detection in Robotic Scientific Laboratories
by: Lin, Shiwei, et al.
Published: (2025)
by: Lin, Shiwei, et al.
Published: (2025)
RelationVLM: Making Large Vision-Language Models Understand Visual Relations
by: Huang, Zhipeng, et al.
Published: (2024)
by: Huang, Zhipeng, et al.
Published: (2024)
K-Sort Eval: Efficient Preference Evaluation for Visual Generation via Corrected VLM-as-a-Judge
by: Li, Zhikai, et al.
Published: (2026)
by: Li, Zhikai, et al.
Published: (2026)
Evaluating Visual and Cultural Interpretation: The K-Viscuit Benchmark with Human-VLM Collaboration
by: Park, ChaeHun, et al.
Published: (2024)
by: Park, ChaeHun, et al.
Published: (2024)
CogVLM: Visual Expert for Pretrained Language Models
by: Wang, Weihan, et al.
Published: (2023)
by: Wang, Weihan, et al.
Published: (2023)
Weakly-supervised VLM-guided Partial Contrastive Learning for Visual Language Navigation
by: Wang, Ruoyu, et al.
Published: (2025)
by: Wang, Ruoyu, et al.
Published: (2025)
SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
by: Zhang, Yuan, et al.
Published: (2024)
by: Zhang, Yuan, et al.
Published: (2024)
MUSE-VL: Modeling Unified VLM through Semantic Discrete Encoding
by: Xie, Rongchang, et al.
Published: (2024)
by: Xie, Rongchang, et al.
Published: (2024)
SuperDisco: Super-Class Discovery Improves Visual Recognition for the Long-Tail
by: Du, Yingjun, et al.
Published: (2023)
by: Du, Yingjun, et al.
Published: (2023)
CREG: Compass Relational Evidence Graph for Characterizing Directional Structure in VLM Spatial-Reasoning Attribution
by: Tan, Kaizhen, et al.
Published: (2026)
by: Tan, Kaizhen, et al.
Published: (2026)
HERO: Rethinking Visual Token Early Dropping in High-Resolution Large Vision-Language Models
by: Li, Xu, et al.
Published: (2025)
by: Li, Xu, et al.
Published: (2025)
Visual Perception by Large Language Model's Weights
by: Ma, Feipeng, et al.
Published: (2024)
by: Ma, Feipeng, et al.
Published: (2024)
Advancing Reliable Test-Time Adaptation of Vision-Language Models under Visual Variations
by: Liang, Yiwen, et al.
Published: (2025)
by: Liang, Yiwen, et al.
Published: (2025)
VACoT: Rethinking Visual Data Augmentation with VLMs
by: Xu, Zhengzhuo, et al.
Published: (2025)
by: Xu, Zhengzhuo, et al.
Published: (2025)
AdaNAT: Exploring Adaptive Policy for Token-Based Image Generation
by: Ni, Zanlin, et al.
Published: (2024)
by: Ni, Zanlin, et al.
Published: (2024)
Rethinking Model Selection in VLM Through the Lens of Gromov-Wasserstein Distance
by: Li, Muyang, et al.
Published: (2026)
by: Li, Muyang, et al.
Published: (2026)
Zero-Reference Joint Low-Light Enhancement and Deblurring via Visual Autoregressive Modeling with VLM-Derived Modulation
by: Dong, Wei, et al.
Published: (2025)
by: Dong, Wei, et al.
Published: (2025)
VLM-Assisted Continual learning for Visual Question Answering in Self-Driving
by: Lin, Yuxin, et al.
Published: (2025)
by: Lin, Yuxin, et al.
Published: (2025)
PhysVLM: Enabling Visual Language Models to Understand Robotic Physical Reachability
by: Zhou, Weijie, et al.
Published: (2025)
by: Zhou, Weijie, et al.
Published: (2025)
CogVLM2: Visual Language Models for Image and Video Understanding
by: Hong, Wenyi, et al.
Published: (2024)
by: Hong, Wenyi, et al.
Published: (2024)
PhysVLM-AVR: Active Visual Reasoning for Multimodal Large Language Models in Physical Environments
by: Zhou, Weijie, et al.
Published: (2025)
by: Zhou, Weijie, et al.
Published: (2025)
LensVLM: Selective Context Expansion for Compressed Visual Representation of Text
by: Xie, Roy, et al.
Published: (2026)
by: Xie, Roy, et al.
Published: (2026)
VLM-KD: Knowledge Distillation from VLM for Long-Tail Visual Recognition
by: Zhang, Zaiwei, et al.
Published: (2024)
by: Zhang, Zaiwei, et al.
Published: (2024)
Video-RAG: Visually-aligned Retrieval-Augmented Long Video Comprehension
by: Luo, Yongdong, et al.
Published: (2024)
by: Luo, Yongdong, et al.
Published: (2024)
HLGFA: High-Low Resolution Guided Feature Alignment for Unsupervised Anomaly Detection
by: Zhou, Han, et al.
Published: (2026)
by: Zhou, Han, et al.
Published: (2026)
YOLOv13: Real-Time Object Detection with Hypergraph-Enhanced Adaptive Visual Perception
by: Lei, Mengqi, et al.
Published: (2025)
by: Lei, Mengqi, et al.
Published: (2025)
NaturalVLM: Leveraging Fine-grained Natural Language for Affordance-Guided Visual Manipulation
by: Xu, Ran, et al.
Published: (2024)
by: Xu, Ran, et al.
Published: (2024)
See, Act, Adapt: Active Perception for Unsupervised Cross-Domain Visual Adaptation via Personalized VLM-Guided Agent
by: Tang, Tianci, et al.
Published: (2026)
by: Tang, Tianci, et al.
Published: (2026)
Beyond Intermediate States: Explaining Visual Redundancy through Language
by: Yang, Dingchen, et al.
Published: (2025)
by: Yang, Dingchen, et al.
Published: (2025)
SparseVILA: Decoupling Visual Sparsity for Efficient VLM Inference
by: Khaki, Samir, et al.
Published: (2025)
by: Khaki, Samir, et al.
Published: (2025)
Similar Items
-
ESP-Zero: Unsupervised enhancement of zero-shot classification for Extremely Sparse Point cloud
by: Han, Jiayi, et al.
Published: (2024) -
SLIM: Let LLM Learn More and Forget Less with Soft LoRA and Identity Mixture
by: Han, Jiayi, et al.
Published: (2024) -
Rethinking Precision of Pseudo Label: Test-Time Adaptation via Complementary Learning
by: Han, Jiayi, et al.
Published: (2023) -
Efficient Multi-modal Large Language Models via Visual Token Grouping
by: Huang, Minbin, et al.
Published: (2024) -
AdaGen: Learning Adaptive Policy for Image Synthesis
by: Ni, Zanlin, et al.
Published: (2026)