Saved in:
| Main Authors: | Ma, Yiming, Yang, Hongkun, Wang, Lionel Z., Chen, Bin, Xian, Weizhi, Teng, Jianzhi |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2603.01111 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Attention in Space: Functional Roles of VLM Heads for Spatial Reasoning
by: Ma, Xueqi, et al.
Published: (2026)
by: Ma, Xueqi, et al.
Published: (2026)
Fine-Grained VLM Fine-tuning via Latent Hierarchical Adapter Learning
by: Zhao, Yumiao, et al.
Published: (2025)
by: Zhao, Yumiao, et al.
Published: (2025)
HOLa: Zero-Shot HOI Detection with Low-Rank Decomposed VLM Feature Adaptation
by: Lei, Qinqian, et al.
Published: (2025)
by: Lei, Qinqian, et al.
Published: (2025)
DOCR-Inspector: Fine-Grained and Automated Evaluation of Document Parsing with VLM
by: Zhang, Qintong, et al.
Published: (2025)
by: Zhang, Qintong, et al.
Published: (2025)
Think-Before-Draw: Decomposing Emotion Semantics & Fine-Grained Controllable Expressive Talking Head Generation
by: Shi, Hanlei, et al.
Published: (2025)
by: Shi, Hanlei, et al.
Published: (2025)
Head-Aware Visual Cropping: Enhancing Fine-Grained VQA with Attention-Guided Subimage
by: Xie, Junfei, et al.
Published: (2026)
by: Xie, Junfei, et al.
Published: (2026)
FacialFlowNet: Advancing Facial Optical Flow Estimation with a Diverse Dataset and a Decomposed Model
by: Lu, Jianzhi, et al.
Published: (2024)
by: Lu, Jianzhi, et al.
Published: (2024)
HieroAction: Hierarchically Guided VLM for Fine-Grained Action Analysis
by: Wu, Junhao, et al.
Published: (2025)
by: Wu, Junhao, et al.
Published: (2025)
XR-VLM: Cross-Relationship Modeling with Multi-part Prompts and Visual Features for Fine-Grained Recognition
by: Wang, Chuanming, et al.
Published: (2025)
by: Wang, Chuanming, et al.
Published: (2025)
AURA: A Fine-Grained Benchmark and Decomposed Metric for Audio-Visual Reasoning
by: Galougah, Siminfar Samakoush, et al.
Published: (2025)
by: Galougah, Siminfar Samakoush, et al.
Published: (2025)
Decomposed Vision-Language Alignment for Fine-Grained Open-Vocabulary Segmentation
by: Wang, Chenhao, et al.
Published: (2026)
by: Wang, Chenhao, et al.
Published: (2026)
A Large-Scale Remote Sensing Dataset and VLM-based Algorithm for Fine-Grained Road Hierarchy Classification
by: Han, Ting, et al.
Published: (2026)
by: Han, Ting, et al.
Published: (2026)
Lightning Fast Caching-based Parallel Denoising Prediction for Accelerating Talking Head Generation
by: Long, Jianzhi, et al.
Published: (2025)
by: Long, Jianzhi, et al.
Published: (2025)
Domain Adaptation of Attention Heads for Zero-shot Anomaly Detection
by: Jeong, Kiyoon, et al.
Published: (2025)
by: Jeong, Kiyoon, et al.
Published: (2025)
Robust Multiview Multimodal Driver Monitoring System Using Masked Multi-Head Self-Attention
by: Ma, Yiming, et al.
Published: (2023)
by: Ma, Yiming, et al.
Published: (2023)
Synchronized and Fine-Grained Head for Skeleton-Based Ambiguous Action Recognition
by: Huang, Hao, et al.
Published: (2024)
by: Huang, Hao, et al.
Published: (2024)
Beyond Graph Model: Reliable VLM Fine-Tuning via Random Graph Adapter
by: Jiang, Bo, et al.
Published: (2025)
by: Jiang, Bo, et al.
Published: (2025)
SaRA: High-Efficient Diffusion Model Fine-tuning with Progressive Sparse Low-Rank Adaptation
by: Hu, Teng, et al.
Published: (2024)
by: Hu, Teng, et al.
Published: (2024)
From Attenuation to Attention: Variational Information Flow Manipulation for Fine-Grained Visual Perception
by: Zhu, Jilong, et al.
Published: (2026)
by: Zhu, Jilong, et al.
Published: (2026)
Head-wise Adaptive Rotary Positional Encoding for Fine-Grained Image Generation
by: Li, Jiaye, et al.
Published: (2025)
by: Li, Jiaye, et al.
Published: (2025)
Toward Fine-Grained Facial Control in 3D Talking Head Generation
by: Xie, Shaoyang, et al.
Published: (2026)
by: Xie, Shaoyang, et al.
Published: (2026)
Toward Reliable VLM: A Fine-Grained Benchmark and Framework for Exposure, Bias, and Inference in Korean Street Views
by: Wang, Xiaonan, et al.
Published: (2025)
by: Wang, Xiaonan, et al.
Published: (2025)
DoRA: Weight-Decomposed Low-Rank Adaptation
by: Liu, Shih-Yang, et al.
Published: (2024)
by: Liu, Shih-Yang, et al.
Published: (2024)
LoLDU: Low-Rank Adaptation via Lower-Diag-Upper Decomposition for Parameter-Efficient Fine-Tuning
by: Shi, Yiming, et al.
Published: (2024)
by: Shi, Yiming, et al.
Published: (2024)
Reassessing the Role of Supervised Fine-Tuning: An Empirical Study in VLM Reasoning
by: Yu, Yongcan, et al.
Published: (2025)
by: Yu, Yongcan, et al.
Published: (2025)
Mask-Guided Attention Regulation for Anatomically Consistent Counterfactual CXR Synthesis
by: Zhang, Zichun, et al.
Published: (2026)
by: Zhang, Zichun, et al.
Published: (2026)
D-Attn: Decomposed Attention for Large Vision-and-Language Models
by: Kuo, Chia-Wen, et al.
Published: (2025)
by: Kuo, Chia-Wen, et al.
Published: (2025)
Fine-Grained Open-Vocabulary Object Detection with Fined-Grained Prompts: Task, Dataset and Benchmark
by: Liu, Ying, et al.
Published: (2025)
by: Liu, Ying, et al.
Published: (2025)
TryOn-Adapter: Efficient Fine-Grained Clothing Identity Adaptation for High-Fidelity Virtual Try-On
by: Xing, Jiazheng, et al.
Published: (2024)
by: Xing, Jiazheng, et al.
Published: (2024)
DePT: Decomposed Prompt Tuning for Parameter-Efficient Fine-tuning
by: Shi, Zhengxiang, et al.
Published: (2023)
by: Shi, Zhengxiang, et al.
Published: (2023)
Fine-Grained Prototypes Distillation for Few-Shot Object Detection
by: Wang, Zichen, et al.
Published: (2024)
by: Wang, Zichen, et al.
Published: (2024)
ORION: ORthonormal Text Encoding for Universal VLM AdaptatION
by: Chakraborty, Omprakash, et al.
Published: (2026)
by: Chakraborty, Omprakash, et al.
Published: (2026)
DA-HFNet: Progressive Fine-Grained Forgery Image Detection and Localization Based on Dual Attention
by: Liu, Yang, et al.
Published: (2024)
by: Liu, Yang, et al.
Published: (2024)
Knowledge Transfer and Domain Adaptation for Fine-Grained Remote Sensing Image Segmentation
by: Zhang, Shun, et al.
Published: (2024)
by: Zhang, Shun, et al.
Published: (2024)
VLM-FO1: Bridging the Gap Between High-Level Reasoning and Fine-Grained Perception in VLMs
by: Liu, Peng, et al.
Published: (2025)
by: Liu, Peng, et al.
Published: (2025)
De-fine: Decomposing and Refining Visual Programs with Auto-Feedback
by: Gao, Minghe, et al.
Published: (2023)
by: Gao, Minghe, et al.
Published: (2023)
LES-Talker: Fine-Grained Emotion Editing for Talking Head Generation in Linear Emotion Space
by: Feng, Guanwen, et al.
Published: (2024)
by: Feng, Guanwen, et al.
Published: (2024)
Domain Adaptation of VLM for Soccer Video Understanding
by: Jiang, Tiancheng, et al.
Published: (2025)
by: Jiang, Tiancheng, et al.
Published: (2025)
CropVLM: Learning to Zoom for Fine-Grained Vision-Language Perception
by: Carvalho, Miguel, et al.
Published: (2025)
by: Carvalho, Miguel, et al.
Published: (2025)
Modality-Collaborative Low-Rank Decomposers for Few-Shot Video Domain Adaptation
by: Wanyan, Yuyang, et al.
Published: (2025)
by: Wanyan, Yuyang, et al.
Published: (2025)
Similar Items
-
Attention in Space: Functional Roles of VLM Heads for Spatial Reasoning
by: Ma, Xueqi, et al.
Published: (2026) -
Fine-Grained VLM Fine-tuning via Latent Hierarchical Adapter Learning
by: Zhao, Yumiao, et al.
Published: (2025) -
HOLa: Zero-Shot HOI Detection with Low-Rank Decomposed VLM Feature Adaptation
by: Lei, Qinqian, et al.
Published: (2025) -
DOCR-Inspector: Fine-Grained and Automated Evaluation of Document Parsing with VLM
by: Zhang, Qintong, et al.
Published: (2025) -
Think-Before-Draw: Decomposing Emotion Semantics & Fine-Grained Controllable Expressive Talking Head Generation
by: Shi, Hanlei, et al.
Published: (2025)