Difference Vector Equalization for Robust Fine-tuning of Vision-Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Suzuki, Satoshi, Yamaguchi, Shin'ya, Takeda, Shoichiro, Yamane, Taiga, Makishima, Naoki, Kawata, Naotaka, Ihori, Mana, Tanaka, Tomohiro, Orihashi, Shota, Masumura, Ryo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Few-shot Personalization via In-Context Learning for Speech Emotion Recognition based on Speech-Language Model
by: Ihori, Mana, et al.
Published: (2025)
by: Ihori, Mana, et al.
Published: (2025)
MSMVD: Exploiting Multi-scale Image Features via Multi-scale BEV Features for Multi-view Pedestrian Detection
by: Yamane, Taiga, et al.
Published: (2025)
by: Yamane, Taiga, et al.
Published: (2025)
Joint Modeling of Big Five and HEXACO for Multimodal Apparent Personality-trait Recognition
by: Masumura, Ryo, et al.
Published: (2025)
by: Masumura, Ryo, et al.
Published: (2025)
MVTrajecter: Multi-View Pedestrian Tracking with Trajectory Motion Cost and Trajectory Appearance Cost
by: Yamane, Taiga, et al.
Published: (2025)
by: Yamane, Taiga, et al.
Published: (2025)
MVAFormer: RGB-based Multi-View Spatio-Temporal Action Recognition with Transformer
by: Yamane, Taiga, et al.
Published: (2025)
by: Yamane, Taiga, et al.
Published: (2025)
MultiModal Fine-tuning with Synthetic Captions
by: Enomoto, Shohei, et al.
Published: (2026)
by: Enomoto, Shohei, et al.
Published: (2026)
Learning Robust Convolutional Neural Networks with Relevant Feature Focusing via Explanations
by: Adachi, Kazuki, et al.
Published: (2022)
by: Adachi, Kazuki, et al.
Published: (2022)
Explanation Bottleneck Models
by: Yamaguchi, Shin'ya, et al.
Published: (2024)
by: Yamaguchi, Shin'ya, et al.
Published: (2024)
Adaptive Random Feature Regularization on Fine-tuning Deep Neural Networks
by: Yamaguchi, Shin'ya, et al.
Published: (2024)
by: Yamaguchi, Shin'ya, et al.
Published: (2024)
Uniformity First: Uniformity-aware Test-time Adaptation of Vision-language Models against Image Corruption
by: Adachi, Kazuki, et al.
Published: (2025)
by: Adachi, Kazuki, et al.
Published: (2025)
Covariance-aware Feature Alignment with Pre-computed Source Statistics for Test-time Adaptation to Multiple Image Corruptions
by: Adachi, Kazuki, et al.
Published: (2022)
by: Adachi, Kazuki, et al.
Published: (2022)
Rationale-Enhanced Decoding for Multi-modal Chain-of-Thought
by: Yamaguchi, Shin'ya, et al.
Published: (2025)
by: Yamaguchi, Shin'ya, et al.
Published: (2025)
Test-time Adaptation Meets Image Enhancement: Improving Accuracy via Uncertainty-aware Logit Switching
by: Enomoto, Shohei, et al.
Published: (2024)
by: Enomoto, Shohei, et al.
Published: (2024)
Factor-Conditioned Speaking-Style Captioning
by: Ando, Atsushi, et al.
Published: (2024)
by: Ando, Atsushi, et al.
Published: (2024)
WAVELET ANALYSIS METHOD FOR DEFECTS DETECTION IN CFRP COMPOSITES WITH FULLY NON-CONTACT LAMB WAVES PROPAGATION
by: Lea Lecointre, et al.
Published: (2023)
by: Lea Lecointre, et al.
Published: (2023)
Zero-shot Concept Bottleneck Models
by: Yamaguchi, Shin'ya, et al.
Published: (2025)
by: Yamaguchi, Shin'ya, et al.
Published: (2025)
Test-time Adaptation for Regression by Subspace Alignment
by: Adachi, Kazuki, et al.
Published: (2024)
by: Adachi, Kazuki, et al.
Published: (2024)
Test-time Similarity Modification for Person Re-identification toward Temporal Distribution Shift
by: Adachi, Kazuki, et al.
Published: (2024)
by: Adachi, Kazuki, et al.
Published: (2024)
Toward Data Efficient Model Merging between Different Datasets without Performance Degradation
by: Yamada, Masanori, et al.
Published: (2023)
by: Yamada, Masanori, et al.
Published: (2023)
Parallel In-context Learning for Large Vision Language Models
by: Yamaguchi, Shin'ya, et al.
Published: (2026)
by: Yamaguchi, Shin'ya, et al.
Published: (2026)
All-in-One ASR: Unifying Encoder-Decoder Models of CTC, Attention, and Transducer in Dual-Mode ASR
by: Moriya, Takafumi, et al.
Published: (2025)
by: Moriya, Takafumi, et al.
Published: (2025)
Dense Landslides Triggered by a Large Earthquake Reduced Evapotranspiration in a Hilly, Forested Catchment
by: Shin'ya Katsura, et al.
Published: (2025)
by: Shin'ya Katsura, et al.
Published: (2025)
Understanding Pure Textual Reasoning for Blind Image Quality Assessment
by: Li, Yuan, et al.
Published: (2026)
by: Li, Yuan, et al.
Published: (2026)
Loading loss-cone distributions in particle simulations
by: Zenitani, Seiji, et al.
Published: (2023)
by: Zenitani, Seiji, et al.
Published: (2023)
Transfer Learning with Pre-trained Conditional Generative Models
by: Yamaguchi, Shin'ya, et al.
Published: (2022)
by: Yamaguchi, Shin'ya, et al.
Published: (2022)
Post-pre-training for Modality Alignment in Vision-Language Foundation Models
by: Yamaguchi, Shin'ya, et al.
Published: (2025)
by: Yamaguchi, Shin'ya, et al.
Published: (2025)
When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars
by: Higuchi, Rei, et al.
Published: (2025)
by: Higuchi, Rei, et al.
Published: (2025)
Self-Supervised Neural Architecture Search for Multimodal Deep Neural Networks
by: Suzuki, Shota, et al.
Published: (2025)
by: Suzuki, Shota, et al.
Published: (2025)
Condensation in a Membrane Oxygenator Analyzed by Bronchoscopy and Ultrasonography
by: Shin‐ichiro Nagao, et al.
Published: (2026)
by: Shin‐ichiro Nagao, et al.
Published: (2026)
Approximate Maintenance of Maximum Subarray Sum in the Sliding Window Model
by: Suzuki, Ryo, et al.
Published: (2026)
by: Suzuki, Ryo, et al.
Published: (2026)
Evaluating Time-Series Training Dataset through Lens of Spectrum in Deep State Space Models
by: Kanai, Sekitoshi, et al.
Published: (2024)
by: Kanai, Sekitoshi, et al.
Published: (2024)
Frequency‐Dependent Effects of Transcranial Alternating Current Stimulation on Cerebro–Cerebellar Connectivity and Sensorimotor Performance
by: Makoto Suzuki, et al.
Published: (2026)
by: Makoto Suzuki, et al.
Published: (2026)
Quaternary volcanic activity of Hudson and Lautaro volcanoes, Chilean Patagonia: New constraints from K-Ar ages
by: Yuji Orihashi
Published: (2004)
by: Yuji Orihashi
Published: (2004)
Feeding, growth and grazing rates of Gyrodinium dominans determined experimentally
by: Nakamura, Yasuo, et al.
Published: (2013)
by: Nakamura, Yasuo, et al.
Published: (2013)
Female–female competition in two giant water bug species
by: Shin‐ya Ohba, et al.
Published: (2025)
by: Shin‐ya Ohba, et al.
Published: (2025)
Lossless Vocabulary Reduction for Auto-Regressive Language Models
by: Chijiwa, Daiki, et al.
Published: (2025)
by: Chijiwa, Daiki, et al.
Published: (2025)
Cyclic Equalizability Characterized by Parikh Vectors
by: Thongjarast, Sarunyu, et al.
Published: (2026)
by: Thongjarast, Sarunyu, et al.
Published: (2026)
Primary localized cutaneous amyloidosis associated with atopic dermatitis treated successfully with nemolizumab
by: Takeshi Fukumoto, et al.
Published: (2024)
by: Takeshi Fukumoto, et al.
Published: (2024)
Transformers as Measure-Theoretic Associative Memory: A Statistical Perspective and Minimax Optimality
by: Kawata, Ryotaro, et al.
Published: (2026)
by: Kawata, Ryotaro, et al.
Published: (2026)
IPCD: Intrinsic Point-Cloud Decomposition
by: Sato, Shogo, et al.
Published: (2025)
by: Sato, Shogo, et al.
Published: (2025)
Similar Items
-
Few-shot Personalization via In-Context Learning for Speech Emotion Recognition based on Speech-Language Model
by: Ihori, Mana, et al.
Published: (2025) -
MSMVD: Exploiting Multi-scale Image Features via Multi-scale BEV Features for Multi-view Pedestrian Detection
by: Yamane, Taiga, et al.
Published: (2025) -
Joint Modeling of Big Five and HEXACO for Multimodal Apparent Personality-trait Recognition
by: Masumura, Ryo, et al.
Published: (2025) -
MVTrajecter: Multi-View Pedestrian Tracking with Trajectory Motion Cost and Trajectory Appearance Cost
by: Yamane, Taiga, et al.
Published: (2025) -
MVAFormer: RGB-based Multi-View Spatio-Temporal Action Recognition with Transformer
by: Yamane, Taiga, et al.
Published: (2025)