SpectraDINO: Bridging the Spectral Gap in Vision Foundation Models via Lightweight Adapters
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Nalcakan, Yagiz, Ju, Hyeongjin, Park, Incheol, Yeo, Sanghyeop, Jin, Youngwan, Kim, Shiho |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
RPT-SR: Regional Prior attention Transformer for infrared image Super-Resolution
von: Jin, Youngwan, et al.
Veröffentlicht: (2026)
von: Jin, Youngwan, et al.
Veröffentlicht: (2026)
Pix2Next: Leveraging Vision Foundation Models for RGB to NIR Image Translation
von: Jin, Youngwan, et al.
Veröffentlicht: (2024)
von: Jin, Youngwan, et al.
Veröffentlicht: (2024)
RASMD: RGB And SWIR Multispectral Driving Dataset for Robust Perception in Adverse Conditions
von: Jin, Youngwan, et al.
Veröffentlicht: (2025)
von: Jin, Youngwan, et al.
Veröffentlicht: (2025)
VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding
von: Kim, Kangsan, et al.
Veröffentlicht: (2024)
von: Kim, Kangsan, et al.
Veröffentlicht: (2024)
Surgical-DINO: Adapter Learning of Foundation Models for Depth Estimation in Endoscopic Surgery
von: Cui, Beilei, et al.
Veröffentlicht: (2024)
von: Cui, Beilei, et al.
Veröffentlicht: (2024)
Driving with DINO: Vision Foundation Features as a Unified Bridge for Sim-to-Real Generation in Autonomous Driving
von: Chen, Xuyang, et al.
Veröffentlicht: (2026)
von: Chen, Xuyang, et al.
Veröffentlicht: (2026)
Visualizing the loss landscape of Self-supervised Vision Transformer
von: Lee, Youngwan, et al.
Veröffentlicht: (2024)
von: Lee, Youngwan, et al.
Veröffentlicht: (2024)
T-Gated Adapter: A Lightweight Temporal Adapter for Vision-Language Medical Segmentation
von: Khadka, Pranjal
Veröffentlicht: (2026)
von: Khadka, Pranjal
Veröffentlicht: (2026)
GuiDINO: Rethinking Vision Foundation Model in Medical Image Segmentation
von: Liang, Zhuonan, et al.
Veröffentlicht: (2026)
von: Liang, Zhuonan, et al.
Veröffentlicht: (2026)
HoliSafe: Holistic Safety Benchmarking and Modeling for Vision-Language Model
von: Lee, Youngwan, et al.
Veröffentlicht: (2025)
von: Lee, Youngwan, et al.
Veröffentlicht: (2025)
Bridging the Domain Gap for Flight-Ready Spaceborne Vision
von: Park, Tae Ha, et al.
Veröffentlicht: (2024)
von: Park, Tae Ha, et al.
Veröffentlicht: (2024)
MA-EgoQA: Question Answering over Egocentric Videos from Multiple Embodied Agents
von: Kim, Kangsan, et al.
Veröffentlicht: (2026)
von: Kim, Kangsan, et al.
Veröffentlicht: (2026)
SIGMA: Bridging Structural and Distributional Gaps for Vision Foundation Model Adaptation
von: Xiong, Lingyu, et al.
Veröffentlicht: (2026)
von: Xiong, Lingyu, et al.
Veröffentlicht: (2026)
DINO-R1: Incentivizing Reasoning Capability in Vision Foundation Models
von: Pan, Chenbin, et al.
Veröffentlicht: (2025)
von: Pan, Chenbin, et al.
Veröffentlicht: (2025)
Inv-Adapter: ID Customization Generation via Image Inversion and Lightweight Adapter
von: Xing, Peng, et al.
Veröffentlicht: (2024)
von: Xing, Peng, et al.
Veröffentlicht: (2024)
DINO-CoDT: Multi-class Collaborative Detection and Tracking with Vision Foundation Models
von: He, Xunjie, et al.
Veröffentlicht: (2025)
von: He, Xunjie, et al.
Veröffentlicht: (2025)
MonoDINO-DETR: Depth-Enhanced Monocular 3D Object Detection Using a Vision Foundation Model
von: Kim, Jihyeok, et al.
Veröffentlicht: (2025)
von: Kim, Jihyeok, et al.
Veröffentlicht: (2025)
Earth-Adapter: Bridge the Geospatial Domain Gaps with Mixture of Frequency Adaptation
von: Hu, Xiaoxing, et al.
Veröffentlicht: (2025)
von: Hu, Xiaoxing, et al.
Veröffentlicht: (2025)
VLPose: Bridging the Domain Gap in Pose Estimation with Language-Vision Tuning
von: Li, Jingyao, et al.
Veröffentlicht: (2024)
von: Li, Jingyao, et al.
Veröffentlicht: (2024)
KOALA: Empirical Lessons Toward Memory-Efficient and Fast Diffusion Models for Text-to-Image Synthesis
von: Lee, Youngwan, et al.
Veröffentlicht: (2023)
von: Lee, Youngwan, et al.
Veröffentlicht: (2023)
SIA: Enhancing Safety via Intent Awareness for Vision-Language Models
von: Na, Youngjin, et al.
Veröffentlicht: (2025)
von: Na, Youngjin, et al.
Veröffentlicht: (2025)
SpectraIrisPAD: Leveraging Vision Foundation Models for Spectrally Conditioned Multispectral Iris Presentation Attack Detection
von: Ramachandra, Raghavendra, et al.
Veröffentlicht: (2025)
von: Ramachandra, Raghavendra, et al.
Veröffentlicht: (2025)
SSDA: Bridging Spectral and Structural Gaps via Dual Adaptation for Vision-Based Time Series Forecasting
von: Zhang, Mingrui, et al.
Veröffentlicht: (2026)
von: Zhang, Mingrui, et al.
Veröffentlicht: (2026)
DINO-Foresight: Looking into the Future with DINO
von: Karypidis, Efstathios, et al.
Veröffentlicht: (2024)
von: Karypidis, Efstathios, et al.
Veröffentlicht: (2024)
MultihopSpatial: Multi-hop Compositional Spatial Reasoning Benchmark for Vision-Language Model
von: Lee, Youngwan, et al.
Veröffentlicht: (2026)
von: Lee, Youngwan, et al.
Veröffentlicht: (2026)
AD-DINO: Attention-Dynamic DINO for Distance-Aware Embodied Reference Understanding
von: Guo, Hao, et al.
Veröffentlicht: (2024)
von: Guo, Hao, et al.
Veröffentlicht: (2024)
Back to the Features: DINO as a Foundation for Video World Models
von: Baldassarre, Federico, et al.
Veröffentlicht: (2025)
von: Baldassarre, Federico, et al.
Veröffentlicht: (2025)
DBT-DINO: Towards Foundation model based analysis of Digital Breast Tomosynthesis
von: Dorfner, Felix J., et al.
Veröffentlicht: (2025)
von: Dorfner, Felix J., et al.
Veröffentlicht: (2025)
DINO-Tok: Adapting DINO for Visual Tokenizers
von: Jia, Mingkai, et al.
Veröffentlicht: (2025)
von: Jia, Mingkai, et al.
Veröffentlicht: (2025)
Talking to DINO: Bridging Self-Supervised Vision Backbones with Language for Open-Vocabulary Segmentation
von: Barsellotti, Luca, et al.
Veröffentlicht: (2024)
von: Barsellotti, Luca, et al.
Veröffentlicht: (2024)
Generalizable Disaster Damage Assessment via Change Detection with Vision Foundation Model
von: Ahn, Kyeongjin, et al.
Veröffentlicht: (2024)
von: Ahn, Kyeongjin, et al.
Veröffentlicht: (2024)
ReferDINO: Referring Video Object Segmentation with Visual Grounding Foundations
von: Liang, Tianming, et al.
Veröffentlicht: (2025)
von: Liang, Tianming, et al.
Veröffentlicht: (2025)
VideoPatchCore: An Effective Method to Memorize Normality for Video Anomaly Detection
von: Ahn, Sunghyun, et al.
Veröffentlicht: (2024)
von: Ahn, Sunghyun, et al.
Veröffentlicht: (2024)
Optimal Transport Adapter Tuning for Bridging Modality Gaps in Few-Shot Remote Sensing Scene Classification
von: Ji, Zhong, et al.
Veröffentlicht: (2025)
von: Ji, Zhong, et al.
Veröffentlicht: (2025)
EndoDINO: A Foundation Model for GI Endoscopy
von: Dermyer, Patrick, et al.
Veröffentlicht: (2025)
von: Dermyer, Patrick, et al.
Veröffentlicht: (2025)
BGG: Bridging the Geometric Gap between Cross-View images by Vision Foundation Model Adaptation for Geo-Localization
von: Wang, Wei, et al.
Veröffentlicht: (2026)
von: Wang, Wei, et al.
Veröffentlicht: (2026)
Evaluating Stenosis Detection with Grounding DINO, YOLO, and DINO-DETR
von: Ansari, Muhammad Musab
Veröffentlicht: (2025)
von: Ansari, Muhammad Musab
Veröffentlicht: (2025)
DINO in the Room: Leveraging 2D Foundation Models for 3D Segmentation
von: Knaebel, Karim, et al.
Veröffentlicht: (2025)
von: Knaebel, Karim, et al.
Veröffentlicht: (2025)
Hyperspectral Adapter for Semantic Segmentation with Vision Foundation Models
von: Hurtado, Juana Valeria, et al.
Veröffentlicht: (2025)
von: Hurtado, Juana Valeria, et al.
Veröffentlicht: (2025)
HeGraphAdapter: Tuning Multi-Modal Vision-Language Models with Heterogeneous Graph Adapter
von: Zhao, Yumiao, et al.
Veröffentlicht: (2024)
von: Zhao, Yumiao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
RPT-SR: Regional Prior attention Transformer for infrared image Super-Resolution
von: Jin, Youngwan, et al.
Veröffentlicht: (2026) -
Pix2Next: Leveraging Vision Foundation Models for RGB to NIR Image Translation
von: Jin, Youngwan, et al.
Veröffentlicht: (2024) -
RASMD: RGB And SWIR Multispectral Driving Dataset for Robust Perception in Adverse Conditions
von: Jin, Youngwan, et al.
Veröffentlicht: (2025) -
VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding
von: Kim, Kangsan, et al.
Veröffentlicht: (2024) -
Surgical-DINO: Adapter Learning of Foundation Models for Depth Estimation in Endoscopic Surgery
von: Cui, Beilei, et al.
Veröffentlicht: (2024)