Decoupling Vision and Language: Codebook Anchored Visual Adaptation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wu, Jason, Zhao, Tianchen, Liu, Chang, Cai, Jiarui, Zhang, Zheng, Li, Zhuowei, Singh, Aaditya, Xu, Xiang, Srivastava, Mani, Wu, Jonathan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Implementing Adaptations for Vision AutoRegressive Model
von: Shaikh, Kaif, et al.
Veröffentlicht: (2025)
von: Shaikh, Kaif, et al.
Veröffentlicht: (2025)
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
von: Raoufi, Behnam, et al.
Veröffentlicht: (2025)
von: Raoufi, Behnam, et al.
Veröffentlicht: (2025)
Lifelong Learning in Vision-Language Models: Enhanced EWC with Cross-Modal Knowledge Retention
von: Durrani, Hamza Ahmed, et al.
Veröffentlicht: (2026)
von: Durrani, Hamza Ahmed, et al.
Veröffentlicht: (2026)
Optimal Transport-Guided Source-Free Adaptation for Face Anti-Spoofing
von: Li, Zhuowei, et al.
Veröffentlicht: (2025)
von: Li, Zhuowei, et al.
Veröffentlicht: (2025)
Butter: Frequency Consistency and Hierarchical Fusion for Autonomous Driving Object Detection
von: Lin, Xiaojian, et al.
Veröffentlicht: (2025)
von: Lin, Xiaojian, et al.
Veröffentlicht: (2025)
Salient Concept-Aware Generative Data Augmentation
von: Zhao, Tianchen, et al.
Veröffentlicht: (2025)
von: Zhao, Tianchen, et al.
Veröffentlicht: (2025)
SpectralCA: Bi-Directional Cross-Attention for Next-Generation UAV Hyperspectral Vision
von: Brovko, D. V.
Veröffentlicht: (2025)
von: Brovko, D. V.
Veröffentlicht: (2025)
Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning
von: Ji, Binbin, et al.
Veröffentlicht: (2025)
von: Ji, Binbin, et al.
Veröffentlicht: (2025)
Object detection in adverse weather conditions for autonomous vehicles using Instruct Pix2Pix
von: Gurbindo, Unai, et al.
Veröffentlicht: (2025)
von: Gurbindo, Unai, et al.
Veröffentlicht: (2025)
Predictive Modeling of Maritime Radar Data Using Transformer Architecture
von: Qesaraku, Bjorna, et al.
Veröffentlicht: (2025)
von: Qesaraku, Bjorna, et al.
Veröffentlicht: (2025)
VLM-VPI: A Vision-Language Reasoning Framework for Improving Automated Vehicle-Pedestrian Interactions
von: Pu, Qingwen, et al.
Veröffentlicht: (2026)
von: Pu, Qingwen, et al.
Veröffentlicht: (2026)
FedWCM: Unleashing the Potential of Momentum-based Federated Learning in Long-Tailed Scenarios
von: Li, Tianle, et al.
Veröffentlicht: (2025)
von: Li, Tianle, et al.
Veröffentlicht: (2025)
Motion Attribution for Video Generation
von: Wu, Xindi, et al.
Veröffentlicht: (2026)
von: Wu, Xindi, et al.
Veröffentlicht: (2026)
Geo2Sound: A Scalable Geo-Aligned Framework for Soundscape Generation from Satellite Imagery
von: Wu, Kunlin, et al.
Veröffentlicht: (2026)
von: Wu, Kunlin, et al.
Veröffentlicht: (2026)
Selection, Not Fusion: Radar-Modulated State Space Models for Radar-Camera Depth Estimation
von: Hou, Zhangcheng, et al.
Veröffentlicht: (2026)
von: Hou, Zhangcheng, et al.
Veröffentlicht: (2026)
Smooth regularization for efficient video recognition
von: Goldman, Gil, et al.
Veröffentlicht: (2025)
von: Goldman, Gil, et al.
Veröffentlicht: (2025)
Neuromorphic Monocular Depth Estimation with Uncertainty Modeling
von: Bergkvist, Viktor, et al.
Veröffentlicht: (2026)
von: Bergkvist, Viktor, et al.
Veröffentlicht: (2026)
Sparse Concept Bottleneck Models: Gumbel Tricks in Contrastive Learning
von: Semenov, Andrei, et al.
Veröffentlicht: (2024)
von: Semenov, Andrei, et al.
Veröffentlicht: (2024)
TAG-Head: Time-Aligned Graph Head for Plug-and-Play Fine-grained Action Recognition
von: Hassan, Imtiaz Ul, et al.
Veröffentlicht: (2026)
von: Hassan, Imtiaz Ul, et al.
Veröffentlicht: (2026)
DejaVid: Encoder-Agnostic Learned Temporal Matching for Video Classification
von: Ho, Darryl, et al.
Veröffentlicht: (2025)
von: Ho, Darryl, et al.
Veröffentlicht: (2025)
High-Frequency Semantics and Geometric Priors for End-to-End Detection Transformers in Challenging UAV Imagery
von: Peng, Hongxing, et al.
Veröffentlicht: (2025)
von: Peng, Hongxing, et al.
Veröffentlicht: (2025)
PhysicsNeRF: Physics-Guided 3D Reconstruction from Sparse Views
von: Barhdadi, Mohamed Rayan, et al.
Veröffentlicht: (2025)
von: Barhdadi, Mohamed Rayan, et al.
Veröffentlicht: (2025)
Deep Learning-based Depth Estimation Methods from Monocular Image and Videos: A Comprehensive Survey
von: Rajapaksha, Uchitha, et al.
Veröffentlicht: (2024)
von: Rajapaksha, Uchitha, et al.
Veröffentlicht: (2024)
Rethinking Visual Intelligence: Insights from Video Pretraining
von: Acuaviva, Pablo, et al.
Veröffentlicht: (2025)
von: Acuaviva, Pablo, et al.
Veröffentlicht: (2025)
NV3D: Leveraging Spatial Shape Through Normal Vector-based 3D Object Detection
von: Chaowakarn, Krittin, et al.
Veröffentlicht: (2025)
von: Chaowakarn, Krittin, et al.
Veröffentlicht: (2025)
Perceptual Flow Network for Visually Grounded Reasoning
von: Li, Yangfu, et al.
Veröffentlicht: (2026)
von: Li, Yangfu, et al.
Veröffentlicht: (2026)
CCVA-FL: Cross-Client Variations Adaptive Federated Learning for Medical Imaging
von: Gupta, Sunny, et al.
Veröffentlicht: (2024)
von: Gupta, Sunny, et al.
Veröffentlicht: (2024)
Taming the Tail: Leveraging Asymmetric Loss and Pade Approximation to Overcome Medical Image Long-Tailed Class Imbalance
von: Kashyap, Pankhi, et al.
Veröffentlicht: (2024)
von: Kashyap, Pankhi, et al.
Veröffentlicht: (2024)
Few-Shot Learning of a Graph-Based Neural Network Model Without Backpropagation
von: Lapin, Mykyta, et al.
Veröffentlicht: (2025)
von: Lapin, Mykyta, et al.
Veröffentlicht: (2025)
CARScenes: Semantic VLM Dataset for Safe Autonomous Driving
von: He, Yuankai, et al.
Veröffentlicht: (2025)
von: He, Yuankai, et al.
Veröffentlicht: (2025)
WildRoadBench: A Wild Aerial Road-Damage Grounding Benchmark for Vision-Language Models and Autonomous Agents
von: Liu, Bingnan, et al.
Veröffentlicht: (2026)
von: Liu, Bingnan, et al.
Veröffentlicht: (2026)
Interpreting Structured Perturbations in Image Protection Methods for Diffusion Models
von: Martin, Michael R., et al.
Veröffentlicht: (2025)
von: Martin, Michael R., et al.
Veröffentlicht: (2025)
FeedbackSTS-Det: Sparse Frames-Based Spatio-Temporal Semantic Feedback Network for Moving Infrared Small Target Detection
von: Huang, Yian, et al.
Veröffentlicht: (2026)
von: Huang, Yian, et al.
Veröffentlicht: (2026)
SERA-H: Beyond Native Sentinel Spatial Limits for High-Resolution Canopy Height Mapping
von: Boudras, Thomas, et al.
Veröffentlicht: (2025)
von: Boudras, Thomas, et al.
Veröffentlicht: (2025)
MaSC: A Masked Similarity Metric for Evaluating Concept-Driven Generation
von: Bartkowiak, Patryk, et al.
Veröffentlicht: (2026)
von: Bartkowiak, Patryk, et al.
Veröffentlicht: (2026)
Adapting SAM with Dynamic Similarity Graphs for Few-Shot Parameter-Efficient Small Dense Object Detection: A Case Study of Chickpea Pods in Field Conditions
von: Jiang, Xintong, et al.
Veröffentlicht: (2025)
von: Jiang, Xintong, et al.
Veröffentlicht: (2025)
CG-HOI: Contact-Guided 3D Human-Object Interaction Generation
von: Diller, Christian, et al.
Veröffentlicht: (2023)
von: Diller, Christian, et al.
Veröffentlicht: (2023)
Sign language recognition based on deep learning and low-cost handcrafted descriptors
von: Carneiro, Alvaro Leandro Cavalcante, et al.
Veröffentlicht: (2024)
von: Carneiro, Alvaro Leandro Cavalcante, et al.
Veröffentlicht: (2024)
Closed-Loop Neural Activation Control in Vision-Language-Action Models
von: Babu, Abhijith, et al.
Veröffentlicht: (2026)
von: Babu, Abhijith, et al.
Veröffentlicht: (2026)
Dream to Fly: Model-Based Reinforcement Learning for Vision-Based Drone Flight
von: Romero, Angel, et al.
Veröffentlicht: (2025)
von: Romero, Angel, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Implementing Adaptations for Vision AutoRegressive Model
von: Shaikh, Kaif, et al.
Veröffentlicht: (2025) -
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
von: Raoufi, Behnam, et al.
Veröffentlicht: (2025) -
Lifelong Learning in Vision-Language Models: Enhanced EWC with Cross-Modal Knowledge Retention
von: Durrani, Hamza Ahmed, et al.
Veröffentlicht: (2026) -
Optimal Transport-Guided Source-Free Adaptation for Face Anti-Spoofing
von: Li, Zhuowei, et al.
Veröffentlicht: (2025) -
Butter: Frequency Consistency and Hierarchical Fusion for Autonomous Driving Object Detection
von: Lin, Xiaojian, et al.
Veröffentlicht: (2025)