CrossVL: Complexity-Aware Feature Routing and Paired Curriculum for Cross-View Vision-Language Detection
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Liu, Zhipeng, Luo, Chunbo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Cross-Domain Generalization Limits of Vision Foundation Models in Facial Deepfake Detection
von: Delibasoglu, Ibrahim
Veröffentlicht: (2026)
von: Delibasoglu, Ibrahim
Veröffentlicht: (2026)
Anomaly-Aware Vision-Language Adapters for Zero-Shot Anomaly Detection
von: Aqeel, Muhammad, et al.
Veröffentlicht: (2026)
von: Aqeel, Muhammad, et al.
Veröffentlicht: (2026)
Multi-scale Quaternion CNN and BiGRU with Cross Self-attention Feature Fusion for Fault Diagnosis of Bearing
von: Liu, Huanbai, et al.
Veröffentlicht: (2024)
von: Liu, Huanbai, et al.
Veröffentlicht: (2024)
MFAF: An EVA02-Based Multi-scale Frequency Attention Fusion Method for Cross-View Geo-Localization
von: Liu, YiTong, et al.
Veröffentlicht: (2025)
von: Liu, YiTong, et al.
Veröffentlicht: (2025)
NVIDIA Nemotron Nano V2 VL
von: NVIDIA, et al.
Veröffentlicht: (2025)
von: NVIDIA, et al.
Veröffentlicht: (2025)
HalluRNN: Mitigating Hallucinations via Recurrent Cross-Layer Reasoning in Large Vision-Language Models
von: Yu, Le, et al.
Veröffentlicht: (2025)
von: Yu, Le, et al.
Veröffentlicht: (2025)
Exposing and Addressing Cross-Task Inconsistency in Unified Vision-Language Models
von: Maharana, Adyasha, et al.
Veröffentlicht: (2023)
von: Maharana, Adyasha, et al.
Veröffentlicht: (2023)
COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training
von: Kim, Sanghwan, et al.
Veröffentlicht: (2024)
von: Kim, Sanghwan, et al.
Veröffentlicht: (2024)
ROCKET-2: Steering Visuomotor Policy via Cross-View Goal Alignment
von: Cai, Shaofei, et al.
Veröffentlicht: (2025)
von: Cai, Shaofei, et al.
Veröffentlicht: (2025)
Visual Sync: Multi-Camera Synchronization via Cross-View Object Motion
von: Liu, Shaowei, et al.
Veröffentlicht: (2025)
von: Liu, Shaowei, et al.
Veröffentlicht: (2025)
Feature Purified Transformer With Cross-level Feature Guiding Decoder For Multi-class OOD and Anomaly Deteciton
von: Lin, Jerry Chun-Wei, et al.
Veröffentlicht: (2024)
von: Lin, Jerry Chun-Wei, et al.
Veröffentlicht: (2024)
CLASH: A Benchmark for Cross-Modal Contradiction Detection
von: Popordanoska, Teodora, et al.
Veröffentlicht: (2025)
von: Popordanoska, Teodora, et al.
Veröffentlicht: (2025)
Sparse Autoencoders Learn Monosemantic Features in Vision-Language Models
von: Pach, Mateusz, et al.
Veröffentlicht: (2025)
von: Pach, Mateusz, et al.
Veröffentlicht: (2025)
VIFO: Visual Feature Empowered Multivariate Time Series Forecasting with Cross-Modal Fusion
von: Wang, Yanlong, et al.
Veröffentlicht: (2025)
von: Wang, Yanlong, et al.
Veröffentlicht: (2025)
ReasoningTrack: Chain-of-Thought Reasoning for Long-term Vision-Language Tracking
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
Domain-Invariant Per-Frame Feature Extraction for Cross-Domain Imitation Learning with Visual Observations
von: Kim, Minung, et al.
Veröffentlicht: (2025)
von: Kim, Minung, et al.
Veröffentlicht: (2025)
X-AVDT: Audio-Visual Cross-Attention for Robust Deepfake Detection
von: Kim, Youngseo, et al.
Veröffentlicht: (2026)
von: Kim, Youngseo, et al.
Veröffentlicht: (2026)
Transferability-Guided Cross-Domain Cross-Task Transfer Learning
von: Tan, Yang, et al.
Veröffentlicht: (2022)
von: Tan, Yang, et al.
Veröffentlicht: (2022)
Multi-Layer Feature Fusion with Cross-Channel Attention-Based U-Net for Kidney Tumor Segmentation
von: Neha, Fnu, et al.
Veröffentlicht: (2024)
von: Neha, Fnu, et al.
Veröffentlicht: (2024)
R1-SyntheticVL: Is Synthetic Data from Generative Models Ready for Multimodal Large Language Model?
von: Zhang, Jingyi, et al.
Veröffentlicht: (2026)
von: Zhang, Jingyi, et al.
Veröffentlicht: (2026)
VERA: Explainable Video Anomaly Detection via Verbalized Learning of Vision-Language Models
von: Ye, Muchao, et al.
Veröffentlicht: (2024)
von: Ye, Muchao, et al.
Veröffentlicht: (2024)
Cross Dataset Analysis and Network Architecture Repair for Autonomous Car Lane Detection
von: Ganeriwala, Parth, et al.
Veröffentlicht: (2024)
von: Ganeriwala, Parth, et al.
Veröffentlicht: (2024)
CNC: Cross-modal Normality Constraint for Unsupervised Multi-class Anomaly Detection
von: Wang, Xiaolei, et al.
Veröffentlicht: (2024)
von: Wang, Xiaolei, et al.
Veröffentlicht: (2024)
Robust Building Damage Detection in Cross-Disaster Settings Using Domain Adaptation
von: Mouradi, Asmae, et al.
Veröffentlicht: (2026)
von: Mouradi, Asmae, et al.
Veröffentlicht: (2026)
Soft Task-Aware Routing of Experts for Equivariant Representation Learning
von: Jeon, Jaebyeong, et al.
Veröffentlicht: (2025)
von: Jeon, Jaebyeong, et al.
Veröffentlicht: (2025)
Vision-Language Meets the Skeleton: Progressively Distillation with Cross-Modal Knowledge for 3D Action Representation Learning
von: Chen, Yang, et al.
Veröffentlicht: (2024)
von: Chen, Yang, et al.
Veröffentlicht: (2024)
AgrI Challenge: A Data-Centric AI Competition for Cross-Team Validation in Agricultural Vision
von: Brahimi, Mohammed, et al.
Veröffentlicht: (2026)
von: Brahimi, Mohammed, et al.
Veröffentlicht: (2026)
Zero-shot Vision-Language Reranking for Cross-View Geolocalization
von: Erzurumlu, Yunus Talha, et al.
Veröffentlicht: (2026)
von: Erzurumlu, Yunus Talha, et al.
Veröffentlicht: (2026)
Uncertainty-Aware Test-Time Adaptation for Cross-Region Spatio-Temporal Fusion of Land Surface Temperature
von: Bouaziz, Sofiane, et al.
Veröffentlicht: (2026)
von: Bouaziz, Sofiane, et al.
Veröffentlicht: (2026)
Spurious Feature Eraser: Stabilizing Test-Time Adaptation for Vision-Language Foundation Model
von: Ma, Huan, et al.
Veröffentlicht: (2024)
von: Ma, Huan, et al.
Veröffentlicht: (2024)
Harnessing Vision-Language Models for Time Series Anomaly Detection
von: He, Zelin, et al.
Veröffentlicht: (2025)
von: He, Zelin, et al.
Veröffentlicht: (2025)
A-VL: Adaptive Attention for Large Vision-Language Models
von: Zhang, Junyang, et al.
Veröffentlicht: (2024)
von: Zhang, Junyang, et al.
Veröffentlicht: (2024)
How Culturally Aware are Vision-Language Models?
von: Burda-Lassen, Olena, et al.
Veröffentlicht: (2024)
von: Burda-Lassen, Olena, et al.
Veröffentlicht: (2024)
Adaptive Prompt Tuning: Vision Guided Prompt Tuning with Cross-Attention for Fine-Grained Few-Shot Learning
von: Brouwer, Eric, et al.
Veröffentlicht: (2024)
von: Brouwer, Eric, et al.
Veröffentlicht: (2024)
Collaborative Temporal Feature Generation via Critic-Free Reinforcement Learning for Cross-User Sensor-Based Activity Recognition
von: Ye, Xiaozhou, et al.
Veröffentlicht: (2026)
von: Ye, Xiaozhou, et al.
Veröffentlicht: (2026)
Exploring Curriculum Learning for Vision-Language Tasks: A Study on Small-Scale Multimodal Training
von: Saha, Rohan, et al.
Veröffentlicht: (2024)
von: Saha, Rohan, et al.
Veröffentlicht: (2024)
Collision-Aware Vision-Language Learning for End-to-End Driving with Multimodal Infraction Datasets
von: Koran, Alex, et al.
Veröffentlicht: (2026)
von: Koran, Alex, et al.
Veröffentlicht: (2026)
Mitigating Object Hallucinations in Vision-Language Models through Region-Aware Attention Recalibration
von: Xu, Yuanzhi, et al.
Veröffentlicht: (2026)
von: Xu, Yuanzhi, et al.
Veröffentlicht: (2026)
Evaluating Hallucination in Large Vision-Language Models based on Context-Aware Object Similarities
von: Datta, Shounak, et al.
Veröffentlicht: (2025)
von: Datta, Shounak, et al.
Veröffentlicht: (2025)
BUSTR: Breast Ultrasound Text Reporting with a Descriptor-Aware Vision-Language Model
von: Mohammed, Rawa, et al.
Veröffentlicht: (2025)
von: Mohammed, Rawa, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Cross-Domain Generalization Limits of Vision Foundation Models in Facial Deepfake Detection
von: Delibasoglu, Ibrahim
Veröffentlicht: (2026) -
Anomaly-Aware Vision-Language Adapters for Zero-Shot Anomaly Detection
von: Aqeel, Muhammad, et al.
Veröffentlicht: (2026) -
Multi-scale Quaternion CNN and BiGRU with Cross Self-attention Feature Fusion for Fault Diagnosis of Bearing
von: Liu, Huanbai, et al.
Veröffentlicht: (2024) -
MFAF: An EVA02-Based Multi-scale Frequency Attention Fusion Method for Cross-View Geo-Localization
von: Liu, YiTong, et al.
Veröffentlicht: (2025) -
NVIDIA Nemotron Nano V2 VL
von: NVIDIA, et al.
Veröffentlicht: (2025)