Libra: Building Decoupled Vision System on Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xu, Yifan, Yang, Xiaoshan, Song, Yaguang, Xu, Changsheng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Pilot: Building the Federated Multimodal Instruction Tuning Framework
von: Xiong, Baochen, et al.
Veröffentlicht: (2025)
von: Xiong, Baochen, et al.
Veröffentlicht: (2025)
PivotMerge: Bridging Heterogeneous Multimodal Pre-training via Post-Alignment Model Merging
von: Shao, Zibo, et al.
Veröffentlicht: (2026)
von: Shao, Zibo, et al.
Veröffentlicht: (2026)
Towards Domain-Generalized Open-Vocabulary Object Detection: A Progressive Domain-invariant Cross-modal Alignment Method
von: Xu, Xiaoran, et al.
Veröffentlicht: (2026)
von: Xu, Xiaoran, et al.
Veröffentlicht: (2026)
A Comprehensive Review of Few-shot Action Recognition
von: Wanyan, Yuyang, et al.
Veröffentlicht: (2024)
von: Wanyan, Yuyang, et al.
Veröffentlicht: (2024)
OneRef: Unified One-tower Expression Grounding and Segmentation with Mask Referring Modeling
von: Xiao, Linhui, et al.
Veröffentlicht: (2024)
von: Xiao, Linhui, et al.
Veröffentlicht: (2024)
NavMorph: A Self-Evolving World Model for Vision-and-Language Navigation in Continuous Environments
von: Yao, Xuan, et al.
Veröffentlicht: (2025)
von: Yao, Xuan, et al.
Veröffentlicht: (2025)
Towards Visual Grounding: A Survey
von: Xiao, Linhui, et al.
Veröffentlicht: (2024)
von: Xiao, Linhui, et al.
Veröffentlicht: (2024)
HiVG: Hierarchical Multimodal Fine-grained Modulation for Visual Grounding
von: Xiao, Linhui, et al.
Veröffentlicht: (2024)
von: Xiao, Linhui, et al.
Veröffentlicht: (2024)
Modality-Collaborative Low-Rank Decomposers for Few-Shot Video Domain Adaptation
von: Wanyan, Yuyang, et al.
Veröffentlicht: (2025)
von: Wanyan, Yuyang, et al.
Veröffentlicht: (2025)
Hierarchical Dual-Subspace Decoupling for Continual Learning in Vision-Language Models
von: Qin, Mengxin, et al.
Veröffentlicht: (2026)
von: Qin, Mengxin, et al.
Veröffentlicht: (2026)
Fast-Slow Test-Time Adaptation for Online Vision-and-Language Navigation
von: Gao, Junyu, et al.
Veröffentlicht: (2023)
von: Gao, Junyu, et al.
Veröffentlicht: (2023)
CLIP-VG: Self-paced Curriculum Adapting of CLIP for Visual Grounding
von: Xiao, Linhui, et al.
Veröffentlicht: (2023)
von: Xiao, Linhui, et al.
Veröffentlicht: (2023)
Conjugated Semantic Pool Improves OOD Detection with Pre-trained Vision-Language Models
von: Chen, Mengyuan, et al.
Veröffentlicht: (2024)
von: Chen, Mengyuan, et al.
Veröffentlicht: (2024)
VL-Calibration: Decoupled Confidence Calibration for Large Vision-Language Models Reasoning
von: Xiao, Wenyi, et al.
Veröffentlicht: (2026)
von: Xiao, Wenyi, et al.
Veröffentlicht: (2026)
Text-Guided Attention is All You Need for Zero-Shot Robustness in Vision-Language Models
von: Yu, Lu, et al.
Veröffentlicht: (2024)
von: Yu, Lu, et al.
Veröffentlicht: (2024)
TCP:Textual-based Class-aware Prompt tuning for Visual-Language Model
von: Yao, Hantao, et al.
Veröffentlicht: (2023)
von: Yao, Hantao, et al.
Veröffentlicht: (2023)
General Covariant Action Modeling: Constructing Generalized Manifolds via Spatio-Temporal Decoupling
von: Lyu, Huaihai, et al.
Veröffentlicht: (2026)
von: Lyu, Huaihai, et al.
Veröffentlicht: (2026)
Decoupled Similarity for Task-Aware Token Pruning in Large Vision-Language Models
von: Ma, Kexin, et al.
Veröffentlicht: (2026)
von: Ma, Kexin, et al.
Veröffentlicht: (2026)
Spatiotemporal Decoupling for Efficient Vision-Based Occupancy Forecasting
von: Xu, Jingyi, et al.
Veröffentlicht: (2024)
von: Xu, Jingyi, et al.
Veröffentlicht: (2024)
VisualSimpleQA: A Benchmark for Decoupled Evaluation of Large Vision-Language Models in Fact-Seeking Question Answering
von: Wang, Yanling, et al.
Veröffentlicht: (2025)
von: Wang, Yanling, et al.
Veröffentlicht: (2025)
Fine-tuning Pre-trained Vision-Language Models in a Human-Annotation-Free Manner
von: Wang, Qian-Wei, et al.
Veröffentlicht: (2026)
von: Wang, Qian-Wei, et al.
Veröffentlicht: (2026)
Unveiling Typographic Deceptions: Insights of the Typographic Vulnerability in Large Vision-Language Model
von: Cheng, Hao, et al.
Veröffentlicht: (2024)
von: Cheng, Hao, et al.
Veröffentlicht: (2024)
DepthLM: Metric Depth From Vision Language Models
von: Cai, Zhipeng, et al.
Veröffentlicht: (2025)
von: Cai, Zhipeng, et al.
Veröffentlicht: (2025)
DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models
von: Yao, Linli, et al.
Veröffentlicht: (2024)
von: Yao, Linli, et al.
Veröffentlicht: (2024)
VEAttack: Downstream-agnostic Vision Encoder Attack against Large Vision Language Models
von: Mei, Hefei, et al.
Veröffentlicht: (2025)
von: Mei, Hefei, et al.
Veröffentlicht: (2025)
Jagle: Building a Large-Scale Japanese Multimodal Post-Training Dataset for Vision-Language Models
von: Sugiura, Issa, et al.
Veröffentlicht: (2026)
von: Sugiura, Issa, et al.
Veröffentlicht: (2026)
SEP: Self-Enhanced Prompt Tuning for Visual-Language Model
von: Yao, Hantao, et al.
Veröffentlicht: (2024)
von: Yao, Hantao, et al.
Veröffentlicht: (2024)
Language Guided Concept Bottleneck Models for Interpretable Continual Learning
von: Yu, Lu, et al.
Veröffentlicht: (2025)
von: Yu, Lu, et al.
Veröffentlicht: (2025)
Dual Cluster Contrastive learning for Object Re-Identification
von: Yao, Hantao, et al.
Veröffentlicht: (2021)
von: Yao, Hantao, et al.
Veröffentlicht: (2021)
Exploring the Distinctiveness and Fidelity of the Descriptions Generated by Large Vision-Language Models
von: Huang, Yuhang, et al.
Veröffentlicht: (2024)
von: Huang, Yuhang, et al.
Veröffentlicht: (2024)
Mitigating Hallucinations in Large Vision-Language Models via DPO: On-Policy Data Hold the Key
von: Yang, Zhihe, et al.
Veröffentlicht: (2025)
von: Yang, Zhihe, et al.
Veröffentlicht: (2025)
Unified Generation and Self-Verification for Vision-Language Models via Advantage Decoupled Preference Optimization
von: Qiu, Xinyu, et al.
Veröffentlicht: (2026)
von: Qiu, Xinyu, et al.
Veröffentlicht: (2026)
3D Vision and Language Pretraining with Large-Scale Synthetic Data
von: Yang, Dejie, et al.
Veröffentlicht: (2024)
von: Yang, Dejie, et al.
Veröffentlicht: (2024)
Q-VLM: Post-training Quantization for Large Vision-Language Models
von: Wang, Changyuan, et al.
Veröffentlicht: (2024)
von: Wang, Changyuan, et al.
Veröffentlicht: (2024)
MemLens: Benchmarking Multimodal Long-Term Memory in Large Vision-Language Models
von: Ren, Xiyu, et al.
Veröffentlicht: (2026)
von: Ren, Xiyu, et al.
Veröffentlicht: (2026)
MPCAR: Multi-Perspective Contextual Augmentation for Enhanced Visual Reasoning in Large Vision-Language Models
von: Rahman, Amirul, et al.
Veröffentlicht: (2025)
von: Rahman, Amirul, et al.
Veröffentlicht: (2025)
Decoupling the components of geometric understanding in Vision Language Models
von: Kosoy, Eliza, et al.
Veröffentlicht: (2025)
von: Kosoy, Eliza, et al.
Veröffentlicht: (2025)
Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization
von: Jin, Yang, et al.
Veröffentlicht: (2024)
von: Jin, Yang, et al.
Veröffentlicht: (2024)
Open-Vocabulary Semantic Segmentation with Uncertainty Alignment for Robotic Scene Understanding in Indoor Building Environments
von: Xu, Yifan, et al.
Veröffentlicht: (2025)
von: Xu, Yifan, et al.
Veröffentlicht: (2025)
FTibSuite: A Comprehensive Resource Suite for Tibetan Vision-Language Modeling
von: Xu, Guixian, et al.
Veröffentlicht: (2026)
von: Xu, Guixian, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Pilot: Building the Federated Multimodal Instruction Tuning Framework
von: Xiong, Baochen, et al.
Veröffentlicht: (2025) -
PivotMerge: Bridging Heterogeneous Multimodal Pre-training via Post-Alignment Model Merging
von: Shao, Zibo, et al.
Veröffentlicht: (2026) -
Towards Domain-Generalized Open-Vocabulary Object Detection: A Progressive Domain-invariant Cross-modal Alignment Method
von: Xu, Xiaoran, et al.
Veröffentlicht: (2026) -
A Comprehensive Review of Few-shot Action Recognition
von: Wanyan, Yuyang, et al.
Veröffentlicht: (2024) -
OneRef: Unified One-tower Expression Grounding and Segmentation with Mask Referring Modeling
von: Xiao, Linhui, et al.
Veröffentlicht: (2024)