Firebolt-VL: Efficient Vision-Language Understanding with Cross-Modality Modulation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Trinh, Quoc-Huy, Abdullahi, Mustapha, Zhao, Bo, Jha, Debesh |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Viper-F1: Fast and Fine-Grained Multimodal Understanding with Cross-Modal State-Space Modulation
von: Trinh, Quoc-Huy
Veröffentlicht: (2025)
von: Trinh, Quoc-Huy
Veröffentlicht: (2025)
PRS-Med: Position Reasoning Segmentation in Medical Imaging
von: Trinh, Quoc-Huy, et al.
Veröffentlicht: (2025)
von: Trinh, Quoc-Huy, et al.
Veröffentlicht: (2025)
SAM-EG: Segment Anything Model with Egde Guidance framework for efficient Polyp Segmentation
von: Trinh, Quoc-Huy, et al.
Veröffentlicht: (2024)
von: Trinh, Quoc-Huy, et al.
Veröffentlicht: (2024)
Investigating Zero-Shot Diagnostic Pathology in Vision-Language Models with Efficient Prompt Design
von: Sharma, Vasudev, et al.
Veröffentlicht: (2025)
von: Sharma, Vasudev, et al.
Veröffentlicht: (2025)
Simple Token-Efficient Vision-Language Model for Case-level Pathology Synoptic Report Generation
von: Yang, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Yang, Zhiyuan, et al.
Veröffentlicht: (2026)
SDAR-VL: Stable and Efficient Block-wise Diffusion for Vision-Language Understanding
von: Cheng, Shuang, et al.
Veröffentlicht: (2025)
von: Cheng, Shuang, et al.
Veröffentlicht: (2025)
HaploVL: A Single-Transformer Baseline for Multi-Modal Understanding
von: Yang, Rui, et al.
Veröffentlicht: (2025)
von: Yang, Rui, et al.
Veröffentlicht: (2025)
PGDS: Pose-Guidance Deep Supervision for Mitigating Clothes-Changing in Person Re-Identification
von: Trinh, Quoc-Huy, et al.
Veröffentlicht: (2023)
von: Trinh, Quoc-Huy, et al.
Veröffentlicht: (2023)
BcQLM: Efficient Vision-Language Understanding with Distilled Q-Gated Cross-Modal Fusion
von: Xiang, Sike, et al.
Veröffentlicht: (2025)
von: Xiang, Sike, et al.
Veröffentlicht: (2025)
ADEM-VL: Adaptive and Embedded Fusion for Efficient Vision-Language Tuning
von: Hao, Zhiwei, et al.
Veröffentlicht: (2024)
von: Hao, Zhiwei, et al.
Veröffentlicht: (2024)
GraphVL: Graph-Enhanced Semantic Modeling via Vision-Language Models for Generalized Class Discovery
von: Solanki, Bhupendra, et al.
Veröffentlicht: (2024)
von: Solanki, Bhupendra, et al.
Veröffentlicht: (2024)
TinyChemVL: Advancing Chemical Vision-Language Models via Efficient Visual Token Reduction and Complex Reaction Tasks
von: Zhao, Xuanle, et al.
Veröffentlicht: (2025)
von: Zhao, Xuanle, et al.
Veröffentlicht: (2025)
Distilling Vision-Language Pretraining for Efficient Cross-Modal Retrieval
von: Jang, Young Kyun, et al.
Veröffentlicht: (2024)
von: Jang, Young Kyun, et al.
Veröffentlicht: (2024)
MOOZY: A Patient-First Foundation Model for Computational Pathology
von: Kotp, Yousef, et al.
Veröffentlicht: (2026)
von: Kotp, Yousef, et al.
Veröffentlicht: (2026)
Vim4Path: Self-Supervised Vision Mamba for Histopathology Images
von: Nasiri-Sarvi, Ali, et al.
Veröffentlicht: (2024)
von: Nasiri-Sarvi, Ali, et al.
Veröffentlicht: (2024)
SRMA-Mamba: Spatial Reverse Mamba Attention Network for Pathological Liver Segmentation in MRI Volumes
von: Zeng, Jun, et al.
Veröffentlicht: (2025)
von: Zeng, Jun, et al.
Veröffentlicht: (2025)
FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding
von: Liu, Zheng, et al.
Veröffentlicht: (2025)
von: Liu, Zheng, et al.
Veröffentlicht: (2025)
LEO-VL: Efficient Scene Representation for Scalable 3D Vision-Language Learning
von: Huang, Jiangyong, et al.
Veröffentlicht: (2025)
von: Huang, Jiangyong, et al.
Veröffentlicht: (2025)
VEN-VL: A Visual Ensemble MoE Framework for Effective and Efficient Multi-Modal Understanding
von: Wu, Yinghao, et al.
Veröffentlicht: (2026)
von: Wu, Yinghao, et al.
Veröffentlicht: (2026)
Rice-VL: Evaluating Vision-Language Models for Cultural Understanding Across ASEAN Countries
von: Pranav, Tushar, et al.
Veröffentlicht: (2025)
von: Pranav, Tushar, et al.
Veröffentlicht: (2025)
EarthVL: A Progressive Earth Vision-Language Understanding and Generation Framework
von: Wang, Junjue, et al.
Veröffentlicht: (2026)
von: Wang, Junjue, et al.
Veröffentlicht: (2026)
ZipVL: Efficient Large Vision-Language Models with Dynamic Token Sparsification
von: He, Yefei, et al.
Veröffentlicht: (2024)
von: He, Yefei, et al.
Veröffentlicht: (2024)
Youtu-VL: Unleashing Visual Potential via Unified Vision-Language Supervision
von: Wei, Zhixiang, et al.
Veröffentlicht: (2026)
von: Wei, Zhixiang, et al.
Veröffentlicht: (2026)
Thermo-VL: Extending Vision-Language Models to Thermal Infrared Perception
von: Thushara, Rusiru, et al.
Veröffentlicht: (2026)
von: Thushara, Rusiru, et al.
Veröffentlicht: (2026)
VL-Reader: Vision and Language Reconstructor is an Effective Scene Text Recognizer
von: Zhong, Humen, et al.
Veröffentlicht: (2024)
von: Zhong, Humen, et al.
Veröffentlicht: (2024)
3VL: Using Trees to Improve Vision-Language Models' Interpretability
von: Yellinek, Nir, et al.
Veröffentlicht: (2023)
von: Yellinek, Nir, et al.
Veröffentlicht: (2023)
VL4Gaze: Unleashing Vision-Language Models for Gaze Following
von: Wang, Shijing, et al.
Veröffentlicht: (2025)
von: Wang, Shijing, et al.
Veröffentlicht: (2025)
ConsensusDrop: Fusing Visual and Cross-Modal Saliency for Efficient Vision Language Models
von: Parikh, Dhruv, et al.
Veröffentlicht: (2026)
von: Parikh, Dhruv, et al.
Veröffentlicht: (2026)
From SAM to DINOv2: Towards Distilling Foundation Models to Lightweight Baselines for Generalized Polyp Segmentation
von: Agnihotri, Shivanshu, et al.
Veröffentlicht: (2025)
von: Agnihotri, Shivanshu, et al.
Veröffentlicht: (2025)
MeDocVL: A Visual Language Model for Medical Document Understanding and Parsing
von: Wang, Wenjie, et al.
Veröffentlicht: (2026)
von: Wang, Wenjie, et al.
Veröffentlicht: (2026)
RotCAtt-TransUNet++: Novel Deep Neural Network for Sophisticated Cardiac Segmentation
von: Nguyen-Le, Quoc-Bao, et al.
Veröffentlicht: (2024)
von: Nguyen-Le, Quoc-Bao, et al.
Veröffentlicht: (2024)
From Classification to Cross-Modal Understanding: Leveraging Vision-Language Models for Fine-Grained Renal Pathology
von: Guo, Zhenhao, et al.
Veröffentlicht: (2025)
von: Guo, Zhenhao, et al.
Veröffentlicht: (2025)
DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
von: Wu, Zhiyu, et al.
Veröffentlicht: (2024)
von: Wu, Zhiyu, et al.
Veröffentlicht: (2024)
SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding
von: Li, Hao, et al.
Veröffentlicht: (2024)
von: Li, Hao, et al.
Veröffentlicht: (2024)
DynamicVL: Benchmarking Multimodal Large Language Models for Dynamic City Understanding
von: Xuan, Weihao, et al.
Veröffentlicht: (2025)
von: Xuan, Weihao, et al.
Veröffentlicht: (2025)
2DMamba: Efficient State Space Model for Image Representation with Applications on Giga-Pixel Whole Slide Image Classification
von: Zhang, Jingwei, et al.
Veröffentlicht: (2024)
von: Zhang, Jingwei, et al.
Veröffentlicht: (2024)
UMind-VL: A Generalist Ultrasound Vision-Language Model for Unified Grounded Perception and Comprehensive Interpretation
von: Chen, Dengbo, et al.
Veröffentlicht: (2025)
von: Chen, Dengbo, et al.
Veröffentlicht: (2025)
In the Era of Prompt Learning with Vision-Language Models
von: Jha, Ankit
Veröffentlicht: (2024)
von: Jha, Ankit
Veröffentlicht: (2024)
CoME-VL: Scaling Complementary Multi-Encoder Vision-Language Learning
von: Deria, Ankan, et al.
Veröffentlicht: (2026)
von: Deria, Ankan, et al.
Veröffentlicht: (2026)
DiffusionVL: Translating Any Autoregressive Models into Diffusion Vision Language Models
von: Zeng, Lunbin, et al.
Veröffentlicht: (2025)
von: Zeng, Lunbin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Viper-F1: Fast and Fine-Grained Multimodal Understanding with Cross-Modal State-Space Modulation
von: Trinh, Quoc-Huy
Veröffentlicht: (2025) -
PRS-Med: Position Reasoning Segmentation in Medical Imaging
von: Trinh, Quoc-Huy, et al.
Veröffentlicht: (2025) -
SAM-EG: Segment Anything Model with Egde Guidance framework for efficient Polyp Segmentation
von: Trinh, Quoc-Huy, et al.
Veröffentlicht: (2024) -
Investigating Zero-Shot Diagnostic Pathology in Vision-Language Models with Efficient Prompt Design
von: Sharma, Vasudev, et al.
Veröffentlicht: (2025) -
Simple Token-Efficient Vision-Language Model for Case-level Pathology Synoptic Report Generation
von: Yang, Zhiyuan, et al.
Veröffentlicht: (2026)