Viper-F1: Fast and Fine-Grained Multimodal Understanding with Cross-Modal State-Space Modulation
Fuente:
arXiv
Saved in:
| Main Author: | Trinh, Quoc-Huy |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Firebolt-VL: Efficient Vision-Language Understanding with Cross-Modality Modulation
by: Trinh, Quoc-Huy, et al.
Published: (2026)
by: Trinh, Quoc-Huy, et al.
Published: (2026)
2DMamba: Efficient State Space Model for Image Representation with Applications on Giga-Pixel Whole Slide Image Classification
by: Zhang, Jingwei, et al.
Published: (2024)
by: Zhang, Jingwei, et al.
Published: (2024)
From Classification to Cross-Modal Understanding: Leveraging Vision-Language Models for Fine-Grained Renal Pathology
by: Guo, Zhenhao, et al.
Published: (2025)
by: Guo, Zhenhao, et al.
Published: (2025)
Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space
by: Trinh, Quoc-Huy, et al.
Published: (2026)
by: Trinh, Quoc-Huy, et al.
Published: (2026)
Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding
by: Kim, Namho, et al.
Published: (2025)
by: Kim, Namho, et al.
Published: (2025)
MOOZY: A Patient-First Foundation Model for Computational Pathology
by: Kotp, Yousef, et al.
Published: (2026)
by: Kotp, Yousef, et al.
Published: (2026)
MotionSight: Boosting Fine-Grained Motion Understanding in Multimodal LLMs
by: Du, Yipeng, et al.
Published: (2025)
by: Du, Yipeng, et al.
Published: (2025)
Advancing Fine-Grained Visual Understanding with Multi-Scale Alignment in Multi-Modal Models
by: Wang, Wei, et al.
Published: (2024)
by: Wang, Wei, et al.
Published: (2024)
Investigating Zero-Shot Diagnostic Pathology in Vision-Language Models with Efficient Prompt Design
by: Sharma, Vasudev, et al.
Published: (2025)
by: Sharma, Vasudev, et al.
Published: (2025)
PRS-Med: Position Reasoning Segmentation in Medical Imaging
by: Trinh, Quoc-Huy, et al.
Published: (2025)
by: Trinh, Quoc-Huy, et al.
Published: (2025)
Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space
by: Yu, Shiyao, et al.
Published: (2025)
by: Yu, Shiyao, et al.
Published: (2025)
RotCAtt-TransUNet++: Novel Deep Neural Network for Sophisticated Cardiac Segmentation
by: Nguyen-Le, Quoc-Bao, et al.
Published: (2024)
by: Nguyen-Le, Quoc-Bao, et al.
Published: (2024)
Simple Token-Efficient Vision-Language Model for Case-level Pathology Synoptic Report Generation
by: Yang, Zhiyuan, et al.
Published: (2026)
by: Yang, Zhiyuan, et al.
Published: (2026)
Probing Fine-Grained Action Understanding and Cross-View Generalization of Foundation Models
by: Ponbagavathi, Thinesh Thiyakesan, et al.
Published: (2024)
by: Ponbagavathi, Thinesh Thiyakesan, et al.
Published: (2024)
MMFuser: Multimodal Multi-Layer Feature Fuser for Fine-Grained Vision-Language Understanding
by: Cao, Yue, et al.
Published: (2024)
by: Cao, Yue, et al.
Published: (2024)
Cross Modal Fine-Grained Alignment via Granularity-Aware and Region-Uncertain Modeling
by: Liu, Jiale, et al.
Published: (2025)
by: Liu, Jiale, et al.
Published: (2025)
CoF: Coarse to Fine-Grained Image Understanding for Multi-modal Large Language Models
by: Wang, Yeyuan, et al.
Published: (2024)
by: Wang, Yeyuan, et al.
Published: (2024)
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding
by: Peng, Yi-Xing, et al.
Published: (2025)
by: Peng, Yi-Xing, et al.
Published: (2025)
Multimodal Contextualized Support for Enhancing Video Retrieval System
by: Nguyen-Le, Quoc-Bao, et al.
Published: (2024)
by: Nguyen-Le, Quoc-Bao, et al.
Published: (2024)
MC-MKE: A Fine-Grained Multimodal Knowledge Editing Benchmark Emphasizing Modality Consistency
by: Zhang, Junzhe, et al.
Published: (2024)
by: Zhang, Junzhe, et al.
Published: (2024)
VMambaMorph: a Multi-Modality Deformable Image Registration Framework based on Visual State Space Model with Cross-Scan Module
by: Wang, Ziyang, et al.
Published: (2024)
by: Wang, Ziyang, et al.
Published: (2024)
Mixture-of-Modality-Experts with Holistic Token Learning for Fine-Grained Multimodal Visual Analytics in Driver Action Recognition
by: Liu, Tianyi, et al.
Published: (2026)
by: Liu, Tianyi, et al.
Published: (2026)
FG$^2$: Fine-Grained Cross-View Localization by Fine-Grained Feature Matching
by: Xia, Zimin, et al.
Published: (2025)
by: Xia, Zimin, et al.
Published: (2025)
Fast-then-Fine: A Two-Stage Framework with Multi-Granular Representation for Cross-Modal Retrieval in Remote Sensing
by: Chen, Xi, et al.
Published: (2026)
by: Chen, Xi, et al.
Published: (2026)
LLaVA-ST: A Multimodal Large Language Model for Fine-Grained Spatial-Temporal Understanding
by: Li, Hongyu, et al.
Published: (2025)
by: Li, Hongyu, et al.
Published: (2025)
Fine-Grained Scene Image Classification with Modality-Agnostic Adapter
by: Wang, Yiqun, et al.
Published: (2024)
by: Wang, Yiqun, et al.
Published: (2024)
Learning by Aligning 2D Skeleton Sequences and Multi-Modality Fusion
by: Tran, Quoc-Huy, et al.
Published: (2023)
by: Tran, Quoc-Huy, et al.
Published: (2023)
Vim4Path: Self-Supervised Vision Mamba for Histopathology Images
by: Nasiri-Sarvi, Ali, et al.
Published: (2024)
by: Nasiri-Sarvi, Ali, et al.
Published: (2024)
FaNe: Towards Fine-Grained Cross-Modal Contrast with False-Negative Reduction and Text-Conditioned Sparse Attention
by: Zhang, Peng, et al.
Published: (2025)
by: Zhang, Peng, et al.
Published: (2025)
Interactive State Space Model with Cross-Modal Local Scanning for Depth Super-Resolution
by: Wu, Chen, et al.
Published: (2026)
by: Wu, Chen, et al.
Published: (2026)
NeIn: Telling What You Don't Want
by: Bui, Nhat-Tan, et al.
Published: (2024)
by: Bui, Nhat-Tan, et al.
Published: (2024)
TimeViper: A Hybrid Mamba-Transformer Vision-Language Model for Efficient Long Video Understanding
by: Xu, Boshen, et al.
Published: (2025)
by: Xu, Boshen, et al.
Published: (2025)
Text-guided Fine-Grained Video Anomaly Understanding
by: Gu, Jihao, et al.
Published: (2025)
by: Gu, Jihao, et al.
Published: (2025)
UniM-OV3D: Uni-Modality Open-Vocabulary 3D Scene Understanding with Fine-Grained Feature Representation
by: He, Qingdong, et al.
Published: (2024)
by: He, Qingdong, et al.
Published: (2024)
SAM-EG: Segment Anything Model with Egde Guidance framework for efficient Polyp Segmentation
by: Trinh, Quoc-Huy, et al.
Published: (2024)
by: Trinh, Quoc-Huy, et al.
Published: (2024)
Cross-Modal Guidance for Fast Diffusion-Based Computed Tomography
by: Efimov, Timofey, et al.
Published: (2026)
by: Efimov, Timofey, et al.
Published: (2026)
OmniMamba: Efficient and Unified Multimodal Understanding and Generation via State Space Models
by: Zou, Jialv, et al.
Published: (2025)
by: Zou, Jialv, et al.
Published: (2025)
Rethinking Cross-Modal Interaction in Multimodal Diffusion Transformers
by: Lv, Zhengyao, et al.
Published: (2025)
by: Lv, Zhengyao, et al.
Published: (2025)
KDAS: Knowledge Distillation via Attention Supervision Framework for Polyp Segmentation
by: Trinh, Quoc-Huy, et al.
Published: (2023)
by: Trinh, Quoc-Huy, et al.
Published: (2023)
FineState-Bench: A Comprehensive Benchmark for Fine-Grained State Control in GUI Agents
by: Ji, Fengxian, et al.
Published: (2025)
by: Ji, Fengxian, et al.
Published: (2025)
Similar Items
-
Firebolt-VL: Efficient Vision-Language Understanding with Cross-Modality Modulation
by: Trinh, Quoc-Huy, et al.
Published: (2026) -
2DMamba: Efficient State Space Model for Image Representation with Applications on Giga-Pixel Whole Slide Image Classification
by: Zhang, Jingwei, et al.
Published: (2024) -
From Classification to Cross-Modal Understanding: Leveraging Vision-Language Models for Fine-Grained Renal Pathology
by: Guo, Zhenhao, et al.
Published: (2025) -
Beyond Medical Diagnostics: How Medical Multimodal Large Language Models Think in Space
by: Trinh, Quoc-Huy, et al.
Published: (2026) -
Multimodal Alignment with Cross-Attentive GRUs for Fine-Grained Video Understanding
by: Kim, Namho, et al.
Published: (2025)