EffiMiniVLM: A Compact Dual-Encoder Regression Framework
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Khor, Yin-Loon, Wong, Yi-Jie, Hum, Yan Chai |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
EffiPerception: an Efficient Framework for Various Perception Tasks
von: Xiang, Xinhao, et al.
Veröffentlicht: (2024)
von: Xiang, Xinhao, et al.
Veröffentlicht: (2024)
EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models
von: Wang, Zekun, et al.
Veröffentlicht: (2025)
von: Wang, Zekun, et al.
Veröffentlicht: (2025)
REO-VLM: Transforming VLM to Meet Regression Challenges in Earth Observation
von: Xue, Xizhe, et al.
Veröffentlicht: (2024)
von: Xue, Xizhe, et al.
Veröffentlicht: (2024)
DAVE: A VLM Vision Encoder for Document Understanding and Web Agents
von: Huang, Brandon, et al.
Veröffentlicht: (2025)
von: Huang, Brandon, et al.
Veröffentlicht: (2025)
GeoDANO: Geometric VLM with Domain Agnostic Vision Encoder
von: Cho, Seunghyuk, et al.
Veröffentlicht: (2025)
von: Cho, Seunghyuk, et al.
Veröffentlicht: (2025)
Penguin-VL: Exploring the Efficiency Limits of VLM with LLM-based Vision Encoders
von: Zhang, Boqiang, et al.
Veröffentlicht: (2026)
von: Zhang, Boqiang, et al.
Veröffentlicht: (2026)
EffiVED:Efficient Video Editing via Text-instruction Diffusion Models
von: Zhang, Zhenghao, et al.
Veröffentlicht: (2024)
von: Zhang, Zhenghao, et al.
Veröffentlicht: (2024)
Bodhi VLM: Privacy-Alignment Modeling for Hierarchical Visual Representations in Vision Backbones and VLM Encoders via Bottom-Up and Top-Down Feature Search
von: Ma, Bo, et al.
Veröffentlicht: (2026)
von: Ma, Bo, et al.
Veröffentlicht: (2026)
Dual Associated Encoder for Face Restoration
von: Tsai, Yu-Ju, et al.
Veröffentlicht: (2023)
von: Tsai, Yu-Ju, et al.
Veröffentlicht: (2023)
MoiréNet: A Compact Dual-Domain Network for Image Demoiréing
von: Guo, Shuwei, et al.
Veröffentlicht: (2025)
von: Guo, Shuwei, et al.
Veröffentlicht: (2025)
Siamese-Driven Optimization for Low-Resolution Image Latent Embedding in Image Captioning
von: Tan, Jing Jie, et al.
Veröffentlicht: (2025)
von: Tan, Jing Jie, et al.
Veröffentlicht: (2025)
Dual-Domain Perspective on Degradation-Aware Fusion: A VLM-Guided Robust Infrared and Visible Image Fusion Framework
von: Zhang, Tianpei, et al.
Veröffentlicht: (2025)
von: Zhang, Tianpei, et al.
Veröffentlicht: (2025)
EffiComm: Bandwidth Efficient Multi Agent Communication
von: Yazgan, Melih, et al.
Veröffentlicht: (2025)
von: Yazgan, Melih, et al.
Veröffentlicht: (2025)
Towards Comprehensive Interactive Change Understanding in Remote Sensing: A Large-scale Dataset and Dual-granularity Enhanced VLM
von: Xue, Junxiao, et al.
Veröffentlicht: (2025)
von: Xue, Junxiao, et al.
Veröffentlicht: (2025)
DUET-VLM: Dual stage Unified Efficient Token reduction for VLM Training and Inference
von: Singh, Aditya Kumar, et al.
Veröffentlicht: (2026)
von: Singh, Aditya Kumar, et al.
Veröffentlicht: (2026)
A Dual Process VLA: Efficient Robotic Manipulation Leveraging VLM
von: Han, ByungOk, et al.
Veröffentlicht: (2024)
von: Han, ByungOk, et al.
Veröffentlicht: (2024)
DuMo: Dual Encoder Modulation Network for Precise Concept Erasure
von: Han, Feng, et al.
Veröffentlicht: (2025)
von: Han, Feng, et al.
Veröffentlicht: (2025)
SEDS: Semantically Enhanced Dual-Stream Encoder for Sign Language Retrieval
von: Jiang, Longtao, et al.
Veröffentlicht: (2024)
von: Jiang, Longtao, et al.
Veröffentlicht: (2024)
Vehicle Detection Performance in Nordic Region
von: Mokayed, Hamam, et al.
Veröffentlicht: (2024)
von: Mokayed, Hamam, et al.
Veröffentlicht: (2024)
CogVLM: Visual Expert for Pretrained Language Models
von: Wang, Weihan, et al.
Veröffentlicht: (2023)
von: Wang, Weihan, et al.
Veröffentlicht: (2023)
Graph Domain Adaptation with Dual-branch Encoder and Two-level Alignment for Whole Slide Image-based Survival Prediction
von: Shou, Yuntao, et al.
Veröffentlicht: (2024)
von: Shou, Yuntao, et al.
Veröffentlicht: (2024)
OmniSAT: Compact Action Token, Faster Auto Regression
von: Lyu, Huaihai, et al.
Veröffentlicht: (2025)
von: Lyu, Huaihai, et al.
Veröffentlicht: (2025)
Benchmarking and Enhancing VLM for Compressed Image Understanding
von: Zhang, Zifu, et al.
Veröffentlicht: (2025)
von: Zhang, Zifu, et al.
Veröffentlicht: (2025)
Dual-Prompt CLIP with Hybrid Visual Encoders for Occluded Person Re-Identification
von: Ji, Zhangjian, et al.
Veröffentlicht: (2026)
von: Ji, Zhangjian, et al.
Veröffentlicht: (2026)
AGE-Net: Spectral--Spatial Fusion and Anatomical Graph Reasoning with Evidential Ordinal Regression for Knee Osteoarthritis Grading
von: Li, Xiaoyang, et al.
Veröffentlicht: (2026)
von: Li, Xiaoyang, et al.
Veröffentlicht: (2026)
FDCE-Net: Underwater Image Enhancement with Embedding Frequency and Dual Color Encoder
von: Cheng, Zheng, et al.
Veröffentlicht: (2024)
von: Cheng, Zheng, et al.
Veröffentlicht: (2024)
Contrastive Pretraining with Dual Visual Encoders for Gloss-Free Sign Language Translation
von: Sincan, Ozge Mercanoglu, et al.
Veröffentlicht: (2025)
von: Sincan, Ozge Mercanoglu, et al.
Veröffentlicht: (2025)
From Representational Complementarity to Dual Systems: Synergizing VLM and Vision-Only Backbones for End-to-End Driving
von: Ang, Sining, et al.
Veröffentlicht: (2026)
von: Ang, Sining, et al.
Veröffentlicht: (2026)
Precision Synthesis of Multi-Tracer PET via VLM-Modulated Rectified Flow for Stratifying Mild Cognitive Impairment
von: Liu, Tuo, et al.
Veröffentlicht: (2026)
von: Liu, Tuo, et al.
Veröffentlicht: (2026)
Is Micro-expression Ethnic Leaning?
von: Khor, Huai-Qian, et al.
Veröffentlicht: (2025)
von: Khor, Huai-Qian, et al.
Veröffentlicht: (2025)
Infused Suppression Of Magnification Artefacts For Micro-AU Detection
von: Khor, Huai-Qian, et al.
Veröffentlicht: (2025)
von: Khor, Huai-Qian, et al.
Veröffentlicht: (2025)
CogVLM2: Visual Language Models for Image and Video Understanding
von: Hong, Wenyi, et al.
Veröffentlicht: (2024)
von: Hong, Wenyi, et al.
Veröffentlicht: (2024)
MiniMax-Remover: Taming Bad Noise Helps Video Object Removal
von: Zi, Bojia, et al.
Veröffentlicht: (2025)
von: Zi, Bojia, et al.
Veröffentlicht: (2025)
Language-Image Alignment with Fixed Text Encoders
von: Yang, Jingfeng, et al.
Veröffentlicht: (2025)
von: Yang, Jingfeng, et al.
Veröffentlicht: (2025)
An Attentive Dual-Encoder Framework Leveraging Multimodal Visual and Semantic Information for Automatic OSAHS Diagnosis
von: Wei, Yingchen, et al.
Veröffentlicht: (2024)
von: Wei, Yingchen, et al.
Veröffentlicht: (2024)
TransVLM: A Vision-Language Framework and Benchmark for Detecting Any Shot Transitions
von: Chen, Ce, et al.
Veröffentlicht: (2026)
von: Chen, Ce, et al.
Veröffentlicht: (2026)
Slot-VLM: SlowFast Slots for Video-Language Modeling
von: Xu, Jiaqi, et al.
Veröffentlicht: (2024)
von: Xu, Jiaqi, et al.
Veröffentlicht: (2024)
Q-VLM: Post-training Quantization for Large Vision-Language Models
von: Wang, Changyuan, et al.
Veröffentlicht: (2024)
von: Wang, Changyuan, et al.
Veröffentlicht: (2024)
IMC-Net: A Lightweight Content-Conditioned Encoder with Multi-Pass Processing for Image Classification
von: Li, YiZhou
Veröffentlicht: (2025)
von: Li, YiZhou
Veröffentlicht: (2025)
VLM-KD: Knowledge Distillation from VLM for Long-Tail Visual Recognition
von: Zhang, Zaiwei, et al.
Veröffentlicht: (2024)
von: Zhang, Zaiwei, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
EffiPerception: an Efficient Framework for Various Perception Tasks
von: Xiang, Xinhao, et al.
Veröffentlicht: (2024) -
EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models
von: Wang, Zekun, et al.
Veröffentlicht: (2025) -
REO-VLM: Transforming VLM to Meet Regression Challenges in Earth Observation
von: Xue, Xizhe, et al.
Veröffentlicht: (2024) -
DAVE: A VLM Vision Encoder for Document Understanding and Web Agents
von: Huang, Brandon, et al.
Veröffentlicht: (2025) -
GeoDANO: Geometric VLM with Domain Agnostic Vision Encoder
von: Cho, Seunghyuk, et al.
Veröffentlicht: (2025)