MOVE: A Mixture-of-Vision-Encoders Approach for Domain-Focused Vision-Language Processing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Skripkin, Matvey, Goncharova, Elizaveta, Tarasov, Dmitrii, Kuznetsov, Andrey |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
OmniFusion Technical Report
von: Goncharova, Elizaveta, et al.
Veröffentlicht: (2024)
von: Goncharova, Elizaveta, et al.
Veröffentlicht: (2024)
Context-Dependent Affordance Computation in Vision-Language Models
von: Farzulla, Murad
Veröffentlicht: (2026)
von: Farzulla, Murad
Veröffentlicht: (2026)
A Hybrid Vision-Language Architecture for Automated Defect Reasoning and Report Generation in Industrial Inspection
von: Malikussaid, et al.
Veröffentlicht: (2026)
von: Malikussaid, et al.
Veröffentlicht: (2026)
Vision-based Situational Graphs Exploiting Fiducial Markers for the Integration of Semantic Entities
von: Tourani, Ali, et al.
Veröffentlicht: (2023)
von: Tourani, Ali, et al.
Veröffentlicht: (2023)
Image Reconstruction as a Tool for Feature Analysis
von: Allakhverdov, Eduard, et al.
Veröffentlicht: (2025)
von: Allakhverdov, Eduard, et al.
Veröffentlicht: (2025)
TowerVision: Understanding and Improving Multilinguality in Vision-Language Models
von: Viveiros, André G., et al.
Veröffentlicht: (2025)
von: Viveiros, André G., et al.
Veröffentlicht: (2025)
ABot-Claw: A Foundation for Persistent, Cooperative, and Self-Evolving Robotic Agents
von: Huo, Dongjie, et al.
Veröffentlicht: (2026)
von: Huo, Dongjie, et al.
Veröffentlicht: (2026)
A Survey on Vision-Language-Action Models for Embodied AI
von: Ma, Yueen, et al.
Veröffentlicht: (2024)
von: Ma, Yueen, et al.
Veröffentlicht: (2024)
StratXplore: Strategic Novelty-seeking and Instruction-aligned Exploration for Vision and Language Navigation
von: Gopinathan, Muraleekrishna, et al.
Veröffentlicht: (2024)
von: Gopinathan, Muraleekrishna, et al.
Veröffentlicht: (2024)
Spatially-Aware Speaker for Vision-and-Language Navigation Instruction Generation
von: Gopinathan, Muraleekrishna, et al.
Veröffentlicht: (2024)
von: Gopinathan, Muraleekrishna, et al.
Veröffentlicht: (2024)
UAV-assisted Visual SLAM Generating Reconstructed 3D Scene Graphs in GPS-denied Environments
von: Radwan, Ahmed, et al.
Veröffentlicht: (2024)
von: Radwan, Ahmed, et al.
Veröffentlicht: (2024)
Parameter-efficient fine-tuning (PEFT) of Vision Foundation Models for Atypical Mitotic Figure Classification
von: Ramchandani, Lavish, et al.
Veröffentlicht: (2025)
von: Ramchandani, Lavish, et al.
Veröffentlicht: (2025)
MerNav: A Highly Generalizable Memory-Execute-Review Framework for Zero-Shot Object Goal Navigation
von: Qi, Dekang, et al.
Veröffentlicht: (2026)
von: Qi, Dekang, et al.
Veröffentlicht: (2026)
FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models
von: Fu, Tianyu, et al.
Veröffentlicht: (2024)
von: Fu, Tianyu, et al.
Veröffentlicht: (2024)
Multi-Agent Object Detection Framework Based on Raspberry Pi YOLO Detector and Slack-Ollama Natural Language Interface
von: Kalušev, Vladimir, et al.
Veröffentlicht: (2026)
von: Kalušev, Vladimir, et al.
Veröffentlicht: (2026)
FlexDoc: Parameterized Sampling for Diverse Multilingual Synthetic Documents for Training Document Understanding Models
von: Dua, Karan, et al.
Veröffentlicht: (2025)
von: Dua, Karan, et al.
Veröffentlicht: (2025)
K-MetBench: A Multi-Dimensional Benchmark for Fine-Grained Evaluation of Expert Reasoning, Locality, and Multimodality in Meteorology
von: Kim, Soyeon, et al.
Veröffentlicht: (2026)
von: Kim, Soyeon, et al.
Veröffentlicht: (2026)
Universal Adversarial Attack on Aligned Multimodal LLMs
von: Rahmatullaev, Temurbek, et al.
Veröffentlicht: (2025)
von: Rahmatullaev, Temurbek, et al.
Veröffentlicht: (2025)
MORQA: Benchmarking Evaluation Metrics for Medical Open-Ended Question Answering
von: Yim, Wen-wai, et al.
Veröffentlicht: (2025)
von: Yim, Wen-wai, et al.
Veröffentlicht: (2025)
ExpReS-VLA: Specializing Vision-Language-Action Models Through Experience Replay and Retrieval
von: Syed, Shahram Najam, et al.
Veröffentlicht: (2025)
von: Syed, Shahram Najam, et al.
Veröffentlicht: (2025)
Unpacking Hateful Memes: Presupposed Context and False Claims
von: Cai, Weibin, et al.
Veröffentlicht: (2025)
von: Cai, Weibin, et al.
Veröffentlicht: (2025)
Flex: End-to-End Text-Instructed Visual Navigation from Foundation Model Features
von: Chahine, Makram, et al.
Veröffentlicht: (2024)
von: Chahine, Makram, et al.
Veröffentlicht: (2024)
OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language Mapping
von: Li, Danyang, et al.
Veröffentlicht: (2025)
von: Li, Danyang, et al.
Veröffentlicht: (2025)
Taking Flight with Dialogue: Enabling Natural Language Control for PX4-based Drone Agent
von: Lim, Shoon Kit, et al.
Veröffentlicht: (2025)
von: Lim, Shoon Kit, et al.
Veröffentlicht: (2025)
SUGARCREPE++ Dataset: Vision-Language Model Sensitivity to Semantic and Lexical Alterations
von: Dumpala, Sri Harsha, et al.
Veröffentlicht: (2024)
von: Dumpala, Sri Harsha, et al.
Veröffentlicht: (2024)
Beyond RGB: Leveraging Vision Transformers for Thermal Weapon Segmentation
von: Kambhatla, Akhila, et al.
Veröffentlicht: (2025)
von: Kambhatla, Akhila, et al.
Veröffentlicht: (2025)
Autonomous Navigation and Collision Avoidance for Mobile Robots: Classification and Review
von: de Carvalho, Marcus Vinicius Leal, et al.
Veröffentlicht: (2024)
von: de Carvalho, Marcus Vinicius Leal, et al.
Veröffentlicht: (2024)
When Less is Enough: Adaptive Token Reduction for Efficient Image Representation
von: Allakhverdov, Eduard, et al.
Veröffentlicht: (2025)
von: Allakhverdov, Eduard, et al.
Veröffentlicht: (2025)
Waking Up Blind: Cold-Start Optimization of Supervision-Free Agentic Trajectories for Grounded Visual Perception
von: Bajpai, Ashutosh, et al.
Veröffentlicht: (2026)
von: Bajpai, Ashutosh, et al.
Veröffentlicht: (2026)
Clarification as Supervision: Reinforcement Learning for Vision-Language Interfaces
von: Gkountouras, John, et al.
Veröffentlicht: (2025)
von: Gkountouras, John, et al.
Veröffentlicht: (2025)
MVTamperBench: Evaluating Robustness of Vision-Language Models
von: Agarwal, Amit, et al.
Veröffentlicht: (2024)
von: Agarwal, Amit, et al.
Veröffentlicht: (2024)
ReaderLM-v2: Small Language Model for HTML to Markdown and JSON
von: Wang, Feng, et al.
Veröffentlicht: (2025)
von: Wang, Feng, et al.
Veröffentlicht: (2025)
Advanced Long-term Earth System Forecasting
von: Wu, Hao, et al.
Veröffentlicht: (2025)
von: Wu, Hao, et al.
Veröffentlicht: (2025)
Ego-Motion Aware Target Prediction Module for Robust Multi-Object Tracking
von: Mahdian, Navid, et al.
Veröffentlicht: (2024)
von: Mahdian, Navid, et al.
Veröffentlicht: (2024)
ThinkJEPA: Empowering Latent World Models with Large Vision-Language Reasoning Model
von: Zhang, Haichao, et al.
Veröffentlicht: (2026)
von: Zhang, Haichao, et al.
Veröffentlicht: (2026)
WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation
von: Niu, Yuwei, et al.
Veröffentlicht: (2025)
von: Niu, Yuwei, et al.
Veröffentlicht: (2025)
The MSR-Video to Text Dataset with Clean Annotations
von: Chen, Haoran, et al.
Veröffentlicht: (2021)
von: Chen, Haoran, et al.
Veröffentlicht: (2021)
Beyond RNNs: Benchmarking Attention-Based Image Captioning Models
von: Yanambakkam, Hemanth Teja, et al.
Veröffentlicht: (2025)
von: Yanambakkam, Hemanth Teja, et al.
Veröffentlicht: (2025)
RACAS: Controlling Diverse Robots With a Single Agentic System
von: Ashley, Dylan R., et al.
Veröffentlicht: (2026)
von: Ashley, Dylan R., et al.
Veröffentlicht: (2026)
ATAAT: Adaptive Threat-Aware Adversarial Tuning Framework against Backdoor Attacks on Vision-Language-Action Models
von: Chen, Kewei, et al.
Veröffentlicht: (2026)
von: Chen, Kewei, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
OmniFusion Technical Report
von: Goncharova, Elizaveta, et al.
Veröffentlicht: (2024) -
Context-Dependent Affordance Computation in Vision-Language Models
von: Farzulla, Murad
Veröffentlicht: (2026) -
A Hybrid Vision-Language Architecture for Automated Defect Reasoning and Report Generation in Industrial Inspection
von: Malikussaid, et al.
Veröffentlicht: (2026) -
Vision-based Situational Graphs Exploiting Fiducial Markers for the Integration of Semantic Entities
von: Tourani, Ali, et al.
Veröffentlicht: (2023) -
Image Reconstruction as a Tool for Feature Analysis
von: Allakhverdov, Eduard, et al.
Veröffentlicht: (2025)