BREEN: Bridge Data-Efficient Encoder-Free Multimodal Learning with Learnable Queries
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Tianle, Rao, Yongming, Hu, Winston, Cheng, Yu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Efficient Encoder-Free Fourier-based 3D Large Multimodal Model
von: Mei, Guofeng, et al.
Veröffentlicht: (2026)
von: Mei, Guofeng, et al.
Veröffentlicht: (2026)
MM-IQ: Benchmarking Human-Like Abstraction and Reasoning in Multimodal Models
von: Cai, Huanqia, et al.
Veröffentlicht: (2025)
von: Cai, Huanqia, et al.
Veröffentlicht: (2025)
Investigating Redundancy in Multimodal Large Language Models with Multiple Vision Encoders
von: Wang, Yizhou, et al.
Veröffentlicht: (2025)
von: Wang, Yizhou, et al.
Veröffentlicht: (2025)
ScalSelect: Scalable Training-Free Multimodal Data Selection for Efficient Visual Instruction Tuning
von: Wu, Changti, et al.
Veröffentlicht: (2026)
von: Wu, Changti, et al.
Veröffentlicht: (2026)
Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation
von: Zhang, Jihai, et al.
Veröffentlicht: (2025)
von: Zhang, Jihai, et al.
Veröffentlicht: (2025)
Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning
von: Ou, Siqu, et al.
Veröffentlicht: (2025)
von: Ou, Siqu, et al.
Veröffentlicht: (2025)
Insight-V++: Towards Advanced Long-Chain Visual Reasoning with Multimodal Large Language Models
von: Dong, Yuhao, et al.
Veröffentlicht: (2026)
von: Dong, Yuhao, et al.
Veröffentlicht: (2026)
Zoom-shot: Fast and Efficient Unsupervised Zero-Shot Transfer of CLIP to Vision Encoders with Multimodal Loss
von: Shipard, Jordan, et al.
Veröffentlicht: (2024)
von: Shipard, Jordan, et al.
Veröffentlicht: (2024)
Q-Zoom: Query-Aware Adaptive Perception for Efficient Multimodal Large Language Models
von: Shi, Yuheng, et al.
Veröffentlicht: (2026)
von: Shi, Yuheng, et al.
Veröffentlicht: (2026)
Efficient Learnable Collaborative Attention for Single Image Super-Resolution
von: Zheng, Yigang Zhao Chaowei, et al.
Veröffentlicht: (2024)
von: Zheng, Yigang Zhao Chaowei, et al.
Veröffentlicht: (2024)
Visual Encoders for Data-Efficient Imitation Learning in Modern Video Games
von: Schäfer, Lukas, et al.
Veröffentlicht: (2023)
von: Schäfer, Lukas, et al.
Veröffentlicht: (2023)
Localizing Events in Videos with Multimodal Queries
von: Zhang, Gengyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Gengyuan, et al.
Veröffentlicht: (2024)
EVEv2: Improved Baselines for Encoder-Free Vision-Language Models
von: Diao, Haiwen, et al.
Veröffentlicht: (2025)
von: Diao, Haiwen, et al.
Veröffentlicht: (2025)
FROST-Drive: Scalable and Efficient End-to-End Driving with a Frozen Vision Encoder
von: Dong, Zeyu, et al.
Veröffentlicht: (2026)
von: Dong, Zeyu, et al.
Veröffentlicht: (2026)
Explaining How Visual, Textual and Multimodal Encoders Share Concepts
von: Cornet, Clément, et al.
Veröffentlicht: (2025)
von: Cornet, Clément, et al.
Veröffentlicht: (2025)
QMoP: Query Guided Mixture-of-Projector for Efficient Visual Token Compression
von: Li, Zhongyang, et al.
Veröffentlicht: (2026)
von: Li, Zhongyang, et al.
Veröffentlicht: (2026)
CtrlSynth: Controllable Image Text Synthesis for Data-Efficient Multimodal Learning
von: Cao, Qingqing, et al.
Veröffentlicht: (2024)
von: Cao, Qingqing, et al.
Veröffentlicht: (2024)
Data-Free Group-Wise Fully Quantized Winograd Convolution via Learnable Scales
von: Pan, Shuokai, et al.
Veröffentlicht: (2024)
von: Pan, Shuokai, et al.
Veröffentlicht: (2024)
Fusion to Enhance: Fusion Visual Encoder to Enhance Multimodal Language Model
von: She, Yifei, et al.
Veröffentlicht: (2025)
von: She, Yifei, et al.
Veröffentlicht: (2025)
SAEN-BGS: Energy-Efficient Spiking AutoEncoder Network for Background Subtraction
von: Zhang, Zhixuan, et al.
Veröffentlicht: (2025)
von: Zhang, Zhixuan, et al.
Veröffentlicht: (2025)
Efficient Egocentric Action Recognition with Multimodal Data
von: Calzavara, Marco, et al.
Veröffentlicht: (2025)
von: Calzavara, Marco, et al.
Veröffentlicht: (2025)
PathFound: An Agentic Multimodal Model Activating Evidence-seeking Pathological Diagnosis
von: Hua, Shengyi, et al.
Veröffentlicht: (2025)
von: Hua, Shengyi, et al.
Veröffentlicht: (2025)
DocSAM: Unified Document Image Segmentation via Query Decomposition and Heterogeneous Mixed Learning
von: Li, Xiao-Hui, et al.
Veröffentlicht: (2025)
von: Li, Xiao-Hui, et al.
Veröffentlicht: (2025)
Category Query Learning for Human-Object Interaction Classification
von: Xie, Chi, et al.
Veröffentlicht: (2023)
von: Xie, Chi, et al.
Veröffentlicht: (2023)
BridgeDiff: Bridging Human Observations and Flat-Garment Synthesis for Virtual Try-Off
von: Liu, Shuang, et al.
Veröffentlicht: (2026)
von: Liu, Shuang, et al.
Veröffentlicht: (2026)
DaMO: A Data-Efficient Multimodal Orchestrator for Temporal Reasoning with Video LLMs
von: Chiu, Bo-Cheng, et al.
Veröffentlicht: (2025)
von: Chiu, Bo-Cheng, et al.
Veröffentlicht: (2025)
Single Image Reflection Separation via Dual Prior Interaction Transformer
von: Huang, Yue, et al.
Veröffentlicht: (2025)
von: Huang, Yue, et al.
Veröffentlicht: (2025)
M$^2$IV: Towards Efficient and Fine-grained Multimodal In-Context Learning via Representation Engineering
von: Li, Yanshu, et al.
Veröffentlicht: (2025)
von: Li, Yanshu, et al.
Veröffentlicht: (2025)
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion
von: Chen, Zhuokun, et al.
Veröffentlicht: (2024)
von: Chen, Zhuokun, et al.
Veröffentlicht: (2024)
LFTR: Learning-Free Token Reduction for Multimodal Large Language Models
von: Zhao, Zihui, et al.
Veröffentlicht: (2025)
von: Zhao, Zihui, et al.
Veröffentlicht: (2025)
Dynamic Object Queries for Transformer-based Incremental Object Detection
von: Zhang, Jichuan, et al.
Veröffentlicht: (2024)
von: Zhang, Jichuan, et al.
Veröffentlicht: (2024)
Mamba-VMR: Multimodal Query Augmentation via Generated Videos for Precise Temporal Grounding
von: Sun, Yunzhuo, et al.
Veröffentlicht: (2026)
von: Sun, Yunzhuo, et al.
Veröffentlicht: (2026)
VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding
von: Lin, Kuanwei, et al.
Veröffentlicht: (2026)
von: Lin, Kuanwei, et al.
Veröffentlicht: (2026)
Efficient Low-Resolution Face Recognition via Bridge Distillation
von: Ge, Shiming, et al.
Veröffentlicht: (2024)
von: Ge, Shiming, et al.
Veröffentlicht: (2024)
SEATrack: Simple, Efficient, and Adaptive Multimodal Tracker
von: Su, Junbin, et al.
Veröffentlicht: (2026)
von: Su, Junbin, et al.
Veröffentlicht: (2026)
Learning to Pose Problems: Reasoning-Driven and Solver-Adaptive Data Synthesis
von: Wei, Yongxian, et al.
Veröffentlicht: (2025)
von: Wei, Yongxian, et al.
Veröffentlicht: (2025)
MIRe: Enhancing Multimodal Queries Representation via Fusion-Free Modality Interaction for Multimodal Retrieval
von: Ju, Yeong-Joon, et al.
Veröffentlicht: (2024)
von: Ju, Yeong-Joon, et al.
Veröffentlicht: (2024)
Efficient Learning for Product Attributes with Compact Multimodal Models
von: Kulkarni, Mandar
Veröffentlicht: (2025)
von: Kulkarni, Mandar
Veröffentlicht: (2025)
Teacher Encoder-Student Decoder Denoising Guided Segmentation Network for Anomaly Detection
von: Song, Shixuan, et al.
Veröffentlicht: (2025)
von: Song, Shixuan, et al.
Veröffentlicht: (2025)
Towards Label-Free Brain Tumor Segmentation: Unsupervised Learning with Multimodal MRI
von: Comas-Quiles, Gerard, et al.
Veröffentlicht: (2025)
von: Comas-Quiles, Gerard, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Efficient Encoder-Free Fourier-based 3D Large Multimodal Model
von: Mei, Guofeng, et al.
Veröffentlicht: (2026) -
MM-IQ: Benchmarking Human-Like Abstraction and Reasoning in Multimodal Models
von: Cai, Huanqia, et al.
Veröffentlicht: (2025) -
Investigating Redundancy in Multimodal Large Language Models with Multiple Vision Encoders
von: Wang, Yizhou, et al.
Veröffentlicht: (2025) -
ScalSelect: Scalable Training-Free Multimodal Data Selection for Efficient Visual Instruction Tuning
von: Wu, Changti, et al.
Veröffentlicht: (2026) -
Are Unified Vision-Language Models Necessary: Generalization Across Understanding and Generation
von: Zhang, Jihai, et al.
Veröffentlicht: (2025)