ImagebindDC: Compressing Multi-modal Data with Imagebind-based Condensation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Min, Yue, Wang, Shaobo, Li, Jiaze, Niu, Tianle, Fan, Junxin, Miao, Yongliang, Yang, Lijin, Zhang, Linfeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VideoCompressa: Data-Efficient Video Understanding via Joint Temporal Compression and Spatial Reconstruction
von: Wang, Shaobo, et al.
Veröffentlicht: (2025)
von: Wang, Shaobo, et al.
Veröffentlicht: (2025)
Efficient Multi-modal Large Language Models via Progressive Consistency Distillation
von: Wen, Zichen, et al.
Veröffentlicht: (2025)
von: Wen, Zichen, et al.
Veröffentlicht: (2025)
Efficient Large Multi-modal Models via Visual Context Compression
von: Chen, Jieneng, et al.
Veröffentlicht: (2024)
von: Chen, Jieneng, et al.
Veröffentlicht: (2024)
Socratic-Geo: Synthetic Data Generation and Geometric Reasoning via Multi-Agent Interaction
von: Jiao, Zhengbo, et al.
Veröffentlicht: (2026)
von: Jiao, Zhengbo, et al.
Veröffentlicht: (2026)
GaitMA: Pose-guided Multi-modal Feature Fusion for Gait Recognition
von: Min, Fanxu, et al.
Veröffentlicht: (2024)
von: Min, Fanxu, et al.
Veröffentlicht: (2024)
TemCoCo: Temporally Consistent Multi-modal Video Fusion with Visual-Semantic Collaboration
von: Gong, Meiqi, et al.
Veröffentlicht: (2025)
von: Gong, Meiqi, et al.
Veröffentlicht: (2025)
Compute Only 16 Tokens in One Timestep: Accelerating Diffusion Transformers with Cluster-Driven Feature Caching
von: Zheng, Zhixin, et al.
Veröffentlicht: (2025)
von: Zheng, Zhixin, et al.
Veröffentlicht: (2025)
DAMA: Data- and Model-aware Alignment of Multi-modal LLMs
von: Lu, Jinda, et al.
Veröffentlicht: (2025)
von: Lu, Jinda, et al.
Veröffentlicht: (2025)
LLMControl: Grounded Control of Text-to-Image Diffusion-based Synthesis with Multimodal LLMs
von: Wang, Jiaze, et al.
Veröffentlicht: (2025)
von: Wang, Jiaze, et al.
Veröffentlicht: (2025)
Mastering Collaborative Multi-modal Data Selection: A Focus on Informativeness, Uniqueness, and Representativeness
von: Yu, Qifan, et al.
Veröffentlicht: (2024)
von: Yu, Qifan, et al.
Veröffentlicht: (2024)
Instance Data Condensation for Image Super-Resolution
von: Peng, Tianhao, et al.
Veröffentlicht: (2025)
von: Peng, Tianhao, et al.
Veröffentlicht: (2025)
Implementing blind navigation through multi-modal sensing and gait guidance
von: Yan, Feifan, et al.
Veröffentlicht: (2025)
von: Yan, Feifan, et al.
Veröffentlicht: (2025)
Enhancing Incomplete Multi-modal Brain Tumor Segmentation with Intra-modal Asymmetry and Inter-modal Dependency
von: Liu, Weide, et al.
Veröffentlicht: (2024)
von: Liu, Weide, et al.
Veröffentlicht: (2024)
Unified Batch Normalization: Identifying and Alleviating the Feature Condensation in Batch Normalization and a Unified Framework
von: Wang, Shaobo, et al.
Veröffentlicht: (2023)
von: Wang, Shaobo, et al.
Veröffentlicht: (2023)
Multi-modal Crowd Counting via a Broker Modality
von: Meng, Haoliang, et al.
Veröffentlicht: (2024)
von: Meng, Haoliang, et al.
Veröffentlicht: (2024)
OmniSelect: Dynamic Modality-Aware Token Compression for Efficient Omni-modal Large Language Models
von: Yang, Morunliu, et al.
Veröffentlicht: (2026)
von: Yang, Morunliu, et al.
Veröffentlicht: (2026)
FusionEdit: Semantic Fusion and Attention Modulation for Training-Free Image Editing
von: Lai, Yongwen, et al.
Veröffentlicht: (2026)
von: Lai, Yongwen, et al.
Veröffentlicht: (2026)
What and Where to Adapt: Structure-Semantics Co-Tuning for Machine Vision Compression via Synergistic Adapters
von: Liu, Shaobo, et al.
Veröffentlicht: (2026)
von: Liu, Shaobo, et al.
Veröffentlicht: (2026)
LLMRA: Multi-modal Large Language Model based Restoration Assistant
von: Jin, Xiaoyu, et al.
Veröffentlicht: (2024)
von: Jin, Xiaoyu, et al.
Veröffentlicht: (2024)
SAM2Grasp: Resolve Multi-modal Grasping via Prompt-conditioned Temporal Action Prediction
von: Wu, Shengkai, et al.
Veröffentlicht: (2025)
von: Wu, Shengkai, et al.
Veröffentlicht: (2025)
VideoFusion: A Spatio-Temporal Collaborative Network for Multi-modal Video Fusion
von: Tang, Linfeng, et al.
Veröffentlicht: (2025)
von: Tang, Linfeng, et al.
Veröffentlicht: (2025)
CT-CLIP: A Multi-modal Fusion Framework for Robust Apple Leaf Disease Recognition in Complex Environments
von: Liu, Lemin, et al.
Veröffentlicht: (2025)
von: Liu, Lemin, et al.
Veröffentlicht: (2025)
Medical Large Vision Language Models with Multi-Image Visual Ability
von: Yang, Xikai, et al.
Veröffentlicht: (2025)
von: Yang, Xikai, et al.
Veröffentlicht: (2025)
DRUPI: Dataset Reduction Using Privileged Information
von: Wang, Shaobo, et al.
Veröffentlicht: (2024)
von: Wang, Shaobo, et al.
Veröffentlicht: (2024)
Judge, Then Drive: A Critic-Centric Vision Language Action Framework for Autonomous Driving
von: Yang, Lijin, et al.
Veröffentlicht: (2026)
von: Yang, Lijin, et al.
Veröffentlicht: (2026)
Towards Continual Egocentric Activity Recognition: A Multi-modal Egocentric Activity Dataset for Continual Learning
von: Xu, Linfeng, et al.
Veröffentlicht: (2023)
von: Xu, Linfeng, et al.
Veröffentlicht: (2023)
Not All Samples Should Be Utilized Equally: Towards Understanding and Improving Dataset Distillation
von: Wang, Shaobo, et al.
Veröffentlicht: (2024)
von: Wang, Shaobo, et al.
Veröffentlicht: (2024)
Self-supervised 3D Patient Modeling with Multi-modal Attentive Fusion
von: Zheng, Meng, et al.
Veröffentlicht: (2024)
von: Zheng, Meng, et al.
Veröffentlicht: (2024)
RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought
von: Lu, Yi, et al.
Veröffentlicht: (2025)
von: Lu, Yi, et al.
Veröffentlicht: (2025)
SparseVoxFormer: Sparse Voxel-based Transformer for Multi-modal 3D Object Detection
von: Son, Hyeongseok, et al.
Veröffentlicht: (2025)
von: Son, Hyeongseok, et al.
Veröffentlicht: (2025)
Learning Robust Anymodal Segmentor with Unimodal and Cross-modal Distillation
von: Zheng, Xu, et al.
Veröffentlicht: (2024)
von: Zheng, Xu, et al.
Veröffentlicht: (2024)
FMM-Attack: A Flow-based Multi-modal Adversarial Attack on Video-based LLMs
von: Li, Jinmin, et al.
Veröffentlicht: (2024)
von: Li, Jinmin, et al.
Veröffentlicht: (2024)
Knowledge Condensation and Reasoning for Knowledge-based VQA
von: Hao, Dongze, et al.
Veröffentlicht: (2024)
von: Hao, Dongze, et al.
Veröffentlicht: (2024)
Self-paced Multi-grained Cross-modal Interaction Modeling for Referring Expression Comprehension
von: Miao, Peihan, et al.
Veröffentlicht: (2022)
von: Miao, Peihan, et al.
Veröffentlicht: (2022)
Multi-modal Crowd Counting via Modal Emulation
von: Wang, Chenhao, et al.
Veröffentlicht: (2024)
von: Wang, Chenhao, et al.
Veröffentlicht: (2024)
SciVerse: Unveiling the Knowledge Comprehension and Visual Reasoning of LMMs on Multi-modal Scientific Problems
von: Guo, Ziyu, et al.
Veröffentlicht: (2025)
von: Guo, Ziyu, et al.
Veröffentlicht: (2025)
Multi-modal Mutual-Guidance Conditional Prompt Learning for Vision-Language Models
von: Yang, Shijun, et al.
Veröffentlicht: (2025)
von: Yang, Shijun, et al.
Veröffentlicht: (2025)
Enhancing Weakly Supervised Semantic Segmentation with Multi-modal Foundation Models: An End-to-End Approach
von: Ravanbakhsh, Elham, et al.
Veröffentlicht: (2024)
von: Ravanbakhsh, Elham, et al.
Veröffentlicht: (2024)
Video Compression Commander: Plug-and-Play Inference Acceleration for Video Large Language Models
von: Liu, Xuyang, et al.
Veröffentlicht: (2025)
von: Liu, Xuyang, et al.
Veröffentlicht: (2025)
Multi-modal Fusion based Q-distribution Prediction for Controlled Nuclear Fusion
von: Wang, Shiao, et al.
Veröffentlicht: (2024)
von: Wang, Shiao, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
VideoCompressa: Data-Efficient Video Understanding via Joint Temporal Compression and Spatial Reconstruction
von: Wang, Shaobo, et al.
Veröffentlicht: (2025) -
Efficient Multi-modal Large Language Models via Progressive Consistency Distillation
von: Wen, Zichen, et al.
Veröffentlicht: (2025) -
Efficient Large Multi-modal Models via Visual Context Compression
von: Chen, Jieneng, et al.
Veröffentlicht: (2024) -
Socratic-Geo: Synthetic Data Generation and Geometric Reasoning via Multi-Agent Interaction
von: Jiao, Zhengbo, et al.
Veröffentlicht: (2026) -
GaitMA: Pose-guided Multi-modal Feature Fusion for Gait Recognition
von: Min, Fanxu, et al.
Veröffentlicht: (2024)