MOBIUS: Big-to-Mobile Universal Instance Segmentation via Multi-modal Bottleneck Fusion and Calibrated Decoder Pruning
Fuente:
arXiv
Saved in:
| Main Authors: | Segu, Mattia, Gazulla, Marta Tintore, Xian, Yongqin, Van Gool, Luc, Tombari, Federico |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Know Your Neighbors: Improving Single-View Reconstruction via Spatial Vision-Language Reasoning
by: Li, Rui, et al.
Published: (2024)
by: Li, Rui, et al.
Published: (2024)
Lost in Translation? Vocabulary Alignment for Source-Free Adaptation in Open-Vocabulary Semantic Segmentation
by: Mazzucco, Silvio, et al.
Published: (2025)
by: Mazzucco, Silvio, et al.
Published: (2025)
Shifting the Breaking Point of Flow Matching for Multi-Instance Editing
by: Zaccagnino, Carmine, et al.
Published: (2026)
by: Zaccagnino, Carmine, et al.
Published: (2026)
Matching Anything by Segmenting Anything
by: Li, Siyuan, et al.
Published: (2024)
by: Li, Siyuan, et al.
Published: (2024)
UniDepth: Universal Monocular Metric Depth Estimation
by: Piccinelli, Luigi, et al.
Published: (2024)
by: Piccinelli, Luigi, et al.
Published: (2024)
Walker: Self-supervised Multiple Object Tracking by Walking on Temporal Appearance Graphs
by: Segu, Mattia, et al.
Published: (2024)
by: Segu, Mattia, et al.
Published: (2024)
UniK3D: Universal Camera Monocular 3D Estimation
by: Piccinelli, Luigi, et al.
Published: (2025)
by: Piccinelli, Luigi, et al.
Published: (2025)
UniDepthV2: Universal Monocular Metric Depth Estimation Made Simpler
by: Piccinelli, Luigi, et al.
Published: (2025)
by: Piccinelli, Luigi, et al.
Published: (2025)
Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos
by: Plizzari, Chiara, et al.
Published: (2025)
by: Plizzari, Chiara, et al.
Published: (2025)
Samba: Synchronized Set-of-Sequences Modeling for Multiple Object Tracking
by: Segu, Mattia, et al.
Published: (2024)
by: Segu, Mattia, et al.
Published: (2024)
Self-supervised Shape Completion via Involution and Implicit Correspondences
by: Liu, Mengya, et al.
Published: (2024)
by: Liu, Mengya, et al.
Published: (2024)
UIP2P: Unsupervised Instruction-based Image Editing via Edit Reversibility Constraint
by: Simsar, Enis, et al.
Published: (2024)
by: Simsar, Enis, et al.
Published: (2024)
LIME: Localized Image Editing via Attention Regularization in Diffusion Models
by: Simsar, Enis, et al.
Published: (2023)
by: Simsar, Enis, et al.
Published: (2023)
Text-Conditioned Resampler For Long Form Video Understanding
by: Korbar, Bruno, et al.
Published: (2023)
by: Korbar, Bruno, et al.
Published: (2023)
Language-Guided Instance-Aware Domain-Adaptive Panoptic Segmentation
by: Mansour, Elham Amin, et al.
Published: (2024)
by: Mansour, Elham Amin, et al.
Published: (2024)
SLAck: Semantic, Location, and Appearance Aware Open-Vocabulary Tracking
by: Li, Siyuan, et al.
Published: (2024)
by: Li, Siyuan, et al.
Published: (2024)
PALM: Predicting Actions through Language Models
by: Kim, Sanghwan, et al.
Published: (2023)
by: Kim, Sanghwan, et al.
Published: (2023)
Semantic Library Adaptation: LoRA Retrieval and Fusion for Open-Vocabulary Semantic Segmentation
by: Qorbani, Reza, et al.
Published: (2025)
by: Qorbani, Reza, et al.
Published: (2025)
Learning to Prompt with Text Only Supervision for Vision-Language Models
by: Khattak, Muhammad Uzair, et al.
Published: (2024)
by: Khattak, Muhammad Uzair, et al.
Published: (2024)
One2Any: One-Reference 6D Pose Estimation for Any Object
by: Liu, Mengya, et al.
Published: (2025)
by: Liu, Mengya, et al.
Published: (2025)
Towards Real-Time Open-Vocabulary Video Instance Segmentation
by: Yan, Bin, et al.
Published: (2024)
by: Yan, Bin, et al.
Published: (2024)
Bayesian Self-Training for Semi-Supervised 3D Segmentation
by: Unal, Ozan, et al.
Published: (2024)
by: Unal, Ozan, et al.
Published: (2024)
Shapley Pruning for Neural Network Compression
by: Adamczewski, Kamil, et al.
Published: (2024)
by: Adamczewski, Kamil, et al.
Published: (2024)
InseRF: Text-Driven Generative Object Insertion in Neural 3D Scenes
by: Shahbazi, Mohamad, et al.
Published: (2024)
by: Shahbazi, Mohamad, et al.
Published: (2024)
Optimizing against Infeasible Inclusions from Data for Semantic Segmentation through Morphology
by: Basu, Shamik, et al.
Published: (2024)
by: Basu, Shamik, et al.
Published: (2024)
Benchmarking Multi-modal Semantic Segmentation under Sensor Failures: Missing and Noisy Modality Robustness
by: Liao, Chenfei, et al.
Published: (2025)
by: Liao, Chenfei, et al.
Published: (2025)
Splat-SLAM: Globally Optimized RGB-only SLAM with 3D Gaussians
by: Sandström, Erik, et al.
Published: (2024)
by: Sandström, Erik, et al.
Published: (2024)
Condition-Invariant Semantic Segmentation
by: Sakaridis, Christos, et al.
Published: (2023)
by: Sakaridis, Christos, et al.
Published: (2023)
VGGRPO: Towards World-Consistent Video Generation with 4D Latent Reward
by: An, Zhaochong, et al.
Published: (2026)
by: An, Zhaochong, et al.
Published: (2026)
CAFuser: Condition-Aware Multimodal Fusion for Robust Semantic Perception of Driving Scenes
by: Broedermann, Tim, et al.
Published: (2024)
by: Broedermann, Tim, et al.
Published: (2024)
A Simple and Generalist Approach for Panoptic Segmentation
by: Prisadnikov, Nedyalko, et al.
Published: (2024)
by: Prisadnikov, Nedyalko, et al.
Published: (2024)
Four Ways to Improve Verbo-visual Fusion for Dense 3D Visual Grounding
by: Unal, Ozan, et al.
Published: (2023)
by: Unal, Ozan, et al.
Published: (2023)
Human Pose Descriptions and Subject-Focused Attention for Improved Zero-Shot Transfer in Human-Centric Classification Tasks
by: Khan, Muhammad Saif Ullah, et al.
Published: (2024)
by: Khan, Muhammad Saif Ullah, et al.
Published: (2024)
Toward a Diffusion-Based Generalist for Dense Vision Tasks
by: Fan, Yue, et al.
Published: (2024)
by: Fan, Yue, et al.
Published: (2024)
Cross-View Multi-Modal Segmentation @ Ego-Exo4D Challenges 2025
by: Fu, Yuqian, et al.
Published: (2025)
by: Fu, Yuqian, et al.
Published: (2025)
Reducing Unimodal Bias in Multi-Modal Semantic Segmentation with Multi-Scale Functional Entropy Regularization
by: Zheng, Xu, et al.
Published: (2025)
by: Zheng, Xu, et al.
Published: (2025)
DGFusion: Depth-Guided Sensor Fusion for Robust Semantic Perception
by: Broedermannn, Tim, et al.
Published: (2025)
by: Broedermannn, Tim, et al.
Published: (2025)
CalibNet: Dual-branch Cross-modal Calibration for RGB-D Salient Instance Segmentation
by: Pei, Jialun, et al.
Published: (2023)
by: Pei, Jialun, et al.
Published: (2023)
Training-Free Dataset Pruning for Instance Segmentation
by: Dai, Yalun, et al.
Published: (2025)
by: Dai, Yalun, et al.
Published: (2025)
OVI-MAP:Open-Vocabulary Instance-Semantic Mapping
by: Deng, Zilong, et al.
Published: (2026)
by: Deng, Zilong, et al.
Published: (2026)
Similar Items
-
Know Your Neighbors: Improving Single-View Reconstruction via Spatial Vision-Language Reasoning
by: Li, Rui, et al.
Published: (2024) -
Lost in Translation? Vocabulary Alignment for Source-Free Adaptation in Open-Vocabulary Semantic Segmentation
by: Mazzucco, Silvio, et al.
Published: (2025) -
Shifting the Breaking Point of Flow Matching for Multi-Instance Editing
by: Zaccagnino, Carmine, et al.
Published: (2026) -
Matching Anything by Segmenting Anything
by: Li, Siyuan, et al.
Published: (2024) -
UniDepth: Universal Monocular Metric Depth Estimation
by: Piccinelli, Luigi, et al.
Published: (2024)