OmniSegmentor: A Flexible Multi-Modal Learning Framework for Semantic Segmentation
Fuente:
arXiv
Saved in:
| Main Authors: | Yin, Bo-Wen, Cao, Jiao-Long, Zhang, Xuying, Chen, Yuming, Cheng, Ming-Ming, Hou, Qibin |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DFormerv2: Geometry Self-Attention for RGBD Semantic Segmentation
by: Yin, Bo-Wen, et al.
Published: (2025)
by: Yin, Bo-Wen, et al.
Published: (2025)
DFormer: Rethinking RGBD Representation Learning for Semantic Segmentation
by: Yin, Bowen, et al.
Published: (2023)
by: Yin, Bowen, et al.
Published: (2023)
Revisiting Efficient Semantic Segmentation: Learning Offsets for Better Spatial and Class Feature Alignment
by: Zhang, Shi-Chen, et al.
Published: (2025)
by: Zhang, Shi-Chen, et al.
Published: (2025)
Referring Camouflaged Object Detection
by: Zhang, Xuying, et al.
Published: (2023)
by: Zhang, Xuying, et al.
Published: (2023)
Cascade-CLIP: Cascaded Vision-Language Embeddings Alignment for Zero-Shot Semantic Segmentation
by: Li, Yunheng, et al.
Published: (2024)
by: Li, Yunheng, et al.
Published: (2024)
YOLO-MS: Rethinking Multi-Scale Representation Learning for Real-time Object Detection
by: Chen, Yuming, et al.
Published: (2023)
by: Chen, Yuming, et al.
Published: (2023)
AR-1-to-3: Single Image to Consistent 3D Object Generation via Next-View Prediction
by: Zhang, Xuying, et al.
Published: (2025)
by: Zhang, Xuying, et al.
Published: (2025)
GeoWorld: Unlocking the Potential of Geometry Models to Facilitate High-Fidelity 3D Scene Generation
by: Wan, Yuhao, et al.
Published: (2025)
by: Wan, Yuhao, et al.
Published: (2025)
Zone Evaluation: Revealing Spatial Bias in Object Detection
by: Zheng, Zhaohui, et al.
Published: (2023)
by: Zheng, Zhaohui, et al.
Published: (2023)
CrossKD: Cross-Head Knowledge Distillation for Object Detection
by: Wang, Jiabao, et al.
Published: (2023)
by: Wang, Jiabao, et al.
Published: (2023)
VGGT-Segmentor: Geometry-Enhanced Cross-View Segmentation
by: Gao, Yulu, et al.
Published: (2026)
by: Gao, Yulu, et al.
Published: (2026)
High-Quality Mask Tuning Matters for Open-Vocabulary Segmentation
by: Zeng, Quan-Sheng, et al.
Published: (2024)
by: Zeng, Quan-Sheng, et al.
Published: (2024)
SM3Det: A Unified Model for Multi-Modal Remote Sensing Object Detection
by: Li, Yuxuan, et al.
Published: (2024)
by: Li, Yuxuan, et al.
Published: (2024)
Multi-Token Enhancing for Vision Representation Learning
by: Li, Zhong-Yu, et al.
Published: (2024)
by: Li, Zhong-Yu, et al.
Published: (2024)
Integrating Extra Modality Helps Segmentor Find Camouflaged Objects Well
by: Fang, Chengyu, et al.
Published: (2025)
by: Fang, Chengyu, et al.
Published: (2025)
TempSamp-R1: Effective Temporal Sampling with Reinforcement Fine-Tuning for Video LLMs
by: Li, Yunheng, et al.
Published: (2025)
by: Li, Yunheng, et al.
Published: (2025)
Unifying Heterogeneous Multi-Modal Remote Sensing Detection Via Language-Pivoted Pretraining
by: Li, Yuxuan, et al.
Published: (2026)
by: Li, Yuxuan, et al.
Published: (2026)
StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video Generation
by: Zhou, Yupeng, et al.
Published: (2024)
by: Zhou, Yupeng, et al.
Published: (2024)
Traffic Scene Parsing through the TSP6K Dataset
by: Jiang, Peng-Tao, et al.
Published: (2023)
by: Jiang, Peng-Tao, et al.
Published: (2023)
MS-NeRF: Multi-Space Neural Radiance Fields
by: Yin, Ze-Xin, et al.
Published: (2023)
by: Yin, Ze-Xin, et al.
Published: (2023)
Making Training-Free Diffusion Segmentors Scale with the Generative Power
by: Meng, Benyuan, et al.
Published: (2026)
by: Meng, Benyuan, et al.
Published: (2026)
MedSeg-R: Medical Image Segmentation with Clinical Reasoning
by: Shao, Hao, et al.
Published: (2025)
by: Shao, Hao, et al.
Published: (2025)
Learning Robust Anymodal Segmentor with Unimodal and Cross-modal Distillation
by: Zheng, Xu, et al.
Published: (2024)
by: Zheng, Xu, et al.
Published: (2024)
Re-Aligning Language to Visual Objects with an Agentic Workflow
by: Chen, Yuming, et al.
Published: (2025)
by: Chen, Yuming, et al.
Published: (2025)
Rethinking Token-Level Policy Optimization for Multimodal Chain-of-Thought
by: Li, Yunheng, et al.
Published: (2026)
by: Li, Yunheng, et al.
Published: (2026)
KAC: Kolmogorov-Arnold Classifier for Continual Learning
by: Hu, Yusong, et al.
Published: (2025)
by: Hu, Yusong, et al.
Published: (2025)
TAR3D: Creating High-Quality 3D Assets via Next-Part Prediction
by: Zhang, Xuying, et al.
Published: (2024)
by: Zhang, Xuying, et al.
Published: (2024)
ControlSR: Taming Diffusion Models for Consistent Real-World Image Super Resolution
by: Wan, Yuhao, et al.
Published: (2024)
by: Wan, Yuhao, et al.
Published: (2024)
Unbiased Region-Language Alignment for Open-Vocabulary Dense Prediction
by: Li, Yunheng, et al.
Published: (2024)
by: Li, Yunheng, et al.
Published: (2024)
Sora Generates Videos with Stunning Geometrical Consistency
by: Li, Xuanyi, et al.
Published: (2024)
by: Li, Xuanyi, et al.
Published: (2024)
Towards Universal Video MLLMs with Attribute-Structured and Quality-Verified Instructions
by: Li, Yunheng, et al.
Published: (2026)
by: Li, Yunheng, et al.
Published: (2026)
MemorySAM: Memorize Modalities and Semantics with Segment Anything Model 2 for Multi-modal Semantic Segmentation
by: Liao, Chenfei, et al.
Published: (2025)
by: Liao, Chenfei, et al.
Published: (2025)
BiXFormer: A Robust Framework for Maximizing Modality Effectiveness in Multi-Modal Semantic Segmentation
by: Chen, Jialei, et al.
Published: (2025)
by: Chen, Jialei, et al.
Published: (2025)
Mixture of Style Experts for Diverse Image Stylization
by: Zhu, Shihao, et al.
Published: (2026)
by: Zhu, Shihao, et al.
Published: (2026)
Enhancing Representations through Heterogeneous Self-Supervised Learning
by: Li, Zhong-Yu, et al.
Published: (2023)
by: Li, Zhong-Yu, et al.
Published: (2023)
Multi-Scale Representations by Varying Window Attention for Semantic Segmentation
by: Yan, Haotian, et al.
Published: (2024)
by: Yan, Haotian, et al.
Published: (2024)
SRFormerV2: Taking a Closer Look at Permuted Self-Attention for Image Super-Resolution
by: Zhou, Yupeng, et al.
Published: (2023)
by: Zhou, Yupeng, et al.
Published: (2023)
MetaSeg: Content-Aware Meta-Net for Omni-Supervised Semantic Segmentation
by: Jiang, Shenwang, et al.
Published: (2024)
by: Jiang, Shenwang, et al.
Published: (2024)
Language-guided Scale-aware MedSegmentor for Lesion Segmentation in Medical Imaging
by: Ouyang, Shuyi, et al.
Published: (2024)
by: Ouyang, Shuyi, et al.
Published: (2024)
Learning Generalized and Flexible Trajectory Models from Omni-Semantic Supervision
by: Zhu, Yuanshao, et al.
Published: (2025)
by: Zhu, Yuanshao, et al.
Published: (2025)
Similar Items
-
DFormerv2: Geometry Self-Attention for RGBD Semantic Segmentation
by: Yin, Bo-Wen, et al.
Published: (2025) -
DFormer: Rethinking RGBD Representation Learning for Semantic Segmentation
by: Yin, Bowen, et al.
Published: (2023) -
Revisiting Efficient Semantic Segmentation: Learning Offsets for Better Spatial and Class Feature Alignment
by: Zhang, Shi-Chen, et al.
Published: (2025) -
Referring Camouflaged Object Detection
by: Zhang, Xuying, et al.
Published: (2023) -
Cascade-CLIP: Cascaded Vision-Language Embeddings Alignment for Zero-Shot Semantic Segmentation
by: Li, Yunheng, et al.
Published: (2024)