OmniBind: Large-scale Omni Multimodal Representation via Binding Spaces
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Zehan, Zhang, Ziang, Zhang, Hang, Liu, Luping, Huang, Rongjie, Cheng, Xize, Zhao, Hengshuang, Zhao, Zhou |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
OmniBind: Teach to Build Unequal-Scale Modality Interaction for Omni-Bind of All
von: Lyu, Yuanhuiyi, et al.
Veröffentlicht: (2024)
von: Lyu, Yuanhuiyi, et al.
Veröffentlicht: (2024)
FreeBind: Free Lunch in Unified Multimodal Space via Knowledge Fusion
von: Wang, Zehan, et al.
Veröffentlicht: (2024)
von: Wang, Zehan, et al.
Veröffentlicht: (2024)
OmniSep: Unified Omni-Modality Sound Separation with Query-Mixup
von: Cheng, Xize, et al.
Veröffentlicht: (2024)
von: Cheng, Xize, et al.
Veröffentlicht: (2024)
Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers
von: Huang, Haifeng, et al.
Veröffentlicht: (2023)
von: Huang, Haifeng, et al.
Veröffentlicht: (2023)
GenSpace: Benchmarking Spatially-Aware Image Generation
von: Wang, Zehan, et al.
Veröffentlicht: (2025)
von: Wang, Zehan, et al.
Veröffentlicht: (2025)
Orient Anything: Learning Robust Object Orientation Estimation from Rendering 3D Models
von: Wang, Zehan, et al.
Veröffentlicht: (2024)
von: Wang, Zehan, et al.
Veröffentlicht: (2024)
Depth Anything with Any Prior
von: Wang, Zehan, et al.
Veröffentlicht: (2025)
von: Wang, Zehan, et al.
Veröffentlicht: (2025)
OmniMamba: Efficient and Unified Multimodal Understanding and Generation via State Space Models
von: Zou, Jialv, et al.
Veröffentlicht: (2025)
von: Zou, Jialv, et al.
Veröffentlicht: (2025)
R1-Omni: Explainable Omni-Multimodal Emotion Recognition with Reinforcement Learning
von: Zhao, Jiaxing, et al.
Veröffentlicht: (2025)
von: Zhao, Jiaxing, et al.
Veröffentlicht: (2025)
OmniAD: Detect and Understand Industrial Anomaly via Multimodal Reasoning
von: Zhao, Shifang, et al.
Veröffentlicht: (2025)
von: Zhao, Shifang, et al.
Veröffentlicht: (2025)
Omni Survey for Multimodality Analysis in Visual Object Tracking
von: Tang, Zhangyong, et al.
Veröffentlicht: (2025)
von: Tang, Zhangyong, et al.
Veröffentlicht: (2025)
UniBind: LLM-Augmented Unified and Balanced Representation Space to Bind Them All
von: Lyu, Yuanhuiyi, et al.
Veröffentlicht: (2024)
von: Lyu, Yuanhuiyi, et al.
Veröffentlicht: (2024)
OA-CNNs: Omni-Adaptive Sparse CNNs for 3D Semantic Segmentation
von: Peng, Bohao, et al.
Veröffentlicht: (2024)
von: Peng, Bohao, et al.
Veröffentlicht: (2024)
OmniNFT: Modality-wise Omni Diffusion Reinforcement for Joint Audio-Video Generation
von: Zhang, Guohui, et al.
Veröffentlicht: (2026)
von: Zhang, Guohui, et al.
Veröffentlicht: (2026)
OmniSTVG: Toward Spatio-Temporal Omni-Object Video Grounding
von: Yao, Jiali, et al.
Veröffentlicht: (2025)
von: Yao, Jiali, et al.
Veröffentlicht: (2025)
Omni-Scene: Omni-Gaussian Representation for Ego-Centric Sparse-View Scene Reconstruction
von: Wei, Dongxu, et al.
Veröffentlicht: (2024)
von: Wei, Dongxu, et al.
Veröffentlicht: (2024)
OmniFysics: Towards Physical Intelligence Evolution via Omni-Modal Signal Processing and Network Optimization
von: Han, Minghao, et al.
Veröffentlicht: (2026)
von: Han, Minghao, et al.
Veröffentlicht: (2026)
OmniVGGT: Omni-Modality Driven Visual Geometry Grounded Transformer
von: Peng, Haosong, et al.
Veröffentlicht: (2025)
von: Peng, Haosong, et al.
Veröffentlicht: (2025)
OmniRe: Omni Urban Scene Reconstruction
von: Chen, Ziyu, et al.
Veröffentlicht: (2024)
von: Chen, Ziyu, et al.
Veröffentlicht: (2024)
OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding
von: Zhao, Ruixiang, et al.
Veröffentlicht: (2026)
von: Zhao, Ruixiang, et al.
Veröffentlicht: (2026)
OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering
von: Jia, Yiduo, et al.
Veröffentlicht: (2026)
von: Jia, Yiduo, et al.
Veröffentlicht: (2026)
OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding
von: Xi, Dianbing, et al.
Veröffentlicht: (2025)
von: Xi, Dianbing, et al.
Veröffentlicht: (2025)
DreamOmni2: Multimodal Instruction-based Editing and Generation
von: Xia, Bin, et al.
Veröffentlicht: (2025)
von: Xia, Bin, et al.
Veröffentlicht: (2025)
OmniEvent: Unified Event Representation Learning
von: Yan, Weiqi, et al.
Veröffentlicht: (2025)
von: Yan, Weiqi, et al.
Veröffentlicht: (2025)
DreamVideo-Omni: Omni-Motion Controlled Multi-Subject Video Customization with Latent Identity Reinforcement Learning
von: Wei, Yujie, et al.
Veröffentlicht: (2026)
von: Wei, Yujie, et al.
Veröffentlicht: (2026)
Tele-Omni: a Unified Multimodal Framework for Video Generation and Editing
von: Liu, Jialun, et al.
Veröffentlicht: (2026)
von: Liu, Jialun, et al.
Veröffentlicht: (2026)
OmniSelect: Dynamic Modality-Aware Token Compression for Efficient Omni-modal Large Language Models
von: Yang, Morunliu, et al.
Veröffentlicht: (2026)
von: Yang, Morunliu, et al.
Veröffentlicht: (2026)
Context Unrolling in Omni Models
von: Yang, Ceyuan, et al.
Veröffentlicht: (2026)
von: Yang, Ceyuan, et al.
Veröffentlicht: (2026)
ThinkOmni: Lifting Textual Reasoning to Omni-modal Scenarios via Guidance Decoding
von: Guan, Yiran, et al.
Veröffentlicht: (2026)
von: Guan, Yiran, et al.
Veröffentlicht: (2026)
Kling-Omni Technical Report
von: Kling Team, et al.
Veröffentlicht: (2025)
von: Kling Team, et al.
Veröffentlicht: (2025)
OmniVaT: Single Domain Generalization for Multimodal Visual-Tactile Learning
von: Qiu, Liuxiang, et al.
Veröffentlicht: (2026)
von: Qiu, Liuxiang, et al.
Veröffentlicht: (2026)
Omni$^2$: Unifying Omnidirectional Image Generation and Editing in an Omni Model
von: Yang, Liu, et al.
Veröffentlicht: (2025)
von: Yang, Liu, et al.
Veröffentlicht: (2025)
OmniAudio: Generating Spatial Audio from 360-Degree Video
von: Liu, Huadai, et al.
Veröffentlicht: (2025)
von: Liu, Huadai, et al.
Veröffentlicht: (2025)
OmniVCus: Feedforward Subject-driven Video Customization with Multimodal Control Conditions
von: Cai, Yuanhao, et al.
Veröffentlicht: (2025)
von: Cai, Yuanhao, et al.
Veröffentlicht: (2025)
OmniScience: A Large-scale Multi-modal Dataset for Scientific Image Understanding
von: Tao, Haoyi, et al.
Veröffentlicht: (2026)
von: Tao, Haoyi, et al.
Veröffentlicht: (2026)
OralGPT-Omni: A Versatile Dental Multimodal Large Language Model
von: Hao, Jing, et al.
Veröffentlicht: (2025)
von: Hao, Jing, et al.
Veröffentlicht: (2025)
HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context
von: Yang, Qize, et al.
Veröffentlicht: (2025)
von: Yang, Qize, et al.
Veröffentlicht: (2025)
Omni-Weather: A Unified Multimodal Model for Weather Radar Understanding and Generation
von: Zhou, Zhiwang, et al.
Veröffentlicht: (2025)
von: Zhou, Zhiwang, et al.
Veröffentlicht: (2025)
Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation
von: Liao, Chao, et al.
Veröffentlicht: (2025)
von: Liao, Chao, et al.
Veröffentlicht: (2025)
ViT-Lens: Towards Omni-modal Representations
von: Lei, Weixian, et al.
Veröffentlicht: (2023)
von: Lei, Weixian, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
OmniBind: Teach to Build Unequal-Scale Modality Interaction for Omni-Bind of All
von: Lyu, Yuanhuiyi, et al.
Veröffentlicht: (2024) -
FreeBind: Free Lunch in Unified Multimodal Space via Knowledge Fusion
von: Wang, Zehan, et al.
Veröffentlicht: (2024) -
OmniSep: Unified Omni-Modality Sound Separation with Query-Mixup
von: Cheng, Xize, et al.
Veröffentlicht: (2024) -
Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers
von: Huang, Haifeng, et al.
Veröffentlicht: (2023) -
GenSpace: Benchmarking Spatially-Aware Image Generation
von: Wang, Zehan, et al.
Veröffentlicht: (2025)