OmniBind: Teach to Build Unequal-Scale Modality Interaction for Omni-Bind of All
Fuente:
arXiv
Guardado en:
| Autores principales: | Lyu, Yuanhuiyi, Zheng, Xu, Kim, Dahun, Wang, Lin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
OmniBind: Large-scale Omni Multimodal Representation via Binding Spaces
por: Wang, Zehan, et al.
Publicado: (2024)
por: Wang, Zehan, et al.
Publicado: (2024)
UniBind: LLM-Augmented Unified and Balanced Representation Space to Bind Them All
por: Lyu, Yuanhuiyi, et al.
Publicado: (2024)
por: Lyu, Yuanhuiyi, et al.
Publicado: (2024)
EventBind: Learning a Unified Representation to Bind Them All for Event-based Open-world Understanding
por: Zhou, Jiazhou, et al.
Publicado: (2023)
por: Zhou, Jiazhou, et al.
Publicado: (2023)
Learning Modality-agnostic Representation for Semantic Segmentation from Any Modalities
por: Zheng, Xu, et al.
Publicado: (2024)
por: Zheng, Xu, et al.
Publicado: (2024)
Centering the Value of Every Modality: Towards Efficient and Resilient Modality-agnostic Semantic Segmentation
por: Zheng, Xu, et al.
Publicado: (2024)
por: Zheng, Xu, et al.
Publicado: (2024)
OmniSAM: Omnidirectional Segment Anything Model for UDA in Panoramic Semantic Segmentation
por: Zhong, Ding, et al.
Publicado: (2025)
por: Zhong, Ding, et al.
Publicado: (2025)
T-Rex-Omni: Integrating Negative Visual Prompt in Generic Object Detection
por: Zhou, Jiazhou, et al.
Publicado: (2025)
por: Zhou, Jiazhou, et al.
Publicado: (2025)
OmniVGGT: Omni-Modality Driven Visual Geometry Grounded Transformer
por: Peng, Haosong, et al.
Publicado: (2025)
por: Peng, Haosong, et al.
Publicado: (2025)
OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering
por: Jia, Yiduo, et al.
Publicado: (2026)
por: Jia, Yiduo, et al.
Publicado: (2026)
Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction
por: He, Chaoqun, et al.
Publicado: (2026)
por: He, Chaoqun, et al.
Publicado: (2026)
MAGIC++: Efficient and Resilient Modality-Agnostic Semantic Segmentation via Hierarchical Modality Selection
por: Zheng, Xu, et al.
Publicado: (2024)
por: Zheng, Xu, et al.
Publicado: (2024)
OmniSelect: Dynamic Modality-Aware Token Compression for Efficient Omni-modal Large Language Models
por: Yang, Morunliu, et al.
Publicado: (2026)
por: Yang, Morunliu, et al.
Publicado: (2026)
Image Anything: Towards Reasoning-coherent and Training-free Multi-modal Image Generation
por: Lyu, Yuanhuiyi, et al.
Publicado: (2024)
por: Lyu, Yuanhuiyi, et al.
Publicado: (2024)
OmniNFT: Modality-wise Omni Diffusion Reinforcement for Joint Audio-Video Generation
por: Zhang, Guohui, et al.
Publicado: (2026)
por: Zhang, Guohui, et al.
Publicado: (2026)
OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM
por: Ye, Hanrong, et al.
Publicado: (2025)
por: Ye, Hanrong, et al.
Publicado: (2025)
HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context
por: Yang, Qize, et al.
Publicado: (2025)
por: Yang, Qize, et al.
Publicado: (2025)
Capybara-OMNI: An Efficient Paradigm for Building Omni-Modal Language Models
por: Ji, Xingguang, et al.
Publicado: (2025)
por: Ji, Xingguang, et al.
Publicado: (2025)
InteractiveOmni: A Unified Omni-modal Model for Audio-Visual Multi-turn Dialogue
por: Tong, Wenwen, et al.
Publicado: (2025)
por: Tong, Wenwen, et al.
Publicado: (2025)
OmniRe: Omni Urban Scene Reconstruction
por: Chen, Ziyu, et al.
Publicado: (2024)
por: Chen, Ziyu, et al.
Publicado: (2024)
LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning
por: Dai, Yifan, et al.
Publicado: (2026)
por: Dai, Yifan, et al.
Publicado: (2026)
OmniFysics: Towards Physical Intelligence Evolution via Omni-Modal Signal Processing and Network Optimization
por: Han, Minghao, et al.
Publicado: (2026)
por: Han, Minghao, et al.
Publicado: (2026)
Chasing Day and Night: Towards Robust and Efficient All-Day Object Detection Guided by an Event Camera
por: Cao, Jiahang, et al.
Publicado: (2023)
por: Cao, Jiahang, et al.
Publicado: (2023)
ExACT: Language-guided Conceptual Reasoning and Uncertainty Estimation for Event-based Action Recognition and More
por: Zhou, Jiazhou, et al.
Publicado: (2024)
por: Zhou, Jiazhou, et al.
Publicado: (2024)
OmniSep: Unified Omni-Modality Sound Separation with Query-Mixup
por: Cheng, Xize, et al.
Publicado: (2024)
por: Cheng, Xize, et al.
Publicado: (2024)
Is Extending Modality The Right Path Towards Omni-Modality?
por: Zhu, Tinghui, et al.
Publicado: (2025)
por: Zhu, Tinghui, et al.
Publicado: (2025)
OmniGAIA: Towards Native Omni-Modal AI Agents
por: Li, Xiaoxi, et al.
Publicado: (2026)
por: Li, Xiaoxi, et al.
Publicado: (2026)
OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding
por: Xi, Dianbing, et al.
Publicado: (2025)
por: Xi, Dianbing, et al.
Publicado: (2025)
FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs
por: Chen, Qian, et al.
Publicado: (2026)
por: Chen, Qian, et al.
Publicado: (2026)
OmniCLIP: Adapting CLIP for Video Recognition with Spatial-Temporal Omni-Scale Feature Learning
por: Liu, Mushui, et al.
Publicado: (2024)
por: Liu, Mushui, et al.
Publicado: (2024)
Reducing Unimodal Bias in Multi-Modal Semantic Segmentation with Multi-Scale Functional Entropy Regularization
por: Zheng, Xu, et al.
Publicado: (2025)
por: Zheng, Xu, et al.
Publicado: (2025)
Modality Unified Attack for Omni-Modality Person Re-Identification
por: Bian, Yuan, et al.
Publicado: (2025)
por: Bian, Yuan, et al.
Publicado: (2025)
Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation
por: Liao, Chao, et al.
Publicado: (2025)
por: Liao, Chao, et al.
Publicado: (2025)
Omni$^2$: Unifying Omnidirectional Image Generation and Editing in an Omni Model
por: Yang, Liu, et al.
Publicado: (2025)
por: Yang, Liu, et al.
Publicado: (2025)
OmniStyle: Filtering High Quality Style Transfer Data at Scale
por: Wang, Ye, et al.
Publicado: (2025)
por: Wang, Ye, et al.
Publicado: (2025)
Omni-Reward: Towards Generalist Omni-Modal Reward Modeling with Free-Form Preferences
por: Jin, Zhuoran, et al.
Publicado: (2025)
por: Jin, Zhuoran, et al.
Publicado: (2025)
ViT-Lens: Initiating Omni-Modal Exploration through 3D Insights
por: Lei, Weixian, et al.
Publicado: (2023)
por: Lei, Weixian, et al.
Publicado: (2023)
OmniTransfer: All-in-one Framework for Spatio-temporal Video Transfer
por: Zhang, Pengze, et al.
Publicado: (2026)
por: Zhang, Pengze, et al.
Publicado: (2026)
OmniSat: Self-Supervised Modality Fusion for Earth Observation
por: Astruc, Guillaume, et al.
Publicado: (2024)
por: Astruc, Guillaume, et al.
Publicado: (2024)
OmniGCD: Abstracting Generalized Category Discovery for Modality Agnosticism
por: Shipard, Jordan, et al.
Publicado: (2026)
por: Shipard, Jordan, et al.
Publicado: (2026)
Omni-Captioner: Data Pipeline, Models, and Benchmark for Omni Detailed Perception
por: Ma, Ziyang, et al.
Publicado: (2025)
por: Ma, Ziyang, et al.
Publicado: (2025)
Ejemplares similares
-
OmniBind: Large-scale Omni Multimodal Representation via Binding Spaces
por: Wang, Zehan, et al.
Publicado: (2024) -
UniBind: LLM-Augmented Unified and Balanced Representation Space to Bind Them All
por: Lyu, Yuanhuiyi, et al.
Publicado: (2024) -
EventBind: Learning a Unified Representation to Bind Them All for Event-based Open-world Understanding
por: Zhou, Jiazhou, et al.
Publicado: (2023) -
Learning Modality-agnostic Representation for Semantic Segmentation from Any Modalities
por: Zheng, Xu, et al.
Publicado: (2024) -
Centering the Value of Every Modality: Towards Efficient and Resilient Modality-agnostic Semantic Segmentation
por: Zheng, Xu, et al.
Publicado: (2024)