OmniEncoder: See, Hear, and Feel Continuous Motion Like Humans With One Encoder
Fuente:
arXiv
Saved in:
| Main Authors: | Bai, Detao, Yao, Shimin, Chen, Weixuan, Lai, Chengen, Li, Yuanming, Ma, Zhiheng, Wei, Xihan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HumanOmni-Speaker: Identifying Who said What and When
by: Bai, Detao, et al.
Published: (2026)
by: Bai, Detao, et al.
Published: (2026)
HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context
by: Yang, Qize, et al.
Published: (2025)
by: Yang, Qize, et al.
Published: (2025)
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding
by: Zhao, Jiaxing, et al.
Published: (2025)
by: Zhao, Jiaxing, et al.
Published: (2025)
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization
by: Bai, Detao, et al.
Published: (2025)
by: Bai, Detao, et al.
Published: (2025)
Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis
by: Yang, Qize, et al.
Published: (2025)
by: Yang, Qize, et al.
Published: (2025)
Encoder-Free Human Motion Understanding via Structured Motion Descriptions
by: Zhang, Yao, et al.
Published: (2026)
by: Zhang, Yao, et al.
Published: (2026)
R1-Omni: Explainable Omni-Multimodal Emotion Recognition with Reinforcement Learning
by: Zhao, Jiaxing, et al.
Published: (2025)
by: Zhao, Jiaxing, et al.
Published: (2025)
Omni-Attribute: Open-vocabulary Attribute Encoder for Visual Concept Personalization
by: Chen, Tsai-Shien, et al.
Published: (2025)
by: Chen, Tsai-Shien, et al.
Published: (2025)
Omni-C: Compressing Heterogeneous Modalities into a Single Dense Encoder
by: Lau, Kin Wai, et al.
Published: (2026)
by: Lau, Kin Wai, et al.
Published: (2026)
Can Visual Encoder Learn to See Arrows?
by: Terashita, Naoyuki, et al.
Published: (2025)
by: Terashita, Naoyuki, et al.
Published: (2025)
Decodable and Sample Invariant Continuous Object Encoder
by: Yuan, Dehao, et al.
Published: (2023)
by: Yuan, Dehao, et al.
Published: (2023)
Candidate Set Re-ranking for Composed Image Retrieval with Dual Multi-modal Encoder
by: Liu, Zheyuan, et al.
Published: (2023)
by: Liu, Zheyuan, et al.
Published: (2023)
Utonia: Toward One Encoder for All Point Clouds
by: Zhang, Yujia, et al.
Published: (2026)
by: Zhang, Yujia, et al.
Published: (2026)
Exploiting the Semantic Knowledge of Pre-trained Text-Encoders for Continual Learning
by: Yu, Lu, et al.
Published: (2024)
by: Yu, Lu, et al.
Published: (2024)
UNIT: Unifying Image and Text Recognition in One Vision Encoder
by: Zhu, Yi, et al.
Published: (2024)
by: Zhu, Yi, et al.
Published: (2024)
See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding
by: Sun, Boyuan, et al.
Published: (2026)
by: Sun, Boyuan, et al.
Published: (2026)
OmniMotion: Multimodal Motion Generation with Continuous Masked Autoregression
by: Li, Zhe, et al.
Published: (2025)
by: Li, Zhe, et al.
Published: (2025)
Efficient Universal Perception Encoder
by: Zhu, Chenchen, et al.
Published: (2026)
by: Zhu, Chenchen, et al.
Published: (2026)
Language-Image Alignment with Fixed Text Encoders
by: Yang, Jingfeng, et al.
Published: (2025)
by: Yang, Jingfeng, et al.
Published: (2025)
Seeing through Unclear Glass: Occlusion Removal with One Shot
by: Li, Qiang, et al.
Published: (2025)
by: Li, Qiang, et al.
Published: (2025)
OneEncoder: A Lightweight Framework for Progressive Alignment of Modalities
by: Faye, Bilal, et al.
Published: (2024)
by: Faye, Bilal, et al.
Published: (2024)
Perception Encoder: The best visual embeddings are not at the output of the network
by: Bolya, Daniel, et al.
Published: (2025)
by: Bolya, Daniel, et al.
Published: (2025)
Beyond the Encoder: Joint Encoder-Decoder Contrastive Pre-Training Improves Dense Prediction
by: Quetin, Sébastien, et al.
Published: (2025)
by: Quetin, Sébastien, et al.
Published: (2025)
Securely Fine-tuning Pre-trained Encoders Against Adversarial Examples
by: Zhou, Ziqi, et al.
Published: (2024)
by: Zhou, Ziqi, et al.
Published: (2024)
Encoder-Only Image Registration
by: Chen, Xiang, et al.
Published: (2025)
by: Chen, Xiang, et al.
Published: (2025)
Multiscale Encoder and Omni-Dimensional Dynamic Convolution Enrichment in nnU-Net for Brain Tumor Segmentation
by: Mistry, Sahaj K., et al.
Published: (2024)
by: Mistry, Sahaj K., et al.
Published: (2024)
One Layer Is Enough: Adapting Pretrained Visual Encoders for Image Generation
by: Gao, Yuan, et al.
Published: (2025)
by: Gao, Yuan, et al.
Published: (2025)
OneVision-Encoder: Codec-Aligned Sparsity as a Foundational Principle for Multimodal Intelligence
by: Tang, Feilong, et al.
Published: (2026)
by: Tang, Feilong, et al.
Published: (2026)
Unified Map Prior Encoder for Mapping and Planning
by: Zhang, Zongzheng, et al.
Published: (2026)
by: Zhang, Zongzheng, et al.
Published: (2026)
Uni-Encoder Meets Multi-Encoders: Representation Before Fusion for Brain Tumor Segmentation with Missing Modalities
by: Song, Peibo, et al.
Published: (2026)
by: Song, Peibo, et al.
Published: (2026)
Appearance Blur-driven AutoEncoder and Motion-guided Memory Module for Video Anomaly Detection
by: Lyu, Jiahao, et al.
Published: (2024)
by: Lyu, Jiahao, et al.
Published: (2024)
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
by: Liu, Zhiheng, et al.
Published: (2026)
by: Liu, Zhiheng, et al.
Published: (2026)
Human-like Content Analysis for Generative AI with Language-Grounded Sparse Encoders
by: Tang, Yiming, et al.
Published: (2025)
by: Tang, Yiming, et al.
Published: (2025)
VENI: Variational Encoder for Natural Illumination
by: Walker, Paul, et al.
Published: (2026)
by: Walker, Paul, et al.
Published: (2026)
Image Generation with a Sphere Encoder
by: Yue, Kaiyu, et al.
Published: (2026)
by: Yue, Kaiyu, et al.
Published: (2026)
Dual Associated Encoder for Face Restoration
by: Tsai, Yu-Ju, et al.
Published: (2023)
by: Tsai, Yu-Ju, et al.
Published: (2023)
Distribution Matching Variational AutoEncoder
by: Ye, Sen, et al.
Published: (2025)
by: Ye, Sen, et al.
Published: (2025)
3D Gaussian Point Encoders
by: James, Jim, et al.
Published: (2025)
by: James, Jim, et al.
Published: (2025)
Unified Multimodal Models as Auto-Encoders
by: Yan, Zhiyuan, et al.
Published: (2025)
by: Yan, Zhiyuan, et al.
Published: (2025)
Text-Guided Semantic Image Encoder
by: Thirukovalluru, Raghuveer, et al.
Published: (2025)
by: Thirukovalluru, Raghuveer, et al.
Published: (2025)
Similar Items
-
HumanOmni-Speaker: Identifying Who said What and When
by: Bai, Detao, et al.
Published: (2026) -
HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context
by: Yang, Qize, et al.
Published: (2025) -
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding
by: Zhao, Jiaxing, et al.
Published: (2025) -
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization
by: Bai, Detao, et al.
Published: (2025) -
Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis
by: Yang, Qize, et al.
Published: (2025)