MMOne: Representing Multiple Modalities in One Scene
Fuente:
arXiv
Saved in:
| Main Authors: | Gu, Zhifeng, Wang, Bing |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Taming Diffusion for Dataset Distillation with High Representativeness
by: Zhao, Lin, et al.
Published: (2025)
by: Zhao, Lin, et al.
Published: (2025)
MultiOOD: Scaling Out-of-Distribution Detection for Multiple Modalities
by: Dong, Hao, et al.
Published: (2024)
by: Dong, Hao, et al.
Published: (2024)
MMGDreamer: Mixed-Modality Graph for Geometry-Controllable 3D Indoor Scene Generation
by: Yang, Zhifei, et al.
Published: (2025)
by: Yang, Zhifei, et al.
Published: (2025)
Self-Supervised Multi-Frame Neural Scene Flow
by: Liu, Dongrui, et al.
Published: (2024)
by: Liu, Dongrui, et al.
Published: (2024)
Beyond Unimodal Learning: The Importance of Integrating Multiple Modalities for Lifelong Learning
by: Sarfraz, Fahad, et al.
Published: (2024)
by: Sarfraz, Fahad, et al.
Published: (2024)
OneLLM: One Framework to Align All Modalities with Language
by: Han, Jiaming, et al.
Published: (2023)
by: Han, Jiaming, et al.
Published: (2023)
FedDiff: Diffusion Model Driven Federated Learning for Multi-Modal and Multi-Clients
by: Li, DaiXun, et al.
Published: (2023)
by: Li, DaiXun, et al.
Published: (2023)
EchoScene: Indoor Scene Generation via Information Echo over Scene Graph Diffusion
by: Zhai, Guangyao, et al.
Published: (2024)
by: Zhai, Guangyao, et al.
Published: (2024)
Representing Online Handwriting for Recognition in Large Vision-Language Models
by: Fadeeva, Anastasiia, et al.
Published: (2024)
by: Fadeeva, Anastasiia, et al.
Published: (2024)
Pixel-level Scene Understanding in One Token: Visual States Need What-is-Where Composition
by: Lee, Seokmin, et al.
Published: (2026)
by: Lee, Seokmin, et al.
Published: (2026)
Instruct 4D-to-4D: Editing 4D Scenes as Pseudo-3D Scenes Using 2D Diffusion
by: Mou, Linzhan, et al.
Published: (2024)
by: Mou, Linzhan, et al.
Published: (2024)
Attention over Scene Graphs: Indoor Scene Representations Toward CSAI Classification
by: Barros, Artur, et al.
Published: (2025)
by: Barros, Artur, et al.
Published: (2025)
SceneTok: A Compressed, Diffusable Token Space for 3D Scenes
by: Asim, Mohammad, et al.
Published: (2026)
by: Asim, Mohammad, et al.
Published: (2026)
BYE: Build Your Encoder with One Sequence of Exploration Data for Long-Term Dynamic Scene Understanding
by: Huang, Chenguang, et al.
Published: (2024)
by: Huang, Chenguang, et al.
Published: (2024)
Generate Any Scene: Scene Graph Driven Data Synthesis for Visual Generation Training
by: Gao, Ziqi, et al.
Published: (2024)
by: Gao, Ziqi, et al.
Published: (2024)
Multi-Modal Video Feature Extraction for Popularity Prediction
by: Liu, Haixu, et al.
Published: (2025)
by: Liu, Haixu, et al.
Published: (2025)
Explaining and Mitigating the Modality Gap in Contrastive Multimodal Learning
by: Yaras, Can, et al.
Published: (2024)
by: Yaras, Can, et al.
Published: (2024)
Deep Multimodal Learning with Missing Modality: A Survey
by: Wu, Renjie, et al.
Published: (2024)
by: Wu, Renjie, et al.
Published: (2024)
Holistic Order Prediction in Natural Scenes
by: Musacchio, Pierre, et al.
Published: (2025)
by: Musacchio, Pierre, et al.
Published: (2025)
One-Prompt-One-Story: Free-Lunch Consistent Text-to-Image Generation Using a Single Prompt
by: Liu, Tao, et al.
Published: (2025)
by: Liu, Tao, et al.
Published: (2025)
MMCORE: MultiModal COnnection with Representation Aligned Latent Embeddings
by: Li, Zijie, et al.
Published: (2026)
by: Li, Zijie, et al.
Published: (2026)
Interleaved-Modal Chain-of-Thought
by: Gao, Jun, et al.
Published: (2024)
by: Gao, Jun, et al.
Published: (2024)
Learnable Expansion of Graph Operators for Multi-Modal Feature Fusion
by: Ding, Dexuan, et al.
Published: (2024)
by: Ding, Dexuan, et al.
Published: (2024)
Multiple Object Stitching for Unsupervised Representation Learning
by: Shen, Chengchao, et al.
Published: (2025)
by: Shen, Chengchao, et al.
Published: (2025)
MM-PoE: Multiple Choice Reasoning via. Process of Elimination using Multi-Modal Models
by: Chakrabarty, Sayak, et al.
Published: (2024)
by: Chakrabarty, Sayak, et al.
Published: (2024)
PhysInOne: Visual Physics Learning and Reasoning in One Suite
by: Zhou, Siyuan, et al.
Published: (2026)
by: Zhou, Siyuan, et al.
Published: (2026)
Do Vision and Language Encoders Represent the World Similarly?
by: Maniparambil, Mayug, et al.
Published: (2024)
by: Maniparambil, Mayug, et al.
Published: (2024)
D$^3$epth: Self-Supervised Depth Estimation with Dynamic Mask in Dynamic Scenes
by: Chen, Siyu, et al.
Published: (2024)
by: Chen, Siyu, et al.
Published: (2024)
A Survey on Cache Methods in Diffusion Models: Toward Efficient Multi-Modal Generation
by: Liu, Jiacheng, et al.
Published: (2025)
by: Liu, Jiacheng, et al.
Published: (2025)
Diffusion for Out-of-Distribution Detection on Road Scenes and Beyond
by: Galesso, Silvio, et al.
Published: (2024)
by: Galesso, Silvio, et al.
Published: (2024)
Pixel-Wise Recognition for Holistic Surgical Scene Understanding
by: Ayobi, Nicolás, et al.
Published: (2024)
by: Ayobi, Nicolás, et al.
Published: (2024)
Moving Off-the-Grid: Scene-Grounded Video Representations
by: van Steenkiste, Sjoerd, et al.
Published: (2024)
by: van Steenkiste, Sjoerd, et al.
Published: (2024)
Step-level Denoising-time Diffusion Alignment with Multiple Objectives
by: Zhang, Qi, et al.
Published: (2026)
by: Zhang, Qi, et al.
Published: (2026)
Visual Explanations of Image-Text Representations via Multi-Modal Information Bottleneck Attribution
by: Wang, Ying, et al.
Published: (2023)
by: Wang, Ying, et al.
Published: (2023)
(Almost) Free Modality Stitching of Foundation Models
by: Singh, Jaisidh, et al.
Published: (2025)
by: Singh, Jaisidh, et al.
Published: (2025)
End-to-End Multi-Modal Diffusion Mamba
by: Lu, Chunhao, et al.
Published: (2025)
by: Lu, Chunhao, et al.
Published: (2025)
GSNeRF: Generalizable Semantic Neural Radiance Fields with Enhanced 3D Scene Understanding
by: Chou, Zi-Ting, et al.
Published: (2024)
by: Chou, Zi-Ting, et al.
Published: (2024)
Multi-Modal Adapter for Vision-Language Models
by: Seputis, Dominykas, et al.
Published: (2024)
by: Seputis, Dominykas, et al.
Published: (2024)
LLMs Can Evolve Continually on Modality for X-Modal Reasoning
by: Yu, Jiazuo, et al.
Published: (2024)
by: Yu, Jiazuo, et al.
Published: (2024)
The Scene Language: Representing Scenes with Programs, Words, and Embeddings
by: Zhang, Yunzhi, et al.
Published: (2024)
by: Zhang, Yunzhi, et al.
Published: (2024)
Similar Items
-
Taming Diffusion for Dataset Distillation with High Representativeness
by: Zhao, Lin, et al.
Published: (2025) -
MultiOOD: Scaling Out-of-Distribution Detection for Multiple Modalities
by: Dong, Hao, et al.
Published: (2024) -
MMGDreamer: Mixed-Modality Graph for Geometry-Controllable 3D Indoor Scene Generation
by: Yang, Zhifei, et al.
Published: (2025) -
Self-Supervised Multi-Frame Neural Scene Flow
by: Liu, Dongrui, et al.
Published: (2024) -
Beyond Unimodal Learning: The Importance of Integrating Multiple Modalities for Lifelong Learning
by: Sarfraz, Fahad, et al.
Published: (2024)