Explore the Limits of Omni-modal Pretraining at Scale
Fuente:
arXiv
Guardado en:
| Autores principales: | Zhang, Yiyuan, Li, Handong, Liu, Jing, Yue, Xiangyu |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
InteractiveVideo: User-Centric Controllable Video Generation with Synergistic Multimodal Instructions
por: Zhang, Yiyuan, et al.
Publicado: (2024)
por: Zhang, Yiyuan, et al.
Publicado: (2024)
Multimodal Long Video Modeling Based on Temporal Dynamic Context
por: Hao, Haoran, et al.
Publicado: (2025)
por: Hao, Haoran, et al.
Publicado: (2025)
OneLLM: One Framework to Align All Modalities with Language
por: Han, Jiaming, et al.
Publicado: (2023)
por: Han, Jiaming, et al.
Publicado: (2023)
OmniGAIA: Towards Native Omni-Modal AI Agents
por: Li, Xiaoxi, et al.
Publicado: (2026)
por: Li, Xiaoxi, et al.
Publicado: (2026)
OmniEvalKit: A Modular, Lightweight Toolbox for Evaluating Large Language Model and its Omni-Extensions
por: Zhang, Yi-Kai, et al.
Publicado: (2024)
por: Zhang, Yi-Kai, et al.
Publicado: (2024)
Audio-Omni: Extending Multi-modal Understanding to Versatile Audio Generation and Editing
por: Tian, Zeyue, et al.
Publicado: (2026)
por: Tian, Zeyue, et al.
Publicado: (2026)
Enhancing multimodal cooperation via sample-level modality valuation
por: Wei, Yake, et al.
Publicado: (2023)
por: Wei, Yake, et al.
Publicado: (2023)
MAVEN: Multi-modal Attention for Valence-Arousal Emotion Network
por: Ahire, Vrushank, et al.
Publicado: (2025)
por: Ahire, Vrushank, et al.
Publicado: (2025)
Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation
por: Liu, Che, et al.
Publicado: (2026)
por: Liu, Che, et al.
Publicado: (2026)
RAP: Retrieval-Augmented Personalization for Multimodal Large Language Models
por: Hao, Haoran, et al.
Publicado: (2024)
por: Hao, Haoran, et al.
Publicado: (2024)
GoodDrag: Towards Good Practices for Drag Editing with Diffusion Models
por: Zhang, Zewei, et al.
Publicado: (2024)
por: Zhang, Zewei, et al.
Publicado: (2024)
MultiWay-Adapater: Adapting large-scale multi-modal models for scalable image-text retrieval
por: Long, Zijun, et al.
Publicado: (2023)
por: Long, Zijun, et al.
Publicado: (2023)
SMPLer: Taming Transformers for Monocular 3D Human Shape and Pose Estimation
por: Xu, Xiangyu, et al.
Publicado: (2024)
por: Xu, Xiangyu, et al.
Publicado: (2024)
How Do Images Align and Complement LiDAR? Towards a Harmonized Multi-modal 3D Panoptic Segmentation
por: Pan, Yining, et al.
Publicado: (2025)
por: Pan, Yining, et al.
Publicado: (2025)
Scaling Spatial Intelligence with Multimodal Foundation Models
por: Cai, Zhongang, et al.
Publicado: (2025)
por: Cai, Zhongang, et al.
Publicado: (2025)
DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video Generation
por: Cai, Minghong, et al.
Publicado: (2024)
por: Cai, Minghong, et al.
Publicado: (2024)
Instant3D: Instant Text-to-3D Generation
por: Li, Ming, et al.
Publicado: (2023)
por: Li, Ming, et al.
Publicado: (2023)
Art2Music: Generating Music for Art Images with Multi-modal Feeling Alignment
por: Hong, Jiaying, et al.
Publicado: (2025)
por: Hong, Jiaying, et al.
Publicado: (2025)
MGM-Omni: Scaling Omni LLMs to Personalized Long-Horizon Speech
por: Wang, Chengyao, et al.
Publicado: (2025)
por: Wang, Chengyao, et al.
Publicado: (2025)
Scaling Up Your Kernels: Large Kernel Design in ConvNets towards Universal Representations
por: Zhang, Yiyuan, et al.
Publicado: (2024)
por: Zhang, Yiyuan, et al.
Publicado: (2024)
Size Matters: Reconstructing Real-Scale 3D Models from Monocular Images for Food Portion Estimation
por: Vinod, Gautham, et al.
Publicado: (2026)
por: Vinod, Gautham, et al.
Publicado: (2026)
LayerT2V: A Unified Multi-Layer Video Generation Framework
por: Li, Guangzhao, et al.
Publicado: (2025)
por: Li, Guangzhao, et al.
Publicado: (2025)
Long-tailed Medical Diagnosis with Relation-aware Representation Learning and Iterative Classifier Calibration
por: Pan, Li, et al.
Publicado: (2025)
por: Pan, Li, et al.
Publicado: (2025)
Pay Less Attention to Deceptive Artifacts: Robust Detection of Compressed Deepfakes on Online Social Networks
por: Li, Manyi, et al.
Publicado: (2025)
por: Li, Manyi, et al.
Publicado: (2025)
VSTAR: Generative Temporal Nursing for Longer Dynamic Video Synthesis
por: Li, Yumeng, et al.
Publicado: (2024)
por: Li, Yumeng, et al.
Publicado: (2024)
A Systematic Review on Long-Tailed Learning
por: Zhang, Chongsheng, et al.
Publicado: (2024)
por: Zhang, Chongsheng, et al.
Publicado: (2024)
MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory Guidance
por: Li, Quanhao, et al.
Publicado: (2025)
por: Li, Quanhao, et al.
Publicado: (2025)
Vlogger: Make Your Dream A Vlog
por: Zhuang, Shaobin, et al.
Publicado: (2024)
por: Zhuang, Shaobin, et al.
Publicado: (2024)
UniF$^2$ace: A Unified Fine-grained Face Understanding and Generation Model
por: Li, Junzhe, et al.
Publicado: (2025)
por: Li, Junzhe, et al.
Publicado: (2025)
David and Goliath: Small One-step Model Beats Large Diffusion with Score Post-training
por: Luo, Weijian, et al.
Publicado: (2024)
por: Luo, Weijian, et al.
Publicado: (2024)
PAND: Prompt-Aware Neighborhood Distillation for Lightweight Fine-Grained Visual Classification
por: Luo, Qiuming, et al.
Publicado: (2026)
por: Luo, Qiuming, et al.
Publicado: (2026)
STIV: Scalable Text and Image Conditioned Video Generation
por: Lin, Zongyu, et al.
Publicado: (2024)
por: Lin, Zongyu, et al.
Publicado: (2024)
Infinite Video Understanding
por: Zhang, Dell, et al.
Publicado: (2025)
por: Zhang, Dell, et al.
Publicado: (2025)
PlanLLM: Video Procedure Planning with Refinable Large Language Models
por: Yang, Dejie, et al.
Publicado: (2024)
por: Yang, Dejie, et al.
Publicado: (2024)
Boosting Facial Action Unit Detection Through Jointly Learning Facial Landmark Detection and Domain Separation and Reconstruction
por: Shang, Ziqiao, et al.
Publicado: (2023)
por: Shang, Ziqiao, et al.
Publicado: (2023)
Dual Memory Networks: A Versatile Adaptation Approach for Vision-Language Models
por: Zhang, Yabin, et al.
Publicado: (2024)
por: Zhang, Yabin, et al.
Publicado: (2024)
SGAT4PASS: Spherical Geometry-Aware Transformer for PAnoramic Semantic Segmentation
por: Li, Xuewei, et al.
Publicado: (2023)
por: Li, Xuewei, et al.
Publicado: (2023)
Are We Making Progress in Multimodal Domain Generalization? A Comprehensive Benchmark Study
por: Dong, Hao, et al.
Publicado: (2026)
por: Dong, Hao, et al.
Publicado: (2026)
Reducing Hallucinations in Vision-Language Models via Latent Space Steering
por: Liu, Sheng, et al.
Publicado: (2024)
por: Liu, Sheng, et al.
Publicado: (2024)
HOIN: High-Order Implicit Neural Representations
por: Chen, Yang, et al.
Publicado: (2024)
por: Chen, Yang, et al.
Publicado: (2024)
Ejemplares similares
-
InteractiveVideo: User-Centric Controllable Video Generation with Synergistic Multimodal Instructions
por: Zhang, Yiyuan, et al.
Publicado: (2024) -
Multimodal Long Video Modeling Based on Temporal Dynamic Context
por: Hao, Haoran, et al.
Publicado: (2025) -
OneLLM: One Framework to Align All Modalities with Language
por: Han, Jiaming, et al.
Publicado: (2023) -
OmniGAIA: Towards Native Omni-Modal AI Agents
por: Li, Xiaoxi, et al.
Publicado: (2026) -
OmniEvalKit: A Modular, Lightweight Toolbox for Evaluating Large Language Model and its Omni-Extensions
por: Zhang, Yi-Kai, et al.
Publicado: (2024)