WorldGPT: Empowering LLM as Multimodal World Model
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ge, Zhiqi, Huang, Hongzhe, Zhou, Mingze, Li, Juncheng, Wang, Guoming, Tang, Siliang, Zhuang, Yueting |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
InstructVid2Vid: Controllable Video Editing with Natural Language Instructions
von: Qin, Bosheng, et al.
Veröffentlicht: (2023)
von: Qin, Bosheng, et al.
Veröffentlicht: (2023)
LiveK12Bench: Have Large Multimodal Models Truly Conquered High School-level Examinations?
von: Wang, Xiaohan, et al.
Veröffentlicht: (2026)
von: Wang, Xiaohan, et al.
Veröffentlicht: (2026)
Modeling Human Responses to Multimodal AI Content
von: Shen, Zhiqi, et al.
Veröffentlicht: (2025)
von: Shen, Zhiqi, et al.
Veröffentlicht: (2025)
Learning Primitive Embodied World Models: Towards Scalable Robotic Learning
von: Sun, Qiao, et al.
Veröffentlicht: (2025)
von: Sun, Qiao, et al.
Veröffentlicht: (2025)
Semantic Item Graph Enhancement for Multimodal Recommendation
von: Zhang, Xiaoxiong, et al.
Veröffentlicht: (2025)
von: Zhang, Xiaoxiong, et al.
Veröffentlicht: (2025)
LLMER: Crafting Interactive Extended Reality Worlds with JSON Data Generated by Large Language Models
von: Chen, Jiangong, et al.
Veröffentlicht: (2025)
von: Chen, Jiangong, et al.
Veröffentlicht: (2025)
Prolonged Reasoning Is Not All You Need: Certainty-Based Adaptive Routing for Efficient LLM/MLLM Reasoning
von: Lu, Jinghui, et al.
Veröffentlicht: (2025)
von: Lu, Jinghui, et al.
Veröffentlicht: (2025)
PRISM-XR: Empowering Privacy-Aware XR Collaboration with Multimodal Large Language Models
von: Chen, Jiangong, et al.
Veröffentlicht: (2026)
von: Chen, Jiangong, et al.
Veröffentlicht: (2026)
FashionReGen: LLM-Empowered Fashion Report Generation
von: Ding, Yujuan, et al.
Veröffentlicht: (2024)
von: Ding, Yujuan, et al.
Veröffentlicht: (2024)
GMFVAD: Using Grained Multi-modal Feature to Improve Video Anomaly Detection
von: Dai, Guangyu, et al.
Veröffentlicht: (2025)
von: Dai, Guangyu, et al.
Veröffentlicht: (2025)
LLM-Guided Semantic Relational Reasoning for Multimodal Intent Recognition
von: Zhou, Qianrui, et al.
Veröffentlicht: (2025)
von: Zhou, Qianrui, et al.
Veröffentlicht: (2025)
MaLoRA: Gated Modality LoRA for Key-Space Alignment in Multimodal LLM Fine-Tuning
von: Zheng, Xinhan, et al.
Veröffentlicht: (2025)
von: Zheng, Xinhan, et al.
Veröffentlicht: (2025)
Distilling Vision-Language Foundation Models: A Data-Free Approach via Prompt Diversification
von: Xuan, Yunyi, et al.
Veröffentlicht: (2024)
von: Xuan, Yunyi, et al.
Veröffentlicht: (2024)
A Survey on Multimodal Benchmarks: In the Era of Large AI Models
von: Li, Lin, et al.
Veröffentlicht: (2024)
von: Li, Lin, et al.
Veröffentlicht: (2024)
QMAVIS: Long Video-Audio Understanding using Fusion of Large Multimodal Models
von: Lin, Zixing, et al.
Veröffentlicht: (2026)
von: Lin, Zixing, et al.
Veröffentlicht: (2026)
Differential Multimodal Transformers
von: Li, Jerry, et al.
Veröffentlicht: (2025)
von: Li, Jerry, et al.
Veröffentlicht: (2025)
KEN: Knowledge Augmentation and Emotion Guidance Network for Multimodal Fake News Detection
von: Zhu, Peican, et al.
Veröffentlicht: (2025)
von: Zhu, Peican, et al.
Veröffentlicht: (2025)
Personalized Image Generation with Large Multimodal Models
von: Xu, Yiyan, et al.
Veröffentlicht: (2024)
von: Xu, Yiyan, et al.
Veröffentlicht: (2024)
SEER: Semantic Enhancement and Emotional Reasoning Network for Multimodal Fake News Detection
von: Zhu, Peican, et al.
Veröffentlicht: (2025)
von: Zhu, Peican, et al.
Veröffentlicht: (2025)
Why We Feel: Breaking Boundaries in Emotional Reasoning with Multimodal Large Language Models
von: Lin, Yuxiang, et al.
Veröffentlicht: (2025)
von: Lin, Yuxiang, et al.
Veröffentlicht: (2025)
Image is All You Need to Empower Large-scale Diffusion Models for In-Domain Generation
von: Cao, Pu, et al.
Veröffentlicht: (2023)
von: Cao, Pu, et al.
Veröffentlicht: (2023)
Synthesizing Sentiment-Controlled Feedback For Multimodal Text and Image Data
von: Kumar, Puneet, et al.
Veröffentlicht: (2024)
von: Kumar, Puneet, et al.
Veröffentlicht: (2024)
Enhancing Multimodal Retrieval via Complementary Information Extraction and Alignment
von: Zeng, Delong, et al.
Veröffentlicht: (2026)
von: Zeng, Delong, et al.
Veröffentlicht: (2026)
Has Multimodal Learning Delivered Universal Intelligence in Healthcare? A Comprehensive Survey
von: Lin, Qika, et al.
Veröffentlicht: (2024)
von: Lin, Qika, et al.
Veröffentlicht: (2024)
GeoGuess: Multimodal Reasoning based on Hierarchy of Visual Information in Street View
von: Cheng, Fenghua, et al.
Veröffentlicht: (2025)
von: Cheng, Fenghua, et al.
Veröffentlicht: (2025)
Can Large Language Models Help Multimodal Language Analysis? MMLA: A Comprehensive Benchmark
von: Zhang, Hanlei, et al.
Veröffentlicht: (2025)
von: Zhang, Hanlei, et al.
Veröffentlicht: (2025)
Robust Modality-incomplete Anomaly Detection: A Modality-instructive Framework with Benchmark
von: Miao, Bingchen, et al.
Veröffentlicht: (2024)
von: Miao, Bingchen, et al.
Veröffentlicht: (2024)
MixEval-X: Any-to-Any Evaluations from Real-World Data Mixtures
von: Ni, Jinjie, et al.
Veröffentlicht: (2024)
von: Ni, Jinjie, et al.
Veröffentlicht: (2024)
Emotion-LLaMA: Multimodal Emotion Recognition and Reasoning with Instruction Tuning
von: Cheng, Zebang, et al.
Veröffentlicht: (2024)
von: Cheng, Zebang, et al.
Veröffentlicht: (2024)
RA-BLIP: Multimodal Adaptive Retrieval-Augmented Bootstrapping Language-Image Pre-training
von: Ding, Muhe, et al.
Veröffentlicht: (2024)
von: Ding, Muhe, et al.
Veröffentlicht: (2024)
PSA-MF: Personality-Sentiment Aligned Multi-Level Fusion for Multimodal Sentiment Analysis
von: Xie, Heng, et al.
Veröffentlicht: (2025)
von: Xie, Heng, et al.
Veröffentlicht: (2025)
Tri-Subspaces Disentanglement for Multimodal Sentiment Analysis
von: Meng, Chunlei, et al.
Veröffentlicht: (2026)
von: Meng, Chunlei, et al.
Veröffentlicht: (2026)
SynthGuard: An Open Platform for Detecting AI-Generated Multimedia with Multimodal LLMs
von: Desai, Shail, et al.
Veröffentlicht: (2025)
von: Desai, Shail, et al.
Veröffentlicht: (2025)
OmniMER: Auxiliary-Enhanced LLM Adaptation for Indonesian Multimodal Emotion Recognition
von: Yan, Xueming, et al.
Veröffentlicht: (2025)
von: Yan, Xueming, et al.
Veröffentlicht: (2025)
WorldGPT: A Sora-Inspired Video AI Agent as Rich World Models from Text and Image Inputs
von: Yang, Deshun, et al.
Veröffentlicht: (2024)
von: Yang, Deshun, et al.
Veröffentlicht: (2024)
Interpretable Multimodal Misinformation Detection with Logic Reasoning
von: Liu, Hui, et al.
Veröffentlicht: (2023)
von: Liu, Hui, et al.
Veröffentlicht: (2023)
When Harmful Content Gets Camouflaged: Unveiling Perception Failure of LVLMs with CamHarmTI
von: Li, Yanhui, et al.
Veröffentlicht: (2025)
von: Li, Yanhui, et al.
Veröffentlicht: (2025)
Multimodal Emotion Recognition by Fusing Video Semantic in MOOC Learning Scenarios
von: Zhang, Yuan, et al.
Veröffentlicht: (2024)
von: Zhang, Yuan, et al.
Veröffentlicht: (2024)
RW-Post: Auditable Evidence-Grounded Multimodal Fact-Checking in the Wild
von: Xu, Danni, et al.
Veröffentlicht: (2026)
von: Xu, Danni, et al.
Veröffentlicht: (2026)
HiQuE: Hierarchical Question Embedding Network for Multimodal Depression Detection
von: Jung, Juho, et al.
Veröffentlicht: (2024)
von: Jung, Juho, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
InstructVid2Vid: Controllable Video Editing with Natural Language Instructions
von: Qin, Bosheng, et al.
Veröffentlicht: (2023) -
LiveK12Bench: Have Large Multimodal Models Truly Conquered High School-level Examinations?
von: Wang, Xiaohan, et al.
Veröffentlicht: (2026) -
Modeling Human Responses to Multimodal AI Content
von: Shen, Zhiqi, et al.
Veröffentlicht: (2025) -
Learning Primitive Embodied World Models: Towards Scalable Robotic Learning
von: Sun, Qiao, et al.
Veröffentlicht: (2025) -
Semantic Item Graph Enhancement for Multimodal Recommendation
von: Zhang, Xiaoxiong, et al.
Veröffentlicht: (2025)