MicroVerse: A Preliminary Exploration Toward a Micro-World Simulation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Rongsheng, Wu, Minghao, Zhou, Hongru, Yu, Zhihan, Cai, Zhenyang, Chen, Junying, Wang, Benyou |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MedGen: Unlocking Medical Video Generation by Scaling Granularly-annotated Medical Videos
von: Wang, Rongsheng, et al.
Veröffentlicht: (2025)
von: Wang, Rongsheng, et al.
Veröffentlicht: (2025)
Exploring Compositional Generalization of Multimodal LLMs for Medical Imaging
von: Cai, Zhenyang, et al.
Veröffentlicht: (2024)
von: Cai, Zhenyang, et al.
Veröffentlicht: (2024)
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
von: Chen, Junying, et al.
Veröffentlicht: (2025)
von: Chen, Junying, et al.
Veröffentlicht: (2025)
ShizhenGPT: Towards Multimodal LLMs for Traditional Chinese Medicine
von: Chen, Junying, et al.
Veröffentlicht: (2025)
von: Chen, Junying, et al.
Veröffentlicht: (2025)
HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale
von: Chen, Junying, et al.
Veröffentlicht: (2024)
von: Chen, Junying, et al.
Veröffentlicht: (2024)
UrbanVerse: Scaling Urban Simulation by Watching City-Tour Videos
von: Liu, Mingxuan, et al.
Veröffentlicht: (2025)
von: Liu, Mingxuan, et al.
Veröffentlicht: (2025)
LongLLaVA: Scaling Multi-modal LLMs to 1000 Images Efficiently via a Hybrid Architecture
von: Wang, Xidong, et al.
Veröffentlicht: (2024)
von: Wang, Xidong, et al.
Veröffentlicht: (2024)
A Benchmark for Incremental Micro-expression Recognition
von: Lai, Zhengqin, et al.
Veröffentlicht: (2025)
von: Lai, Zhengqin, et al.
Veröffentlicht: (2025)
Virgo: A Preliminary Exploration on Reproducing o1-like MLLM
von: Du, Yifan, et al.
Veröffentlicht: (2025)
von: Du, Yifan, et al.
Veröffentlicht: (2025)
DreamSAC: Learning Hamiltonian World Models via Symmetry Exploration
von: Tang, Jinzhou, et al.
Veröffentlicht: (2026)
von: Tang, Jinzhou, et al.
Veröffentlicht: (2026)
ShareVerse: Multi-Agent Consistent Video Generation for Shared World Modeling
von: Zhu, Jiayi, et al.
Veröffentlicht: (2026)
von: Zhu, Jiayi, et al.
Veröffentlicht: (2026)
Can Multimodal LLMs See Materials Clearly? A Multimodal Benchmark on Materials Characterization
von: Lai, Zhengzhao, et al.
Veröffentlicht: (2025)
von: Lai, Zhengzhao, et al.
Veröffentlicht: (2025)
VerseCrafter: Dynamic Realistic Video World Model with 4D Geometric Control
von: Zheng, Sixiao, et al.
Veröffentlicht: (2026)
von: Zheng, Sixiao, et al.
Veröffentlicht: (2026)
TalkVerse: Democratizing Minute-Long Audio-Driven Video Generation
von: Wang, Zhenzhi, et al.
Veröffentlicht: (2025)
von: Wang, Zhenzhi, et al.
Veröffentlicht: (2025)
WorldEdit: Towards Open-World Image Editing with a Knowledge-Informed Benchmark
von: Lin, Wang, et al.
Veröffentlicht: (2026)
von: Lin, Wang, et al.
Veröffentlicht: (2026)
MA-Bench: Towards Fine-grained Micro-Action Understanding
von: Li, Kun, et al.
Veröffentlicht: (2026)
von: Li, Kun, et al.
Veröffentlicht: (2026)
Beyond ViT Tokens: Masked-Diffusion Pretrained Convolutional Pathology Foundation Model for Cell-Level Dense Prediction
von: Chen, Weiming, et al.
Veröffentlicht: (2026)
von: Chen, Weiming, et al.
Veröffentlicht: (2026)
A Preliminary Exploration Towards General Image Restoration
von: Kong, Xiangtao, et al.
Veröffentlicht: (2024)
von: Kong, Xiangtao, et al.
Veröffentlicht: (2024)
Prototype Learning for Micro-gesture Classification
von: Chen, Guoliang, et al.
Veröffentlicht: (2024)
von: Chen, Guoliang, et al.
Veröffentlicht: (2024)
DRRNet: Macro-Micro Feature Fusion and Dual Reverse Refinement for Camouflaged Object Detection
von: Sun, Jianlin, et al.
Veröffentlicht: (2025)
von: Sun, Jianlin, et al.
Veröffentlicht: (2025)
DeepVerse: 4D Autoregressive Video Generation as a World Model
von: Chen, Junyi, et al.
Veröffentlicht: (2025)
von: Chen, Junyi, et al.
Veröffentlicht: (2025)
NutritionVerse: Empirical Study of Various Dietary Intake Estimation Approaches
von: Tai, Chi-en Amy, et al.
Veröffentlicht: (2023)
von: Tai, Chi-en Amy, et al.
Veröffentlicht: (2023)
Adv-CPG: A Customized Portrait Generation Framework with Facial Adversarial Attacks
von: Wang, Junying, et al.
Veröffentlicht: (2025)
von: Wang, Junying, et al.
Veröffentlicht: (2025)
Boosting Generalizability towards Zero-Shot Cross-Dataset Single-Image Indoor Depth by Meta-Initialization
von: Wu, Cho-Ying, et al.
Veröffentlicht: (2024)
von: Wu, Cho-Ying, et al.
Veröffentlicht: (2024)
Sekai: A Video Dataset towards World Exploration
von: Li, Zhen, et al.
Veröffentlicht: (2025)
von: Li, Zhen, et al.
Veröffentlicht: (2025)
MM-Gesture: Towards Precise Micro-Gesture Recognition through Multimodal Fusion
von: Gu, Jihao, et al.
Veröffentlicht: (2025)
von: Gu, Jihao, et al.
Veröffentlicht: (2025)
Latent Bias Alignment for High-Fidelity Diffusion Inversion in Real-World Image Reconstruction and Manipulation
von: Chen, Weiming, et al.
Veröffentlicht: (2026)
von: Chen, Weiming, et al.
Veröffentlicht: (2026)
MetaScope: Optics-Driven Neural Network for Ultra-Micro Metalens Endoscopy
von: Li, Wuyang, et al.
Veröffentlicht: (2025)
von: Li, Wuyang, et al.
Veröffentlicht: (2025)
EditWorld: Simulating World Dynamics for Instruction-Following Image Editing
von: Yang, Ling, et al.
Veröffentlicht: (2024)
von: Yang, Ling, et al.
Veröffentlicht: (2024)
TalkVid: A Large-Scale Diversified Dataset for Audio-Driven Talking Head Synthesis
von: Chen, Shunian, et al.
Veröffentlicht: (2025)
von: Chen, Shunian, et al.
Veröffentlicht: (2025)
MicroAUNet: Boundary-Enhanced Multi-scale Fusion with Knowledge Distillation for Colonoscopy Polyp Image Segmentation
von: Wang, Ziyi, et al.
Veröffentlicht: (2025)
von: Wang, Ziyi, et al.
Veröffentlicht: (2025)
PanoWorld: Towards Spatial Supersensing in 360$^\circ$ Panorama World
von: Wang, Changpeng, et al.
Veröffentlicht: (2026)
von: Wang, Changpeng, et al.
Veröffentlicht: (2026)
Towards Efficient Diffusion-Based Image Editing with Instant Attention Masks
von: Zou, Siyu, et al.
Veröffentlicht: (2024)
von: Zou, Siyu, et al.
Veröffentlicht: (2024)
Temporal and Spatial Feature Fusion Framework for Dynamic Micro Expression Recognition
von: Liu, Feng, et al.
Veröffentlicht: (2025)
von: Liu, Feng, et al.
Veröffentlicht: (2025)
Exploring the Limits of Semantic Image Compression at Micro-bits per Pixel
von: Dotzel, Jordan, et al.
Veröffentlicht: (2024)
von: Dotzel, Jordan, et al.
Veröffentlicht: (2024)
Micro-expression Recognition Based on Dual-branch Feature Extraction and Fusion
von: Zhang, Mingjie, et al.
Veröffentlicht: (2026)
von: Zhang, Mingjie, et al.
Veröffentlicht: (2026)
Generative Semantic Coding for Ultra-Low Bitrate Visual Communication and Analysis
von: Chen, Weiming, et al.
Veröffentlicht: (2025)
von: Chen, Weiming, et al.
Veröffentlicht: (2025)
Runge-Kutta Approximation and Decoupled Attention for Rectified Flow Inversion and Semantic Editing
von: Chen, Weiming, et al.
Veröffentlicht: (2025)
von: Chen, Weiming, et al.
Veröffentlicht: (2025)
Towards Multi-dimensional Explanation Alignment for Medical Classification
von: Hu, Lijie, et al.
Veröffentlicht: (2024)
von: Hu, Lijie, et al.
Veröffentlicht: (2024)
ModaVerse: Efficiently Transforming Modalities with LLMs
von: Wang, Xinyu, et al.
Veröffentlicht: (2024)
von: Wang, Xinyu, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
MedGen: Unlocking Medical Video Generation by Scaling Granularly-annotated Medical Videos
von: Wang, Rongsheng, et al.
Veröffentlicht: (2025) -
Exploring Compositional Generalization of Multimodal LLMs for Medical Imaging
von: Cai, Zhenyang, et al.
Veröffentlicht: (2024) -
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
von: Chen, Junying, et al.
Veröffentlicht: (2025) -
ShizhenGPT: Towards Multimodal LLMs for Traditional Chinese Medicine
von: Chen, Junying, et al.
Veröffentlicht: (2025) -
HuatuoGPT-Vision, Towards Injecting Medical Visual Knowledge into Multimodal LLMs at Scale
von: Chen, Junying, et al.
Veröffentlicht: (2024)