Saved in:
| Main Authors: | Wang, Xinyu, Zhuang, Bohan, Wu, Qi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2401.06395 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Are Large Vision Language Models Good Game Players?
by: Wang, Xinyu, et al.
Published: (2025)
by: Wang, Xinyu, et al.
Published: (2025)
MammothModa: Multi-Modal Large Language Model
by: She, Qi, et al.
Published: (2024)
by: She, Qi, et al.
Published: (2024)
Streaming Video Diffusion: Online Video Editing with Diffusion Models
by: Chen, Feng, et al.
Published: (2024)
by: Chen, Feng, et al.
Published: (2024)
EfficientDM: Efficient Quantization-Aware Fine-Tuning of Low-Bit Diffusion Models
by: He, Yefei, et al.
Published: (2023)
by: He, Yefei, et al.
Published: (2023)
ModaLink: Unifying Modalities for Efficient Image-to-PointCloud Place Recognition
by: Xie, Weidong, et al.
Published: (2024)
by: Xie, Weidong, et al.
Published: (2024)
CogniVerse: Revolutionizing Multi-Modal Retrieval-Augmented Generation with Cognitive Reflection and Geometric Reasoning
by: Fang, Xiang, et al.
Published: (2026)
by: Fang, Xiang, et al.
Published: (2026)
FlashAR: Efficient Post-Training Acceleration for Autoregressive Image Generation
by: Zhou, Junkang, et al.
Published: (2026)
by: Zhou, Junkang, et al.
Published: (2026)
FlashBlock: Attention Caching for Efficient Long-Context Block Diffusion
by: Chen, Zhuokun, et al.
Published: (2026)
by: Chen, Zhuokun, et al.
Published: (2026)
InfiniMotion: Mamba Boosts Memory in Transformer for Arbitrary Long Motion Generation
by: Zhang, Zeyu, et al.
Published: (2024)
by: Zhang, Zeyu, et al.
Published: (2024)
VideoVerse: Does Your T2V Generator Have World Model Capability to Synthesize Videos?
by: Wang, Zeqing, et al.
Published: (2025)
by: Wang, Zeqing, et al.
Published: (2025)
LongVLM: Efficient Long Video Understanding via Large Language Models
by: Weng, Yuetian, et al.
Published: (2024)
by: Weng, Yuetian, et al.
Published: (2024)
Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion
by: Chen, Zhuokun, et al.
Published: (2024)
by: Chen, Zhuokun, et al.
Published: (2024)
Exploiting Modality-Specific Features For Multi-Modal Manipulation Detection And Grounding
by: Wang, Jiazhen, et al.
Published: (2023)
by: Wang, Jiazhen, et al.
Published: (2023)
Motion Mamba: Efficient and Long Sequence Motion Generation
by: Zhang, Zeyu, et al.
Published: (2024)
by: Zhang, Zeyu, et al.
Published: (2024)
The SkatingVerse Workshop & Challenge: Methods and Results
by: Zhao, Jian, et al.
Published: (2024)
by: Zhao, Jian, et al.
Published: (2024)
1st Place Solution to the 1st SkatingVerse Challenge
by: Sun, Tao, et al.
Published: (2024)
by: Sun, Tao, et al.
Published: (2024)
WonderVerse: Extendable 3D Scene Generation with Video Generative Models
by: Feng, Hao, et al.
Published: (2025)
by: Feng, Hao, et al.
Published: (2025)
Efficient Stitchable Task Adaptation
by: He, Haoyu, et al.
Published: (2023)
by: He, Haoyu, et al.
Published: (2023)
RAPID^3: Tri-Level Reinforced Acceleration Policies for Diffusion Transformer
by: Zhao, Wangbo, et al.
Published: (2025)
by: Zhao, Wangbo, et al.
Published: (2025)
ComboVerse: Compositional 3D Assets Creation Using Spatially-Aware Diffusion Guidance
by: Chen, Yongwei, et al.
Published: (2024)
by: Chen, Yongwei, et al.
Published: (2024)
Frequency-Aware Autoregressive Modeling for Efficient High-Resolution Image Synthesis
by: Chen, Zhuokun, et al.
Published: (2025)
by: Chen, Zhuokun, et al.
Published: (2025)
An Empirical Study on How Video-LLMs Answer Video Questions
by: Gou, Chenhui, et al.
Published: (2025)
by: Gou, Chenhui, et al.
Published: (2025)
PSA: Pyramid Sparse Attention for Efficient Video Understanding and Generation
by: Li, Xiaolong, et al.
Published: (2025)
by: Li, Xiaolong, et al.
Published: (2025)
EchoingPixels: Cross-Modal Adaptive Token Reduction for Efficient Audio-Visual LLMs
by: Gong, Chao, et al.
Published: (2025)
by: Gong, Chao, et al.
Published: (2025)
ShotVerse: Advancing Cinematic Camera Control for Text-Driven Multi-Shot Video Creation
by: Yang, Songlin, et al.
Published: (2026)
by: Yang, Songlin, et al.
Published: (2026)
MammothModa2: A Unified AR-Diffusion Framework for Multimodal Understanding and Generation
by: Shen, Tao, et al.
Published: (2025)
by: Shen, Tao, et al.
Published: (2025)
MicroVerse: A Preliminary Exploration Toward a Micro-World Simulation
by: Wang, Rongsheng, et al.
Published: (2026)
by: Wang, Rongsheng, et al.
Published: (2026)
ChartVerse: Scaling Chart Reasoning via Reliable Programmatic Synthesis from Scratch
by: Liu, Zheng, et al.
Published: (2026)
by: Liu, Zheng, et al.
Published: (2026)
EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning
by: Ju, Xuan, et al.
Published: (2025)
by: Ju, Xuan, et al.
Published: (2025)
NeoVerse: Enhancing 4D World Model with in-the-wild Monocular Videos
by: Yang, Yuxue, et al.
Published: (2026)
by: Yang, Yuxue, et al.
Published: (2026)
CanoVerse: 3D Object Scalable Canonicalization and Dataset for Generation and Pose
by: Jin, Li, et al.
Published: (2026)
by: Jin, Li, et al.
Published: (2026)
BeetleVerse: A Study on Taxonomic Classification of Ground Beetles
by: Rayeed, S M, et al.
Published: (2025)
by: Rayeed, S M, et al.
Published: (2025)
CineVerse: Consistent Keyframe Synthesis for Cinematic Scene Composition
by: Phung, Quynh, et al.
Published: (2025)
by: Phung, Quynh, et al.
Published: (2025)
MirrorVerse: Pushing Diffusion Models to Realistically Reflect the World
by: Dhiman, Ankit, et al.
Published: (2025)
by: Dhiman, Ankit, et al.
Published: (2025)
MapKD: Unlocking Prior Knowledge with Cross-Modal Distillation for Efficient Online HD Map Construction
by: Yan, Ziyang, et al.
Published: (2025)
by: Yan, Ziyang, et al.
Published: (2025)
UniVerse: Unleashing the Scene Prior of Video Diffusion Models for Robust Radiance Field Reconstruction
by: Cao, Jin, et al.
Published: (2025)
by: Cao, Jin, et al.
Published: (2025)
GarmentDiffusion: 3D Garment Sewing Pattern Generation with Multimodal Diffusion Transformers
by: Li, Xinyu, et al.
Published: (2025)
by: Li, Xinyu, et al.
Published: (2025)
UniVerse-1: Unified Audio-Video Generation via Stitching of Experts
by: Wang, Duomin, et al.
Published: (2025)
by: Wang, Duomin, et al.
Published: (2025)
DeepVerse: 4D Autoregressive Video Generation as a World Model
by: Chen, Junyi, et al.
Published: (2025)
by: Chen, Junyi, et al.
Published: (2025)
Object-aware Inversion and Reassembly for Image Editing
by: Yang, Zhen, et al.
Published: (2023)
by: Yang, Zhen, et al.
Published: (2023)
Similar Items
-
Are Large Vision Language Models Good Game Players?
by: Wang, Xinyu, et al.
Published: (2025) -
MammothModa: Multi-Modal Large Language Model
by: She, Qi, et al.
Published: (2024) -
Streaming Video Diffusion: Online Video Editing with Diffusion Models
by: Chen, Feng, et al.
Published: (2024) -
EfficientDM: Efficient Quantization-Aware Fine-Tuning of Low-Bit Diffusion Models
by: He, Yefei, et al.
Published: (2023) -
ModaLink: Unifying Modalities for Efficient Image-to-PointCloud Place Recognition
by: Xie, Weidong, et al.
Published: (2024)