Salvato in:
| Autori principali: | Wu, Weijia, Zhu, Zeyu, Shou, Mike Zheng |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2503.07314 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Multi-human Interactive Talking Dataset
di: Zhu, Zeyu, et al.
Pubblicazione: (2025)
di: Zhu, Zeyu, et al.
Pubblicazione: (2025)
MovieBench: A Hierarchical Movie Level Dataset for Long Video Generation
di: Wu, Weijia, et al.
Pubblicazione: (2024)
di: Wu, Weijia, et al.
Pubblicazione: (2024)
DoraCycle: Domain-Oriented Adaptation of Unified Generative Model in Multimodal Cycles
di: Zhao, Rui, et al.
Pubblicazione: (2025)
di: Zhao, Rui, et al.
Pubblicazione: (2025)
UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning
di: Mao, Weijia, et al.
Pubblicazione: (2025)
di: Mao, Weijia, et al.
Pubblicazione: (2025)
Mitty: Diffusion-based Human-to-Robot Video Generation
di: Song, Yiren, et al.
Pubblicazione: (2025)
di: Song, Yiren, et al.
Pubblicazione: (2025)
Long-Context Autoregressive Video Modeling with Next-Frame Prediction
di: Gu, Yuchao, et al.
Pubblicazione: (2025)
di: Gu, Yuchao, et al.
Pubblicazione: (2025)
UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths
di: Mao, Weijia, et al.
Pubblicazione: (2025)
di: Mao, Weijia, et al.
Pubblicazione: (2025)
The Image as Its Own Reward: Reinforcement Learning with Adversarial Reward for Image Generation
di: Mao, Weijia, et al.
Pubblicazione: (2025)
di: Mao, Weijia, et al.
Pubblicazione: (2025)
FakeVLM-R1: Internalizing Physical Laws via CoT for Synthetic Image Detection
di: Zhu, Leqi, et al.
Pubblicazione: (2026)
di: Zhu, Leqi, et al.
Pubblicazione: (2026)
DiffuMask: Synthesizing Images with Pixel-level Annotations for Semantic Segmentation Using Diffusion Models
di: Wu, Weijia, et al.
Pubblicazione: (2023)
di: Wu, Weijia, et al.
Pubblicazione: (2023)
Paper2Video: Automatic Video Generation from Scientific Papers
di: Zhu, Zeyu, et al.
Pubblicazione: (2025)
di: Zhu, Zeyu, et al.
Pubblicazione: (2025)
MakeAnything: Harnessing Diffusion Transformers for Multi-Domain Procedural Sequence Generation
di: Song, Yiren, et al.
Pubblicazione: (2025)
di: Song, Yiren, et al.
Pubblicazione: (2025)
P-Flow: Prompting Visual Effects Generation
di: Zhao, Rui, et al.
Pubblicazione: (2026)
di: Zhao, Rui, et al.
Pubblicazione: (2026)
D-AR: Diffusion via Autoregressive Models
di: Gao, Ziteng, et al.
Pubblicazione: (2025)
di: Gao, Ziteng, et al.
Pubblicazione: (2025)
Soap2Soap: Long Cinematic Video Remaking via Multi-Agent Collaboration
di: Song, Yiren, et al.
Pubblicazione: (2026)
di: Song, Yiren, et al.
Pubblicazione: (2026)
Towards Automated Movie Trailer Generation
di: Argaw, Dawit Mureja, et al.
Pubblicazione: (2024)
di: Argaw, Dawit Mureja, et al.
Pubblicazione: (2024)
VLog: Video-Language Models by Generative Retrieval of Narration Vocabulary
di: Lin, Kevin Qinghong, et al.
Pubblicazione: (2025)
di: Lin, Kevin Qinghong, et al.
Pubblicazione: (2025)
Paragraph-to-Image Generation with Information-Enriched Diffusion Model
di: Wu, Weijia, et al.
Pubblicazione: (2023)
di: Wu, Weijia, et al.
Pubblicazione: (2023)
PANDA: Towards Generalist Video Anomaly Detection via Agentic AI Engineer
di: Yang, Zhiwei, et al.
Pubblicazione: (2025)
di: Yang, Zhiwei, et al.
Pubblicazione: (2025)
LayerTracer: Cognitive-Aligned Layered SVG Synthesis via Diffusion Transformer
di: Song, Yiren, et al.
Pubblicazione: (2025)
di: Song, Yiren, et al.
Pubblicazione: (2025)
TPDiff: Temporal Pyramid Video Diffusion Model
di: Ran, Lingmin, et al.
Pubblicazione: (2025)
di: Ran, Lingmin, et al.
Pubblicazione: (2025)
Ego-centric Predictive Model Conditioned on Hand Trajectories
di: Zhang, Binjie, et al.
Pubblicazione: (2025)
di: Zhang, Binjie, et al.
Pubblicazione: (2025)
LaV-CoT: Language-Aware Visual CoT with Multi-Aspect Reward Optimization for Real-World Multilingual VQA
di: Huang, Jing, et al.
Pubblicazione: (2025)
di: Huang, Jing, et al.
Pubblicazione: (2025)
DraCo: Draft as CoT for Text-to-Image Preview and Rare Concept Generation
di: Jiang, Dongzhi, et al.
Pubblicazione: (2025)
di: Jiang, Dongzhi, et al.
Pubblicazione: (2025)
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing
di: Zheng, Junjie, et al.
Pubblicazione: (2025)
di: Zheng, Junjie, et al.
Pubblicazione: (2025)
RingID: Rethinking Tree-Ring Watermarking for Enhanced Multi-Key Identification
di: Ci, Hai, et al.
Pubblicazione: (2024)
di: Ci, Hai, et al.
Pubblicazione: (2024)
OmniPSD: Layered PSD Generation with Diffusion Transformer
di: Liu, Cheng, et al.
Pubblicazione: (2025)
di: Liu, Cheng, et al.
Pubblicazione: (2025)
X-Humanoid: Robotize Human Videos to Generate Humanoid Videos at Scale
di: Yang, Pei, et al.
Pubblicazione: (2025)
di: Yang, Pei, et al.
Pubblicazione: (2025)
StreamingEffect: Real-Time Human-Centric Video Effect Generation
di: Song, Yiren, et al.
Pubblicazione: (2026)
di: Song, Yiren, et al.
Pubblicazione: (2026)
IDProtector: An Adversarial Noise Encoder to Protect Against ID-Preserving Image Generation
di: Song, Yiren, et al.
Pubblicazione: (2024)
di: Song, Yiren, et al.
Pubblicazione: (2024)
Edit Transfer: Learning Image Editing via Vision In-Context Relations
di: Chen, Lan, et al.
Pubblicazione: (2025)
di: Chen, Lan, et al.
Pubblicazione: (2025)
The Consistency Critic: Correcting Inconsistencies in Generated Images via Reference-Guided Attentive Alignment
di: Ouyang, Ziheng, et al.
Pubblicazione: (2025)
di: Ouyang, Ziheng, et al.
Pubblicazione: (2025)
Empowering Lightweight MLLMs with Reasoning via Long CoT SFT
di: Ou, Linyu, et al.
Pubblicazione: (2025)
di: Ou, Linyu, et al.
Pubblicazione: (2025)
VideoAgent2: Enhancing the LLM-Based Agent System for Long-Form Video Understanding by Uncertainty-Aware CoT
di: Zhi, Zhuo, et al.
Pubblicazione: (2025)
di: Zhi, Zhuo, et al.
Pubblicazione: (2025)
SurvAgent: Hierarchical CoT-Enhanced Case Banking and Dichotomy-Based Multi-Agent System for Multimodal Survival Prediction
di: Huang, Guolin, et al.
Pubblicazione: (2025)
di: Huang, Guolin, et al.
Pubblicazione: (2025)
VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video Prompting
di: Ilaslan, Muhammet Furkan, et al.
Pubblicazione: (2024)
di: Ilaslan, Muhammet Furkan, et al.
Pubblicazione: (2024)
MC-CoT: A Modular Collaborative CoT Framework for Zero-shot Medical-VQA with LLM and MLLM Integration
di: Wei, Lai, et al.
Pubblicazione: (2024)
di: Wei, Lai, et al.
Pubblicazione: (2024)
Parrot Captions Teach CLIP to Spot Text
di: Lin, Yiqi, et al.
Pubblicazione: (2023)
di: Lin, Yiqi, et al.
Pubblicazione: (2023)
Benchmarking Multimodal CoT Reward Model Stepwise by Visual Program
di: Gao, Minghe, et al.
Pubblicazione: (2025)
di: Gao, Minghe, et al.
Pubblicazione: (2025)
Multi-Modal Generative Embedding Model
di: Ma, Feipeng, et al.
Pubblicazione: (2024)
di: Ma, Feipeng, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Multi-human Interactive Talking Dataset
di: Zhu, Zeyu, et al.
Pubblicazione: (2025) -
MovieBench: A Hierarchical Movie Level Dataset for Long Video Generation
di: Wu, Weijia, et al.
Pubblicazione: (2024) -
DoraCycle: Domain-Oriented Adaptation of Unified Generative Model in Multimodal Cycles
di: Zhao, Rui, et al.
Pubblicazione: (2025) -
UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning
di: Mao, Weijia, et al.
Pubblicazione: (2025) -
Mitty: Diffusion-based Human-to-Robot Video Generation
di: Song, Yiren, et al.
Pubblicazione: (2025)