A Versatile Multimodal Agent for Multimedia Content Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Daoan, Yao, Wenlin, Wang, Xiaoyang, Hu, Yebowen, Luo, Jiebo, Yu, Dong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image Inputs
von: Zhang, Daoan, et al.
Veröffentlicht: (2024)
von: Zhang, Daoan, et al.
Veröffentlicht: (2024)
LAST: LeArning to Think in Space and Time for Generalist Vision-Language Models
von: Wang, Shuai, et al.
Veröffentlicht: (2025)
von: Wang, Shuai, et al.
Veröffentlicht: (2025)
MIRA: Multimodal Iterative Reasoning Agent for Image Editing
von: Zeng, Ziyun, et al.
Veröffentlicht: (2025)
von: Zeng, Ziyun, et al.
Veröffentlicht: (2025)
Learning Brain Tumor Representation in 3D High-Resolution MR Images via Interpretable State Space Models
von: Hu, Qingqiao, et al.
Veröffentlicht: (2024)
von: Hu, Qingqiao, et al.
Veröffentlicht: (2024)
UniVA: Universal Video Agent towards Open-Source Next-Generation Video Generalist
von: Liang, Zhengyang, et al.
Veröffentlicht: (2025)
von: Liang, Zhengyang, et al.
Veröffentlicht: (2025)
WorldGenBench: A World-Knowledge-Integrated Benchmark for Reasoning-Driven Text-to-Image Generation
von: Zhang, Daoan, et al.
Veröffentlicht: (2025)
von: Zhang, Daoan, et al.
Veröffentlicht: (2025)
Chain-of-Thought Prompting for Demographic Inference with Large Multimodal Models
von: Yu, Yongsheng, et al.
Veröffentlicht: (2024)
von: Yu, Yongsheng, et al.
Veröffentlicht: (2024)
Ouroboros-Diffusion: Exploring Consistent Content Generation in Tuning-free Long Video Diffusion
von: Chen, Jingyuan, et al.
Veröffentlicht: (2025)
von: Chen, Jingyuan, et al.
Veröffentlicht: (2025)
GaussianStyle: Gaussian Head Avatar via StyleGAN
von: Liu, Pinxin, et al.
Veröffentlicht: (2024)
von: Liu, Pinxin, et al.
Veröffentlicht: (2024)
FINEMATCH: Aspect-based Fine-grained Image and Text Mismatch Detection and Correction
von: Hua, Hang, et al.
Veröffentlicht: (2024)
von: Hua, Hang, et al.
Veröffentlicht: (2024)
VisualActBench: Can VLMs See and Act like a Human?
von: Zhang, Daoan, et al.
Veröffentlicht: (2025)
von: Zhang, Daoan, et al.
Veröffentlicht: (2025)
MementoGUI: Learning Agentic Multimodal Memory Control for Long-Horizon GUI Agents
von: Zeng, Ziyun, et al.
Veröffentlicht: (2026)
von: Zeng, Ziyun, et al.
Veröffentlicht: (2026)
Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey
von: Wang, Yaoting, et al.
Veröffentlicht: (2025)
von: Wang, Yaoting, et al.
Veröffentlicht: (2025)
On Path to Multimodal Generalist: General-Level and General-Bench
von: Fei, Hao, et al.
Veröffentlicht: (2025)
von: Fei, Hao, et al.
Veröffentlicht: (2025)
Robust and Calibrated Detection of Authentic Multimedia Content
von: Hashmi, Sarim, et al.
Veröffentlicht: (2025)
von: Hashmi, Sarim, et al.
Veröffentlicht: (2025)
OmniV2V: Versatile Video Generation and Editing via Dynamic Content Manipulation
von: Liang, Sen, et al.
Veröffentlicht: (2025)
von: Liang, Sen, et al.
Veröffentlicht: (2025)
On Inductive Biases That Enable Generalization of Diffusion Transformers
von: An, Jie, et al.
Veröffentlicht: (2024)
von: An, Jie, et al.
Veröffentlicht: (2024)
Diffusion Transformer-to-Mamba Distillation for High-Resolution Image Generation
von: Yao, Yuan, et al.
Veröffentlicht: (2025)
von: Yao, Yuan, et al.
Veröffentlicht: (2025)
Latent-Reframe: Enabling Camera Control for Video Diffusion Model without Training
von: Zhou, Zhenghong, et al.
Veröffentlicht: (2024)
von: Zhou, Zhenghong, et al.
Veröffentlicht: (2024)
SurgLLM: A Versatile Large Multimodal Model with Spatial Focus and Temporal Awareness for Surgical Video Understanding
von: Chen, Zhen, et al.
Veröffentlicht: (2025)
von: Chen, Zhen, et al.
Veröffentlicht: (2025)
RMPL: Relation-aware Multi-task Progressive Learning with Stage-wise Training for Multimedia Event Extraction
von: Jin, Yongkang, et al.
Veröffentlicht: (2026)
von: Jin, Yongkang, et al.
Veröffentlicht: (2026)
Bridging SFT and DPO for Diffusion Model Alignment with Self-Sampling Preference Optimization
von: Zhang, Daoan, et al.
Veröffentlicht: (2024)
von: Zhang, Daoan, et al.
Veröffentlicht: (2024)
Aurora: Unified Video Editing with a Tool-Using Agent
von: Yu, Yongsheng, et al.
Veröffentlicht: (2026)
von: Yu, Yongsheng, et al.
Veröffentlicht: (2026)
OmniPaint: Mastering Object-Oriented Editing via Disentangled Insertion-Removal Inpainting
von: Yu, Yongsheng, et al.
Veröffentlicht: (2025)
von: Yu, Yongsheng, et al.
Veröffentlicht: (2025)
Benchmarking In-the-wild Multimodal Disease Recognition and A Versatile Baseline
von: Wei, Tianqi, et al.
Veröffentlicht: (2024)
von: Wei, Tianqi, et al.
Veröffentlicht: (2024)
PixelDiT: Pixel Diffusion Transformers for Image Generation
von: Yu, Yongsheng, et al.
Veröffentlicht: (2025)
von: Yu, Yongsheng, et al.
Veröffentlicht: (2025)
PixPerfect: Seamless Latent Diffusion Local Editing with Discriminative Pixel-Space Refinement
von: Zheng, Haitian, et al.
Veröffentlicht: (2025)
von: Zheng, Haitian, et al.
Veröffentlicht: (2025)
Evolver: Chain-of-Evolution Prompting to Boost Large Multimodal Models for Hateful Meme Detection
von: Huang, Jinfa, et al.
Veröffentlicht: (2024)
von: Huang, Jinfa, et al.
Veröffentlicht: (2024)
DocRefine: An Intelligent Framework for Scientific Document Understanding and Content Optimization based on Multimodal Large Model Agents
von: Qian, Kun, et al.
Veröffentlicht: (2025)
von: Qian, Kun, et al.
Veröffentlicht: (2025)
Structured 3D Latents for Scalable and Versatile 3D Generation
von: Xiang, Jianfeng, et al.
Veröffentlicht: (2024)
von: Xiang, Jianfeng, et al.
Veröffentlicht: (2024)
SurgVisAgent: Multimodal Agentic Model for Versatile Surgical Visual Enhancement
von: Lei, Zeyu, et al.
Veröffentlicht: (2025)
von: Lei, Zeyu, et al.
Veröffentlicht: (2025)
End-to-end Open-vocabulary Video Visual Relationship Detection using Multi-modal Prompting
von: Wang, Yongqi, et al.
Veröffentlicht: (2024)
von: Wang, Yongqi, et al.
Veröffentlicht: (2024)
Simulate, Refocus and Ensemble: An Attention-Refocusing Scheme for Domain Generalization
von: Wang, Ziyi, et al.
Veröffentlicht: (2025)
von: Wang, Ziyi, et al.
Veröffentlicht: (2025)
GIR-Bench: Versatile Benchmark for Generating Images with Reasoning
von: Li, Hongxiang, et al.
Veröffentlicht: (2025)
von: Li, Hongxiang, et al.
Veröffentlicht: (2025)
OmniMotion-X: Versatile Multimodal Whole-Body Motion Generation
von: Xu, Guowei, et al.
Veröffentlicht: (2025)
von: Xu, Guowei, et al.
Veröffentlicht: (2025)
VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models
von: Huang, Ziqi, et al.
Veröffentlicht: (2024)
von: Huang, Ziqi, et al.
Veröffentlicht: (2024)
Versatile Transition Generation with Image-to-Video Diffusion
von: Yang, Zuhao, et al.
Veröffentlicht: (2025)
von: Yang, Zuhao, et al.
Veröffentlicht: (2025)
FDIM: A Feature-distance-based Generic Video Quality Metric for Versatile Codecs
von: Wang, Jiayi, et al.
Veröffentlicht: (2026)
von: Wang, Jiayi, et al.
Veröffentlicht: (2026)
SpecEyes: Accelerating Agentic Multimodal LLMs via Speculative Perception and Planning
von: Huang, Haoyu, et al.
Veröffentlicht: (2026)
von: Huang, Haoyu, et al.
Veröffentlicht: (2026)
How Far Are Video Models from True Multimodal Reasoning?
von: Zhang, Xiaotian, et al.
Veröffentlicht: (2026)
von: Zhang, Xiaotian, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image Inputs
von: Zhang, Daoan, et al.
Veröffentlicht: (2024) -
LAST: LeArning to Think in Space and Time for Generalist Vision-Language Models
von: Wang, Shuai, et al.
Veröffentlicht: (2025) -
MIRA: Multimodal Iterative Reasoning Agent for Image Editing
von: Zeng, Ziyun, et al.
Veröffentlicht: (2025) -
Learning Brain Tumor Representation in 3D High-Resolution MR Images via Interpretable State Space Models
von: Hu, Qingqiao, et al.
Veröffentlicht: (2024) -
UniVA: Universal Video Agent towards Open-Source Next-Generation Video Generalist
von: Liang, Zhengyang, et al.
Veröffentlicht: (2025)