Vlogger: Make Your Dream A Vlog
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhuang, Shaobin, Li, Kunchang, Chen, Xinyuan, Wang, Yaohui, Liu, Ziwei, Qiao, Yu, Wang, Yali |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
LMVD: A Large-Scale Multimodal Vlog Dataset for Depression Detection in the Wild
di: He, Lang, et al.
Pubblicazione: (2024)
di: He, Lang, et al.
Pubblicazione: (2024)
Discover Your Neighbors: Advanced Stable Test-Time Adaptation in Dynamic World
di: Jiang, Qinting, et al.
Pubblicazione: (2024)
di: Jiang, Qinting, et al.
Pubblicazione: (2024)
TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning
di: Zeng, Xiangyu, et al.
Pubblicazione: (2024)
di: Zeng, Xiangyu, et al.
Pubblicazione: (2024)
VDC-Agent: When Video Detailed Captioners Evolve Themselves via Agentic Self-Reflection
di: Wang, Qiang, et al.
Pubblicazione: (2025)
di: Wang, Qiang, et al.
Pubblicazione: (2025)
IG Captioner: Information Gain Captioners are Strong Zero-shot Classifiers
di: Yang, Chenglin, et al.
Pubblicazione: (2023)
di: Yang, Chenglin, et al.
Pubblicazione: (2023)
Are We Making Progress in Multimodal Domain Generalization? A Comprehensive Benchmark Study
di: Dong, Hao, et al.
Pubblicazione: (2026)
di: Dong, Hao, et al.
Pubblicazione: (2026)
Make VLM Recognize Visual Hallucination on Cartoon Character Image with Pose Information
di: Kim, Bumsoo, et al.
Pubblicazione: (2024)
di: Kim, Bumsoo, et al.
Pubblicazione: (2024)
MUSES: 3D-Controllable Image Generation via Multi-Modal Agent Collaboration
di: Ding, Yanbo, et al.
Pubblicazione: (2024)
di: Ding, Yanbo, et al.
Pubblicazione: (2024)
CVSearch: Empowering Multimodal LLMs with Cognitive Visual Search for High-Resolution Image Perception
di: Li, Liupeng, et al.
Pubblicazione: (2026)
di: Li, Liupeng, et al.
Pubblicazione: (2026)
Human-Centric Foundation Models: Perception, Generation and Agentic Modeling
di: Tang, Shixiang, et al.
Pubblicazione: (2025)
di: Tang, Shixiang, et al.
Pubblicazione: (2025)
MotionCtrl: A Unified and Flexible Motion Controller for Video Generation
di: Wang, Zhouxia, et al.
Pubblicazione: (2023)
di: Wang, Zhouxia, et al.
Pubblicazione: (2023)
HaloQuest: A Visual Hallucination Dataset for Advancing Multimodal Reasoning
di: Wang, Zhecan, et al.
Pubblicazione: (2024)
di: Wang, Zhecan, et al.
Pubblicazione: (2024)
Deciphering Personalization: Towards Fine-Grained Explainability in Natural Language for Personalized Image Generation Models
di: Wang, Haoming, et al.
Pubblicazione: (2025)
di: Wang, Haoming, et al.
Pubblicazione: (2025)
Vision-Language Meets the Skeleton: Progressively Distillation with Cross-Modal Knowledge for 3D Action Representation Learning
di: Chen, Yang, et al.
Pubblicazione: (2024)
di: Chen, Yang, et al.
Pubblicazione: (2024)
HOIN: High-Order Implicit Neural Representations
di: Chen, Yang, et al.
Pubblicazione: (2024)
di: Chen, Yang, et al.
Pubblicazione: (2024)
V-Stylist: Video Stylization via Collaboration and Reflection of MLLM Agents
di: Yue, Zhengrong, et al.
Pubblicazione: (2025)
di: Yue, Zhengrong, et al.
Pubblicazione: (2025)
Towards Multi-Task Multi-Modal Models: A Video Generative Perspective
di: Yu, Lijun
Pubblicazione: (2024)
di: Yu, Lijun
Pubblicazione: (2024)
POINTS: Improving Your Vision-language Model with Affordable Strategies
di: Liu, Yuan, et al.
Pubblicazione: (2024)
di: Liu, Yuan, et al.
Pubblicazione: (2024)
Meta-CoT: Enhancing Granularity and Generalization in Image Editing
di: Zhang, Shiyi, et al.
Pubblicazione: (2026)
di: Zhang, Shiyi, et al.
Pubblicazione: (2026)
Progressive Confident Masking Attention Network for Audio-Visual Segmentation
di: Wang, Yuxuan, et al.
Pubblicazione: (2024)
di: Wang, Yuxuan, et al.
Pubblicazione: (2024)
Enhancing multimodal cooperation via sample-level modality valuation
di: Wei, Yake, et al.
Pubblicazione: (2023)
di: Wei, Yake, et al.
Pubblicazione: (2023)
Boosting Facial Action Unit Detection Through Jointly Learning Facial Landmark Detection and Domain Separation and Reconstruction
di: Shang, Ziqiao, et al.
Pubblicazione: (2023)
di: Shang, Ziqiao, et al.
Pubblicazione: (2023)
Scaling Spatial Intelligence with Multimodal Foundation Models
di: Cai, Zhongang, et al.
Pubblicazione: (2025)
di: Cai, Zhongang, et al.
Pubblicazione: (2025)
RMAdapter: Reconstruction-based Multi-Modal Adapter for Vision-Language Models
di: Lin, Xiang, et al.
Pubblicazione: (2025)
di: Lin, Xiang, et al.
Pubblicazione: (2025)
Data or Language Supervision: What Makes CLIP Better than DINO?
di: Liu, Yiming, et al.
Pubblicazione: (2025)
di: Liu, Yiming, et al.
Pubblicazione: (2025)
MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory Guidance
di: Li, Quanhao, et al.
Pubblicazione: (2025)
di: Li, Quanhao, et al.
Pubblicazione: (2025)
SGAT4PASS: Spherical Geometry-Aware Transformer for PAnoramic Semantic Segmentation
di: Li, Xuewei, et al.
Pubblicazione: (2023)
di: Li, Xuewei, et al.
Pubblicazione: (2023)
STIV: Scalable Text and Image Conditioned Video Generation
di: Lin, Zongyu, et al.
Pubblicazione: (2024)
di: Lin, Zongyu, et al.
Pubblicazione: (2024)
COSMIC: Clique-Oriented Semantic Multi-space Integration for Robust CLIP Test-Time Adaptation
di: Huang, Fanding, et al.
Pubblicazione: (2025)
di: Huang, Fanding, et al.
Pubblicazione: (2025)
FlashMotion: Few-Step Controllable Video Generation with Trajectory Guidance
di: Li, Quanhao, et al.
Pubblicazione: (2026)
di: Li, Quanhao, et al.
Pubblicazione: (2026)
Deep ReLU Networks Have Surprisingly Simple Polytopes
di: Fan, Feng-Lei, et al.
Pubblicazione: (2023)
di: Fan, Feng-Lei, et al.
Pubblicazione: (2023)
PlanLLM: Video Procedure Planning with Refinable Large Language Models
di: Yang, Dejie, et al.
Pubblicazione: (2024)
di: Yang, Dejie, et al.
Pubblicazione: (2024)
Who Brings the Frisbee: Probing Hidden Hallucination Factors in Large Vision-Language Model via Causality Analysis
di: Huang, Po-Hsuan, et al.
Pubblicazione: (2024)
di: Huang, Po-Hsuan, et al.
Pubblicazione: (2024)
X-Prompt: Towards Universal In-Context Image Generation in Auto-Regressive Vision Language Foundation Models
di: Sun, Zeyi, et al.
Pubblicazione: (2024)
di: Sun, Zeyi, et al.
Pubblicazione: (2024)
Scalable Object Relation Encoding for Better 3D Spatial Reasoning in Large Language Models
di: Zhou, Shengli, et al.
Pubblicazione: (2026)
di: Zhou, Shengli, et al.
Pubblicazione: (2026)
Explore the Limits of Omni-modal Pretraining at Scale
di: Zhang, Yiyuan, et al.
Pubblicazione: (2024)
di: Zhang, Yiyuan, et al.
Pubblicazione: (2024)
Reducing Hallucinations in Vision-Language Models via Latent Space Steering
di: Liu, Sheng, et al.
Pubblicazione: (2024)
di: Liu, Sheng, et al.
Pubblicazione: (2024)
OneLLM: One Framework to Align All Modalities with Language
di: Han, Jiaming, et al.
Pubblicazione: (2023)
di: Han, Jiaming, et al.
Pubblicazione: (2023)
Long-tailed Medical Diagnosis with Relation-aware Representation Learning and Iterative Classifier Calibration
di: Pan, Li, et al.
Pubblicazione: (2025)
di: Pan, Li, et al.
Pubblicazione: (2025)
COMODO: Cross-Modal Video-to-IMU Distillation for Efficient Egocentric Human Activity Recognition
di: Chen, Baiyu, et al.
Pubblicazione: (2025)
di: Chen, Baiyu, et al.
Pubblicazione: (2025)
Documenti analoghi
-
LMVD: A Large-Scale Multimodal Vlog Dataset for Depression Detection in the Wild
di: He, Lang, et al.
Pubblicazione: (2024) -
Discover Your Neighbors: Advanced Stable Test-Time Adaptation in Dynamic World
di: Jiang, Qinting, et al.
Pubblicazione: (2024) -
TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning
di: Zeng, Xiangyu, et al.
Pubblicazione: (2024) -
VDC-Agent: When Video Detailed Captioners Evolve Themselves via Agentic Self-Reflection
di: Wang, Qiang, et al.
Pubblicazione: (2025) -
IG Captioner: Information Gain Captioners are Strong Zero-shot Classifiers
di: Yang, Chenglin, et al.
Pubblicazione: (2023)