Openstory++: A Large-scale Dataset and Benchmark for Instance-aware Open-domain Visual Storytelling
Fuente:
arXiv
Salvato in:
| Autori principali: | Ye, Zilyu, Liu, Jinxiu, Peng, Ruotian, Cao, Jinjin, Chen, Zhiyang, Zhang, Yiyang, Xuan, Ziwei, Zhou, Mingyuan, Shen, Xiaoqian, Elhoseiny, Mohamed, Liu, Qi, Qi, Guo-Jun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
StoryGPT-V: Large Language Models as Consistent Story Visualizers
di: Shen, Xiaoqian, et al.
Pubblicazione: (2023)
di: Shen, Xiaoqian, et al.
Pubblicazione: (2023)
Schedule On the Fly: Diffusion Time Prediction for Faster and Better Image Generation
di: Ye, Zilyu, et al.
Pubblicazione: (2024)
di: Ye, Zilyu, et al.
Pubblicazione: (2024)
Affective Visual Dialog: A Large-Scale Benchmark for Emotional Reasoning Based on Visually Grounded Conversations
di: Haydarov, Kilichbek, et al.
Pubblicazione: (2023)
di: Haydarov, Kilichbek, et al.
Pubblicazione: (2023)
VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image Understanding
di: Li, Xiang, et al.
Pubblicazione: (2024)
di: Li, Xiang, et al.
Pubblicazione: (2024)
UltrAvatar: A Realistic Animatable 3D Avatar Diffusion Model with Authenticity Guided Textures
di: Zhou, Mingyuan, et al.
Pubblicazione: (2024)
di: Zhou, Mingyuan, et al.
Pubblicazione: (2024)
Vgent: Graph-based Retrieval-Reasoning-Augmented Generation For Long Video Understanding
di: Shen, Xiaoqian, et al.
Pubblicazione: (2025)
di: Shen, Xiaoqian, et al.
Pubblicazione: (2025)
PRIMEdit: Probability Redistribution for Instance-aware Multi-object Video Editing with Benchmark Dataset
di: Teodoro, Samuel, et al.
Pubblicazione: (2024)
di: Teodoro, Samuel, et al.
Pubblicazione: (2024)
Fine-Grained Open-Vocabulary Object Detection with Fined-Grained Prompts: Task, Dataset and Benchmark
di: Liu, Ying, et al.
Pubblicazione: (2025)
di: Liu, Ying, et al.
Pubblicazione: (2025)
InstructRL4Pix: Training Diffusion for Image Editing by Reinforcement Learning
di: Li, Tiancheng, et al.
Pubblicazione: (2024)
di: Li, Tiancheng, et al.
Pubblicazione: (2024)
Multi-scale 2D Temporal Map Diffusion Models for Natural Language Video Localization
di: Zhang, Chongzhi, et al.
Pubblicazione: (2024)
di: Zhang, Chongzhi, et al.
Pubblicazione: (2024)
MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens
di: Ataallah, Kirolos, et al.
Pubblicazione: (2024)
di: Ataallah, Kirolos, et al.
Pubblicazione: (2024)
Improving Time Series Forecasting via Instance-aware Post-hoc Revision
di: Liu, Zhiding, et al.
Pubblicazione: (2025)
di: Liu, Zhiding, et al.
Pubblicazione: (2025)
Context-aware Visual Storytelling with Visual Prefix Tuning and Contrastive Learning
di: Song, Yingjin, et al.
Pubblicazione: (2024)
di: Song, Yingjin, et al.
Pubblicazione: (2024)
Visual Instance-aware Prompt Tuning
di: Xiao, Xi, et al.
Pubblicazione: (2025)
di: Xiao, Xi, et al.
Pubblicazione: (2025)
Intelligent Grimm -- Open-ended Visual Storytelling via Latent Diffusion Models
di: Liu, Chang, et al.
Pubblicazione: (2023)
di: Liu, Chang, et al.
Pubblicazione: (2023)
Improving Generalized Visual Grounding with Instance-aware Joint Learning
di: Dai, Ming, et al.
Pubblicazione: (2025)
di: Dai, Ming, et al.
Pubblicazione: (2025)
Hard-aware Instance Adaptive Self-training for Unsupervised Cross-domain Semantic Segmentation
di: Zhu, Chuang, et al.
Pubblicazione: (2023)
di: Zhu, Chuang, et al.
Pubblicazione: (2023)
MULTISCRIPT: Multimodal Script Learning for Supporting Open Domain Everyday Tasks
di: Qi, Jingyuan, et al.
Pubblicazione: (2023)
di: Qi, Jingyuan, et al.
Pubblicazione: (2023)
iMotion-LLM: Instruction-Conditioned Trajectory Generation
di: Felemban, Abdulwahab, et al.
Pubblicazione: (2024)
di: Felemban, Abdulwahab, et al.
Pubblicazione: (2024)
Convergence analysis of a balancing domain decomposition method for an elliptic optimal control problem with HDG discretizations
di: Liu, Sijing, et al.
Pubblicazione: (2025)
di: Liu, Sijing, et al.
Pubblicazione: (2025)
Multilingual KokoroChat: A Multi-LLM Ensemble Translation Method for Creating a Multilingual Counseling Dialogue Dataset
di: Suzuki, Ryoma, et al.
Pubblicazione: (2026)
di: Suzuki, Ryoma, et al.
Pubblicazione: (2026)
OpenTracer: A Dynamic Transaction Trace Analyzer for Smart Contract Invariant Generation and Beyond
di: Chen, Zhiyang, et al.
Pubblicazione: (2024)
di: Chen, Zhiyang, et al.
Pubblicazione: (2024)
CoT3DRef: Chain-of-Thoughts Data-Efficient 3D Visual Grounding
di: Abdelrahman, Eslam, et al.
Pubblicazione: (2023)
di: Abdelrahman, Eslam, et al.
Pubblicazione: (2023)
Privacy Risks in the Storytelling of Open Government Data: A Study from the Perspective of User Cognitive Reasoning
di: Ruili Geng, et al.
Pubblicazione: (2024)
di: Ruili Geng, et al.
Pubblicazione: (2024)
A balancing domain decomposition by constraints preconditioner for a hybridizable discontinuous Galerkin discretization of an elliptic optimal control problem
di: Liu, Sijing, et al.
Pubblicazione: (2025)
di: Liu, Sijing, et al.
Pubblicazione: (2025)
OpenInsGaussian: Open-vocabulary Instance Gaussian Segmentation with Context-aware Cross-view Fusion
di: Huang, Tianyu, et al.
Pubblicazione: (2025)
di: Huang, Tianyu, et al.
Pubblicazione: (2025)
InfLVG: Reinforce Inference-Time Consistent Long Video Generation with GRPO
di: Fang, Xueji, et al.
Pubblicazione: (2025)
di: Fang, Xueji, et al.
Pubblicazione: (2025)
Zoom-Zero: Reinforced Coarse-to-Fine Video Understanding via Temporal Zoom-in
di: Shen, Xiaoqian, et al.
Pubblicazione: (2025)
di: Shen, Xiaoqian, et al.
Pubblicazione: (2025)
UniHGKR: Unified Instruction-aware Heterogeneous Knowledge Retrievers
di: Min, Dehai, et al.
Pubblicazione: (2024)
di: Min, Dehai, et al.
Pubblicazione: (2024)
When Images Speak Louder: Mitigating Language Bias-induced Hallucinations in VLMs through Cross-Modal Guidance
di: Cao, Jinjin, et al.
Pubblicazione: (2025)
di: Cao, Jinjin, et al.
Pubblicazione: (2025)
Class Agnostic Instance-level Descriptor for Visual Instance Search
di: Sun, Qi-Ying, et al.
Pubblicazione: (2025)
di: Sun, Qi-Ying, et al.
Pubblicazione: (2025)
VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs?
di: Gado, Mohamed, et al.
Pubblicazione: (2025)
di: Gado, Mohamed, et al.
Pubblicazione: (2025)
3DCoMPaT$^{++}$: An improved Large-scale 3D Vision Dataset for Compositional Recognition
di: Slim, Habib, et al.
Pubblicazione: (2023)
di: Slim, Habib, et al.
Pubblicazione: (2023)
Aligning by Misaligning: Boundary-aware Curriculum Learning for Multimodal Alignment
di: Ye, Hua, et al.
Pubblicazione: (2025)
di: Ye, Hua, et al.
Pubblicazione: (2025)
M-MiniGPT4: Multilingual VLLM Alignment via Translated Data
di: Han, Seung Hun, et al.
Pubblicazione: (2026)
di: Han, Seung Hun, et al.
Pubblicazione: (2026)
From Semantics, Scene to Instance-awareness: Distilling Foundation Model for Grounded Open-vocabulary Situation Recognition
di: Cai, Chen, et al.
Pubblicazione: (2025)
di: Cai, Chen, et al.
Pubblicazione: (2025)
MaaSO: SLO-aware Orchestration of Heterogeneous Model Instances for MaaS
di: Xuan, Mo, et al.
Pubblicazione: (2025)
di: Xuan, Mo, et al.
Pubblicazione: (2025)
Data Augmentation Integrating Dialogue Flow and Style to Adapt Spoken Dialogue Systems to Low-Resource User Groups
di: Qi, Zhiyang, et al.
Pubblicazione: (2024)
di: Qi, Zhiyang, et al.
Pubblicazione: (2024)
Enhancing Dialogue Generation in Werewolf Game Through Situation Analysis and Persuasion Strategies
di: Qi, Zhiyang, et al.
Pubblicazione: (2024)
di: Qi, Zhiyang, et al.
Pubblicazione: (2024)
The asymptoticity of extremal length in Teichmüller space
di: Lyu, Zhiyang, et al.
Pubblicazione: (2025)
di: Lyu, Zhiyang, et al.
Pubblicazione: (2025)
Documenti analoghi
-
StoryGPT-V: Large Language Models as Consistent Story Visualizers
di: Shen, Xiaoqian, et al.
Pubblicazione: (2023) -
Schedule On the Fly: Diffusion Time Prediction for Faster and Better Image Generation
di: Ye, Zilyu, et al.
Pubblicazione: (2024) -
Affective Visual Dialog: A Large-Scale Benchmark for Emotional Reasoning Based on Visually Grounded Conversations
di: Haydarov, Kilichbek, et al.
Pubblicazione: (2023) -
VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image Understanding
di: Li, Xiang, et al.
Pubblicazione: (2024) -
UltrAvatar: A Realistic Animatable 3D Avatar Diffusion Model with Authenticity Guided Textures
di: Zhou, Mingyuan, et al.
Pubblicazione: (2024)