ContextualStory: Consistent Visual Storytelling with Spatially-Enhanced and Storyline Context
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zheng, Sixiao, Fu, Yanwei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Intelligent Director: An Automatic Framework for Dynamic Visual Composition using ChatGPT
von: Zheng, Sixiao, et al.
Veröffentlicht: (2024)
von: Zheng, Sixiao, et al.
Veröffentlicht: (2024)
DreamStory: Open-Domain Story Visualization by LLM-Guided Multi-Subject Consistent Diffusion
von: He, Huiguo, et al.
Veröffentlicht: (2024)
von: He, Huiguo, et al.
Veröffentlicht: (2024)
VidCRAFT3: Camera, Object, and Lighting Control for Image-to-Video Generation
von: Zheng, Sixiao, et al.
Veröffentlicht: (2025)
von: Zheng, Sixiao, et al.
Veröffentlicht: (2025)
Causal-Story: Local Causal Attention Utilizing Parameter-Efficient Tuning For Visual Story Synthesis
von: Song, Tianyi, et al.
Veröffentlicht: (2023)
von: Song, Tianyi, et al.
Veröffentlicht: (2023)
Hierarchical Knowledge Graphs for Story Understanding in Visual Narratives
von: Chen, Yi-Chun
Veröffentlicht: (2025)
von: Chen, Yi-Chun
Veröffentlicht: (2025)
Controllable Audio-Visual Viewpoint Generation from 360° Spatial Information
von: Marinoni, Christian, et al.
Veröffentlicht: (2025)
von: Marinoni, Christian, et al.
Veröffentlicht: (2025)
Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech
von: Liu, Rui, et al.
Veröffentlicht: (2024)
von: Liu, Rui, et al.
Veröffentlicht: (2024)
AesopAgent: Agent-driven Evolutionary System on Story-to-Video Production
von: Wang, Jiuniu, et al.
Veröffentlicht: (2024)
von: Wang, Jiuniu, et al.
Veröffentlicht: (2024)
Detecting Misinformation in Multimedia Content through Cross-Modal Entity Consistency: A Dual Learning Approach
von: Fu, Zhe, et al.
Veröffentlicht: (2024)
von: Fu, Zhe, et al.
Veröffentlicht: (2024)
Detecting Multimodal Situations with Insufficient Context and Abstaining from Baseless Predictions
von: Liu, Junzhang, et al.
Veröffentlicht: (2024)
von: Liu, Junzhang, et al.
Veröffentlicht: (2024)
Diversify, Contextualize, and Adapt: Efficient Entropy Modeling for Neural Image Codec
von: Kim, Jun-Hyuk, et al.
Veröffentlicht: (2024)
von: Kim, Jun-Hyuk, et al.
Veröffentlicht: (2024)
Taming Modality Entanglement in Continual Audio-Visual Segmentation
von: Hong, Yuyang, et al.
Veröffentlicht: (2025)
von: Hong, Yuyang, et al.
Veröffentlicht: (2025)
ContextDet: Temporal Action Detection with Adaptive Context Aggregation
von: Wang, Ning, et al.
Veröffentlicht: (2024)
von: Wang, Ning, et al.
Veröffentlicht: (2024)
Enhancing Visible-Infrared Person Re-identification with Modality- and Instance-aware Visual Prompt Learning
von: Wu, Ruiqi, et al.
Veröffentlicht: (2024)
von: Wu, Ruiqi, et al.
Veröffentlicht: (2024)
Enhanced Multimodal Hate Video Detection via Channel-wise and Modality-wise Fusion
von: Zhang, Yinghui, et al.
Veröffentlicht: (2025)
von: Zhang, Yinghui, et al.
Veröffentlicht: (2025)
Audio-Guided Visual Perception for Audio-Visual Navigation
von: Wang, Yi, et al.
Veröffentlicht: (2025)
von: Wang, Yi, et al.
Veröffentlicht: (2025)
Understanding Temporal Logic Consistency in Video-Language Models through Cross-Modal Attention Discriminability
von: Li, Chengzhi, et al.
Veröffentlicht: (2025)
von: Li, Chengzhi, et al.
Veröffentlicht: (2025)
SFANet: Spatial-Frequency Attention Network for Deepfake Detection
von: Ahire, Vrushank, et al.
Veröffentlicht: (2025)
von: Ahire, Vrushank, et al.
Veröffentlicht: (2025)
VidCtx: Context-aware Video Question Answering with Image Models
von: Goulas, Andreas, et al.
Veröffentlicht: (2024)
von: Goulas, Andreas, et al.
Veröffentlicht: (2024)
Decoupled Audio-Visual Dataset Distillation
von: Li, Wenyuan, et al.
Veröffentlicht: (2025)
von: Li, Wenyuan, et al.
Veröffentlicht: (2025)
From Satellite to Street: A Hybrid Framework Integrating Stable Diffusion and PanoGAN for Consistent Cross-View Synthesis
von: Bajbaa, Khawlah, et al.
Veröffentlicht: (2025)
von: Bajbaa, Khawlah, et al.
Veröffentlicht: (2025)
Autoregressive Image Generation with Linear Complexity: A Spatial-Aware Decay Perspective
von: Mao, Yuxin, et al.
Veröffentlicht: (2025)
von: Mao, Yuxin, et al.
Veröffentlicht: (2025)
Audio Visual Segmentation Through Text Embeddings
von: Lee, Kyungbok, et al.
Veröffentlicht: (2025)
von: Lee, Kyungbok, et al.
Veröffentlicht: (2025)
Attributes-aware Visual Emotion Representation Learning
von: Maharjan, Rahul Singh, et al.
Veröffentlicht: (2025)
von: Maharjan, Rahul Singh, et al.
Veröffentlicht: (2025)
LLaVA-MORE: A Comparative Study of LLMs and Visual Backbones for Enhanced Visual Instruction Tuning
von: Cocchi, Federico, et al.
Veröffentlicht: (2025)
von: Cocchi, Federico, et al.
Veröffentlicht: (2025)
Storybooth: Training-free Multi-Subject Consistency for Improved Visual Storytelling
von: Singh, Jaskirat, et al.
Veröffentlicht: (2025)
von: Singh, Jaskirat, et al.
Veröffentlicht: (2025)
VideoPainter: Any-length Video Inpainting and Editing with Plug-and-Play Context Control
von: Bian, Yuxuan, et al.
Veröffentlicht: (2025)
von: Bian, Yuxuan, et al.
Veröffentlicht: (2025)
EgoBlind: Towards Egocentric Visual Assistance for the Blind
von: Xiao, Junbin, et al.
Veröffentlicht: (2025)
von: Xiao, Junbin, et al.
Veröffentlicht: (2025)
Label-anticipated Event Disentanglement for Audio-Visual Video Parsing
von: Zhou, Jinxing, et al.
Veröffentlicht: (2024)
von: Zhou, Jinxing, et al.
Veröffentlicht: (2024)
GiVE: Guiding Visual Encoder to Perceive Overlooked Information
von: Li, Junjie, et al.
Veröffentlicht: (2024)
von: Li, Junjie, et al.
Veröffentlicht: (2024)
Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation
von: Yu, Lijun, et al.
Veröffentlicht: (2023)
von: Yu, Lijun, et al.
Veröffentlicht: (2023)
Learning Trimodal Relation for Audio-Visual Question Answering with Missing Modality
von: Park, Kyu Ri, et al.
Veröffentlicht: (2024)
von: Park, Kyu Ri, et al.
Veröffentlicht: (2024)
AdaCoder: Adaptive Prompt Compression for Programmatic Visual Question Answering
von: Ukai, Mahiro, et al.
Veröffentlicht: (2024)
von: Ukai, Mahiro, et al.
Veröffentlicht: (2024)
JPEG AI Image Compression Visual Artifacts: Detection Methods and Dataset
von: Tsereh, Daria, et al.
Veröffentlicht: (2024)
von: Tsereh, Daria, et al.
Veröffentlicht: (2024)
CrypticBio: A Large Multimodal Dataset for Visually Confusing Biodiversity
von: Manolache, Georgiana, et al.
Veröffentlicht: (2025)
von: Manolache, Georgiana, et al.
Veröffentlicht: (2025)
Learning to Mask and Permute Visual Tokens for Vision Transformer Pre-Training
von: Baraldi, Lorenzo, et al.
Veröffentlicht: (2023)
von: Baraldi, Lorenzo, et al.
Veröffentlicht: (2023)
MLLM-VADStory: Domain Knowledge-Driven Multimodal LLMs for Video Ad Storyline Insights
von: Yang, Jasmine, et al.
Veröffentlicht: (2026)
von: Yang, Jasmine, et al.
Veröffentlicht: (2026)
Boosting Audio Visual Question Answering via Key Semantic-Aware Cues
von: Li, Guangyao, et al.
Veröffentlicht: (2024)
von: Li, Guangyao, et al.
Veröffentlicht: (2024)
Can I Trust Your Answer? Visually Grounded Video Question Answering
von: Xiao, Junbin, et al.
Veröffentlicht: (2023)
von: Xiao, Junbin, et al.
Veröffentlicht: (2023)
PROVE: A Perceptual RemOVal cohErence Benchmark for Visual Media
von: Li, Fuhao, et al.
Veröffentlicht: (2026)
von: Li, Fuhao, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Intelligent Director: An Automatic Framework for Dynamic Visual Composition using ChatGPT
von: Zheng, Sixiao, et al.
Veröffentlicht: (2024) -
DreamStory: Open-Domain Story Visualization by LLM-Guided Multi-Subject Consistent Diffusion
von: He, Huiguo, et al.
Veröffentlicht: (2024) -
VidCRAFT3: Camera, Object, and Lighting Control for Image-to-Video Generation
von: Zheng, Sixiao, et al.
Veröffentlicht: (2025) -
Causal-Story: Local Causal Attention Utilizing Parameter-Efficient Tuning For Visual Story Synthesis
von: Song, Tianyi, et al.
Veröffentlicht: (2023) -
Hierarchical Knowledge Graphs for Story Understanding in Visual Narratives
von: Chen, Yi-Chun
Veröffentlicht: (2025)