Intelligent Director: An Automatic Framework for Dynamic Visual Composition using ChatGPT
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zheng, Sixiao, Huo, Jingyang, Wang, Yu, Fu, Yanwei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ContextualStory: Consistent Visual Storytelling with Spatially-Enhanced and Storyline Context
von: Zheng, Sixiao, et al.
Veröffentlicht: (2024)
von: Zheng, Sixiao, et al.
Veröffentlicht: (2024)
VidCRAFT3: Camera, Object, and Lighting Control for Image-to-Video Generation
von: Zheng, Sixiao, et al.
Veröffentlicht: (2025)
von: Zheng, Sixiao, et al.
Veröffentlicht: (2025)
How Good is ChatGPT at Audiovisual Deepfake Detection: A Comparative Study of ChatGPT, AI Models and Human Perception
von: Shahzad, Sahibzada Adil, et al.
Veröffentlicht: (2024)
von: Shahzad, Sahibzada Adil, et al.
Veröffentlicht: (2024)
Audio-Guided Visual Perception for Audio-Visual Navigation
von: Wang, Yi, et al.
Veröffentlicht: (2025)
von: Wang, Yi, et al.
Veröffentlicht: (2025)
XEmoGPT: An Explainable Multimodal Emotion Recognition Framework with Cue-Level Perception and Reasoning
von: Zhang, Hanwen, et al.
Veröffentlicht: (2026)
von: Zhang, Hanwen, et al.
Veröffentlicht: (2026)
Taming Modality Entanglement in Continual Audio-Visual Segmentation
von: Hong, Yuyang, et al.
Veröffentlicht: (2025)
von: Hong, Yuyang, et al.
Veröffentlicht: (2025)
ScreenWriter: Automatic Screenplay Generation and Movie Summarisation
von: Mahon, Louis, et al.
Veröffentlicht: (2024)
von: Mahon, Louis, et al.
Veröffentlicht: (2024)
PROVE: A Perceptual RemOVal cohErence Benchmark for Visual Media
von: Li, Fuhao, et al.
Veröffentlicht: (2026)
von: Li, Fuhao, et al.
Veröffentlicht: (2026)
EARTalking: End-to-end GPT-style Autoregressive Talking Head Synthesis with Frame-wise Control
von: Weng, Yuzhe, et al.
Veröffentlicht: (2026)
von: Weng, Yuzhe, et al.
Veröffentlicht: (2026)
SoccerHigh: A Benchmark Dataset for Automatic Soccer Video Summarization
von: Díaz-Juan, Artur, et al.
Veröffentlicht: (2025)
von: Díaz-Juan, Artur, et al.
Veröffentlicht: (2025)
Automatic Recognition of Food Ingestion Environment from the AIM-2 Wearable Sensor
von: Huang, Yuning, et al.
Veröffentlicht: (2024)
von: Huang, Yuning, et al.
Veröffentlicht: (2024)
VDE Bench: Evaluating The Capability of Image Editing Models to Modify Visual Documents
von: Yi, Hongzhu, et al.
Veröffentlicht: (2026)
von: Yi, Hongzhu, et al.
Veröffentlicht: (2026)
EgoBlind: Towards Egocentric Visual Assistance for the Blind
von: Xiao, Junbin, et al.
Veröffentlicht: (2025)
von: Xiao, Junbin, et al.
Veröffentlicht: (2025)
Label-anticipated Event Disentanglement for Audio-Visual Video Parsing
von: Zhou, Jinxing, et al.
Veröffentlicht: (2024)
von: Zhou, Jinxing, et al.
Veröffentlicht: (2024)
Enhancing Visible-Infrared Person Re-identification with Modality- and Instance-aware Visual Prompt Learning
von: Wu, Ruiqi, et al.
Veröffentlicht: (2024)
von: Wu, Ruiqi, et al.
Veröffentlicht: (2024)
Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation
von: Yu, Lijun, et al.
Veröffentlicht: (2023)
von: Yu, Lijun, et al.
Veröffentlicht: (2023)
A User-Friendly Framework for Generating Model-Preferred Prompts in Text-to-Image Synthesis
von: Hei, Nailei, et al.
Veröffentlicht: (2024)
von: Hei, Nailei, et al.
Veröffentlicht: (2024)
Decoupled Audio-Visual Dataset Distillation
von: Li, Wenyuan, et al.
Veröffentlicht: (2025)
von: Li, Wenyuan, et al.
Veröffentlicht: (2025)
SEPS: Semantic-enhanced Patch Slimming Framework for fine-grained cross-modal alignment
von: Mao, Xinyu, et al.
Veröffentlicht: (2025)
von: Mao, Xinyu, et al.
Veröffentlicht: (2025)
Causal-Story: Local Causal Attention Utilizing Parameter-Efficient Tuning For Visual Story Synthesis
von: Song, Tianyi, et al.
Veröffentlicht: (2023)
von: Song, Tianyi, et al.
Veröffentlicht: (2023)
Audio Visual Segmentation Through Text Embeddings
von: Lee, Kyungbok, et al.
Veröffentlicht: (2025)
von: Lee, Kyungbok, et al.
Veröffentlicht: (2025)
Attributes-aware Visual Emotion Representation Learning
von: Maharjan, Rahul Singh, et al.
Veröffentlicht: (2025)
von: Maharjan, Rahul Singh, et al.
Veröffentlicht: (2025)
LLM-EvRep: Learning an LLM-Compatible Event Representation Using a Self-Supervised Framework
von: Yu, Zongyou, et al.
Veröffentlicht: (2025)
von: Yu, Zongyou, et al.
Veröffentlicht: (2025)
Q-Ponder: A Unified Training Pipeline for Reasoning-based Visual Quality Assessment
von: Cai, Zhuoxuan, et al.
Veröffentlicht: (2025)
von: Cai, Zhuoxuan, et al.
Veröffentlicht: (2025)
Hierarchical Knowledge Graphs for Story Understanding in Visual Narratives
von: Chen, Yi-Chun
Veröffentlicht: (2025)
von: Chen, Yi-Chun
Veröffentlicht: (2025)
DreamStory: Open-Domain Story Visualization by LLM-Guided Multi-Subject Consistent Diffusion
von: He, Huiguo, et al.
Veröffentlicht: (2024)
von: He, Huiguo, et al.
Veröffentlicht: (2024)
GiVE: Guiding Visual Encoder to Perceive Overlooked Information
von: Li, Junjie, et al.
Veröffentlicht: (2024)
von: Li, Junjie, et al.
Veröffentlicht: (2024)
CrypticBio: A Large Multimodal Dataset for Visually Confusing Biodiversity
von: Manolache, Georgiana, et al.
Veröffentlicht: (2025)
von: Manolache, Georgiana, et al.
Veröffentlicht: (2025)
Learning Trimodal Relation for Audio-Visual Question Answering with Missing Modality
von: Park, Kyu Ri, et al.
Veröffentlicht: (2024)
von: Park, Kyu Ri, et al.
Veröffentlicht: (2024)
Learning to Mask and Permute Visual Tokens for Vision Transformer Pre-Training
von: Baraldi, Lorenzo, et al.
Veröffentlicht: (2023)
von: Baraldi, Lorenzo, et al.
Veröffentlicht: (2023)
Controllable Audio-Visual Viewpoint Generation from 360° Spatial Information
von: Marinoni, Christian, et al.
Veröffentlicht: (2025)
von: Marinoni, Christian, et al.
Veröffentlicht: (2025)
AdaCoder: Adaptive Prompt Compression for Programmatic Visual Question Answering
von: Ukai, Mahiro, et al.
Veröffentlicht: (2024)
von: Ukai, Mahiro, et al.
Veröffentlicht: (2024)
JPEG AI Image Compression Visual Artifacts: Detection Methods and Dataset
von: Tsereh, Daria, et al.
Veröffentlicht: (2024)
von: Tsereh, Daria, et al.
Veröffentlicht: (2024)
VIA: Unified Spatiotemporal Video Adaptation Framework for Global and Local Video Editing
von: Gu, Jing, et al.
Veröffentlicht: (2024)
von: Gu, Jing, et al.
Veröffentlicht: (2024)
Exploring Category-level Articulated Object Pose Tracking on SE(3) Manifolds
von: Meng, Xianhui, et al.
Veröffentlicht: (2025)
von: Meng, Xianhui, et al.
Veröffentlicht: (2025)
Boosting Audio Visual Question Answering via Key Semantic-Aware Cues
von: Li, Guangyao, et al.
Veröffentlicht: (2024)
von: Li, Guangyao, et al.
Veröffentlicht: (2024)
Can I Trust Your Answer? Visually Grounded Video Question Answering
von: Xiao, Junbin, et al.
Veröffentlicht: (2023)
von: Xiao, Junbin, et al.
Veröffentlicht: (2023)
Crab$^{+}$: A Scalable and Unified Audio-Visual Scene Understanding Model with Explicit Cooperation
von: Cai, Dongnuan, et al.
Veröffentlicht: (2026)
von: Cai, Dongnuan, et al.
Veröffentlicht: (2026)
Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation
von: Liu, Che, et al.
Veröffentlicht: (2026)
von: Liu, Che, et al.
Veröffentlicht: (2026)
Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech
von: Liu, Rui, et al.
Veröffentlicht: (2024)
von: Liu, Rui, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
ContextualStory: Consistent Visual Storytelling with Spatially-Enhanced and Storyline Context
von: Zheng, Sixiao, et al.
Veröffentlicht: (2024) -
VidCRAFT3: Camera, Object, and Lighting Control for Image-to-Video Generation
von: Zheng, Sixiao, et al.
Veröffentlicht: (2025) -
How Good is ChatGPT at Audiovisual Deepfake Detection: A Comparative Study of ChatGPT, AI Models and Human Perception
von: Shahzad, Sahibzada Adil, et al.
Veröffentlicht: (2024) -
Audio-Guided Visual Perception for Audio-Visual Navigation
von: Wang, Yi, et al.
Veröffentlicht: (2025) -
XEmoGPT: An Explainable Multimodal Emotion Recognition Framework with Cue-Level Perception and Reasoning
von: Zhang, Hanwen, et al.
Veröffentlicht: (2026)