LLMs Behind the Scenes: Enabling Narrative Scene Illustration
Fuente:
arXiv
Saved in:
| Main Authors: | Roemmele, Melissa, Chung, John Joon Young, Kim, Taewook, Sun, Yuqian, Calderwood, Alex, Kreminski, Max |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Phraselette: A Poet's Procedural Palette
by: Calderwood, Alex, et al.
Published: (2025)
by: Calderwood, Alex, et al.
Published: (2025)
Toyteller: AI-powered Visual Storytelling Through Toy-Playing with Character Symbols
by: Chung, John Joon Young, et al.
Published: (2025)
by: Chung, John Joon Young, et al.
Published: (2025)
Drama Llama: An LLM-Powered Storylets Framework for Authorable Responsiveness in Interactive Narrative
by: Sun, Yuqian, et al.
Published: (2025)
by: Sun, Yuqian, et al.
Published: (2025)
Modifying Large Language Model Post-Training for Diverse Creative Writing
by: Chung, John Joon Young, et al.
Published: (2025)
by: Chung, John Joon Young, et al.
Published: (2025)
Design Techniques for LLM-Powered Interactive Storytelling: A Case Study of the Dramamancer System
by: Wang, Tiffany, et al.
Published: (2026)
by: Wang, Tiffany, et al.
Published: (2026)
Scaffolding Recursive Divergence and Convergence in Story Ideation
by: Kim, Taewook, et al.
Published: (2025)
by: Kim, Taewook, et al.
Published: (2025)
Elsewise: Authoring AI-Based Interactive Narrative with Possibility Space Visualization
by: Wang, Yi, et al.
Published: (2025)
by: Wang, Yi, et al.
Published: (2025)
SceneCritic: A Symbolic Evaluator for 3D Indoor Scene Synthesis
by: Sengupta, Kathakoli, et al.
Published: (2026)
by: Sengupta, Kathakoli, et al.
Published: (2026)
Int3DNet: Scene-Motion Cross Attention Network for 3D Intention Prediction in Mixed Reality
by: Ha, Taewook, et al.
Published: (2026)
by: Ha, Taewook, et al.
Published: (2026)
Lumos : Empowering Multimodal LLMs with Scene Text Recognition
by: Shenoy, Ashish, et al.
Published: (2024)
by: Shenoy, Ashish, et al.
Published: (2024)
Patchview: LLM-Powered Worldbuilding with Generative Dust and Magnet Visualization
by: Chung, John Joon Young, et al.
Published: (2024)
by: Chung, John Joon Young, et al.
Published: (2024)
SceneAlign: Aligning Multimodal Reasoning to Scene Graphs in Complex Visual Scenes
by: Wang, Chuhan, et al.
Published: (2026)
by: Wang, Chuhan, et al.
Published: (2026)
DenseAnnotate: Enabling Scalable Dense Caption Collection for Images and 3D Scenes via Spoken Descriptions
by: Lin, Xiaoyu, et al.
Published: (2025)
by: Lin, Xiaoyu, et al.
Published: (2025)
MSG-Chart: Multimodal Scene Graph for ChartQA
by: Dai, Yue, et al.
Published: (2024)
by: Dai, Yue, et al.
Published: (2024)
DivScene: Towards Open-Vocabulary Object Navigation with Large Vision Language Models in Diverse Scenes
by: Wang, Zhaowei, et al.
Published: (2024)
by: Wang, Zhaowei, et al.
Published: (2024)
SceneLinker: Compositional 3D Scene Generation via Semantic Scene Graph from RGB Sequences
by: Kim, Seok-Young, et al.
Published: (2026)
by: Kim, Seok-Young, et al.
Published: (2026)
Learning 3D Scene Analogies with Neural Contextual Scene Maps
by: Kim, Junho, et al.
Published: (2025)
by: Kim, Junho, et al.
Published: (2025)
LED: A Benchmark for Evaluating Layout Error Detection in Document Analysis
by: Heo, Inbum, et al.
Published: (2026)
by: Heo, Inbum, et al.
Published: (2026)
VideoStudio: Generating Consistent-Content and Multi-Scene Videos
by: Long, Fuchen, et al.
Published: (2024)
by: Long, Fuchen, et al.
Published: (2024)
Interleaved Scene Graphs for Interleaved Text-and-Image Generation Assessment
by: Chen, Dongping, et al.
Published: (2024)
by: Chen, Dongping, et al.
Published: (2024)
SceneMI: Motion In-betweening for Modeling Human-Scene Interactions
by: Hwang, Inwoo, et al.
Published: (2025)
by: Hwang, Inwoo, et al.
Published: (2025)
BiMa: Towards Biases Mitigation for Text-Video Retrieval via Scene Element Guidance
by: Le, Huy, et al.
Published: (2025)
by: Le, Huy, et al.
Published: (2025)
Generative Visual Commonsense Answering and Explaining with Generative Scene Graph Constructing
by: Yuan, Fan, et al.
Published: (2025)
by: Yuan, Fan, et al.
Published: (2025)
Open World Scene Graph Generation using Vision Language Models
by: Dutta, Amartya, et al.
Published: (2025)
by: Dutta, Amartya, et al.
Published: (2025)
Behind the Veil: Enhanced Indoor 3D Scene Reconstruction with Occluded Surfaces Completion
by: Sun, Su, et al.
Published: (2024)
by: Sun, Su, et al.
Published: (2024)
PoSh: Using Scene Graphs To Guide LLMs-as-a-Judge For Detailed Image Descriptions
by: Ananthram, Amith, et al.
Published: (2025)
by: Ananthram, Amith, et al.
Published: (2025)
Finding 3D Scene Analogies with Multimodal Foundation Models
by: Kim, Junho, et al.
Published: (2025)
by: Kim, Junho, et al.
Published: (2025)
Behind Maya: Building a Multilingual Vision Language Model
by: Alam, Nahid, et al.
Published: (2025)
by: Alam, Nahid, et al.
Published: (2025)
Bharat Scene Text: A Novel Comprehensive Dataset and Benchmark for Indian Language Scene Text Understanding
by: De, Anik, et al.
Published: (2025)
by: De, Anik, et al.
Published: (2025)
Global-Local Tree Search in VLMs for 3D Indoor Scene Generation
by: Deng, Wei, et al.
Published: (2025)
by: Deng, Wei, et al.
Published: (2025)
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding
by: Wang, Yan, et al.
Published: (2025)
by: Wang, Yan, et al.
Published: (2025)
Zero-Shot Scene Understanding with Multimodal Large Language Models for Automated Vehicles
by: Elhenawy, Mohammed, et al.
Published: (2025)
by: Elhenawy, Mohammed, et al.
Published: (2025)
LiteraryTaste: A Preference Dataset for Creative Writing Personalization
by: Chung, John Joon Young, et al.
Published: (2025)
by: Chung, John Joon Young, et al.
Published: (2025)
Geometry-Aware Scene Configurations for Novel View Synthesis
by: Kim, Minkwan, et al.
Published: (2025)
by: Kim, Minkwan, et al.
Published: (2025)
Scene-Graph ViT: End-to-End Open-Vocabulary Visual Relationship Detection
by: Salzmann, Tim, et al.
Published: (2024)
by: Salzmann, Tim, et al.
Published: (2024)
Towards Holistic Surgical Scene Graph
by: Shin, Jongmin, et al.
Published: (2025)
by: Shin, Jongmin, et al.
Published: (2025)
Seeing Speech and Sound: Distinguishing and Locating Audios in Visual Scenes
by: Ryu, Hyeonggon, et al.
Published: (2025)
by: Ryu, Hyeonggon, et al.
Published: (2025)
PaperBanana: Automating Academic Illustration for AI Scientists
by: Zhu, Dawei, et al.
Published: (2026)
by: Zhu, Dawei, et al.
Published: (2026)
T$^3$-S2S: Training-free Triplet Tuning for Sketch to Scene Synthesis in Controllable Concept Art Generation
by: Sun, Zhenhong, et al.
Published: (2024)
by: Sun, Zhenhong, et al.
Published: (2024)
General Scene Adaptation for Vision-and-Language Navigation
by: Hong, Haodong, et al.
Published: (2025)
by: Hong, Haodong, et al.
Published: (2025)
Similar Items
-
Phraselette: A Poet's Procedural Palette
by: Calderwood, Alex, et al.
Published: (2025) -
Toyteller: AI-powered Visual Storytelling Through Toy-Playing with Character Symbols
by: Chung, John Joon Young, et al.
Published: (2025) -
Drama Llama: An LLM-Powered Storylets Framework for Authorable Responsiveness in Interactive Narrative
by: Sun, Yuqian, et al.
Published: (2025) -
Modifying Large Language Model Post-Training for Diverse Creative Writing
by: Chung, John Joon Young, et al.
Published: (2025) -
Design Techniques for LLM-Powered Interactive Storytelling: A Case Study of the Dramamancer System
by: Wang, Tiffany, et al.
Published: (2026)