TaleDiffusion: Multi-Character Story Generation with Dialogue Rendering
Fuente:
arXiv
Saved in:
| Main Authors: | Banerjee, Ayan, Llados, Josep, Pal, Umapada, Dutta, Anjan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CraftSVG: Multi-Object Text-to-SVG Synthesis via Layout Guided Diffusion
by: Banerjee, Ayan, et al.
Published: (2024)
by: Banerjee, Ayan, et al.
Published: (2024)
DocRevive: A Unified Pipeline for Document Text Restoration
by: Purkayastha, Kunal, et al.
Published: (2026)
by: Purkayastha, Kunal, et al.
Published: (2026)
GraphKD: Exploring Knowledge Distillation Towards Document Object Detection with Structured Graph Creation
by: Banerjee, Ayan, et al.
Published: (2024)
by: Banerjee, Ayan, et al.
Published: (2024)
CountLoop: Training-Free High-Instance Image Generation via Iterative Agent Guidance
by: Mondal, Anindya, et al.
Published: (2025)
by: Mondal, Anindya, et al.
Published: (2025)
Diving into the Depths of Spotting Text in Multi-Domain Noisy Scenes
by: Das, Alloy, et al.
Published: (2023)
by: Das, Alloy, et al.
Published: (2023)
CraftGraffiti: Exploring Human Identity with Custom Graffiti Art via Facial-Preserving Diffusion Models
by: Banerjee, Ayan, et al.
Published: (2025)
by: Banerjee, Ayan, et al.
Published: (2025)
FastTextSpotter: A High-Efficiency Transformer for Multilingual Scene Text Spotting
by: Das, Alloy, et al.
Published: (2024)
by: Das, Alloy, et al.
Published: (2024)
FASTER: A Font-Agnostic Scene Text Editing and Rendering Framework
by: Das, Alloy, et al.
Published: (2023)
by: Das, Alloy, et al.
Published: (2023)
SketchGPT: Autoregressive Modeling for Sketch Generation and Recognition
by: Tiwari, Adarsh, et al.
Published: (2024)
by: Tiwari, Adarsh, et al.
Published: (2024)
Correlation Weighted Prototype-based Self-Supervised One-Shot Segmentation of Medical Images
by: Manna, Siladittya, et al.
Published: (2024)
by: Manna, Siladittya, et al.
Published: (2024)
Towards Generative Class Prompt Learning for Fine-grained Visual Recognition
by: Chattopadhyay, Soumitri, et al.
Published: (2024)
by: Chattopadhyay, Soumitri, et al.
Published: (2024)
DRG-Font: Dynamic Reference-Guided Few-shot Font Generation via Contrastive Style-Content Disentanglement
by: Chakraborty, Rejoy, et al.
Published: (2026)
by: Chakraborty, Rejoy, et al.
Published: (2026)
Multi-scale Attention Guided Pose Transfer
by: Roy, Prasun, et al.
Published: (2022)
by: Roy, Prasun, et al.
Published: (2022)
Human Knowledge Integrated Multi-modal Learning for Single Source Domain Generalization
by: Banerjee, Ayan, et al.
Published: (2026)
by: Banerjee, Ayan, et al.
Published: (2026)
The OCR Quest for Generalization: Learning to recognize low-resource alphabets with model editing
by: Rodríguez, Adrià Molina, et al.
Published: (2025)
by: Rodríguez, Adrià Molina, et al.
Published: (2025)
GeoContrastNet: Contrastive Key-Value Edge Learning for Language-Agnostic Document Understanding
by: Biescas, Nil, et al.
Published: (2024)
by: Biescas, Nil, et al.
Published: (2024)
A CNN Based Framework for Unistroke Numeral Recognition in Air-Writing
by: Roy, Prasun, et al.
Published: (2023)
by: Roy, Prasun, et al.
Published: (2023)
Towards Robust Cross-Dataset Object Detection Generalization under Domain Specificity
by: Chakraborty, Ritabrata, et al.
Published: (2026)
by: Chakraborty, Ritabrata, et al.
Published: (2026)
MIO : Mutual Information Optimization using Self-Supervised Binary Contrastive Learning
by: Manna, Siladittya, et al.
Published: (2021)
by: Manna, Siladittya, et al.
Published: (2021)
DistilDoc: Knowledge Distillation for Visually-Rich Document Applications
by: Van Landeghem, Jordy, et al.
Published: (2024)
by: Van Landeghem, Jordy, et al.
Published: (2024)
Semantically Consistent Person Image Generation
by: Roy, Prasun, et al.
Published: (2023)
by: Roy, Prasun, et al.
Published: (2023)
Fetch-A-Set: A Large-Scale OCR-Free Benchmark for Historical Document Retrieval
by: Molina, Adrià, et al.
Published: (2024)
by: Molina, Adrià, et al.
Published: (2024)
Scene Aware Person Image Generation through Global Contextual Conditioning
by: Roy, Prasun, et al.
Published: (2022)
by: Roy, Prasun, et al.
Published: (2022)
Character-Centered Dialogue Generation from Scene-Level Prompts
by: Kang, Taewon, et al.
Published: (2025)
by: Kang, Taewon, et al.
Published: (2025)
Exploring Mutual Cross-Modal Attention for Context-Aware Human Affordance Generation
by: Roy, Prasun, et al.
Published: (2025)
by: Roy, Prasun, et al.
Published: (2025)
Reliability-Aware Weighted Multi-Scale Spatio-Temporal Maps for Heart Rate Monitoring
by: Bairagi, Arpan, et al.
Published: (2026)
by: Bairagi, Arpan, et al.
Published: (2026)
StoryMaker: Towards Holistic Consistent Characters in Text-to-image Generation
by: Zhou, Zhengguang, et al.
Published: (2024)
by: Zhou, Zhengguang, et al.
Published: (2024)
MDIW-13: a New Multi-Lingual and Multi-Script Database and Benchmark for Script Identification
by: Ferrer, Miguel A., et al.
Published: (2024)
by: Ferrer, Miguel A., et al.
Published: (2024)
STEFANN: Scene Text Editor using Font Adaptive Neural Network
by: Roy, Prasun, et al.
Published: (2019)
by: Roy, Prasun, et al.
Published: (2019)
TaleForge: Interactive Multimodal System for Personalized Story Creation
by: Nguyen, Minh-Loi, et al.
Published: (2025)
by: Nguyen, Minh-Loi, et al.
Published: (2025)
Tails Tell Tales: Chapter-Wide Manga Transcriptions with Character Names
by: Sachdeva, Ragav, et al.
Published: (2024)
by: Sachdeva, Ragav, et al.
Published: (2024)
Decorrelation-based Self-Supervised Visual Representation Learning for Writer Identification
by: Maitra, Arkadip, et al.
Published: (2024)
by: Maitra, Arkadip, et al.
Published: (2024)
StoryWeaver: A Unified World Model for Knowledge-Enhanced Story Character Customization
by: Zhang, Jinlu, et al.
Published: (2024)
by: Zhang, Jinlu, et al.
Published: (2024)
Conformal uncertainty quantification to evaluate predictive fairness of foundation AI model for skin lesion classes across patient demographics
by: Bhattacharyya, Swarnava, et al.
Published: (2025)
by: Bhattacharyya, Swarnava, et al.
Published: (2025)
Persistent Story World Simulation with Continuous Character Customization
by: Zhang, Jinlu, et al.
Published: (2026)
by: Zhang, Jinlu, et al.
Published: (2026)
A Transformer Based Handwriting Recognition System Jointly Using Online and Offline Features
by: Lodh, Ayush, et al.
Published: (2025)
by: Lodh, Ayush, et al.
Published: (2025)
InfinityStory: Unlimited Video Generation with World Consistency and Character-Aware Shot Transitions
by: Elmoghany, Mohamed, et al.
Published: (2026)
by: Elmoghany, Mohamed, et al.
Published: (2026)
Masked Generative Story Transformer with Character Guidance and Caption Augmentation
by: Papadimitriou, Christos, et al.
Published: (2024)
by: Papadimitriou, Christos, et al.
Published: (2024)
Position and Rotation Invariant Sign Language Recognition from 3D Kinect Data with Recurrent Neural Networks
by: Roy, Prasun, et al.
Published: (2020)
by: Roy, Prasun, et al.
Published: (2020)
TIPS: Text-Induced Pose Synthesis
by: Roy, Prasun, et al.
Published: (2022)
by: Roy, Prasun, et al.
Published: (2022)
Similar Items
-
CraftSVG: Multi-Object Text-to-SVG Synthesis via Layout Guided Diffusion
by: Banerjee, Ayan, et al.
Published: (2024) -
DocRevive: A Unified Pipeline for Document Text Restoration
by: Purkayastha, Kunal, et al.
Published: (2026) -
GraphKD: Exploring Knowledge Distillation Towards Document Object Detection with Structured Graph Creation
by: Banerjee, Ayan, et al.
Published: (2024) -
CountLoop: Training-Free High-Instance Image Generation via Iterative Agent Guidance
by: Mondal, Anindya, et al.
Published: (2025) -
Diving into the Depths of Spotting Text in Multi-Domain Noisy Scenes
by: Das, Alloy, et al.
Published: (2023)