Scene Aware Person Image Generation through Global Contextual Conditioning
Fuente:
arXiv
Saved in:
| Main Authors: | Roy, Prasun, Ghosh, Subhankar, Bhattacharya, Saumik, Pal, Umapada, Blumenstein, Michael |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Semantically Consistent Person Image Generation
by: Roy, Prasun, et al.
Published: (2023)
by: Roy, Prasun, et al.
Published: (2023)
Exploring Mutual Cross-Modal Attention for Context-Aware Human Affordance Generation
by: Roy, Prasun, et al.
Published: (2025)
by: Roy, Prasun, et al.
Published: (2025)
TIPS: Text-Induced Pose Synthesis
by: Roy, Prasun, et al.
Published: (2022)
by: Roy, Prasun, et al.
Published: (2022)
STEFANN: Scene Text Editor using Font Adaptive Neural Network
by: Roy, Prasun, et al.
Published: (2019)
by: Roy, Prasun, et al.
Published: (2019)
Multi-scale Attention Guided Pose Transfer
by: Roy, Prasun, et al.
Published: (2022)
by: Roy, Prasun, et al.
Published: (2022)
d-Sketch: Improving Visual Fidelity of Sketch-to-Image Translation with Pretrained Latent Diffusion Models without Retraining
by: Roy, Prasun, et al.
Published: (2025)
by: Roy, Prasun, et al.
Published: (2025)
FASTER: A Font-Agnostic Scene Text Editing and Rendering Framework
by: Das, Alloy, et al.
Published: (2023)
by: Das, Alloy, et al.
Published: (2023)
Effects of Degradations on Deep Neural Network Architectures
by: Roy, Prasun, et al.
Published: (2018)
by: Roy, Prasun, et al.
Published: (2018)
DRG-Font: Dynamic Reference-Guided Few-shot Font Generation via Contrastive Style-Content Disentanglement
by: Chakraborty, Rejoy, et al.
Published: (2026)
by: Chakraborty, Rejoy, et al.
Published: (2026)
A CNN Based Framework for Unistroke Numeral Recognition in Air-Writing
by: Roy, Prasun, et al.
Published: (2023)
by: Roy, Prasun, et al.
Published: (2023)
Position and Rotation Invariant Sign Language Recognition from 3D Kinect Data with Recurrent Neural Networks
by: Roy, Prasun, et al.
Published: (2020)
by: Roy, Prasun, et al.
Published: (2020)
Correlation Weighted Prototype-based Self-Supervised One-Shot Segmentation of Medical Images
by: Manna, Siladittya, et al.
Published: (2024)
by: Manna, Siladittya, et al.
Published: (2024)
MIO : Mutual Information Optimization using Self-Supervised Binary Contrastive Learning
by: Manna, Siladittya, et al.
Published: (2021)
by: Manna, Siladittya, et al.
Published: (2021)
FastTextSpotter: A High-Efficiency Transformer for Multilingual Scene Text Spotting
by: Das, Alloy, et al.
Published: (2024)
by: Das, Alloy, et al.
Published: (2024)
Modality-Aware Shot Relating and Comparing for Video Scene Detection
by: Tan, Jiawei, et al.
Published: (2024)
by: Tan, Jiawei, et al.
Published: (2024)
CIV-DG: Conditional Instrumental Variables for Domain Generalization in Medical Imaging
by: Bai, Shaojin, et al.
Published: (2026)
by: Bai, Shaojin, et al.
Published: (2026)
Single Image Dehazing Using Scene Depth Ordering
by: Ling, Pengyang, et al.
Published: (2024)
by: Ling, Pengyang, et al.
Published: (2024)
SceneDreamer360: Text-Driven 3D-Consistent Scene Generation with Panoramic Gaussian Splatting
by: Li, Wenrui, et al.
Published: (2024)
by: Li, Wenrui, et al.
Published: (2024)
EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering
by: Zhou, Sheng, et al.
Published: (2025)
by: Zhou, Sheng, et al.
Published: (2025)
SOSControl: Enhancing Human Motion Generation through Saliency-Aware Symbolic Orientation and Timing Control
by: Au, Ho Yin, et al.
Published: (2025)
by: Au, Ho Yin, et al.
Published: (2025)
ContextBLIP: Doubly Contextual Alignment for Contrastive Image Retrieval from Linguistically Complex Descriptions
by: Lin, Honglin, et al.
Published: (2024)
by: Lin, Honglin, et al.
Published: (2024)
Decorrelation-based Self-Supervised Visual Representation Learning for Writer Identification
by: Maitra, Arkadip, et al.
Published: (2024)
by: Maitra, Arkadip, et al.
Published: (2024)
Scene Graph Generation with Role-Playing Large Language Models
by: Chen, Guikun, et al.
Published: (2024)
by: Chen, Guikun, et al.
Published: (2024)
Multiverse Through Deepfakes: The MultiFakeVerse Dataset of Person-Centric Visual and Conceptual Manipulations
by: Gupta, Parul, et al.
Published: (2025)
by: Gupta, Parul, et al.
Published: (2025)
Noisy-Correspondence Learning for Text-to-Image Person Re-identification
by: Qin, Yang, et al.
Published: (2023)
by: Qin, Yang, et al.
Published: (2023)
AdaptaGen: Domain-Specific Image Generation through Hierarchical Semantic Optimization Framework
by: Zhang, Suoxiang, et al.
Published: (2025)
by: Zhang, Suoxiang, et al.
Published: (2025)
Creatively Upscaling Images with Global-Regional Priors
by: Qian, Yurui, et al.
Published: (2025)
by: Qian, Yurui, et al.
Published: (2025)
ROGLE: Robust Global-Local Alignment with Automated Region Supervision for Text-Based Person Search
by: Xie, Zequn, et al.
Published: (2026)
by: Xie, Zequn, et al.
Published: (2026)
Unbiased Video Scene Graph Generation via Visual and Semantic Dual Debiasing
by: Li, Yanjun, et al.
Published: (2025)
by: Li, Yanjun, et al.
Published: (2025)
HiScene: Creating Hierarchical 3D Scenes with Isometric View Generation
by: Dong, Wenqi, et al.
Published: (2025)
by: Dong, Wenqi, et al.
Published: (2025)
HiGS: Hierarchical Generative Scene Framework for Multi-Step Associative Semantic Spatial Composition
by: Hong, Jiacheng, et al.
Published: (2025)
by: Hong, Jiacheng, et al.
Published: (2025)
Dual Mutual Learning Network with Global-local Awareness for RGB-D Salient Object Detection
by: Yi, Kang, et al.
Published: (2025)
by: Yi, Kang, et al.
Published: (2025)
Beyond Walking: A Large-Scale Image-Text Benchmark for Text-based Person Anomaly Search
by: Yang, Shuyu, et al.
Published: (2024)
by: Yang, Shuyu, et al.
Published: (2024)
Generating Attribute-Aware Human Motions from Textual Prompt
by: Wang, Xinghan, et al.
Published: (2025)
by: Wang, Xinghan, et al.
Published: (2025)
"Humor, Art, or Misinformation?": A Multimodal Dataset for Intent-Aware Synthetic Image Detection
by: Skoularikis, Anastasios, et al.
Published: (2025)
by: Skoularikis, Anastasios, et al.
Published: (2025)
Alfie: Democratising RGBA Image Generation With No $$$
by: Quattrini, Fabio, et al.
Published: (2024)
by: Quattrini, Fabio, et al.
Published: (2024)
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
by: Chen, Liyang, et al.
Published: (2025)
by: Chen, Liyang, et al.
Published: (2025)
CalliReader: Contextualizing Chinese Calligraphy via an Embedding-Aligned Vision-Language Model
by: Luo, Yuxuan, et al.
Published: (2025)
by: Luo, Yuxuan, et al.
Published: (2025)
TALDS-Net: Task-Aware Adaptive Local Descriptors Selection for Few-shot Image Classification
by: Qiao, Qian, et al.
Published: (2023)
by: Qiao, Qian, et al.
Published: (2023)
SmartFreeEdit: Mask-Free Spatial-Aware Image Editing with Complex Instruction Understanding
by: Sun, Qianqian, et al.
Published: (2025)
by: Sun, Qianqian, et al.
Published: (2025)
Similar Items
-
Semantically Consistent Person Image Generation
by: Roy, Prasun, et al.
Published: (2023) -
Exploring Mutual Cross-Modal Attention for Context-Aware Human Affordance Generation
by: Roy, Prasun, et al.
Published: (2025) -
TIPS: Text-Induced Pose Synthesis
by: Roy, Prasun, et al.
Published: (2022) -
STEFANN: Scene Text Editor using Font Adaptive Neural Network
by: Roy, Prasun, et al.
Published: (2019) -
Multi-scale Attention Guided Pose Transfer
by: Roy, Prasun, et al.
Published: (2022)