An Image is Worth Multiple Words: Discovering Object Level Concepts using Multi-Concept Prompt Learning
Fuente:
arXiv
Saved in:
| Main Authors: | Jin, Chen, Tanno, Ryutaro, Saseendran, Amrutha, Diethe, Tom, Teare, Philip |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Diffusion Instruction Tuning
by: Jin, Chen, et al.
Published: (2025)
by: Jin, Chen, et al.
Published: (2025)
CoRefine: Confidence-Guided Self-Refinement for Adaptive Test-Time Compute
by: Jin, Chen, et al.
Published: (2026)
by: Jin, Chen, et al.
Published: (2026)
DeCoRe: Decoding by Contrasting Retrieval Heads to Mitigate Hallucinations
by: Gema, Aryo Pradipta, et al.
Published: (2024)
by: Gema, Aryo Pradipta, et al.
Published: (2024)
PartComposer: Learning and Composing Part-Level Concepts from Single-Image Examples
by: Liu, Junyu, et al.
Published: (2025)
by: Liu, Junyu, et al.
Published: (2025)
T$^3$-S2S: Training-free Triplet Tuning for Sketch to Scene Synthesis in Controllable Concept Art Generation
by: Sun, Zhenhong, et al.
Published: (2024)
by: Sun, Zhenhong, et al.
Published: (2024)
Segment Anyword: Mask Prompt Inversion for Open-Set Grounded Segmentation
by: Liu, Zhihua, et al.
Published: (2025)
by: Liu, Zhihua, et al.
Published: (2025)
ComfyGen: Prompt-Adaptive Workflows for Text-to-Image Generation
by: Gal, Rinon, et al.
Published: (2024)
by: Gal, Rinon, et al.
Published: (2024)
Movie Weaver: Tuning-Free Multi-Concept Video Personalization with Anchored Prompts
by: Liang, Feng, et al.
Published: (2025)
by: Liang, Feng, et al.
Published: (2025)
PALP: Prompt Aligned Personalization of Text-to-Image Models
by: Arar, Moab, et al.
Published: (2024)
by: Arar, Moab, et al.
Published: (2024)
An Object is Worth 64x64 Pixels: Generating 3D Object via Image Diffusion
by: Yan, Xingguang, et al.
Published: (2024)
by: Yan, Xingguang, et al.
Published: (2024)
CustomSketching: Sketch Concept Extraction for Sketch-based Image Synthesis and Editing
by: Xiao, Chufeng, et al.
Published: (2024)
by: Xiao, Chufeng, et al.
Published: (2024)
From Words to Worlds: Transforming One-line Prompt into Immersive Multi-modal Digital Stories with Communicative LLM Agent
by: Sohn, Samuel S., et al.
Published: (2024)
by: Sohn, Samuel S., et al.
Published: (2024)
IP-Composer: Semantic Composition of Visual Concepts
by: Dorfman, Sara, et al.
Published: (2025)
by: Dorfman, Sara, et al.
Published: (2025)
CAP: Evaluation of Persuasive and Creative Image Generation
by: Aghazadeh, Aysan, et al.
Published: (2024)
by: Aghazadeh, Aysan, et al.
Published: (2024)
Design2GarmentCode: Turning Design Concepts to Tangible Garments Through Program Synthesis
by: Zhou, Feng, et al.
Published: (2024)
by: Zhou, Feng, et al.
Published: (2024)
MagiCapture: High-Resolution Multi-Concept Portrait Customization
by: Hyung, Junha, et al.
Published: (2023)
by: Hyung, Junha, et al.
Published: (2023)
Global Position Aware Group Choreography using Large Language Model
by: Pang, Haozhou, et al.
Published: (2025)
by: Pang, Haozhou, et al.
Published: (2025)
CharNeRF: 3D Character Generation from Concept Art
by: Chu, Eddy, et al.
Published: (2024)
by: Chu, Eddy, et al.
Published: (2024)
Object-level Visual Prompts for Compositional Image Generation
by: Parmar, Gaurav, et al.
Published: (2025)
by: Parmar, Gaurav, et al.
Published: (2025)
ORACLE: Orchestrate NPC Daily Activities using Contrastive Learning with Transformer-CVAE
by: Hong, Seong-Eun, et al.
Published: (2026)
by: Hong, Seong-Eun, et al.
Published: (2026)
Dynamic Concepts Personalization from Single Videos
by: Abdal, Rameen, et al.
Published: (2025)
by: Abdal, Rameen, et al.
Published: (2025)
Multi-LoRA Composition for Image Generation
by: Zhong, Ming, et al.
Published: (2024)
by: Zhong, Ming, et al.
Published: (2024)
Image Generation Models: A Technical History
by: Shirvani, Rouzbeh
Published: (2026)
by: Shirvani, Rouzbeh
Published: (2026)
Real-time Image-based Lighting of Glints
by: Kneiphof, Tom, et al.
Published: (2025)
by: Kneiphof, Tom, et al.
Published: (2025)
Grounding Language in Multi-Perspective Referential Communication
by: Tang, Zineng, et al.
Published: (2024)
by: Tang, Zineng, et al.
Published: (2024)
Learning Explicit Contact for Implicit Reconstruction of Hand-held Objects from Monocular Images
by: Hu, Junxing, et al.
Published: (2023)
by: Hu, Junxing, et al.
Published: (2023)
FIFA: Unified Faithfulness Evaluation Framework for Text-to-Video and Video-to-Text Generation
by: Jing, Liqiang, et al.
Published: (2025)
by: Jing, Liqiang, et al.
Published: (2025)
Is this chart lying to me? Automating the detection of misleading visualizations
by: Tonglet, Jonathan, et al.
Published: (2025)
by: Tonglet, Jonathan, et al.
Published: (2025)
A Study of the Framework and Real-World Applications of Language Embedding for 3D Scene Understanding
by: Zaouali, Mahmoud Chick, et al.
Published: (2025)
by: Zaouali, Mahmoud Chick, et al.
Published: (2025)
TexGS-VolVis: Expressive Scene Editing for Volume Visualization via Textured Gaussian Splatting
by: Tang, Kaiyuan, et al.
Published: (2025)
by: Tang, Kaiyuan, et al.
Published: (2025)
Is Your World Simulator a Good Story Presenter? A Consecutive Events-Based Benchmark for Future Long Video Generation
by: Wang, Yiping, et al.
Published: (2024)
by: Wang, Yiping, et al.
Published: (2024)
FlairGPT: Repurposing LLMs for Interior Designs
by: Littlefair, Gabrielle, et al.
Published: (2025)
by: Littlefair, Gabrielle, et al.
Published: (2025)
Co-Layout: LLM-driven Co-optimization for Interior Layout
by: Xiang, Chucheng, et al.
Published: (2025)
by: Xiang, Chucheng, et al.
Published: (2025)
Learning to Infer Generative Template Programs for Visual Concepts
by: Jones, R. Kenny, et al.
Published: (2024)
by: Jones, R. Kenny, et al.
Published: (2024)
GarVerseLOD: High-Fidelity 3D Garment Reconstruction from a Single In-the-Wild Image using a Dataset with Levels of Details
by: Luo, Zhongjin, et al.
Published: (2024)
by: Luo, Zhongjin, et al.
Published: (2024)
Nested Attention: Semantic-aware Attention Values for Concept Personalization
by: Patashnik, Or, et al.
Published: (2025)
by: Patashnik, Or, et al.
Published: (2025)
Alterbute: Editing Intrinsic Attributes of Objects in Images
by: Reiss, Tal, et al.
Published: (2026)
by: Reiss, Tal, et al.
Published: (2026)
MM-Conv: A Multi-modal Conversational Dataset for Virtual Humans
by: Deichler, Anna, et al.
Published: (2024)
by: Deichler, Anna, et al.
Published: (2024)
Gesture2Text: A Generalizable Decoder for Word-Gesture Keyboards in XR Through Trajectory Coarse Discretization and Pre-training
by: Shen, Junxiao, et al.
Published: (2024)
by: Shen, Junxiao, et al.
Published: (2024)
TelePhysics: Physics-Grounded Multi-Object Scene Generation from a Single Image with Real-Time Interaction
by: Zhang, Xin, et al.
Published: (2026)
by: Zhang, Xin, et al.
Published: (2026)
Similar Items
-
Diffusion Instruction Tuning
by: Jin, Chen, et al.
Published: (2025) -
CoRefine: Confidence-Guided Self-Refinement for Adaptive Test-Time Compute
by: Jin, Chen, et al.
Published: (2026) -
DeCoRe: Decoding by Contrasting Retrieval Heads to Mitigate Hallucinations
by: Gema, Aryo Pradipta, et al.
Published: (2024) -
PartComposer: Learning and Composing Part-Level Concepts from Single-Image Examples
by: Liu, Junyu, et al.
Published: (2025) -
T$^3$-S2S: Training-free Triplet Tuning for Sketch to Scene Synthesis in Controllable Concept Art Generation
by: Sun, Zhenhong, et al.
Published: (2024)