Language-Grounded Multi-Domain Image Translation via Semantic Difference Guidance
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ryu, Jongwon, Park, Joonhyung, Han, Jaeho, Kim, Yeong-Seok, Kim, Hye-rin, Yoon, Sunjae, Kim, Junyeong |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Appearance Matching Adapter for Exemplar-based Semantic Image Synthesis in-the-Wild
von: Jin, Siyoon, et al.
Veröffentlicht: (2024)
von: Jin, Siyoon, et al.
Veröffentlicht: (2024)
Visual Funnel: Resolving Contextual Blindness in Multimodal Large Language Models
von: Jung, Woojun, et al.
Veröffentlicht: (2025)
von: Jung, Woojun, et al.
Veröffentlicht: (2025)
Unveiling the Response of Large Vision-Language Models to Visually Absent Tokens
von: Kim, Sohee, et al.
Veröffentlicht: (2025)
von: Kim, Sohee, et al.
Veröffentlicht: (2025)
MEGA-GUI: Multi-stage Enhanced Grounding Agents for GUI Elements
von: Kwak, SeokJoo, et al.
Veröffentlicht: (2025)
von: Kwak, SeokJoo, et al.
Veröffentlicht: (2025)
Selective Query-guided Debiasing for Video Corpus Moment Retrieval
von: Yoon, Sunjae, et al.
Veröffentlicht: (2022)
von: Yoon, Sunjae, et al.
Veröffentlicht: (2022)
Contrastive Language Prompting to Ease False Positives in Medical Anomaly Detection
von: Park, YeongHyeon, et al.
Veröffentlicht: (2024)
von: Park, YeongHyeon, et al.
Veröffentlicht: (2024)
GranAlign: Granularity-Aware Alignment Framework for Zero-Shot Video Moment Retrieval
von: Jeon, Mingyu, et al.
Veröffentlicht: (2026)
von: Jeon, Mingyu, et al.
Veröffentlicht: (2026)
CollabVR: Collaborative Video Reasoning with Vision-Language and Video Generation Models
von: Kim, Joowon, et al.
Veröffentlicht: (2026)
von: Kim, Joowon, et al.
Veröffentlicht: (2026)
Progress by Pieces: Test-Time Scaling for Autoregressive Image Generation
von: Park, Joonhyung, et al.
Veröffentlicht: (2025)
von: Park, Joonhyung, et al.
Veröffentlicht: (2025)
Robust Grounding with MLLMs Against Occlusion and Small Objects via Language-Guided Semantic Cues
von: Park, Beomchan, et al.
Veröffentlicht: (2026)
von: Park, Beomchan, et al.
Veröffentlicht: (2026)
Finetuning Pre-trained Model with Limited Data for LiDAR-based 3D Object Detection by Bridging Domain Gaps
von: Jang, Jiyun, et al.
Veröffentlicht: (2024)
von: Jang, Jiyun, et al.
Veröffentlicht: (2024)
Empirical Analysis of Anomaly Detection on Hyperspectral Imaging Using Dimension Reduction Methods
von: Kim, Dongeon, et al.
Veröffentlicht: (2024)
von: Kim, Dongeon, et al.
Veröffentlicht: (2024)
Diffusion Model Compression for Image-to-Image Translation
von: Kim, Geonung, et al.
Veröffentlicht: (2024)
von: Kim, Geonung, et al.
Veröffentlicht: (2024)
QEVA: A Reference-Free Evaluation Metric for Narrative Video Summarization with Multimodal Question Answering
von: Jung, Woojun, et al.
Veröffentlicht: (2026)
von: Jung, Woojun, et al.
Veröffentlicht: (2026)
Merge-Friendly Post-Training Quantization for Multi-Target Domain Adaptation
von: Shin, Juncheol, et al.
Veröffentlicht: (2025)
von: Shin, Juncheol, et al.
Veröffentlicht: (2025)
HEAR: Hearing Enhanced Audio Response for Video-grounded Dialogue
von: Yoon, Sunjae, et al.
Veröffentlicht: (2023)
von: Yoon, Sunjae, et al.
Veröffentlicht: (2023)
Coarse-to-Fine: Progressive Image Compression for Semantically Hierarchical Classification
von: Kim, Jungwoo, et al.
Veröffentlicht: (2026)
von: Kim, Jungwoo, et al.
Veröffentlicht: (2026)
Adaptive Latent Diffusion Model for 3D Medical Image to Image Translation: Multi-modal Magnetic Resonance Imaging Study
von: Kim, Jonghun, et al.
Veröffentlicht: (2023)
von: Kim, Jonghun, et al.
Veröffentlicht: (2023)
Semantic-Aware Reconstruction Error for Detecting AI-Generated Images
von: Kang, Ju Yeon, et al.
Veröffentlicht: (2025)
von: Kang, Ju Yeon, et al.
Veröffentlicht: (2025)
Object Re-identification via Spatial-temporal Fusion Networks and Causal Identity Matching
von: Kim, Hye-Geun, et al.
Veröffentlicht: (2024)
von: Kim, Hye-Geun, et al.
Veröffentlicht: (2024)
Point to Span: Zero-Shot Moment Retrieval for Navigating Unseen Hour-Long Videos
von: Jeon, Mingyu, et al.
Veröffentlicht: (2025)
von: Jeon, Mingyu, et al.
Veröffentlicht: (2025)
SC-Pro: Training-Free Framework for Defending Unsafe Image Synthesis Attack
von: Park, Junha, et al.
Veröffentlicht: (2025)
von: Park, Junha, et al.
Veröffentlicht: (2025)
FRAG: Frequency Adapting Group for Diffusion Video Editing
von: Yoon, Sunjae, et al.
Veröffentlicht: (2024)
von: Yoon, Sunjae, et al.
Veröffentlicht: (2024)
SCANet: Scene Complexity Aware Network for Weakly-Supervised Video Moment Retrieval
von: Yoon, Sunjae, et al.
Veröffentlicht: (2023)
von: Yoon, Sunjae, et al.
Veröffentlicht: (2023)
Variation-Aware Semantic Image Synthesis
von: Xu, Mingle, et al.
Veröffentlicht: (2023)
von: Xu, Mingle, et al.
Veröffentlicht: (2023)
The Pragmatic Persona: Discovering LLM Persona through Bridging Inference
von: Yang, Jisoo, et al.
Veröffentlicht: (2026)
von: Yang, Jisoo, et al.
Veröffentlicht: (2026)
Remote Sensing Large Vision-Language Model: Semantic-augmented Multi-level Alignment and Semantic-aware Expert Modeling
von: Park, Sungjune, et al.
Veröffentlicht: (2025)
von: Park, Sungjune, et al.
Veröffentlicht: (2025)
From Prompts to Deployment: Auto-Curated Domain-Specific Dataset Generation via Diffusion Models
von: Yoon, Dongsik, et al.
Veröffentlicht: (2026)
von: Yoon, Dongsik, et al.
Veröffentlicht: (2026)
Domain-generalizable Face Anti-Spoofing with Patch-based Multi-tasking and Artifact Pattern Conversion
von: Jung, Seungjin, et al.
Veröffentlicht: (2026)
von: Jung, Seungjin, et al.
Veröffentlicht: (2026)
Revisiting Reliability in the Reasoning-based Pose Estimation Benchmark
von: Kim, Junsu, et al.
Veröffentlicht: (2025)
von: Kim, Junsu, et al.
Veröffentlicht: (2025)
Generalizable Human Gaussian Splatting via Multi-view Semantic Consistency
von: Kim, Jingi, et al.
Veröffentlicht: (2026)
von: Kim, Jingi, et al.
Veröffentlicht: (2026)
Difference Inversion: Interpolate and Isolate the Difference with Token Consistency for Image Analogy Generation
von: Kim, Hyunsoo, et al.
Veröffentlicht: (2025)
von: Kim, Hyunsoo, et al.
Veröffentlicht: (2025)
Feature Attenuation of Defective Representation Can Resolve Incomplete Masking on Anomaly Detection
von: Park, YeongHyeon, et al.
Veröffentlicht: (2024)
von: Park, YeongHyeon, et al.
Veröffentlicht: (2024)
PruNeRF: Segment-Centric Dataset Pruning via 3D Spatial Consistency
von: Jung, Yeonsung, et al.
Veröffentlicht: (2024)
von: Jung, Yeonsung, et al.
Veröffentlicht: (2024)
Activation Quantization of Vision Encoders Needs Prefixing Registers
von: Kim, Seunghyeon, et al.
Veröffentlicht: (2025)
von: Kim, Seunghyeon, et al.
Veröffentlicht: (2025)
PeLiCal: Targetless Extrinsic Calibration via Penetrating Lines for RGB-D Cameras with Limited Co-visibility
von: Shin, Jaeho, et al.
Veröffentlicht: (2024)
von: Shin, Jaeho, et al.
Veröffentlicht: (2024)
ESD: Expected Squared Difference as a Tuning-Free Trainable Calibration Measure
von: Yoon, Hee Suk, et al.
Veröffentlicht: (2023)
von: Yoon, Hee Suk, et al.
Veröffentlicht: (2023)
IG-FIQA: Improving Face Image Quality Assessment through Intra-class Variance Guidance robust to Inaccurate Pseudo-Labels
von: Kim, Minsoo, et al.
Veröffentlicht: (2024)
von: Kim, Minsoo, et al.
Veröffentlicht: (2024)
Linear Algebraic Approaches to Neuroimaging Data Compression: A Comparative Analysis of Matrix and Tensor Decomposition Methods for High-Dimensional Medical Images
von: Kim, Jaeho, et al.
Veröffentlicht: (2025)
von: Kim, Jaeho, et al.
Veröffentlicht: (2025)
Is it safe to cross? Interpretable Risk Assessment with GPT-4V for Safety-Aware Street Crossing
von: Hwang, Hochul, et al.
Veröffentlicht: (2024)
von: Hwang, Hochul, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Appearance Matching Adapter for Exemplar-based Semantic Image Synthesis in-the-Wild
von: Jin, Siyoon, et al.
Veröffentlicht: (2024) -
Visual Funnel: Resolving Contextual Blindness in Multimodal Large Language Models
von: Jung, Woojun, et al.
Veröffentlicht: (2025) -
Unveiling the Response of Large Vision-Language Models to Visually Absent Tokens
von: Kim, Sohee, et al.
Veröffentlicht: (2025) -
MEGA-GUI: Multi-stage Enhanced Grounding Agents for GUI Elements
von: Kwak, SeokJoo, et al.
Veröffentlicht: (2025) -
Selective Query-guided Debiasing for Video Corpus Moment Retrieval
von: Yoon, Sunjae, et al.
Veröffentlicht: (2022)