Saved in:
| Main Authors: | Miyamoto, Mizuki, Morita, Ryugo, Zhou, Jinjia |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2411.10183 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Block based Adaptive Compressive Sensing with Sampling Rate Control
by: Iwama, Kosuke, et al.
Published: (2024)
by: Iwama, Kosuke, et al.
Published: (2024)
Audio-Visual Driven Compression for Low-Bitrate Talking Head Videos
by: Takahashi, Riku, et al.
Published: (2025)
by: Takahashi, Riku, et al.
Published: (2025)
Edge-based Denoising Image Compression
by: Morita, Ryugo, et al.
Published: (2024)
by: Morita, Ryugo, et al.
Published: (2024)
Bidirectional Learned Facial Animation Codec for Low Bitrate Talking Head Videos
by: Takahashi, Riku, et al.
Published: (2025)
by: Takahashi, Riku, et al.
Published: (2025)
TKG-DM: Training-free Chroma Key Content Generation Diffusion Model
by: Morita, Ryugo, et al.
Published: (2024)
by: Morita, Ryugo, et al.
Published: (2024)
Visual question answering-based image-finding generation for pulmonary nodules on chest CT from structured annotations
by: Nagao, Maiko, et al.
Published: (2026)
by: Nagao, Maiko, et al.
Published: (2026)
TAUE: Training-free Noise Transplant and Cultivation Diffusion Model
by: Nagai, Daichi, et al.
Published: (2025)
by: Nagai, Daichi, et al.
Published: (2025)
Visual question answering: from early developments to recent advances -- a survey
by: Huynh, Ngoc Dung, et al.
Published: (2025)
by: Huynh, Ngoc Dung, et al.
Published: (2025)
Multi-perspective Contrastive Logit Distillation
by: Wang, Qi, et al.
Published: (2024)
by: Wang, Qi, et al.
Published: (2024)
TopKD: Top-scaled Knowledge Distillation
by: Wang, Qi, et al.
Published: (2025)
by: Wang, Qi, et al.
Published: (2025)
LGTM: Training-Free Light-Guided Text-to-Image Diffusion Model via Initial Noise Manipulation
by: Morita, Ryugo, et al.
Published: (2026)
by: Morita, Ryugo, et al.
Published: (2026)
LimeCross: Context-Conditioned Layered Image Editing with Structural Consistency
by: Morita, Ryugo, et al.
Published: (2026)
by: Morita, Ryugo, et al.
Published: (2026)
KG-CMI: Knowledge graph enhanced cross-Mamba interaction for medical visual question answering
by: Zheng, Xianyao, et al.
Published: (2026)
by: Zheng, Xianyao, et al.
Published: (2026)
Exploring text-to-image generation for historical document image retrieval
by: Cote, Melissa, et al.
Published: (2025)
by: Cote, Melissa, et al.
Published: (2025)
Coarse Semantic Injection for LLM-Conditioned Structured Indoor Prediction
by: Zhu, Shuliang, et al.
Published: (2026)
by: Zhu, Shuliang, et al.
Published: (2026)
PointT2I: LLM-based text-to-image generation via keypoints
by: Lee, Taekyung, et al.
Published: (2025)
by: Lee, Taekyung, et al.
Published: (2025)
CMI-MTL: Cross-Mamba interaction based multi-task learning for medical visual question answering
by: Jin, Qiangguo, et al.
Published: (2025)
by: Jin, Qiangguo, et al.
Published: (2025)
Consistent text-to-image generation via scene de-contextualization
by: Tang, Song, et al.
Published: (2025)
by: Tang, Song, et al.
Published: (2025)
IMAGE-ALCHEMY: Advancing subject fidelity in personalised text-to-image generation
by: Tiwari, Amritanshu, et al.
Published: (2025)
by: Tiwari, Amritanshu, et al.
Published: (2025)
Adaptive Sampling Scheduler
by: Wang, Qi, et al.
Published: (2025)
by: Wang, Qi, et al.
Published: (2025)
Investigation to answer three key questions concerning plant pest identification and development of a practical identification framework
by: Wayama, Ryosuke, et al.
Published: (2024)
by: Wayama, Ryosuke, et al.
Published: (2024)
Trinity Detector:text-assisted and attention mechanisms based spectral fusion for diffusion generation image detection
by: Song, Jiawei, et al.
Published: (2024)
by: Song, Jiawei, et al.
Published: (2024)
Dark Miner: Defend against undesirable generation for text-to-image diffusion models
by: Meng, Zheling, et al.
Published: (2024)
by: Meng, Zheling, et al.
Published: (2024)
Several questions of visual generation in 2024
by: Gu, Shuyang
Published: (2024)
by: Gu, Shuyang
Published: (2024)
Decomposed evaluations of geographic disparities in text-to-image models
by: Sureddy, Abhishek, et al.
Published: (2024)
by: Sureddy, Abhishek, et al.
Published: (2024)
Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No!
by: Imam, Mohamed Fazli, et al.
Published: (2025)
by: Imam, Mohamed Fazli, et al.
Published: (2025)
Anchor Learning with Potential Cluster Constraints for Multi-view Clustering
by: Chen, Yawei, et al.
Published: (2024)
by: Chen, Yawei, et al.
Published: (2024)
Performance evaluation of deep learning models for image analysis: considerations for visual control and statistical metrics
by: Bertram, Christof A., et al.
Published: (2026)
by: Bertram, Christof A., et al.
Published: (2026)
Efficient scene text image super-resolution with semantic guidance
by: TomyEnrique, LeoWu, et al.
Published: (2024)
by: TomyEnrique, LeoWu, et al.
Published: (2024)
Metrics that matter: Evaluating image quality metrics for medical image generation
by: Deo, Yash, et al.
Published: (2025)
by: Deo, Yash, et al.
Published: (2025)
Learn from Balance: Rectifying Knowledge Transfer for Long-Tailed Scenarios
by: Huang, Xinlei, et al.
Published: (2024)
by: Huang, Xinlei, et al.
Published: (2024)
TurboEdit: Instant text-based image editing
by: Wu, Zongze, et al.
Published: (2024)
by: Wu, Zongze, et al.
Published: (2024)
Understanding implementation pitfalls of distance-based metrics for image segmentation
by: Podobnik, Gasper, et al.
Published: (2024)
by: Podobnik, Gasper, et al.
Published: (2024)
Quantifying the uncertainty of model-based synthetic image quality metrics
by: Bench, Ciaran, et al.
Published: (2025)
by: Bench, Ciaran, et al.
Published: (2025)
Scene-Adaptive Person Search via Bilateral Modulations
by: Jiang, Yimin, et al.
Published: (2024)
by: Jiang, Yimin, et al.
Published: (2024)
Unsupervised Domain Adaptive Person Search via Dual Self-Calibration
by: Qi, Linfeng, et al.
Published: (2024)
by: Qi, Linfeng, et al.
Published: (2024)
ImagenHub: Standardizing the evaluation of conditional image generation models
by: Ku, Max, et al.
Published: (2023)
by: Ku, Max, et al.
Published: (2023)
Learning text-to-video retrieval from image captioning
by: Ventura, Lucas, et al.
Published: (2024)
by: Ventura, Lucas, et al.
Published: (2024)
Hyper-parameter tuning for text guided image editing
by: Zhang, Shiwen
Published: (2024)
by: Zhang, Shiwen
Published: (2024)
LEAST: "Local" text-conditioned image style transfer
by: Singh, Silky, et al.
Published: (2024)
by: Singh, Silky, et al.
Published: (2024)
Similar Items
-
Block based Adaptive Compressive Sensing with Sampling Rate Control
by: Iwama, Kosuke, et al.
Published: (2024) -
Audio-Visual Driven Compression for Low-Bitrate Talking Head Videos
by: Takahashi, Riku, et al.
Published: (2025) -
Edge-based Denoising Image Compression
by: Morita, Ryugo, et al.
Published: (2024) -
Bidirectional Learned Facial Animation Codec for Low Bitrate Talking Head Videos
by: Takahashi, Riku, et al.
Published: (2025) -
TKG-DM: Training-free Chroma Key Content Generation Diffusion Model
by: Morita, Ryugo, et al.
Published: (2024)