Saved in:
| Main Authors: | Cai, Kaiwen, Lu, Chris Xiaoxuan, Zhao, Xingyu, Huang, Xiaowei |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2307.07336 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Self-Adapting Large Visual-Language Models to Edge Devices across Visual Modalities
by: Cai, Kaiwen, et al.
Published: (2024)
by: Cai, Kaiwen, et al.
Published: (2024)
FlexIP: Dynamic Control of Preservation and Personality for Customized Image Generation
by: Huang, Linyan, et al.
Published: (2025)
by: Huang, Linyan, et al.
Published: (2025)
RichControl: Structure- and Appearance-Rich Training-Free Spatial Control for Text-to-Image Generation
by: Pang, Lexi, et al.
Published: (2025)
by: Pang, Lexi, et al.
Published: (2025)
Robust 3D Object Detection from LiDAR-Radar Point Clouds via Cross-Modal Feature Augmentation
by: Deng, Jianning, et al.
Published: (2023)
by: Deng, Jianning, et al.
Published: (2023)
Image Augmentation with Controlled Diffusion for Weakly-Supervised Semantic Segmentation
by: Wu, Wangyu, et al.
Published: (2023)
by: Wu, Wangyu, et al.
Published: (2023)
AeroReformer: Aerial Referring Transformer for UAV-based Referring Image Segmentation
by: Li, Rui, et al.
Published: (2025)
by: Li, Rui, et al.
Published: (2025)
milliFlow: Scene Flow Estimation on mmWave Radar Point Cloud for Human Motion Sensing
by: Ding, Fangqiang, et al.
Published: (2023)
by: Ding, Fangqiang, et al.
Published: (2023)
High-quality Image Dehazing with Diffusion Model
by: Yu, Hu, et al.
Published: (2023)
by: Yu, Hu, et al.
Published: (2023)
TAIJI: Textual Anchoring for Immunizing Jailbreak Images in Vision Language Models
by: Yin, Xiangyu, et al.
Published: (2025)
by: Yin, Xiangyu, et al.
Published: (2025)
Fragile by Design: On the Limits of Adversarial Defenses in Personalized Generation
by: Chen, Zhen, et al.
Published: (2025)
by: Chen, Zhen, et al.
Published: (2025)
SIDA: Social Media Image Deepfake Detection, Localization and Explanation with Large Multimodal Model
by: Huang, Zhenglin, et al.
Published: (2024)
by: Huang, Zhenglin, et al.
Published: (2024)
ProTIP: Probabilistic Robustness Verification on Text-to-Image Diffusion Models against Stochastic Perturbation
by: Zhang, Yi, et al.
Published: (2024)
by: Zhang, Yi, et al.
Published: (2024)
CPCL: Cross-Modal Prototypical Contrastive Learning for Weakly Supervised Text-based Person Retrieval
by: Zhao, Xinpeng, et al.
Published: (2024)
by: Zhao, Xinpeng, et al.
Published: (2024)
Training-free Test-time Improvement for Explainable Medical Image Classification
by: He, Hangzhou, et al.
Published: (2025)
by: He, Hangzhou, et al.
Published: (2025)
RadarOcc: Robust 3D Occupancy Prediction with 4D Imaging Radar
by: Ding, Fangqiang, et al.
Published: (2024)
by: Ding, Fangqiang, et al.
Published: (2024)
HINT: Composed Image Retrieval with Dual-path Compositional Contextualized Network
by: Zhang, Mingyu, et al.
Published: (2026)
by: Zhang, Mingyu, et al.
Published: (2026)
Data-Efficient Generalization for Zero-shot Composed Image Retrieval
by: Chen, Zining, et al.
Published: (2025)
by: Chen, Zining, et al.
Published: (2025)
ThermoHands: A Benchmark for 3D Hand Pose Estimation from Egocentric Thermal Images
by: Ding, Fangqiang, et al.
Published: (2024)
by: Ding, Fangqiang, et al.
Published: (2024)
SemiGDA: Generative Dual-distribution Alignment for Semi-Supervised Medical Image Segmentation
by: Huang, Kaiwen, et al.
Published: (2026)
by: Huang, Kaiwen, et al.
Published: (2026)
Seeing Through Words: Controlling Visual Retrieval Quality with Language Models
by: Lu, Jianglin, et al.
Published: (2026)
by: Lu, Jianglin, et al.
Published: (2026)
Adaptive Confidence Multi-View Hashing for Multimedia Retrieval
by: Zhu, Jian, et al.
Published: (2023)
by: Zhu, Jian, et al.
Published: (2023)
Enhancing Descriptive Image Quality Assessment with A Large-scale Multi-modal Dataset
by: You, Zhiyuan, et al.
Published: (2024)
by: You, Zhiyuan, et al.
Published: (2024)
GFRRN: Explore the Gaps in Single Image Reflection Removal
by: Chen, Yu, et al.
Published: (2026)
by: Chen, Yu, et al.
Published: (2026)
ChatSearch: a Dataset and a Generative Retrieval Model for General Conversational Image Retrieval
by: Zhao, Zijia, et al.
Published: (2024)
by: Zhao, Zijia, et al.
Published: (2024)
Sketch-Guided Scene Image Generation
by: Zhang, Tianyu, et al.
Published: (2024)
by: Zhang, Tianyu, et al.
Published: (2024)
Text-driven Multiplanar Visual Interaction for Semi-supervised Medical Image Segmentation
by: Huang, Kaiwen, et al.
Published: (2025)
by: Huang, Kaiwen, et al.
Published: (2025)
Domain-RAG: Retrieval-Guided Compositional Image Generation for Cross-Domain Few-Shot Object Detection
by: Li, Yu, et al.
Published: (2025)
by: Li, Yu, et al.
Published: (2025)
Towards Context-aware Convolutional Network for Image Restoration
by: Hao, Fangwei, et al.
Published: (2024)
by: Hao, Fangwei, et al.
Published: (2024)
RaTrack: Moving Object Detection and Tracking with 4D Radar Point Cloud
by: Pan, Zhijun, et al.
Published: (2023)
by: Pan, Zhijun, et al.
Published: (2023)
Uncertainty-aware Cross-training for Semi-supervised Medical Image Segmentation
by: Huang, Kaiwen, et al.
Published: (2025)
by: Huang, Kaiwen, et al.
Published: (2025)
Robustness-Guided Image Synthesis for Data-Free Quantization
by: Bai, Jianhong, et al.
Published: (2023)
by: Bai, Jianhong, et al.
Published: (2023)
Hybrid-Learning Video Moment Retrieval across Multi-Domain Labels
by: Cai, Weitong, et al.
Published: (2024)
by: Cai, Weitong, et al.
Published: (2024)
Accelerating Masked Image Generation by Learning Latent Controlled Dynamics
by: Zhu, Kaiwen, et al.
Published: (2026)
by: Zhu, Kaiwen, et al.
Published: (2026)
Learned Image Compression with Dictionary-based Entropy Model
by: Lu, Jingbo, et al.
Published: (2025)
by: Lu, Jingbo, et al.
Published: (2025)
Click to Grasp: Zero-Shot Precise Manipulation via Visual Diffusion Descriptors
by: Tsagkas, Nikolaos, et al.
Published: (2024)
by: Tsagkas, Nikolaos, et al.
Published: (2024)
Image Denoising Using Global and Local Circulant Representation
by: Kong, Zhaoming, et al.
Published: (2025)
by: Kong, Zhaoming, et al.
Published: (2025)
Rethinking Cross-Generator Image Forgery Detection through DINOv3
by: Huang, Zhenglin, et al.
Published: (2025)
by: Huang, Zhenglin, et al.
Published: (2025)
Singpath-VL Technical Report
by: Qiu, Zhen, et al.
Published: (2026)
by: Qiu, Zhen, et al.
Published: (2026)
Identifying and Solving Conditional Image Leakage in Image-to-Video Diffusion Model
by: Zhao, Min, et al.
Published: (2024)
by: Zhao, Min, et al.
Published: (2024)
REST3D: Reconstructing Physically Stable 3D Scenes from a Single Image
by: Ma, Xiaoxuan, et al.
Published: (2026)
by: Ma, Xiaoxuan, et al.
Published: (2026)
Similar Items
-
Self-Adapting Large Visual-Language Models to Edge Devices across Visual Modalities
by: Cai, Kaiwen, et al.
Published: (2024) -
FlexIP: Dynamic Control of Preservation and Personality for Customized Image Generation
by: Huang, Linyan, et al.
Published: (2025) -
RichControl: Structure- and Appearance-Rich Training-Free Spatial Control for Text-to-Image Generation
by: Pang, Lexi, et al.
Published: (2025) -
Robust 3D Object Detection from LiDAR-Radar Point Clouds via Cross-Modal Feature Augmentation
by: Deng, Jianning, et al.
Published: (2023) -
Image Augmentation with Controlled Diffusion for Weakly-Supervised Semantic Segmentation
by: Wu, Wangyu, et al.
Published: (2023)