CoSimGen: Controllable Diffusion Model for Simultaneous Image and Mask Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Bose, Rupak, Nwoye, Chinedu Innocent, Bhat, Aditya, Padoy, Nicolas |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SimGen: A Diffusion-Based Framework for Simultaneous Surgical Image and Segmentation Mask Generation
by: Bhat, Aditya, et al.
Published: (2025)
by: Bhat, Aditya, et al.
Published: (2025)
SurgiTrack: Fine-Grained Multi-Class Multi-Tool Tracking in Surgical Videos
by: Nwoye, Chinedu Innocent, et al.
Published: (2024)
by: Nwoye, Chinedu Innocent, et al.
Published: (2024)
Feature Mixing Approach for Detecting Intraoperative Adverse Events in Laparoscopic Roux-en-Y Gastric Bypass Surgery
by: Bose, Rupak, et al.
Published: (2025)
by: Bose, Rupak, et al.
Published: (2025)
Surgical Text-to-Image Generation
by: Nwoye, Chinedu Innocent, et al.
Published: (2024)
by: Nwoye, Chinedu Innocent, et al.
Published: (2024)
State-Change Learning for Prediction of Future Events in Endoscopic Videos
by: Sharma, Saurav, et al.
Published: (2025)
by: Sharma, Saurav, et al.
Published: (2025)
CholecTrack20: A Multi-Perspective Tracking Dataset for Surgical Tools
by: Nwoye, Chinedu Innocent, et al.
Published: (2023)
by: Nwoye, Chinedu Innocent, et al.
Published: (2023)
SkinDualGen: Prompt-Driven Diffusion for Simultaneous Image-Mask Generation in Skin Lesions
by: Xu, Zhaobin
Published: (2025)
by: Xu, Zhaobin
Published: (2025)
SurgLaVi: Large-Scale Hierarchical Dataset for Surgical Vision-Language Representation Learning
by: Perez, Alejandra, et al.
Published: (2025)
by: Perez, Alejandra, et al.
Published: (2025)
DExTeR: Weakly Semi-Supervised Object Detection with Class and Instance Experts for Medical Imaging
by: Meyer, Adrien, et al.
Published: (2026)
by: Meyer, Adrien, et al.
Published: (2026)
HiCoGen: Hierarchical Compositional Text-to-Image Generation in Diffusion Models via Reinforcement Learning
by: Yang, Hongji, et al.
Published: (2025)
by: Yang, Hongji, et al.
Published: (2025)
fine-CLIP: Enhancing Zero-Shot Fine-Grained Surgical Action Recognition with Vision-Language Models
by: Sharma, Saurav, et al.
Published: (2025)
by: Sharma, Saurav, et al.
Published: (2025)
SimGen: Simulator-conditioned Driving Scene Generation
by: Zhou, Yunsong, et al.
Published: (2024)
by: Zhou, Yunsong, et al.
Published: (2024)
DiffAtlas: GenAI-fying Atlas Segmentation via Image-Mask Diffusion
by: Zhang, Hantao, et al.
Published: (2025)
by: Zhang, Hantao, et al.
Published: (2025)
On the Role of Depth in Surgical Vision Foundation Models: An Empirical Study of RGB-D Pre-training
by: Han, John J., et al.
Published: (2026)
by: Han, John J., et al.
Published: (2026)
Overcoming Dimensional Collapse in Self-supervised Contrastive Learning for Medical Image Segmentation
by: Hassanpour, Jamshid, et al.
Published: (2024)
by: Hassanpour, Jamshid, et al.
Published: (2024)
PiCo: Enhancing Text-Image Alignment with Improved Noise Selection and Precise Mask Control in Diffusion Models
by: Xie, Chang, et al.
Published: (2025)
by: Xie, Chang, et al.
Published: (2025)
EmoGen: Emotional Image Content Generation with Text-to-Image Diffusion Models
by: Yang, Jingyuan, et al.
Published: (2024)
by: Yang, Jingyuan, et al.
Published: (2024)
End-to-End Learning of Multi-Organ Implicit Surfaces from 3D Medical Imaging Data
by: Zarin, Farahdiba, et al.
Published: (2025)
by: Zarin, Farahdiba, et al.
Published: (2025)
GenMask: Adapting DiT for Segmentation via Direct Mask Generation
by: Yang, Yuhuan, et al.
Published: (2026)
by: Yang, Yuhuan, et al.
Published: (2026)
SUREON: A Benchmark and Vision-Language-Model for Surgical Reasoning
by: Perez, Alejandra, et al.
Published: (2026)
by: Perez, Alejandra, et al.
Published: (2026)
Accelerating Inference of Masked Image Generators via Reinforcement Learning
by: Subbaraman, Pranav, et al.
Published: (2025)
by: Subbaraman, Pranav, et al.
Published: (2025)
GenTron: Diffusion Transformers for Image and Video Generation
by: Chen, Shoufa, et al.
Published: (2023)
by: Chen, Shoufa, et al.
Published: (2023)
SelfPose3d: Self-Supervised Multi-Person Multi-View 3d Pose Estimation
by: Srivastav, Vinkle, et al.
Published: (2024)
by: Srivastav, Vinkle, et al.
Published: (2024)
Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation
by: Li, Shufan, et al.
Published: (2025)
by: Li, Shufan, et al.
Published: (2025)
CycleSAM: Few-Shot Surgical Scene Segmentation with Cycle- and Scene-Consistent Feature Matching
by: Murali, Aditya, et al.
Published: (2024)
by: Murali, Aditya, et al.
Published: (2024)
Optimizing Latent Graph Representations of Surgical Scenes for Zero-Shot Domain Transfer
by: Satyanaik, Siddhant, et al.
Published: (2024)
by: Satyanaik, Siddhant, et al.
Published: (2024)
Harnessing Diffusion-Generated Synthetic Images for Fair Image Classification
by: Basu, Abhipsa, et al.
Published: (2025)
by: Basu, Abhipsa, et al.
Published: (2025)
UltraSam: A Foundation Model for Ultrasound using Large Open-Access Segmentation Datasets
by: Meyer, Adrien, et al.
Published: (2024)
by: Meyer, Adrien, et al.
Published: (2024)
GenMix: Effective Data Augmentation with Generative Diffusion Model Image Editing
by: Islam, Khawar, et al.
Published: (2024)
by: Islam, Khawar, et al.
Published: (2024)
EpiMask: Leveraging Epipolar Distance Based Masks in Cross-Attention for Satellite Image Matching
by: Deshmukh, Rahul, et al.
Published: (2026)
by: Deshmukh, Rahul, et al.
Published: (2026)
SSEditor: Controllable Mask-to-Scene Generation with Diffusion Model
by: Zheng, Haowen, et al.
Published: (2024)
by: Zheng, Haowen, et al.
Published: (2024)
Mask-ControlNet: Higher-Quality Image Generation with An Additional Mask Prompt
by: Huang, Zhiqi, et al.
Published: (2024)
by: Huang, Zhiqi, et al.
Published: (2024)
CoMoGen: COntrollable MOtion Dynamics and Interactions with Mask-Guided Video GENeration
by: Meric, Adil, et al.
Published: (2026)
by: Meric, Adil, et al.
Published: (2026)
RadGazeGen: Radiomics and Gaze-guided Medical Image Generation using Diffusion Models
by: Bhattacharya, Moinak, et al.
Published: (2024)
by: Bhattacharya, Moinak, et al.
Published: (2024)
UniGenDet: A Unified Generative-Discriminative Framework for Co-Evolutionary Image Generation and Generated Image Detection
by: Zhang, Yanran, et al.
Published: (2026)
by: Zhang, Yanran, et al.
Published: (2026)
Reinforcement Learning Meets Masked Generative Models: Mask-GRPO for Text-to-Image Generation
by: Luo, Yifu, et al.
Published: (2025)
by: Luo, Yifu, et al.
Published: (2025)
Learning from Synchronization: Self-Supervised Uncalibrated Multi-View Person Association in Challenging Scenes
by: Chen, Keqi, et al.
Published: (2025)
by: Chen, Keqi, et al.
Published: (2025)
Simultaneous Image-to-Zero and Zero-to-Noise: Diffusion Models with Analytical Image Attenuation
by: Huang, Yuhang, et al.
Published: (2023)
by: Huang, Yuhang, et al.
Published: (2023)
GenRec: Unifying Video Generation and Recognition with Diffusion Models
by: Weng, Zejia, et al.
Published: (2024)
by: Weng, Zejia, et al.
Published: (2024)
StarGen: A Spatiotemporal Autoregression Framework with Video Diffusion Model for Scalable and Controllable Scene Generation
by: Zhai, Shangjin, et al.
Published: (2025)
by: Zhai, Shangjin, et al.
Published: (2025)
Similar Items
-
SimGen: A Diffusion-Based Framework for Simultaneous Surgical Image and Segmentation Mask Generation
by: Bhat, Aditya, et al.
Published: (2025) -
SurgiTrack: Fine-Grained Multi-Class Multi-Tool Tracking in Surgical Videos
by: Nwoye, Chinedu Innocent, et al.
Published: (2024) -
Feature Mixing Approach for Detecting Intraoperative Adverse Events in Laparoscopic Roux-en-Y Gastric Bypass Surgery
by: Bose, Rupak, et al.
Published: (2025) -
Surgical Text-to-Image Generation
by: Nwoye, Chinedu Innocent, et al.
Published: (2024) -
State-Change Learning for Prediction of Future Events in Endoscopic Videos
by: Sharma, Saurav, et al.
Published: (2025)