Surgical Triplet Recognition via Diffusion Model
Fuente:
arXiv
Saved in:
| Main Authors: | Liu, Daochang, Hu, Axel, Shah, Mubarak, Xu, Chang |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Exploring Local Memorization in Diffusion Models via Bright Ending Attention
by: Chen, Chen, et al.
Published: (2024)
by: Chen, Chen, et al.
Published: (2024)
Enhancing Privacy-Utility Trade-offs to Mitigate Memorization in Diffusion Models
by: Chen, Chen, et al.
Published: (2025)
by: Chen, Chen, et al.
Published: (2025)
Investigating Memorization in Video Diffusion Models
by: Chen, Chen, et al.
Published: (2024)
by: Chen, Chen, et al.
Published: (2024)
CollagePrompt: A Benchmark for Budget-Friendly Visual Recognition with GPT-4V
by: Xu, Siyu, et al.
Published: (2024)
by: Xu, Siyu, et al.
Published: (2024)
Generative Physical AI in Vision: A Survey
by: Liu, Daochang, et al.
Published: (2025)
by: Liu, Daochang, et al.
Published: (2025)
Diffexplainer: Towards Cross-modal Global Explanations with Diffusion Models
by: Pennisi, Matteo, et al.
Published: (2024)
by: Pennisi, Matteo, et al.
Published: (2024)
GVD: Guiding Video Diffusion Model for Scalable Video Distillation
by: Li, Kunyang, et al.
Published: (2025)
by: Li, Kunyang, et al.
Published: (2025)
Towards Memorization-Free Diffusion Models
by: Chen, Chen, et al.
Published: (2024)
by: Chen, Chen, et al.
Published: (2024)
Diffusion Models in Vision: A Survey
by: Croitoru, Florinel-Alin, et al.
Published: (2022)
by: Croitoru, Florinel-Alin, et al.
Published: (2022)
Draft-and-Target Sampling for Video Generation Policy
by: Zhang, Qikang, et al.
Published: (2026)
by: Zhang, Qikang, et al.
Published: (2026)
Weakly-Supervised Spatiotemporal Anomaly Detection
by: Gianchandani, Urvi, et al.
Published: (2026)
by: Gianchandani, Urvi, et al.
Published: (2026)
Not All Tokens are Guided Equal: Improving Guidance in Visual Autoregressive Models
by: Nguyen, Ky Dan, et al.
Published: (2025)
by: Nguyen, Ky Dan, et al.
Published: (2025)
Multi-Level Heterogeneous Knowledge Transfer Network on Forward Scattering Center Model for Limited Samples SAR ATR
by: Zhao, Chenxi, et al.
Published: (2025)
by: Zhao, Chenxi, et al.
Published: (2025)
Surgical-LLaVA: Toward Surgical Scenario Understanding via Large Language and Vision Models
by: Jin, Juseong, et al.
Published: (2024)
by: Jin, Juseong, et al.
Published: (2024)
Curriculum Direct Preference Optimization for Diffusion and Consistency Models
by: Croitoru, Florinel-Alin, et al.
Published: (2024)
by: Croitoru, Florinel-Alin, et al.
Published: (2024)
Robin3D: Improving 3D Large Language Model via Robust Instruction Tuning
by: Kang, Weitai, et al.
Published: (2024)
by: Kang, Weitai, et al.
Published: (2024)
Safety Without Semantic Disruptions: Editing-free Safe Image Generation via Context-preserving Dual Latent Reconstruction
by: Vice, Jordan, et al.
Published: (2024)
by: Vice, Jordan, et al.
Published: (2024)
USAD: End-to-End Human Activity Recognition via Diffusion Model with Spatiotemporal Attention
by: Xiao, Hang, et al.
Published: (2025)
by: Xiao, Hang, et al.
Published: (2025)
Stabilizing Temporal Inference Dynamics for Online Surgical Phase Recognition
by: Liu, Yang, et al.
Published: (2026)
by: Liu, Yang, et al.
Published: (2026)
Text-guided 3D Human Motion Generation with Keyframe-based Parallel Skip Transformer
by: Geng, Zichen, et al.
Published: (2024)
by: Geng, Zichen, et al.
Published: (2024)
LoViT: Long Video Transformer for Surgical Phase Recognition
by: Liu, Yang, et al.
Published: (2023)
by: Liu, Yang, et al.
Published: (2023)
Holistic Surgical Phase Recognition with Hierarchical Input Dependent State Space Models
by: Wu, Haoyang, et al.
Published: (2025)
by: Wu, Haoyang, et al.
Published: (2025)
Curriculum-DPO++: Direct Preference Optimization via Data and Model Curricula for Text-to-Image Generation
by: Croitoru, Florinel-Alin, et al.
Published: (2026)
by: Croitoru, Florinel-Alin, et al.
Published: (2026)
MuST: Multi-Scale Transformers for Surgical Phase Recognition
by: Pérez, Alejandra, et al.
Published: (2024)
by: Pérez, Alejandra, et al.
Published: (2024)
Image Synthesis with Class-Aware Semantic Diffusion Models for Surgical Scene Segmentation
by: Zhou, Yihang, et al.
Published: (2024)
by: Zhou, Yihang, et al.
Published: (2024)
Repairing Catastrophic-Neglect in Text-to-Image Diffusion Models via Attention-Guided Feature Enhancement
by: Chang, Zhiyuan, et al.
Published: (2024)
by: Chang, Zhiyuan, et al.
Published: (2024)
Any-to-Any Learning in Computational Pathology via Triplet Multimodal Pretraining
by: Sun, Qichen, et al.
Published: (2025)
by: Sun, Qichen, et al.
Published: (2025)
Latent-based Diffusion Model for Long-tailed Recognition
by: Han, Pengxiao, et al.
Published: (2024)
by: Han, Pengxiao, et al.
Published: (2024)
DEXTER: Diffusion-Guided EXplanations with TExtual Reasoning for Vision Models
by: Carnemolla, Simone, et al.
Published: (2025)
by: Carnemolla, Simone, et al.
Published: (2025)
DSTED: Decoupling Temporal Stabilization and Discriminative Enhancement for Surgical Workflow Recognition
by: Chen, Yueyao, et al.
Published: (2025)
by: Chen, Yueyao, et al.
Published: (2025)
TrackDiffusion: Tracklet-Conditioned Video Generation via Diffusion Models
by: Li, Pengxiang, et al.
Published: (2023)
by: Li, Pengxiang, et al.
Published: (2023)
You Don't Need All That Attention: Surgical Memorization Mitigation in Text-to-Image Diffusion Models
by: Zhao, Kairan, et al.
Published: (2026)
by: Zhao, Kairan, et al.
Published: (2026)
Compress Guidance in Conditional Diffusion Sampling
by: Dinh, Anh-Dung, et al.
Published: (2024)
by: Dinh, Anh-Dung, et al.
Published: (2024)
SafeR-CLIP: Mitigating NSFW Content in Vision-Language Models While Preserving Pre-Trained Knowledge
by: Yousaf, Adeel, et al.
Published: (2025)
by: Yousaf, Adeel, et al.
Published: (2025)
HecVL: Hierarchical Video-Language Pretraining for Zero-shot Surgical Phase Recognition
by: Yuan, Kun, et al.
Published: (2024)
by: Yuan, Kun, et al.
Published: (2024)
End to End AI System for Surgical Gesture Sequence Recognition and Clinical Outcome Prediction
by: Li, Xi, et al.
Published: (2025)
by: Li, Xi, et al.
Published: (2025)
Meta-Entity Driven Triplet Mining for Aligning Medical Vision-Language Models
by: Ozturk, Saban, et al.
Published: (2025)
by: Ozturk, Saban, et al.
Published: (2025)
Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models
by: Liu, Huijie, et al.
Published: (2025)
by: Liu, Huijie, et al.
Published: (2025)
HumaniBench: A Human-Centric Framework for Large Multimodal Models Evaluation
by: Raza, Shaina, et al.
Published: (2025)
by: Raza, Shaina, et al.
Published: (2025)
Beyond Dataset Distillation: Lossless Dataset Concentration via Diffusion-Assisted Distribution Alignment
by: Liu, Tongfei, et al.
Published: (2026)
by: Liu, Tongfei, et al.
Published: (2026)
Similar Items
-
Exploring Local Memorization in Diffusion Models via Bright Ending Attention
by: Chen, Chen, et al.
Published: (2024) -
Enhancing Privacy-Utility Trade-offs to Mitigate Memorization in Diffusion Models
by: Chen, Chen, et al.
Published: (2025) -
Investigating Memorization in Video Diffusion Models
by: Chen, Chen, et al.
Published: (2024) -
CollagePrompt: A Benchmark for Budget-Friendly Visual Recognition with GPT-4V
by: Xu, Siyu, et al.
Published: (2024) -
Generative Physical AI in Vision: A Survey
by: Liu, Daochang, et al.
Published: (2025)