S^2Former-OR: Single-Stage Bi-Modal Transformer for Scene Graph Generation in OR
Fuente:
arXiv
Saved in:
| Main Authors: | Pei, Jialun, Guo, Diandian, Zhang, Jingyang, Lin, Manxi, Jin, Yueming, Heng, Pheng-Ann |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Tri-modal Confluence with Temporal Dynamics for Scene Graph Generation in Operating Rooms
by: Guo, Diandian, et al.
Published: (2024)
by: Guo, Diandian, et al.
Published: (2024)
Surgical Workflow Recognition and Blocking Effectiveness Detection in Laparoscopic Liver Resections with Pringle Maneuver
by: Guo, Diandian, et al.
Published: (2024)
by: Guo, Diandian, et al.
Published: (2024)
SurgLQA: Scalable Long-Horizon Surgical Video Question Answering
by: Guo, Diandian, et al.
Published: (2026)
by: Guo, Diandian, et al.
Published: (2026)
Synergistic Bleeding Region and Point Detection in Laparoscopic Surgical Videos
by: Pei, Jialun, et al.
Published: (2025)
by: Pei, Jialun, et al.
Published: (2025)
Attenuation-Resilient Alternating Optimization for Laparoscopic Liver Landmark Detection
by: Liu, Lanqing, et al.
Published: (2026)
by: Liu, Lanqing, et al.
Published: (2026)
Benchmarking Endoscopic Surgical Image Restoration and Beyond
by: Pei, Jialun, et al.
Published: (2025)
by: Pei, Jialun, et al.
Published: (2025)
CalibNet: Dual-branch Cross-modal Calibration for RGB-D Salient Instance Segmentation
by: Pei, Jialun, et al.
Published: (2023)
by: Pei, Jialun, et al.
Published: (2023)
Towards Synchronous Memorizability and Generalizability with Site-Modulated Diffusion Replay for Cross-Site Continual Segmentation
by: Xu, Dunyuan, et al.
Published: (2024)
by: Xu, Dunyuan, et al.
Published: (2024)
Depth-Driven Geometric Prompt Learning for Laparoscopic Liver Landmark Detection
by: Pei, Jialun, et al.
Published: (2024)
by: Pei, Jialun, et al.
Published: (2024)
Topology-Constrained Learning for Efficient Laparoscopic Liver Landmark Detection
by: Cui, Ruize, et al.
Published: (2025)
by: Cui, Ruize, et al.
Published: (2025)
Comprehensive Generative Replay for Task-Incremental Segmentation with Concurrent Appearance and Semantic Forgetting
by: Li, Wei, et al.
Published: (2024)
by: Li, Wei, et al.
Published: (2024)
Epicardium Prompt-guided Real-time Cardiac Ultrasound Frame-to-volume Registration
by: Lei, Long, et al.
Published: (2024)
by: Lei, Long, et al.
Published: (2024)
SceneDecorator: Towards Scene-Oriented Story Generation with Scene Planning and Scene Consistency
by: Song, Quanjian, et al.
Published: (2025)
by: Song, Quanjian, et al.
Published: (2025)
Landmark-Free Preoperative-to-Intraoperative Registration in Laparoscopic Liver Resection
by: Zhou, Jun, et al.
Published: (2025)
by: Zhou, Jun, et al.
Published: (2025)
MEJO: MLLM-Engaged Surgical Triplet Recognition via Inter- and Intra-Task Joint Optimization
by: Zhang, Yiyi, et al.
Published: (2025)
by: Zhang, Yiyi, et al.
Published: (2025)
Preoperative-to-intraoperative Liver Registration for Laparoscopic Surgery via Latent-Grounded Correspondence Constraints
by: Cui, Ruize, et al.
Published: (2026)
by: Cui, Ruize, et al.
Published: (2026)
Unifying Physically-Informed Weather Priors in A Single Model for Image Restoration Across Multiple Adverse Weather Conditions
by: Xu, Jiaqi, et al.
Published: (2026)
by: Xu, Jiaqi, et al.
Published: (2026)
SiMA-Hand: Boosting 3D Hand-Mesh Reconstruction by Single-to-Multi-View Adaptation
by: Wang, Yinqiao, et al.
Published: (2024)
by: Wang, Yinqiao, et al.
Published: (2024)
Does Engram Do Memory Retrieval in Autoregressive Image Generation?
by: Wang, Jinghao, et al.
Published: (2026)
by: Wang, Jinghao, et al.
Published: (2026)
Perceive and Calibrate: Analyzing and Enhancing Robustness of Medical Multi-Modal Large Language Models
by: XU, Dunyuan, et al.
Published: (2025)
by: XU, Dunyuan, et al.
Published: (2025)
3DSAM-adapter: Holistic adaptation of SAM from 2D to 3D for promptable tumor segmentation
by: Gong, Shizhan, et al.
Published: (2023)
by: Gong, Shizhan, et al.
Published: (2023)
Rethinking End-to-End 2D to 3D Scene Segmentation in Gaussian Splatting
by: Zhu, Runsong, et al.
Published: (2025)
by: Zhu, Runsong, et al.
Published: (2025)
Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis
by: Yu, Yang, et al.
Published: (2026)
by: Yu, Yang, et al.
Published: (2026)
Adaptive Negative Evidential Deep Learning for Open-set Semi-supervised Learning
by: Yu, Yang, et al.
Published: (2023)
by: Yu, Yang, et al.
Published: (2023)
Improved mmFormer for Liver Fibrosis Staging via Missing-Modality Compensation
by: Zhang, Zhejia, et al.
Published: (2025)
by: Zhang, Zhejia, et al.
Published: (2025)
DeBiFormer: Vision Transformer with Deformable Agent Bi-level Routing Attention
by: Long, Nguyen Huu Bao, et al.
Published: (2024)
by: Long, Nguyen Huu Bao, et al.
Published: (2024)
Deep Omni-supervised Learning for Rib Fracture Detection from Chest Radiology Images
by: Chai, Zhizhong, et al.
Published: (2023)
by: Chai, Zhizhong, et al.
Published: (2023)
Vanishing-Point-Guided Video Semantic Segmentation of Driving Scenes
by: Guo, Diandian, et al.
Published: (2024)
by: Guo, Diandian, et al.
Published: (2024)
ModalFormer: Multimodal Transformer for Low-Light Image Enhancement
by: Brateanu, Alexandru, et al.
Published: (2025)
by: Brateanu, Alexandru, et al.
Published: (2025)
SignVTCL: Multi-Modal Continuous Sign Language Recognition Enhanced by Visual-Textual Contrastive Learning
by: Chen, Hao, et al.
Published: (2024)
by: Chen, Hao, et al.
Published: (2024)
Unlocking Positive Transfer in Incrementally Learning Surgical Instruments: A Self-reflection Hierarchical Prompt Framework
by: Zhu, Yu, et al.
Published: (2026)
by: Zhu, Yu, et al.
Published: (2026)
MM-Mixing: Multi-Modal Mixing Alignment for 3D Understanding
by: Wang, Jiaze, et al.
Published: (2024)
by: Wang, Jiaze, et al.
Published: (2024)
Generalized Unbiased Scene Graph Generation
by: Lyu, Xinyu, et al.
Published: (2023)
by: Lyu, Xinyu, et al.
Published: (2023)
Informative Scene Graph Generation via Debiasing
by: Gao, Lianli, et al.
Published: (2023)
by: Gao, Lianli, et al.
Published: (2023)
Spatio-Temporal Representation Decoupling and Enhancement for Federated Instrument Segmentation in Surgical Videos
by: Fang, Zheng, et al.
Published: (2025)
by: Fang, Zheng, et al.
Published: (2025)
UniHOPE: A Unified Approach for Hand-Only and Hand-Object Pose Estimation
by: Wang, Yinqiao, et al.
Published: (2025)
by: Wang, Yinqiao, et al.
Published: (2025)
Evaluation Study on SAM 2 for Class-agnostic Instance-level Segmentation
by: Pei, Jialun, et al.
Published: (2024)
by: Pei, Jialun, et al.
Published: (2024)
ATLAS: Agentic or Latent Visual Reasoning? One Word is Enough for Both
by: Guo, Ziyu, et al.
Published: (2026)
by: Guo, Ziyu, et al.
Published: (2026)
Rethinking Detecting Salient and Camouflaged Objects in Unconstrained Scenes
by: Zhou, Zhangjun, et al.
Published: (2024)
by: Zhou, Zhangjun, et al.
Published: (2024)
Decoupling Feature Representations of Ego and Other Modalities for Incomplete Multi-modal Brain Tumor Segmentation
by: Yang, Kaixiang, et al.
Published: (2024)
by: Yang, Kaixiang, et al.
Published: (2024)
Similar Items
-
Tri-modal Confluence with Temporal Dynamics for Scene Graph Generation in Operating Rooms
by: Guo, Diandian, et al.
Published: (2024) -
Surgical Workflow Recognition and Blocking Effectiveness Detection in Laparoscopic Liver Resections with Pringle Maneuver
by: Guo, Diandian, et al.
Published: (2024) -
SurgLQA: Scalable Long-Horizon Surgical Video Question Answering
by: Guo, Diandian, et al.
Published: (2026) -
Synergistic Bleeding Region and Point Detection in Laparoscopic Surgical Videos
by: Pei, Jialun, et al.
Published: (2025) -
Attenuation-Resilient Alternating Optimization for Laparoscopic Liver Landmark Detection
by: Liu, Lanqing, et al.
Published: (2026)