MEJO: MLLM-Engaged Surgical Triplet Recognition via Inter- and Intra-Task Joint Optimization
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Yiyi, Yuan, Yuchen, Zheng, Ying, Pei, Jialun, Li, Jinpeng, Li, Zheng, Heng, Pheng-Ann |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Surgical Workflow Recognition and Blocking Effectiveness Detection in Laparoscopic Liver Resections with Pringle Maneuver
by: Guo, Diandian, et al.
Published: (2024)
by: Guo, Diandian, et al.
Published: (2024)
SurgLQA: Scalable Long-Horizon Surgical Video Question Answering
by: Guo, Diandian, et al.
Published: (2026)
by: Guo, Diandian, et al.
Published: (2026)
Synergistic Bleeding Region and Point Detection in Laparoscopic Surgical Videos
by: Pei, Jialun, et al.
Published: (2025)
by: Pei, Jialun, et al.
Published: (2025)
Unlocking Positive Transfer in Incrementally Learning Surgical Instruments: A Self-reflection Hierarchical Prompt Framework
by: Zhu, Yu, et al.
Published: (2026)
by: Zhu, Yu, et al.
Published: (2026)
Benchmarking Endoscopic Surgical Image Restoration and Beyond
by: Pei, Jialun, et al.
Published: (2025)
by: Pei, Jialun, et al.
Published: (2025)
Depth-Driven Geometric Prompt Learning for Laparoscopic Liver Landmark Detection
by: Pei, Jialun, et al.
Published: (2024)
by: Pei, Jialun, et al.
Published: (2024)
Medical Large Vision Language Models with Multi-Image Visual Ability
by: Yang, Xikai, et al.
Published: (2025)
by: Yang, Xikai, et al.
Published: (2025)
From Learning to Unlearning: Biomedical Security Protection in Multimodal Large Language Models
by: Xu, Dunyuan, et al.
Published: (2025)
by: Xu, Dunyuan, et al.
Published: (2025)
Attenuation-Resilient Alternating Optimization for Laparoscopic Liver Landmark Detection
by: Liu, Lanqing, et al.
Published: (2026)
by: Liu, Lanqing, et al.
Published: (2026)
S^2Former-OR: Single-Stage Bi-Modal Transformer for Scene Graph Generation in OR
by: Pei, Jialun, et al.
Published: (2024)
by: Pei, Jialun, et al.
Published: (2024)
Topology-Constrained Learning for Efficient Laparoscopic Liver Landmark Detection
by: Cui, Ruize, et al.
Published: (2025)
by: Cui, Ruize, et al.
Published: (2025)
Tri-modal Confluence with Temporal Dynamics for Scene Graph Generation in Operating Rooms
by: Guo, Diandian, et al.
Published: (2024)
by: Guo, Diandian, et al.
Published: (2024)
Landmark-Free Preoperative-to-Intraoperative Registration in Laparoscopic Liver Resection
by: Zhou, Jun, et al.
Published: (2025)
by: Zhou, Jun, et al.
Published: (2025)
Perceive and Calibrate: Analyzing and Enhancing Robustness of Medical Multi-Modal Large Language Models
by: XU, Dunyuan, et al.
Published: (2025)
by: XU, Dunyuan, et al.
Published: (2025)
Med-Evo: Test-time Self-evolution for Medical Multimodal Large Language Models
by: Xu, Dunyuan, et al.
Published: (2026)
by: Xu, Dunyuan, et al.
Published: (2026)
Preoperative-to-intraoperative Liver Registration for Laparoscopic Surgery via Latent-Grounded Correspondence Constraints
by: Cui, Ruize, et al.
Published: (2026)
by: Cui, Ruize, et al.
Published: (2026)
Adapting 2D Multi-Modal Large Language Model for 3D CT Image Analysis
by: Yu, Yang, et al.
Published: (2026)
by: Yu, Yang, et al.
Published: (2026)
Surgical Triplet Recognition via Diffusion Model
by: Liu, Daochang, et al.
Published: (2024)
by: Liu, Daochang, et al.
Published: (2024)
Decomposing and Fusing Intra- and Inter-Sensor Spatio-Temporal Signal for Multi-Sensor Wearable Human Activity Recognition
by: Xie, Haoyu, et al.
Published: (2025)
by: Xie, Haoyu, et al.
Published: (2025)
Comprehensive Generative Replay for Task-Incremental Segmentation with Concurrent Appearance and Semantic Forgetting
by: Li, Wei, et al.
Published: (2024)
by: Li, Wei, et al.
Published: (2024)
MedHallTune: An Instruction-Tuning Benchmark for Mitigating Medical Hallucination in Vision-Language Models
by: Yan, Qiao, et al.
Published: (2025)
by: Yan, Qiao, et al.
Published: (2025)
CalibNet: Dual-branch Cross-modal Calibration for RGB-D Salient Instance Segmentation
by: Pei, Jialun, et al.
Published: (2023)
by: Pei, Jialun, et al.
Published: (2023)
SignVTCL: Multi-Modal Continuous Sign Language Recognition Enhanced by Visual-Textual Contrastive Learning
by: Chen, Hao, et al.
Published: (2024)
by: Chen, Hao, et al.
Published: (2024)
Rethinking Facial Expression Recognition in the Era of Multimodal Large Language Models: Benchmark, Datasets, and Beyond
by: Zhang, Fan, et al.
Published: (2025)
by: Zhang, Fan, et al.
Published: (2025)
3DSAM-adapter: Holistic adaptation of SAM from 2D to 3D for promptable tumor segmentation
by: Gong, Shizhan, et al.
Published: (2023)
by: Gong, Shizhan, et al.
Published: (2023)
Multi-scale Spatio-temporal Transformer-based Imbalanced Longitudinal Learning for Glaucoma Forecasting from Irregular Time Series Images
by: Yang, Xikai, et al.
Published: (2024)
by: Yang, Xikai, et al.
Published: (2024)
Dose Prediction Driven Radiotherapy Paramters Regression via Intra- and Inter-Relation Modeling
by: Cui, Jiaqi, et al.
Published: (2024)
by: Cui, Jiaqi, et al.
Published: (2024)
Toward Reliable AR-Guided Surgical Navigation: Interactive Deformation Modeling with Data-Driven Biomechanics and Prompts
by: Han, Zheng, et al.
Published: (2025)
by: Han, Zheng, et al.
Published: (2025)
Spatio-Temporal Representation Decoupling and Enhancement for Federated Instrument Segmentation in Surgical Videos
by: Fang, Zheng, et al.
Published: (2025)
by: Fang, Zheng, et al.
Published: (2025)
VersaViT: Enhancing MLLM Vision Backbones via Task-Guided Optimization
by: Liu, Yikun, et al.
Published: (2026)
by: Liu, Yikun, et al.
Published: (2026)
Fuzzy-aware Loss for Source-free Domain Adaptation in Visual Emotion Recognition
by: Zheng, Ying, et al.
Published: (2025)
by: Zheng, Ying, et al.
Published: (2025)
Epicardium Prompt-guided Real-time Cardiac Ultrasound Frame-to-volume Registration
by: Lei, Long, et al.
Published: (2024)
by: Lei, Long, et al.
Published: (2024)
Memory-Efficient Prompt Tuning for Incremental Histopathology Classification
by: Zhu, Yu, et al.
Published: (2024)
by: Zhu, Yu, et al.
Published: (2024)
The Coherence Trap: When MLLM-Crafted Narratives Exploit Manipulated Visual Contexts
by: Zhang, Yuchen, et al.
Published: (2025)
by: Zhang, Yuchen, et al.
Published: (2025)
Evolution-Inspired Sample Competition for Deep Neural Network Optimization
by: Zheng, Ying, et al.
Published: (2026)
by: Zheng, Ying, et al.
Published: (2026)
Jointly Modeling Inter- & Intra-Modality Dependencies for Multi-modal Learning
by: Madaan, Divyam, et al.
Published: (2024)
by: Madaan, Divyam, et al.
Published: (2024)
SurgPub-Video: A Comprehensive Surgical Video Dataset for Enhanced Surgical Intelligence in Vision-Language Model
by: Li, Yaoqian, et al.
Published: (2025)
by: Li, Yaoqian, et al.
Published: (2025)
IIP-Transformer: Intra-Inter-Part Transformer for Skeleton-Based Action Recognition
by: Wang, Qingtian, et al.
Published: (2021)
by: Wang, Qingtian, et al.
Published: (2021)
Detail-Enhanced Intra- and Inter-modal Interaction for Audio-Visual Emotion Recognition
by: Shi, Tong, et al.
Published: (2024)
by: Shi, Tong, et al.
Published: (2024)
SuPRA: Surgical Phase Recognition and Anticipation for Intra-Operative Planning
by: Boels, Maxence, et al.
Published: (2024)
by: Boels, Maxence, et al.
Published: (2024)
Similar Items
-
Surgical Workflow Recognition and Blocking Effectiveness Detection in Laparoscopic Liver Resections with Pringle Maneuver
by: Guo, Diandian, et al.
Published: (2024) -
SurgLQA: Scalable Long-Horizon Surgical Video Question Answering
by: Guo, Diandian, et al.
Published: (2026) -
Synergistic Bleeding Region and Point Detection in Laparoscopic Surgical Videos
by: Pei, Jialun, et al.
Published: (2025) -
Unlocking Positive Transfer in Incrementally Learning Surgical Instruments: A Self-reflection Hierarchical Prompt Framework
by: Zhu, Yu, et al.
Published: (2026) -
Benchmarking Endoscopic Surgical Image Restoration and Beyond
by: Pei, Jialun, et al.
Published: (2025)