Saved in:
| Main Authors: | Liu, Chen, Wu, Haitao, Wang, Kafeng, Huang, Weiran |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2509.16702 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DisCo3D: Distilling Multi-View Consistency for 3D Scene Editing
by: Chi, Yufeng, et al.
Published: (2025)
by: Chi, Yufeng, et al.
Published: (2025)
RCR-AF: Enhancing Model Generalization via Rademacher Complexity Reduction Activation Function
by: Yu, Yunrui, et al.
Published: (2025)
by: Yu, Yunrui, et al.
Published: (2025)
PersonificationNet: Making customized subject act like a person
by: Guo, Tianchu, et al.
Published: (2024)
by: Guo, Tianchu, et al.
Published: (2024)
Robust and Efficient Adversarial Defense in SNNs via Image Purification and Joint Detection
by: Chen, Weiran, et al.
Published: (2024)
by: Chen, Weiran, et al.
Published: (2024)
MIMO: A medical vision language model with visual referring multimodal input and pixel grounding multimodal output
by: Chen, Yanyuan, et al.
Published: (2025)
by: Chen, Yanyuan, et al.
Published: (2025)
EZIGen: Enhancing zero-shot personalized image generation with precise subject encoding and decoupled guidance
by: Duan, Zicheng, et al.
Published: (2024)
by: Duan, Zicheng, et al.
Published: (2024)
Quasi-multimodal-based pathophysiological feature learning for retinal disease diagnosis
by: Zhang, Lu, et al.
Published: (2026)
by: Zhang, Lu, et al.
Published: (2026)
AutoEval-Video: An Automatic Benchmark for Assessing Large Vision Language Models in Open-Ended Video Question Answering
by: Chen, Xiuyuan, et al.
Published: (2023)
by: Chen, Xiuyuan, et al.
Published: (2023)
On the robustness of multimodal language model towards distractions
by: Liu, Ming, et al.
Published: (2025)
by: Liu, Ming, et al.
Published: (2025)
OTMatch: Improving Semi-Supervised Learning with Optimal Transport
by: Tan, Zhiquan, et al.
Published: (2023)
by: Tan, Zhiquan, et al.
Published: (2023)
TransMed: Large Language Models Enhance Vision Transformer for Biomedical Image Classification
by: Zheng, Kaipeng, et al.
Published: (2023)
by: Zheng, Kaipeng, et al.
Published: (2023)
Bias-constrained multimodal intelligence for equitable and reliable clinical AI
by: Li, Cheng, et al.
Published: (2026)
by: Li, Cheng, et al.
Published: (2026)
AI-powered multimodal modeling of personalized hemodynamics in aortic stenosis
by: Ozturk, Caglar, et al.
Published: (2024)
by: Ozturk, Caglar, et al.
Published: (2024)
Joint-octamamba:an octa joint segmentation network based on feature enhanced mamba
by: Liu, Chuang, et al.
Published: (2025)
by: Liu, Chuang, et al.
Published: (2025)
Buffer replay enhances the robustness of multimodal learning under missing-modality
by: Zhu, Hongye, et al.
Published: (2025)
by: Zhu, Hongye, et al.
Published: (2025)
Robust feature knowledge distillation for enhanced performance of lightweight crack segmentation models
by: Chen, Zhaohui, et al.
Published: (2024)
by: Chen, Zhaohui, et al.
Published: (2024)
Scarecrow monitoring system:employing mobilenet ssd for enhanced animal supervision
by: VS, Balaji, et al.
Published: (2024)
by: VS, Balaji, et al.
Published: (2024)
SpectraFlow: Unifying Structural Pretraining and Frequency Adaptation for Medical Image Segmentation
by: Chen, Zhiquan, et al.
Published: (2026)
by: Chen, Zhiquan, et al.
Published: (2026)
Med-DisSeg: Dispersion-Driven Representation Learning for Fine-Grained Medical Image Segmentation
by: Chen, Zhiquan, et al.
Published: (2026)
by: Chen, Zhiquan, et al.
Published: (2026)
Supervised makeup transfer with a curated dataset: Decoupling identity and makeup features for enhanced transformation
by: Pan, Qihe, et al.
Published: (2026)
by: Pan, Qihe, et al.
Published: (2026)
DA-Font: Few-Shot Font Generation via Dual-Attention Hybrid Integration
by: Chen, Weiran, et al.
Published: (2025)
by: Chen, Weiran, et al.
Published: (2025)
DM-FNet: Unified multimodal medical image fusion via diffusion process-trained encoder-decoder
by: He, Dan, et al.
Published: (2025)
by: He, Dan, et al.
Published: (2025)
RhythmFormer: Extracting Patterned rPPG Signals based on Periodic Sparse Attention
by: Zou, Bochao, et al.
Published: (2024)
by: Zou, Bochao, et al.
Published: (2024)
Generalized Category Discovery via Reciprocal Learning and Class-Wise Distribution Regularization
by: Liu, Duo, et al.
Published: (2025)
by: Liu, Duo, et al.
Published: (2025)
From Image to Video, what do we need in multimodal LLMs?
by: Huang, Suyuan, et al.
Published: (2024)
by: Huang, Suyuan, et al.
Published: (2024)
Information Flow in Self-Supervised Learning
by: Tan, Zhiquan, et al.
Published: (2023)
by: Tan, Zhiquan, et al.
Published: (2023)
SHAP-CAT: A interpretable multi-modal framework enhancing WSI classification via virtual staining and shapley-value-based multimodal fusion
by: Wang, Jun, et al.
Published: (2024)
by: Wang, Jun, et al.
Published: (2024)
Attacks on multimodal models
by: Iablochnikov, Viacheslav, et al.
Published: (2024)
by: Iablochnikov, Viacheslav, et al.
Published: (2024)
MS-Diffusion: Multi-subject Zero-shot Image Personalization with Layout Guidance
by: Wang, Xierui, et al.
Published: (2024)
by: Wang, Xierui, et al.
Published: (2024)
Attend to Evidence: Evidence-Anchored Spatial Attention Supervision for Multimodal RLVR
by: Hu, Ruina, et al.
Published: (2026)
by: Hu, Ruina, et al.
Published: (2026)
MimicTalk: Mimicking a personalized and expressive 3D talking face in minutes
by: Ye, Zhenhui, et al.
Published: (2024)
by: Ye, Zhenhui, et al.
Published: (2024)
A multimodal vision foundation model for generalizable knee pathology
by: Yu, Kang, et al.
Published: (2026)
by: Yu, Kang, et al.
Published: (2026)
*: Improving the 3D detector by introducing Voxel2Pillar feature encoding and extracting multi-scale features
by: Li, Xusheng, et al.
Published: (2024)
by: Li, Xusheng, et al.
Published: (2024)
SAMIR, an efficient registration framework via robust feature learning from SAM
by: He, Yue, et al.
Published: (2025)
by: He, Yue, et al.
Published: (2025)
MiVE: Multiscale Vision-language features for reference-guided video Editing
by: Wang, Tong, et al.
Published: (2026)
by: Wang, Tong, et al.
Published: (2026)
TinyGPT-V: Efficient Multimodal Large Language Model via Small Backbones
by: Yuan, Zhengqing, et al.
Published: (2023)
by: Yuan, Zhengqing, et al.
Published: (2023)
Improved Iterative Refinement for Chart-to-Code Generation via Structured Instruction
by: Xu, Chengzhi, et al.
Published: (2025)
by: Xu, Chengzhi, et al.
Published: (2025)
CAD-feature enhanced machine learning for manufacturing effort estimation on sheet metal bending parts
by: Ballegeer, Matteo, et al.
Published: (2026)
by: Ballegeer, Matteo, et al.
Published: (2026)
LDRFusion: A LiDAR-Dominant multimodal refinement framework for 3D object detection
by: Wang, Jijun, et al.
Published: (2025)
by: Wang, Jijun, et al.
Published: (2025)
From Points to Clouds: Learning Robust Semantic Distributions for Multi-modal Prompts
by: Li, Weiran, et al.
Published: (2025)
by: Li, Weiran, et al.
Published: (2025)
Similar Items
-
DisCo3D: Distilling Multi-View Consistency for 3D Scene Editing
by: Chi, Yufeng, et al.
Published: (2025) -
RCR-AF: Enhancing Model Generalization via Rademacher Complexity Reduction Activation Function
by: Yu, Yunrui, et al.
Published: (2025) -
PersonificationNet: Making customized subject act like a person
by: Guo, Tianchu, et al.
Published: (2024) -
Robust and Efficient Adversarial Defense in SNNs via Image Purification and Joint Detection
by: Chen, Weiran, et al.
Published: (2024) -
MIMO: A medical vision language model with visual referring multimodal input and pixel grounding multimodal output
by: Chen, Yanyuan, et al.
Published: (2025)