Saved in:
| Main Authors: | Wang, Zitong, Shen, Zijun, Xu, Haohao, Luo, Zhengjie, Wu, Weibin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2603.10210 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PA-FAS: Towards Interpretable and Generalizable Multimodal Face Anti-Spoofing via Path-Augmented Reinforcement Learning
by: Ma, Yingjie, et al.
Published: (2025)
by: Ma, Yingjie, et al.
Published: (2025)
XD-MAP: Cross-Modal Domain Adaptation via Semantic Parametric Maps for Scalable Training Data Generation
by: Bieder, Frank, et al.
Published: (2026)
by: Bieder, Frank, et al.
Published: (2026)
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption
by: Fan, Tiehan, et al.
Published: (2024)
by: Fan, Tiehan, et al.
Published: (2024)
Intuitive Axial Augmentation Using Polar-Sine-Based Piecewise Distortion for Medical Slice-Wise Segmentation
by: Zhang, Yiqin, et al.
Published: (2024)
by: Zhang, Yiqin, et al.
Published: (2024)
IFAdapter: Instance Feature Control for Grounded Text-to-Image Generation
by: Wu, Yinwei, et al.
Published: (2024)
by: Wu, Yinwei, et al.
Published: (2024)
Cross-modality Guidance-aided Multi-modal Learning with Dual Attention for MRI Brain Tumor Grading
by: Xu, Dunyuan, et al.
Published: (2024)
by: Xu, Dunyuan, et al.
Published: (2024)
Taming Hallucinations: Boosting MLLMs' Video Understanding via Counterfactual Video Generation
by: Huang, Zhe, et al.
Published: (2025)
by: Huang, Zhe, et al.
Published: (2025)
Analyzing and Boosting the Power of Fine-Grained Visual Recognition for Multi-modal Large Language Models
by: He, Hulingxiao, et al.
Published: (2025)
by: He, Hulingxiao, et al.
Published: (2025)
HAISTA-NET: Human Assisted Instance Segmentation Through Attention
by: Korkmaz, Muhammed, et al.
Published: (2023)
by: Korkmaz, Muhammed, et al.
Published: (2023)
VISOR: Agentic Visual Retrieval-Augmented Generation via Iterative Search and Over-horizon Reasoning
by: Shen, Yucheng, et al.
Published: (2026)
by: Shen, Yucheng, et al.
Published: (2026)
SCA: Improve Semantic Consistent in Unrestricted Adversarial Attacks via DDPM Inversion
by: Pan, Zihao, et al.
Published: (2024)
by: Pan, Zihao, et al.
Published: (2024)
Boosting Overlapping Organoid Instance Segmentation Using Pseudo-Label Unmixing and Synthesis-Assisted Learning
by: Huang, Gui, et al.
Published: (2026)
by: Huang, Gui, et al.
Published: (2026)
SAC-MIL: Spatial-Aware Correlated Multiple Instance Learning for Histopathology Whole Slide Image Classification
by: Bai, Yu, et al.
Published: (2025)
by: Bai, Yu, et al.
Published: (2025)
Magic-Boost: Boost 3D Generation with Multi-View Conditioned Diffusion
by: Yang, Fan, et al.
Published: (2024)
by: Yang, Fan, et al.
Published: (2024)
GPT-4 Enhanced Multimodal Grounding for Autonomous Driving: Leveraging Cross-Modal Attention with Large Language Models
by: Liao, Haicheng, et al.
Published: (2023)
by: Liao, Haicheng, et al.
Published: (2023)
Boost Self-Supervised Dataset Distillation via Parameterization, Predefined Augmentation, and Approximation
by: Yu, Sheng-Feng, et al.
Published: (2025)
by: Yu, Sheng-Feng, et al.
Published: (2025)
Single-Teacher View Augmentation: Boosting Knowledge Distillation via Angular Diversity
by: Yu, Seonghoon, et al.
Published: (2025)
by: Yu, Seonghoon, et al.
Published: (2025)
RAAP: Retrieval-Augmented Affordance Prediction with Cross-Image Action Alignment
by: Zhuang, Qiyuan, et al.
Published: (2026)
by: Zhuang, Qiyuan, et al.
Published: (2026)
Text3DAug -- Prompted Instance Augmentation for LiDAR Perception
by: Reichardt, Laurenz, et al.
Published: (2024)
by: Reichardt, Laurenz, et al.
Published: (2024)
InstanceDiffusion: Instance-level Control for Image Generation
by: Wang, Xudong, et al.
Published: (2024)
by: Wang, Xudong, et al.
Published: (2024)
Fair Lung Disease Diagnosis from Chest CT via Gender-Adversarial Attention Multiple Instance Learning
by: Parikh, Aditya, et al.
Published: (2026)
by: Parikh, Aditya, et al.
Published: (2026)
MSLoRA: Multi-Scale Low-Rank Adaptation via Attention Reweighting
by: Yang, Xu, et al.
Published: (2025)
by: Yang, Xu, et al.
Published: (2025)
DMS-Net:Dual-Modal Multi-Scale Siamese Network for Binocular Fundus Image Classification
by: Huo, Guohao, et al.
Published: (2025)
by: Huo, Guohao, et al.
Published: (2025)
Semi-supervised Semantic Segmentation for Remote Sensing Images via Multi-scale Uncertainty Consistency and Cross-Teacher-Student Attention
by: Wang, Shanwen, et al.
Published: (2025)
by: Wang, Shanwen, et al.
Published: (2025)
DraftAttention: Fast Video Diffusion via Low-Resolution Attention Guidance
by: Shen, Xuan, et al.
Published: (2025)
by: Shen, Xuan, et al.
Published: (2025)
Generalized Class Discovery in Instance Segmentation
by: Hoang, Cuong Manh, et al.
Published: (2025)
by: Hoang, Cuong Manh, et al.
Published: (2025)
RaDL: Relation-aware Disentangled Learning for Multi-Instance Text-to-Image Generation
by: Park, Geon, et al.
Published: (2025)
by: Park, Geon, et al.
Published: (2025)
Visual Instance-aware Prompt Tuning
by: Xiao, Xi, et al.
Published: (2025)
by: Xiao, Xi, et al.
Published: (2025)
Object-level Cross-view Geo-localization with Location Enhancement and Multi-Head Cross Attention
by: Huang, Zheyang, et al.
Published: (2025)
by: Huang, Zheyang, et al.
Published: (2025)
Mamba-VMR: Multimodal Query Augmentation via Generated Videos for Precise Temporal Grounding
by: Sun, Yunzhuo, et al.
Published: (2026)
by: Sun, Yunzhuo, et al.
Published: (2026)
Boosting Generalizability towards Zero-Shot Cross-Dataset Single-Image Indoor Depth by Meta-Initialization
by: Wu, Cho-Ying, et al.
Published: (2024)
by: Wu, Cho-Ying, et al.
Published: (2024)
Semantic Image Synthesis via Class-Adaptive Cross-Attention
by: Fontanini, Tomaso, et al.
Published: (2023)
by: Fontanini, Tomaso, et al.
Published: (2023)
Disco: Densely-overlapping Cell Instance Segmentation via Adjacency-aware Collaborative Coloring
by: Sun, Rui, et al.
Published: (2026)
by: Sun, Rui, et al.
Published: (2026)
Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought Reasoning
by: He, Hulingxiao, et al.
Published: (2026)
by: He, Hulingxiao, et al.
Published: (2026)
Wavelet-based Global-Local Interaction Network with Cross-Attention for Multi-View Diabetic Retinopathy Detection
by: Hu, Yongting, et al.
Published: (2025)
by: Hu, Yongting, et al.
Published: (2025)
Robust Latent Matters: Boosting Image Generation with Sampling Error Synthesis
by: Qiu, Kai, et al.
Published: (2025)
by: Qiu, Kai, et al.
Published: (2025)
Pedestrian Attribute Recognition via Hierarchical Cross-Modality HyperGraph Learning
by: Wang, Xiao, et al.
Published: (2025)
by: Wang, Xiao, et al.
Published: (2025)
NutriScreener: Retrieval-Augmented Multi-Pose Graph Attention Network for Malnourishment Screening
by: Khan, Misaal, et al.
Published: (2025)
by: Khan, Misaal, et al.
Published: (2025)
MRG: A Multi-Robot Manufacturing Digital Scene Generation Method Using Multi-Instance Point Cloud Registration
by: Han, Songjie, et al.
Published: (2025)
by: Han, Songjie, et al.
Published: (2025)
CMHANet: A Cross-Modal Hybrid Attention Network for Point Cloud Registration
by: Zhang, Dongxu, et al.
Published: (2026)
by: Zhang, Dongxu, et al.
Published: (2026)
Similar Items
-
PA-FAS: Towards Interpretable and Generalizable Multimodal Face Anti-Spoofing via Path-Augmented Reinforcement Learning
by: Ma, Yingjie, et al.
Published: (2025) -
XD-MAP: Cross-Modal Domain Adaptation via Semantic Parametric Maps for Scalable Training Data Generation
by: Bieder, Frank, et al.
Published: (2026) -
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption
by: Fan, Tiehan, et al.
Published: (2024) -
Intuitive Axial Augmentation Using Polar-Sine-Based Piecewise Distortion for Medical Slice-Wise Segmentation
by: Zhang, Yiqin, et al.
Published: (2024) -
IFAdapter: Instance Feature Control for Grounded Text-to-Image Generation
by: Wu, Yinwei, et al.
Published: (2024)