Delta-K: Boosting Multi-Instance Generation via Cross-Attention Augmentation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Zitong, Shen, Zijun, Xu, Haohao, Luo, Zhengjie, Wu, Weibin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
PA-FAS: Towards Interpretable and Generalizable Multimodal Face Anti-Spoofing via Path-Augmented Reinforcement Learning
von: Ma, Yingjie, et al.
Veröffentlicht: (2025)
von: Ma, Yingjie, et al.
Veröffentlicht: (2025)
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption
von: Fan, Tiehan, et al.
Veröffentlicht: (2024)
von: Fan, Tiehan, et al.
Veröffentlicht: (2024)
Cross-modality Guidance-aided Multi-modal Learning with Dual Attention for MRI Brain Tumor Grading
von: Xu, Dunyuan, et al.
Veröffentlicht: (2024)
von: Xu, Dunyuan, et al.
Veröffentlicht: (2024)
IFAdapter: Instance Feature Control for Grounded Text-to-Image Generation
von: Wu, Yinwei, et al.
Veröffentlicht: (2024)
von: Wu, Yinwei, et al.
Veröffentlicht: (2024)
Taming Hallucinations: Boosting MLLMs' Video Understanding via Counterfactual Video Generation
von: Huang, Zhe, et al.
Veröffentlicht: (2025)
von: Huang, Zhe, et al.
Veröffentlicht: (2025)
XD-MAP: Cross-Modal Domain Adaptation via Semantic Parametric Maps for Scalable Training Data Generation
von: Bieder, Frank, et al.
Veröffentlicht: (2026)
von: Bieder, Frank, et al.
Veröffentlicht: (2026)
Intuitive Axial Augmentation Using Polar-Sine-Based Piecewise Distortion for Medical Slice-Wise Segmentation
von: Zhang, Yiqin, et al.
Veröffentlicht: (2024)
von: Zhang, Yiqin, et al.
Veröffentlicht: (2024)
HAISTA-NET: Human Assisted Instance Segmentation Through Attention
von: Korkmaz, Muhammed, et al.
Veröffentlicht: (2023)
von: Korkmaz, Muhammed, et al.
Veröffentlicht: (2023)
VISOR: Agentic Visual Retrieval-Augmented Generation via Iterative Search and Over-horizon Reasoning
von: Shen, Yucheng, et al.
Veröffentlicht: (2026)
von: Shen, Yucheng, et al.
Veröffentlicht: (2026)
Boosting Overlapping Organoid Instance Segmentation Using Pseudo-Label Unmixing and Synthesis-Assisted Learning
von: Huang, Gui, et al.
Veröffentlicht: (2026)
von: Huang, Gui, et al.
Veröffentlicht: (2026)
Magic-Boost: Boost 3D Generation with Multi-View Conditioned Diffusion
von: Yang, Fan, et al.
Veröffentlicht: (2024)
von: Yang, Fan, et al.
Veröffentlicht: (2024)
Boost Self-Supervised Dataset Distillation via Parameterization, Predefined Augmentation, and Approximation
von: Yu, Sheng-Feng, et al.
Veröffentlicht: (2025)
von: Yu, Sheng-Feng, et al.
Veröffentlicht: (2025)
Single-Teacher View Augmentation: Boosting Knowledge Distillation via Angular Diversity
von: Yu, Seonghoon, et al.
Veröffentlicht: (2025)
von: Yu, Seonghoon, et al.
Veröffentlicht: (2025)
Text3DAug -- Prompted Instance Augmentation for LiDAR Perception
von: Reichardt, Laurenz, et al.
Veröffentlicht: (2024)
von: Reichardt, Laurenz, et al.
Veröffentlicht: (2024)
Fair Lung Disease Diagnosis from Chest CT via Gender-Adversarial Attention Multiple Instance Learning
von: Parikh, Aditya, et al.
Veröffentlicht: (2026)
von: Parikh, Aditya, et al.
Veröffentlicht: (2026)
GPT-4 Enhanced Multimodal Grounding for Autonomous Driving: Leveraging Cross-Modal Attention with Large Language Models
von: Liao, Haicheng, et al.
Veröffentlicht: (2023)
von: Liao, Haicheng, et al.
Veröffentlicht: (2023)
Semi-supervised Semantic Segmentation for Remote Sensing Images via Multi-scale Uncertainty Consistency and Cross-Teacher-Student Attention
von: Wang, Shanwen, et al.
Veröffentlicht: (2025)
von: Wang, Shanwen, et al.
Veröffentlicht: (2025)
MSLoRA: Multi-Scale Low-Rank Adaptation via Attention Reweighting
von: Yang, Xu, et al.
Veröffentlicht: (2025)
von: Yang, Xu, et al.
Veröffentlicht: (2025)
SAC-MIL: Spatial-Aware Correlated Multiple Instance Learning for Histopathology Whole Slide Image Classification
von: Bai, Yu, et al.
Veröffentlicht: (2025)
von: Bai, Yu, et al.
Veröffentlicht: (2025)
SCA: Improve Semantic Consistent in Unrestricted Adversarial Attacks via DDPM Inversion
von: Pan, Zihao, et al.
Veröffentlicht: (2024)
von: Pan, Zihao, et al.
Veröffentlicht: (2024)
Object-level Cross-view Geo-localization with Location Enhancement and Multi-Head Cross Attention
von: Huang, Zheyang, et al.
Veröffentlicht: (2025)
von: Huang, Zheyang, et al.
Veröffentlicht: (2025)
RAAP: Retrieval-Augmented Affordance Prediction with Cross-Image Action Alignment
von: Zhuang, Qiyuan, et al.
Veröffentlicht: (2026)
von: Zhuang, Qiyuan, et al.
Veröffentlicht: (2026)
RaDL: Relation-aware Disentangled Learning for Multi-Instance Text-to-Image Generation
von: Park, Geon, et al.
Veröffentlicht: (2025)
von: Park, Geon, et al.
Veröffentlicht: (2025)
DraftAttention: Fast Video Diffusion via Low-Resolution Attention Guidance
von: Shen, Xuan, et al.
Veröffentlicht: (2025)
von: Shen, Xuan, et al.
Veröffentlicht: (2025)
Generalized Class Discovery in Instance Segmentation
von: Hoang, Cuong Manh, et al.
Veröffentlicht: (2025)
von: Hoang, Cuong Manh, et al.
Veröffentlicht: (2025)
Semantic Image Synthesis via Class-Adaptive Cross-Attention
von: Fontanini, Tomaso, et al.
Veröffentlicht: (2023)
von: Fontanini, Tomaso, et al.
Veröffentlicht: (2023)
Mamba-VMR: Multimodal Query Augmentation via Generated Videos for Precise Temporal Grounding
von: Sun, Yunzhuo, et al.
Veröffentlicht: (2026)
von: Sun, Yunzhuo, et al.
Veröffentlicht: (2026)
Visual Instance-aware Prompt Tuning
von: Xiao, Xi, et al.
Veröffentlicht: (2025)
von: Xiao, Xi, et al.
Veröffentlicht: (2025)
InstanceDiffusion: Instance-level Control for Image Generation
von: Wang, Xudong, et al.
Veröffentlicht: (2024)
von: Wang, Xudong, et al.
Veröffentlicht: (2024)
Boosting Generalizability towards Zero-Shot Cross-Dataset Single-Image Indoor Depth by Meta-Initialization
von: Wu, Cho-Ying, et al.
Veröffentlicht: (2024)
von: Wu, Cho-Ying, et al.
Veröffentlicht: (2024)
NutriScreener: Retrieval-Augmented Multi-Pose Graph Attention Network for Malnourishment Screening
von: Khan, Misaal, et al.
Veröffentlicht: (2025)
von: Khan, Misaal, et al.
Veröffentlicht: (2025)
Analyzing and Boosting the Power of Fine-Grained Visual Recognition for Multi-modal Large Language Models
von: He, Hulingxiao, et al.
Veröffentlicht: (2025)
von: He, Hulingxiao, et al.
Veröffentlicht: (2025)
Disco: Densely-overlapping Cell Instance Segmentation via Adjacency-aware Collaborative Coloring
von: Sun, Rui, et al.
Veröffentlicht: (2026)
von: Sun, Rui, et al.
Veröffentlicht: (2026)
MRG: A Multi-Robot Manufacturing Digital Scene Generation Method Using Multi-Instance Point Cloud Registration
von: Han, Songjie, et al.
Veröffentlicht: (2025)
von: Han, Songjie, et al.
Veröffentlicht: (2025)
DMS-Net:Dual-Modal Multi-Scale Siamese Network for Binocular Fundus Image Classification
von: Huo, Guohao, et al.
Veröffentlicht: (2025)
von: Huo, Guohao, et al.
Veröffentlicht: (2025)
Robust Latent Matters: Boosting Image Generation with Sampling Error Synthesis
von: Qiu, Kai, et al.
Veröffentlicht: (2025)
von: Qiu, Kai, et al.
Veröffentlicht: (2025)
Pedestrian Attribute Recognition via Hierarchical Cross-Modality HyperGraph Learning
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
CMHANet: A Cross-Modal Hybrid Attention Network for Point Cloud Registration
von: Zhang, Dongxu, et al.
Veröffentlicht: (2026)
von: Zhang, Dongxu, et al.
Veröffentlicht: (2026)
Visual Delta Generator with Large Multi-modal Models for Semi-supervised Composed Image Retrieval
von: Jang, Young Kyun, et al.
Veröffentlicht: (2024)
von: Jang, Young Kyun, et al.
Veröffentlicht: (2024)
Enhancing Traffic Sign Recognition with Tailored Data Augmentation: Addressing Class Imbalance and Instance Scarcity
von: Alsiyeu, Ulan, et al.
Veröffentlicht: (2024)
von: Alsiyeu, Ulan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
PA-FAS: Towards Interpretable and Generalizable Multimodal Face Anti-Spoofing via Path-Augmented Reinforcement Learning
von: Ma, Yingjie, et al.
Veröffentlicht: (2025) -
InstanceCap: Improving Text-to-Video Generation via Instance-aware Structured Caption
von: Fan, Tiehan, et al.
Veröffentlicht: (2024) -
Cross-modality Guidance-aided Multi-modal Learning with Dual Attention for MRI Brain Tumor Grading
von: Xu, Dunyuan, et al.
Veröffentlicht: (2024) -
IFAdapter: Instance Feature Control for Grounded Text-to-Image Generation
von: Wu, Yinwei, et al.
Veröffentlicht: (2024) -
Taming Hallucinations: Boosting MLLMs' Video Understanding via Counterfactual Video Generation
von: Huang, Zhe, et al.
Veröffentlicht: (2025)