StochCA: A Novel Approach for Exploiting Pretrained Models with Cross-Attention
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Seo, Seungwon, Lee, Suho, Hwang, Sangheum |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Focusing Where Vision Matters: Selective Training for Large Vision Language Models via Visual Information Gain
von: Lee, Seulbi, et al.
Veröffentlicht: (2026)
von: Lee, Seulbi, et al.
Veröffentlicht: (2026)
Enhanced OoD Detection through Cross-Modal Alignment of Multi-Modal Representations
von: Kim, Jeonghyeon, et al.
Veröffentlicht: (2025)
von: Kim, Jeonghyeon, et al.
Veröffentlicht: (2025)
Localized Concept Erasure in Text-to-Image Diffusion Models via High-Level Representation Misdirection
von: Lee, Uichan, et al.
Veröffentlicht: (2026)
von: Lee, Uichan, et al.
Veröffentlicht: (2026)
Reflexive Guidance: Improving OoDD in Vision-Language Models via Self-Guided Image-Adaptive Concept Generation
von: Kim, Jihyo, et al.
Veröffentlicht: (2024)
von: Kim, Jihyo, et al.
Veröffentlicht: (2024)
GTA: Guided Transfer of Spatial Attention from Object-Centric Representations
von: Seo, SeokHyun, et al.
Veröffentlicht: (2024)
von: Seo, SeokHyun, et al.
Veröffentlicht: (2024)
Parallel Rescaling: Rebalancing Consistency Guidance for Personalized Diffusion Models
von: Chae, JungWoo, et al.
Veröffentlicht: (2025)
von: Chae, JungWoo, et al.
Veröffentlicht: (2025)
CA^2ST: Cross-Attention in Audio, Space, and Time for Holistic Video Recognition
von: Lee, Jongseo, et al.
Veröffentlicht: (2025)
von: Lee, Jongseo, et al.
Veröffentlicht: (2025)
On the Nature of Attention Sink that Shapes Decoding Strategy in Omni-LLMs
von: Yoo, Suho, et al.
Veröffentlicht: (2026)
von: Yoo, Suho, et al.
Veröffentlicht: (2026)
CA-YOLO: Cross Attention Empowered YOLO for Biomimetic Localization
von: Zhang, Zhen, et al.
Veröffentlicht: (2026)
von: Zhang, Zhen, et al.
Veröffentlicht: (2026)
ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention
von: Liu, Wenjie, et al.
Veröffentlicht: (2026)
von: Liu, Wenjie, et al.
Veröffentlicht: (2026)
CrossFuse: A Novel Cross Attention Mechanism based Infrared and Visible Image Fusion Approach
von: Li, Hui, et al.
Veröffentlicht: (2024)
von: Li, Hui, et al.
Veröffentlicht: (2024)
MoCA: Identity-Preserving Text-to-Video Generation via Mixture of Cross Attention
von: Xie, Qi, et al.
Veröffentlicht: (2025)
von: Xie, Qi, et al.
Veröffentlicht: (2025)
StochSync: Stochastic Diffusion Synchronization for Image Generation in Arbitrary Spaces
von: Yeo, Kyeongmin, et al.
Veröffentlicht: (2025)
von: Yeo, Kyeongmin, et al.
Veröffentlicht: (2025)
CA-IDD: Cross-Attention Guided Identity-Conditional Diffusion for Identity-Consistent Face Swapping
von: Rana, Md Shohel, et al.
Veröffentlicht: (2026)
von: Rana, Md Shohel, et al.
Veröffentlicht: (2026)
Improving Calibration in Test-Time Prompt Tuning for Vision-Language Models via Data-Free Flatness-Aware Prompt Pretraining
von: Jang, Hyeonseo, et al.
Veröffentlicht: (2026)
von: Jang, Hyeonseo, et al.
Veröffentlicht: (2026)
CA-Stream: Attention-based pooling for interpretable image recognition
von: Torres, Felipe, et al.
Veröffentlicht: (2024)
von: Torres, Felipe, et al.
Veröffentlicht: (2024)
Towards Scalable Human-aligned Benchmark for Text-guided Image Editing
von: Ryu, Suho, et al.
Veröffentlicht: (2025)
von: Ryu, Suho, et al.
Veröffentlicht: (2025)
SyncMask: Synchronized Attentional Masking for Fashion-centric Vision-Language Pretraining
von: Song, Chull Hwan, et al.
Veröffentlicht: (2024)
von: Song, Chull Hwan, et al.
Veröffentlicht: (2024)
EditCrafter: Tuning-free High-Resolution Image Editing via Pretrained Diffusion Model
von: Kim, Kunho, et al.
Veröffentlicht: (2026)
von: Kim, Kunho, et al.
Veröffentlicht: (2026)
Bridging Audio and Vision: Zero-Shot Audiovisual Segmentation by Connecting Pretrained Models
von: Lee, Seung-jae, et al.
Veröffentlicht: (2025)
von: Lee, Seung-jae, et al.
Veröffentlicht: (2025)
CA-LoRA: Concept-Aware LoRA for Domain-Aligned Segmentation Dataset Generation
von: Park, Minho, et al.
Veröffentlicht: (2025)
von: Park, Minho, et al.
Veröffentlicht: (2025)
CAST: Cross-Attention in Space and Time for Video Action Recognition
von: Lee, Dongho, et al.
Veröffentlicht: (2023)
von: Lee, Dongho, et al.
Veröffentlicht: (2023)
Integrating Query-aware Segmentation and Cross-Attention for Robust VQA
von: Choi, Wonjun, et al.
Veröffentlicht: (2024)
von: Choi, Wonjun, et al.
Veröffentlicht: (2024)
MoCA: Mixture-of-Components Attention for Scalable Compositional 3D Generation
von: Li, Zhiqi, et al.
Veröffentlicht: (2025)
von: Li, Zhiqi, et al.
Veröffentlicht: (2025)
DroneKey++: A Size Prior-free Method and New Benchmark for Drone 3D Pose Estimation from Sequential Images
von: Hwang, Seo-Bin, et al.
Veröffentlicht: (2026)
von: Hwang, Seo-Bin, et al.
Veröffentlicht: (2026)
NeighborMAE: Exploiting Spatial Dependencies between Neighboring Earth Observation Images in Masked Autoencoders Pretraining
von: Zeng, Liang, et al.
Veröffentlicht: (2026)
von: Zeng, Liang, et al.
Veröffentlicht: (2026)
Y-CA-Net: A Convolutional Attention Based Network for Volumetric Medical Image Segmentation
von: Sharif, Muhammad Hamza, et al.
Veröffentlicht: (2024)
von: Sharif, Muhammad Hamza, et al.
Veröffentlicht: (2024)
RoCA: Robust Cross-Domain End-to-End Autonomous Driving
von: Yasarla, Rajeev, et al.
Veröffentlicht: (2025)
von: Yasarla, Rajeev, et al.
Veröffentlicht: (2025)
When Surfaces Lie: Exploiting Wrinkle-Induced Attention Shift to Attack Vision-Language Models
von: Hu, Chengyin, et al.
Veröffentlicht: (2026)
von: Hu, Chengyin, et al.
Veröffentlicht: (2026)
AttentionDrag: Exploiting Latent Correlation Knowledge in Pre-trained Diffusion Models for Image Editing
von: Yang, Biao, et al.
Veröffentlicht: (2025)
von: Yang, Biao, et al.
Veröffentlicht: (2025)
Automatic Channel Pruning for Multi-Head Attention
von: Lee, Eunho, et al.
Veröffentlicht: (2024)
von: Lee, Eunho, et al.
Veröffentlicht: (2024)
State-Space Hierarchical Compression with Gated Attention and Learnable Sampling for Hour-Long Video Understanding in Large Multimodal Models
von: Kim, Geewook, et al.
Veröffentlicht: (2025)
von: Kim, Geewook, et al.
Veröffentlicht: (2025)
Masked Autoregressive Model for Weather Forecasting
von: Kim, Doyi, et al.
Veröffentlicht: (2024)
von: Kim, Doyi, et al.
Veröffentlicht: (2024)
Foreground-Covering Prototype Generation and Matching for SAM-Aided Few-Shot Segmentation
von: Park, Suho, et al.
Veröffentlicht: (2025)
von: Park, Suho, et al.
Veröffentlicht: (2025)
Aligned Novel View Image and Geometry Synthesis via Cross-modal Attention Instillation
von: Kwak, Min-Seop, et al.
Veröffentlicht: (2025)
von: Kwak, Min-Seop, et al.
Veröffentlicht: (2025)
DroneKey: Drone 3D Pose Estimation in Image Sequences using Gated Key-representation and Pose-adaptive Learning
von: Hwang, Seo-Bin, et al.
Veröffentlicht: (2025)
von: Hwang, Seo-Bin, et al.
Veröffentlicht: (2025)
Adapting Pretrained ViTs with Convolution Injector for Visuo-Motor Control
von: Hwang, Dongyoon, et al.
Veröffentlicht: (2024)
von: Hwang, Dongyoon, et al.
Veröffentlicht: (2024)
Attention-driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models without Fine-Tuning
von: Xu, Hai-Ming, et al.
Veröffentlicht: (2024)
von: Xu, Hai-Ming, et al.
Veröffentlicht: (2024)
Compact Attention: Exploiting Structured Spatio-Temporal Sparsity for Fast Video Generation
von: Li, Qirui, et al.
Veröffentlicht: (2025)
von: Li, Qirui, et al.
Veröffentlicht: (2025)
Autoregressive Sign Language Production: A Gloss-Free Approach with Discrete Representations
von: Hwang, Eui Jun, et al.
Veröffentlicht: (2023)
von: Hwang, Eui Jun, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Focusing Where Vision Matters: Selective Training for Large Vision Language Models via Visual Information Gain
von: Lee, Seulbi, et al.
Veröffentlicht: (2026) -
Enhanced OoD Detection through Cross-Modal Alignment of Multi-Modal Representations
von: Kim, Jeonghyeon, et al.
Veröffentlicht: (2025) -
Localized Concept Erasure in Text-to-Image Diffusion Models via High-Level Representation Misdirection
von: Lee, Uichan, et al.
Veröffentlicht: (2026) -
Reflexive Guidance: Improving OoDD in Vision-Language Models via Self-Guided Image-Adaptive Concept Generation
von: Kim, Jihyo, et al.
Veröffentlicht: (2024) -
GTA: Guided Transfer of Spatial Attention from Object-Centric Representations
von: Seo, SeokHyun, et al.
Veröffentlicht: (2024)