Elliptical Attention
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Nielsen, Stefan K., Abdullaev, Laziz U., Teo, Rachel S. Y., Nguyen, Tan M. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Unveiling the Hidden Structure of Self-Attention via Kernel Principal Component Analysis
von: Teo, Rachel S. Y., et al.
Veröffentlicht: (2024)
von: Teo, Rachel S. Y., et al.
Veröffentlicht: (2024)
MomentumSMoE: Integrating Momentum into Sparse Mixture of Experts
von: Teo, Rachel S. Y., et al.
Veröffentlicht: (2024)
von: Teo, Rachel S. Y., et al.
Veröffentlicht: (2024)
MoLEx: Mixture of Layer Experts for Finetuning with Sparse Upcycling
von: Teo, Rachel S. Y., et al.
Veröffentlicht: (2025)
von: Teo, Rachel S. Y., et al.
Veröffentlicht: (2025)
Revisiting Transformers with Insights from Image Filtering and Boosting
von: Abdullaev, Laziz U., et al.
Veröffentlicht: (2025)
von: Abdullaev, Laziz U., et al.
Veröffentlicht: (2025)
The Blessing and Curse of Dimensionality in Safety Alignment
von: Teo, Rachel S. Y., et al.
Veröffentlicht: (2025)
von: Teo, Rachel S. Y., et al.
Veröffentlicht: (2025)
A Primal-Dual Framework for Transformers and Neural Networks
von: Nguyen, Tan M., et al.
Veröffentlicht: (2024)
von: Nguyen, Tan M., et al.
Veröffentlicht: (2024)
BitNet b1.58 Reloaded: State-of-the-art Performance Also on Smaller Networks
von: Nielsen, Jacob, et al.
Veröffentlicht: (2024)
von: Nielsen, Jacob, et al.
Veröffentlicht: (2024)
A Novel Framework for Automated Explain Vision Model Using Vision-Language Models
von: Nguyen, Phu-Vinh, et al.
Veröffentlicht: (2025)
von: Nguyen, Phu-Vinh, et al.
Veröffentlicht: (2025)
LatentLLM: Attention-Aware Joint Tensor Compression
von: Koike-Akino, Toshiaki, et al.
Veröffentlicht: (2025)
von: Koike-Akino, Toshiaki, et al.
Veröffentlicht: (2025)
Divide & Bind Your Attention for Improved Generative Semantic Nursing
von: Li, Yumeng, et al.
Veröffentlicht: (2023)
von: Li, Yumeng, et al.
Veröffentlicht: (2023)
Out-of-Distribution Detection with Attention Head Masking for Multimodal Document Classification
von: Constantinou, Christos, et al.
Veröffentlicht: (2024)
von: Constantinou, Christos, et al.
Veröffentlicht: (2024)
AttnLRP: Attention-Aware Layer-Wise Relevance Propagation for Transformers
von: Achtibat, Reduan, et al.
Veröffentlicht: (2024)
von: Achtibat, Reduan, et al.
Veröffentlicht: (2024)
Voila-A: Aligning Vision-Language Models with User's Gaze Attention
von: Yan, Kun, et al.
Veröffentlicht: (2023)
von: Yan, Kun, et al.
Veröffentlicht: (2023)
Federated Document Visual Question Answering: A Pilot Study
von: Nguyen, Khanh, et al.
Veröffentlicht: (2024)
von: Nguyen, Khanh, et al.
Veröffentlicht: (2024)
OSCaR: Object State Captioning and State Change Representation
von: Nguyen, Nguyen, et al.
Veröffentlicht: (2024)
von: Nguyen, Nguyen, et al.
Veröffentlicht: (2024)
A Reproduction Study: The Kernel PCA Interpretation of Self-Attention Fails Under Scrutiny
von: Sarıtaş, Karahan, et al.
Veröffentlicht: (2025)
von: Sarıtaş, Karahan, et al.
Veröffentlicht: (2025)
Zero-Shot Visual Reasoning by Vision-Language Models: Benchmarking and Analysis
von: Nagar, Aishik, et al.
Veröffentlicht: (2024)
von: Nagar, Aishik, et al.
Veröffentlicht: (2024)
The LLM Bottleneck: Why Open-Source Vision LLMs Struggle with Hierarchical Visual Recognition
von: Tan, Yuwen, et al.
Veröffentlicht: (2025)
von: Tan, Yuwen, et al.
Veröffentlicht: (2025)
CP-MoE: Consistency-Preserving Mixture-of-Experts for Continual Learning
von: Liu, Yang, et al.
Veröffentlicht: (2026)
von: Liu, Yang, et al.
Veröffentlicht: (2026)
MHA2MLA-VLM: Enabling DeepSeek's Economical Multi-Head Latent Attention across Vision-Language Models
von: Fan, Xiaoran, et al.
Veröffentlicht: (2026)
von: Fan, Xiaoran, et al.
Veröffentlicht: (2026)
Match & Choose: Model Selection Framework for Fine-tuning Text-to-Image Diffusion Models
von: Lewandowski, Basile, et al.
Veröffentlicht: (2025)
von: Lewandowski, Basile, et al.
Veröffentlicht: (2025)
Pre-trained Vision-Language Models Learn Discoverable Visual Concepts
von: Zang, Yuan, et al.
Veröffentlicht: (2024)
von: Zang, Yuan, et al.
Veröffentlicht: (2024)
Rethinking Training Dynamics in Scale-wise Autoregressive Generation
von: Zhou, Gengze, et al.
Veröffentlicht: (2025)
von: Zhou, Gengze, et al.
Veröffentlicht: (2025)
GUI-AIMA: Aligning Intrinsic Multimodal Attention with a Context Anchor for GUI Grounding
von: Zhou, Shijie, et al.
Veröffentlicht: (2025)
von: Zhou, Shijie, et al.
Veröffentlicht: (2025)
Can Visual Encoder Learn to See Arrows?
von: Terashita, Naoyuki, et al.
Veröffentlicht: (2025)
von: Terashita, Naoyuki, et al.
Veröffentlicht: (2025)
Logical Closed Loop: Uncovering Object Hallucinations in Large Vision-Language Models
von: Wu, Junfei, et al.
Veröffentlicht: (2024)
von: Wu, Junfei, et al.
Veröffentlicht: (2024)
The Best of Both Worlds: Integrating Language Models and Diffusion Models for Video Generation
von: Yin, Aoxiong, et al.
Veröffentlicht: (2025)
von: Yin, Aoxiong, et al.
Veröffentlicht: (2025)
CausalChaos! Dataset for Comprehensive Causal Action Question Answering Over Longer Causal Chains Grounded in Dynamic Visual Scenes
von: Parmar, Paritosh, et al.
Veröffentlicht: (2024)
von: Parmar, Paritosh, et al.
Veröffentlicht: (2024)
LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention
von: Zhang, Renrui, et al.
Veröffentlicht: (2023)
von: Zhang, Renrui, et al.
Veröffentlicht: (2023)
ZigMa: A DiT-style Zigzag Mamba Diffusion Model
von: Hu, Vincent Tao, et al.
Veröffentlicht: (2024)
von: Hu, Vincent Tao, et al.
Veröffentlicht: (2024)
Counterfactual Segmentation Reasoning: Diagnosing and Mitigating Pixel-Grounding Hallucination
von: Li, Xinzhuo, et al.
Veröffentlicht: (2025)
von: Li, Xinzhuo, et al.
Veröffentlicht: (2025)
BiCLIP: Domain Canonicalization via Structured Geometric Transformation
von: Mantini, Pranav, et al.
Veröffentlicht: (2026)
von: Mantini, Pranav, et al.
Veröffentlicht: (2026)
Multimodal Needle in a Haystack: Benchmarking Long-Context Capability of Multimodal Large Language Models
von: Wang, Hengyi, et al.
Veröffentlicht: (2024)
von: Wang, Hengyi, et al.
Veröffentlicht: (2024)
ColorBench: Can VLMs See and Understand the Colorful World? A Comprehensive Benchmark for Color Perception, Reasoning, and Robustness
von: Liang, Yijun, et al.
Veröffentlicht: (2025)
von: Liang, Yijun, et al.
Veröffentlicht: (2025)
mDPO: Conditional Preference Optimization for Multimodal Large Language Models
von: Wang, Fei, et al.
Veröffentlicht: (2024)
von: Wang, Fei, et al.
Veröffentlicht: (2024)
Improving Prediction Performance and Model Interpretability through Attention Mechanisms from Basic and Applied Research Perspectives
von: Kitada, Shunsuke
Veröffentlicht: (2023)
von: Kitada, Shunsuke
Veröffentlicht: (2023)
Many-Shot In-Context Learning in Multimodal Foundation Models
von: Jiang, Yixing, et al.
Veröffentlicht: (2024)
von: Jiang, Yixing, et al.
Veröffentlicht: (2024)
Not (yet) the whole story: Evaluating Visual Storytelling Requires More than Measuring Coherence, Grounding, and Repetition
von: Surikuchi, Aditya K, et al.
Veröffentlicht: (2024)
von: Surikuchi, Aditya K, et al.
Veröffentlicht: (2024)
Natural Language Generation from Visual Events: State-of-the-Art and Key Open Questions
von: Surikuchi, Aditya K, et al.
Veröffentlicht: (2025)
von: Surikuchi, Aditya K, et al.
Veröffentlicht: (2025)
Physical Property Understanding from Language-Embedded Feature Fields
von: Zhai, Albert J., et al.
Veröffentlicht: (2024)
von: Zhai, Albert J., et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Unveiling the Hidden Structure of Self-Attention via Kernel Principal Component Analysis
von: Teo, Rachel S. Y., et al.
Veröffentlicht: (2024) -
MomentumSMoE: Integrating Momentum into Sparse Mixture of Experts
von: Teo, Rachel S. Y., et al.
Veröffentlicht: (2024) -
MoLEx: Mixture of Layer Experts for Finetuning with Sparse Upcycling
von: Teo, Rachel S. Y., et al.
Veröffentlicht: (2025) -
Revisiting Transformers with Insights from Image Filtering and Boosting
von: Abdullaev, Laziz U., et al.
Veröffentlicht: (2025) -
The Blessing and Curse of Dimensionality in Safety Alignment
von: Teo, Rachel S. Y., et al.
Veröffentlicht: (2025)