Fibottention: Inceptive Visual Representation Learning with Diverse Attention Across Heads
Fuente:
arXiv
Saved in:
| Main Authors: | Rahimian, Ali K., Govind, Manish K., Maity, Subhajit, Reilly, Dominick, Kümmerle, Christian, Das, Srijan, Dutta, Aritra |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models
by: Reilly, Dominick, et al.
Published: (2025)
by: Reilly, Dominick, et al.
Published: (2025)
UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models
by: Govind, Manish Kumar, et al.
Published: (2026)
by: Govind, Manish Kumar, et al.
Published: (2026)
From My View to Yours: Ego-to-Exo Transfer in VLMs for Understanding Activities of Daily Living
by: Reilly, Dominick, et al.
Published: (2025)
by: Reilly, Dominick, et al.
Published: (2025)
Kolmogorov-Arnold Attention: Is Learnable Attention Better For Vision Transformers?
by: Maity, Subhajit, et al.
Published: (2025)
by: Maity, Subhajit, et al.
Published: (2025)
LLAVIDAL: A Large LAnguage VIsion Model for Daily Activities of Living
by: Reilly, Dominick, et al.
Published: (2024)
by: Reilly, Dominick, et al.
Published: (2024)
SKI Models: Skeleton Induced Vision-Language Embeddings for Understanding Activities of Daily Living
by: Sinha, Arkaprava, et al.
Published: (2025)
by: Sinha, Arkaprava, et al.
Published: (2025)
Inceptive Transformers: Enhancing Contextual Representations through Multi-Scale Feature Learning Across Domains and Languages
by: Shahriar, Asif, et al.
Published: (2025)
by: Shahriar, Asif, et al.
Published: (2025)
Self-supervised Auxiliary Learning for Texture and Model-based Hybrid Robust and Fair Featuring in Face Analysis
by: Reddy, Shukesh, et al.
Published: (2024)
by: Reddy, Shukesh, et al.
Published: (2024)
Fusion-SSAT: Unleashing the Potential of Self-supervised Auxiliary Task by Feature Fusion for Generalized Deepfake Detection
by: Reddy, Shukesh, et al.
Published: (2026)
by: Reddy, Shukesh, et al.
Published: (2026)
Introducing Gating and Context into Temporal Action Detection
by: Reka, Aglind, et al.
Published: (2024)
by: Reka, Aglind, et al.
Published: (2024)
Recent Advancement in 3D Biometrics using Monocular Camera
by: Mukherjee, Aritra, et al.
Published: (2024)
by: Mukherjee, Aritra, et al.
Published: (2024)
Robust Multi-Modal Image Stitching for Improved Scene Understanding
by: Dutta, Aritra, et al.
Published: (2023)
by: Dutta, Aritra, et al.
Published: (2023)
Transferable Tactile Transformers for Representation Learning Across Diverse Sensors and Tasks
by: Zhao, Jialiang, et al.
Published: (2024)
by: Zhao, Jialiang, et al.
Published: (2024)
HAViT: Historical Attention Vision Transformer
by: Banik, Swarnendu, et al.
Published: (2026)
by: Banik, Swarnendu, et al.
Published: (2026)
Selective LoRA for Visual Tokens and Attention Heads
by: Luo, Tiange, et al.
Published: (2025)
by: Luo, Tiange, et al.
Published: (2025)
Intelligent Image Sensing for Crime Analysis: A ML Approach towards Enhanced Violence Detection and Investigation
by: Dutta, Aritra, et al.
Published: (2025)
by: Dutta, Aritra, et al.
Published: (2025)
Prompt-based Dynamic Token Pruning for Efficient Segmentation of Medical Images
by: Dutta, Pallabi, et al.
Published: (2025)
by: Dutta, Pallabi, et al.
Published: (2025)
Unveiling Visual Perception in Language Models: An Attention Head Analysis Approach
by: Bi, Jing, et al.
Published: (2024)
by: Bi, Jing, et al.
Published: (2024)
GaussianHeads: End-to-End Learning of Drivable Gaussian Head Avatars from Coarse-to-fine Representations
by: Teotia, Kartik, et al.
Published: (2024)
by: Teotia, Kartik, et al.
Published: (2024)
Towards Multi-modal Transformers in Federated Learning
by: Sun, Guangyu, et al.
Published: (2024)
by: Sun, Guangyu, et al.
Published: (2024)
Learning 1D Causal Visual Representation with De-focus Attention Networks
by: Tao, Chenxin, et al.
Published: (2024)
by: Tao, Chenxin, et al.
Published: (2024)
Beyond Pixels: Semi-Supervised Semantic Segmentation with a Multi-scale Patch-based Multi-Label Classifier
by: Howlader, Prantik, et al.
Published: (2024)
by: Howlader, Prantik, et al.
Published: (2024)
TICON: A Slide-Level Tile Contextualizer for Histopathology Representation Learning
by: Belagali, Varun, et al.
Published: (2025)
by: Belagali, Varun, et al.
Published: (2025)
DistilDoc: Knowledge Distillation for Visually-Rich Document Applications
by: Van Landeghem, Jordy, et al.
Published: (2024)
by: Van Landeghem, Jordy, et al.
Published: (2024)
Not All Attention Heads Are What You Need: Refining CLIP's Image Representation with Attention Ablation
by: Lin, Feng, et al.
Published: (2025)
by: Lin, Feng, et al.
Published: (2025)
Efficient Image Generation with Variadic Attention Heads
by: Walton, Steven, et al.
Published: (2022)
by: Walton, Steven, et al.
Published: (2022)
Leveraging Deep Learning with Multi-Head Attention for Accurate Extraction of Medicine from Handwritten Prescriptions
by: Ali, Usman, et al.
Published: (2024)
by: Ali, Usman, et al.
Published: (2024)
MoH: Multi-Head Attention as Mixture-of-Head Attention
by: Jin, Peng, et al.
Published: (2024)
by: Jin, Peng, et al.
Published: (2024)
Head-Aware Visual Cropping: Enhancing Fine-Grained VQA with Attention-Guided Subimage
by: Xie, Junfei, et al.
Published: (2026)
by: Xie, Junfei, et al.
Published: (2026)
Self-Rewarded Multimodal Coherent Reasoning Across Diverse Visual Domains
by: Zhang, Jesen, et al.
Published: (2025)
by: Zhang, Jesen, et al.
Published: (2025)
Diverse Image Generation with Diffusion Models and Cross Class Label Learning for Polyp Classification
by: Sharma, Vanshali, et al.
Published: (2025)
by: Sharma, Vanshali, et al.
Published: (2025)
Investigating the Benefits of Projection Head for Representation Learning
by: Xue, Yihao, et al.
Published: (2024)
by: Xue, Yihao, et al.
Published: (2024)
TruthLens:A Training-Free Paradigm for DeepFake Detection
by: Chakraborty, Ritabrata, et al.
Published: (2025)
by: Chakraborty, Ritabrata, et al.
Published: (2025)
DiverseDiT: Towards Diverse Representation Learning in Diffusion Transformers
by: Yang, Mengping, et al.
Published: (2026)
by: Yang, Mengping, et al.
Published: (2026)
Robust Q-Learning under Corrupted Rewards
by: Maity, Sreejeet, et al.
Published: (2024)
by: Maity, Sreejeet, et al.
Published: (2024)
Spectral State Space Model for Rotation-Invariant Visual Representation Learning
by: Dastani, Sahar, et al.
Published: (2025)
by: Dastani, Sahar, et al.
Published: (2025)
CellGenNet: A Knowledge-Distilled Framework for Robust Cell Segmentation in Cancer Tissues
by: Ray, Srijan, et al.
Published: (2025)
by: Ray, Srijan, et al.
Published: (2025)
Corruption-Tolerant Asynchronous Q-Learning with Near-Optimal Rates
by: Maity, Sreejeet, et al.
Published: (2025)
by: Maity, Sreejeet, et al.
Published: (2025)
Recovering Simultaneously Structured Data via Non-Convex Iteratively Reweighted Least Squares
by: Kümmerle, Christian, et al.
Published: (2023)
by: Kümmerle, Christian, et al.
Published: (2023)
An Exposition of Pathfinding Strategies Within Lightning Network Clients
by: Saraswathi, Sindura, et al.
Published: (2024)
by: Saraswathi, Sindura, et al.
Published: (2024)
Similar Items
-
VisCoP: Visual Probing for Video Domain Adaptation of Vision Language Models
by: Reilly, Dominick, et al.
Published: (2025) -
UniLACT: Depth-Aware RGB Latent Action Learning for Vision-Language-Action Models
by: Govind, Manish Kumar, et al.
Published: (2026) -
From My View to Yours: Ego-to-Exo Transfer in VLMs for Understanding Activities of Daily Living
by: Reilly, Dominick, et al.
Published: (2025) -
Kolmogorov-Arnold Attention: Is Learnable Attention Better For Vision Transformers?
by: Maity, Subhajit, et al.
Published: (2025) -
LLAVIDAL: A Large LAnguage VIsion Model for Daily Activities of Living
by: Reilly, Dominick, et al.
Published: (2024)