Cross-Attention Fusion of Visual and Geometric Features for Large Vocabulary Arabic Lipreading
Fuente:
arXiv
Guardado en:
| Autores principales: | Daou, Samar, Ben-Hamadou, Achraf, Rekik, Ahmed, Kallel, Abdelaziz |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Towards Open-Vocabulary Audio-Visual Event Localization
por: Zhou, Jinxing, et al.
Publicado: (2024)
por: Zhou, Jinxing, et al.
Publicado: (2024)
Where Does Vision Meet Language? Understanding and Refining Visual Fusion in MLLMs via Contrastive Attention
por: Song, Shezheng, et al.
Publicado: (2026)
por: Song, Shezheng, et al.
Publicado: (2026)
Vocabulary Hijacking in LVLMs: Unveiling Critical Attention Heads by Excluding Inert Tokens to Mitigate Hallucination
por: Chen, Yangneng, et al.
Publicado: (2026)
por: Chen, Yangneng, et al.
Publicado: (2026)
Cross-Modal Binary Attention: An Energy-Efficient Fusion Framework for Audio-Visual Learning
por: Saleh, Mohamed, et al.
Publicado: (2026)
por: Saleh, Mohamed, et al.
Publicado: (2026)
DepthGait: Multi-Scale Cross-Level Feature Fusion of RGB-Derived Depth and Silhouette Sequences for Robust Gait Recognition
por: Li, Xinzhu, et al.
Publicado: (2025)
por: Li, Xinzhu, et al.
Publicado: (2025)
Scalable Image Coding for Humans and Machines Using Feature Fusion Network
por: Shindo, Takahiro, et al.
Publicado: (2024)
por: Shindo, Takahiro, et al.
Publicado: (2024)
Multi-Modal Image Fusion via Intervention-Stable Feature Learning
por: Wang, Xue, et al.
Publicado: (2026)
por: Wang, Xue, et al.
Publicado: (2026)
Joint Flow And Feature Refinement Using Attention For Video Restoration
por: Merugu, Ranjith, et al.
Publicado: (2025)
por: Merugu, Ranjith, et al.
Publicado: (2025)
MSCT: Differential Cross-Modal Attention for Deepfake Detection
por: Wei, Fangda, et al.
Publicado: (2026)
por: Wei, Fangda, et al.
Publicado: (2026)
AVFF: Audio-Visual Feature Fusion for Video Deepfake Detection
por: Oorloff, Trevine, et al.
Publicado: (2024)
por: Oorloff, Trevine, et al.
Publicado: (2024)
Embedded Heterogeneous Attention Transformer for Cross-lingual Image Captioning
por: Song, Zijie, et al.
Publicado: (2023)
por: Song, Zijie, et al.
Publicado: (2023)
CLIP Brings Better Features to Visual Aesthetics Learners
por: Xu, Liwu, et al.
Publicado: (2023)
por: Xu, Liwu, et al.
Publicado: (2023)
MindTuner: Cross-Subject Visual Decoding with Visual Fingerprint and Semantic Correction
por: Gong, Zixuan, et al.
Publicado: (2024)
por: Gong, Zixuan, et al.
Publicado: (2024)
Probabilistic Temporal Masked Attention for Cross-view Online Action Detection
por: Xie, Liping, et al.
Publicado: (2025)
por: Xie, Liping, et al.
Publicado: (2025)
Exploring Mutual Cross-Modal Attention for Context-Aware Human Affordance Generation
por: Roy, Prasun, et al.
Publicado: (2025)
por: Roy, Prasun, et al.
Publicado: (2025)
SFFNet: Synergistic Feature Fusion Network With Dual-Domain Edge Enhancement for UAV Image Object Detection
por: Zhang, Wenfeng, et al.
Publicado: (2026)
por: Zhang, Wenfeng, et al.
Publicado: (2026)
Towards Open-Vocabulary Remote Sensing Image Semantic Segmentation
por: Ye, Chengyang, et al.
Publicado: (2024)
por: Ye, Chengyang, et al.
Publicado: (2024)
TextRefiner: Internal Visual Feature as Efficient Refiner for Vision-Language Models Prompt Tuning
por: Xie, Jingjing, et al.
Publicado: (2024)
por: Xie, Jingjing, et al.
Publicado: (2024)
Teacher-Guided Pseudo Supervision and Cross-Modal Alignment for Audio-Visual Video Parsing
por: Chen, Yaru, et al.
Publicado: (2025)
por: Chen, Yaru, et al.
Publicado: (2025)
Decompose and Transfer: CoT-Prompting Enhanced Alignment for Open-Vocabulary Temporal Action Detection
por: Zhu, Sa, et al.
Publicado: (2026)
por: Zhu, Sa, et al.
Publicado: (2026)
Art2Mus: Artwork-to-Music Generation via Visual Conditioning and Large-Scale Cross-Modal Alignment
por: Rinaldi, Ivan, et al.
Publicado: (2026)
por: Rinaldi, Ivan, et al.
Publicado: (2026)
Segmentation-Based Attention Entropy: Detecting and Mitigating Object Hallucinations in Large Vision-Language Models
por: Song, Jiale, et al.
Publicado: (2026)
por: Song, Jiale, et al.
Publicado: (2026)
Interactive Spatial-Frequency Fusion Mamba for Multi-Modal Image Fusion
por: Zhu, Yixin, et al.
Publicado: (2026)
por: Zhu, Yixin, et al.
Publicado: (2026)
Incorporating Visual Experts to Resolve the Information Loss in Multimodal Large Language Models
por: He, Xin, et al.
Publicado: (2024)
por: He, Xin, et al.
Publicado: (2024)
DT-UFC: Universal Large Model Feature Coding via Peaky-to-Balanced Distribution Transformation
por: Gao, Changsheng, et al.
Publicado: (2025)
por: Gao, Changsheng, et al.
Publicado: (2025)
DeepMoLM: Leveraging Visual and Geometric Structural Information for Molecule-Text Modeling
por: Lan, Jing, et al.
Publicado: (2026)
por: Lan, Jing, et al.
Publicado: (2026)
LLM-based Fusion of Multi-modal Features for Commercial Memorability Prediction
por: Pramov, Aleksandar
Publicado: (2025)
por: Pramov, Aleksandar
Publicado: (2025)
Med-Banana-50K: A Cross-modality Large-Scale Dataset for Text-guided Medical Image Editing
por: Chen, Zhihui, et al.
Publicado: (2025)
por: Chen, Zhihui, et al.
Publicado: (2025)
XEmbodied: A Foundation Model with Enhanced Geometric and Physical Cues for Large-Scale Embodied Environments
por: Qian, Kangan, et al.
Publicado: (2026)
por: Qian, Kangan, et al.
Publicado: (2026)
Multiscale Feature Importance-based Bit Allocation for End-to-End Feature Coding for Machines
por: Liu, Junle, et al.
Publicado: (2025)
por: Liu, Junle, et al.
Publicado: (2025)
Efficient Object-centric Representation Learning with Pre-trained Geometric Prior
por: Khac, Phúc H. Le, et al.
Publicado: (2024)
por: Khac, Phúc H. Le, et al.
Publicado: (2024)
CL2CM: Improving Cross-Lingual Cross-Modal Retrieval via Cross-Lingual Knowledge Transfer
por: Wang, Yabing, et al.
Publicado: (2023)
por: Wang, Yabing, et al.
Publicado: (2023)
AttentionBender: Manipulating Cross-Attention in Video Diffusion Transformers as a Creative Probe
por: Cole, Adam, et al.
Publicado: (2026)
por: Cole, Adam, et al.
Publicado: (2026)
Multi-scale Attention Guided Pose Transfer
por: Roy, Prasun, et al.
Publicado: (2022)
por: Roy, Prasun, et al.
Publicado: (2022)
Securing Social Media Against Deepfakes using Identity, Behavioral, and Geometric Signatures
por: Farooq, Muhammad Umar, et al.
Publicado: (2024)
por: Farooq, Muhammad Umar, et al.
Publicado: (2024)
Cross Modification Attention Based Deliberation Model for Image Captioning
por: Lian, Zheng, et al.
Publicado: (2021)
por: Lian, Zheng, et al.
Publicado: (2021)
MoRAG -- Multi-Fusion Retrieval Augmented Generation for Human Motion
por: Kalakonda, Sai Shashank, et al.
Publicado: (2024)
por: Kalakonda, Sai Shashank, et al.
Publicado: (2024)
Anchoring Emotions in Text: Robust Multimodal Fusion for Mimicry Intensity Estimation
por: Zhu, Lingsi, et al.
Publicado: (2026)
por: Zhu, Lingsi, et al.
Publicado: (2026)
Depth and Image Fusion for Road Obstacle Detection Using Stereo Camera
por: Perezyabov, Oleg, et al.
Publicado: (2025)
por: Perezyabov, Oleg, et al.
Publicado: (2025)
Tile Classification Based Viewport Prediction with Multi-modal Fusion Transformer
por: Zhang, Zhihao, et al.
Publicado: (2023)
por: Zhang, Zhihao, et al.
Publicado: (2023)
Ejemplares similares
-
Towards Open-Vocabulary Audio-Visual Event Localization
por: Zhou, Jinxing, et al.
Publicado: (2024) -
Where Does Vision Meet Language? Understanding and Refining Visual Fusion in MLLMs via Contrastive Attention
por: Song, Shezheng, et al.
Publicado: (2026) -
Vocabulary Hijacking in LVLMs: Unveiling Critical Attention Heads by Excluding Inert Tokens to Mitigate Hallucination
por: Chen, Yangneng, et al.
Publicado: (2026) -
Cross-Modal Binary Attention: An Energy-Efficient Fusion Framework for Audio-Visual Learning
por: Saleh, Mohamed, et al.
Publicado: (2026) -
DepthGait: Multi-Scale Cross-Level Feature Fusion of RGB-Derived Depth and Silhouette Sequences for Robust Gait Recognition
por: Li, Xinzhu, et al.
Publicado: (2025)