Tailored Design of Audio-Visual Speech Recognition Models using Branchformers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gimeno-Gómez, David, Martínez-Hinarejos, Carlos-D. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
AnnoTheia: A Semi-Automatic Annotation Toolkit for Audio-Visual Speech Technologies
von: Acosta-Triana, José-M., et al.
Veröffentlicht: (2024)
von: Acosta-Triana, José-M., et al.
Veröffentlicht: (2024)
Comparison of Conventional Hybrid and CTC/Attention Decoders for Continuous Visual Speech Recognition
von: Gimeno-Gómez, David, et al.
Veröffentlicht: (2024)
von: Gimeno-Gómez, David, et al.
Veröffentlicht: (2024)
Evaluation of End-to-End Continuous Spanish Lipreading in Different Data Conditions
von: Gimeno-Gómez, David, et al.
Veröffentlicht: (2025)
von: Gimeno-Gómez, David, et al.
Veröffentlicht: (2025)
Unveiling Interpretability in Self-Supervised Speech Representations for Parkinson's Diagnosis
von: Gimeno-Gómez, David, et al.
Veröffentlicht: (2024)
von: Gimeno-Gómez, David, et al.
Veröffentlicht: (2024)
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition
von: Nguyen, Thai-Binh, et al.
Veröffentlicht: (2025)
von: Nguyen, Thai-Binh, et al.
Veröffentlicht: (2025)
Designing Practical Models for Isolated Word Visual Speech Recognition
von: Panagos, Iason Ioannis, et al.
Veröffentlicht: (2025)
von: Panagos, Iason Ioannis, et al.
Veröffentlicht: (2025)
Reading Between the Frames: Multi-Modal Depression Detection in Videos from Non-Verbal Cues
von: Gimeno-Gómez, David, et al.
Veröffentlicht: (2024)
von: Gimeno-Gómez, David, et al.
Veröffentlicht: (2024)
Evaluation of Audio-Visual Alignments in Visually Grounded Speech Models
von: Khorrami, Khazar, et al.
Veröffentlicht: (2021)
von: Khorrami, Khazar, et al.
Veröffentlicht: (2021)
Phoneme-Level Visual Speech Recognition via Point-Visual Fusion and Language Model Reconstruction
von: Teng, Matthew Kit Khinn, et al.
Veröffentlicht: (2025)
von: Teng, Matthew Kit Khinn, et al.
Veröffentlicht: (2025)
SyncVSR: Data-Efficient Visual Speech Recognition with End-to-End Crossmodal Audio Token Synchronization
von: Ahn, Young Jin, et al.
Veröffentlicht: (2024)
von: Ahn, Young Jin, et al.
Veröffentlicht: (2024)
CNVSRC 2023: The First Chinese Continuous Visual Speech Recognition Challenge
von: Chen, Chen, et al.
Veröffentlicht: (2024)
von: Chen, Chen, et al.
Veröffentlicht: (2024)
Grounding Language Models for Visual Entity Recognition
von: Xiao, Zilin, et al.
Veröffentlicht: (2024)
von: Xiao, Zilin, et al.
Veröffentlicht: (2024)
Visually Grounded Speech Models for Low-resource Languages and Cognitive Modelling
von: Nortje, Leanne
Veröffentlicht: (2024)
von: Nortje, Leanne
Veröffentlicht: (2024)
Lightweight Operations for Visual Speech Recognition
von: Panagos, Iason Ioannis, et al.
Veröffentlicht: (2025)
von: Panagos, Iason Ioannis, et al.
Veröffentlicht: (2025)
Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation
von: Goncalves, Lucas, et al.
Veröffentlicht: (2024)
von: Goncalves, Lucas, et al.
Veröffentlicht: (2024)
SynesLM: A Unified Approach for Audio-visual Speech Recognition and Translation via Language Model and Synthetic Data
von: Lu, Yichen, et al.
Veröffentlicht: (2024)
von: Lu, Yichen, et al.
Veröffentlicht: (2024)
video-SALMONN 2: Caption-Enhanced Audio-Visual Large Language Models
von: Tang, Changli, et al.
Veröffentlicht: (2025)
von: Tang, Changli, et al.
Veröffentlicht: (2025)
Finer: Investigating and Enhancing Fine-Grained Visual Concept Recognition in Large Vision Language Models
von: Kim, Jeonghwan, et al.
Veröffentlicht: (2024)
von: Kim, Jeonghwan, et al.
Veröffentlicht: (2024)
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation
von: Tan, Weiting, et al.
Veröffentlicht: (2025)
von: Tan, Weiting, et al.
Veröffentlicht: (2025)
LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning
von: Dai, Yifan, et al.
Veröffentlicht: (2026)
von: Dai, Yifan, et al.
Veröffentlicht: (2026)
Towards Generative Class Prompt Learning for Fine-grained Visual Recognition
von: Chattopadhyay, Soumitri, et al.
Veröffentlicht: (2024)
von: Chattopadhyay, Soumitri, et al.
Veröffentlicht: (2024)
You May Speak Freely: Improving the Fine-Grained Visual Recognition Capabilities of Multimodal Large Language Models with Answer Extraction
von: Lawrence, Logan, et al.
Veröffentlicht: (2025)
von: Lawrence, Logan, et al.
Veröffentlicht: (2025)
Integrating Fine-Grained Audio-Visual Evidence for Robust Multimodal Emotion Reasoning
von: Zhao, Zhixian, et al.
Veröffentlicht: (2026)
von: Zhao, Zhixian, et al.
Veröffentlicht: (2026)
AutoPresent: Designing Structured Visuals from Scratch
von: Ge, Jiaxin, et al.
Veröffentlicht: (2025)
von: Ge, Jiaxin, et al.
Veröffentlicht: (2025)
Exploring the Design Space of Visual Context Representation in Video MLLMs
von: Du, Yifan, et al.
Veröffentlicht: (2024)
von: Du, Yifan, et al.
Veröffentlicht: (2024)
Guiding Medical Vision-Language Models with Explicit Visual Prompts: Framework Design and Comprehensive Exploration of Prompt Variations
von: Zhu, Kangyu, et al.
Veröffentlicht: (2025)
von: Zhu, Kangyu, et al.
Veröffentlicht: (2025)
Evaluating Vision-Language Models for Emotion Recognition
von: Bhattacharyya, Sree, et al.
Veröffentlicht: (2025)
von: Bhattacharyya, Sree, et al.
Veröffentlicht: (2025)
OLKAVS: An Open Large-Scale Korean Audio-Visual Speech Dataset
von: Park, Jeongkyun, et al.
Veröffentlicht: (2023)
von: Park, Jeongkyun, et al.
Veröffentlicht: (2023)
Robust Audiovisual Speech Recognition Models with Mixture-of-Experts
von: Wu, Yihan, et al.
Veröffentlicht: (2024)
von: Wu, Yihan, et al.
Veröffentlicht: (2024)
Zero-Shot Action Recognition in Surveillance Videos
von: Pereira, Joao, et al.
Veröffentlicht: (2024)
von: Pereira, Joao, et al.
Veröffentlicht: (2024)
Design as Desired: Utilizing Visual Question Answering for Multimodal Pre-training
von: Su, Tongkun, et al.
Veröffentlicht: (2024)
von: Su, Tongkun, et al.
Veröffentlicht: (2024)
VISTA: Visual Integrated System for Tailored Automation in Math Problem Generation Using LLM
von: Lee, Jeongwoo, et al.
Veröffentlicht: (2024)
von: Lee, Jeongwoo, et al.
Veröffentlicht: (2024)
video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models
von: Sun, Guangzhi, et al.
Veröffentlicht: (2024)
von: Sun, Guangzhi, et al.
Veröffentlicht: (2024)
VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
von: Cheng, Zesen, et al.
Veröffentlicht: (2024)
von: Cheng, Zesen, et al.
Veröffentlicht: (2024)
AV-Dialog: Spoken Dialogue Models with Audio-Visual Input
von: Chen, Tuochao, et al.
Veröffentlicht: (2025)
von: Chen, Tuochao, et al.
Veröffentlicht: (2025)
Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models
von: Hu, Yushi, et al.
Veröffentlicht: (2024)
von: Hu, Yushi, et al.
Veröffentlicht: (2024)
Visual Representations inside the Language Model
von: Liu, Benlin, et al.
Veröffentlicht: (2025)
von: Liu, Benlin, et al.
Veröffentlicht: (2025)
Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalities
von: Zhou, Ziwei, et al.
Veröffentlicht: (2025)
von: Zhou, Ziwei, et al.
Veröffentlicht: (2025)
GlyphPattern: An Abstract Pattern Recognition Benchmark for Vision-Language Models
von: Wu, Zixuan, et al.
Veröffentlicht: (2024)
von: Wu, Zixuan, et al.
Veröffentlicht: (2024)
Texture or Semantics? Vision-Language Models Get Lost in Font Recognition
von: Li, Zhecheng, et al.
Veröffentlicht: (2025)
von: Li, Zhecheng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
AnnoTheia: A Semi-Automatic Annotation Toolkit for Audio-Visual Speech Technologies
von: Acosta-Triana, José-M., et al.
Veröffentlicht: (2024) -
Comparison of Conventional Hybrid and CTC/Attention Decoders for Continuous Visual Speech Recognition
von: Gimeno-Gómez, David, et al.
Veröffentlicht: (2024) -
Evaluation of End-to-End Continuous Spanish Lipreading in Different Data Conditions
von: Gimeno-Gómez, David, et al.
Veröffentlicht: (2025) -
Unveiling Interpretability in Self-Supervised Speech Representations for Parkinson's Diagnosis
von: Gimeno-Gómez, David, et al.
Veröffentlicht: (2024) -
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition
von: Nguyen, Thai-Binh, et al.
Veröffentlicht: (2025)