AnnoTheia: A Semi-Automatic Annotation Toolkit for Audio-Visual Speech Technologies
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Acosta-Triana, José-M., Gimeno-Gómez, David, Martínez-Hinarejos, Carlos-D. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Tailored Design of Audio-Visual Speech Recognition Models using Branchformers
von: Gimeno-Gómez, David, et al.
Veröffentlicht: (2024)
von: Gimeno-Gómez, David, et al.
Veröffentlicht: (2024)
Comparison of Conventional Hybrid and CTC/Attention Decoders for Continuous Visual Speech Recognition
von: Gimeno-Gómez, David, et al.
Veröffentlicht: (2024)
von: Gimeno-Gómez, David, et al.
Veröffentlicht: (2024)
Evaluation of End-to-End Continuous Spanish Lipreading in Different Data Conditions
von: Gimeno-Gómez, David, et al.
Veröffentlicht: (2025)
von: Gimeno-Gómez, David, et al.
Veröffentlicht: (2025)
Unveiling Interpretability in Self-Supervised Speech Representations for Parkinson's Diagnosis
von: Gimeno-Gómez, David, et al.
Veröffentlicht: (2024)
von: Gimeno-Gómez, David, et al.
Veröffentlicht: (2024)
Reading Between the Frames: Multi-Modal Depression Detection in Videos from Non-Verbal Cues
von: Gimeno-Gómez, David, et al.
Veröffentlicht: (2024)
von: Gimeno-Gómez, David, et al.
Veröffentlicht: (2024)
Evaluation of Audio-Visual Alignments in Visually Grounded Speech Models
von: Khorrami, Khazar, et al.
Veröffentlicht: (2021)
von: Khorrami, Khazar, et al.
Veröffentlicht: (2021)
ViCocktail: Automated Multi-Modal Data Collection for Vietnamese Audio-Visual Speech Recognition
von: Nguyen, Thai-Binh, et al.
Veröffentlicht: (2025)
von: Nguyen, Thai-Binh, et al.
Veröffentlicht: (2025)
SemiHVision: Enhancing Medical Multimodal Models with a Semi-Human Annotated Dataset and Fine-Tuned Instruction Generation
von: Wang, Junda, et al.
Veröffentlicht: (2024)
von: Wang, Junda, et al.
Veröffentlicht: (2024)
Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation
von: Goncalves, Lucas, et al.
Veröffentlicht: (2024)
von: Goncalves, Lucas, et al.
Veröffentlicht: (2024)
EyeTheia: A Lightweight and Accessible Eye-Tracking Toolbox
von: Pather, Stevenson, et al.
Veröffentlicht: (2026)
von: Pather, Stevenson, et al.
Veröffentlicht: (2026)
SyncVSR: Data-Efficient Visual Speech Recognition with End-to-End Crossmodal Audio Token Synchronization
von: Ahn, Young Jin, et al.
Veröffentlicht: (2024)
von: Ahn, Young Jin, et al.
Veröffentlicht: (2024)
A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations
von: Lan, Tian, et al.
Veröffentlicht: (2025)
von: Lan, Tian, et al.
Veröffentlicht: (2025)
Phoneme-Level Visual Speech Recognition via Point-Visual Fusion and Language Model Reconstruction
von: Teng, Matthew Kit Khinn, et al.
Veröffentlicht: (2025)
von: Teng, Matthew Kit Khinn, et al.
Veröffentlicht: (2025)
Audio Description Generation in the Era of LLMs and VLMs: A Review of Transferable Generative AI Technologies
von: Gao, Yingqiang, et al.
Veröffentlicht: (2024)
von: Gao, Yingqiang, et al.
Veröffentlicht: (2024)
Visually Grounded Speech Models for Low-resource Languages and Cognitive Modelling
von: Nortje, Leanne
Veröffentlicht: (2024)
von: Nortje, Leanne
Veröffentlicht: (2024)
CNVSRC 2023: The First Chinese Continuous Visual Speech Recognition Challenge
von: Chen, Chen, et al.
Veröffentlicht: (2024)
von: Chen, Chen, et al.
Veröffentlicht: (2024)
OLKAVS: An Open Large-Scale Korean Audio-Visual Speech Dataset
von: Park, Jeongkyun, et al.
Veröffentlicht: (2023)
von: Park, Jeongkyun, et al.
Veröffentlicht: (2023)
LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning
von: Dai, Yifan, et al.
Veröffentlicht: (2026)
von: Dai, Yifan, et al.
Veröffentlicht: (2026)
Anno-incomplete Multi-dataset Detection
von: Xu, Yiran, et al.
Veröffentlicht: (2024)
von: Xu, Yiran, et al.
Veröffentlicht: (2024)
Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation
von: Tan, Weiting, et al.
Veröffentlicht: (2025)
von: Tan, Weiting, et al.
Veröffentlicht: (2025)
Vision-Braille: A Curriculum Learning Toolkit and Braille-Chinese Corpus for Braille Translation
von: Wu, Alan, et al.
Veröffentlicht: (2024)
von: Wu, Alan, et al.
Veröffentlicht: (2024)
Integrating Fine-Grained Audio-Visual Evidence for Robust Multimodal Emotion Reasoning
von: Zhao, Zhixian, et al.
Veröffentlicht: (2026)
von: Zhao, Zhixian, et al.
Veröffentlicht: (2026)
video-SALMONN 2: Caption-Enhanced Audio-Visual Large Language Models
von: Tang, Changli, et al.
Veröffentlicht: (2025)
von: Tang, Changli, et al.
Veröffentlicht: (2025)
NYC-Indoor-VPR: A Long-Term Indoor Visual Place Recognition Dataset with Semi-Automatic Annotation
von: Sheng, Diwei, et al.
Veröffentlicht: (2024)
von: Sheng, Diwei, et al.
Veröffentlicht: (2024)
Automatic Layout Planning for Visually-Rich Documents with Instruction-Following Models
von: Zhu, Wanrong, et al.
Veröffentlicht: (2024)
von: Zhu, Wanrong, et al.
Veröffentlicht: (2024)
Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalities
von: Zhou, Ziwei, et al.
Veröffentlicht: (2025)
von: Zhou, Ziwei, et al.
Veröffentlicht: (2025)
Designing Practical Models for Isolated Word Visual Speech Recognition
von: Panagos, Iason Ioannis, et al.
Veröffentlicht: (2025)
von: Panagos, Iason Ioannis, et al.
Veröffentlicht: (2025)
Towards Automatic Evaluation for Image Transcreation
von: Khanuja, Simran, et al.
Veröffentlicht: (2024)
von: Khanuja, Simran, et al.
Veröffentlicht: (2024)
Query-centric Audio-Visual Cognition Network for Moment Retrieval, Segmentation and Step-Captioning
von: Tu, Yunbin, et al.
Veröffentlicht: (2024)
von: Tu, Yunbin, et al.
Veröffentlicht: (2024)
Docling: An Efficient Open-Source Toolkit for AI-driven Document Conversion
von: Livathinos, Nikolaos, et al.
Veröffentlicht: (2025)
von: Livathinos, Nikolaos, et al.
Veröffentlicht: (2025)
DermaBench: A Clinician-Annotated Benchmark Dataset for Dermatology Visual Question Answering and Reasoning
von: Yilmaz, Abdurrahim, et al.
Veröffentlicht: (2026)
von: Yilmaz, Abdurrahim, et al.
Veröffentlicht: (2026)
Theia: Distilling Diverse Vision Foundation Models for Robot Learning
von: Shang, Jinghuan, et al.
Veröffentlicht: (2024)
von: Shang, Jinghuan, et al.
Veröffentlicht: (2024)
Lightweight Operations for Visual Speech Recognition
von: Panagos, Iason Ioannis, et al.
Veröffentlicht: (2025)
von: Panagos, Iason Ioannis, et al.
Veröffentlicht: (2025)
A SAM based Tool for Semi-Automatic Food Annotation
von: Rahman, Lubnaa Abdur, et al.
Veröffentlicht: (2024)
von: Rahman, Lubnaa Abdur, et al.
Veröffentlicht: (2024)
Disability Representations: Finding Biases in Automatic Image Generation
von: Tevissen, Yannis
Veröffentlicht: (2024)
von: Tevissen, Yannis
Veröffentlicht: (2024)
Knowledge-Aware Reasoning over Multimodal Semi-structured Tables
von: Mathur, Suyash Vardhan, et al.
Veröffentlicht: (2024)
von: Mathur, Suyash Vardhan, et al.
Veröffentlicht: (2024)
Improving Text-to-Image Consistency via Automatic Prompt Optimization
von: Mañas, Oscar, et al.
Veröffentlicht: (2024)
von: Mañas, Oscar, et al.
Veröffentlicht: (2024)
VideoCoT: A Video Chain-of-Thought Dataset with Active Annotation Tool
von: Wang, Yan, et al.
Veröffentlicht: (2024)
von: Wang, Yan, et al.
Veröffentlicht: (2024)
Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions
von: Shen, Yijun, et al.
Veröffentlicht: (2025)
von: Shen, Yijun, et al.
Veröffentlicht: (2025)
AV-Dialog: Spoken Dialogue Models with Audio-Visual Input
von: Chen, Tuochao, et al.
Veröffentlicht: (2025)
von: Chen, Tuochao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Tailored Design of Audio-Visual Speech Recognition Models using Branchformers
von: Gimeno-Gómez, David, et al.
Veröffentlicht: (2024) -
Comparison of Conventional Hybrid and CTC/Attention Decoders for Continuous Visual Speech Recognition
von: Gimeno-Gómez, David, et al.
Veröffentlicht: (2024) -
Evaluation of End-to-End Continuous Spanish Lipreading in Different Data Conditions
von: Gimeno-Gómez, David, et al.
Veröffentlicht: (2025) -
Unveiling Interpretability in Self-Supervised Speech Representations for Parkinson's Diagnosis
von: Gimeno-Gómez, David, et al.
Veröffentlicht: (2024) -
Reading Between the Frames: Multi-Modal Depression Detection in Videos from Non-Verbal Cues
von: Gimeno-Gómez, David, et al.
Veröffentlicht: (2024)