Vision-Speech Models: Teaching Speech Models to Converse about Images
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Royer, Amélie, Böhle, Moritz, de Marmiesse, Gabriel, Mazaré, Laurent, Zeghidour, Neil, Défossez, Alexandre, Pérez, Patrick |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
CASA: Cross-Attention over Self-Attention for Efficient Vision-Language Fusion
von: Böhle, Moritz, et al.
Veröffentlicht: (2025)
von: Böhle, Moritz, et al.
Veröffentlicht: (2025)
High-Fidelity Simultaneous Speech-To-Speech Translation
von: Labiausse, Tom, et al.
Veröffentlicht: (2025)
von: Labiausse, Tom, et al.
Veröffentlicht: (2025)
Aligning Spoken Dialogue Models from User Interactions
von: Wu, Anne, et al.
Veröffentlicht: (2025)
von: Wu, Anne, et al.
Veröffentlicht: (2025)
Streaming Sequence-to-Sequence Learning with Delayed Streams Modeling
von: Zeghidour, Neil, et al.
Veröffentlicht: (2025)
von: Zeghidour, Neil, et al.
Veröffentlicht: (2025)
Moshi: a speech-text foundation model for real-time dialogue
von: Défossez, Alexandre, et al.
Veröffentlicht: (2024)
von: Défossez, Alexandre, et al.
Veröffentlicht: (2024)
B-cos Alignment for Inherently Interpretable CNNs and Vision Transformers
von: Böhle, Moritz, et al.
Veröffentlicht: (2023)
von: Böhle, Moritz, et al.
Veröffentlicht: (2023)
Simultaneous Speech-to-Speech Translation Without Aligned Data
von: Labiausse, Tom, et al.
Veröffentlicht: (2026)
von: Labiausse, Tom, et al.
Veröffentlicht: (2026)
Optimising for Interpretability: Convolutional Dynamic Alignment Networks
von: Böhle, Moritz, et al.
Veröffentlicht: (2021)
von: Böhle, Moritz, et al.
Veröffentlicht: (2021)
Towards Better Understanding Attribution Methods
von: Rao, Sukrut, et al.
Veröffentlicht: (2022)
von: Rao, Sukrut, et al.
Veröffentlicht: (2022)
MultiMedEval: A Benchmark and a Toolkit for Evaluating Medical Vision-Language Models
von: Royer, Corentin, et al.
Veröffentlicht: (2024)
von: Royer, Corentin, et al.
Veröffentlicht: (2024)
How to Probe: Simple Yet Effective Techniques for Improving Post-hoc Explanations
von: Gairola, Siddhartha, et al.
Veröffentlicht: (2025)
von: Gairola, Siddhartha, et al.
Veröffentlicht: (2025)
Codebook Transfer with Part-of-Speech for Vector-Quantized Image Modeling
von: Zhang, Baoquan, et al.
Veröffentlicht: (2024)
von: Zhang, Baoquan, et al.
Veröffentlicht: (2024)
Studying How to Efficiently and Effectively Guide Models with Explanations
von: Rao, Sukrut, et al.
Veröffentlicht: (2023)
von: Rao, Sukrut, et al.
Veröffentlicht: (2023)
Better Understanding Differences in Attribution Methods via Systematic Evaluations
von: Rao, Sukrut, et al.
Veröffentlicht: (2023)
von: Rao, Sukrut, et al.
Veröffentlicht: (2023)
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding
von: Zhao, Jiaxing, et al.
Veröffentlicht: (2025)
von: Zhao, Jiaxing, et al.
Veröffentlicht: (2025)
Multimodal Large Language Models for Image, Text, and Speech Data Augmentation: A Survey
von: Sapkota, Ranjan, et al.
Veröffentlicht: (2025)
von: Sapkota, Ranjan, et al.
Veröffentlicht: (2025)
MoshiRAG: Asynchronous Knowledge Retrieval for Full-Duplex Speech Language Models
von: Chien, Chung-Ming, et al.
Veröffentlicht: (2026)
von: Chien, Chung-Ming, et al.
Veröffentlicht: (2026)
Three Pillars improving Vision Foundation Model Distillation for Lidar
von: Puy, Gilles, et al.
Veröffentlicht: (2023)
von: Puy, Gilles, et al.
Veröffentlicht: (2023)
ConvoFusion: Multi-Modal Conversational Diffusion for Co-Speech Gesture Synthesis
von: Mughal, Muhammad Hamza, et al.
Veröffentlicht: (2024)
von: Mughal, Muhammad Hamza, et al.
Veröffentlicht: (2024)
VideoFDB: Evaluating Full-Duplex Vision-Speech Capabilities in Conversational Agents
von: Mazumdar, Amrita, et al.
Veröffentlicht: (2026)
von: Mazumdar, Amrita, et al.
Veröffentlicht: (2026)
Vector search with small radiuses
von: Szilvasy, Gergely, et al.
Veröffentlicht: (2024)
von: Szilvasy, Gergely, et al.
Veröffentlicht: (2024)
InkSight: Offline-to-Online Handwriting Conversion by Teaching Vision-Language Models to Read and Write
von: Mitrevski, Blagoj, et al.
Veröffentlicht: (2024)
von: Mitrevski, Blagoj, et al.
Veröffentlicht: (2024)
LiveGesture Streamable Co-Speech Gesture Generation Model
von: Saleem, Muhammad Usama, et al.
Veröffentlicht: (2026)
von: Saleem, Muhammad Usama, et al.
Veröffentlicht: (2026)
B-cosification: Transforming Deep Neural Networks to be Inherently Interpretable
von: Arya, Shreyash, et al.
Veröffentlicht: (2024)
von: Arya, Shreyash, et al.
Veröffentlicht: (2024)
Discover-then-Name: Task-Agnostic Concept Bottlenecks via Automated Concept Discovery
von: Rao, Sukrut, et al.
Veröffentlicht: (2024)
von: Rao, Sukrut, et al.
Veröffentlicht: (2024)
SViQA: A Unified Speech-Vision Multimodal Model for Textless Visual Question Answering
von: Li, Bingxin
Veröffentlicht: (2025)
von: Li, Bingxin
Veröffentlicht: (2025)
SilVar-Med: A Speech-Driven Visual Language Model for Explainable Abnormality Detection in Medical Imaging
von: Pham, Tan-Hanh, et al.
Veröffentlicht: (2025)
von: Pham, Tan-Hanh, et al.
Veröffentlicht: (2025)
Leveraging WaveNet for Dynamic Listening Head Modeling from Speech
von: Nguyen, Minh-Duc, et al.
Veröffentlicht: (2024)
von: Nguyen, Minh-Duc, et al.
Veröffentlicht: (2024)
video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models
von: Sun, Guangzhi, et al.
Veröffentlicht: (2024)
von: Sun, Guangzhi, et al.
Veröffentlicht: (2024)
An Effective Training Framework for Light-Weight Automatic Speech Recognition Models
von: Hannan, Abdul, et al.
Veröffentlicht: (2025)
von: Hannan, Abdul, et al.
Veröffentlicht: (2025)
SpeechAct: Towards Generating Whole-body Motion from Speech
von: Zhang, Jinsong, et al.
Veröffentlicht: (2023)
von: Zhang, Jinsong, et al.
Veröffentlicht: (2023)
EchoMask: Speech-Queried Attention-based Mask Modeling for Holistic Co-Speech Motion Generation
von: Zhang, Xiangyue, et al.
Veröffentlicht: (2025)
von: Zhang, Xiangyue, et al.
Veröffentlicht: (2025)
SpeechForensics: Audio-Visual Speech Representation Learning for Face Forgery Detection
von: Liang, Yachao, et al.
Veröffentlicht: (2025)
von: Liang, Yachao, et al.
Veröffentlicht: (2025)
Good Teachers Explain: Explanation-Enhanced Knowledge Distillation
von: Parchami-Araghi, Amin, et al.
Veröffentlicht: (2024)
von: Parchami-Araghi, Amin, et al.
Veröffentlicht: (2024)
Unified Speech Recognition: A Single Model for Auditory, Visual, and Audiovisual Inputs
von: Haliassos, Alexandros, et al.
Veröffentlicht: (2024)
von: Haliassos, Alexandros, et al.
Veröffentlicht: (2024)
ARTalk: Speech-Driven 3D Head Animation via Autoregressive Model
von: Chu, Xuangeng, et al.
Veröffentlicht: (2025)
von: Chu, Xuangeng, et al.
Veröffentlicht: (2025)
RadVLM: A Multitask Conversational Vision-Language Model for Radiology
von: Deperrois, Nicolas, et al.
Veröffentlicht: (2025)
von: Deperrois, Nicolas, et al.
Veröffentlicht: (2025)
Visually Grounded Speech Models for Low-resource Languages and Cognitive Modelling
von: Nortje, Leanne
Veröffentlicht: (2024)
von: Nortje, Leanne
Veröffentlicht: (2024)
Speech-to-LaTeX: New Models and Datasets for Converting Spoken Equations and Sentences
von: Korzh, Dmitrii, et al.
Veröffentlicht: (2025)
von: Korzh, Dmitrii, et al.
Veröffentlicht: (2025)
AudioFace: Language-Assisted Speech-Driven Facial Animation with Multimodal Language Models
von: Zheng, Kai, et al.
Veröffentlicht: (2026)
von: Zheng, Kai, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
CASA: Cross-Attention over Self-Attention for Efficient Vision-Language Fusion
von: Böhle, Moritz, et al.
Veröffentlicht: (2025) -
High-Fidelity Simultaneous Speech-To-Speech Translation
von: Labiausse, Tom, et al.
Veröffentlicht: (2025) -
Aligning Spoken Dialogue Models from User Interactions
von: Wu, Anne, et al.
Veröffentlicht: (2025) -
Streaming Sequence-to-Sequence Learning with Delayed Streams Modeling
von: Zeghidour, Neil, et al.
Veröffentlicht: (2025) -
Moshi: a speech-text foundation model for real-time dialogue
von: Défossez, Alexandre, et al.
Veröffentlicht: (2024)