Real-Time Generation of Game Video Commentary with Multimodal LLMs: Pause-Aware Decoding Approaches
Fuente:
arXiv
Saved in:
| Main Authors: | Afzal, Anum, Saito, Yuki, Takamura, Hiroya, Sudoh, Katsuhito, Takamichi, Shinnosuke, Neubig, Graham, Matthes, Florian, Ishigaki, Tatsuya |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Multimodal Task Interference: A Benchmark and Analysis of History-Target Mismatch in Multimodal LLMs
by: Kawarada, Masayuki, et al.
Published: (2026)
by: Kawarada, Masayuki, et al.
Published: (2026)
Prompting for Numerical Sequences: A Case Study on Market Comment Generation
by: Kawarada, Masayuki, et al.
Published: (2024)
by: Kawarada, Masayuki, et al.
Published: (2024)
Can Smaller LLMs do better? Unlocking Cross-Domain Potential through Parameter-Efficient Fine-Tuning for Text Summarization
by: Afzal, Anum, et al.
Published: (2025)
by: Afzal, Anum, et al.
Published: (2025)
LLMs Are Zero-Shot Context-Aware Simultaneous Translators
by: Koshkin, Roman, et al.
Published: (2024)
by: Koshkin, Roman, et al.
Published: (2024)
FActBench: A Benchmark for Fine-grained Automatic Evaluation of LLM-Generated Text in the Medical Domain
by: Afzal, Anum, et al.
Published: (2025)
by: Afzal, Anum, et al.
Published: (2025)
JaccDiv: A Metric and Benchmark for Quantifying Diversity of Generated Marketing Text in the Music Industry
by: Afzal, Anum, et al.
Published: (2025)
by: Afzal, Anum, et al.
Published: (2025)
Towards Optimizing and Evaluating a Retrieval Augmented QA Chatbot using LLMs with Human in the Loop
by: Afzal, Anum, et al.
Published: (2024)
by: Afzal, Anum, et al.
Published: (2024)
A Comparative Study of Demonstration Selection for Practical Large Language Models-based Next POI Prediction
by: Nishida, Ryo, et al.
Published: (2026)
by: Nishida, Ryo, et al.
Published: (2026)
Knowing Before Saying: LLM Representations Encode Information About Chain-of-Thought Success Before Completion
by: Afzal, Anum, et al.
Published: (2025)
by: Afzal, Anum, et al.
Published: (2025)
AdaptEval: Evaluating Large Language Models on Domain Adaptation for Text Summarization
by: Afzal, Anum, et al.
Published: (2024)
by: Afzal, Anum, et al.
Published: (2024)
Spatial-CLAP: Learning Spatially-Aware audio--text Embeddings for Multi-Source Conditions
by: Seki, Kentaro, et al.
Published: (2025)
by: Seki, Kentaro, et al.
Published: (2025)
Top-down string-to-dependency Neural Machine Translation
by: Kondo, Shuhei, et al.
Published: (2026)
by: Kondo, Shuhei, et al.
Published: (2026)
Automated Essay Scoring Using Grammatical Variety and Errors with Multi-Task Learning and Item Response Theory
by: Doi, Kosuke, et al.
Published: (2024)
by: Doi, Kosuke, et al.
Published: (2024)
TransLLaMa: LLM-based Simultaneous Translation System
by: Koshkin, Roman, et al.
Published: (2024)
by: Koshkin, Roman, et al.
Published: (2024)
Towards Optimizing a Retrieval Augmented Generation using Large Language Model on Academic Data
by: Afzal, Anum, et al.
Published: (2024)
by: Afzal, Anum, et al.
Published: (2024)
HOME-KGQA: A Benchmark Dataset for Multimodal Knowledge Graph Question Answering on Household Daily Activities
by: Egami, Shusaku, et al.
Published: (2026)
by: Egami, Shusaku, et al.
Published: (2026)
RELATE: Subjective evaluation dataset for automatic evaluation of relevance between text and audio
by: Kanamori, Yusuke, et al.
Published: (2025)
by: Kanamori, Yusuke, et al.
Published: (2025)
Subspace Representations for Soft Set Operations and Sentence Similarities
by: Ishibashi, Yoichi, et al.
Published: (2022)
by: Ishibashi, Yoichi, et al.
Published: (2022)
An Automatic Quality Metric for Evaluating Simultaneous Interpretation
by: Makinae, Mana, et al.
Published: (2024)
by: Makinae, Mana, et al.
Published: (2024)
Spatial Voice Conversion: Voice Conversion Preserving Spatial Information and Non-target Signals
by: Seki, Kentaro, et al.
Published: (2024)
by: Seki, Kentaro, et al.
Published: (2024)
Building speech corpus with diverse voice characteristics for its prompt-based representation
by: Watanabe, Aya, et al.
Published: (2024)
by: Watanabe, Aya, et al.
Published: (2024)
J-CHAT: Japanese Large-scale Spoken Dialogue Corpus for Spoken Dialogue Language Modeling
by: Nakata, Wataru, et al.
Published: (2024)
by: Nakata, Wataru, et al.
Published: (2024)
JVNV: A Corpus of Japanese Emotional Speech with Verbal Content and Nonverbal Expressions
by: Xin, Detai, et al.
Published: (2023)
by: Xin, Detai, et al.
Published: (2023)
Word Order in English-Japanese Simultaneous Interpretation: Analyses and Evaluation using Chunk-wise Monotonic Translation
by: Doi, Kosuke, et al.
Published: (2024)
by: Doi, Kosuke, et al.
Published: (2024)
Analysing the Language of Neural Audio Codecs
by: Park, Joonyong, et al.
Published: (2025)
by: Park, Joonyong, et al.
Published: (2025)
QuantumBench: A Benchmark for Quantum Problem Solving
by: Minami, Shunya, et al.
Published: (2025)
by: Minami, Shunya, et al.
Published: (2025)
Noise-Robust Voice Conversion by Conditional Denoising Training Using Latent Variables of Recording Quality and Environment
by: Igarashi, Takuto, et al.
Published: (2024)
by: Igarashi, Takuto, et al.
Published: (2024)
SRC4VC: Smartphone-Recorded Corpus for Voice Conversion Benchmark
by: Saito, Yuki, et al.
Published: (2024)
by: Saito, Yuki, et al.
Published: (2024)
AudioBERTScore: Objective Evaluation of Environmental Sound Synthesis Based on Similarity of Audio embedding Sequences
by: Kishi, Minoru, et al.
Published: (2025)
by: Kishi, Minoru, et al.
Published: (2025)
Who Laughs with Whom? Disentangling Influential Factors in Humor Preferences across User Clusters and LLMs
by: Murakami, Soichiro, et al.
Published: (2026)
by: Murakami, Soichiro, et al.
Published: (2026)
Agent AI for Finance
by: Chen, Chung-Chi, et al.
Published: (2025)
by: Chen, Chung-Chi, et al.
Published: (2025)
Voice Conversion for Likability Control via Automated Rating of Speech Synthesis Corpora
by: Suda, Hitoshi, et al.
Published: (2025)
by: Suda, Hitoshi, et al.
Published: (2025)
Who Finds This Voice Attractive? A Large-Scale Experiment Using In-the-Wild Data
by: Suda, Hitoshi, et al.
Published: (2024)
by: Suda, Hitoshi, et al.
Published: (2024)
Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models
by: Kando, Shunsuke, et al.
Published: (2025)
by: Kando, Shunsuke, et al.
Published: (2025)
LLMs for Legal Subsumption in German Employment Contracts
by: Wardas, Oliver, et al.
Published: (2025)
by: Wardas, Oliver, et al.
Published: (2025)
Grounding Multilingual Multimodal LLMs With Cultural Knowledge
by: Nyandwi, Jean de Dieu, et al.
Published: (2025)
by: Nyandwi, Jean de Dieu, et al.
Published: (2025)
Reading Comprehension using Entity-based Memory Network
by: Wang, Xun, et al.
Published: (2016)
by: Wang, Xun, et al.
Published: (2016)
NAIST-SIC-Aligned: an Aligned English-Japanese Simultaneous Interpretation Corpus
by: Zhao, Jinming, et al.
Published: (2023)
by: Zhao, Jinming, et al.
Published: (2023)
SpeechBERTScore: Reference-Aware Automatic Evaluation of Speech Generation Leveraging NLP Evaluation Metrics
by: Saeki, Takaaki, et al.
Published: (2024)
by: Saeki, Takaaki, et al.
Published: (2024)
Improving Health Question Answering with Reliable and Time-Aware Evidence Retrieval
by: Vladika, Juraj, et al.
Published: (2024)
by: Vladika, Juraj, et al.
Published: (2024)
Similar Items
-
Multimodal Task Interference: A Benchmark and Analysis of History-Target Mismatch in Multimodal LLMs
by: Kawarada, Masayuki, et al.
Published: (2026) -
Prompting for Numerical Sequences: A Case Study on Market Comment Generation
by: Kawarada, Masayuki, et al.
Published: (2024) -
Can Smaller LLMs do better? Unlocking Cross-Domain Potential through Parameter-Efficient Fine-Tuning for Text Summarization
by: Afzal, Anum, et al.
Published: (2025) -
LLMs Are Zero-Shot Context-Aware Simultaneous Translators
by: Koshkin, Roman, et al.
Published: (2024) -
FActBench: A Benchmark for Fine-grained Automatic Evaluation of LLM-Generated Text in the Medical Domain
by: Afzal, Anum, et al.
Published: (2025)