Exploring the Potential of Multimodal LLM with Knowledge-Intensive Multimodal ASR
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Minghan, Wang, Yuxia, Vu, Thuy-Trang, Shareghi, Ehsan, Haffari, Gholamreza |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Conversational SimulMT: Efficient Simultaneous Translation with Large Language Models
by: Wang, Minghan, et al.
Published: (2024)
by: Wang, Minghan, et al.
Published: (2024)
Towards Inference-time Scaling for Continuous Space Reasoning
by: Wang, Minghan, et al.
Published: (2025)
by: Wang, Minghan, et al.
Published: (2025)
SpeechDialogueFactory: Generating High-Quality Speech Dialogue Data to Accelerate Your Speech-LLM Development
by: Wang, Minghan, et al.
Published: (2025)
by: Wang, Minghan, et al.
Published: (2025)
Discrete Minds in a Continuous World: Do Language Models Know Time Passes?
by: Wang, Minghan, et al.
Published: (2025)
by: Wang, Minghan, et al.
Published: (2025)
GTS: Inference-Time Scaling of Latent Reasoning with a Learnable Gaussian Thought Sampler
by: Wang, Minghan, et al.
Published: (2026)
by: Wang, Minghan, et al.
Published: (2026)
Simultaneous Machine Translation with Large Language Models
by: Wang, Minghan, et al.
Published: (2023)
by: Wang, Minghan, et al.
Published: (2023)
Audio Is the Achilles' Heel: Red Teaming Audio Large Multimodal Models
by: Yang, Hao, et al.
Published: (2024)
by: Yang, Hao, et al.
Published: (2024)
Towards Probing Speech-Specific Risks in Large Multimodal Models: A Taxonomy, Benchmark, and Insights
by: Yang, Hao, et al.
Published: (2024)
by: Yang, Hao, et al.
Published: (2024)
Beyond Imitation: Recovering Dense Rewards from Demonstrations
by: Li, Jiangnan, et al.
Published: (2025)
by: Li, Jiangnan, et al.
Published: (2025)
SituatedThinker: Grounding LLM Reasoning with Real-World through Situated Thinking
by: Liu, Junnan, et al.
Published: (2025)
by: Liu, Junnan, et al.
Published: (2025)
The Best of Both Worlds: Bridging Quality and Diversity in Data Selection with Bipartite Graph
by: Wu, Minghao, et al.
Published: (2024)
by: Wu, Minghao, et al.
Published: (2024)
Mixture-of-Skills: Learning to Optimize Data Usage for Fine-Tuning Large Language Models
by: Wu, Minghao, et al.
Published: (2024)
by: Wu, Minghao, et al.
Published: (2024)
Resurfacing Paralinguistic Awareness in Large Audio Language Models
by: Yang, Hao, et al.
Published: (2026)
by: Yang, Hao, et al.
Published: (2026)
Jigsaw Puzzles: Splitting Harmful Questions to Jailbreak Large Language Models
by: Yang, Hao, et al.
Published: (2024)
by: Yang, Hao, et al.
Published: (2024)
Active Continual Learning: On Balancing Knowledge Retention and Learnability
by: Vu, Thuy-Trang, et al.
Published: (2023)
by: Vu, Thuy-Trang, et al.
Published: (2023)
AIPO: Learning to Reason from Active Interaction
by: Liu, Junnan, et al.
Published: (2026)
by: Liu, Junnan, et al.
Published: (2026)
Reshaping Representation Space to Balance the Safety and Over-rejection in Large Audio Language Models
by: Yang, Hao, et al.
Published: (2025)
by: Yang, Hao, et al.
Published: (2025)
Adapting Large Language Models for Document-Level Machine Translation
by: Wu, Minghao, et al.
Published: (2024)
by: Wu, Minghao, et al.
Published: (2024)
MAPLE: Multi-Agent Adaptive Planning with Long-Term Memory for Table Reasoning
by: Bai, Ye, et al.
Published: (2025)
by: Bai, Ye, et al.
Published: (2025)
Extending LLMs to New Languages: A Case Study of Llama and Persian Adaptation
by: Sani, Samin Mahdizadeh, et al.
Published: (2024)
by: Sani, Samin Mahdizadeh, et al.
Published: (2024)
SCAR: Data Selection via Style Consistency-Aware Response Ranking for Efficient Instruction-Tuning of Large Language Models
by: Li, Zhuang, et al.
Published: (2024)
by: Li, Zhuang, et al.
Published: (2024)
CONGRAD:Conflicting Gradient Filtering for Multilingual Preference Alignment
by: Li, Jiangnan, et al.
Published: (2025)
by: Li, Jiangnan, et al.
Published: (2025)
Direct Evaluation of Chain-of-Thought in Multi-hop Reasoning with Knowledge Graphs
by: Nguyen, Minh-Vuong, et al.
Published: (2024)
by: Nguyen, Minh-Vuong, et al.
Published: (2024)
Continual Learning for Large Language Models: A Survey
by: Wu, Tongtong, et al.
Published: (2024)
by: Wu, Tongtong, et al.
Published: (2024)
Proverbs Run in Pairs: Evaluating Proverb Translation Capability of Large Language Model
by: Wang, Minghan, et al.
Published: (2025)
by: Wang, Minghan, et al.
Published: (2025)
Discourse Graph Guided Document Translation with Large Language Models
by: Pham, Viet-Thanh, et al.
Published: (2025)
by: Pham, Viet-Thanh, et al.
Published: (2025)
Rethinking STS and NLI in Large Language Models
by: Wang, Yuxia, et al.
Published: (2023)
by: Wang, Yuxia, et al.
Published: (2023)
Planning for Success: Exploring LLM Long-term Planning Capabilities in Table Understanding
by: Nguyen, Thi-Nhung, et al.
Published: (2025)
by: Nguyen, Thi-Nhung, et al.
Published: (2025)
Equipping Language Models with Tool Use Capability for Tabular Data Analysis in Finance
by: Theuma, Adrian, et al.
Published: (2024)
by: Theuma, Adrian, et al.
Published: (2024)
Modelling Political Coalition Negotiations Using LLM-based Agents
by: Moghimifar, Farhad, et al.
Published: (2024)
by: Moghimifar, Farhad, et al.
Published: (2024)
Cube Bench: A Benchmark for Spatial Visual Reasoning in MLLMs
by: Anand, Dhruv, et al.
Published: (2025)
by: Anand, Dhruv, et al.
Published: (2025)
One STEP at a time: Language Agents are Stepwise Planners
by: Nguyen, Minh, et al.
Published: (2024)
by: Nguyen, Minh, et al.
Published: (2024)
Evaluating LLM-based Approaches to Legal Citation Prediction: Domain-specific Pre-training, Fine-tuning, or RAG? A Benchmark and an Australian Law Case Study
by: Han, Jiuzhou, et al.
Published: (2024)
by: Han, Jiuzhou, et al.
Published: (2024)
Improving Cross-Domain Low-Resource Text Generation through LLM Post-Editing: A Programmer-Interpreter Approach
by: Li, Zhuang, et al.
Published: (2024)
by: Li, Zhuang, et al.
Published: (2024)
Towards Uncertainty-Aware Language Agent
by: Han, Jiuzhou, et al.
Published: (2024)
by: Han, Jiuzhou, et al.
Published: (2024)
Reward Engineering for Generating Semi-structured Explanation
by: Han, Jiuzhou, et al.
Published: (2023)
by: Han, Jiuzhou, et al.
Published: (2023)
TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law
by: Hui, Zheng, et al.
Published: (2025)
by: Hui, Zheng, et al.
Published: (2025)
Can Machines Resonate with Humans? Evaluating the Emotional and Empathic Comprehension of LMs
by: Manzoor, Muhammad Arslan, et al.
Published: (2024)
by: Manzoor, Muhammad Arslan, et al.
Published: (2024)
Multimodal Reranking for Knowledge-Intensive Visual Question Answering
by: Wen, Haoyang, et al.
Published: (2024)
by: Wen, Haoyang, et al.
Published: (2024)
Importance-Aware Data Augmentation for Document-Level Neural Machine Translation
by: Wu, Minghao, et al.
Published: (2024)
by: Wu, Minghao, et al.
Published: (2024)
Similar Items
-
Conversational SimulMT: Efficient Simultaneous Translation with Large Language Models
by: Wang, Minghan, et al.
Published: (2024) -
Towards Inference-time Scaling for Continuous Space Reasoning
by: Wang, Minghan, et al.
Published: (2025) -
SpeechDialogueFactory: Generating High-Quality Speech Dialogue Data to Accelerate Your Speech-LLM Development
by: Wang, Minghan, et al.
Published: (2025) -
Discrete Minds in a Continuous World: Do Language Models Know Time Passes?
by: Wang, Minghan, et al.
Published: (2025) -
GTS: Inference-Time Scaling of Latent Reasoning with a Learnable Gaussian Thought Sampler
by: Wang, Minghan, et al.
Published: (2026)