Learning to Unify Audio, Visual and Text for Audio-Enhanced Multilingual Visual Answer Localization
Fuente:
arXiv
Saved in:
| Main Authors: | Wen, Zhibin, Li, Bin |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Node-Based Editing for Multimodal Generation of Text, Audio, Image, and Video
by: Kyaw, Alexander Htet, et al.
Published: (2025)
by: Kyaw, Alexander Htet, et al.
Published: (2025)
Emotion-Driven Personalized Recommendation for AI-Generated Content Using Multi-Modal Sentiment and Intent Analysis
by: Hu, Zheqi, et al.
Published: (2025)
by: Hu, Zheqi, et al.
Published: (2025)
Do LLMs Provide Consistent Answers to Health-Related Questions across Languages?
by: Schlicht, Ipek Baris, et al.
Published: (2025)
by: Schlicht, Ipek Baris, et al.
Published: (2025)
InfoCIR: Multimedia Analysis for Composed Image Retrieval
by: Dravilas, Ioannis, et al.
Published: (2026)
by: Dravilas, Ioannis, et al.
Published: (2026)
Muse-it: A Tool for Analyzing Music Discourse on Reddit
by: Agarwala, Jatin, et al.
Published: (2025)
by: Agarwala, Jatin, et al.
Published: (2025)
Generative AI for Video Trailer Synthesis: From Extractive Heuristics to Autoregressive Creativity
by: Dharmaratnakar, Abhishek, et al.
Published: (2026)
by: Dharmaratnakar, Abhishek, et al.
Published: (2026)
Towards Unified Multi-Modal Personalization: Large Vision-Language Models for Generative Recommendation and Beyond
by: Wei, Tianxin, et al.
Published: (2024)
by: Wei, Tianxin, et al.
Published: (2024)
WildVis: Open Source Visualizer for Million-Scale Chat Logs in the Wild
by: Deng, Yuntian, et al.
Published: (2024)
by: Deng, Yuntian, et al.
Published: (2024)
User Perception of Attention Visualizations: Effects on Interpretability Across Evidence-Based Medical Documents
by: Carvallo, Andrés, et al.
Published: (2025)
by: Carvallo, Andrés, et al.
Published: (2025)
NExT-Search: Rebuilding User Feedback Ecosystem for Generative AI Search
by: Dai, Sunhao, et al.
Published: (2025)
by: Dai, Sunhao, et al.
Published: (2025)
Visualization for Recommendation Explainability: A Survey and New Perspectives
by: Chatti, Mohamed Amine, et al.
Published: (2023)
by: Chatti, Mohamed Amine, et al.
Published: (2023)
Improving Multi-Domain Task-Oriented Dialogue System with Offline Reinforcement Learning
by: Prajapat, Dharmendra, et al.
Published: (2024)
by: Prajapat, Dharmendra, et al.
Published: (2024)
From Surface Learning to Deep Understanding: A Grounded AI Tutoring System for Moodle
by: Ostrowska, Anna, et al.
Published: (2026)
by: Ostrowska, Anna, et al.
Published: (2026)
ClinicalTrialsHub: Bridging Registries and Literature for Comprehensive Clinical Trial Access
by: Park, Jiwoo, et al.
Published: (2025)
by: Park, Jiwoo, et al.
Published: (2025)
Development and Evaluation of Dental Image Exchange and Management System: A User-Centered Perspective
by: Rahimi, B, et al.
Published: (2022)
by: Rahimi, B, et al.
Published: (2022)
Chasing RATs: Tracing Reading for and as Creative Activity
by: Liu, Sophia, et al.
Published: (2026)
by: Liu, Sophia, et al.
Published: (2026)
Towards Context-Aware Adaptation in Extended Reality: A Design Space for XR Interfaces and an Adaptive Placement Strategy
by: Davari, Shakiba, et al.
Published: (2024)
by: Davari, Shakiba, et al.
Published: (2024)
MetaDesigner: Advancing Artistic Typography Through AI-Driven, User-Centric, and Multilingual WordArt Synthesis
by: He, Jun-Yan, et al.
Published: (2024)
by: He, Jun-Yan, et al.
Published: (2024)
LAVE: LLM-Powered Agent Assistance and Language Augmentation for Video Editing
by: Wang, Bryan, et al.
Published: (2024)
by: Wang, Bryan, et al.
Published: (2024)
Task-Oriented Dialog Systems for the Senegalese Wolof Language
by: Mbaye, Derguene, et al.
Published: (2024)
by: Mbaye, Derguene, et al.
Published: (2024)
Leveraging Large Language Models for Hybrid Workplace Decision Support
by: Kim, Yujin, et al.
Published: (2024)
by: Kim, Yujin, et al.
Published: (2024)
Retail-GPT: leveraging Retrieval Augmented Generation (RAG) for building E-commerce Chat Assistants
by: de Freitas, Bruno Amaral Teixeira, et al.
Published: (2024)
by: de Freitas, Bruno Amaral Teixeira, et al.
Published: (2024)
Detecting Deceptive Dark Patterns in E-commerce Platforms
by: Ramteke, Arya, et al.
Published: (2024)
by: Ramteke, Arya, et al.
Published: (2024)
Language Modelling Approaches to Adaptive Machine Translation
by: Moslem, Yasmin
Published: (2024)
by: Moslem, Yasmin
Published: (2024)
What should I wear to a party in a Greek taverna? Evaluation for Conversational Agents in the Fashion Domain
by: Maronikolakis, Antonis, et al.
Published: (2024)
by: Maronikolakis, Antonis, et al.
Published: (2024)
Rewriting Conversational Utterances with Instructed Large Language Models
by: Galimzhanova, Elnara, et al.
Published: (2024)
by: Galimzhanova, Elnara, et al.
Published: (2024)
Agentic Chunking and Bayesian De-chunking of AI Generated Fuzzy Cognitive Maps: A Model of the Thucydides Trap
by: Panda, Akash Kumar, et al.
Published: (2026)
by: Panda, Akash Kumar, et al.
Published: (2026)
Semantic Interaction for Narrative Map Sensemaking: An Insight-based Evaluation
by: Keith-Norambuena, Brian Felipe, et al.
Published: (2026)
by: Keith-Norambuena, Brian Felipe, et al.
Published: (2026)
Much of Geospatial Web Search Is Beyond Traditional GIS
by: Ilyankou, Ilya, et al.
Published: (2026)
by: Ilyankou, Ilya, et al.
Published: (2026)
Modelling and Classifying the Components of a Literature Review
by: Bolaños, Francisco, et al.
Published: (2025)
by: Bolaños, Francisco, et al.
Published: (2025)
EHR-MCP: Real-world Evaluation of Clinical Information Retrieval by Large Language Models via Model Context Protocol
by: Masayoshi, Kanato, et al.
Published: (2025)
by: Masayoshi, Kanato, et al.
Published: (2025)
SteerEval: A Framework for Evaluating Steerability with Natural Language Profiles for Recommendation
by: Zhou, Joyce, et al.
Published: (2026)
by: Zhou, Joyce, et al.
Published: (2026)
Causal Autoencoder-like Generation of Feedback Fuzzy Cognitive Maps with an LLM Agent
by: Panda, Akash Kumar, et al.
Published: (2025)
by: Panda, Akash Kumar, et al.
Published: (2025)
The Agentic Leash: Extracting Causal Feedback Fuzzy Cognitive Maps with LLMs
by: Panda, Akash Kumar, et al.
Published: (2025)
by: Panda, Akash Kumar, et al.
Published: (2025)
Are the confidence scores of reviewers consistent with the review content? Evidence from top conference proceedings in AI
by: Wu, Wenqing, et al.
Published: (2025)
by: Wu, Wenqing, et al.
Published: (2025)
Surfacing Isolated Learners with Outcome-Independent Mediation of Feedback between Teachers and Students Using AI
by: Park, Junsoo, et al.
Published: (2026)
by: Park, Junsoo, et al.
Published: (2026)
Counterfactual Reasoning Using Predicted Latent Personality Dimensions for Optimizing Persuasion Outcome
by: Zeng, Donghuo, et al.
Published: (2024)
by: Zeng, Donghuo, et al.
Published: (2024)
DailyLLM: Context-Aware Activity Log Generation Using Multi-Modal Sensors and LLMs
by: Tian, Ye, et al.
Published: (2025)
by: Tian, Ye, et al.
Published: (2025)
Reverse Region-to-Entity Annotation for Pixel-Level Visual Entity Linking
by: Xu, Zhengfei, et al.
Published: (2024)
by: Xu, Zhengfei, et al.
Published: (2024)
Three Modalities, Two Design Probes, One Prototype, and No Vision: Experience-Based Co-Design of a Multi-modal 3D Data Visualization Tool
by: Kamath, Sanchita S., et al.
Published: (2026)
by: Kamath, Sanchita S., et al.
Published: (2026)
Similar Items
-
Node-Based Editing for Multimodal Generation of Text, Audio, Image, and Video
by: Kyaw, Alexander Htet, et al.
Published: (2025) -
Emotion-Driven Personalized Recommendation for AI-Generated Content Using Multi-Modal Sentiment and Intent Analysis
by: Hu, Zheqi, et al.
Published: (2025) -
Do LLMs Provide Consistent Answers to Health-Related Questions across Languages?
by: Schlicht, Ipek Baris, et al.
Published: (2025) -
InfoCIR: Multimedia Analysis for Composed Image Retrieval
by: Dravilas, Ioannis, et al.
Published: (2026) -
Muse-it: A Tool for Analyzing Music Discourse on Reddit
by: Agarwala, Jatin, et al.
Published: (2025)