Salvato in:
| Autori principali: | Jang, Jihyoung, Bae, Minwook, Kim, Minji, Hakkani-Tur, Dilek, Kim, Hyounghun |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2506.00421 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
AQuA: Toward Strategic Response Generation for Ambiguous Visual Questions
di: Jang, Jihyoung, et al.
Pubblicazione: (2026)
di: Jang, Jihyoung, et al.
Pubblicazione: (2026)
MIRAGE: A Benchmark for Multimodal Information-Seeking and Reasoning in Agricultural Expert-Guided Conversations
di: Dongre, Vardhan, et al.
Pubblicazione: (2025)
di: Dongre, Vardhan, et al.
Pubblicazione: (2025)
Mixed-Session Conversation with Egocentric Memory
di: Jang, Jihyoung, et al.
Pubblicazione: (2024)
di: Jang, Jihyoung, et al.
Pubblicazione: (2024)
Collective Critics for Creative Story Generation
di: Bae, Minwook, et al.
Pubblicazione: (2024)
di: Bae, Minwook, et al.
Pubblicazione: (2024)
Visually Guided Decoding: Gradient-Free Hard Prompt Inversion with Language Models
di: Kim, Donghoon, et al.
Pubblicazione: (2025)
di: Kim, Donghoon, et al.
Pubblicazione: (2025)
ChatEXAONEPath: An Expert-level Multimodal Large Language Model for Histopathology Using Whole Slide Images
di: Kim, Sangwook, et al.
Pubblicazione: (2025)
di: Kim, Sangwook, et al.
Pubblicazione: (2025)
Towards Conversational Medical AI with Eyes, Ears and a Voice
di: Shah, Meet, et al.
Pubblicazione: (2026)
di: Shah, Meet, et al.
Pubblicazione: (2026)
Look & Mark: Leveraging Radiologist Eye Fixations and Bounding boxes in Multimodal Large Language Models for Chest X-ray Report Generation
di: Kim, Yunsoo, et al.
Pubblicazione: (2025)
di: Kim, Yunsoo, et al.
Pubblicazione: (2025)
Simulating User Agents for Embodied Conversational-AI
di: Philipov, Daniel, et al.
Pubblicazione: (2024)
di: Philipov, Daniel, et al.
Pubblicazione: (2024)
Leveraging the Power of MLLMs for Gloss-Free Sign Language Translation
di: Kim, Jungeun, et al.
Pubblicazione: (2024)
di: Kim, Jungeun, et al.
Pubblicazione: (2024)
ESREAL: Exploiting Semantic Reconstruction to Mitigate Hallucinations in Vision-Language Models
di: Kim, Minchan, et al.
Pubblicazione: (2024)
di: Kim, Minchan, et al.
Pubblicazione: (2024)
Goal Alignment in LLM-Based User Simulators for Conversational AI
di: Mehri, Shuhaib, et al.
Pubblicazione: (2025)
di: Mehri, Shuhaib, et al.
Pubblicazione: (2025)
Know Your Mistakes: Towards Preventing Overreliance on Task-Oriented Conversational AI Through Accountability Modeling
di: Dey, Suvodip, et al.
Pubblicazione: (2025)
di: Dey, Suvodip, et al.
Pubblicazione: (2025)
Question Generation for Assessing Early Literacy Reading Comprehension
di: Yang, Xiaocheng, et al.
Pubblicazione: (2025)
di: Yang, Xiaocheng, et al.
Pubblicazione: (2025)
ReVision: A Dataset and Baseline VLM for Privacy-Preserving Task-Oriented Visual Instruction Rewriting
di: Mishra, Abhijit, et al.
Pubblicazione: (2025)
di: Mishra, Abhijit, et al.
Pubblicazione: (2025)
Jailbreaking Multimodal Large Language Models using Multi-Clip Video
di: Kang, Choongwon, et al.
Pubblicazione: (2026)
di: Kang, Choongwon, et al.
Pubblicazione: (2026)
Revealing the Inherent Instructability of Pre-Trained Language Models
di: An, Seokhyun, et al.
Pubblicazione: (2024)
di: An, Seokhyun, et al.
Pubblicazione: (2024)
OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality Interaction
di: Zhang, Haonan, et al.
Pubblicazione: (2025)
di: Zhang, Haonan, et al.
Pubblicazione: (2025)
v1: Learning to Point Visual Tokens for Multimodal Grounded Reasoning
di: Chung, Jiwan, et al.
Pubblicazione: (2025)
di: Chung, Jiwan, et al.
Pubblicazione: (2025)
Enhancing Human-Computer Interaction in Chest X-ray Analysis using Vision and Language Model with Eye Gaze Patterns
di: Kim, Yunsoo, et al.
Pubblicazione: (2024)
di: Kim, Yunsoo, et al.
Pubblicazione: (2024)
LITTA: Late-Interaction and Test-Time Alignment for Visually-Grounded Multimodal Retrieval
di: Kim, Seonok
Pubblicazione: (2026)
di: Kim, Seonok
Pubblicazione: (2026)
CollEX -- A Multimodal Agentic RAG System Enabling Interactive Exploration of Scientific Collections
di: Schneider, Florian, et al.
Pubblicazione: (2025)
di: Schneider, Florian, et al.
Pubblicazione: (2025)
SpecEyes: Accelerating Agentic Multimodal LLMs via Speculative Perception and Planning
di: Huang, Haoyu, et al.
Pubblicazione: (2026)
di: Huang, Haoyu, et al.
Pubblicazione: (2026)
Personalized Scientific Figure Caption Generation: An Empirical Study on Author-Specific Writing Style Transfer
di: Kim, Jaeyoung, et al.
Pubblicazione: (2025)
di: Kim, Jaeyoung, et al.
Pubblicazione: (2025)
Do LLMs Encode Functional Importance of Reasoning Tokens?
di: Singh, Janvijay, et al.
Pubblicazione: (2026)
di: Singh, Janvijay, et al.
Pubblicazione: (2026)
Neural Networks for Learnable and Scalable Influence Estimation of Instruction Fine-Tuning Data
di: Agarwal, Ishika, et al.
Pubblicazione: (2025)
di: Agarwal, Ishika, et al.
Pubblicazione: (2025)
Embodied Multi-Agent Coordination by Aligning World Models Through Dialogue
di: Dongre, Vardhan, et al.
Pubblicazione: (2026)
di: Dongre, Vardhan, et al.
Pubblicazione: (2026)
Map the Flow: Revealing Hidden Pathways of Information in VideoLLMs
di: Kim, Minji, et al.
Pubblicazione: (2025)
di: Kim, Minji, et al.
Pubblicazione: (2025)
Asking Multimodal Clarifying Questions in Mixed-Initiative Conversational Search
di: Yuan, Yifei, et al.
Pubblicazione: (2024)
di: Yuan, Yifei, et al.
Pubblicazione: (2024)
Toward Multimodal Conversational AI for Age-Related Macular Degeneration
di: Gu, Ran, et al.
Pubblicazione: (2026)
di: Gu, Ran, et al.
Pubblicazione: (2026)
Dialog Flow Induction for Constrainable LLM-Based Chatbots
di: Agrawal, Stuti, et al.
Pubblicazione: (2024)
di: Agrawal, Stuti, et al.
Pubblicazione: (2024)
From Fact to Judgment: Investigating the Impact of Task Framing on LLM Conviction in Dialogue Systems
di: Rabbani, Parisa, et al.
Pubblicazione: (2025)
di: Rabbani, Parisa, et al.
Pubblicazione: (2025)
ReIn: Conversational Error Recovery with Reasoning Inception
di: Kim, Takyoung, et al.
Pubblicazione: (2026)
di: Kim, Takyoung, et al.
Pubblicazione: (2026)
Evaluating Multimodal Generative AI with Korean Educational Standards
di: Park, Sanghee, et al.
Pubblicazione: (2025)
di: Park, Sanghee, et al.
Pubblicazione: (2025)
VideoJudge: Bootstrapping Enables Scalable Supervision of MLLM-as-a-Judge for Video Understanding
di: Waheed, Abdul, et al.
Pubblicazione: (2025)
di: Waheed, Abdul, et al.
Pubblicazione: (2025)
LLMs Behind the Scenes: Enabling Narrative Scene Illustration
di: Roemmele, Melissa, et al.
Pubblicazione: (2025)
di: Roemmele, Melissa, et al.
Pubblicazione: (2025)
User Preference Modeling for Conversational LLM Agents: Weak Rewards from Retrieval-Augmented Interaction
di: Hao, Yuren, et al.
Pubblicazione: (2026)
di: Hao, Yuren, et al.
Pubblicazione: (2026)
MemEye: A Visual-Centric Evaluation Framework for Multimodal Agent Memory
di: Guo, Minghao, et al.
Pubblicazione: (2026)
di: Guo, Minghao, et al.
Pubblicazione: (2026)
Bridging the Missing-Modality Gap: Improving Text-Only Calibration of Vision Language Models
di: Kim, Mingyeong, et al.
Pubblicazione: (2026)
di: Kim, Mingyeong, et al.
Pubblicazione: (2026)
DepthFocus: Controllable Depth Estimation for See-Through Scenes
di: Min, Junhong, et al.
Pubblicazione: (2025)
di: Min, Junhong, et al.
Pubblicazione: (2025)
Documenti analoghi
-
AQuA: Toward Strategic Response Generation for Ambiguous Visual Questions
di: Jang, Jihyoung, et al.
Pubblicazione: (2026) -
MIRAGE: A Benchmark for Multimodal Information-Seeking and Reasoning in Agricultural Expert-Guided Conversations
di: Dongre, Vardhan, et al.
Pubblicazione: (2025) -
Mixed-Session Conversation with Egocentric Memory
di: Jang, Jihyoung, et al.
Pubblicazione: (2024) -
Collective Critics for Creative Story Generation
di: Bae, Minwook, et al.
Pubblicazione: (2024) -
Visually Guided Decoding: Gradient-Free Hard Prompt Inversion with Language Models
di: Kim, Donghoon, et al.
Pubblicazione: (2025)