Router-Suggest: Dynamic Routing for Multimodal Auto-Completion in Visually-Grounded Dialogs
Fuente:
arXiv
Saved in:
| Main Authors: | Mishra, Sandeep, Budagam, Devichand, Mandal, Anubhab, Santra, Bishal, Goyal, Pawan, Gupta, Manish |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Chat-Ghosting: A Comparative Study of Methods for Auto-Completion in Dialog Systems
by: Mishra, Sandeep, et al.
Published: (2025)
by: Mishra, Sandeep, et al.
Published: (2025)
DocQAC: Adaptive Trie-Guided Decoding for Effective In-Document Query Auto-Completion
by: Mehta, Rahul, et al.
Published: (2026)
by: Mehta, Rahul, et al.
Published: (2026)
Hierarchical Prompting Taxonomy: A Universal Evaluation Framework for Large Language Models Aligned with Human Cognitive Principles
by: Budagam, Devichand, et al.
Published: (2024)
by: Budagam, Devichand, et al.
Published: (2024)
Two Causal Principles for Improving Visual Dialog
by: Qi, Jiaxin, et al.
Published: (2019)
by: Qi, Jiaxin, et al.
Published: (2019)
ERVQA: A Dataset to Benchmark the Readiness of Large Vision Language Models in Hospital Environments
by: Ray, Sourjyadip, et al.
Published: (2024)
by: Ray, Sourjyadip, et al.
Published: (2024)
SCULPT: Systematic Tuning of Long Prompts
by: Kumar, Shanu, et al.
Published: (2024)
by: Kumar, Shanu, et al.
Published: (2024)
OralBBNet: Spatially Guided Dental Segmentation of Panoramic X-Rays with Bounding Box Priors
by: Budagam, Devichand, et al.
Published: (2024)
by: Budagam, Devichand, et al.
Published: (2024)
ReFeR: Improving Evaluation and Reasoning through Hierarchy of Models
by: Narsupalli, Yaswanth, et al.
Published: (2024)
by: Narsupalli, Yaswanth, et al.
Published: (2024)
Long Dialog Summarization: An Analysis
by: Mullick, Ankan, et al.
Published: (2024)
by: Mullick, Ankan, et al.
Published: (2024)
Intent Detection and Entity Extraction from BioMedical Literature
by: Mullick, Ankan, et al.
Published: (2024)
by: Mullick, Ankan, et al.
Published: (2024)
v1: Learning to Point Visual Tokens for Multimodal Grounded Reasoning
by: Chung, Jiwan, et al.
Published: (2025)
by: Chung, Jiwan, et al.
Published: (2025)
MOTOR: Multimodal Optimal Transport via Grounded Retrieval in Medical Visual Question Answering
by: Shaaban, Mai A., et al.
Published: (2025)
by: Shaaban, Mai A., et al.
Published: (2025)
ECIS-VQG: Generation of Entity-centric Information-seeking Questions from Videos
by: Phukan, Arpan, et al.
Published: (2024)
by: Phukan, Arpan, et al.
Published: (2024)
LatentRouter: Can We Choose the Right Multimodal Model Before Seeing Its Answer?
by: Cheng, Xueqi, et al.
Published: (2026)
by: Cheng, Xueqi, et al.
Published: (2026)
YesBut: A High-Quality Annotated Multimodal Dataset for evaluating Satire Comprehension capability of Vision-Language Models
by: Nandy, Abhilash, et al.
Published: (2024)
by: Nandy, Abhilash, et al.
Published: (2024)
Knowledge-Aware Reasoning over Multimodal Semi-structured Tables
by: Mathur, Suyash Vardhan, et al.
Published: (2024)
by: Mathur, Suyash Vardhan, et al.
Published: (2024)
AutoPresent: Designing Structured Visuals from Scratch
by: Ge, Jiaxin, et al.
Published: (2025)
by: Ge, Jiaxin, et al.
Published: (2025)
HistoryBankQA: Multilingual Temporal Question Answering on Historical Events
by: Mandal, Biswadip, et al.
Published: (2025)
by: Mandal, Biswadip, et al.
Published: (2025)
AV-Dialog: Spoken Dialogue Models with Audio-Visual Input
by: Chen, Tuochao, et al.
Published: (2025)
by: Chen, Tuochao, et al.
Published: (2025)
Visual Text Matters: Improving Text-KVQA with Visual Text Entity Knowledge-aware Large Multimodal Assistant
by: Penamakuri, Abhirama Subramanyam, et al.
Published: (2024)
by: Penamakuri, Abhirama Subramanyam, et al.
Published: (2024)
Text Takes Over: A Study of Modality Bias in Multimodal Intent Detection
by: Mullick, Ankan, et al.
Published: (2025)
by: Mullick, Ankan, et al.
Published: (2025)
Grounding Partially-Defined Events in Multimodal Data
by: Sanders, Kate, et al.
Published: (2024)
by: Sanders, Kate, et al.
Published: (2024)
Towards Visual Text Grounding of Multimodal Large Language Model
by: Li, Ming, et al.
Published: (2025)
by: Li, Ming, et al.
Published: (2025)
Arch-Router: Aligning LLM Routing with Human Preferences
by: Tran, Co, et al.
Published: (2025)
by: Tran, Co, et al.
Published: (2025)
Bootstrapping Action-Grounded Visual Dynamics in Unified Vision-Language Models
by: Qiu, Yifu, et al.
Published: (2025)
by: Qiu, Yifu, et al.
Published: (2025)
Grounding Language Models for Visual Entity Recognition
by: Xiao, Zilin, et al.
Published: (2024)
by: Xiao, Zilin, et al.
Published: (2024)
Diffusion Is Your Friend in Show, Suggest and Tell
by: Hu, Jia Cheng, et al.
Published: (2025)
by: Hu, Jia Cheng, et al.
Published: (2025)
Enhancing Visual Dialog State Tracking through Iterative Object-Entity Alignment in Multi-Round Conversations
by: Pang, Wei, et al.
Published: (2024)
by: Pang, Wei, et al.
Published: (2024)
GeoRouter: Dynamic Paradigm Routing for Worldwide Image Geolocalization
by: Jia, Pengyue, et al.
Published: (2026)
by: Jia, Pengyue, et al.
Published: (2026)
HierRouter: Coordinated Routing of Specialized Large Language Models via Reinforcement Learning
by: Gupta, Nikunj, et al.
Published: (2025)
by: Gupta, Nikunj, et al.
Published: (2025)
DRAGON: A Benchmark for Evidence-Grounded Visual Reasoning over Diagrams
by: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Published: (2026)
by: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Published: (2026)
Groma: Localized Visual Tokenization for Grounding Multimodal Large Language Models
by: Ma, Chuofan, et al.
Published: (2024)
by: Ma, Chuofan, et al.
Published: (2024)
Weakly-Supervised 3D Visual Grounding based on Visual Language Alignment
by: Xu, Xiaoxu, et al.
Published: (2023)
by: Xu, Xiaoxu, et al.
Published: (2023)
Composite Sketch+Text Queries for Retrieving Objects with Elusive Names and Complex Interactions
by: Gatti, Prajwal, et al.
Published: (2025)
by: Gatti, Prajwal, et al.
Published: (2025)
SparrowVQE: Visual Question Explanation for Course Content Understanding
by: Li, Jialu, et al.
Published: (2024)
by: Li, Jialu, et al.
Published: (2024)
Plug-and-Play Grounding of Reasoning in Multimodal Large Language Models
by: Chen, Jiaxing, et al.
Published: (2024)
by: Chen, Jiaxing, et al.
Published: (2024)
LLMs as Bridges: Reformulating Grounded Multimodal Named Entity Recognition
by: Li, Jinyuan, et al.
Published: (2024)
by: Li, Jinyuan, et al.
Published: (2024)
Pre-Training Multimodal Hallucination Detectors with Corrupted Grounding Data
by: Whitehead, Spencer, et al.
Published: (2024)
by: Whitehead, Spencer, et al.
Published: (2024)
Olympus: A Universal Task Router for Computer Vision Tasks
by: Lin, Yuanze, et al.
Published: (2024)
by: Lin, Yuanze, et al.
Published: (2024)
Affective Visual Dialog: A Large-Scale Benchmark for Emotional Reasoning Based on Visually Grounded Conversations
by: Haydarov, Kilichbek, et al.
Published: (2023)
by: Haydarov, Kilichbek, et al.
Published: (2023)
Similar Items
-
Chat-Ghosting: A Comparative Study of Methods for Auto-Completion in Dialog Systems
by: Mishra, Sandeep, et al.
Published: (2025) -
DocQAC: Adaptive Trie-Guided Decoding for Effective In-Document Query Auto-Completion
by: Mehta, Rahul, et al.
Published: (2026) -
Hierarchical Prompting Taxonomy: A Universal Evaluation Framework for Large Language Models Aligned with Human Cognitive Principles
by: Budagam, Devichand, et al.
Published: (2024) -
Two Causal Principles for Improving Visual Dialog
by: Qi, Jiaxin, et al.
Published: (2019) -
ERVQA: A Dataset to Benchmark the Readiness of Large Vision Language Models in Hospital Environments
by: Ray, Sourjyadip, et al.
Published: (2024)