Q-Router: Agentic Video Quality Assessment with Expert Model Routing and Artifact Localization
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Xing, Shuo, Dey, Soumik, Wu, Mingyang, Mishra, Ashirbad, Ravipati, Naveen, Li, Binbin, Wu, Hansi, Tu, Zhengzhong |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
ConsID-Gen: View-Consistent and Identity-Preserving Image-to-Video Generation
par: Wu, Mingyang, et autres
Publié: (2026)
par: Wu, Mingyang, et autres
Publié: (2026)
LLMDistill4Ads: Using Cross-Encoders to Distill from LLM Signals for Advertiser Keyphrase Recommendations at eBay
par: Dey, Soumik, et autres
Publié: (2025)
par: Dey, Soumik, et autres
Publié: (2025)
BroadGen: A Framework for Generating Effective and Efficient Advertiser Broad Match Keyphrase Recommendations
par: Mishra, Ashirbad, et autres
Publié: (2025)
par: Mishra, Ashirbad, et autres
Publié: (2025)
Batch Speculative Decoding Done Right
par: Zhang, Ranran Haoran, et autres
Publié: (2025)
par: Zhang, Ranran Haoran, et autres
Publié: (2025)
To Judge or not to Judge: Using LLM Judgements for Advertiser Keyphrase Relevance at eBay
par: Dey, Soumik, et autres
Publié: (2025)
par: Dey, Soumik, et autres
Publié: (2025)
Subjective and Objective Quality Assessment of Banding Artifacts on Compressed Videos
par: Zheng, Qi, et autres
Publié: (2025)
par: Zheng, Qi, et autres
Publié: (2025)
Graphite: A Graph-based Extreme Multi-Label Short Text Classifier for Keyphrase Recommendation
par: Mishra, Ashirbad, et autres
Publié: (2024)
par: Mishra, Ashirbad, et autres
Publié: (2024)
Video Quality Assessment: A Comprehensive Survey
par: Zheng, Qi, et autres
Publié: (2024)
par: Zheng, Qi, et autres
Publié: (2024)
Middleman Bias in Advertising: Aligning Relevance of Keyphrase Recommendations with Search
par: Dey, Soumik, et autres
Publié: (2025)
par: Dey, Soumik, et autres
Publié: (2025)
VISTAv2: World Imagination for Indoor Vision-and-Language Navigation
par: Huang, Yanjia, et autres
Publié: (2025)
par: Huang, Yanjia, et autres
Publié: (2025)
GraphEx: A Graph-based Extraction Method for Advertiser Keyphrase Recommendation
par: Mishra, Ashirbad, et autres
Publié: (2024)
par: Mishra, Ashirbad, et autres
Publié: (2024)
Artifact-Aware Evaluation for High-Quality Video Generation
par: Zhu, Chen, et autres
Publié: (2026)
par: Zhu, Chen, et autres
Publié: (2026)
DecAlign: Hierarchical Cross-Modal Alignment for Decoupled Multimodal Representation Learning
par: Qian, Chengxuan, et autres
Publié: (2025)
par: Qian, Chengxuan, et autres
Publié: (2025)
From Lazy to Prolific: Tackling Missing Labels in Open Vocabulary Extreme Classification by Positive-Unlabeled Sequence Learning
par: Zhang, Ranran Haoran, et autres
Publié: (2024)
par: Zhang, Ranran Haoran, et autres
Publié: (2024)
VideoRouter: Query-Adaptive Dual Routing for Efficient Long-Video Understanding
par: Lin, Kuanwei, et autres
Publié: (2026)
par: Lin, Kuanwei, et autres
Publié: (2026)
Demystifying the Visual Quality Paradox in Multimodal Large Language Models
par: Xing, Shuo, et autres
Publié: (2025)
par: Xing, Shuo, et autres
Publié: (2025)
LEHA-CVQAD: Dataset To Enable Generalized Video Quality Assessment of Compression Artifacts
par: Gushchin, Aleksandr, et autres
Publié: (2025)
par: Gushchin, Aleksandr, et autres
Publié: (2025)
Router-Suggest: Dynamic Routing for Multimodal Auto-Completion in Visually-Grounded Dialogs
par: Mishra, Sandeep, et autres
Publié: (2026)
par: Mishra, Sandeep, et autres
Publié: (2026)
The Pulse of Motion: Measuring Physical Frame Rate from Visual Dynamics
par: Gao, Xiangbo, et autres
Publié: (2026)
par: Gao, Xiangbo, et autres
Publié: (2026)
MMHU: A Massive-Scale Multimodal Benchmark for Human Behavior Understanding
par: Li, Renjie, et autres
Publié: (2025)
par: Li, Renjie, et autres
Publié: (2025)
VersusQ: Pairwise Margin Reasoning for Generalizable Video Quality Assessment
par: Meng, Shibei, et autres
Publié: (2026)
par: Meng, Shibei, et autres
Publié: (2026)
Q-Doc: Benchmarking Document Image Quality Assessment Capabilities in Multi-modal Large Language Models
par: Huang, Jiaxi, et autres
Publié: (2025)
par: Huang, Jiaxi, et autres
Publié: (2025)
Q-Bench-Video: Benchmarking the Video Quality Understanding of LMMs
par: Zhang, Zicheng, et autres
Publié: (2024)
par: Zhang, Zicheng, et autres
Publié: (2024)
ArtifactWorld: Scaling 3D Gaussian Splatting Artifact Restoration via Video Generation Models
par: Wang, Xinliang, et autres
Publié: (2026)
par: Wang, Xinliang, et autres
Publié: (2026)
Routers in Vision Mixture of Experts: An Empirical Study
par: Liu, Tianlin, et autres
Publié: (2024)
par: Liu, Tianlin, et autres
Publié: (2024)
GeoRouter: Dynamic Paradigm Routing for Worldwide Image Geolocalization
par: Jia, Pengyue, et autres
Publié: (2026)
par: Jia, Pengyue, et autres
Publié: (2026)
PISCO: Precise Video Instance Insertion with Sparse Control
par: Gao, Xiangbo, et autres
Publié: (2026)
par: Gao, Xiangbo, et autres
Publié: (2026)
Multi-Branch Collaborative Learning Network for Video Quality Assessment in Industrial Video Search
par: Tang, Hengzhu, et autres
Publié: (2025)
par: Tang, Hengzhu, et autres
Publié: (2025)
4KAgent: Agentic Any Image to 4K Super-Resolution
par: Zuo, Yushen, et autres
Publié: (2025)
par: Zuo, Yushen, et autres
Publié: (2025)
4KLSDB: A Large-Scale Dataset for 4K Image Restoration and Generation
par: Zhu, Zihao, et autres
Publié: (2026)
par: Zhu, Zihao, et autres
Publié: (2026)
MCMoE: Completing Missing Modalities with Mixture of Experts for Incomplete Multimodal Action Quality Assessment
par: Xu, Huangbiao, et autres
Publié: (2025)
par: Xu, Huangbiao, et autres
Publié: (2025)
Delta Forcing: Trust Region Steering for Interactive Autoregressive Video Generation
par: Wu, Yuheng, et autres
Publié: (2026)
par: Wu, Yuheng, et autres
Publié: (2026)
YOLOv1 to YOLOv11: A Comprehensive Survey of Real-Time Object Detection Innovations and Challenges
par: Kotthapalli, Manikanta, et autres
Publié: (2025)
par: Kotthapalli, Manikanta, et autres
Publié: (2025)
OpenEMMA: Open-Source Multimodal Model for End-to-End Autonomous Driving
par: Xing, Shuo, et autres
Publié: (2024)
par: Xing, Shuo, et autres
Publié: (2024)
AgenticIQA: An Agentic Framework for Adaptive and Interpretable Image Quality Assessment
par: Zhu, Hanwei, et autres
Publié: (2025)
par: Zhu, Hanwei, et autres
Publié: (2025)
Q-Adapt: Adapting LMM for Visual Quality Assessment with Progressive Instruction Tuning
par: Lu, Yiting, et autres
Publié: (2025)
par: Lu, Yiting, et autres
Publié: (2025)
Background Fades, Foreground Leads: Curriculum-Guided Background Pruning for Efficient Foreground-Centric Collaborative Perception
par: Wu, Yuheng, et autres
Publié: (2025)
par: Wu, Yuheng, et autres
Publié: (2025)
3D4D: An Interactive, Editable, 4D World Model via 3D Video Generation
par: He, Yunhong, et autres
Publié: (2025)
par: He, Yunhong, et autres
Publié: (2025)
Let the Abyss Stare Back Adaptive Falsification for Autonomous Scientific Discovery
par: Li, Peiran, et autres
Publié: (2026)
par: Li, Peiran, et autres
Publié: (2026)
Region-R1: Reinforcing Query-Side Region Cropping for Multi-Modal Re-Ranking
par: Hu, Chan-Wei, et autres
Publié: (2026)
par: Hu, Chan-Wei, et autres
Publié: (2026)
Documents similaires
-
ConsID-Gen: View-Consistent and Identity-Preserving Image-to-Video Generation
par: Wu, Mingyang, et autres
Publié: (2026) -
LLMDistill4Ads: Using Cross-Encoders to Distill from LLM Signals for Advertiser Keyphrase Recommendations at eBay
par: Dey, Soumik, et autres
Publié: (2025) -
BroadGen: A Framework for Generating Effective and Efficient Advertiser Broad Match Keyphrase Recommendations
par: Mishra, Ashirbad, et autres
Publié: (2025) -
Batch Speculative Decoding Done Right
par: Zhang, Ranran Haoran, et autres
Publié: (2025) -
To Judge or not to Judge: Using LLM Judgements for Advertiser Keyphrase Relevance at eBay
par: Dey, Soumik, et autres
Publié: (2025)