Omni-Judge: Can Omni-LLMs Serve as Human-Aligned Judges for Text-Conditioned Audio-Video Generation?
Fuente:
arXiv
Saved in:
| Main Authors: | Liang, Susan, Huang, Chao, Bellos, Filippos, Tang, Yolo Yunlong, Shen, Qianxiang, Bi, Jing, Song, Luchuan, Zhang, Zeliang, Corso, Jason, Xu, Chenliang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
When to Think and When to Look: Uncertainty-Guided Lookback
by: Bi, Jing, et al.
Published: (2025)
by: Bi, Jing, et al.
Published: (2025)
Empowering LLMs with Pseudo-Untrimmed Videos for Audio-Visual Temporal Understanding
by: Tang, Yolo Yunlong, et al.
Published: (2024)
by: Tang, Yolo Yunlong, et al.
Published: (2024)
Towards Consistent Long-Term Pose Generation
by: Li, Yayuan, et al.
Published: (2025)
by: Li, Yayuan, et al.
Published: (2025)
Can VLMs Truly Forget? Benchmarking Training-Free Visual Concept Unlearning
by: Tan, Zhangyun, et al.
Published: (2026)
by: Tan, Zhangyun, et al.
Published: (2026)
Mistake Attribution: Fine-Grained Mistake Understanding in Egocentric Videos
by: Li, Yayuan, et al.
Published: (2025)
by: Li, Yayuan, et al.
Published: (2025)
TDMM-LM: Bridging Facial Understanding and Animation via Language Models
by: Song, Luchuan, et al.
Published: (2026)
by: Song, Luchuan, et al.
Published: (2026)
VITRO: Vocabulary Inversion for Time-series Representation Optimization
by: Bellos, Filippos, et al.
Published: (2024)
by: Bellos, Filippos, et al.
Published: (2024)
VidComposition: Can MLLMs Analyze Compositions in Compiled Videos?
by: Tang, Yolo Y., et al.
Published: (2024)
by: Tang, Yolo Y., et al.
Published: (2024)
Deep TPC: Temporal-Prior Conditioning for Time Series Forecasting
by: Bellos, Filippos, et al.
Published: (2026)
by: Bellos, Filippos, et al.
Published: (2026)
Video-R4: Reinforcing Text-Rich Video Reasoning with Visual Rumination
by: Tang, Yolo Y., et al.
Published: (2025)
by: Tang, Yolo Y., et al.
Published: (2025)
Substantial, Decomposable, and Invisible: Visual Context Misalignment in Instructional Videos for Physical Tasks
by: Li, Yayuan, et al.
Published: (2026)
by: Li, Yayuan, et al.
Published: (2026)
Towards Effective Human-in-the-Loop Assistive AI Agents
by: Bellos, Filippos, et al.
Published: (2025)
by: Bellos, Filippos, et al.
Published: (2025)
V2Xum-LLM: Cross-Modal Video Summarization with Temporal Prompt Instruction Tuning
by: Hua, Hang, et al.
Published: (2024)
by: Hua, Hang, et al.
Published: (2024)
OmniCLIP: Adapting CLIP for Video Recognition with Spatial-Temporal Omni-Scale Feature Learning
by: Liu, Mushui, et al.
Published: (2024)
by: Liu, Mushui, et al.
Published: (2024)
Omni-Embed-Audio: Leveraging Multimodal LLMs for Robust Audio-Text Retrieval
by: Yoo, HaeJun, et al.
Published: (2026)
by: Yoo, HaeJun, et al.
Published: (2026)
EAGLE: Egocentric AGgregated Language-video Engine
by: Bi, Jing, et al.
Published: (2024)
by: Bi, Jing, et al.
Published: (2024)
OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs
by: Li, Caorui, et al.
Published: (2025)
by: Li, Caorui, et al.
Published: (2025)
Intentional Gesture: Deliver Your Intentions with Gestures for Speech
by: Liu, Pinxin, et al.
Published: (2025)
by: Liu, Pinxin, et al.
Published: (2025)
Omni2Sound: Towards Unified Video-Text-to-Audio Generation
by: Dai, Yusheng, et al.
Published: (2026)
by: Dai, Yusheng, et al.
Published: (2026)
SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers
by: Fei, Zhengcong, et al.
Published: (2025)
by: Fei, Zhengcong, et al.
Published: (2025)
OmniNFT: Modality-wise Omni Diffusion Reinforcement for Joint Audio-Video Generation
by: Zhang, Guohui, et al.
Published: (2026)
by: Zhang, Guohui, et al.
Published: (2026)
Rethinking Audio-Visual Adversarial Vulnerability from Temporal and Modality Perspectives
by: Zhang, Zeliang, et al.
Published: (2025)
by: Zhang, Zeliang, et al.
Published: (2025)
CaRDiff: Video Salient Object Ranking Chain of Thought Reasoning for Saliency Prediction with Diffusion
by: Tang, Yolo Yunlong, et al.
Published: (2024)
by: Tang, Yolo Yunlong, et al.
Published: (2024)
MMPerspective: Do MLLMs Understand Perspective? A Comprehensive Benchmark for Perspective Perception, Reasoning, and Robustness
by: Tang, Yolo Y., et al.
Published: (2025)
by: Tang, Yolo Y., et al.
Published: (2025)
OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text
by: Pian, Weiguo, et al.
Published: (2026)
by: Pian, Weiguo, et al.
Published: (2026)
TextToon: Real-Time Text Toonify Head Avatar from Single Video
by: Song, Luchuan, et al.
Published: (2024)
by: Song, Luchuan, et al.
Published: (2024)
OmniAudio: Generating Spatial Audio from 360-Degree Video
by: Liu, Huadai, et al.
Published: (2025)
by: Liu, Huadai, et al.
Published: (2025)
SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models
by: Xie, Tianyu, et al.
Published: (2026)
by: Xie, Tianyu, et al.
Published: (2026)
OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding
by: Zhao, Ruixiang, et al.
Published: (2026)
by: Zhao, Ruixiang, et al.
Published: (2026)
DreamID-Omni: Unified Framework for Controllable Human-Centric Audio-Video Generation
by: Guo, Xu, et al.
Published: (2026)
by: Guo, Xu, et al.
Published: (2026)
OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation
by: Zhou, Donghao, et al.
Published: (2026)
by: Zhou, Donghao, et al.
Published: (2026)
ZeroSep: Separate Anything in Audio with Zero Training
by: Huang, Chao, et al.
Published: (2025)
by: Huang, Chao, et al.
Published: (2025)
Video Understanding with Large Language Models: A Survey
by: Tang, Yolo Y., et al.
Published: (2023)
by: Tang, Yolo Y., et al.
Published: (2023)
OmniRetriever: Any-to-Any Audio-Video-Text Retrieval via Fusion-as-Teacher Distillation
by: Liu, Yunze, et al.
Published: (2026)
by: Liu, Yunze, et al.
Published: (2026)
Omni-Embed-Nemotron: A Unified Multimodal Retrieval Model for Text, Image, Audio, and Video
by: Xu, Mengyao, et al.
Published: (2025)
by: Xu, Mengyao, et al.
Published: (2025)
OmniVox: Zero-Shot Emotion Recognition with Omni-LLMs
by: Murzaku, John, et al.
Published: (2025)
by: Murzaku, John, et al.
Published: (2025)
Can Sound Replace Vision in LLaVA With Token Substitution?
by: Vosoughi, Ali, et al.
Published: (2025)
by: Vosoughi, Ali, et al.
Published: (2025)
OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding
by: Xi, Dianbing, et al.
Published: (2025)
by: Xi, Dianbing, et al.
Published: (2025)
Do More Details Always Introduce More Hallucinations in LVLM-based Image Captioning?
by: Feng, Mingqian, et al.
Published: (2024)
by: Feng, Mingqian, et al.
Published: (2024)
OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model
by: Li, Maomao, et al.
Published: (2026)
by: Li, Maomao, et al.
Published: (2026)
Similar Items
-
When to Think and When to Look: Uncertainty-Guided Lookback
by: Bi, Jing, et al.
Published: (2025) -
Empowering LLMs with Pseudo-Untrimmed Videos for Audio-Visual Temporal Understanding
by: Tang, Yolo Yunlong, et al.
Published: (2024) -
Towards Consistent Long-Term Pose Generation
by: Li, Yayuan, et al.
Published: (2025) -
Can VLMs Truly Forget? Benchmarking Training-Free Visual Concept Unlearning
by: Tan, Zhangyun, et al.
Published: (2026) -
Mistake Attribution: Fine-Grained Mistake Understanding in Egocentric Videos
by: Li, Yayuan, et al.
Published: (2025)