HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhao, Jiaxing, Yang, Qize, Peng, Yixing, Bai, Detao, Yao, Shimin, Sun, Boyuan, Chen, Xiang, Fu, Shenghao, chen, Weixuan, Wei, Xihan, Bo, Liefeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HumanOmni-Speaker: Identifying Who said What and When
von: Bai, Detao, et al.
Veröffentlicht: (2026)
von: Bai, Detao, et al.
Veröffentlicht: (2026)
HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context
von: Yang, Qize, et al.
Veröffentlicht: (2025)
von: Yang, Qize, et al.
Veröffentlicht: (2025)
OmniEncoder: See, Hear, and Feel Continuous Motion Like Humans With One Encoder
von: Bai, Detao, et al.
Veröffentlicht: (2026)
von: Bai, Detao, et al.
Veröffentlicht: (2026)
Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis
von: Yang, Qize, et al.
Veröffentlicht: (2025)
von: Yang, Qize, et al.
Veröffentlicht: (2025)
R1-Omni: Explainable Omni-Multimodal Emotion Recognition with Reinforcement Learning
von: Zhao, Jiaxing, et al.
Veröffentlicht: (2025)
von: Zhao, Jiaxing, et al.
Veröffentlicht: (2025)
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding
von: Peng, Yi-Xing, et al.
Veröffentlicht: (2025)
von: Peng, Yi-Xing, et al.
Veröffentlicht: (2025)
CoGenAV: Versatile Audio-Visual Representation Learning via Contrastive-Generative Synchronization
von: Bai, Detao, et al.
Veröffentlicht: (2025)
von: Bai, Detao, et al.
Veröffentlicht: (2025)
LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding
von: Sun, Boyuan, et al.
Veröffentlicht: (2025)
von: Sun, Boyuan, et al.
Veröffentlicht: (2025)
LOVE-R1: Advancing Long Video Understanding with an Adaptive Zoom-in Mechanism via Multi-Step Reasoning
von: Fu, Shenghao, et al.
Veröffentlicht: (2025)
von: Fu, Shenghao, et al.
Veröffentlicht: (2025)
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness
von: Zhao, Jiaxing, et al.
Veröffentlicht: (2025)
von: Zhao, Jiaxing, et al.
Veröffentlicht: (2025)
Frozen-DETR: Enhancing DETR with Image Understanding from Frozen Foundation Models
von: Fu, Shenghao, et al.
Veröffentlicht: (2024)
von: Fu, Shenghao, et al.
Veröffentlicht: (2024)
HumanVideo-MME: Benchmarking MLLMs for Human-Centric Video Understanding
von: Cai, Yuxuan, et al.
Veröffentlicht: (2025)
von: Cai, Yuxuan, et al.
Veröffentlicht: (2025)
GaussianIP: Identity-Preserving Realistic 3D Human Generation via Human-Centric Diffusion Prior
von: Tang, Zichen, et al.
Veröffentlicht: (2025)
von: Tang, Zichen, et al.
Veröffentlicht: (2025)
LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs
von: Sun, Boyuan, et al.
Veröffentlicht: (2025)
von: Sun, Boyuan, et al.
Veröffentlicht: (2025)
A Hierarchical Semantic Distillation Framework for Open-Vocabulary Object Detection
von: Fu, Shenghao, et al.
Veröffentlicht: (2025)
von: Fu, Shenghao, et al.
Veröffentlicht: (2025)
OmniHuman: A Large-scale Dataset and Benchmark for Human-Centric Video Generation
von: Zhu, Lei, et al.
Veröffentlicht: (2026)
von: Zhu, Lei, et al.
Veröffentlicht: (2026)
See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding
von: Sun, Boyuan, et al.
Veröffentlicht: (2026)
von: Sun, Boyuan, et al.
Veröffentlicht: (2026)
ViSpeak: Visual Instruction Feedback in Streaming Videos
von: Fu, Shenghao, et al.
Veröffentlicht: (2025)
von: Fu, Shenghao, et al.
Veröffentlicht: (2025)
LongInsightBench: A Comprehensive Benchmark for Evaluating Omni-Modal Models on Human-Centric Long-Video Understanding
von: Han, ZhaoYang, et al.
Veröffentlicht: (2025)
von: Han, ZhaoYang, et al.
Veröffentlicht: (2025)
DreamID-Omni: Unified Framework for Controllable Human-Centric Audio-Video Generation
von: Guo, Xu, et al.
Veröffentlicht: (2026)
von: Guo, Xu, et al.
Veröffentlicht: (2026)
LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models
von: Fu, Shenghao, et al.
Veröffentlicht: (2025)
von: Fu, Shenghao, et al.
Veröffentlicht: (2025)
UniPortrait: A Unified Framework for Identity-Preserving Single- and Multi-Human Image Personalization
von: He, Junjie, et al.
Veröffentlicht: (2024)
von: He, Junjie, et al.
Veröffentlicht: (2024)
HumanVBench: Probing Human-Centric Video Understanding in MLLMs with Automatically Synthesized Benchmarks
von: Zhou, Ting, et al.
Veröffentlicht: (2024)
von: Zhou, Ting, et al.
Veröffentlicht: (2024)
OmniTalker: One-shot Real-time Text-Driven Talking Audio-Video Generation With Multimodal Style Mimicking
von: Wang, Zhongjian, et al.
Veröffentlicht: (2025)
von: Wang, Zhongjian, et al.
Veröffentlicht: (2025)
Lyra: An Efficient and Speech-Centric Framework for Omni-Cognition
von: Zhong, Zhisheng, et al.
Veröffentlicht: (2024)
von: Zhong, Zhisheng, et al.
Veröffentlicht: (2024)
Towards Fine-grained Interactive Segmentation in Images and Videos
von: Yao, Yuan, et al.
Veröffentlicht: (2025)
von: Yao, Yuan, et al.
Veröffentlicht: (2025)
IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation
von: Li, Yuan-Ming, et al.
Veröffentlicht: (2025)
von: Li, Yuan-Ming, et al.
Veröffentlicht: (2025)
Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation
von: Cheng, Shihao, et al.
Veröffentlicht: (2026)
von: Cheng, Shihao, et al.
Veröffentlicht: (2026)
MIMO: Controllable Character Video Synthesis with Spatial Decomposed Modeling
von: Men, Yifang, et al.
Veröffentlicht: (2024)
von: Men, Yifang, et al.
Veröffentlicht: (2024)
SpeechAlign: Aligning Speech Generation to Human Preferences
von: Zhang, Dong, et al.
Veröffentlicht: (2024)
von: Zhang, Dong, et al.
Veröffentlicht: (2024)
Data Augmentation in Human-Centric Vision
von: Jiang, Wentao, et al.
Veröffentlicht: (2024)
von: Jiang, Wentao, et al.
Veröffentlicht: (2024)
Social Matrix and the Theory of Human Soul Evolution The Evolution of Life Forms from Survival Fear to Free Play
von: chen, can
Veröffentlicht: (2026)
von: chen, can
Veröffentlicht: (2026)
OmniPT: Unleashing the Potential of Large Vision Language Models for Pedestrian Tracking and Understanding
von: Fu, Teng, et al.
Veröffentlicht: (2025)
von: Fu, Teng, et al.
Veröffentlicht: (2025)
ShapeLLM-Omni: A Native Multimodal LLM for 3D Generation and Understanding
von: Ye, Junliang, et al.
Veröffentlicht: (2025)
von: Ye, Junliang, et al.
Veröffentlicht: (2025)
OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation
von: Zhou, Donghao, et al.
Veröffentlicht: (2026)
von: Zhou, Donghao, et al.
Veröffentlicht: (2026)
CogGuide: Human-Like Guidance for Zero-Shot Omni-Modal Reasoning
von: Shou, Zhou-Peng, et al.
Veröffentlicht: (2025)
von: Shou, Zhou-Peng, et al.
Veröffentlicht: (2025)
Controllable and Expressive One-Shot Video Head Swapping
von: Ji, Chaonan, et al.
Veröffentlicht: (2025)
von: Ji, Chaonan, et al.
Veröffentlicht: (2025)
Fewer Tokens and Fewer Videos: Extending Video Understanding Abilities in Large Vision-Language Models
von: Chen, Shimin, et al.
Veröffentlicht: (2024)
von: Chen, Shimin, et al.
Veröffentlicht: (2024)
OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding
von: Xi, Dianbing, et al.
Veröffentlicht: (2025)
von: Xi, Dianbing, et al.
Veröffentlicht: (2025)
OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs
von: Li, Caorui, et al.
Veröffentlicht: (2025)
von: Li, Caorui, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
HumanOmni-Speaker: Identifying Who said What and When
von: Bai, Detao, et al.
Veröffentlicht: (2026) -
HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context
von: Yang, Qize, et al.
Veröffentlicht: (2025) -
OmniEncoder: See, Hear, and Feel Continuous Motion Like Humans With One Encoder
von: Bai, Detao, et al.
Veröffentlicht: (2026) -
Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis
von: Yang, Qize, et al.
Veröffentlicht: (2025) -
R1-Omni: Explainable Omni-Multimodal Emotion Recognition with Reinforcement Learning
von: Zhao, Jiaxing, et al.
Veröffentlicht: (2025)