Lyra: An Efficient and Speech-Centric Framework for Omni-Cognition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhong, Zhisheng, Wang, Chengyao, Liu, Yuqi, Yang, Senqiao, Tang, Longxiang, Zhang, Yuechen, Li, Jingyao, Qu, Tianyuan, Li, Yanwei, Chen, Yukang, Yu, Shaozuo, Wu, Sitong, Lo, Eric, Liu, Shu, Jia, Jiaya |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MGM-Omni: Scaling Omni LLMs to Personalized Long-Horizon Speech
von: Wang, Chengyao, et al.
Veröffentlicht: (2025)
von: Wang, Chengyao, et al.
Veröffentlicht: (2025)
TraveLLaMA: A Multimodal Travel Assistant with Large-Scale Dataset and Structured Reasoning
von: Chu, Meng, et al.
Veröffentlicht: (2025)
von: Chu, Meng, et al.
Veröffentlicht: (2025)
MiniMind-O Technical Report: An Open Small-Scale Speech-Native Omni Model
von: Gong, Jingyao
Veröffentlicht: (2026)
von: Gong, Jingyao
Veröffentlicht: (2026)
VisionZip: Longer is Better but Not Necessary in Vision Language Models
von: Yang, Senqiao, et al.
Veröffentlicht: (2024)
von: Yang, Senqiao, et al.
Veröffentlicht: (2024)
Learning Switchable Priors for Neural Image Compression
von: Zhang, Haotian, et al.
Veröffentlicht: (2025)
von: Zhang, Haotian, et al.
Veröffentlicht: (2025)
Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
von: Li, Yanwei, et al.
Veröffentlicht: (2024)
von: Li, Yanwei, et al.
Veröffentlicht: (2024)
DreamOmni: Unified Image Generation and Editing
von: Xia, Bin, et al.
Veröffentlicht: (2024)
von: Xia, Bin, et al.
Veröffentlicht: (2024)
MOODv2: Masked Image Modeling for Out-of-Distribution Detection
von: Li, Jingyao, et al.
Veröffentlicht: (2024)
von: Li, Jingyao, et al.
Veröffentlicht: (2024)
LongInsightBench: A Comprehensive Benchmark for Evaluating Omni-Modal Models on Human-Centric Long-Video Understanding
von: Han, ZhaoYang, et al.
Veröffentlicht: (2025)
von: Han, ZhaoYang, et al.
Veröffentlicht: (2025)
Does Your Vision-Language Model Get Lost in the Long Video Sampling Dilemma?
von: Qu, Tianyuan, et al.
Veröffentlicht: (2025)
von: Qu, Tianyuan, et al.
Veröffentlicht: (2025)
Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation
von: Cheng, Shihao, et al.
Veröffentlicht: (2026)
von: Cheng, Shihao, et al.
Veröffentlicht: (2026)
AVID: A Benchmark for Omni-Modal Audio-Visual Inconsistency Understanding via Agent-Driven Construction
von: Chen, Zixuan, et al.
Veröffentlicht: (2026)
von: Chen, Zixuan, et al.
Veröffentlicht: (2026)
OmniDPO: A Preference Optimization Framework to Address Omni-Modal Hallucination
von: Chen, Junzhe, et al.
Veröffentlicht: (2025)
von: Chen, Junzhe, et al.
Veröffentlicht: (2025)
VisionReasoner: Unified Reasoning-Integrated Visual Perception via Reinforcement Learning
von: Liu, Yuqi, et al.
Veröffentlicht: (2025)
von: Liu, Yuqi, et al.
Veröffentlicht: (2025)
SpeechEE: A Novel Benchmark for Speech Event Extraction
von: Wang, Bin, et al.
Veröffentlicht: (2024)
von: Wang, Bin, et al.
Veröffentlicht: (2024)
SFE-Net: Harnessing Biological Principles of Differential Gene Expression for Improved Feature Selection in Deep Learning Networks
von: Li, Yuqi, et al.
Veröffentlicht: (2024)
von: Li, Yuqi, et al.
Veröffentlicht: (2024)
DreamOmni2: Multimodal Instruction-based Editing and Generation
von: Xia, Bin, et al.
Veröffentlicht: (2025)
von: Xia, Bin, et al.
Veröffentlicht: (2025)
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues
von: Zhou, Dingkun, et al.
Veröffentlicht: (2025)
von: Zhou, Dingkun, et al.
Veröffentlicht: (2025)
Exploring Category-level Articulated Object Pose Tracking on SE(3) Manifolds
von: Meng, Xianhui, et al.
Veröffentlicht: (2025)
von: Meng, Xianhui, et al.
Veröffentlicht: (2025)
ZO-ASR: Zeroth-Order Fine-Tuning of Speech Foundation Models without Back-Propagation
von: Peng, Yuezhang, et al.
Veröffentlicht: (2025)
von: Peng, Yuezhang, et al.
Veröffentlicht: (2025)
HOP: Heterogeneous Topology-based Multimodal Entanglement for Co-Speech Gesture Generation
von: Cheng, Hongye, et al.
Veröffentlicht: (2025)
von: Cheng, Hongye, et al.
Veröffentlicht: (2025)
Explore the Limits of Omni-modal Pretraining at Scale
von: Zhang, Yiyuan, et al.
Veröffentlicht: (2024)
von: Zhang, Yiyuan, et al.
Veröffentlicht: (2024)
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
von: Chen, Liyang, et al.
Veröffentlicht: (2025)
von: Chen, Liyang, et al.
Veröffentlicht: (2025)
StyleSpeaker: Audio-Enhanced Fine-Grained Style Modeling for Speech-Driven 3D Facial Animation
von: Yang, An, et al.
Veröffentlicht: (2025)
von: Yang, An, et al.
Veröffentlicht: (2025)
FutureOmni: Evaluating Future Forecasting from Omni-Modal Context for Multimodal LLMs
von: Chen, Qian, et al.
Veröffentlicht: (2026)
von: Chen, Qian, et al.
Veröffentlicht: (2026)
Mining the Social Fabric: Unveiling Communities for Fake News Detection in Short Videos
von: Gong, Haisong, et al.
Veröffentlicht: (2025)
von: Gong, Haisong, et al.
Veröffentlicht: (2025)
Design-MLLM: A Reinforcement Alignment Framework for Verifiable and Aesthetic Interior Design
von: Yang, Yuxuan, et al.
Veröffentlicht: (2026)
von: Yang, Yuxuan, et al.
Veröffentlicht: (2026)
MTAVG-Bench: A Diagnostic Benchmark for Multi-Talker Dialogue-Centric Audio-Video Generation
von: Zhou, Yang-Hao, et al.
Veröffentlicht: (2026)
von: Zhou, Yang-Hao, et al.
Veröffentlicht: (2026)
FinCall-Surprise: A Large Scale Multi-modal Benchmark for Earning Surprise Prediction
von: Shu, Dong, et al.
Veröffentlicht: (2025)
von: Shu, Dong, et al.
Veröffentlicht: (2025)
ConvBench: A Multi-Turn Conversation Evaluation Benchmark with Hierarchical Capability for Large Vision-Language Models
von: Liu, Shuo, et al.
Veröffentlicht: (2024)
von: Liu, Shuo, et al.
Veröffentlicht: (2024)
OmniTrace: A Unified Framework for Generation-Time Attribution in Omni-Modal LLMs
von: Yan, Qianqi, et al.
Veröffentlicht: (2026)
von: Yan, Qianqi, et al.
Veröffentlicht: (2026)
VSpeechLM: A Visual Speech Language Model for Visual Text-to-Speech Task
von: Wang, Yuyue, et al.
Veröffentlicht: (2025)
von: Wang, Yuyue, et al.
Veröffentlicht: (2025)
SVLA: A Unified Speech-Vision-Language Assistant with Multimodal Reasoning and Speech Generation
von: Huynh, Ngoc Dung, et al.
Veröffentlicht: (2025)
von: Huynh, Ngoc Dung, et al.
Veröffentlicht: (2025)
MusiCRS: Benchmarking Audio-Centric Conversational Recommendation
von: Surana, Rohan, et al.
Veröffentlicht: (2025)
von: Surana, Rohan, et al.
Veröffentlicht: (2025)
RoSMM: A Robust and Secure Multi-Modal Watermarking Framework for Diffusion Models
von: Fang, ZhongLi, et al.
Veröffentlicht: (2025)
von: Fang, ZhongLi, et al.
Veröffentlicht: (2025)
Learned Image Compression with Hierarchical Progressive Context Modeling
von: Li, Yuqi, et al.
Veröffentlicht: (2025)
von: Li, Yuqi, et al.
Veröffentlicht: (2025)
OmniForcing: Unleashing Real-time Joint Audio-Visual Generation
von: Su, Yaofeng, et al.
Veröffentlicht: (2026)
von: Su, Yaofeng, et al.
Veröffentlicht: (2026)
Memory-Anchored Multimodal Reasoning for Explainable Video Forensics
von: Chen, Chen, et al.
Veröffentlicht: (2025)
von: Chen, Chen, et al.
Veröffentlicht: (2025)
M3ST-DTI: A multi-task learning model for drug-target interactions based on multi-modal features and multi-stage alignment
von: Li, Xiangyu, et al.
Veröffentlicht: (2025)
von: Li, Xiangyu, et al.
Veröffentlicht: (2025)
Private Speech Classification without Collapse: Stabilized DP Training and Offline Distillation
von: Wen, Yadi, et al.
Veröffentlicht: (2026)
von: Wen, Yadi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
MGM-Omni: Scaling Omni LLMs to Personalized Long-Horizon Speech
von: Wang, Chengyao, et al.
Veröffentlicht: (2025) -
TraveLLaMA: A Multimodal Travel Assistant with Large-Scale Dataset and Structured Reasoning
von: Chu, Meng, et al.
Veröffentlicht: (2025) -
MiniMind-O Technical Report: An Open Small-Scale Speech-Native Omni Model
von: Gong, Jingyao
Veröffentlicht: (2026) -
VisionZip: Longer is Better but Not Necessary in Vision Language Models
von: Yang, Senqiao, et al.
Veröffentlicht: (2024) -
Learning Switchable Priors for Neural Image Compression
von: Zhang, Haotian, et al.
Veröffentlicht: (2025)