UniM: A Unified Any-to-Any Interleaved Multimodal Benchmark
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Yanlin, Guo, Minghui, Zhang, Kaiwen, Zhang, Shize, Zhao, Yiran, Li, Haodong, Zhou, Congyue, Zheng, Weijie, Yan, Yushen, Wu, Shengqiong, Ji, Wei, Cui, Lei, Wei, Furu, Fei, Hao, Lee, Mong-Li, Hsu, Wynne |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
NExT-GPT: Any-to-Any Multimodal LLM
di: Wu, Shengqiong, et al.
Pubblicazione: (2023)
di: Wu, Shengqiong, et al.
Pubblicazione: (2023)
Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition
di: Fei, Hao, et al.
Pubblicazione: (2024)
di: Fei, Hao, et al.
Pubblicazione: (2024)
Orthogonal Spatial-temporal Distributional Transfer for 4D Generation
di: Liu, Wei, et al.
Pubblicazione: (2026)
di: Liu, Wei, et al.
Pubblicazione: (2026)
FormFactory: An Interactive Benchmarking Suite for Multimodal Form-Filling Agents
di: Li, Bobo, et al.
Pubblicazione: (2025)
di: Li, Bobo, et al.
Pubblicazione: (2025)
PanoSent: A Panoptic Sextuple Extraction Benchmark for Multimodal Conversational Aspect-based Sentiment Analysis
di: Luo, Meng, et al.
Pubblicazione: (2024)
di: Luo, Meng, et al.
Pubblicazione: (2024)
Synergizing Understanding and Generation with Interleaved Analyzing-Drafting Thinking
di: Wu, Shengqiong, et al.
Pubblicazione: (2026)
di: Wu, Shengqiong, et al.
Pubblicazione: (2026)
AR-Omni: A Unified Autoregressive Model for Any-to-Any Generation
di: Cheng, Dongjie, et al.
Pubblicazione: (2026)
di: Cheng, Dongjie, et al.
Pubblicazione: (2026)
FysicsWorld: A Unified Full-Modality Benchmark for Any-to-Any Understanding, Generation, and Reasoning
di: Jiang, Yue, et al.
Pubblicazione: (2025)
di: Jiang, Yue, et al.
Pubblicazione: (2025)
AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea
di: Yu, Qifan, et al.
Pubblicazione: (2024)
di: Yu, Qifan, et al.
Pubblicazione: (2024)
TRUST-VL: An Explainable News Assistant for General Multimodal Misinformation Detection
di: Yan, Zehong, et al.
Pubblicazione: (2025)
di: Yan, Zehong, et al.
Pubblicazione: (2025)
Unveiling the Cognitive Compass: Theory-of-Mind-Guided Multimodal Emotion Reasoning
di: Luo, Meng, et al.
Pubblicazione: (2026)
di: Luo, Meng, et al.
Pubblicazione: (2026)
Any2Caption:Interpreting Any Condition to Caption for Controllable Video Generation
di: Wu, Shengqiong, et al.
Pubblicazione: (2025)
di: Wu, Shengqiong, et al.
Pubblicazione: (2025)
SNIFFER: Multimodal Large Language Model for Explainable Out-of-Context Misinformation Detection
di: Qi, Peng, et al.
Pubblicazione: (2024)
di: Qi, Peng, et al.
Pubblicazione: (2024)
Mitigating GenAI-powered Evidence Pollution for Out-of-Context Multimodal Misinformation Detection
di: Yan, Zehong, et al.
Pubblicazione: (2025)
di: Yan, Zehong, et al.
Pubblicazione: (2025)
Multi-Part Object Representations via Graph Structures and Co-Part Discovery
di: Foo, Alex, et al.
Pubblicazione: (2025)
di: Foo, Alex, et al.
Pubblicazione: (2025)
Multi-Modal Continual Learning via Cross-Modality Adapters and Representation Alignment with Knowledge Preservation
di: Chee, Evelyn, et al.
Pubblicazione: (2025)
di: Chee, Evelyn, et al.
Pubblicazione: (2025)
Any2Any: Unified Arbitrary Modality Translation for Remote Sensing
di: Chen, Haoyang, et al.
Pubblicazione: (2026)
di: Chen, Haoyang, et al.
Pubblicazione: (2026)
Animate Any Character in Any World
di: Wang, Yitong, et al.
Pubblicazione: (2025)
di: Wang, Yitong, et al.
Pubblicazione: (2025)
AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling
di: Zhan, Jun, et al.
Pubblicazione: (2024)
di: Zhan, Jun, et al.
Pubblicazione: (2024)
Any2Any: Incomplete Multimodal Retrieval with Conformal Prediction
di: Li, Po-han, et al.
Pubblicazione: (2024)
di: Li, Po-han, et al.
Pubblicazione: (2024)
Multimodal Crystal Flow: Any-to-Any Modality Generation for Unified Crystal Modeling
di: Seong, Kiyoung, et al.
Pubblicazione: (2026)
di: Seong, Kiyoung, et al.
Pubblicazione: (2026)
Motion Anything: Any to Motion Generation
di: Zhang, Zeyu, et al.
Pubblicazione: (2025)
di: Zhang, Zeyu, et al.
Pubblicazione: (2025)
Spider: Any-to-Many Multimodal LLM
di: Lai, Jinxiang, et al.
Pubblicazione: (2024)
di: Lai, Jinxiang, et al.
Pubblicazione: (2024)
Evidence-Based Temporal Fact Verification
di: Barik, Anab Maulana, et al.
Pubblicazione: (2024)
di: Barik, Anab Maulana, et al.
Pubblicazione: (2024)
ChronoFact: Timeline-based Temporal Fact Verification
di: Barik, Anab Maulana, et al.
Pubblicazione: (2024)
di: Barik, Anab Maulana, et al.
Pubblicazione: (2024)
InstructSAM: Segment Any Instance with Any Instructions
di: Yuan, Yuqian, et al.
Pubblicazione: (2026)
di: Yuan, Yuqian, et al.
Pubblicazione: (2026)
Unifying Sequences, Structures, and Descriptions for Any-to-Any Protein Generation with the Large Multimodal Model HelixProtX
di: Chen, Zhiyuan, et al.
Pubblicazione: (2024)
di: Chen, Zhiyuan, et al.
Pubblicazione: (2024)
LEAF-Mamba: Local Emphatic and Adaptive Fusion State Space Model for RGB-D Salient Object Detection
di: Wu, Lanhu, et al.
Pubblicazione: (2025)
di: Wu, Lanhu, et al.
Pubblicazione: (2025)
AnyPcc: Compressing Any Point Cloud with a Single Universal Model
di: Wang, Kangli, et al.
Pubblicazione: (2025)
di: Wang, Kangli, et al.
Pubblicazione: (2025)
LogicReward: Incentivizing LLM Reasoning via Step-Wise Logical Supervision
di: Xu, Jundong, et al.
Pubblicazione: (2025)
di: Xu, Jundong, et al.
Pubblicazione: (2025)
Test-Time Adaptation by Causal Trimming
di: Liu, Yingnan, et al.
Pubblicazione: (2025)
di: Liu, Yingnan, et al.
Pubblicazione: (2025)
AnyTSR: Any-Scale Thermal Super-Resolution for UAV
di: Li, Mengyuan, et al.
Pubblicazione: (2025)
di: Li, Mengyuan, et al.
Pubblicazione: (2025)
UniLoc: Towards Universal Place Recognition Using Any Single Modality
di: Xia, Yan, et al.
Pubblicazione: (2024)
di: Xia, Yan, et al.
Pubblicazione: (2024)
Towards the first axion search results of the Any Light Particle Search II experiment
di: Wei, Li-Wei
Pubblicazione: (2024)
di: Wei, Li-Wei
Pubblicazione: (2024)
Modeling Unified Semantic Discourse Structure for High-quality Headline Generation
di: Xu, Minghui, et al.
Pubblicazione: (2024)
di: Xu, Minghui, et al.
Pubblicazione: (2024)
StitchFusion: Weaving Any Visual Modalities to Enhance Multimodal Semantic Segmentation
di: Li, Bingyu, et al.
Pubblicazione: (2024)
di: Li, Bingyu, et al.
Pubblicazione: (2024)
Track Any Motions under Any Disturbances
di: Zhang, Zhikai, et al.
Pubblicazione: (2025)
di: Zhang, Zhikai, et al.
Pubblicazione: (2025)
Cornfigurator: Automated Planning for Any-to-Any Multimodal Model Serving
di: Ma, Jeff J., et al.
Pubblicazione: (2025)
di: Ma, Jeff J., et al.
Pubblicazione: (2025)
Symbolic Representation for Any-to-Any Generative Tasks
di: Chen, Jiaqi, et al.
Pubblicazione: (2025)
di: Chen, Jiaqi, et al.
Pubblicazione: (2025)
AnyFit: Controllable Virtual Try-on for Any Combination of Attire Across Any Scenario
di: Li, Yuhan, et al.
Pubblicazione: (2024)
di: Li, Yuhan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
NExT-GPT: Any-to-Any Multimodal LLM
di: Wu, Shengqiong, et al.
Pubblicazione: (2023) -
Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition
di: Fei, Hao, et al.
Pubblicazione: (2024) -
Orthogonal Spatial-temporal Distributional Transfer for 4D Generation
di: Liu, Wei, et al.
Pubblicazione: (2026) -
FormFactory: An Interactive Benchmarking Suite for Multimodal Form-Filling Agents
di: Li, Bobo, et al.
Pubblicazione: (2025) -
PanoSent: A Panoptic Sextuple Extraction Benchmark for Multimodal Conversational Aspect-based Sentiment Analysis
di: Luo, Meng, et al.
Pubblicazione: (2024)