Agent-Omni: Test-Time Multimodal Reasoning via Model Coordination for Understanding Anything
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lin, Huawei, Shi, Yunzhi, Geng, Tong, Zhao, Weijie, Wang, Wei, Singh, Ravender Pal |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
UniGuardian: A Unified Defense for Detecting Prompt Injection, Backdoor Attacks and Adversarial Attacks in Large Language Models
von: Lin, Huawei, et al.
Veröffentlicht: (2025)
von: Lin, Huawei, et al.
Veröffentlicht: (2025)
Token-wise Influential Training Data Retrieval for Large Language Models
von: Lin, Huawei, et al.
Veröffentlicht: (2024)
von: Lin, Huawei, et al.
Veröffentlicht: (2024)
VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation
von: Lin, Huawei, et al.
Veröffentlicht: (2025)
von: Lin, Huawei, et al.
Veröffentlicht: (2025)
Understanding and Steering the Cognitive Behaviors of Reasoning Models at Test-Time
von: Zhang, Zhenyu, et al.
Veröffentlicht: (2025)
von: Zhang, Zhenyu, et al.
Veröffentlicht: (2025)
ROMA: Real-time Omni-Multimodal Assistant with Interactive Streaming Understanding
von: Tian, Xueyun, et al.
Veröffentlicht: (2026)
von: Tian, Xueyun, et al.
Veröffentlicht: (2026)
RapidUn: Influence-Driven Parameter Reweighting for Efficient Large Language Model Unlearning
von: Zhao, Guoshenghui, et al.
Veröffentlicht: (2025)
von: Zhao, Guoshenghui, et al.
Veröffentlicht: (2025)
Precedent-Informed Reasoning: Mitigating Overthinking in Large Reasoning Models via Test-Time Precedent Learning
von: Wang, Qianyue, et al.
Veröffentlicht: (2026)
von: Wang, Qianyue, et al.
Veröffentlicht: (2026)
Test-Time Detoxification without Training or Learning Anything
von: Saglam, Baturay, et al.
Veröffentlicht: (2026)
von: Saglam, Baturay, et al.
Veröffentlicht: (2026)
Towards Thinking-Optimal Scaling of Test-Time Compute for LLM Reasoning
von: Yang, Wenkai, et al.
Veröffentlicht: (2025)
von: Yang, Wenkai, et al.
Veröffentlicht: (2025)
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision
von: Xi, Zhiheng, et al.
Veröffentlicht: (2024)
von: Xi, Zhiheng, et al.
Veröffentlicht: (2024)
Collaborative Multi-Agent Test-Time Reinforcement Learning for Reasoning
von: Hu, Zhiyuan, et al.
Veröffentlicht: (2026)
von: Hu, Zhiyuan, et al.
Veröffentlicht: (2026)
OmniFusion: Simultaneous Multilingual Multimodal Translations via Modular Fusion
von: Koneru, Sai, et al.
Veröffentlicht: (2025)
von: Koneru, Sai, et al.
Veröffentlicht: (2025)
Efficient Task Adaptation in Large Language Models via Selective Parameter Optimization
von: Wan, Weijie, et al.
Veröffentlicht: (2026)
von: Wan, Weijie, et al.
Veröffentlicht: (2026)
Towards Collaborative Intelligence: Propagating Intentions and Reasoning for Multi-Agent Coordination with Large Language Models
von: Qiu, Xihe, et al.
Veröffentlicht: (2024)
von: Qiu, Xihe, et al.
Veröffentlicht: (2024)
OmniCode: A Benchmark for Evaluating Software Engineering Agents
von: Sonwane, Atharv, et al.
Veröffentlicht: (2026)
von: Sonwane, Atharv, et al.
Veröffentlicht: (2026)
OmniTrace: A Unified Framework for Generation-Time Attribution in Omni-Modal LLMs
von: Yan, Qianqi, et al.
Veröffentlicht: (2026)
von: Yan, Qianqi, et al.
Veröffentlicht: (2026)
Test-Time Matching: Unlocking Compositional Reasoning in Multimodal Models
von: Zhu, Yinglun, et al.
Veröffentlicht: (2025)
von: Zhu, Yinglun, et al.
Veröffentlicht: (2025)
OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM
von: Ye, Hanrong, et al.
Veröffentlicht: (2025)
von: Ye, Hanrong, et al.
Veröffentlicht: (2025)
The Evolution of Thought: Tracking LLM Overthinking via Reasoning Dynamics Analysis
von: Wei, Zihao, et al.
Veröffentlicht: (2025)
von: Wei, Zihao, et al.
Veröffentlicht: (2025)
OpenOmni: A Collaborative Open Source Tool for Building Future-Ready Multimodal Conversational Agents
von: Sun, Qiang, et al.
Veröffentlicht: (2024)
von: Sun, Qiang, et al.
Veröffentlicht: (2024)
From Large to Small: Transferring CUDA Optimization Expertise via Reasoning Graph
von: Gong, Junfeng, et al.
Veröffentlicht: (2025)
von: Gong, Junfeng, et al.
Veröffentlicht: (2025)
Unlocking Multimodal Mathematical Reasoning via Process Reward Model
von: Luo, Ruilin, et al.
Veröffentlicht: (2025)
von: Luo, Ruilin, et al.
Veröffentlicht: (2025)
Bridging the Reasoning Gap in Vietnamese with Small Language Models via Test-Time Scaling
von: Trung, Bui The, et al.
Veröffentlicht: (2026)
von: Trung, Bui The, et al.
Veröffentlicht: (2026)
Does Thinking More always Help? Mirage of Test-Time Scaling in Reasoning Models
von: Ghosal, Soumya Suvra, et al.
Veröffentlicht: (2025)
von: Ghosal, Soumya Suvra, et al.
Veröffentlicht: (2025)
Understanding Performance Collapse in Layer-Pruned Large Language Models via Decision Representation Transitions
von: Shi, Boyu, et al.
Veröffentlicht: (2026)
von: Shi, Boyu, et al.
Veröffentlicht: (2026)
Multi-agent Undercover Gaming: Hallucination Removal via Counterfactual Test for Multimodal Reasoning
von: Liang, Dayong, et al.
Veröffentlicht: (2025)
von: Liang, Dayong, et al.
Veröffentlicht: (2025)
Say Anything but This: When Tokenizer Betrays Reasoning in LLMs
von: Ayoobi, Navid, et al.
Veröffentlicht: (2026)
von: Ayoobi, Navid, et al.
Veröffentlicht: (2026)
Mem-Gallery: Benchmarking Multimodal Long-Term Conversational Memory for MLLM Agents
von: Bei, Yuanchen, et al.
Veröffentlicht: (2026)
von: Bei, Yuanchen, et al.
Veröffentlicht: (2026)
Multimodal Inconsistency Reasoning (MMIR): A New Benchmark for Multimodal Reasoning Models
von: Yan, Qianqi, et al.
Veröffentlicht: (2025)
von: Yan, Qianqi, et al.
Veröffentlicht: (2025)
On Path to Multimodal Historical Reasoning: HistBench and HistAgent
von: Qiu, Jiahao, et al.
Veröffentlicht: (2025)
von: Qiu, Jiahao, et al.
Veröffentlicht: (2025)
ThinkPilot: Steering Reasoning Models via Automated Think-prefixes Optimization
von: Li, Sunzhu, et al.
Veröffentlicht: (2025)
von: Li, Sunzhu, et al.
Veröffentlicht: (2025)
Exploring the Reasoning Abilities of Multimodal Large Language Models (MLLMs): A Comprehensive Survey on Emerging Trends in Multimodal Reasoning
von: Wang, Yiqi, et al.
Veröffentlicht: (2024)
von: Wang, Yiqi, et al.
Veröffentlicht: (2024)
Logical Reasoning with Outcome Reward Models for Test-Time Scaling
von: Thatikonda, Ramya Keerthy, et al.
Veröffentlicht: (2025)
von: Thatikonda, Ramya Keerthy, et al.
Veröffentlicht: (2025)
Understanding the Collapse of LLMs in Model Editing
von: Yang, Wanli, et al.
Veröffentlicht: (2024)
von: Yang, Wanli, et al.
Veröffentlicht: (2024)
MedAdapter: Efficient Test-Time Adaptation of Large Language Models towards Medical Reasoning
von: Shi, Wenqi, et al.
Veröffentlicht: (2024)
von: Shi, Wenqi, et al.
Veröffentlicht: (2024)
LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning
von: Dai, Yifan, et al.
Veröffentlicht: (2026)
von: Dai, Yifan, et al.
Veröffentlicht: (2026)
Retrieval is Not Enough: Enhancing RAG Reasoning through Test-Time Critique and Optimization
von: Wei, Jiaqi, et al.
Veröffentlicht: (2025)
von: Wei, Jiaqi, et al.
Veröffentlicht: (2025)
Test-Time Scaling in Reasoning Models Is Not Effective for Knowledge-Intensive Tasks Yet
von: Zhao, James Xu, et al.
Veröffentlicht: (2025)
von: Zhao, James Xu, et al.
Veröffentlicht: (2025)
Cast a Wider Net: Coordinated Pass@K Policy Optimization for Code Reasoning
von: Li, Yilong, et al.
Veröffentlicht: (2026)
von: Li, Yilong, et al.
Veröffentlicht: (2026)
From No to Know: Taxonomy, Challenges, and Opportunities for Negation Understanding in Multimodal Foundation Models
von: Vatsa, Mayank, et al.
Veröffentlicht: (2025)
von: Vatsa, Mayank, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
UniGuardian: A Unified Defense for Detecting Prompt Injection, Backdoor Attacks and Adversarial Attacks in Large Language Models
von: Lin, Huawei, et al.
Veröffentlicht: (2025) -
Token-wise Influential Training Data Retrieval for Large Language Models
von: Lin, Huawei, et al.
Veröffentlicht: (2024) -
VTBench: Evaluating Visual Tokenizers for Autoregressive Image Generation
von: Lin, Huawei, et al.
Veröffentlicht: (2025) -
Understanding and Steering the Cognitive Behaviors of Reasoning Models at Test-Time
von: Zhang, Zhenyu, et al.
Veröffentlicht: (2025) -
ROMA: Real-time Omni-Multimodal Assistant with Interactive Streaming Understanding
von: Tian, Xueyun, et al.
Veröffentlicht: (2026)