Saved in:
| Main Authors: | Wang, Shijian, Jin, Jiarui, Fu, Runhao, Yan, Zexuan, Wang, Xingjian, Hu, Mengkang, Wang, Eric, Li, Xiaoxi, Zhang, Kangning, Yao, Li, Jiao, Wenxiang, Cheng, Xuelian, Lu, Yuan, Ge, Zongyuan |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2603.27813 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Video-Thinker: Sparking "Thinking with Videos" via Reinforcement Learning
by: Wang, Shijian, et al.
Published: (2025)
by: Wang, Shijian, et al.
Published: (2025)
AgentDisCo: Towards Disentanglement and Collaboration in Open-ended Deep Research Agents
by: Jin, Jiarui, et al.
Published: (2026)
by: Jin, Jiarui, et al.
Published: (2026)
Tool-Genesis: A Task-Driven Tool Creation Benchmark for Self-Evolving Language Agent
by: Xia, Bowei, et al.
Published: (2026)
by: Xia, Bowei, et al.
Published: (2026)
MMSkills: Towards Multimodal Skills for General Visual Agents
by: Zhang, Kangning, et al.
Published: (2026)
by: Zhang, Kangning, et al.
Published: (2026)
GlyphBanana: Advancing Precise Text Rendering Through Agentic Workflows
by: Yan, Zexuan, et al.
Published: (2026)
by: Yan, Zexuan, et al.
Published: (2026)
Agent2World: Learning to Generate Symbolic World Models via Adaptive Multi-Agent Feedback
by: Hu, Mengkang, et al.
Published: (2025)
by: Hu, Mengkang, et al.
Published: (2025)
OmniGAIA: Towards Native Omni-Modal AI Agents
by: Li, Xiaoxi, et al.
Published: (2026)
by: Li, Xiaoxi, et al.
Published: (2026)
SEAgent: Self-Evolving Computer Use Agent with Autonomous Learning from Experience
by: Sun, Zeyi, et al.
Published: (2025)
by: Sun, Zeyi, et al.
Published: (2025)
DeepAgent: A General Reasoning Agent with Scalable Toolsets
by: Li, Xiaoxi, et al.
Published: (2025)
by: Li, Xiaoxi, et al.
Published: (2025)
Synthetic Curriculum Reinforces Compositional Text-to-Image Generation
by: Wang, Shijian, et al.
Published: (2025)
by: Wang, Shijian, et al.
Published: (2025)
ColonAdapter: Geometry Estimation Through Foundation Model Adaptation for Colonoscopy
by: Jiang, Zhiyi, et al.
Published: (2025)
by: Jiang, Zhiyi, et al.
Published: (2025)
MuSLR: Multimodal Symbolic Logical Reasoning
by: Xu, Jundong, et al.
Published: (2025)
by: Xu, Jundong, et al.
Published: (2025)
RoomPlanner: Explicit Layout Planner for Easier LLM-Driven 3D Room Generation
by: Sun, Wenzhuo, et al.
Published: (2025)
by: Sun, Wenzhuo, et al.
Published: (2025)
Cosmic Ray Inter-Station Correlation Variations as Precursors of Geomagnetic Storms: A Statistical Study and Multi-Parameter Early Warning Framework
by: Li, Haoyang, et al.
Published: (2026)
by: Li, Haoyang, et al.
Published: (2026)
TourPlanner: A Competitive Consensus Framework with Constraint-Gated Reinforcement Learning for Travel Planning
by: Wang, Yinuo, et al.
Published: (2026)
by: Wang, Yinuo, et al.
Published: (2026)
MuLan: Multimodal-LLM Agent for Progressive and Interactive Multi-Object Diffusion
by: Li, Sen, et al.
Published: (2024)
by: Li, Sen, et al.
Published: (2024)
Benchmarking Chinese Commonsense Reasoning with a Multi-hop Reasoning Perspective
by: You, Wangjie, et al.
Published: (2025)
by: You, Wangjie, et al.
Published: (2025)
MoRE: Unlocking Scalability in Reinforcement Learning for Quadruped Vision-Language-Action Models
by: Zhao, Han, et al.
Published: (2025)
by: Zhao, Han, et al.
Published: (2025)
Fints: Efficient Inference-Time Personalization for LLMs with Fine-Grained Instance-Tailored Steering
by: Du, Kounianhua, et al.
Published: (2025)
by: Du, Kounianhua, et al.
Published: (2025)
Multimodal retrieval‐augmented generation framework for machine translation
by: Shijian Li
Published: (2025)
by: Shijian Li
Published: (2025)
MindAlign: Bridging EEG, Vision, and Language for Zero-Shot Visual Decoding
by: Chen, Zexuan, et al.
Published: (2026)
by: Chen, Zexuan, et al.
Published: (2026)
UNeMo: Collaborative Visual-Language Reasoning and Navigation via a Multimodal World Model
by: Huang, Changxin, et al.
Published: (2025)
by: Huang, Changxin, et al.
Published: (2025)
Neurosymbolic Framework for Concept-Driven Logical Reasoning in Skeleton-Based Human Action Recognition
by: Ilyas, Talha, et al.
Published: (2026)
by: Ilyas, Talha, et al.
Published: (2026)
Hunting Attributes: Context Prototype-Aware Learning for Weakly Supervised Semantic Segmentation
by: Tang, Feilong, et al.
Published: (2024)
by: Tang, Feilong, et al.
Published: (2024)
Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning
by: Yin, Shaofeng, et al.
Published: (2026)
by: Yin, Shaofeng, et al.
Published: (2026)
Image Reconstruction Using a Mixture Score Function (MSF)
by: Cong, Wenxiang, et al.
Published: (2023)
by: Cong, Wenxiang, et al.
Published: (2023)
Unveiling roles of non‐coding RNAs in cancer through advanced technologies
by: Runhao Wang, et al.
Published: (2025)
by: Runhao Wang, et al.
Published: (2025)
AgentFugue: Agent Scaling for Long-Horizon Tasks through Collective Reasoning
by: Hu, Yuyang, et al.
Published: (2026)
by: Hu, Yuyang, et al.
Published: (2026)
LoopTool: Closing the Data-Training Loop for Robust LLM Tool Calls
by: Zhang, Kangning, et al.
Published: (2025)
by: Zhang, Kangning, et al.
Published: (2025)
FANS -- Formal Answer Selection for Natural Language Math Reasoning Using Lean4
by: Yao, Jiarui, et al.
Published: (2025)
by: Yao, Jiarui, et al.
Published: (2025)
WISE: Weak-Supervision-Guided Step-by-Step Explanations for Multimodal LLMs in Image Classification
by: Jiang, Yiwen, et al.
Published: (2025)
by: Jiang, Yiwen, et al.
Published: (2025)
Message Feedback Interference Cancellation Aided UAMP Iterative Detector for OTFS Systems
by: Li, Xiangxiang, et al.
Published: (2024)
by: Li, Xiangxiang, et al.
Published: (2024)
Cognitive Pivot Points and Visual Anchoring: Unveiling and Rectifying Hallucinations in Multimodal Reasoning Models
by: Qian, Zhe, et al.
Published: (2026)
by: Qian, Zhe, et al.
Published: (2026)
TReMu: Towards Neuro-Symbolic Temporal Reasoning for LLM-Agents with Memory in Multi-Session Dialogues
by: Ge, Yubin, et al.
Published: (2025)
by: Ge, Yubin, et al.
Published: (2025)
Lifting Scheme-Based Implicit Disentanglement of Emotion-Related Facial Dynamics in the Wild
by: Wang, Xingjian, et al.
Published: (2024)
by: Wang, Xingjian, et al.
Published: (2024)
TED: Training-Free Experience Distillation for Multimodal Reasoning
by: Yuan, Shuozhi, et al.
Published: (2026)
by: Yuan, Shuozhi, et al.
Published: (2026)
TimelineReasoner: Advancing Timeline Summarization with Large Reasoning Models
by: Zhang, Liancheng, et al.
Published: (2026)
by: Zhang, Liancheng, et al.
Published: (2026)
Generalized Category Discovery under Domain Shift: A Frequency Domain Perspective
by: Feng, Wei, et al.
Published: (2025)
by: Feng, Wei, et al.
Published: (2025)
Expansion and improvement of ChinaMu by MuT‐seq and chromosome‐level assembly of theMu‐starter genome
by: Lei Liang, et al.
Published: (2024)
by: Lei Liang, et al.
Published: (2024)
X-WebAgentBench: A Multilingual Interactive Web Benchmark for Evaluating Global Agentic System
by: Wang, Peng, et al.
Published: (2025)
by: Wang, Peng, et al.
Published: (2025)
Similar Items
-
Video-Thinker: Sparking "Thinking with Videos" via Reinforcement Learning
by: Wang, Shijian, et al.
Published: (2025) -
AgentDisCo: Towards Disentanglement and Collaboration in Open-ended Deep Research Agents
by: Jin, Jiarui, et al.
Published: (2026) -
Tool-Genesis: A Task-Driven Tool Creation Benchmark for Self-Evolving Language Agent
by: Xia, Bowei, et al.
Published: (2026) -
MMSkills: Towards Multimodal Skills for General Visual Agents
by: Zhang, Kangning, et al.
Published: (2026) -
GlyphBanana: Advancing Precise Text Rendering Through Agentic Workflows
by: Yan, Zexuan, et al.
Published: (2026)