FlowReasoner: Reinforcing Query-Level Meta-Agents
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gao, Hongcheng, Liu, Yue, He, Yufei, Dou, Longxu, Du, Chao, Deng, Zhijie, Hooi, Bryan, Lin, Min, Pang, Tianyu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Meta-Unlearning on Diffusion Models: Preventing Relearning Unlearned Concepts
von: Gao, Hongcheng, et al.
Veröffentlicht: (2024)
von: Gao, Hongcheng, et al.
Veröffentlicht: (2024)
NoisyRollout: Reinforcing Visual Reasoning with Data Augmentation
von: Liu, Xiangyan, et al.
Veröffentlicht: (2025)
von: Liu, Xiangyan, et al.
Veröffentlicht: (2025)
Meta-Reasoner: Dynamic Guidance for Optimized Inference-time Reasoning in Large Language Models
von: Sui, Yuan, et al.
Veröffentlicht: (2025)
von: Sui, Yuan, et al.
Veröffentlicht: (2025)
Zombie Agents: Persistent Control of Self-Evolving LLM Agents via Self-Reinforcing Injections
von: Yang, Xianglin, et al.
Veröffentlicht: (2026)
von: Yang, Xianglin, et al.
Veröffentlicht: (2026)
GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning
von: Liu, Yue, et al.
Veröffentlicht: (2025)
von: Liu, Yue, et al.
Veröffentlicht: (2025)
Just-In-Time Reinforcement Learning: Continual Learning in LLM Agents Without Gradient Updates
von: Li, Yibo, et al.
Veröffentlicht: (2026)
von: Li, Yibo, et al.
Veröffentlicht: (2026)
GuardReasoner: Towards Reasoning-based LLM Safeguards
von: Liu, Yue, et al.
Veröffentlicht: (2025)
von: Liu, Yue, et al.
Veröffentlicht: (2025)
Enhancing Multi-Agent Debate System Performance via Confidence Expression
von: Lin, Zijie, et al.
Veröffentlicht: (2025)
von: Lin, Zijie, et al.
Veröffentlicht: (2025)
Reasoning Does Not Necessarily Improve Role-Playing Ability
von: Feng, Xiachong, et al.
Veröffentlicht: (2025)
von: Feng, Xiachong, et al.
Veröffentlicht: (2025)
TACT: Mitigating Overthinking and Overacting in Coding Agents via Activation Steering
von: Sui, Yuan, et al.
Veröffentlicht: (2026)
von: Sui, Yuan, et al.
Veröffentlicht: (2026)
Chain of Preference Optimization: Improving Chain-of-Thought Reasoning in LLMs
von: Zhang, Xuan, et al.
Veröffentlicht: (2024)
von: Zhang, Xuan, et al.
Veröffentlicht: (2024)
UniGraph: Learning a Unified Cross-Domain Foundation Model for Text-Attributed Graphs
von: He, Yufei, et al.
Veröffentlicht: (2024)
von: He, Yufei, et al.
Veröffentlicht: (2024)
Diffusion Language Models are Super Data Learners
von: Ni, Jinjie, et al.
Veröffentlicht: (2025)
von: Ni, Jinjie, et al.
Veröffentlicht: (2025)
Training Optimal Large Diffusion Language Models
von: Ni, Jinjie, et al.
Veröffentlicht: (2025)
von: Ni, Jinjie, et al.
Veröffentlicht: (2025)
RegMix: Data Mixture as Regression for Language Model Pre-training
von: Liu, Qian, et al.
Veröffentlicht: (2024)
von: Liu, Qian, et al.
Veröffentlicht: (2024)
Reinforcing General Reasoning without Verifiers
von: Zhou, Xiangxin, et al.
Veröffentlicht: (2025)
von: Zhou, Xiangxin, et al.
Veröffentlicht: (2025)
FiDeLiS: Faithful Reasoning in Large Language Model for Knowledge Graph Question Answering
von: Sui, Yuan, et al.
Veröffentlicht: (2024)
von: Sui, Yuan, et al.
Veröffentlicht: (2024)
Can Knowledge Graphs Make Large Language Models More Trustworthy? An Empirical Study Over Open-ended Question Answering
von: Sui, Yuan, et al.
Veröffentlicht: (2024)
von: Sui, Yuan, et al.
Veröffentlicht: (2024)
Conversation for Non-verifiable Learning: Self-Evolving LLMs through Meta-Evaluation
von: Sui, Yuan, et al.
Veröffentlicht: (2026)
von: Sui, Yuan, et al.
Veröffentlicht: (2026)
WebAgentGuard: A Reasoning-Driven Guard Model for Detecting Prompt Injection Attacks in Web Agents
von: Chen, Yulin, et al.
Veröffentlicht: (2026)
von: Chen, Yulin, et al.
Veröffentlicht: (2026)
UniGraph2: Learning a Unified Embedding Space to Bind Multimodal Graphs
von: He, Yufei, et al.
Veröffentlicht: (2025)
von: He, Yufei, et al.
Veröffentlicht: (2025)
Enhancing Numerical Reasoning with the Guidance of Reliable Reasoning Processes
von: Wang, Dingzirui, et al.
Veröffentlicht: (2024)
von: Wang, Dingzirui, et al.
Veröffentlicht: (2024)
Efficient Inference for Large Reasoning Models: A Survey
von: Liu, Yue, et al.
Veröffentlicht: (2025)
von: Liu, Yue, et al.
Veröffentlicht: (2025)
Efficient Detection of LLM-generated Texts with a Bayesian Surrogate Model
von: Miao, Yibo, et al.
Veröffentlicht: (2023)
von: Miao, Yibo, et al.
Veröffentlicht: (2023)
NTSFormer: A Self-Teaching Graph Transformer for Multimodal Isolated Cold-Start Node Classification
von: Hu, Jun, et al.
Veröffentlicht: (2025)
von: Hu, Jun, et al.
Veröffentlicht: (2025)
Towards A Unified View of Answer Calibration for Multi-Step Reasoning
von: Deng, Shumin, et al.
Veröffentlicht: (2023)
von: Deng, Shumin, et al.
Veröffentlicht: (2023)
Scalable Token-Level Hallucination Detection in Large Language Models
von: Min, Rui, et al.
Veröffentlicht: (2026)
von: Min, Rui, et al.
Veröffentlicht: (2026)
Optimizing Anytime Reasoning via Budget Relative Policy Optimization
von: Qi, Penghui, et al.
Veröffentlicht: (2025)
von: Qi, Penghui, et al.
Veröffentlicht: (2025)
MLR-Bench: Evaluating AI Agents on Open-Ended Machine Learning Research
von: Chen, Hui, et al.
Veröffentlicht: (2025)
von: Chen, Hui, et al.
Veröffentlicht: (2025)
A Survey of Table Reasoning with Large Language Models
von: Zhang, Xuanliang, et al.
Veröffentlicht: (2024)
von: Zhang, Xuanliang, et al.
Veröffentlicht: (2024)
Think in Parallel, Answer as One: Logit Averaging for Open-Ended Reasoning
von: Wang, Haonan, et al.
Veröffentlicht: (2025)
von: Wang, Haonan, et al.
Veröffentlicht: (2025)
Variational Reasoning for Language Models
von: Zhou, Xiangxin, et al.
Veröffentlicht: (2025)
von: Zhou, Xiangxin, et al.
Veröffentlicht: (2025)
AliMark: Enhancing Robustness of Sentence-Level Watermarking Against Text Paraphrasing
von: Li, Yuexin, et al.
Veröffentlicht: (2026)
von: Li, Yuexin, et al.
Veröffentlicht: (2026)
Seeing is Believing: Mitigating Hallucination in Large Vision-Language Models via CLIP-Guided Decoding
von: Deng, Ailin, et al.
Veröffentlicht: (2024)
von: Deng, Ailin, et al.
Veröffentlicht: (2024)
Echoless Label-Based Pre-computation for Memory-Efficient Heterogeneous Graph Learning
von: Hu, Jun, et al.
Veröffentlicht: (2025)
von: Hu, Jun, et al.
Veröffentlicht: (2025)
Nonparametric Data Attribution for Diffusion Models
von: Zhao, Yutian, et al.
Veröffentlicht: (2025)
von: Zhao, Yutian, et al.
Veröffentlicht: (2025)
Intriguing Properties of Data Attribution on Diffusion Models
von: Zheng, Xiaosen, et al.
Veröffentlicht: (2023)
von: Zheng, Xiaosen, et al.
Veröffentlicht: (2023)
LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation
von: Zhang, Xuan, et al.
Veröffentlicht: (2024)
von: Zhang, Xuan, et al.
Veröffentlicht: (2024)
AdaMoE: Token-Adaptive Routing with Null Experts for Mixture-of-Experts Language Models
von: Zeng, Zihao, et al.
Veröffentlicht: (2024)
von: Zeng, Zihao, et al.
Veröffentlicht: (2024)
Can Indirect Prompt Injection Attacks Be Detected and Removed?
von: Chen, Yulin, et al.
Veröffentlicht: (2025)
von: Chen, Yulin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Meta-Unlearning on Diffusion Models: Preventing Relearning Unlearned Concepts
von: Gao, Hongcheng, et al.
Veröffentlicht: (2024) -
NoisyRollout: Reinforcing Visual Reasoning with Data Augmentation
von: Liu, Xiangyan, et al.
Veröffentlicht: (2025) -
Meta-Reasoner: Dynamic Guidance for Optimized Inference-time Reasoning in Large Language Models
von: Sui, Yuan, et al.
Veröffentlicht: (2025) -
Zombie Agents: Persistent Control of Self-Evolving LLM Agents via Self-Reinforcing Injections
von: Yang, Xianglin, et al.
Veröffentlicht: (2026) -
GuardReasoner-VL: Safeguarding VLMs via Reinforced Reasoning
von: Liu, Yue, et al.
Veröffentlicht: (2025)