Advancing General-Purpose Reasoning Models with Modular Gradient Surgery
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cai, Min, Liang, Yu, Wang, Longzheng, Wang, Yan, Zhang, Yueyang, Xia, Long, Sun, Zhiyuan, Ye, Xi, Shi, Daiting |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training
von: Liang, Yu, et al.
Veröffentlicht: (2026)
von: Liang, Yu, et al.
Veröffentlicht: (2026)
ReflectRM: Boosting Generative Reward Models via Self-Reflection within a Unified Judgment Framework
von: Qin, Kai, et al.
Veröffentlicht: (2026)
von: Qin, Kai, et al.
Veröffentlicht: (2026)
TRE: Encouraging Exploration in the Trust Region
von: Huang, Chao, et al.
Veröffentlicht: (2026)
von: Huang, Chao, et al.
Veröffentlicht: (2026)
Thinking as Compression: Your Reasoning Model is Secretly a Context Compressor
von: Ma, Guoxin, et al.
Veröffentlicht: (2026)
von: Ma, Guoxin, et al.
Veröffentlicht: (2026)
When Less is More: The LLM Scaling Paradox in Context Compression
von: Guo, Ruishan, et al.
Veröffentlicht: (2026)
von: Guo, Ruishan, et al.
Veröffentlicht: (2026)
GenCRF: Generative Clustering and Reformulation Framework for Enhanced Intent-Driven Information Retrieval
von: Seo, Wonduk, et al.
Veröffentlicht: (2024)
von: Seo, Wonduk, et al.
Veröffentlicht: (2024)
CADDesigner: Conceptual CAD Model Generation with a General-Purpose Agent
von: Fan, Fengxiao, et al.
Veröffentlicht: (2025)
von: Fan, Fengxiao, et al.
Veröffentlicht: (2025)
Nemotron-Cascade: Scaling Cascaded Reinforcement Learning for General-Purpose Reasoning Models
von: Wang, Boxin, et al.
Veröffentlicht: (2025)
von: Wang, Boxin, et al.
Veröffentlicht: (2025)
Bagging-Based Model Merging for Robust General Text Embeddings
von: Zhang, Hengran, et al.
Veröffentlicht: (2026)
von: Zhang, Hengran, et al.
Veröffentlicht: (2026)
MMIDR: Teaching Large Language Model to Interpret Multimodal Misinformation via Knowledge Distillation
von: Wang, Longzheng, et al.
Veröffentlicht: (2024)
von: Wang, Longzheng, et al.
Veröffentlicht: (2024)
ROSD: Reflective On-Policy Self-Distillation for Language Model Reasoning across Domains
von: Zhao, Ziqi, et al.
Veröffentlicht: (2026)
von: Zhao, Ziqi, et al.
Veröffentlicht: (2026)
VisRAG 2.0: Evidence-Guided Multi-Image Reasoning in Visual Retrieval-Augmented Generation
von: Sun, Yubo, et al.
Veröffentlicht: (2025)
von: Sun, Yubo, et al.
Veröffentlicht: (2025)
Learning from Contrasts: Synthesizing Reasoning Paths from Diverse Search Trajectories
von: Liu, Peiyang, et al.
Veröffentlicht: (2026)
von: Liu, Peiyang, et al.
Veröffentlicht: (2026)
Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning
von: Yan, Yibo, et al.
Veröffentlicht: (2025)
von: Yan, Yibo, et al.
Veröffentlicht: (2025)
Structured Outputs Enable General-Purpose LLMs to be Medical Experts
von: Guo, Guangfu, et al.
Veröffentlicht: (2025)
von: Guo, Guangfu, et al.
Veröffentlicht: (2025)
A Modular Multitask Reasoning Framework Integrating Spatio-temporal Models and LLMs
von: Hettige, Kethmi Hirushini, et al.
Veröffentlicht: (2025)
von: Hettige, Kethmi Hirushini, et al.
Veröffentlicht: (2025)
On The Role of Pretrained Language Models in General-Purpose Text Embeddings: A Survey
von: Zhang, Meishan, et al.
Veröffentlicht: (2025)
von: Zhang, Meishan, et al.
Veröffentlicht: (2025)
Generative Evaluation of Complex Reasoning in Large Language Models
von: Lin, Haowei, et al.
Veröffentlicht: (2025)
von: Lin, Haowei, et al.
Veröffentlicht: (2025)
Employing General-Purpose and Biomedical Large Language Models with Advanced Prompt Engineering for Pharmacoepidemiologic Study Design
von: Zhang, Xinyao, et al.
Veröffentlicht: (2026)
von: Zhang, Xinyao, et al.
Veröffentlicht: (2026)
Controlling Thinking Speed in Reasoning Models
von: Lin, Zhengkai, et al.
Veröffentlicht: (2025)
von: Lin, Zhengkai, et al.
Veröffentlicht: (2025)
LLM$\times$MapReduce-V3: Enabling Interactive In-Depth Survey Generation through a MCP-Driven Hierarchically Modular Agent System
von: Chao, Yu, et al.
Veröffentlicht: (2025)
von: Chao, Yu, et al.
Veröffentlicht: (2025)
MedGPT-oss: Training a General-Purpose Vision-Language Model for Biomedicine
von: Zhang, Kai, et al.
Veröffentlicht: (2026)
von: Zhang, Kai, et al.
Veröffentlicht: (2026)
SURGE: On the Potential of Large Language Models as General-Purpose Surrogate Code Executors
von: Lyu, Bohan, et al.
Veröffentlicht: (2025)
von: Lyu, Bohan, et al.
Veröffentlicht: (2025)
Learning to Self-Verify Makes Language Models Better Reasoners
von: Chen, Yuxin, et al.
Veröffentlicht: (2026)
von: Chen, Yuxin, et al.
Veröffentlicht: (2026)
ReasonMed: A 370K Multi-Agent Generated Dataset for Advancing Medical Reasoning
von: Sun, Yu, et al.
Veröffentlicht: (2025)
von: Sun, Yu, et al.
Veröffentlicht: (2025)
Skill-LLM: Repurposing General-Purpose LLMs for Skill Extraction
von: Herandi, Amirhossein, et al.
Veröffentlicht: (2024)
von: Herandi, Amirhossein, et al.
Veröffentlicht: (2024)
Large Language Model-Powered Query-Driven Event Timeline Summarization in Industrial Search
von: Wang, Mingyue, et al.
Veröffentlicht: (2026)
von: Wang, Mingyue, et al.
Veröffentlicht: (2026)
Structure-Enhanced Protein Instruction Tuning: Towards General-Purpose Protein Understanding with LLMs
von: Wu, Wei, et al.
Veröffentlicht: (2024)
von: Wu, Wei, et al.
Veröffentlicht: (2024)
Deep Reasoning in General Purpose Agents via Structured Meta-Cognition
von: Light, Dean, et al.
Veröffentlicht: (2026)
von: Light, Dean, et al.
Veröffentlicht: (2026)
Resprompt: Residual Connection Prompting Advances Multi-Step Reasoning in Large Language Models
von: Jiang, Song, et al.
Veröffentlicht: (2023)
von: Jiang, Song, et al.
Veröffentlicht: (2023)
Reinforced Efficient Reasoning via Semantically Diverse Exploration
von: Zhao, Ziqi, et al.
Veröffentlicht: (2026)
von: Zhao, Ziqi, et al.
Veröffentlicht: (2026)
Beyond Semantic Relevance: Counterfactual Risk Minimization for Robust Retrieval-Augmented Generation
von: Liu, Peiyang, et al.
Veröffentlicht: (2026)
von: Liu, Peiyang, et al.
Veröffentlicht: (2026)
Seed1.5-Thinking: Advancing Superb Reasoning Models with Reinforcement Learning
von: Seed, ByteDance, et al.
Veröffentlicht: (2025)
von: Seed, ByteDance, et al.
Veröffentlicht: (2025)
Unleashing the Power of LLMs in Dense Retrieval with Query Likelihood Modeling
von: Zhang, Hengran, et al.
Veröffentlicht: (2025)
von: Zhang, Hengran, et al.
Veröffentlicht: (2025)
PGMEL: Policy Gradient-based Generative Adversarial Network for Multimodal Entity Linking
von: Pooja, KM, et al.
Veröffentlicht: (2025)
von: Pooja, KM, et al.
Veröffentlicht: (2025)
EMO: Pretraining Mixture of Experts for Emergent Modularity
von: Wang, Ryan, et al.
Veröffentlicht: (2026)
von: Wang, Ryan, et al.
Veröffentlicht: (2026)
Atla Selene Mini: A General Purpose Evaluation Model
von: Alexandru, Andrei, et al.
Veröffentlicht: (2025)
von: Alexandru, Andrei, et al.
Veröffentlicht: (2025)
RAG-Enhanced Large Language Models for Dynamic Content Expiration Prediction in Web Search
von: Chen, Tingyu, et al.
Veröffentlicht: (2026)
von: Chen, Tingyu, et al.
Veröffentlicht: (2026)
LLMs as Scalable, General-Purpose Simulators For Evolving Digital Agent Training
von: Wang, Yiming, et al.
Veröffentlicht: (2025)
von: Wang, Yiming, et al.
Veröffentlicht: (2025)
Concise and Organized Perception Facilitates Reasoning in Large Language Models
von: Liu, Junjie, et al.
Veröffentlicht: (2023)
von: Liu, Junjie, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
ConsistRM: Improving Generative Reward Models via Consistency-Aware Self-Training
von: Liang, Yu, et al.
Veröffentlicht: (2026) -
ReflectRM: Boosting Generative Reward Models via Self-Reflection within a Unified Judgment Framework
von: Qin, Kai, et al.
Veröffentlicht: (2026) -
TRE: Encouraging Exploration in the Trust Region
von: Huang, Chao, et al.
Veröffentlicht: (2026) -
Thinking as Compression: Your Reasoning Model is Secretly a Context Compressor
von: Ma, Guoxin, et al.
Veröffentlicht: (2026) -
When Less is More: The LLM Scaling Paradox in Context Compression
von: Guo, Ruishan, et al.
Veröffentlicht: (2026)