LLM Reasoning as Trajectories: Step-Specific Representation Geometry and Correctness Signals
Fuente:
arXiv
Salvato in:
| Autori principali: | Sun, Lihao, Dong, Hang, Qiao, Bo, Lin, Qingwei, Zhang, Dongmei, Rajmohan, Saravan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
From Reasoning to Answer: Empirical, Attention-Based and Mechanistic Insights into Distilled DeepSeek R1 Models
di: Zhang, Jue, et al.
Pubblicazione: (2025)
di: Zhang, Jue, et al.
Pubblicazione: (2025)
Contrastive Attribution in the Wild: An Interpretability Analysis of LLM Failures on Realistic Benchmarks
di: Tan, Rongyuan, et al.
Pubblicazione: (2026)
di: Tan, Rongyuan, et al.
Pubblicazione: (2026)
Learning to Refine: Self-Refinement of Parallel Reasoning in LLMs
di: Wang, Qibin, et al.
Pubblicazione: (2025)
di: Wang, Qibin, et al.
Pubblicazione: (2025)
AgentGen: Enhancing Planning Abilities for Large Language Model based Agent via Environment and Task Generation
di: Hu, Mengkang, et al.
Pubblicazione: (2024)
di: Hu, Mengkang, et al.
Pubblicazione: (2024)
MEETING DELEGATE: Benchmarking LLMs on Attending Meetings on Our Behalf
di: Hu, Lingxiang, et al.
Pubblicazione: (2025)
di: Hu, Lingxiang, et al.
Pubblicazione: (2025)
Text2Grad: Reinforcement Learning from Natural Language Feedback
di: Wang, Hanyang, et al.
Pubblicazione: (2025)
di: Wang, Hanyang, et al.
Pubblicazione: (2025)
COIN: Chance-Constrained Imitation Learning for Uncertainty-aware Adaptive Resource Oversubscription Policy
di: Wang, Lu, et al.
Pubblicazione: (2024)
di: Wang, Lu, et al.
Pubblicazione: (2024)
Zipage: Maintain High Request Concurrency for LLM Reasoning through Compressed PagedAttention
di: Liao, Mengqi, et al.
Pubblicazione: (2026)
di: Liao, Mengqi, et al.
Pubblicazione: (2026)
Call Me When Necessary: LLMs can Efficiently and Faithfully Reason over Structured Environments
di: Cheng, Sitao, et al.
Pubblicazione: (2024)
di: Cheng, Sitao, et al.
Pubblicazione: (2024)
Self-Evolved Reward Learning for LLMs
di: Huang, Chenghua, et al.
Pubblicazione: (2024)
di: Huang, Chenghua, et al.
Pubblicazione: (2024)
RuAG: Learned-rule-augmented Generation for Large Language Models
di: Zhang, Yudi, et al.
Pubblicazione: (2024)
di: Zhang, Yudi, et al.
Pubblicazione: (2024)
AutoRAG-HP: Automatic Online Hyper-Parameter Tuning for Retrieval-Augmented Generation
di: Fu, Jia, et al.
Pubblicazione: (2024)
di: Fu, Jia, et al.
Pubblicazione: (2024)
EfficientRAG: Efficient Retriever for Multi-Hop Question Answering
di: Zhuang, Ziyuan, et al.
Pubblicazione: (2024)
di: Zhuang, Ziyuan, et al.
Pubblicazione: (2024)
G-KV: Decoding-Time KV Cache Eviction with Global Attention
di: Liao, Mengqi, et al.
Pubblicazione: (2025)
di: Liao, Mengqi, et al.
Pubblicazione: (2025)
AMPO: Active Multi-Preference Optimization for Self-play Preference Selection
di: Gupta, Taneesh, et al.
Pubblicazione: (2025)
di: Gupta, Taneesh, et al.
Pubblicazione: (2025)
REFA: Reference Free Alignment for multi-preference optimization
di: Gupta, Taneesh, et al.
Pubblicazione: (2024)
di: Gupta, Taneesh, et al.
Pubblicazione: (2024)
Navigating the Unknown: A Chat-Based Collaborative Interface for Personalized Exploratory Tasks
di: Peng, Yingzhe, et al.
Pubblicazione: (2024)
di: Peng, Yingzhe, et al.
Pubblicazione: (2024)
DoVer: Intervention-Driven Auto Debugging for LLM Multi-Agent Systems
di: Ma, Ming, et al.
Pubblicazione: (2025)
di: Ma, Ming, et al.
Pubblicazione: (2025)
Pretrain Value, Not Reward: Decoupled Value Policy Optimization
di: Huang, Chenghua, et al.
Pubblicazione: (2025)
di: Huang, Chenghua, et al.
Pubblicazione: (2025)
UFO: A UI-Focused Agent for Windows OS Interaction
di: Zhang, Chaoyun, et al.
Pubblicazione: (2024)
di: Zhang, Chaoyun, et al.
Pubblicazione: (2024)
Distill Not Only Data but Also Rewards: Can Smaller Language Models Surpass Larger Ones?
di: Zhang, Yudi, et al.
Pubblicazione: (2025)
di: Zhang, Yudi, et al.
Pubblicazione: (2025)
The Geometry of Self-Verification in a Task-Specific Reasoning Model
di: Lee, Andrew, et al.
Pubblicazione: (2025)
di: Lee, Andrew, et al.
Pubblicazione: (2025)
Multi-Preference Optimization: Generalizing DPO via Set-Level Contrasts
di: Gupta, Taneesh, et al.
Pubblicazione: (2024)
di: Gupta, Taneesh, et al.
Pubblicazione: (2024)
Everything of Thoughts: Defying the Law of Penrose Triangle for Thought Generation
di: Ding, Ruomeng, et al.
Pubblicazione: (2023)
di: Ding, Ruomeng, et al.
Pubblicazione: (2023)
Offline Reinforcement Learning for LLM Multi-Step Reasoning
di: Wang, Huaijie, et al.
Pubblicazione: (2024)
di: Wang, Huaijie, et al.
Pubblicazione: (2024)
An Advanced Reinforcement Learning Framework for Online Scheduling of Deferrable Workloads in Cloud Computing
di: Dong, Hang, et al.
Pubblicazione: (2024)
di: Dong, Hang, et al.
Pubblicazione: (2024)
Privacy in Action: Towards Realistic Privacy Mitigation and Evaluation for LLM-Powered Agents
di: Wang, Shouju, et al.
Pubblicazione: (2025)
di: Wang, Shouju, et al.
Pubblicazione: (2025)
WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct
di: Luo, Haipeng, et al.
Pubblicazione: (2023)
di: Luo, Haipeng, et al.
Pubblicazione: (2023)
The Vision of Autonomic Computing: Can LLMs Make It a Reality?
di: Zhang, Zhiyang, et al.
Pubblicazione: (2024)
di: Zhang, Zhiyang, et al.
Pubblicazione: (2024)
Supervised Reinforcement Learning: From Expert Trajectories to Step-wise Reasoning
di: Deng, Yihe, et al.
Pubblicazione: (2025)
di: Deng, Yihe, et al.
Pubblicazione: (2025)
Enabling Autonomic Microservice Management through Self-Learning Agents
di: Yu, Fenglin, et al.
Pubblicazione: (2025)
di: Yu, Fenglin, et al.
Pubblicazione: (2025)
STeCa: Step-level Trajectory Calibration for LLM Agent Learning
di: Wang, Hanlin, et al.
Pubblicazione: (2025)
di: Wang, Hanlin, et al.
Pubblicazione: (2025)
ConCISE: Confidence-guided Compression in Step-by-step Efficient Reasoning
di: Qiao, Ziqing, et al.
Pubblicazione: (2025)
di: Qiao, Ziqing, et al.
Pubblicazione: (2025)
AXIS: Efficient Human-Agent-Computer Interaction with API-First LLM-Based Agents
di: Lu, Junting, et al.
Pubblicazione: (2024)
di: Lu, Junting, et al.
Pubblicazione: (2024)
Thought Anchors: Which LLM Reasoning Steps Matter?
di: Bogdan, Paul C., et al.
Pubblicazione: (2025)
di: Bogdan, Paul C., et al.
Pubblicazione: (2025)
StepFly: Agentic Troubleshooting Guide Automation for Incident Diagnosis
di: Mao, Jiayi, et al.
Pubblicazione: (2025)
di: Mao, Jiayi, et al.
Pubblicazione: (2025)
R$^2$PO: Decoupling Training Trajectories from Inference Responses for LLM Reasoning
di: Wang, Jingchu, et al.
Pubblicazione: (2026)
di: Wang, Jingchu, et al.
Pubblicazione: (2026)
Semantic Caching of Contextual Summaries for Efficient Question-Answering with Language Models
di: Couturier, Camille, et al.
Pubblicazione: (2025)
di: Couturier, Camille, et al.
Pubblicazione: (2025)
Assessing LLM Reasoning Steps via Principal Knowledge Grounding
di: Hwang, Hyeon, et al.
Pubblicazione: (2025)
di: Hwang, Hyeon, et al.
Pubblicazione: (2025)
The Reasoning Trap: An Information-Theoretic Bound on Closed-System Multi-Step LLM Reasoning
di: Shin, Kwan Soo
Pubblicazione: (2026)
di: Shin, Kwan Soo
Pubblicazione: (2026)
Documenti analoghi
-
From Reasoning to Answer: Empirical, Attention-Based and Mechanistic Insights into Distilled DeepSeek R1 Models
di: Zhang, Jue, et al.
Pubblicazione: (2025) -
Contrastive Attribution in the Wild: An Interpretability Analysis of LLM Failures on Realistic Benchmarks
di: Tan, Rongyuan, et al.
Pubblicazione: (2026) -
Learning to Refine: Self-Refinement of Parallel Reasoning in LLMs
di: Wang, Qibin, et al.
Pubblicazione: (2025) -
AgentGen: Enhancing Planning Abilities for Large Language Model based Agent via Environment and Task Generation
di: Hu, Mengkang, et al.
Pubblicazione: (2024) -
MEETING DELEGATE: Benchmarking LLMs on Attending Meetings on Our Behalf
di: Hu, Lingxiang, et al.
Pubblicazione: (2025)