The Geometry of Self-Verification in a Task-Specific Reasoning Model
Fuente:
arXiv
Guardado en:
| Autores principales: | Lee, Andrew, Sun, Lihao, Wendler, Chris, Viégas, Fernanda, Wattenberg, Martin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Relational Composition in Neural Networks: A Survey and Call to Action
por: Wattenberg, Martin, et al.
Publicado: (2024)
por: Wattenberg, Martin, et al.
Publicado: (2024)
Dialogue Action Tokens: Steering Language Models in Goal-Directed Dialogue with a Multi-Turn Planner
por: Li, Kenneth, et al.
Publicado: (2024)
por: Li, Kenneth, et al.
Publicado: (2024)
When Bad Data Leads to Good Models
por: Li, Kenneth, et al.
Publicado: (2025)
por: Li, Kenneth, et al.
Publicado: (2025)
Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task
por: Li, Kenneth, et al.
Publicado: (2022)
por: Li, Kenneth, et al.
Publicado: (2022)
Inference-Time Intervention: Eliciting Truthful Answers from a Language Model
por: Li, Kenneth, et al.
Publicado: (2023)
por: Li, Kenneth, et al.
Publicado: (2023)
Shared Global and Local Geometry of Language Model Embeddings
por: Lee, Andrew, et al.
Publicado: (2025)
por: Lee, Andrew, et al.
Publicado: (2025)
Why Can't Transformers Learn Multiplication? Reverse-Engineering Reveals Long-Range Dependency Pitfalls
por: Bai, Xiaoyan, et al.
Publicado: (2025)
por: Bai, Xiaoyan, et al.
Publicado: (2025)
LLM Reasoning as Trajectories: Step-Specific Representation Geometry and Correctness Signals
por: Sun, Lihao, et al.
Publicado: (2026)
por: Sun, Lihao, et al.
Publicado: (2026)
Tensor Product Representation Probes Reveal Shared Structure Across Linear Directions
por: Lee, Andrew, et al.
Publicado: (2026)
por: Lee, Andrew, et al.
Publicado: (2026)
Measuring and Controlling Instruction (In)Stability in Language Model Dialogs
por: Li, Kenneth, et al.
Publicado: (2024)
por: Li, Kenneth, et al.
Publicado: (2024)
What Does it Mean for a Neural Network to Learn a "World Model"?
por: Li, Kenneth, et al.
Publicado: (2025)
por: Li, Kenneth, et al.
Publicado: (2025)
Decomposing Query-Key Feature Interactions Using Contrastive Covariances
por: Lee, Andrew, et al.
Publicado: (2026)
por: Lee, Andrew, et al.
Publicado: (2026)
Learning DAGs from Data with Few Root Causes
por: Misiakos, Panagiotis, et al.
Publicado: (2023)
por: Misiakos, Panagiotis, et al.
Publicado: (2023)
Reasoning on the Manifold: Bidirectional Consistency for Self-Verification in Diffusion Language Models
por: Ruan, Jiaoyang, et al.
Publicado: (2026)
por: Ruan, Jiaoyang, et al.
Publicado: (2026)
Does visualization help AI understand data?
por: Li, Victoria R., et al.
Publicado: (2025)
por: Li, Victoria R., et al.
Publicado: (2025)
Discovering Forbidden Topics in Language Models
por: Rager, Can, et al.
Publicado: (2025)
por: Rager, Can, et al.
Publicado: (2025)
Distill-then-Replace: Efficient Task-Specific Hybrid Attention Model Construction
por: Xia, Xiaojie, et al.
Publicado: (2026)
por: Xia, Xiaojie, et al.
Publicado: (2026)
Reinforcement Inference: Leveraging Uncertainty for Self-Correcting Language Model Reasoning
por: Sun, Xinhai
Publicado: (2026)
por: Sun, Xinhai
Publicado: (2026)
SMART: Self-learning Meta-strategy Agent for Reasoning Tasks
por: Liu, Rongxing, et al.
Publicado: (2024)
por: Liu, Rongxing, et al.
Publicado: (2024)
ThinkSwitch: Context Distillation with LoRA and Weight Interpolation for Specific-Purpose Reasoning Tasks
por: Saini, Dhruv, et al.
Publicado: (2026)
por: Saini, Dhruv, et al.
Publicado: (2026)
Algorithmic Primitives and Compositional Geometry of Reasoning in Language Models
por: Lippl, Samuel, et al.
Publicado: (2025)
por: Lippl, Samuel, et al.
Publicado: (2025)
Think Before You Lie: How Reasoning Leads to Honesty
por: Yuan, Ann, et al.
Publicado: (2026)
por: Yuan, Ann, et al.
Publicado: (2026)
One-Token Verification for Reasoning Correctness Estimation
por: Zhuang, Zhan, et al.
Publicado: (2026)
por: Zhuang, Zhan, et al.
Publicado: (2026)
Introducing Verification Task of Set Consistency with Set-Consistency Energy Networks
por: Song, Mooho, et al.
Publicado: (2025)
por: Song, Mooho, et al.
Publicado: (2025)
ART: Adaptive Reasoning Trees for Explainable Claim Verification
por: Wadhwa, Sahil, et al.
Publicado: (2026)
por: Wadhwa, Sahil, et al.
Publicado: (2026)
Self-Verification Dilemma: Experience-Driven Suppression of Overused Checking in LLM Reasoning
por: Long, Quanyu, et al.
Publicado: (2026)
por: Long, Quanyu, et al.
Publicado: (2026)
ICLR: In-Context Learning of Representations
por: Park, Core Francisco, et al.
Publicado: (2024)
por: Park, Core Francisco, et al.
Publicado: (2024)
Exploring Correlations of Self-Supervised Tasks for Graphs
por: Fang, Taoran, et al.
Publicado: (2024)
por: Fang, Taoran, et al.
Publicado: (2024)
Large Language Models Can Self-Improve At Web Agent Tasks
por: Patel, Ajay, et al.
Publicado: (2024)
por: Patel, Ajay, et al.
Publicado: (2024)
When to Trust the Cheap Check: Weak and Strong Verification for Reasoning
por: Kiyani, Shayan, et al.
Publicado: (2026)
por: Kiyani, Shayan, et al.
Publicado: (2026)
Speculative Coreset Selection for Task-Specific Fine-tuning
por: Zhang, Xiaoyu, et al.
Publicado: (2024)
por: Zhang, Xiaoyu, et al.
Publicado: (2024)
Projected Task-Specific Layers for Multi-Task Reinforcement Learning
por: Roberts, Josselin Somerville, et al.
Publicado: (2023)
por: Roberts, Josselin Somerville, et al.
Publicado: (2023)
A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning
por: Li, Mengqi, et al.
Publicado: (2025)
por: Li, Mengqi, et al.
Publicado: (2025)
LASS-ODE: Scaling ODE Computations to Connect Foundation Models with Dynamical Physical Systems
por: Li, Haoran, et al.
Publicado: (2026)
por: Li, Haoran, et al.
Publicado: (2026)
Forward-Backward Reasoning in Large Language Models for Mathematical Verification
por: Jiang, Weisen, et al.
Publicado: (2023)
por: Jiang, Weisen, et al.
Publicado: (2023)
Binning as a Pretext Task: Improving Self-Supervised Learning in Tabular Domains
por: Lee, Kyungeun, et al.
Publicado: (2024)
por: Lee, Kyungeun, et al.
Publicado: (2024)
Improving LLM Reasoning through Scaling Inference Computation with Collaborative Verification
por: Liang, Zhenwen, et al.
Publicado: (2024)
por: Liang, Zhenwen, et al.
Publicado: (2024)
Parameterized Argumentation-based Reasoning Tasks for Benchmarking Generative Language Models
por: Steging, Cor, et al.
Publicado: (2025)
por: Steging, Cor, et al.
Publicado: (2025)
Principled Understanding of Generalization for Generative Transformer Models in Arithmetic Reasoning Tasks
por: Xu, Xingcheng, et al.
Publicado: (2024)
por: Xu, Xingcheng, et al.
Publicado: (2024)
Information-Theoretic Limits of Safety Verification for Self-Improving Systems
por: Scrivens, Arsenios
Publicado: (2026)
por: Scrivens, Arsenios
Publicado: (2026)
Ejemplares similares
-
Relational Composition in Neural Networks: A Survey and Call to Action
por: Wattenberg, Martin, et al.
Publicado: (2024) -
Dialogue Action Tokens: Steering Language Models in Goal-Directed Dialogue with a Multi-Turn Planner
por: Li, Kenneth, et al.
Publicado: (2024) -
When Bad Data Leads to Good Models
por: Li, Kenneth, et al.
Publicado: (2025) -
Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task
por: Li, Kenneth, et al.
Publicado: (2022) -
Inference-Time Intervention: Eliciting Truthful Answers from a Language Model
por: Li, Kenneth, et al.
Publicado: (2023)