From Leaky Thoughts to Private Reasoning: Controlling What LRMs Say to Themselves
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Puerto, Haritz, Li, Haonan, Han, Xudong, Baldwin, Timothy, Gurevych, Iryna |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Leaky Thoughts: Large Reasoning Models Are Not Private Thinkers
von: Green, Tommaso, et al.
Veröffentlicht: (2025)
von: Green, Tommaso, et al.
Veröffentlicht: (2025)
Are Multilingual LLMs Culturally-Diverse Reasoners? An Investigation into Multicultural Proverbs and Sayings
von: Liu, Chen Cecilia, et al.
Veröffentlicht: (2023)
von: Liu, Chen Cecilia, et al.
Veröffentlicht: (2023)
Code Prompting Elicits Conditional Reasoning Abilities in Text+Code LLMs
von: Puerto, Haritz, et al.
Veröffentlicht: (2024)
von: Puerto, Haritz, et al.
Veröffentlicht: (2024)
Control Illusion: The Failure of Instruction Hierarchies in Large Language Models
von: Geng, Yilin, et al.
Veröffentlicht: (2025)
von: Geng, Yilin, et al.
Veröffentlicht: (2025)
SciCoQA: Quality Assurance for Scientific Paper--Code Alignment
von: Baumgärtner, Tim, et al.
Veröffentlicht: (2026)
von: Baumgärtner, Tim, et al.
Veröffentlicht: (2026)
Fine-Tuning on Diverse Reasoning Chains Drives Within-Inference CoT Refinement in LLMs
von: Puerto, Haritz, et al.
Veröffentlicht: (2024)
von: Puerto, Haritz, et al.
Veröffentlicht: (2024)
CoTJudger: A Graph-Driven Framework for Automatic Evaluation of Chain-of-Thought Efficiency and Redundancy in LRMs
von: Li, Siyi, et al.
Veröffentlicht: (2026)
von: Li, Siyi, et al.
Veröffentlicht: (2026)
Models That Know How Evaluations Are Designed Score Safer
von: Deckenbach, Katharina, et al.
Veröffentlicht: (2026)
von: Deckenbach, Katharina, et al.
Veröffentlicht: (2026)
SpaRC and SpaRP: Spatial Reasoning Characterization and Path Generation for Understanding Spatial Reasoning Capability of Large Language Models
von: Rizvi, Md Imbesat Hassan, et al.
Veröffentlicht: (2024)
von: Rizvi, Md Imbesat Hassan, et al.
Veröffentlicht: (2024)
IRCoder: Intermediate Representations Make Language Models Robust Multilingual Code Generators
von: Paul, Indraneil, et al.
Veröffentlicht: (2024)
von: Paul, Indraneil, et al.
Veröffentlicht: (2024)
Preemptive Detection and Correction of Misaligned Actions in LLM Agents
von: Fang, Haishuo, et al.
Veröffentlicht: (2024)
von: Fang, Haishuo, et al.
Veröffentlicht: (2024)
Towards Privacy-aware Mental Health AI Models: Advances, Challenges, and Opportunities
von: Mandal, Aishik, et al.
Veröffentlicht: (2025)
von: Mandal, Aishik, et al.
Veröffentlicht: (2025)
MAGneT: Coordinated Multi-Agent Generation of Synthetic Multi-Turn Mental Health Counseling Sessions
von: Mandal, Aishik, et al.
Veröffentlicht: (2025)
von: Mandal, Aishik, et al.
Veröffentlicht: (2025)
BARREL: Boundary-Aware Reasoning for Factual and Reliable LRMs
von: Yang, Junxiao, et al.
Veröffentlicht: (2025)
von: Yang, Junxiao, et al.
Veröffentlicht: (2025)
Benchmarking Gender and Political Bias in Large Language Models
von: Yang, Jinrui, et al.
Veröffentlicht: (2025)
von: Yang, Jinrui, et al.
Veröffentlicht: (2025)
Demystifying Instruction Mixing for Fine-tuning Large Language Models
von: Wang, Renxi, et al.
Veröffentlicht: (2023)
von: Wang, Renxi, et al.
Veröffentlicht: (2023)
Stop Unnecessary Reflection: Training LRMs for Efficient Reasoning with Adaptive Reflection and Length Coordinated Penalty
von: Yu, Zewei, et al.
Veröffentlicht: (2026)
von: Yu, Zewei, et al.
Veröffentlicht: (2026)
Multimodal Large Language Models to Support Real-World Fact-Checking
von: Geng, Jiahui, et al.
Veröffentlicht: (2024)
von: Geng, Jiahui, et al.
Veröffentlicht: (2024)
Sensitivity, Performance, Robustness: Deconstructing the Effect of Sociodemographic Prompting
von: Beck, Tilman, et al.
Veröffentlicht: (2023)
von: Beck, Tilman, et al.
Veröffentlicht: (2023)
PeerQA: A Scientific Question Answering Dataset from Peer Reviews
von: Baumgärtner, Tim, et al.
Veröffentlicht: (2025)
von: Baumgärtner, Tim, et al.
Veröffentlicht: (2025)
Factual Self-Awareness in Language Models: Representation, Robustness, and Scaling
von: Tamoyan, Hovhannes, et al.
Veröffentlicht: (2025)
von: Tamoyan, Hovhannes, et al.
Veröffentlicht: (2025)
Stepwise Verification and Remediation of Student Reasoning Errors with Large Language Model Tutors
von: Daheim, Nico, et al.
Veröffentlicht: (2024)
von: Daheim, Nico, et al.
Veröffentlicht: (2024)
DOCE: Finding the Sweet Spot for Execution-Based Code Generation
von: Li, Haau-Sing, et al.
Veröffentlicht: (2024)
von: Li, Haau-Sing, et al.
Veröffentlicht: (2024)
Scaling Up Membership Inference: When and How Attacks Succeed on Large Language Models
von: Puerto, Haritz, et al.
Veröffentlicht: (2024)
von: Puerto, Haritz, et al.
Veröffentlicht: (2024)
Reasoning Models Don't Always Say What They Think
von: Chen, Yanda, et al.
Veröffentlicht: (2025)
von: Chen, Yanda, et al.
Veröffentlicht: (2025)
The Good, the Bad and the Constructive: Automatically Measuring Peer Review's Utility for Authors
von: Sadallah, Abdelrahman, et al.
Veröffentlicht: (2025)
von: Sadallah, Abdelrahman, et al.
Veröffentlicht: (2025)
From Problem-Solving to Teaching Problem-Solving: Aligning LLMs with Pedagogy using Reinforcement Learning
von: Dinucu-Jianu, David, et al.
Veröffentlicht: (2025)
von: Dinucu-Jianu, David, et al.
Veröffentlicht: (2025)
Hypothesis-Driven Feature Manifold Analysis in LLMs via Supervised Multi-Dimensional Scaling
von: Tiblias, Federico, et al.
Veröffentlicht: (2025)
von: Tiblias, Federico, et al.
Veröffentlicht: (2025)
Let LRMs Break Free from Overthinking via Self-Braking Tuning
von: Zhao, Haoran, et al.
Veröffentlicht: (2025)
von: Zhao, Haoran, et al.
Veröffentlicht: (2025)
CORE-T: COherent REtrieval of Tables for Text-to-SQL
von: Soliman, Hassan, et al.
Veröffentlicht: (2026)
von: Soliman, Hassan, et al.
Veröffentlicht: (2026)
SPARE: Single-Pass Annotation with Reference-Guided Evaluation for Automatic Process Supervision and Reward Modelling
von: Rizvi, Md Imbesat Hassan, et al.
Veröffentlicht: (2025)
von: Rizvi, Md Imbesat Hassan, et al.
Veröffentlicht: (2025)
Uncertainty-Aware Decoding with Minimum Bayes Risk
von: Daheim, Nico, et al.
Veröffentlicht: (2025)
von: Daheim, Nico, et al.
Veröffentlicht: (2025)
C-SEO Bench: Does Conversational SEO Work?
von: Puerto, Haritz, et al.
Veröffentlicht: (2025)
von: Puerto, Haritz, et al.
Veröffentlicht: (2025)
A Comprehensive Review of Datasets for Clinical Mental Health AI Systems
von: Mandal, Aishik, et al.
Veröffentlicht: (2025)
von: Mandal, Aishik, et al.
Veröffentlicht: (2025)
Illusion or Algorithm? Investigating Memorization, Emergence, and Symbolic Processing in In-Context Learning
von: Niu, Jingcheng, et al.
Veröffentlicht: (2025)
von: Niu, Jingcheng, et al.
Veröffentlicht: (2025)
Towards Automated Error Discovery: A Study in Conversational AI
von: Petrak, Dominic, et al.
Veröffentlicht: (2025)
von: Petrak, Dominic, et al.
Veröffentlicht: (2025)
Evaluating Evidence Attribution in Generated Fact Checking Explanations
von: Xing, Rui, et al.
Veröffentlicht: (2024)
von: Xing, Rui, et al.
Veröffentlicht: (2024)
Resource-Limited Joint Multimodal Sentiment Reasoning and Classification via Chain-of-Thought Enhancement and Distillation
von: Shangguan, Haonan, et al.
Veröffentlicht: (2025)
von: Shangguan, Haonan, et al.
Veröffentlicht: (2025)
A Survey of Confidence Estimation and Calibration in Large Language Models
von: Geng, Jiahui, et al.
Veröffentlicht: (2023)
von: Geng, Jiahui, et al.
Veröffentlicht: (2023)
Can LLMs Explain Themselves Counterfactually?
von: Dehghanighobadi, Zahra, et al.
Veröffentlicht: (2025)
von: Dehghanighobadi, Zahra, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Leaky Thoughts: Large Reasoning Models Are Not Private Thinkers
von: Green, Tommaso, et al.
Veröffentlicht: (2025) -
Are Multilingual LLMs Culturally-Diverse Reasoners? An Investigation into Multicultural Proverbs and Sayings
von: Liu, Chen Cecilia, et al.
Veröffentlicht: (2023) -
Code Prompting Elicits Conditional Reasoning Abilities in Text+Code LLMs
von: Puerto, Haritz, et al.
Veröffentlicht: (2024) -
Control Illusion: The Failure of Instruction Hierarchies in Large Language Models
von: Geng, Yilin, et al.
Veröffentlicht: (2025) -
SciCoQA: Quality Assurance for Scientific Paper--Code Alignment
von: Baumgärtner, Tim, et al.
Veröffentlicht: (2026)