Knowing Before Saying: LLM Representations Encode Information About Chain-of-Thought Success Before Completion
Fuente:
arXiv
Salvato in:
| Autori principali: | Afzal, Anum, Matthes, Florian, Chechik, Gal, Ziser, Yftah |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
FActBench: A Benchmark for Fine-grained Automatic Evaluation of LLM-Generated Text in the Medical Domain
di: Afzal, Anum, et al.
Pubblicazione: (2025)
di: Afzal, Anum, et al.
Pubblicazione: (2025)
Aligning What LLMs Do and Say: Towards Self-Consistent Explanations
di: Admoni, Sahar, et al.
Pubblicazione: (2025)
di: Admoni, Sahar, et al.
Pubblicazione: (2025)
Can Smaller LLMs do better? Unlocking Cross-Domain Potential through Parameter-Efficient Fine-Tuning for Text Summarization
di: Afzal, Anum, et al.
Pubblicazione: (2025)
di: Afzal, Anum, et al.
Pubblicazione: (2025)
JaccDiv: A Metric and Benchmark for Quantifying Diversity of Generated Marketing Text in the Music Industry
di: Afzal, Anum, et al.
Pubblicazione: (2025)
di: Afzal, Anum, et al.
Pubblicazione: (2025)
More Test-Time Compute Can Hurt: Overestimation Bias in LLM Beam Search
di: Dalal, Gal, et al.
Pubblicazione: (2026)
di: Dalal, Gal, et al.
Pubblicazione: (2026)
AdaptEval: Evaluating Large Language Models on Domain Adaptation for Text Summarization
di: Afzal, Anum, et al.
Pubblicazione: (2024)
di: Afzal, Anum, et al.
Pubblicazione: (2024)
Policy Optimized Text-to-Image Pipeline Design
di: Gadot, Uri, et al.
Pubblicazione: (2025)
di: Gadot, Uri, et al.
Pubblicazione: (2025)
Towards Optimizing and Evaluating a Retrieval Augmented QA Chatbot using LLMs with Human in the Loop
di: Afzal, Anum, et al.
Pubblicazione: (2024)
di: Afzal, Anum, et al.
Pubblicazione: (2024)
The Detection-Extraction Gap: Models Know the Answer Before They Can Say It
di: Wang, Hanyang, et al.
Pubblicazione: (2026)
di: Wang, Hanyang, et al.
Pubblicazione: (2026)
Layer by Layer: Uncovering Where Multi-Task Learning Happens in Instruction-Tuned Large Language Models
di: Zhao, Zheng, et al.
Pubblicazione: (2024)
di: Zhao, Zheng, et al.
Pubblicazione: (2024)
When Knowledge Is Not Free: Cost-Aware Evidence Selection in Retrieval-Augmented Generation
di: Wu, Mingyan, et al.
Pubblicazione: (2026)
di: Wu, Mingyan, et al.
Pubblicazione: (2026)
Understanding Before Reasoning: Enhancing Chain-of-Thought with Iterative Summarization Pre-Prompting
di: Zhu, Dong-Hai, et al.
Pubblicazione: (2025)
di: Zhu, Dong-Hai, et al.
Pubblicazione: (2025)
Chip-Tuning: Classify Before Language Models Say
di: Zhu, Fangwei, et al.
Pubblicazione: (2024)
di: Zhu, Fangwei, et al.
Pubblicazione: (2024)
SnapKV: LLM Knows What You are Looking for Before Generation
di: Li, Yuhong, et al.
Pubblicazione: (2024)
di: Li, Yuhong, et al.
Pubblicazione: (2024)
LLMs Know More About Numbers than They Can Say
di: Yuchi, Fengting, et al.
Pubblicazione: (2026)
di: Yuchi, Fengting, et al.
Pubblicazione: (2026)
Diffusion Language Models Know the Answer Before Decoding
di: Li, Pengxiang, et al.
Pubblicazione: (2025)
di: Li, Pengxiang, et al.
Pubblicazione: (2025)
Know You Before You Speak: User-State Modeling for LLM Personalization in Multi-Turn Conversation
di: Luo, Jiani, et al.
Pubblicazione: (2026)
di: Luo, Jiani, et al.
Pubblicazione: (2026)
Lost in Space? Vision-Language Models Struggle with Relative Camera Pose Estimation
di: Deng, Ken, et al.
Pubblicazione: (2026)
di: Deng, Ken, et al.
Pubblicazione: (2026)
Iterative Multilingual Spectral Attribute Erasure
di: Shao, Shun, et al.
Pubblicazione: (2025)
di: Shao, Shun, et al.
Pubblicazione: (2025)
Real-Time Generation of Game Video Commentary with Multimodal LLMs: Pause-Aware Decoding Approaches
di: Afzal, Anum, et al.
Pubblicazione: (2026)
di: Afzal, Anum, et al.
Pubblicazione: (2026)
Bootstrapping Action-Grounded Visual Dynamics in Unified Vision-Language Models
di: Qiu, Yifu, et al.
Pubblicazione: (2025)
di: Qiu, Yifu, et al.
Pubblicazione: (2025)
CompressKV: Semantic Retrieval Heads Know What Tokens are Not Important Before Generation
di: Lin, Xiaolin, et al.
Pubblicazione: (2025)
di: Lin, Xiaolin, et al.
Pubblicazione: (2025)
LR-DWM: Efficient Watermarking for Diffusion Language Models
di: Raban, Ofek, et al.
Pubblicazione: (2026)
di: Raban, Ofek, et al.
Pubblicazione: (2026)
LLM-as-a-Judge for Privacy Evaluation? Exploring the Alignment of Human and LLM Perceptions of Privacy in Textual Data
di: Meisenbacher, Stephen, et al.
Pubblicazione: (2025)
di: Meisenbacher, Stephen, et al.
Pubblicazione: (2025)
Think Thrice Before You Act: Progressive Thought Refinement in Large Language Models
di: Du, Chengyu, et al.
Pubblicazione: (2024)
di: Du, Chengyu, et al.
Pubblicazione: (2024)
Before and After Temperature: A Distributional View of Creative LLM Generation
di: Parupudi, V. S. Raghu, et al.
Pubblicazione: (2026)
di: Parupudi, V. S. Raghu, et al.
Pubblicazione: (2026)
When Safety Fails Before the Answer: Benchmarking Harmful Behavior Detection in Reasoning Chains
di: Kakkar, Ishita, et al.
Pubblicazione: (2026)
di: Kakkar, Ishita, et al.
Pubblicazione: (2026)
Correcting Hallucinations in News Summaries: Exploration of Self-Correcting LLM Methods with External Knowledge
di: Vladika, Juraj, et al.
Pubblicazione: (2025)
di: Vladika, Juraj, et al.
Pubblicazione: (2025)
Look Before You Leap: Autonomous Exploration for LLM Agents
di: Ye, Ziang, et al.
Pubblicazione: (2026)
di: Ye, Ziang, et al.
Pubblicazione: (2026)
Spectral Editing of Activations for Large Language Model Alignment
di: Qiu, Yifu, et al.
Pubblicazione: (2024)
di: Qiu, Yifu, et al.
Pubblicazione: (2024)
It Is Not About What You Say, It Is About How You Say It: A Surprisingly Simple Approach for Improving Reading Comprehension
di: Shaier, Sagi, et al.
Pubblicazione: (2024)
di: Shaier, Sagi, et al.
Pubblicazione: (2024)
Walk Before You Run! Concise LLM Reasoning via Reinforcement Learning
di: Song, Mingyang, et al.
Pubblicazione: (2025)
di: Song, Mingyang, et al.
Pubblicazione: (2025)
Confidence Before Answering: A Paradigm Shift for Efficient LLM Uncertainty Estimation
di: Li, Changcheng, et al.
Pubblicazione: (2026)
di: Li, Changcheng, et al.
Pubblicazione: (2026)
With Privacy, Size Matters: On the Importance of Dataset Size in Differentially Private Text Rewriting
di: Meisenbacher, Stephen, et al.
Pubblicazione: (2025)
di: Meisenbacher, Stephen, et al.
Pubblicazione: (2025)
Just Rewrite It Again: A Post-Processing Method for Enhanced Semantic Similarity and Privacy Preservation of Differentially Private Rewritten Text
di: Meisenbacher, Stephen, et al.
Pubblicazione: (2024)
di: Meisenbacher, Stephen, et al.
Pubblicazione: (2024)
Thinking Outside of the Differential Privacy Box: A Case Study in Text Privatization with Language Model Prompting
di: Meisenbacher, Stephen, et al.
Pubblicazione: (2024)
di: Meisenbacher, Stephen, et al.
Pubblicazione: (2024)
AGB-DE: A Corpus for the Automated Legal Assessment of Clauses in German Consumer Contracts
di: Braun, Daniel, et al.
Pubblicazione: (2024)
di: Braun, Daniel, et al.
Pubblicazione: (2024)
NLP-KG: A System for Exploratory Search of Scientific Literature in Natural Language Processing
di: Schopf, Tim, et al.
Pubblicazione: (2024)
di: Schopf, Tim, et al.
Pubblicazione: (2024)
From Competition to Collaboration: Designing Sustainable Mechanisms Between LLMs and Online Forums
di: Fono, Niv, et al.
Pubblicazione: (2026)
di: Fono, Niv, et al.
Pubblicazione: (2026)
Before You Interpret the Profile: Validity Scaling for LLM Metacognitive Self-Report
di: Cacioli, Jon-Paul
Pubblicazione: (2026)
di: Cacioli, Jon-Paul
Pubblicazione: (2026)
Documenti analoghi
-
FActBench: A Benchmark for Fine-grained Automatic Evaluation of LLM-Generated Text in the Medical Domain
di: Afzal, Anum, et al.
Pubblicazione: (2025) -
Aligning What LLMs Do and Say: Towards Self-Consistent Explanations
di: Admoni, Sahar, et al.
Pubblicazione: (2025) -
Can Smaller LLMs do better? Unlocking Cross-Domain Potential through Parameter-Efficient Fine-Tuning for Text Summarization
di: Afzal, Anum, et al.
Pubblicazione: (2025) -
JaccDiv: A Metric and Benchmark for Quantifying Diversity of Generated Marketing Text in the Music Industry
di: Afzal, Anum, et al.
Pubblicazione: (2025) -
More Test-Time Compute Can Hurt: Overestimation Bias in LLM Beam Search
di: Dalal, Gal, et al.
Pubblicazione: (2026)