The Truthfulness Spectrum Hypothesis
Fuente:
arXiv
Guardado en:
| Autores principales: | Ying, Zhuofan Josh, Ravfogel, Shauli, Kriegeskorte, Nikolaus, Hase, Peter |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Preserving Task-Relevant Information Under Linear Concept Removal
por: Holstege, Floris, et al.
Publicado: (2025)
por: Holstege, Floris, et al.
Publicado: (2025)
Log-linear Guardedness and its Implications
por: Ravfogel, Shauli, et al.
Publicado: (2022)
por: Ravfogel, Shauli, et al.
Publicado: (2022)
BitFit: Simple Parameter-efficient Fine-tuning for Transformer-based Masked Language-models
por: Ben-Zaken, Elad, et al.
Publicado: (2021)
por: Ben-Zaken, Elad, et al.
Publicado: (2021)
The Topology and Geometry of Neural Representations
por: Lin, Baihan, et al.
Publicado: (2023)
por: Lin, Baihan, et al.
Publicado: (2023)
Discrete Diffusion Models Exploit Asymmetry to Solve Lookahead Planning Tasks
por: Trainin, Itamar, et al.
Publicado: (2026)
por: Trainin, Itamar, et al.
Publicado: (2026)
Kernelized Concept Erasure
por: Ravfogel, Shauli, et al.
Publicado: (2022)
por: Ravfogel, Shauli, et al.
Publicado: (2022)
Linear Adversarial Concept Erasure
por: Ravfogel, Shauli, et al.
Publicado: (2022)
por: Ravfogel, Shauli, et al.
Publicado: (2022)
Gumbel Counterfactual Generation From Language Models
por: Ravfogel, Shauli, et al.
Publicado: (2024)
por: Ravfogel, Shauli, et al.
Publicado: (2024)
A Practical Method for Generating String Counterfactuals
por: Avitan, Matan, et al.
Publicado: (2024)
por: Avitan, Matan, et al.
Publicado: (2024)
Diversity Over Quantity: A Lesson From Few Shot Relation Classification
por: Cohen, Amir DN, et al.
Publicado: (2024)
por: Cohen, Amir DN, et al.
Publicado: (2024)
Transformer brain encoders explain human high-level visual responses
por: Adeli, Hossein, et al.
Publicado: (2025)
por: Adeli, Hossein, et al.
Publicado: (2025)
The Role of Language Imbalance in Cross-lingual Generalisation: Insights from Cloned Language Experiments
por: Schäfer, Anton, et al.
Publicado: (2024)
por: Schäfer, Anton, et al.
Publicado: (2024)
Growing a Neural Network in Breadth, Depth, and Time
por: Butkus, Eivinas, et al.
Publicado: (2026)
por: Butkus, Eivinas, et al.
Publicado: (2026)
Description-Based Text Similarity
por: Ravfogel, Shauli, et al.
Publicado: (2023)
por: Ravfogel, Shauli, et al.
Publicado: (2023)
Representation Surgery: Theory and Practice of Affine Steering
por: Singh, Shashwat, et al.
Publicado: (2024)
por: Singh, Shashwat, et al.
Publicado: (2024)
LEACE: Perfect linear concept erasure in closed form
por: Belrose, Nora, et al.
Publicado: (2023)
por: Belrose, Nora, et al.
Publicado: (2023)
Emergence of Linear Truth Encodings in Language Models
por: Ravfogel, Shauli, et al.
Publicado: (2025)
por: Ravfogel, Shauli, et al.
Publicado: (2025)
Can LLMs Introspect? A Reality Check
por: Singh, Shashwat, et al.
Publicado: (2026)
por: Singh, Shashwat, et al.
Publicado: (2026)
On Affine Homotopy between Language Encoders
por: Chan, Robin SM, et al.
Publicado: (2024)
por: Chan, Robin SM, et al.
Publicado: (2024)
The Unreasonable Effectiveness of Easy Training Data for Hard Tasks
por: Hase, Peter, et al.
Publicado: (2024)
por: Hase, Peter, et al.
Publicado: (2024)
Liquid Networks with Mixture Density Heads for Efficient Imitation Learning
por: Correll, Nikolaus
Publicado: (2026)
por: Correll, Nikolaus
Publicado: (2026)
SpecBound: Adaptive Bounded Self-Speculation with Layer-wise Confidence Calibration
por: Wen, Zhuofan, et al.
Publicado: (2026)
por: Wen, Zhuofan, et al.
Publicado: (2026)
Where is the Truth? The Risk of Getting Confounded in a Continual World
por: Busch, Florian Peter, et al.
Publicado: (2024)
por: Busch, Florian Peter, et al.
Publicado: (2024)
Geometric Factual Recall in Transformers
por: Ravfogel, Shauli, et al.
Publicado: (2026)
por: Ravfogel, Shauli, et al.
Publicado: (2026)
State over Tokens: Characterizing the Role of Reasoning Tokens
por: Levy, Mosh, et al.
Publicado: (2025)
por: Levy, Mosh, et al.
Publicado: (2025)
Truthful Elicitation of Imprecise Forecasts
por: Singh, Anurag, et al.
Publicado: (2025)
por: Singh, Anurag, et al.
Publicado: (2025)
Hypothesis Testing the Circuit Hypothesis in LLMs
por: Shi, Claudia, et al.
Publicado: (2024)
por: Shi, Claudia, et al.
Publicado: (2024)
Pruning Weights but Not Truth: Safeguarding Truthfulness While Pruning LLMs
por: Fu, Yao, et al.
Publicado: (2025)
por: Fu, Yao, et al.
Publicado: (2025)
The Deleuzian Representation Hypothesis
por: Cornet, Clément, et al.
Publicado: (2025)
por: Cornet, Clément, et al.
Publicado: (2025)
Strategic Hypothesis Testing
por: Hossain, Safwan, et al.
Publicado: (2025)
por: Hossain, Safwan, et al.
Publicado: (2025)
Truthfulness of Decision-Theoretic Calibration Measures
por: Qiao, Mingda, et al.
Publicado: (2025)
por: Qiao, Mingda, et al.
Publicado: (2025)
Linear Convergence of Diffusion Models Under the Manifold Hypothesis
por: Potaptchik, Peter, et al.
Publicado: (2024)
por: Potaptchik, Peter, et al.
Publicado: (2024)
Ethnography and Machine Learning: Synergies and New Directions
por: Li, Zhuofan, et al.
Publicado: (2024)
por: Li, Zhuofan, et al.
Publicado: (2024)
Speculative Decoding with CTC-based Draft Model for LLM Inference Acceleration
por: Wen, Zhuofan, et al.
Publicado: (2024)
por: Wen, Zhuofan, et al.
Publicado: (2024)
Implicit Hypothesis Testing and Divergence Preservation in Neural Network Representations
por: Aksoy, Kadircan, et al.
Publicado: (2026)
por: Aksoy, Kadircan, et al.
Publicado: (2026)
Sequence-to-Sequence Models with Attention Mechanistically Map to the Architecture of Human Memory Search
por: Salvatore, Nikolaus, et al.
Publicado: (2025)
por: Salvatore, Nikolaus, et al.
Publicado: (2025)
The Ends Justify the Thoughts: RL-Induced Motivated Reasoning in LLM CoTs
por: Howe, Nikolaus, et al.
Publicado: (2025)
por: Howe, Nikolaus, et al.
Publicado: (2025)
SPoRt -- Safe Policy Ratio: Certified Training and Deployment of Task Policies in Model-Free RL
por: Cloete, Jacques, et al.
Publicado: (2025)
por: Cloete, Jacques, et al.
Publicado: (2025)
TruthRL: Incentivizing Truthful LLMs via Reinforcement Learning
por: Wei, Zhepei, et al.
Publicado: (2025)
por: Wei, Zhepei, et al.
Publicado: (2025)
Truthfulness of Calibration Measures
por: Haghtalab, Nika, et al.
Publicado: (2024)
por: Haghtalab, Nika, et al.
Publicado: (2024)
Ejemplares similares
-
Preserving Task-Relevant Information Under Linear Concept Removal
por: Holstege, Floris, et al.
Publicado: (2025) -
Log-linear Guardedness and its Implications
por: Ravfogel, Shauli, et al.
Publicado: (2022) -
BitFit: Simple Parameter-efficient Fine-tuning for Transformer-based Masked Language-models
por: Ben-Zaken, Elad, et al.
Publicado: (2021) -
The Topology and Geometry of Neural Representations
por: Lin, Baihan, et al.
Publicado: (2023) -
Discrete Diffusion Models Exploit Asymmetry to Solve Lookahead Planning Tasks
por: Trainin, Itamar, et al.
Publicado: (2026)