Gespeichert in:
| Hauptverfasser: | Marks, Samuel, Tegmark, Max |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2310.06824 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Language Models Represent Space and Time
von: Gurnee, Wes, et al.
Veröffentlicht: (2023)
von: Gurnee, Wes, et al.
Veröffentlicht: (2023)
Language Models Use Trigonometry to Do Addition
von: Kantamneni, Subhash, et al.
Veröffentlicht: (2025)
von: Kantamneni, Subhash, et al.
Veröffentlicht: (2025)
Investigating Representation Universality: Case Study on Genealogical Representations
von: Baek, David D., et al.
Veröffentlicht: (2024)
von: Baek, David D., et al.
Veröffentlicht: (2024)
Neural Thermodynamic Laws for Large Language Model Training
von: Liu, Ziming, et al.
Veröffentlicht: (2025)
von: Liu, Ziming, et al.
Veröffentlicht: (2025)
The Geometry of Concepts: Sparse Autoencoder Feature Structure
von: Li, Yuxiao, et al.
Veröffentlicht: (2024)
von: Li, Yuxiao, et al.
Veröffentlicht: (2024)
The Linear Representation Hypothesis and the Geometry of Large Language Models
von: Park, Kiho, et al.
Veröffentlicht: (2023)
von: Park, Kiho, et al.
Veröffentlicht: (2023)
Machine Bullshit: Characterizing the Emergent Disregard for Truth in Large Language Models
von: Liang, Kaiqu, et al.
Veröffentlicht: (2025)
von: Liang, Kaiqu, et al.
Veröffentlicht: (2025)
A Neural Scaling Law from Lottery Ticket Ensembling
von: Liu, Ziming, et al.
Veröffentlicht: (2023)
von: Liu, Ziming, et al.
Veröffentlicht: (2023)
Do Two AI Scientists Agree?
von: Fu, Xinghong, et al.
Veröffentlicht: (2025)
von: Fu, Xinghong, et al.
Veröffentlicht: (2025)
Overthinking the Truth: Understanding how Language Models Process False Demonstrations
von: Halawi, Danny, et al.
Veröffentlicht: (2023)
von: Halawi, Danny, et al.
Veröffentlicht: (2023)
Emergent Structured Representations Support Flexible In-Context Inference in Large Language Models
von: Xu, Ningyu, et al.
Veröffentlicht: (2026)
von: Xu, Ningyu, et al.
Veröffentlicht: (2026)
A Resource Model For Neural Scaling Law
von: Song, Jinyeop, et al.
Veröffentlicht: (2024)
von: Song, Jinyeop, et al.
Veröffentlicht: (2024)
Emergent Hierarchical Structure in Large Language Models: An Information-Theoretic Framework for Multi-Scale Representation
von: Zhang, Yukin, et al.
Veröffentlicht: (2025)
von: Zhang, Yukin, et al.
Veröffentlicht: (2025)
How Do Transformers "Do" Physics? Investigating the Simple Harmonic Oscillator
von: Kantamneni, Subhash, et al.
Veröffentlicht: (2024)
von: Kantamneni, Subhash, et al.
Veröffentlicht: (2024)
Convergent Linear Representations of Emergent Misalignment
von: Soligo, Anna, et al.
Veröffentlicht: (2025)
von: Soligo, Anna, et al.
Veröffentlicht: (2025)
Zero-Shot Referring Expression Comprehension via Vison-Language True/False Verification
von: Liu, Jeffrey, et al.
Veröffentlicht: (2025)
von: Liu, Jeffrey, et al.
Veröffentlicht: (2025)
Context Structure Reshapes the Representational Geometry of Language Models
von: Hosseini, Eghbal A., et al.
Veröffentlicht: (2026)
von: Hosseini, Eghbal A., et al.
Veröffentlicht: (2026)
Measuring Representation Robustness in Large Language Models for Geometry
von: Jawandhia, Vedant, et al.
Veröffentlicht: (2026)
von: Jawandhia, Vedant, et al.
Veröffentlicht: (2026)
Uncovering Emergent Physics Representations Learned In-Context by Large Language Models
von: Song, Yeongwoo, et al.
Veröffentlicht: (2025)
von: Song, Yeongwoo, et al.
Veröffentlicht: (2025)
Liars' Bench: Evaluating Lie Detectors for Language Models
von: Kretschmar, Kieron, et al.
Veröffentlicht: (2025)
von: Kretschmar, Kieron, et al.
Veröffentlicht: (2025)
Are Sparse Autoencoders Useful? A Case Study in Sparse Probing
von: Kantamneni, Subhash, et al.
Veröffentlicht: (2025)
von: Kantamneni, Subhash, et al.
Veröffentlicht: (2025)
TruthX: Alleviating Hallucinations by Editing Large Language Models in Truthful Space
von: Zhang, Shaolei, et al.
Veröffentlicht: (2024)
von: Zhang, Shaolei, et al.
Veröffentlicht: (2024)
The Remarkable Robustness of LLMs: Stages of Inference?
von: Lad, Vedang, et al.
Veröffentlicht: (2024)
von: Lad, Vedang, et al.
Veröffentlicht: (2024)
Scaling Laws For Scalable Oversight
von: Engels, Joshua, et al.
Veröffentlicht: (2025)
von: Engels, Joshua, et al.
Veröffentlicht: (2025)
TruthStance: An Annotated Dataset of Conversations on Truth Social
von: Ameen, Fathima, et al.
Veröffentlicht: (2026)
von: Ameen, Fathima, et al.
Veröffentlicht: (2026)
Unsupervised decoding of encoded reasoning using language model interpretability
von: Fang, Ching, et al.
Veröffentlicht: (2025)
von: Fang, Ching, et al.
Veröffentlicht: (2025)
Truth Forest: Toward Multi-Scale Truthfulness in Large Language Models through Intervention without Tuning
von: Chen, Zhongzhi, et al.
Veröffentlicht: (2023)
von: Chen, Zhongzhi, et al.
Veröffentlicht: (2023)
Steering Evaluation-Aware Language Models to Act Like They Are Deployed
von: Hua, Tim Tian, et al.
Veröffentlicht: (2025)
von: Hua, Tim Tian, et al.
Veröffentlicht: (2025)
TruthEval: A Dataset to Evaluate LLM Truthfulness and Reliability
von: Khatun, Aisha, et al.
Veröffentlicht: (2024)
von: Khatun, Aisha, et al.
Veröffentlicht: (2024)
Unlocking Emergent Modularity in Large Language Models
von: Qiu, Zihan, et al.
Veröffentlicht: (2023)
von: Qiu, Zihan, et al.
Veröffentlicht: (2023)
Emergent Introspective Awareness in Large Language Models
von: Lindsey, Jack
Veröffentlicht: (2026)
von: Lindsey, Jack
Veröffentlicht: (2026)
Ranking Large Language Models without Ground Truth
von: Dhurandhar, Amit, et al.
Veröffentlicht: (2024)
von: Dhurandhar, Amit, et al.
Veröffentlicht: (2024)
Omni Geometry Representation Learning vs Large Language Models for Geospatial Entity Resolution
von: Wijegunarathna, Kalana, et al.
Veröffentlicht: (2025)
von: Wijegunarathna, Kalana, et al.
Veröffentlicht: (2025)
Large Language Models Can Take False First Steps at Inference-time Planning
von: Yan, Haijiang, et al.
Veröffentlicht: (2026)
von: Yan, Haijiang, et al.
Veröffentlicht: (2026)
The Geometry of Refusal in Large Language Models: Concept Cones and Representational Independence
von: Wollschläger, Tom, et al.
Veröffentlicht: (2025)
von: Wollschläger, Tom, et al.
Veröffentlicht: (2025)
Beyond Linear Steering: Unified Multi-Attribute Control for Language Models
von: Oozeer, Narmeen, et al.
Veröffentlicht: (2025)
von: Oozeer, Narmeen, et al.
Veröffentlicht: (2025)
Latent Structure of Affective Representations in Large Language Models
von: Choi, Benjamin J., et al.
Veröffentlicht: (2026)
von: Choi, Benjamin J., et al.
Veröffentlicht: (2026)
The Geometries of Truth Are Orthogonal Across Tasks
von: Azizian, Waiss, et al.
Veröffentlicht: (2025)
von: Azizian, Waiss, et al.
Veröffentlicht: (2025)
Why Cannot Large Language Models Ever Make True Correct Reasoning?
von: Cheng, Jingde
Veröffentlicht: (2025)
von: Cheng, Jingde
Veröffentlicht: (2025)
Towards Reliable Truth-Aligned Uncertainty Estimation in Large Language Models
von: Srey, Ponhvoan, et al.
Veröffentlicht: (2026)
von: Srey, Ponhvoan, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Language Models Represent Space and Time
von: Gurnee, Wes, et al.
Veröffentlicht: (2023) -
Language Models Use Trigonometry to Do Addition
von: Kantamneni, Subhash, et al.
Veröffentlicht: (2025) -
Investigating Representation Universality: Case Study on Genealogical Representations
von: Baek, David D., et al.
Veröffentlicht: (2024) -
Neural Thermodynamic Laws for Large Language Model Training
von: Liu, Ziming, et al.
Veröffentlicht: (2025) -
The Geometry of Concepts: Sparse Autoencoder Feature Structure
von: Li, Yuxiao, et al.
Veröffentlicht: (2024)