Are formal and functional linguistic mechanisms dissociated in language models?
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hanna, Michael, Belinkov, Yonatan, Pezzelle, Sandro |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Have Faith in Faithfulness: Going Beyond Circuit Overlap When Finding Model Mechanisms
von: Hanna, Michael, et al.
Veröffentlicht: (2024)
von: Hanna, Michael, et al.
Veröffentlicht: (2024)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
von: Ashuach, Tomer, et al.
Veröffentlicht: (2025)
von: Ashuach, Tomer, et al.
Veröffentlicht: (2025)
REVS: Unlearning Sensitive Information in Language Models via Rank Editing in the Vocabulary Space
von: Ashuach, Tomer, et al.
Veröffentlicht: (2024)
von: Ashuach, Tomer, et al.
Veröffentlicht: (2024)
Distinguishing Ignorance from Error in LLM Hallucinations
von: Simhi, Adi, et al.
Veröffentlicht: (2024)
von: Simhi, Adi, et al.
Veröffentlicht: (2024)
Constructing Benchmarks and Interventions for Combating Hallucinations in LLMs
von: Simhi, Adi, et al.
Veröffentlicht: (2024)
von: Simhi, Adi, et al.
Veröffentlicht: (2024)
A Dataset for Metaphor Detection in Early Medieval Hebrew Poetry
von: Toker, Michael, et al.
Veröffentlicht: (2024)
von: Toker, Michael, et al.
Veröffentlicht: (2024)
Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer
von: Simhi, Adi, et al.
Veröffentlicht: (2025)
von: Simhi, Adi, et al.
Veröffentlicht: (2025)
Masked by Consensus: Disentangling Privileged Knowledge in LLM Correctness
von: Ashuach, Tomer, et al.
Veröffentlicht: (2026)
von: Ashuach, Tomer, et al.
Veröffentlicht: (2026)
ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMs
von: Simhi, Adi, et al.
Veröffentlicht: (2025)
von: Simhi, Adi, et al.
Veröffentlicht: (2025)
Same Task, Different Circuits: Disentangling Modality-Specific Mechanisms in VLMs
von: Nikankin, Yaniv, et al.
Veröffentlicht: (2025)
von: Nikankin, Yaniv, et al.
Veröffentlicht: (2025)
Arithmetic Without Algorithms: Language Models Solve Math With a Bag of Heuristics
von: Nikankin, Yaniv, et al.
Veröffentlicht: (2024)
von: Nikankin, Yaniv, et al.
Veröffentlicht: (2024)
Reasoning Models Know What's Important, and Encode It in Their Activations
von: Nikankin, Yaniv, et al.
Veröffentlicht: (2026)
von: Nikankin, Yaniv, et al.
Veröffentlicht: (2026)
Old Habits Die Hard: How Conversational History Geometrically Traps LLMs
von: Simhi, Adi, et al.
Veröffentlicht: (2026)
von: Simhi, Adi, et al.
Veröffentlicht: (2026)
HACK: Hallucinations Along Certainty and Knowledge Axes
von: Simhi, Adi, et al.
Veröffentlicht: (2025)
von: Simhi, Adi, et al.
Veröffentlicht: (2025)
ReFACT: Updating Text-to-Image Models by Editing the Text Encoder
von: Arad, Dana, et al.
Veröffentlicht: (2023)
von: Arad, Dana, et al.
Veröffentlicht: (2023)
Diffusion Lens: Interpreting Text Encoders in Text-to-Image Pipelines
von: Toker, Michael, et al.
Veröffentlicht: (2024)
von: Toker, Michael, et al.
Veröffentlicht: (2024)
Human-like fleeting memory improves language learning but impairs reading time prediction in transformer language models
von: Thamma, Abishek, et al.
Veröffentlicht: (2025)
von: Thamma, Abishek, et al.
Veröffentlicht: (2025)
Position-aware Automatic Circuit Discovery
von: Haklay, Tal, et al.
Veröffentlicht: (2025)
von: Haklay, Tal, et al.
Veröffentlicht: (2025)
Learning the meanings of function words from grounded language using a visual question answering model
von: Portelance, Eva, et al.
Veröffentlicht: (2023)
von: Portelance, Eva, et al.
Veröffentlicht: (2023)
PeLLE: Encoder-based language models for Brazilian Portuguese based on open data
von: de Mello, Guilherme Lamartine, et al.
Veröffentlicht: (2024)
von: de Mello, Guilherme Lamartine, et al.
Veröffentlicht: (2024)
Towards interpretable models for language proficiency assessment: Predicting the CEFR level of Estonian learner texts
von: Allkivi, Kais
Veröffentlicht: (2026)
von: Allkivi, Kais
Veröffentlicht: (2026)
LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations
von: Orgad, Hadas, et al.
Veröffentlicht: (2024)
von: Orgad, Hadas, et al.
Veröffentlicht: (2024)
Determination of language families using deep learning
von: Lerner, Peter B.
Veröffentlicht: (2024)
von: Lerner, Peter B.
Veröffentlicht: (2024)
Evaluating an evidence-guided reinforcement learning framework in aligning light-parameter large language models with decision-making cognition in psychiatric clinical reasoning
von: Lin, Xinxin, et al.
Veröffentlicht: (2026)
von: Lin, Xinxin, et al.
Veröffentlicht: (2026)
QFS-Composer: Query-focused summarization pipeline for less resourced languages
von: Đuranović, Vuk, et al.
Veröffentlicht: (2026)
von: Đuranović, Vuk, et al.
Veröffentlicht: (2026)
Generative linguistics contribution to artificial intelligence: Where this contribution lies?
von: Shormani, Mohammed Q.
Veröffentlicht: (2024)
von: Shormani, Mohammed Q.
Veröffentlicht: (2024)
Large Language Models Generate Harmful Content Using a Distinct, Unified Mechanism
von: Orgad, Hadas, et al.
Veröffentlicht: (2026)
von: Orgad, Hadas, et al.
Veröffentlicht: (2026)
Latent Planning Emerges with Scale
von: Hanna, Michael, et al.
Veröffentlicht: (2026)
von: Hanna, Michael, et al.
Veröffentlicht: (2026)
CausalGym: Benchmarking causal interpretability methods on linguistic tasks
von: Arora, Aryaman, et al.
Veröffentlicht: (2024)
von: Arora, Aryaman, et al.
Veröffentlicht: (2024)
What makes a language easy to deep-learn? Deep neural networks and humans similarly benefit from compositional structure
von: Galke, Lukas, et al.
Veröffentlicht: (2023)
von: Galke, Lukas, et al.
Veröffentlicht: (2023)
Tracking linguistic information in transformer-based sentence embeddings through targeted sparsification
von: Nastase, Vivi, et al.
Veröffentlicht: (2024)
von: Nastase, Vivi, et al.
Veröffentlicht: (2024)
Unleashing the potential of prompt engineering for large language models
von: Chen, Banghao, et al.
Veröffentlicht: (2023)
von: Chen, Banghao, et al.
Veröffentlicht: (2023)
Kastor: Fine-tuned Small Language Models for Shape-based Active Relation Extraction
von: Celian, Ringwald, et al.
Veröffentlicht: (2025)
von: Celian, Ringwald, et al.
Veröffentlicht: (2025)
Overcoming the Generalization Limits of SLM Finetuning for Shape-Based Extraction of Datatype and Object Properties
von: Ringwald, Célian, et al.
Veröffentlicht: (2025)
von: Ringwald, Célian, et al.
Veröffentlicht: (2025)
State space models can express n-gram languages
von: Nandakumar, Vinoth, et al.
Veröffentlicht: (2023)
von: Nandakumar, Vinoth, et al.
Veröffentlicht: (2023)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
von: Saji, Alan, et al.
Veröffentlicht: (2025)
von: Saji, Alan, et al.
Veröffentlicht: (2025)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
von: Peters, Sydney, et al.
Veröffentlicht: (2025)
von: Peters, Sydney, et al.
Veröffentlicht: (2025)
Low-resource neural machine translation with morphological modeling
von: Nzeyimana, Antoine
Veröffentlicht: (2024)
von: Nzeyimana, Antoine
Veröffentlicht: (2024)
Reframing linguistic bootstrapping as joint inference using visually-grounded grammar induction models
von: Portelance, Eva, et al.
Veröffentlicht: (2024)
von: Portelance, Eva, et al.
Veröffentlicht: (2024)
Human-interpretable clustering of short-text using large language models
von: Miller, Justin K., et al.
Veröffentlicht: (2024)
von: Miller, Justin K., et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Have Faith in Faithfulness: Going Beyond Circuit Overlap When Finding Model Mechanisms
von: Hanna, Michael, et al.
Veröffentlicht: (2024) -
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
von: Ashuach, Tomer, et al.
Veröffentlicht: (2025) -
REVS: Unlearning Sensitive Information in Language Models via Rank Editing in the Vocabulary Space
von: Ashuach, Tomer, et al.
Veröffentlicht: (2024) -
Distinguishing Ignorance from Error in LLM Hallucinations
von: Simhi, Adi, et al.
Veröffentlicht: (2024) -
Constructing Benchmarks and Interventions for Combating Hallucinations in LLMs
von: Simhi, Adi, et al.
Veröffentlicht: (2024)