DEPTH: Discourse Education through Pre-Training Hierarchically
Fuente:
arXiv
Salvato in:
| Autori principali: | Bamberger, Zachary, Glick, Ofek, Baskin, Chaim, Belinkov, Yonatan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Context-aware Prompt Tuning: Advancing In-Context Learning with Adversarial Methods
di: Blau, Tsachi, et al.
Pubblicazione: (2024)
di: Blau, Tsachi, et al.
Pubblicazione: (2024)
ContraSim -- Analyzing Neural Representations Based on Contrastive Learning
di: Rahamim, Adir, et al.
Pubblicazione: (2023)
di: Rahamim, Adir, et al.
Pubblicazione: (2023)
Fast Forwarding Low-Rank Training
di: Rahamim, Adir, et al.
Pubblicazione: (2024)
di: Rahamim, Adir, et al.
Pubblicazione: (2024)
Back Attention: Understanding and Enhancing Multi-Hop Reasoning in Large Language Models
di: Yu, Zeping, et al.
Pubblicazione: (2025)
di: Yu, Zeping, et al.
Pubblicazione: (2025)
MATCH: Task-Driven Code Evaluation through Contrastive Learning
di: Ghoummaid, Marah, et al.
Pubblicazione: (2025)
di: Ghoummaid, Marah, et al.
Pubblicazione: (2025)
Leveraging Prototypical Representations for Mitigating Social Bias without Demographic Information
di: Iskander, Shadi, et al.
Pubblicazione: (2024)
di: Iskander, Shadi, et al.
Pubblicazione: (2024)
Concept-Best-Matching: Evaluating Compositionality in Emergent Communication
di: Carmeli, Boaz, et al.
Pubblicazione: (2024)
di: Carmeli, Boaz, et al.
Pubblicazione: (2024)
Inside-Out: Hidden Factual Knowledge in LLMs
di: Gekhman, Zorik, et al.
Pubblicazione: (2025)
di: Gekhman, Zorik, et al.
Pubblicazione: (2025)
REVS: Unlearning Sensitive Information in Language Models via Rank Editing in the Vocabulary Space
di: Ashuach, Tomer, et al.
Pubblicazione: (2024)
di: Ashuach, Tomer, et al.
Pubblicazione: (2024)
Are formal and functional linguistic mechanisms dissociated in language models?
di: Hanna, Michael, et al.
Pubblicazione: (2025)
di: Hanna, Michael, et al.
Pubblicazione: (2025)
Hysteresis Activation Function for Efficient Inference
di: Kimhi, Moshe, et al.
Pubblicazione: (2024)
di: Kimhi, Moshe, et al.
Pubblicazione: (2024)
SAEs Are Good for Steering -- If You Select the Right Features
di: Arad, Dana, et al.
Pubblicazione: (2025)
di: Arad, Dana, et al.
Pubblicazione: (2025)
Planted in Pretraining, Swayed by Finetuning: A Case Study on the Origins of Cognitive Biases in LLMs
di: Itzhak, Itay, et al.
Pubblicazione: (2025)
di: Itzhak, Itay, et al.
Pubblicazione: (2025)
LLM4SFC: Sequential Function Chart Generation via Large Language Models
di: Glick, Ofek, et al.
Pubblicazione: (2025)
di: Glick, Ofek, et al.
Pubblicazione: (2025)
Have Faith in Faithfulness: Going Beyond Circuit Overlap When Finding Model Mechanisms
di: Hanna, Michael, et al.
Pubblicazione: (2024)
di: Hanna, Michael, et al.
Pubblicazione: (2024)
ReFACT: Updating Text-to-Image Models by Editing the Text Encoder
di: Arad, Dana, et al.
Pubblicazione: (2023)
di: Arad, Dana, et al.
Pubblicazione: (2023)
Measuring Chain of Thought Faithfulness by Unlearning Reasoning Steps
di: Tutek, Martin, et al.
Pubblicazione: (2025)
di: Tutek, Martin, et al.
Pubblicazione: (2025)
Distinguishing Ignorance from Error in LLM Hallucinations
di: Simhi, Adi, et al.
Pubblicazione: (2024)
di: Simhi, Adi, et al.
Pubblicazione: (2024)
Constructing Benchmarks and Interventions for Combating Hallucinations in LLMs
di: Simhi, Adi, et al.
Pubblicazione: (2024)
di: Simhi, Adi, et al.
Pubblicazione: (2024)
Follow the Flow: On Information Flow Across Textual Tokens in Text-to-Image Models
di: Kaplan, Guy, et al.
Pubblicazione: (2025)
di: Kaplan, Guy, et al.
Pubblicazione: (2025)
Silent Tokens, Loud Effects: Padding in LLMs
di: Himelstein, Rom, et al.
Pubblicazione: (2025)
di: Himelstein, Rom, et al.
Pubblicazione: (2025)
Backward Lens: Projecting Language Model Gradients into the Vocabulary Space
di: Katz, Shahar, et al.
Pubblicazione: (2024)
di: Katz, Shahar, et al.
Pubblicazione: (2024)
From Feelings to Metrics: Understanding and Formalizing How Users Vibe-Test LLMs
di: Itzhak, Itay, et al.
Pubblicazione: (2026)
di: Itzhak, Itay, et al.
Pubblicazione: (2026)
Differentiable Faithfulness Alignment for Cross-Model Circuit Transfer
di: Shao, Shun, et al.
Pubblicazione: (2026)
di: Shao, Shun, et al.
Pubblicazione: (2026)
Arithmetic Without Algorithms: Language Models Solve Math With a Bag of Heuristics
di: Nikankin, Yaniv, et al.
Pubblicazione: (2024)
di: Nikankin, Yaniv, et al.
Pubblicazione: (2024)
Same Task, Different Circuits: Disentangling Modality-Specific Mechanisms in VLMs
di: Nikankin, Yaniv, et al.
Pubblicazione: (2025)
di: Nikankin, Yaniv, et al.
Pubblicazione: (2025)
DeLeaker: Dynamic Inference-Time Reweighting For Semantic Leakage Mitigation in Text-to-Image Models
di: Ventura, Mor, et al.
Pubblicazione: (2025)
di: Ventura, Mor, et al.
Pubblicazione: (2025)
Answer, Assemble, Ace: Understanding How LMs Answer Multiple Choice Questions
di: Wiegreffe, Sarah, et al.
Pubblicazione: (2024)
di: Wiegreffe, Sarah, et al.
Pubblicazione: (2024)
You Had One Job: Per-Task Quantization Using LLMs' Hidden Representations
di: LeVi, Amit, et al.
Pubblicazione: (2025)
di: LeVi, Amit, et al.
Pubblicazione: (2025)
Will it Merge? On The Causes of Model Mergeability
di: Rahamim, Adir, et al.
Pubblicazione: (2026)
di: Rahamim, Adir, et al.
Pubblicazione: (2026)
Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer
di: Simhi, Adi, et al.
Pubblicazione: (2025)
di: Simhi, Adi, et al.
Pubblicazione: (2025)
Fine-Tuning Enhances Existing Mechanisms: A Case Study on Entity Tracking
di: Prakash, Nikhil, et al.
Pubblicazione: (2024)
di: Prakash, Nikhil, et al.
Pubblicazione: (2024)
A Dataset for Metaphor Detection in Early Medieval Hebrew Poetry
di: Toker, Michael, et al.
Pubblicazione: (2024)
di: Toker, Michael, et al.
Pubblicazione: (2024)
Masked by Consensus: Disentangling Privileged Knowledge in LLM Correctness
di: Ashuach, Tomer, et al.
Pubblicazione: (2026)
di: Ashuach, Tomer, et al.
Pubblicazione: (2026)
Reasoning Models Know What's Important, and Encode It in Their Activations
di: Nikankin, Yaniv, et al.
Pubblicazione: (2026)
di: Nikankin, Yaniv, et al.
Pubblicazione: (2026)
Unsupervised Translation of Emergent Communication
di: Levy, Ido, et al.
Pubblicazione: (2025)
di: Levy, Ido, et al.
Pubblicazione: (2025)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
di: Ashuach, Tomer, et al.
Pubblicazione: (2025)
di: Ashuach, Tomer, et al.
Pubblicazione: (2025)
ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMs
di: Simhi, Adi, et al.
Pubblicazione: (2025)
di: Simhi, Adi, et al.
Pubblicazione: (2025)
Diffusion Lens: Interpreting Text Encoders in Text-to-Image Pipelines
di: Toker, Michael, et al.
Pubblicazione: (2024)
di: Toker, Michael, et al.
Pubblicazione: (2024)
Old Habits Die Hard: How Conversational History Geometrically Traps LLMs
di: Simhi, Adi, et al.
Pubblicazione: (2026)
di: Simhi, Adi, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Context-aware Prompt Tuning: Advancing In-Context Learning with Adversarial Methods
di: Blau, Tsachi, et al.
Pubblicazione: (2024) -
ContraSim -- Analyzing Neural Representations Based on Contrastive Learning
di: Rahamim, Adir, et al.
Pubblicazione: (2023) -
Fast Forwarding Low-Rank Training
di: Rahamim, Adir, et al.
Pubblicazione: (2024) -
Back Attention: Understanding and Enhancing Multi-Hop Reasoning in Large Language Models
di: Yu, Zeping, et al.
Pubblicazione: (2025) -
MATCH: Task-Driven Code Evaluation through Contrastive Learning
di: Ghoummaid, Marah, et al.
Pubblicazione: (2025)