Position-aware Automatic Circuit Discovery
Fuente:
arXiv
Salvato in:
| Autori principali: | Haklay, Tal, Orgad, Hadas, Bau, David, Mueller, Aaron, Belinkov, Yonatan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ReFACT: Updating Text-to-Image Models by Editing the Text Encoder
di: Arad, Dana, et al.
Pubblicazione: (2023)
di: Arad, Dana, et al.
Pubblicazione: (2023)
LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations
di: Orgad, Hadas, et al.
Pubblicazione: (2024)
di: Orgad, Hadas, et al.
Pubblicazione: (2024)
Pitfalls in Evaluating Interpretability Agents
di: Haklay, Tal, et al.
Pubblicazione: (2026)
di: Haklay, Tal, et al.
Pubblicazione: (2026)
LLM Questionnaire Completion for Automatic Psychiatric Assessment
di: Rosenman, Gony, et al.
Pubblicazione: (2024)
di: Rosenman, Gony, et al.
Pubblicazione: (2024)
Arithmetic Without Algorithms: Language Models Solve Math With a Bag of Heuristics
di: Nikankin, Yaniv, et al.
Pubblicazione: (2024)
di: Nikankin, Yaniv, et al.
Pubblicazione: (2024)
Automatic identification of diagnosis from hospital discharge letters via weakly-supervised Natural Language Processing
di: Torri, Vittorio, et al.
Pubblicazione: (2024)
di: Torri, Vittorio, et al.
Pubblicazione: (2024)
Mitigating Position-Shift Failures in Text-Based Modular Arithmetic via Position Curriculum and Template Diversity
di: Yudin, Nikolay
Pubblicazione: (2026)
di: Yudin, Nikolay
Pubblicazione: (2026)
Same Task, Different Circuits: Disentangling Modality-Specific Mechanisms in VLMs
di: Nikankin, Yaniv, et al.
Pubblicazione: (2025)
di: Nikankin, Yaniv, et al.
Pubblicazione: (2025)
Zero- and Few-Shot Prompting with LLMs: A Comparative Study with Fine-tuned Models for Bangla Sentiment Analysis
di: Hasan, Md. Arid, et al.
Pubblicazione: (2023)
di: Hasan, Md. Arid, et al.
Pubblicazione: (2023)
$FastDoc$: Domain-Specific Fast Continual Pre-training Technique using Document-Level Metadata and Taxonomy
di: Nandy, Abhilash, et al.
Pubblicazione: (2023)
di: Nandy, Abhilash, et al.
Pubblicazione: (2023)
FairLangProc: A Python package for fairness in NLP
di: Pérez-Peralta, Arturo, et al.
Pubblicazione: (2025)
di: Pérez-Peralta, Arturo, et al.
Pubblicazione: (2025)
A-VERT: Agnostic Verification with Embedding Ranking Targets
di: Aguirre, Nicolás, et al.
Pubblicazione: (2025)
di: Aguirre, Nicolás, et al.
Pubblicazione: (2025)
Pivot Language for Low-Resource Machine Translation
di: Talwar, Abhimanyu, et al.
Pubblicazione: (2025)
di: Talwar, Abhimanyu, et al.
Pubblicazione: (2025)
Unilogit: Robust Machine Unlearning for LLMs Using Uniform-Target Self-Distillation
di: Vasilev, Stefan, et al.
Pubblicazione: (2025)
di: Vasilev, Stefan, et al.
Pubblicazione: (2025)
A Flexible Large Language Models Guardrail Development Methodology Applied to Off-Topic Prompt Detection
di: Chua, Gabriel, et al.
Pubblicazione: (2024)
di: Chua, Gabriel, et al.
Pubblicazione: (2024)
Preserving Empirical Probabilities in BERT for Small-sample Clinical Entity Recognition
di: Rehman, Abdul, et al.
Pubblicazione: (2024)
di: Rehman, Abdul, et al.
Pubblicazione: (2024)
Have Faith in Faithfulness: Going Beyond Circuit Overlap When Finding Model Mechanisms
di: Hanna, Michael, et al.
Pubblicazione: (2024)
di: Hanna, Michael, et al.
Pubblicazione: (2024)
Large Language Models Generate Harmful Content Using a Distinct, Unified Mechanism
di: Orgad, Hadas, et al.
Pubblicazione: (2026)
di: Orgad, Hadas, et al.
Pubblicazione: (2026)
Arithmetic in the Wild: Llama uses Base-10 Addition to Reason About Cyclic Concepts
di: Feucht, Sheridan, et al.
Pubblicazione: (2026)
di: Feucht, Sheridan, et al.
Pubblicazione: (2026)
SentiCSE: A Sentiment-aware Contrastive Sentence Embedding Framework with Sentiment-guided Textual Similarity
di: Kim, Jaemin, et al.
Pubblicazione: (2024)
di: Kim, Jaemin, et al.
Pubblicazione: (2024)
Fine-tuning Large Language Models for Entity Matching
di: Steiner, Aaron, et al.
Pubblicazione: (2024)
di: Steiner, Aaron, et al.
Pubblicazione: (2024)
Cache-to-Cache: Direct Semantic Communication Between Large Language Models
di: Fu, Tianyu, et al.
Pubblicazione: (2025)
di: Fu, Tianyu, et al.
Pubblicazione: (2025)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
di: Ashuach, Tomer, et al.
Pubblicazione: (2025)
di: Ashuach, Tomer, et al.
Pubblicazione: (2025)
Semantic Convergence: Investigating Shared Representations Across Scaled LLMs
di: Son, Daniel, et al.
Pubblicazione: (2025)
di: Son, Daniel, et al.
Pubblicazione: (2025)
Multi-Model Synthetic Training for Mission-Critical Small Language Models
di: Platt, Nolan, et al.
Pubblicazione: (2025)
di: Platt, Nolan, et al.
Pubblicazione: (2025)
Transactional Attention: Semantic Sponsorship for KV-Cache Retention
di: Basu, Abhinaba
Pubblicazione: (2026)
di: Basu, Abhinaba
Pubblicazione: (2026)
Low-Resource Neural Machine Translation Using Recurrent Neural Networks and Transfer Learning: A Case Study on English-to-Igbo
di: Ekle, Ocheme Anthony, et al.
Pubblicazione: (2025)
di: Ekle, Ocheme Anthony, et al.
Pubblicazione: (2025)
Exploring the Effectiveness of Instruction Tuning in Biomedical Language Processing
di: Rohanian, Omid, et al.
Pubblicazione: (2023)
di: Rohanian, Omid, et al.
Pubblicazione: (2023)
Contrasting Linguistic Patterns in Human and LLM-Generated News Text
di: Muñoz-Ortiz, Alberto, et al.
Pubblicazione: (2023)
di: Muñoz-Ortiz, Alberto, et al.
Pubblicazione: (2023)
Nested Named Entity Recognition as Single-Pass Sequence Labeling
di: Muñoz-Ortiz, Alberto, et al.
Pubblicazione: (2025)
di: Muñoz-Ortiz, Alberto, et al.
Pubblicazione: (2025)
Dancing in the syntax forest: fast, accurate and explainable sentiment analysis with SALSA
di: Gómez-Rodríguez, Carlos, et al.
Pubblicazione: (2024)
di: Gómez-Rodríguez, Carlos, et al.
Pubblicazione: (2024)
PaperAudit-Bench: Benchmarking Error Detection in Research Papers for Critical Automated Peer Review
di: Tu, Songjun, et al.
Pubblicazione: (2026)
di: Tu, Songjun, et al.
Pubblicazione: (2026)
Lightweight Transformers for Clinical Natural Language Processing
di: Rohanian, Omid, et al.
Pubblicazione: (2023)
di: Rohanian, Omid, et al.
Pubblicazione: (2023)
Co-NAML-LSTUR: A Combined Model with Attentive Multi-View Learning and Long- and Short-term User Representations for News Recommendation
di: Nguyen, Minh Hoang, et al.
Pubblicazione: (2025)
di: Nguyen, Minh Hoang, et al.
Pubblicazione: (2025)
Diffusion Lens: Interpreting Text Encoders in Text-to-Image Pipelines
di: Toker, Michael, et al.
Pubblicazione: (2024)
di: Toker, Michael, et al.
Pubblicazione: (2024)
Mechanistic Analysis of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning
di: Imanov, Olaf Yunus Laitinen
Pubblicazione: (2026)
di: Imanov, Olaf Yunus Laitinen
Pubblicazione: (2026)
Latent Object Permanence: Topological Phase Transitions, Free-Energy Principles, and Renormalization Group Flows in Deep Transformer Manifolds
di: Alpay, Faruk, et al.
Pubblicazione: (2026)
di: Alpay, Faruk, et al.
Pubblicazione: (2026)
Inference acceleration for large language models using "stairs" assisted greedy generation
di: Grigaliūnas, Domas, et al.
Pubblicazione: (2024)
di: Grigaliūnas, Domas, et al.
Pubblicazione: (2024)
Strategic Doctrine Language Models (sdLM): A Learning-System Framework for Doctrinal Consistency and Geopolitical Forecasting
di: Imanov, Olaf Yunus Laitinen, et al.
Pubblicazione: (2026)
di: Imanov, Olaf Yunus Laitinen, et al.
Pubblicazione: (2026)
Hybrid Gated Flow (HGF): Stabilizing 1.58-bit LLMs via Selective Low-Rank Correction
di: Pizzo, David Alejandro Trejo
Pubblicazione: (2026)
di: Pizzo, David Alejandro Trejo
Pubblicazione: (2026)
Documenti analoghi
-
ReFACT: Updating Text-to-Image Models by Editing the Text Encoder
di: Arad, Dana, et al.
Pubblicazione: (2023) -
LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations
di: Orgad, Hadas, et al.
Pubblicazione: (2024) -
Pitfalls in Evaluating Interpretability Agents
di: Haklay, Tal, et al.
Pubblicazione: (2026) -
LLM Questionnaire Completion for Automatic Psychiatric Assessment
di: Rosenman, Gony, et al.
Pubblicazione: (2024) -
Arithmetic Without Algorithms: Language Models Solve Math With a Bag of Heuristics
di: Nikankin, Yaniv, et al.
Pubblicazione: (2024)