Patches of Nonlinearity: Instruction Vectors in Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Bigoulaeva, Irina, Rohweder, Jonas, Dutta, Subhabrata, Gurevych, Iryna |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Hierarchical Latent Structures in Data Generation Process Unify Mechanistic Phenomena across Scale
von: Rohweder, Jonas, et al.
Veröffentlicht: (2026)
von: Rohweder, Jonas, et al.
Veröffentlicht: (2026)
The Inherent Limits of Pretrained LLMs: The Unexpected Convergence of Instruction Tuning and In-Context Learning Capabilities
von: Bigoulaeva, Irina, et al.
Veröffentlicht: (2025)
von: Bigoulaeva, Irina, et al.
Veröffentlicht: (2025)
Are Emergent Abilities in Large Language Models just In-Context Learning?
von: Lu, Sheng, et al.
Veröffentlicht: (2023)
von: Lu, Sheng, et al.
Veröffentlicht: (2023)
Factual Self-Awareness in Language Models: Representation, Robustness, and Scaling
von: Tamoyan, Hovhannes, et al.
Veröffentlicht: (2025)
von: Tamoyan, Hovhannes, et al.
Veröffentlicht: (2025)
Reward Modeling for Scientific Writing Evaluation
von: Şahinuç, Furkan, et al.
Veröffentlicht: (2026)
von: Şahinuç, Furkan, et al.
Veröffentlicht: (2026)
Expert Preference-based Evaluation of Automated Related Work Generation
von: Şahinuç, Furkan, et al.
Veröffentlicht: (2025)
von: Şahinuç, Furkan, et al.
Veröffentlicht: (2025)
Hypothesis-Driven Feature Manifold Analysis in LLMs via Supervised Multi-Dimensional Scaling
von: Tiblias, Federico, et al.
Veröffentlicht: (2025)
von: Tiblias, Federico, et al.
Veröffentlicht: (2025)
Co-FactChecker: A Framework for Human-AI Collaborative Claim Verification Using Large Reasoning Models
von: Sahnan, Dhruv, et al.
Veröffentlicht: (2026)
von: Sahnan, Dhruv, et al.
Veröffentlicht: (2026)
Subjective Code Preferences in Experts and Large Language Models
von: Mokhova, Anna, et al.
Veröffentlicht: (2026)
von: Mokhova, Anna, et al.
Veröffentlicht: (2026)
Illusion or Algorithm? Investigating Memorization, Emergence, and Symbolic Processing in In-Context Learning
von: Niu, Jingcheng, et al.
Veröffentlicht: (2025)
von: Niu, Jingcheng, et al.
Veröffentlicht: (2025)
Attribute or Abstain: Large Language Models as Long Document Assistants
von: Buchmann, Jan, et al.
Veröffentlicht: (2024)
von: Buchmann, Jan, et al.
Veröffentlicht: (2024)
Robust Utility-Preserving Text Anonymization Based on Large Language Models
von: Yang, Tianyu, et al.
Veröffentlicht: (2024)
von: Yang, Tianyu, et al.
Veröffentlicht: (2024)
Are Large Language Models Good Classifiers? A Study on Edit Intent Classification in Scientific Document Revisions
von: Ruan, Qian, et al.
Veröffentlicht: (2024)
von: Ruan, Qian, et al.
Veröffentlicht: (2024)
Con Instruction: Universal Jailbreaking of Multimodal Large Language Models via Non-Textual Modalities
von: Geng, Jiahui, et al.
Veröffentlicht: (2025)
von: Geng, Jiahui, et al.
Veröffentlicht: (2025)
Token Weighting for Long-Range Language Modeling
von: Helm, Falko, et al.
Veröffentlicht: (2025)
von: Helm, Falko, et al.
Veröffentlicht: (2025)
Differentially Private Steering for Large Language Model Alignment
von: Goel, Anmol, et al.
Veröffentlicht: (2025)
von: Goel, Anmol, et al.
Veröffentlicht: (2025)
How Quantization Shapes Bias in Large Language Models
von: Marcuzzi, Federico, et al.
Veröffentlicht: (2025)
von: Marcuzzi, Federico, et al.
Veröffentlicht: (2025)
ConspirED: A Dataset for Cognitive Traits of Conspiracy Theories and Large Language Model Safety
von: Bates, Luke, et al.
Veröffentlicht: (2025)
von: Bates, Luke, et al.
Veröffentlicht: (2025)
IRCoder: Intermediate Representations Make Language Models Robust Multilingual Code Generators
von: Paul, Indraneil, et al.
Veröffentlicht: (2024)
von: Paul, Indraneil, et al.
Veröffentlicht: (2024)
Multimodal Large Language Models to Support Real-World Fact-Checking
von: Geng, Jiahui, et al.
Veröffentlicht: (2024)
von: Geng, Jiahui, et al.
Veröffentlicht: (2024)
Cultural Learning-Based Culture Adaptation of Language Models
von: Liu, Chen Cecilia, et al.
Veröffentlicht: (2025)
von: Liu, Chen Cecilia, et al.
Veröffentlicht: (2025)
Transforming Scholarly Landscapes: Influence of Large Language Models on Academic Fields beyond Computer Science
von: Pramanick, Aniket, et al.
Veröffentlicht: (2024)
von: Pramanick, Aniket, et al.
Veröffentlicht: (2024)
Automatic Reviewers Fail to Detect Faulty Reasoning in Research Papers: A New Counterfactual Evaluation Framework
von: Dycke, Nils, et al.
Veröffentlicht: (2025)
von: Dycke, Nils, et al.
Veröffentlicht: (2025)
Like a Good Nearest Neighbor: Practical Content Moderation and Text Classification
von: Bates, Luke, et al.
Veröffentlicht: (2023)
von: Bates, Luke, et al.
Veröffentlicht: (2023)
Citation Failure: Definition, Analysis and Efficient Mitigation
von: Buchmann, Jan, et al.
Veröffentlicht: (2025)
von: Buchmann, Jan, et al.
Veröffentlicht: (2025)
GRITHopper: Decomposition-Free Multi-Hop Dense Retrieval
von: Erker, Justus-Jonas, et al.
Veröffentlicht: (2025)
von: Erker, Justus-Jonas, et al.
Veröffentlicht: (2025)
DARA: Decomposition-Alignment-Reasoning Autonomous Language Agent for Question Answering over Knowledge Graphs
von: Fang, Haishuo, et al.
Veröffentlicht: (2024)
von: Fang, Haishuo, et al.
Veröffentlicht: (2024)
Mechanistic Behavior Editing of Language Models
von: Singh, Joykirat, et al.
Veröffentlicht: (2024)
von: Singh, Joykirat, et al.
Veröffentlicht: (2024)
Re3: A Holistic Framework and Dataset for Modeling Collaborative Document Revision
von: Ruan, Qian, et al.
Veröffentlicht: (2024)
von: Ruan, Qian, et al.
Veröffentlicht: (2024)
Turning Logic Against Itself : Probing Model Defenses Through Contrastive Questions
von: Sachdeva, Rachneet, et al.
Veröffentlicht: (2025)
von: Sachdeva, Rachneet, et al.
Veröffentlicht: (2025)
SciCoQA: Quality Assurance for Scientific Paper--Code Alignment
von: Baumgärtner, Tim, et al.
Veröffentlicht: (2026)
von: Baumgärtner, Tim, et al.
Veröffentlicht: (2026)
FUN with Fisher: Improving Generalization of Adapter-Based Cross-lingual Transfer with Scheduled Unfreezing
von: Liu, Chen Cecilia, et al.
Veröffentlicht: (2023)
von: Liu, Chen Cecilia, et al.
Veröffentlicht: (2023)
SpaRC and SpaRP: Spatial Reasoning Characterization and Path Generation for Understanding Spatial Reasoning Capability of Large Language Models
von: Rizvi, Md Imbesat Hassan, et al.
Veröffentlicht: (2024)
von: Rizvi, Md Imbesat Hassan, et al.
Veröffentlicht: (2024)
A Survey of Confidence Estimation and Calibration in Large Language Models
von: Geng, Jiahui, et al.
Veröffentlicht: (2023)
von: Geng, Jiahui, et al.
Veröffentlicht: (2023)
Stepwise Verification and Remediation of Student Reasoning Errors with Large Language Model Tutors
von: Daheim, Nico, et al.
Veröffentlicht: (2024)
von: Daheim, Nico, et al.
Veröffentlicht: (2024)
$\texttt{LM}^\texttt{2}$: A Simple Society of Language Models Solves Complex Reasoning
von: Juneja, Gurusha, et al.
Veröffentlicht: (2024)
von: Juneja, Gurusha, et al.
Veröffentlicht: (2024)
Efficient Performance Tracking: Leveraging Large Language Models for Automated Construction of Scientific Leaderboards
von: Şahinuç, Furkan, et al.
Veröffentlicht: (2024)
von: Şahinuç, Furkan, et al.
Veröffentlicht: (2024)
Holmes: A Benchmark to Assess the Linguistic Competence of Language Models
von: Waldis, Andreas, et al.
Veröffentlicht: (2024)
von: Waldis, Andreas, et al.
Veröffentlicht: (2024)
Self-Rationalization in the Wild: A Large Scale Out-of-Distribution Evaluation on NLI-related tasks
von: Yang, Jing, et al.
Veröffentlicht: (2025)
von: Yang, Jing, et al.
Veröffentlicht: (2025)
Towards Privacy-aware Mental Health AI Models: Advances, Challenges, and Opportunities
von: Mandal, Aishik, et al.
Veröffentlicht: (2025)
von: Mandal, Aishik, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Hierarchical Latent Structures in Data Generation Process Unify Mechanistic Phenomena across Scale
von: Rohweder, Jonas, et al.
Veröffentlicht: (2026) -
The Inherent Limits of Pretrained LLMs: The Unexpected Convergence of Instruction Tuning and In-Context Learning Capabilities
von: Bigoulaeva, Irina, et al.
Veröffentlicht: (2025) -
Are Emergent Abilities in Large Language Models just In-Context Learning?
von: Lu, Sheng, et al.
Veröffentlicht: (2023) -
Factual Self-Awareness in Language Models: Representation, Robustness, and Scaling
von: Tamoyan, Hovhannes, et al.
Veröffentlicht: (2025) -
Reward Modeling for Scientific Writing Evaluation
von: Şahinuç, Furkan, et al.
Veröffentlicht: (2026)