Efficient and accurate steering of Large Language Models through attention-guided feature learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Davarmanesh, Parmida, Wilson, Ashia, Radhakrishnan, Adityanarayanan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Catapults in SGD: spikes in the training loss and their impact on generalization through feature learning
von: Zhu, Libin, et al.
Veröffentlicht: (2023)
von: Zhu, Libin, et al.
Veröffentlicht: (2023)
xRFM: Accurate, scalable, and interpretable feature learning models for tabular data
von: Beaglehole, Daniel, et al.
Veröffentlicht: (2025)
von: Beaglehole, Daniel, et al.
Veröffentlicht: (2025)
Toward universal steering and monitoring of AI models
von: Beaglehole, Daniel, et al.
Veröffentlicht: (2025)
von: Beaglehole, Daniel, et al.
Veröffentlicht: (2025)
Linear Recursive Feature Machines provably recover low-rank matrices
von: Radhakrishnan, Adityanarayanan, et al.
Veröffentlicht: (2024)
von: Radhakrishnan, Adityanarayanan, et al.
Veröffentlicht: (2024)
Context-Scaling versus Task-Scaling in In-Context Learning
von: Abedsoltan, Amirhesam, et al.
Veröffentlicht: (2024)
von: Abedsoltan, Amirhesam, et al.
Veröffentlicht: (2024)
The Weight Gram Matrix Captures Sequential Feature Linearization in Deep Networks
von: Cha, Taehun, et al.
Veröffentlicht: (2026)
von: Cha, Taehun, et al.
Veröffentlicht: (2026)
Quadratic models for understanding catapult dynamics of neural networks
von: Zhu, Libin, et al.
Veröffentlicht: (2022)
von: Zhu, Libin, et al.
Veröffentlicht: (2022)
The Approximate Fisher Influence Function: Faster Estimation of Data Influence in Statistical Models
von: Lev, Omri, et al.
Veröffentlicht: (2024)
von: Lev, Omri, et al.
Veröffentlicht: (2024)
Emergence in non-neural models: grokking modular arithmetic via average gradient outer product
von: Mallinar, Neil, et al.
Veröffentlicht: (2024)
von: Mallinar, Neil, et al.
Veröffentlicht: (2024)
Position: AI Evaluations Should be Grounded on a Theory of Capability
von: Jo, Nathanael, et al.
Veröffentlicht: (2025)
von: Jo, Nathanael, et al.
Veröffentlicht: (2025)
Mean-field underdamped Langevin dynamics and its spacetime discretization
von: Fu, Qiang, et al.
Veröffentlicht: (2023)
von: Fu, Qiang, et al.
Veröffentlicht: (2023)
Semivalue-based data valuation is arbitrary and gameable
von: Diehl, Hannah, et al.
Veröffentlicht: (2025)
von: Diehl, Hannah, et al.
Veröffentlicht: (2025)
From Cross-Validation to SURE: Asymptotic Risk of Tuned Regularized Estimators
von: Adusumilli, Karun, et al.
Veröffentlicht: (2026)
von: Adusumilli, Karun, et al.
Veröffentlicht: (2026)
High-accuracy sampling from constrained spaces with the Metropolis-adjusted Preconditioned Langevin Algorithm
von: Srinivasan, Vishwak, et al.
Veröffentlicht: (2024)
von: Srinivasan, Vishwak, et al.
Veröffentlicht: (2024)
Adaptive Kernel Selection for Stein Variational Gradient Descent
von: Melcher, Moritz, et al.
Veröffentlicht: (2025)
von: Melcher, Moritz, et al.
Veröffentlicht: (2025)
Layered Unlearning for Adversarial Relearning
von: Qian, Timothy, et al.
Veröffentlicht: (2025)
von: Qian, Timothy, et al.
Veröffentlicht: (2025)
An analysis of the combination of feature selection and machine learning methods for an accurate and timely detection of lung cancer
von: Shahriyar, Omid, et al.
Veröffentlicht: (2025)
von: Shahriyar, Omid, et al.
Veröffentlicht: (2025)
Fast sampling from constrained spaces using the Metropolis-adjusted Mirror Langevin algorithm
von: Srinivasan, Vishwak, et al.
Veröffentlicht: (2023)
von: Srinivasan, Vishwak, et al.
Veröffentlicht: (2023)
Adaptive Backtracking Line Search
von: Cavalcanti, Joao V., et al.
Veröffentlicht: (2024)
von: Cavalcanti, Joao V., et al.
Veröffentlicht: (2024)
UCD: Unlearning in LLMs via Contrastive Decoding
von: Suriyakumar, Vinith M., et al.
Veröffentlicht: (2025)
von: Suriyakumar, Vinith M., et al.
Veröffentlicht: (2025)
Clinician input steers frontier AI models toward both accurate and harmful decisions
von: Lopez, Ivan, et al.
Veröffentlicht: (2026)
von: Lopez, Ivan, et al.
Veröffentlicht: (2026)
Pandora's Regret: A Proper Scoring Rule for Evaluating Sequential Search
von: Flores, Gerardo A., et al.
Veröffentlicht: (2026)
von: Flores, Gerardo A., et al.
Veröffentlicht: (2026)
DiffusionWorldViewer: Exposing and Broadening the Worldview Reflected by Generative Text-to-Image Models
von: De Simone, Zoe, et al.
Veröffentlicht: (2023)
von: De Simone, Zoe, et al.
Veröffentlicht: (2023)
The Fast Mixing Mechanism for Differential Privacy
von: Lev, Omri, et al.
Veröffentlicht: (2026)
von: Lev, Omri, et al.
Veröffentlicht: (2026)
The Gaussian Mixing Mechanism: Renyi Differential Privacy via Gaussian Sketches
von: Lev, Omri, et al.
Veröffentlicht: (2025)
von: Lev, Omri, et al.
Veröffentlicht: (2025)
Supervised learning pays attention
von: Craig, Erin, et al.
Veröffentlicht: (2025)
von: Craig, Erin, et al.
Veröffentlicht: (2025)
An extension of linear self-attention for in-context learning
von: Hagiwara, Katsuyuki
Veröffentlicht: (2025)
von: Hagiwara, Katsuyuki
Veröffentlicht: (2025)
In-Context Clustering with Large Language Models
von: Wang, Ying, et al.
Veröffentlicht: (2025)
von: Wang, Ying, et al.
Veröffentlicht: (2025)
Aligning Evaluation with Clinical Priorities: Calibration, Label Shift, and Error Costs
von: Flores, Gerardo A., et al.
Veröffentlicht: (2025)
von: Flores, Gerardo A., et al.
Veröffentlicht: (2025)
Unstable Unlearning: The Hidden Risk of Concept Resurgence in Diffusion Models
von: Suriyakumar, Vinith M., et al.
Veröffentlicht: (2024)
von: Suriyakumar, Vinith M., et al.
Veröffentlicht: (2024)
Redundant feature screening method for human activity recognition based on attention purification mechanism
von: Li, Xiaoyang, et al.
Veröffentlicht: (2025)
von: Li, Xiaoyang, et al.
Veröffentlicht: (2025)
Easydiagnos: a framework for accurate feature selection for automatic diagnosis in smart healthcare
von: Maji, Prasenjit, et al.
Veröffentlicht: (2024)
von: Maji, Prasenjit, et al.
Veröffentlicht: (2024)
DiJiang: Efficient Large Language Models through Compact Kernelization
von: Chen, Hanting, et al.
Veröffentlicht: (2024)
von: Chen, Hanting, et al.
Veröffentlicht: (2024)
Meta-reinforcement learning with minimum attention
von: Gupta, Shashank, et al.
Veröffentlicht: (2025)
von: Gupta, Shashank, et al.
Veröffentlicht: (2025)
Material synthesis through simulations guided by machine learning: a position paper
von: Syed, Usman, et al.
Veröffentlicht: (2024)
von: Syed, Usman, et al.
Veröffentlicht: (2024)
Contextual Linear Activation Steering of Language Models
von: Hsu, Brandon, et al.
Veröffentlicht: (2026)
von: Hsu, Brandon, et al.
Veröffentlicht: (2026)
Language steering in latent space to mitigate unintended code-switching
von: Goncharov, Andrey, et al.
Veröffentlicht: (2025)
von: Goncharov, Andrey, et al.
Veröffentlicht: (2025)
ESM-NBR: fast and accurate nucleic acid-binding residue prediction via protein language model feature representation and multi-task learning
von: Zeng, Wenwu, et al.
Veröffentlicht: (2023)
von: Zeng, Wenwu, et al.
Veröffentlicht: (2023)
Beyond Barren Plateaus: A Scalable Quantum Convolutional Architecture for High-Fidelity Image Classification
von: Delhibabu, Radhakrishnan
Veröffentlicht: (2026)
von: Delhibabu, Radhakrishnan
Veröffentlicht: (2026)
An efficient, accurate, and interpretable machine learning method for computing probability of failure
von: Zhu, Jacob, et al.
Veröffentlicht: (2026)
von: Zhu, Jacob, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Catapults in SGD: spikes in the training loss and their impact on generalization through feature learning
von: Zhu, Libin, et al.
Veröffentlicht: (2023) -
xRFM: Accurate, scalable, and interpretable feature learning models for tabular data
von: Beaglehole, Daniel, et al.
Veröffentlicht: (2025) -
Toward universal steering and monitoring of AI models
von: Beaglehole, Daniel, et al.
Veröffentlicht: (2025) -
Linear Recursive Feature Machines provably recover low-rank matrices
von: Radhakrishnan, Adityanarayanan, et al.
Veröffentlicht: (2024) -
Context-Scaling versus Task-Scaling in In-Context Learning
von: Abedsoltan, Amirhesam, et al.
Veröffentlicht: (2024)