Manifold-Guided Attention Steering
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Ian, Guruprasad, Kapilesh, Sengupta, Raunak, Satish, Ninad, D'Antoni, Loris, Yu, Rose |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
PECAN: A Deterministic Certified Defense Against Backdoor Attacks
di: Zhang, Yuhao, et al.
Pubblicazione: (2023)
di: Zhang, Yuhao, et al.
Pubblicazione: (2023)
Bootstrapping Fuzzers for Compilers of Low-Resource Language Dialects Using Language Models
di: Vaidya, Sairam, et al.
Pubblicazione: (2025)
di: Vaidya, Sairam, et al.
Pubblicazione: (2025)
Continuous Diffusion Models Can Obey Formal Syntax
di: Kim, Jinwoo, et al.
Pubblicazione: (2026)
di: Kim, Jinwoo, et al.
Pubblicazione: (2026)
Learning the Error Patterns of Language Models
di: Kim, Jinwoo, et al.
Pubblicazione: (2026)
di: Kim, Jinwoo, et al.
Pubblicazione: (2026)
Verified Training for Counterfactual Explanation Robustness under Data Shift
di: Meyer, Anna P., et al.
Pubblicazione: (2024)
di: Meyer, Anna P., et al.
Pubblicazione: (2024)
Constrained Adaptive Rejection Sampling
di: Parys, Paweł, et al.
Pubblicazione: (2025)
di: Parys, Paweł, et al.
Pubblicazione: (2025)
Grammar-Aligned Decoding
di: Park, Kanghee, et al.
Pubblicazione: (2024)
di: Park, Kanghee, et al.
Pubblicazione: (2024)
Constrained Sampling for Language Models Should Be Easy: An MCMC Perspective
di: Gonzalez, Emmanuel Anaya, et al.
Pubblicazione: (2025)
di: Gonzalez, Emmanuel Anaya, et al.
Pubblicazione: (2025)
Zero-shot Factual Consistency Evaluation Across Domains
di: Agarwal, Raunak
Pubblicazione: (2024)
di: Agarwal, Raunak
Pubblicazione: (2024)
Verifying Solutions to Semantics-Guided Synthesis Problems
di: Murphy, Charlie, et al.
Pubblicazione: (2024)
di: Murphy, Charlie, et al.
Pubblicazione: (2024)
Fixed-Budget Constrained Best Arm Identification in Grouped Bandits
di: Mukherjee, Raunak, et al.
Pubblicazione: (2026)
di: Mukherjee, Raunak, et al.
Pubblicazione: (2026)
Learning Safe Autonomous Driving Policies Using Predictive Safety Representations
di: Keswani, Mahesh, et al.
Pubblicazione: (2025)
di: Keswani, Mahesh, et al.
Pubblicazione: (2025)
Language-Based Agent Control
di: Zhou, Timothy, et al.
Pubblicazione: (2026)
di: Zhou, Timothy, et al.
Pubblicazione: (2026)
Unrealizability Logic
di: Kim, Jinwoo, et al.
Pubblicazione: (2022)
di: Kim, Jinwoo, et al.
Pubblicazione: (2022)
A One-Layer Decoder-Only Transformer is a Two-Layer RNN: With an Application to Certified Robustness
di: Zhang, Yuhao, et al.
Pubblicazione: (2024)
di: Zhang, Yuhao, et al.
Pubblicazione: (2024)
Flexible and Efficient Grammar-Constrained Decoding
di: Park, Kanghee, et al.
Pubblicazione: (2025)
di: Park, Kanghee, et al.
Pubblicazione: (2025)
Synthesizing Specifications
di: Park, Kanghee, et al.
Pubblicazione: (2023)
di: Park, Kanghee, et al.
Pubblicazione: (2023)
LOUD: Synthesizing Strongest and Weakest Specifications
di: Park, Kanghee, et al.
Pubblicazione: (2024)
di: Park, Kanghee, et al.
Pubblicazione: (2024)
Probabilistic Consensus through Ensemble Validation: A Framework for LLM Reliability
di: Naik, Ninad
Pubblicazione: (2024)
di: Naik, Ninad
Pubblicazione: (2024)
X-SYNTH: Beyond Retrieval -- Enterprise Context Synthesis from Observed Digital Human Attention
di: Raghavan, Guruprasad, et al.
Pubblicazione: (2026)
di: Raghavan, Guruprasad, et al.
Pubblicazione: (2026)
Collaborative Causal Sensemaking: Closing the Complementarity Gap in Human-AI Decision Support
di: Jain, Raunak
Pubblicazione: (2025)
di: Jain, Raunak
Pubblicazione: (2025)
Manifold Steering Reveals the Shared Geometry of Neural Network Representation and Behavior
di: Wurgaft, Daniel, et al.
Pubblicazione: (2026)
di: Wurgaft, Daniel, et al.
Pubblicazione: (2026)
Predicting NCAP Safety Ratings: An Analysis of Vehicle Characteristics and ADAS Features Using Machine Learning
di: Kunwar, Raunak, et al.
Pubblicazione: (2025)
di: Kunwar, Raunak, et al.
Pubblicazione: (2025)
HH-SAE: Discovering and Steering Hierarchical Knowledge of Complex Manifolds
di: Wu, Honghan, et al.
Pubblicazione: (2026)
di: Wu, Honghan, et al.
Pubblicazione: (2026)
Safe Langevin Soft Actor Critic
di: Keswani, Mahesh, et al.
Pubblicazione: (2026)
di: Keswani, Mahesh, et al.
Pubblicazione: (2026)
Spotlight Your Instructions: Instruction-following with Dynamic Attention Steering
di: Venkateswaran, Praveen, et al.
Pubblicazione: (2025)
di: Venkateswaran, Praveen, et al.
Pubblicazione: (2025)
Mitigating Overthinking in Large Reasoning Models via Manifold Steering
di: Huang, Yao, et al.
Pubblicazione: (2025)
di: Huang, Yao, et al.
Pubblicazione: (2025)
Online Convex Optimization with Unbounded Memory
di: Kumar, Raunak, et al.
Pubblicazione: (2022)
di: Kumar, Raunak, et al.
Pubblicazione: (2022)
Dynamic Graph Attention Networks for Travel Time Distribution Prediction in Urban Arterial Roads
di: Yousefzadeh, Nooshin, et al.
Pubblicazione: (2024)
di: Yousefzadeh, Nooshin, et al.
Pubblicazione: (2024)
Curvature-Guided LoRA: Steering in the pretrained NTK subspace
di: Zheng, Frédéric, et al.
Pubblicazione: (2026)
di: Zheng, Frédéric, et al.
Pubblicazione: (2026)
Guided Random Forest and its application to data approximation
di: Gupta, Prashant, et al.
Pubblicazione: (2019)
di: Gupta, Prashant, et al.
Pubblicazione: (2019)
Tell Your Model Where to Attend: Post-hoc Attention Steering for LLMs
di: Zhang, Qingru, et al.
Pubblicazione: (2023)
di: Zhang, Qingru, et al.
Pubblicazione: (2023)
Guided Manifold Alignment with Geometry-Regularized Twin Autoencoders
di: Rhodes, Jake S., et al.
Pubblicazione: (2025)
di: Rhodes, Jake S., et al.
Pubblicazione: (2025)
The Format Tax
di: Lee, Ivan Yee, et al.
Pubblicazione: (2026)
di: Lee, Ivan Yee, et al.
Pubblicazione: (2026)
Semantics of Sets of Programs
di: Kim, Jinwoo, et al.
Pubblicazione: (2024)
di: Kim, Jinwoo, et al.
Pubblicazione: (2024)
ChopChop: a Programmable Framework for Semantically Constraining the Output of Language Models
di: Nagy, Shaan, et al.
Pubblicazione: (2025)
di: Nagy, Shaan, et al.
Pubblicazione: (2025)
Automating Unrealizability Logic: Hoare-Style Proof Synthesis for Infinite Sets of Programs
di: Nagy, Shaan, et al.
Pubblicazione: (2024)
di: Nagy, Shaan, et al.
Pubblicazione: (2024)
Virtualidad crítica en el aula universitaria en la pandemia (y más allá)
di: Maurizia D’Antoni
Pubblicazione: (2020)
di: Maurizia D’Antoni
Pubblicazione: (2020)
VIGOTSKI, DESARROLLO INFANTIL Y APORTES PARA EL AULA COSTARRICENSE
di: Maurizia D'Antoni
Pubblicazione: (2009)
di: Maurizia D'Antoni
Pubblicazione: (2009)
Derechos humanos de la infancia en un trabajo comunal universitario
di: Maurizia D'Antoni
Pubblicazione: (2006)
di: Maurizia D'Antoni
Pubblicazione: (2006)
Documenti analoghi
-
PECAN: A Deterministic Certified Defense Against Backdoor Attacks
di: Zhang, Yuhao, et al.
Pubblicazione: (2023) -
Bootstrapping Fuzzers for Compilers of Low-Resource Language Dialects Using Language Models
di: Vaidya, Sairam, et al.
Pubblicazione: (2025) -
Continuous Diffusion Models Can Obey Formal Syntax
di: Kim, Jinwoo, et al.
Pubblicazione: (2026) -
Learning the Error Patterns of Language Models
di: Kim, Jinwoo, et al.
Pubblicazione: (2026) -
Verified Training for Counterfactual Explanation Robustness under Data Shift
di: Meyer, Anna P., et al.
Pubblicazione: (2024)