Manipulating Transformer-Based Models: Controllability, Steerability, and Robust Interventions
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Alpay, Faruk, Alpay, Taylan |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
NeuroState-Bench: A Human-Calibrated Benchmark for Commitment Integrity in LLM Agent Profiles
par: Jia, Xiao
Publié: (2026)
par: Jia, Xiao
Publié: (2026)
Latent Object Permanence: Topological Phase Transitions, Free-Energy Principles, and Renormalization Group Flows in Deep Transformer Manifolds
par: Alpay, Faruk, et autres
Publié: (2026)
par: Alpay, Faruk, et autres
Publié: (2026)
$ϕ^{\infty}$: Clause Purification, Embedding Realignment, and the Total Suppression of the Em Dash in Autoregressive Language Models
par: Kilictas, Bugra, et autres
Publié: (2025)
par: Kilictas, Bugra, et autres
Publié: (2025)
ART: Adaptive Response Tuning Framework -- A Multi-Agent Tournament-Based Approach to LLM Response Optimization
par: Khan, Omer Jauhar
Publié: (2025)
par: Khan, Omer Jauhar
Publié: (2025)
No Memorization, No Detection: Output Distribution-Based Contamination Detection in Small Language Models
par: Sela, Omer
Publié: (2026)
par: Sela, Omer
Publié: (2026)
Performance Evaluation of Sentiment Analysis on Text and Emoji Data Using End-to-End, Transfer Learning, Distributed and Explainable AI Models
par: Velampalli, Sirisha, et autres
Publié: (2025)
par: Velampalli, Sirisha, et autres
Publié: (2025)
lmfaoooo at SemEval-2026 Task 1: Humor Is an Audience. Preference Modeling for Constrained Humor Generation
par: Tikhonov, Alexey, et autres
Publié: (2026)
par: Tikhonov, Alexey, et autres
Publié: (2026)
From Noise to Diversity: Random Embedding Injection in LLM Reasoning
par: Kim, Heejun, et autres
Publié: (2026)
par: Kim, Heejun, et autres
Publié: (2026)
DeepPersona: A Generative Engine for Scaling Deep Synthetic Personas
par: Wang, Zhen, et autres
Publié: (2025)
par: Wang, Zhen, et autres
Publié: (2025)
Adaptive Multi-Stage Patent Claim Generation with Unified Quality Assessment
par: Liang, Chen-Wei, et autres
Publié: (2026)
par: Liang, Chen-Wei, et autres
Publié: (2026)
Constitution or Collapse? Exploring Constitutional AI with Llama 3-8B
par: Zhang, Xue
Publié: (2025)
par: Zhang, Xue
Publié: (2025)
Generative AI and the Transformation of Software Development Practices
par: Acharya, Vivek
Publié: (2025)
par: Acharya, Vivek
Publié: (2025)
Dynamic Dual-Granularity Skill Bank for Agentic RL
par: Tu, Songjun, et autres
Publié: (2026)
par: Tu, Songjun, et autres
Publié: (2026)
Harnessing non-adversarial robustness in large language models
par: Zhou, Qinghua, et autres
Publié: (2026)
par: Zhou, Qinghua, et autres
Publié: (2026)
Solving Zebra Puzzles Using Constraint-Guided Multi-Agent Systems
par: Berman, Shmuel, et autres
Publié: (2024)
par: Berman, Shmuel, et autres
Publié: (2024)
How much do LLMs learn from negative examples?
par: Hamdan, Shadi, et autres
Publié: (2025)
par: Hamdan, Shadi, et autres
Publié: (2025)
XAutoLM: Efficient Fine-Tuning of Language Models via Meta-Learning and AutoML
par: Estevanell-Valladares, Ernesto L., et autres
Publié: (2025)
par: Estevanell-Valladares, Ernesto L., et autres
Publié: (2025)
Alpay Algebra V: Multi-Layered Semantic Games and Transfinite Fixed-Point Simulation
par: Kilictas, Bugra, et autres
Publié: (2025)
par: Kilictas, Bugra, et autres
Publié: (2025)
Alpay Algebra IV: Symbiotic Semantics and the Fixed-Point Convergence of Observer Embeddings
par: Kilictas, Bugra, et autres
Publié: (2025)
par: Kilictas, Bugra, et autres
Publié: (2025)
The Geometry of Persona: Disentangling Personality from Reasoning in Large Language Models
par: Wang, Zhixiang
Publié: (2025)
par: Wang, Zhixiang
Publié: (2025)
On measuring grounding and generalizing grounding problems
par: Quigley, Daniel, et autres
Publié: (2025)
par: Quigley, Daniel, et autres
Publié: (2025)
AI Agents: Evolution, Architecture, and Real-World Applications
par: Krishnan, Naveen
Publié: (2025)
par: Krishnan, Naveen
Publié: (2025)
Adaptive Minds: Empowering Agents with LoRA-as-Tools
par: Shekar, Pavan C, et autres
Publié: (2025)
par: Shekar, Pavan C, et autres
Publié: (2025)
CLMN: Concept based Language Models via Neural Symbolic Reasoning
par: Yang, Yibo
Publié: (2025)
par: Yang, Yibo
Publié: (2025)
Tool-Genesis: A Task-Driven Tool Creation Benchmark for Self-Evolving Language Agent
par: Xia, Bowei, et autres
Publié: (2026)
par: Xia, Bowei, et autres
Publié: (2026)
Word Overuse and Alignment in Large Language Models: The Influence of Learning from Human Feedback
par: Juzek, Tom S., et autres
Publié: (2025)
par: Juzek, Tom S., et autres
Publié: (2025)
AIPsy-Affect: A Keyword-Free Clinical Stimulus Battery for Mechanistic Interpretability of Emotion in Language Models
par: Keeman, Michael
Publié: (2026)
par: Keeman, Michael
Publié: (2026)
Hybrid Gated Flow (HGF): Stabilizing 1.58-bit LLMs via Selective Low-Rank Correction
par: Pizzo, David Alejandro Trejo
Publié: (2026)
par: Pizzo, David Alejandro Trejo
Publié: (2026)
Measuring and curing reasoning rigidity: from decorative chain-of-thought to genuine faithfulness
par: Basu, Abhinaba, et autres
Publié: (2026)
par: Basu, Abhinaba, et autres
Publié: (2026)
Controlling Long-Horizon Behavior in Language Model Agents with Explicit State Dynamics
par: Subaharan, Sukesh
Publié: (2026)
par: Subaharan, Sukesh
Publié: (2026)
Towards Ontology-Enhanced Representation Learning for Large Language Models
par: Ronzano, Francesco, et autres
Publié: (2024)
par: Ronzano, Francesco, et autres
Publié: (2024)
Communicative Agents for Slideshow Storytelling Video Generation based on LLMs
par: Fan, Jingxing, et autres
Publié: (2025)
par: Fan, Jingxing, et autres
Publié: (2025)
RPRA: Predicting an LLM-Judge for Efficient but Performant Inference
par: Ashley, Dylan R., et autres
Publié: (2026)
par: Ashley, Dylan R., et autres
Publié: (2026)
RACAS: Controlling Diverse Robots With a Single Agentic System
par: Ashley, Dylan R., et autres
Publié: (2026)
par: Ashley, Dylan R., et autres
Publié: (2026)
SafeAnchor: Preventing Cumulative Safety Erosion in Continual Domain Adaptation of Large Language Models
par: Guo, Dongxin, et autres
Publié: (2026)
par: Guo, Dongxin, et autres
Publié: (2026)
ReFactor GNNs: Revisiting Factorisation-based Models from a Message-Passing Perspective
par: Chen, Yihong, et autres
Publié: (2022)
par: Chen, Yihong, et autres
Publié: (2022)
mHC-SSM: Manifold-Constrained Hyper-Connections for State Space Language Models with Stream-Specialized Adapters
par: Mutlu, Abdulvahap, et autres
Publié: (2026)
par: Mutlu, Abdulvahap, et autres
Publié: (2026)
Product-of-Experts Training Reduces Dataset Artifacts in Natural Language Inference
par: Mathew, Aby Mammen
Publié: (2026)
par: Mathew, Aby Mammen
Publié: (2026)
Mitigating Position-Shift Failures in Text-Based Modular Arithmetic via Position Curriculum and Template Diversity
par: Yudin, Nikolay
Publié: (2026)
par: Yudin, Nikolay
Publié: (2026)
Monotonicity as an Architectural Bias for Robust Language Models
par: Cooper, Patrick, et autres
Publié: (2026)
par: Cooper, Patrick, et autres
Publié: (2026)
Documents similaires
-
NeuroState-Bench: A Human-Calibrated Benchmark for Commitment Integrity in LLM Agent Profiles
par: Jia, Xiao
Publié: (2026) -
Latent Object Permanence: Topological Phase Transitions, Free-Energy Principles, and Renormalization Group Flows in Deep Transformer Manifolds
par: Alpay, Faruk, et autres
Publié: (2026) -
$ϕ^{\infty}$: Clause Purification, Embedding Realignment, and the Total Suppression of the Em Dash in Autoregressive Language Models
par: Kilictas, Bugra, et autres
Publié: (2025) -
ART: Adaptive Response Tuning Framework -- A Multi-Agent Tournament-Based Approach to LLM Response Optimization
par: Khan, Omer Jauhar
Publié: (2025) -
No Memorization, No Detection: Output Distribution-Based Contamination Detection in Small Language Models
par: Sela, Omer
Publié: (2026)