Stronger Enforcement of Instruction Hierarchy via Augmented Intermediate Representations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kariyappa, Sanjay, Suh, G. Edward |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SideQuest: Model-Driven KV Cache Management for Long-Horizon Agentic Reasoning
von: Kariyappa, Sanjay, et al.
Veröffentlicht: (2026)
von: Kariyappa, Sanjay, et al.
Veröffentlicht: (2026)
Progressive Inference: Explaining Decoder-Only Sequence Classification Models Using Intermediate Predictions
von: Kariyappa, Sanjay, et al.
Veröffentlicht: (2024)
von: Kariyappa, Sanjay, et al.
Veröffentlicht: (2024)
Mitigating Indirect Prompt Injection via Instruction-Following Intent Analysis
von: Kang, Mintong, et al.
Veröffentlicht: (2025)
von: Kang, Mintong, et al.
Veröffentlicht: (2025)
HIPO: Instruction Hierarchy via Constrained Reinforcement Learning
von: Chen, Keru, et al.
Veröffentlicht: (2026)
von: Chen, Keru, et al.
Veröffentlicht: (2026)
Fine-Tuning Diffusion Models via Intermediate Distribution Shaping
von: Anil, Gautham Govind, et al.
Veröffentlicht: (2025)
von: Anil, Gautham Govind, et al.
Veröffentlicht: (2025)
Hierarchy Representation of Data in Machine Learnings
von: Yegang, Han, et al.
Veröffentlicht: (2023)
von: Yegang, Han, et al.
Veröffentlicht: (2023)
Uncovering the Latent Potential of Deep Intermediate Representations
von: Batra, Arnesh, et al.
Veröffentlicht: (2026)
von: Batra, Arnesh, et al.
Veröffentlicht: (2026)
Privacy Policy Enforcement Guardrails for Data-Sensitive Retrieval-Augmented Generation
von: Zafar, Osama, et al.
Veröffentlicht: (2026)
von: Zafar, Osama, et al.
Veröffentlicht: (2026)
Hyperbolic Residual Quantization: Discrete Representations for Data with Latent Hierarchies
von: Piękos, Piotr, et al.
Veröffentlicht: (2025)
von: Piękos, Piotr, et al.
Veröffentlicht: (2025)
TS-RAG: Retrieval-Augmented Generation based Time Series Foundation Models are Stronger Zero-Shot Forecaster
von: Ning, Kanghui, et al.
Veröffentlicht: (2025)
von: Ning, Kanghui, et al.
Veröffentlicht: (2025)
Zero-Shot Instruction Following in RL via Structured LTL Representations
von: Giuri, Mattia, et al.
Veröffentlicht: (2025)
von: Giuri, Mattia, et al.
Veröffentlicht: (2025)
Zero-Shot Instruction Following in RL via Structured LTL Representations
von: Jackermeier, Mathias, et al.
Veröffentlicht: (2026)
von: Jackermeier, Mathias, et al.
Veröffentlicht: (2026)
Instructional Segment Embedding: Improving LLM Safety with Instruction Hierarchy
von: Wu, Tong, et al.
Veröffentlicht: (2024)
von: Wu, Tong, et al.
Veröffentlicht: (2024)
Data Augmentation for Instruction Following Policies via Trajectory Segmentation
von: Höpner, Niklas, et al.
Veröffentlicht: (2025)
von: Höpner, Niklas, et al.
Veröffentlicht: (2025)
Leveraging Intermediate Representations of Time Series Foundation Models for Anomaly Detection
von: Han, Chan Sik, et al.
Veröffentlicht: (2025)
von: Han, Chan Sik, et al.
Veröffentlicht: (2025)
Encoding Temporal Statistical-space Priors via Augmented Representation
von: Choi, Insu, et al.
Veröffentlicht: (2024)
von: Choi, Insu, et al.
Veröffentlicht: (2024)
LLMBoost: Make Large Language Models Stronger with Boosting
von: Chen, Zehao, et al.
Veröffentlicht: (2025)
von: Chen, Zehao, et al.
Veröffentlicht: (2025)
Next Embedding Prediction Makes World Models Stronger
von: Bredis, George, et al.
Veröffentlicht: (2026)
von: Bredis, George, et al.
Veröffentlicht: (2026)
Norm-Hierarchy Transitions in Representation Learning: When and Why Neural Networks Abandon Shortcuts
von: Khanh, Truong Xuan, et al.
Veröffentlicht: (2026)
von: Khanh, Truong Xuan, et al.
Veröffentlicht: (2026)
Test-Time Adaptation Induces Stronger Accuracy and Agreement-on-the-Line
von: Kim, Eungyeup, et al.
Veröffentlicht: (2023)
von: Kim, Eungyeup, et al.
Veröffentlicht: (2023)
On Stronger Computational Separations Between Multimodal and Unimodal Machine Learning
von: Karchmer, Ari
Veröffentlicht: (2024)
von: Karchmer, Ari
Veröffentlicht: (2024)
When Truthful Representations Flip Under Deceptive Instructions?
von: Long, Xianxuan, et al.
Veröffentlicht: (2025)
von: Long, Xianxuan, et al.
Veröffentlicht: (2025)
Learning Without Augmenting: Unsupervised Time Series Representation Learning via Frame Projections
von: Demirel, Berken Utku, et al.
Veröffentlicht: (2025)
von: Demirel, Berken Utku, et al.
Veröffentlicht: (2025)
Disentangling Policy from Offline Task Representation Learning via Adversarial Data Augmentation
von: Jia, Chengxing, et al.
Veröffentlicht: (2024)
von: Jia, Chengxing, et al.
Veröffentlicht: (2024)
A Stronger Mixture of Low-Rank Experts for Fine-Tuning Foundation Models
von: Sun, Mengyang, et al.
Veröffentlicht: (2025)
von: Sun, Mengyang, et al.
Veröffentlicht: (2025)
Overconfident Errors Need Stronger Correction: Asymmetric Confidence Penalties for Reinforcement Learning
von: Xu, Yuanda, et al.
Veröffentlicht: (2026)
von: Xu, Yuanda, et al.
Veröffentlicht: (2026)
Sparser, Better, Deeper, Stronger: Improving Sparse Training with Exact Orthogonal Initialization
von: Nowak, Aleksandra Irena, et al.
Veröffentlicht: (2024)
von: Nowak, Aleksandra Irena, et al.
Veröffentlicht: (2024)
Can Large Language Models Understand Intermediate Representations in Compilers?
von: Jiang, Hailong, et al.
Veröffentlicht: (2025)
von: Jiang, Hailong, et al.
Veröffentlicht: (2025)
Reverse Thinking Makes LLMs Stronger Reasoners
von: Chen, Justin Chih-Yao, et al.
Veröffentlicht: (2024)
von: Chen, Justin Chih-Yao, et al.
Veröffentlicht: (2024)
Efficient and Robust Knowledge Distillation from A Stronger Teacher Based on Correlation Matching
von: Niu, Wenqi, et al.
Veröffentlicht: (2024)
von: Niu, Wenqi, et al.
Veröffentlicht: (2024)
Nevermind: Instruction Override and Moderation in Large Language Models
von: Kim, Edward
Veröffentlicht: (2024)
von: Kim, Edward
Veröffentlicht: (2024)
Generative Representational Instruction Tuning
von: Muennighoff, Niklas, et al.
Veröffentlicht: (2024)
von: Muennighoff, Niklas, et al.
Veröffentlicht: (2024)
Generating Clinically Realistic EHR Data via a Hierarchy- and Semantics-Guided Transformer
von: Zhou, Guanglin, et al.
Veröffentlicht: (2025)
von: Zhou, Guanglin, et al.
Veröffentlicht: (2025)
Higher Embedding Dimension Creates a Stronger World Model for a Simple Sorting Task
von: Bhalla, Brady, et al.
Veröffentlicht: (2025)
von: Bhalla, Brady, et al.
Veröffentlicht: (2025)
Superior Molecular Representations from Intermediate Encoder Layers
von: Pinto, Luis
Veröffentlicht: (2025)
von: Pinto, Luis
Veröffentlicht: (2025)
Effectively Controlling Reasoning Models through Thinking Intervention
von: Wu, Tong, et al.
Veröffentlicht: (2025)
von: Wu, Tong, et al.
Veröffentlicht: (2025)
Debate Helps Weak Judges Reward Stronger Models
von: Elasky, Ethan, et al.
Veröffentlicht: (2026)
von: Elasky, Ethan, et al.
Veröffentlicht: (2026)
SubstratumGraphEnv: Reinforcement Learning Environment (RLE) for Modeling System Attack Paths
von: Adewunmi, Bahirah, et al.
Veröffentlicht: (2026)
von: Adewunmi, Bahirah, et al.
Veröffentlicht: (2026)
When Attention Collapses: How Degenerate Layers in LLMs Enable Smaller, Stronger Models
von: Sanyal, Sunny, et al.
Veröffentlicht: (2024)
von: Sanyal, Sunny, et al.
Veröffentlicht: (2024)
Multi-Objective Instruction-Aware Representation Learning in Procedural Content Generation RL
von: Kim, Sung-Hyun, et al.
Veröffentlicht: (2025)
von: Kim, Sung-Hyun, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SideQuest: Model-Driven KV Cache Management for Long-Horizon Agentic Reasoning
von: Kariyappa, Sanjay, et al.
Veröffentlicht: (2026) -
Progressive Inference: Explaining Decoder-Only Sequence Classification Models Using Intermediate Predictions
von: Kariyappa, Sanjay, et al.
Veröffentlicht: (2024) -
Mitigating Indirect Prompt Injection via Instruction-Following Intent Analysis
von: Kang, Mintong, et al.
Veröffentlicht: (2025) -
HIPO: Instruction Hierarchy via Constrained Reinforcement Learning
von: Chen, Keru, et al.
Veröffentlicht: (2026) -
Fine-Tuning Diffusion Models via Intermediate Distribution Shaping
von: Anil, Gautham Govind, et al.
Veröffentlicht: (2025)