In-Context Fixation: When Demonstrated Labels Override Semantics in Few-Shot Classification
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Liu, Ming |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
When Persuasion Overrides Truth in Multi-Agent LLM Debates: Introducing a Confidence-Weighted Persuasion Override Rate (CW-POR)
von: Agarwal, Mahak, et al.
Veröffentlicht: (2025)
von: Agarwal, Mahak, et al.
Veröffentlicht: (2025)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
von: Fadli, Samih
Veröffentlicht: (2025)
von: Fadli, Samih
Veröffentlicht: (2025)
The Deterministic Horizon: When Extended Reasoning Fails and Tool Delegation Becomes Necessary
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
Self-Training Doesn't Flatten Language -- It Restructures It: Surface Markers Amplify While Deep Syntax Dies
von: Liu, Ming
Veröffentlicht: (2026)
von: Liu, Ming
Veröffentlicht: (2026)
The Readout Shortcut: Positional Number Copying Dominates Arithmetic CoT Readout in Small Language Models
von: Liu, Ming
Veröffentlicht: (2026)
von: Liu, Ming
Veröffentlicht: (2026)
Language as a Wave Phenomenon: Semantic Phase Locking and Interference in Neural Networks
von: Yıldırım, Alper, et al.
Veröffentlicht: (2025)
von: Yıldırım, Alper, et al.
Veröffentlicht: (2025)
Enhancing Burmese News Classification with Kolmogorov-Arnold Network Head Fine-tuning
von: Aung, Thura, et al.
Veröffentlicht: (2025)
von: Aung, Thura, et al.
Veröffentlicht: (2025)
Model Collapse as Cultural Evolution
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
von: Guo, Dongxin, et al.
Veröffentlicht: (2026)
Decodable but Not Corrected by Fixed Residual-Stream Linear Steering: Evidence from Medical LLM Failure Regimes
von: Liu, Ming
Veröffentlicht: (2026)
von: Liu, Ming
Veröffentlicht: (2026)
Control Reinforcement Learning: Interpretable Token-Level Steering of LLMs via Sparse Autoencoder Features
von: Cho, Seonglae, et al.
Veröffentlicht: (2026)
von: Cho, Seonglae, et al.
Veröffentlicht: (2026)
No Free Swap: Protocol-Dependent Layer Redundancy in Transformers
von: Garcia, Gabriel
Veröffentlicht: (2026)
von: Garcia, Gabriel
Veröffentlicht: (2026)
The Last Word Often Wins: A Format Confound in Chain-of-Thought Corruption Studies
von: Garcia, Gabriel
Veröffentlicht: (2026)
von: Garcia, Gabriel
Veröffentlicht: (2026)
Counterfactual Likelihood Tests for Indirect Influence in Private Reasoning Channels
von: Lorup, Alexander Boesgaard
Veröffentlicht: (2026)
von: Lorup, Alexander Boesgaard
Veröffentlicht: (2026)
TIAR: Trajectory-Informed Advantage Reweighting for LLM Abstention Learning
von: Pan, Muyu, et al.
Veröffentlicht: (2026)
von: Pan, Muyu, et al.
Veröffentlicht: (2026)
Pressure-Testing Deception Probes in LLMs: Scaling, Robustness, and the Geometry of Deceptive Representations
von: Kumar, Sachin
Veröffentlicht: (2026)
von: Kumar, Sachin
Veröffentlicht: (2026)
Structured Prompt Optimization Meets Reinforcement Learning for Global and Local Interpretability over Complex Text
von: Zhou, Tianyang, et al.
Veröffentlicht: (2026)
von: Zhou, Tianyang, et al.
Veröffentlicht: (2026)
AMEL: Accumulated Message Effects on LLM Judgments
von: Temkit, Sid-Ali
Veröffentlicht: (2026)
von: Temkit, Sid-Ali
Veröffentlicht: (2026)
Prototype Transformer: Towards Language Model Architectures Interpretable by Design
von: Yordanov, Yordan, et al.
Veröffentlicht: (2026)
von: Yordanov, Yordan, et al.
Veröffentlicht: (2026)
Forget Attention: Importance-Aware Attention Is All You Need
von: Shin, Soohyeong, et al.
Veröffentlicht: (2026)
von: Shin, Soohyeong, et al.
Veröffentlicht: (2026)
Turning the TIDE: Cross-Architecture Distillation for Diffusion Large Language Models
von: Zhang, Gongbo, et al.
Veröffentlicht: (2026)
von: Zhang, Gongbo, et al.
Veröffentlicht: (2026)
Generalizing Numerical Reasoning in Table Data through Operation Sketches and Self-Supervised Learning
von: Cho, Hanjun, et al.
Veröffentlicht: (2026)
von: Cho, Hanjun, et al.
Veröffentlicht: (2026)
Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy
von: Lan, Guangchen, et al.
Veröffentlicht: (2026)
von: Lan, Guangchen, et al.
Veröffentlicht: (2026)
Graph Memory Transformer (GMT)
von: Zanarini, Nicola, et al.
Veröffentlicht: (2026)
von: Zanarini, Nicola, et al.
Veröffentlicht: (2026)
Synthius-Mem: Brain-Inspired Hallucination-Resistant Persona Memory Achieving 94.4% Memory Accuracy and 99.6% Adversarial Robustness on LoCoMo
von: Gadzhiev, Artem, et al.
Veröffentlicht: (2026)
von: Gadzhiev, Artem, et al.
Veröffentlicht: (2026)
Weakly Supervised Distillation of Hallucination Signals into Transformer Representations
von: Salehmohamed, Shoaib Sadiq, et al.
Veröffentlicht: (2026)
von: Salehmohamed, Shoaib Sadiq, et al.
Veröffentlicht: (2026)
Cognitive Load Limits in Large Language Models: Benchmarking Multi-Hop Reasoning
von: Adapala, Sai Teja Reddy
Veröffentlicht: (2025)
von: Adapala, Sai Teja Reddy
Veröffentlicht: (2025)
Beyond Pass@k: Breadth-Depth Metrics for Reasoning Boundaries
von: Dragoi, Marius, et al.
Veröffentlicht: (2025)
von: Dragoi, Marius, et al.
Veröffentlicht: (2025)
Revisiting Intermediate-Layer Matching in Knowledge Distillation: Layer-Selection Strategy Doesn't Matter (Much)
von: Yu, Zony, et al.
Veröffentlicht: (2025)
von: Yu, Zony, et al.
Veröffentlicht: (2025)
Painless Activation Steering: An Automated, Lightweight Approach for Post-Training Large Language Models
von: Cui, Sasha, et al.
Veröffentlicht: (2025)
von: Cui, Sasha, et al.
Veröffentlicht: (2025)
MaPPO: Maximum a Posteriori Preference Optimization with Prior Knowledge
von: Lan, Guangchen, et al.
Veröffentlicht: (2025)
von: Lan, Guangchen, et al.
Veröffentlicht: (2025)
Dodo: Dynamic Contextual Compression for Decoder-only LMs
von: Qin, Guanghui, et al.
Veröffentlicht: (2023)
von: Qin, Guanghui, et al.
Veröffentlicht: (2023)
CorrSteer: Generation-Time LLM Steering via Correlated Sparse Autoencoder Features
von: Cho, Seonglae, et al.
Veröffentlicht: (2025)
von: Cho, Seonglae, et al.
Veröffentlicht: (2025)
Adapting While Learning: Grounding LLMs for Scientific Problems with Intelligent Tool Usage Adaptation
von: Lyu, Bohan, et al.
Veröffentlicht: (2024)
von: Lyu, Bohan, et al.
Veröffentlicht: (2024)
Harnessing Negative Signals: Reinforcement Distillation from Teacher Data for LLM Reasoning
von: Xu, Shuyao, et al.
Veröffentlicht: (2025)
von: Xu, Shuyao, et al.
Veröffentlicht: (2025)
EasyMath: A 0-shot Math Benchmark for SLMs
von: Karki, Drishya, et al.
Veröffentlicht: (2025)
von: Karki, Drishya, et al.
Veröffentlicht: (2025)
DPO Unchained: Your Training Algorithm is Secretly Disentangled in Human Choice Theory
von: Zhou, Wenxuan, et al.
Veröffentlicht: (2025)
von: Zhou, Wenxuan, et al.
Veröffentlicht: (2025)
Contextual Integrity in LLMs via Reasoning and Reinforcement Learning
von: Lan, Guangchen, et al.
Veröffentlicht: (2025)
von: Lan, Guangchen, et al.
Veröffentlicht: (2025)
Diagnosing and Addressing Pitfalls in KG-RAG Datasets: Toward More Reliable Benchmarking
von: Zhang, Liangliang, et al.
Veröffentlicht: (2025)
von: Zhang, Liangliang, et al.
Veröffentlicht: (2025)
Beyond Hallucinations: A Composite Score for Measuring Reliability in Open-Source Large Language Models
von: Salla, Rohit Kumar, et al.
Veröffentlicht: (2025)
von: Salla, Rohit Kumar, et al.
Veröffentlicht: (2025)
The Instability of Safety: How Random Seeds and Temperature Expose Inconsistent LLM Refusal Behavior
von: Larsen, Erik
Veröffentlicht: (2025)
von: Larsen, Erik
Veröffentlicht: (2025)
Ähnliche Einträge
-
When Persuasion Overrides Truth in Multi-Agent LLM Debates: Introducing a Confidence-Weighted Persuasion Override Rate (CW-POR)
von: Agarwal, Mahak, et al.
Veröffentlicht: (2025) -
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
von: Fadli, Samih
Veröffentlicht: (2025) -
The Deterministic Horizon: When Extended Reasoning Fails and Tool Delegation Becomes Necessary
von: Guo, Dongxin, et al.
Veröffentlicht: (2026) -
Self-Training Doesn't Flatten Language -- It Restructures It: Surface Markers Amplify While Deep Syntax Dies
von: Liu, Ming
Veröffentlicht: (2026) -
The Readout Shortcut: Positional Number Copying Dominates Arithmetic CoT Readout in Small Language Models
von: Liu, Ming
Veröffentlicht: (2026)