The Cylindrical Representation Hypothesis for Language Model Steering
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gao, Lang, Zhang, Jinghui, Liu, Wei, Ji, Fengxian, Wang, Chenxi, Song, Zirui, Ghosh, Akash, Mohamed, Youssef, Nakov, Preslav, Chen, Xiuying |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
When Personalization Tricks Detectors: The Feature-Inversion Trap in Machine-Generated Text Detection
von: Gao, Lang, et al.
Veröffentlicht: (2025)
von: Gao, Lang, et al.
Veröffentlicht: (2025)
Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language Models
von: Gao, Lang, et al.
Veröffentlicht: (2024)
von: Gao, Lang, et al.
Veröffentlicht: (2024)
Word Form Matters: LLMs' Semantic Reconstruction under Typoglycemia
von: Wang, Chenxi, et al.
Veröffentlicht: (2025)
von: Wang, Chenxi, et al.
Veröffentlicht: (2025)
Under the Shadow of Babel: How Language Shapes Reasoning in LLMs
von: Wang, Chenxi, et al.
Veröffentlicht: (2025)
von: Wang, Chenxi, et al.
Veröffentlicht: (2025)
Evaluate Bias without Manual Test Sets: A Concept Representation Perspective for LLMs
von: Gao, Lang, et al.
Veröffentlicht: (2025)
von: Gao, Lang, et al.
Veröffentlicht: (2025)
Rethinking STS and NLI in Large Language Models
von: Wang, Yuxia, et al.
Veröffentlicht: (2023)
von: Wang, Yuxia, et al.
Veröffentlicht: (2023)
Exploring Language Model Generalization in Low-Resource Extractive QA
von: Sengupta, Saptarshi, et al.
Veröffentlicht: (2024)
von: Sengupta, Saptarshi, et al.
Veröffentlicht: (2024)
DUAL-Bench: Measuring Over-Refusal and Robustness in Vision-Language Models
von: Ren, Kaixuan, et al.
Veröffentlicht: (2025)
von: Ren, Kaixuan, et al.
Veröffentlicht: (2025)
ServImage: An Image Generation and Editing Benchmark from Real-world Commercial Imaging Services
von: Ji, Fengxian, et al.
Veröffentlicht: (2026)
von: Ji, Fengxian, et al.
Veröffentlicht: (2026)
Large Language Models are Few-Shot Training Example Generators: A Case Study in Fallacy Recognition
von: Alhindi, Tariq, et al.
Veröffentlicht: (2023)
von: Alhindi, Tariq, et al.
Veröffentlicht: (2023)
Adapting Fake News Detection to the Era of Large Language Models
von: Su, Jinyan, et al.
Veröffentlicht: (2023)
von: Su, Jinyan, et al.
Veröffentlicht: (2023)
UNCERTAINTY-LINE: Length-Invariant Estimation of Uncertainty for Large Language Models
von: Vashurin, Roman, et al.
Veröffentlicht: (2025)
von: Vashurin, Roman, et al.
Veröffentlicht: (2025)
ConspirED: A Dataset for Cognitive Traits of Conspiracy Theories and Large Language Model Safety
von: Bates, Luke, et al.
Veröffentlicht: (2025)
von: Bates, Luke, et al.
Veröffentlicht: (2025)
Multimodal Large Language Models to Support Real-World Fact-Checking
von: Geng, Jiahui, et al.
Veröffentlicht: (2024)
von: Geng, Jiahui, et al.
Veröffentlicht: (2024)
How Does Prefix Matter in Reasoning Model Tuning?
von: Tomar, Raj Vardhan, et al.
Veröffentlicht: (2026)
von: Tomar, Raj Vardhan, et al.
Veröffentlicht: (2026)
Beyond Survival: Evaluating LLMs in Social Deduction Games with Human-Aligned Strategies
von: Song, Zirui, et al.
Veröffentlicht: (2025)
von: Song, Zirui, et al.
Veröffentlicht: (2025)
UnsafeChain: Enhancing Reasoning Model Safety via Hard Cases
von: Tomar, Raj Vardhan, et al.
Veröffentlicht: (2025)
von: Tomar, Raj Vardhan, et al.
Veröffentlicht: (2025)
Adaptive Conformal Prediction for Improving Factuality of Generations by Large Language Models
von: Rubashevskii, Aleksandr, et al.
Veröffentlicht: (2026)
von: Rubashevskii, Aleksandr, et al.
Veröffentlicht: (2026)
DP-Fusion: Token-Level Differentially Private Inference for Large Language Models
von: Thareja, Rushil, et al.
Veröffentlicht: (2025)
von: Thareja, Rushil, et al.
Veröffentlicht: (2025)
CoDet-M4: Detecting Machine-Generated Code in Multi-Lingual, Multi-Generator and Multi-Domain Settings
von: Orel, Daniil, et al.
Veröffentlicht: (2025)
von: Orel, Daniil, et al.
Veröffentlicht: (2025)
From Chaos to Clarity: Claim Normalization to Empower Fact-Checking
von: Sundriyal, Megha, et al.
Veröffentlicht: (2023)
von: Sundriyal, Megha, et al.
Veröffentlicht: (2023)
Toxicity Red-Teaming: Benchmarking LLM Safety in Singapore's Low-Resource Languages
von: Hu, Yujia, et al.
Veröffentlicht: (2025)
von: Hu, Yujia, et al.
Veröffentlicht: (2025)
Injecting Domain-Specific Knowledge into Large Language Models: A Comprehensive Survey
von: Song, Zirui, et al.
Veröffentlicht: (2025)
von: Song, Zirui, et al.
Veröffentlicht: (2025)
Con Instruction: Universal Jailbreaking of Multimodal Large Language Models via Non-Textual Modalities
von: Geng, Jiahui, et al.
Veröffentlicht: (2025)
von: Geng, Jiahui, et al.
Veröffentlicht: (2025)
Curriculum Learning and Pseudo-Labeling Improve the Generalization of Multi-Label Arabic Dialect Identification Models
von: Mekky, Ali, et al.
Veröffentlicht: (2026)
von: Mekky, Ali, et al.
Veröffentlicht: (2026)
A Fano-Style Accuracy Upper Bound for LLM Single-Pass Reasoning in Multi-Hop QA
von: Wan, Kaiyang, et al.
Veröffentlicht: (2025)
von: Wan, Kaiyang, et al.
Veröffentlicht: (2025)
SAE-SSV: Supervised Steering in Sparse Representation Spaces for Reliable Control of Language Models
von: He, Zirui, et al.
Veröffentlicht: (2025)
von: He, Zirui, et al.
Veröffentlicht: (2025)
Linear Representation Transferability Hypothesis: Leveraging Small Models to Steer Large Models
von: Bello, Femi, et al.
Veröffentlicht: (2025)
von: Bello, Femi, et al.
Veröffentlicht: (2025)
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models
von: Xu, Zixiang, et al.
Veröffentlicht: (2025)
von: Xu, Zixiang, et al.
Veröffentlicht: (2025)
TOP-Training: Target-Oriented Pretraining for Medical Extractive Question Answering
von: Sengupta, Saptarshi, et al.
Veröffentlicht: (2023)
von: Sengupta, Saptarshi, et al.
Veröffentlicht: (2023)
A Survey of Confidence Estimation and Calibration in Large Language Models
von: Geng, Jiahui, et al.
Veröffentlicht: (2023)
von: Geng, Jiahui, et al.
Veröffentlicht: (2023)
MuDRiC: Multi-Dialect Reasoning for Arabic Commonsense Validation
von: Elozeiri, Kareem, et al.
Veröffentlicht: (2025)
von: Elozeiri, Kareem, et al.
Veröffentlicht: (2025)
Missci: Reconstructing Fallacies in Misrepresented Science
von: Glockner, Max, et al.
Veröffentlicht: (2024)
von: Glockner, Max, et al.
Veröffentlicht: (2024)
Grounding Fallacies Misrepresenting Scientific Publications in Evidence
von: Glockner, Max, et al.
Veröffentlicht: (2024)
von: Glockner, Max, et al.
Veröffentlicht: (2024)
Do LLMs "Feel"? Emotion Circuits Discovery and Control
von: Wang, Chenxi, et al.
Veröffentlicht: (2025)
von: Wang, Chenxi, et al.
Veröffentlicht: (2025)
Evaluating and Mitigating Bias in AI-Based Medical Text Generation
von: Chen, Xiuying, et al.
Veröffentlicht: (2025)
von: Chen, Xiuying, et al.
Veröffentlicht: (2025)
DyFlow: Dynamic Workflow Framework for Agentic Reasoning
von: Wang, Yanbo, et al.
Veröffentlicht: (2025)
von: Wang, Yanbo, et al.
Veröffentlicht: (2025)
Personalized Steering of Large Language Models: Versatile Steering Vectors Through Bi-directional Preference Optimization
von: Cao, Yuanpu, et al.
Veröffentlicht: (2024)
von: Cao, Yuanpu, et al.
Veröffentlicht: (2024)
Co-FactChecker: A Framework for Human-AI Collaborative Claim Verification Using Large Reasoning Models
von: Sahnan, Dhruv, et al.
Veröffentlicht: (2026)
von: Sahnan, Dhruv, et al.
Veröffentlicht: (2026)
Cross-Cultural Transfer of Commonsense Reasoning in LLMs: Evidence from the Arab World
von: Almheiri, Saeed, et al.
Veröffentlicht: (2025)
von: Almheiri, Saeed, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
When Personalization Tricks Detectors: The Feature-Inversion Trap in Machine-Generated Text Detection
von: Gao, Lang, et al.
Veröffentlicht: (2025) -
Shaping the Safety Boundaries: Understanding and Defending Against Jailbreaks in Large Language Models
von: Gao, Lang, et al.
Veröffentlicht: (2024) -
Word Form Matters: LLMs' Semantic Reconstruction under Typoglycemia
von: Wang, Chenxi, et al.
Veröffentlicht: (2025) -
Under the Shadow of Babel: How Language Shapes Reasoning in LLMs
von: Wang, Chenxi, et al.
Veröffentlicht: (2025) -
Evaluate Bias without Manual Test Sets: A Concept Representation Perspective for LLMs
von: Gao, Lang, et al.
Veröffentlicht: (2025)