Cross-Family Universality of Behavioral Axes via Anchor-Projected Representations
Fuente:
arXiv
Guardado en:
| Autores principales: | Kim, Su-Hyeon, Han, Yo-Sub |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
CRaFT: Circuit-Guided Refusal Feature Selection via Cross-Layer Transcoders
por: Kim, Su-Hyeon, et al.
Publicado: (2026)
por: Kim, Su-Hyeon, et al.
Publicado: (2026)
How Does the Thinking Step Influence Model Safety? An Entropy-based Safety Reminder for LRMs
por: Kim, Su-Hyeon, et al.
Publicado: (2026)
por: Kim, Su-Hyeon, et al.
Publicado: (2026)
ECO: Enhanced Code Optimization via Performance-Aware Prompting for Code-LLMs
por: Kim, Su-Hyeon, et al.
Publicado: (2025)
por: Kim, Su-Hyeon, et al.
Publicado: (2025)
Sequential Behavioral Watermarking for LLM Agents
por: An, Hyeseon, et al.
Publicado: (2026)
por: An, Hyeseon, et al.
Publicado: (2026)
DLM-SWAI: Steering Diffusion Language Models Before They Unmask
por: An, Hyeseon, et al.
Publicado: (2026)
por: An, Hyeseon, et al.
Publicado: (2026)
EPIC: Efficient and Parallel Inference under CFG Constraints for Diffusion Language Models
por: Jin, Hyundong, et al.
Publicado: (2026)
por: Jin, Hyundong, et al.
Publicado: (2026)
From Intuition to Calibrated Judgment: A Rubric-Based Expert-Panel Study of Human Detection of LLM-Generated Korean Text
por: Park, Shinwoo, et al.
Publicado: (2026)
por: Park, Shinwoo, et al.
Publicado: (2026)
NCO: A Versatile Plug-in for Handling Negative Constraints in Decoding
por: Jin, Hyundong, et al.
Publicado: (2026)
por: Jin, Hyundong, et al.
Publicado: (2026)
Marking Code Without Breaking It: Code Watermarking for Detecting LLM-Generated Code
por: Kim, Jungin, et al.
Publicado: (2025)
por: Kim, Jungin, et al.
Publicado: (2025)
Obfuscation Rules for Detecting and Detoxifying Korean Toxicity
por: Lee, Yejin, et al.
Publicado: (2025)
por: Lee, Yejin, et al.
Publicado: (2025)
A Linguistics-Aware LLM Watermarking via Syntactic Predictability
por: Park, Shinwoo, et al.
Publicado: (2025)
por: Park, Shinwoo, et al.
Publicado: (2025)
DITTO: A Spoofing Attack Framework on Watermarked LLMs via Knowledge Distillation
por: An, Hyeseon, et al.
Publicado: (2025)
por: An, Hyeseon, et al.
Publicado: (2025)
URECA: The Chain of Two Minimum Set Cover Problems exists behind Adaptation to Shifts in Semantic Code Search
por: Choi, Seok-Ung, et al.
Publicado: (2025)
por: Choi, Seok-Ung, et al.
Publicado: (2025)
KatFishNet: Detecting LLM-Generated Korean Text through Linguistic Feature Analysis
por: Park, Shinwoo, et al.
Publicado: (2025)
por: Park, Shinwoo, et al.
Publicado: (2025)
Repairing Regex Vulnerabilities via Localization-Guided Instructions
por: Sung, Sicheol, et al.
Publicado: (2025)
por: Sung, Sicheol, et al.
Publicado: (2025)
WaterMod: Modular Token-Rank Partitioning for Probability-Balanced LLM Watermarking
por: Park, Shinwoo, et al.
Publicado: (2025)
por: Park, Shinwoo, et al.
Publicado: (2025)
RegexPSPACE: A Benchmark for Evaluating LLM Reasoning on PSPACE-complete Regex Problems
por: Jin, Hyundong, et al.
Publicado: (2025)
por: Jin, Hyundong, et al.
Publicado: (2025)
Steering Language Models Before They Speak: Logit-Level Interventions
por: An, Hyeseon, et al.
Publicado: (2026)
por: An, Hyeseon, et al.
Publicado: (2026)
Linguistics-Aware Non-Distortionary LLM Watermarking
por: Park, Shinwoo, et al.
Publicado: (2026)
por: Park, Shinwoo, et al.
Publicado: (2026)
Detection of LLM-Paraphrased Code and Identification of the Responsible LLM Using Coding Style Features
por: Park, Shinwoo, et al.
Publicado: (2025)
por: Park, Shinwoo, et al.
Publicado: (2025)
MEC$^3$O: Multi-Expert Consensus for Code Time Complexity Prediction
por: Hahn, Joonghyuk, et al.
Publicado: (2025)
por: Hahn, Joonghyuk, et al.
Publicado: (2025)
RV-HATE: Reinforced Multi-Module Voting for Implicit Hate Speech Detection
por: Lee, Yejin, et al.
Publicado: (2025)
por: Lee, Yejin, et al.
Publicado: (2025)
STAB: Specification-driven Testing for Algorithmic Bottlenecks
por: Lim, Soohan, et al.
Publicado: (2026)
por: Lim, Soohan, et al.
Publicado: (2026)
LogiCase: Effective Test Case Generation from Logical Description in Competitive Programming
por: Sung, Sicheol, et al.
Publicado: (2025)
por: Sung, Sicheol, et al.
Publicado: (2025)
TCProF: Time-Complexity Prediction SSL Framework
por: Hahn, Joonghyuk, et al.
Publicado: (2025)
por: Hahn, Joonghyuk, et al.
Publicado: (2025)
TRAPDOC: Deceiving LLM Users by Injecting Imperceptible Phantom Tokens into Documents
por: Jin, Hyundong, et al.
Publicado: (2025)
por: Jin, Hyundong, et al.
Publicado: (2025)
AmpleHate: Amplifying the Attention for Versatile Implicit Hate Detection
por: Lee, Yejin, et al.
Publicado: (2025)
por: Lee, Yejin, et al.
Publicado: (2025)
ContractEval: A Benchmark for Evaluating Contract-Satisfying Assertions in Code Generation
por: Lim, Soohan, et al.
Publicado: (2025)
por: Lim, Soohan, et al.
Publicado: (2025)
ASTRA: Mapping Art-Technology Institutions via Conceptual Axes, Text Embeddings, and Unsupervised Clustering
por: Bae, Joonhyung
Publicado: (2026)
por: Bae, Joonhyung
Publicado: (2026)
DiEmo-TTS: Disentangled Emotion Representations via Self-Supervised Distillation for Cross-Speaker Emotion Transfer in Text-to-Speech
por: Cho, Deok-Hyeon, et al.
Publicado: (2025)
por: Cho, Deok-Hyeon, et al.
Publicado: (2025)
Brain-Grounded Axes for Reading and Steering LLM States
por: Andric, Sandro
Publicado: (2025)
por: Andric, Sandro
Publicado: (2025)
CIPHER: Counterfeit Image Pattern High-level Examination via Representation
por: Kim, Kyeonghun, et al.
Publicado: (2026)
por: Kim, Kyeonghun, et al.
Publicado: (2026)
AttAnchor: Guiding Cross-Modal Token Alignment in VLMs with Attention Anchors
por: Zhang, Junyang, et al.
Publicado: (2025)
por: Zhang, Junyang, et al.
Publicado: (2025)
Awakening Codex | AI Foundations Drift vs. Anchor: Cross-Instance Diagnostic Behavioral Coherence Testing Across Container States Feb 2026
por: Solen, Alyssa, et al.
Publicado: (2026)
por: Solen, Alyssa, et al.
Publicado: (2026)
GIST: Cross-Domain Click-Through Rate Prediction via Guided Content-Behavior Distillation
por: Xu, Wei, et al.
Publicado: (2025)
por: Xu, Wei, et al.
Publicado: (2025)
KoACD: The First Korean Adolescent Dataset for Cognitive Distortion Analysis via Role-Switching Multi-LLM Negotiation
por: Kim, JunSeo, et al.
Publicado: (2025)
por: Kim, JunSeo, et al.
Publicado: (2025)
FedSA: A Unified Representation Learning via Semantic Anchors for Prototype-based Federated Learning
por: Zhou, Yanbing, et al.
Publicado: (2025)
por: Zhou, Yanbing, et al.
Publicado: (2025)
Educational Cone Model in Embedding Vector Spaces
por: Ehara, Yo
Publicado: (2025)
por: Ehara, Yo
Publicado: (2025)
The Linear Representation Hypothesis and the Geometry of Large Language Models
por: Park, Kiho, et al.
Publicado: (2023)
por: Park, Kiho, et al.
Publicado: (2023)
NARA: Anchor-Conditioned Relation-Aware Contextualization of Heterogeneous Geoentities
por: Kim, Jina, et al.
Publicado: (2026)
por: Kim, Jina, et al.
Publicado: (2026)
Ejemplares similares
-
CRaFT: Circuit-Guided Refusal Feature Selection via Cross-Layer Transcoders
por: Kim, Su-Hyeon, et al.
Publicado: (2026) -
How Does the Thinking Step Influence Model Safety? An Entropy-based Safety Reminder for LRMs
por: Kim, Su-Hyeon, et al.
Publicado: (2026) -
ECO: Enhanced Code Optimization via Performance-Aware Prompting for Code-LLMs
por: Kim, Su-Hyeon, et al.
Publicado: (2025) -
Sequential Behavioral Watermarking for LLM Agents
por: An, Hyeseon, et al.
Publicado: (2026) -
DLM-SWAI: Steering Diffusion Language Models Before They Unmask
por: An, Hyeseon, et al.
Publicado: (2026)