The Hidden Space of Safety: Understanding Preference-Tuned LLMs in Multilingual context
Fuente:
arXiv
Guardado en:
| Autores principales: | Verma, Nikhil, Bharadwaj, Manasa |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
GPT-DETOX: An In-Context Learning-Based Paraphraser for Text Detoxification
por: Pesaranghader, Ali, et al.
Publicado: (2024)
por: Pesaranghader, Ali, et al.
Publicado: (2024)
S3D: A Simple and Cost-Effective Self-Speculative Decoding Scheme for Low-Memory GPUs
por: Zhong, Wei, et al.
Publicado: (2024)
por: Zhong, Wei, et al.
Publicado: (2024)
Understanding Hidden Computations in Chain-of-Thought Reasoning
por: Bharadwaj, Aryasomayajula Ram
Publicado: (2024)
por: Bharadwaj, Aryasomayajula Ram
Publicado: (2024)
SMaRT: Select, Mix, and ReinvenT -- A Strategy Fusion Framework for LLM-Driven Reasoning and Planning
por: Verma, Nikhil, et al.
Publicado: (2025)
por: Verma, Nikhil, et al.
Publicado: (2025)
Quality-Aware Translation Tagging in Multilingual RAG system
por: Moon, Hoyeon, et al.
Publicado: (2025)
por: Moon, Hoyeon, et al.
Publicado: (2025)
OmniReflect: Discovering Transferable Constitutions for LLM agents via Neuro-Symbolic Reflections
por: Bharadwaj, Manasa, et al.
Publicado: (2025)
por: Bharadwaj, Manasa, et al.
Publicado: (2025)
PersonalHomeBench: Evaluating Agents in Personalized Smart Homes
por: Bharadwaj, Manasa, et al.
Publicado: (2026)
por: Bharadwaj, Manasa, et al.
Publicado: (2026)
Concept Space Alignment in Multilingual LLMs
por: Peng, Qiwei, et al.
Publicado: (2024)
por: Peng, Qiwei, et al.
Publicado: (2024)
IndicGenBench: A Multilingual Benchmark to Evaluate Generation Capabilities of LLMs on Indic Languages
por: Singh, Harman, et al.
Publicado: (2024)
por: Singh, Harman, et al.
Publicado: (2024)
In-context Learning vs. Instruction Tuning: The Case of Small and Multilingual Language Models
por: Ponce, David, et al.
Publicado: (2025)
por: Ponce, David, et al.
Publicado: (2025)
Towards Understanding the Fragility of Multilingual LLMs against Fine-Tuning Attacks
por: Poppi, Samuele, et al.
Publicado: (2024)
por: Poppi, Samuele, et al.
Publicado: (2024)
Cross-Attention Speculative Decoding
por: Zhong, Wei, et al.
Publicado: (2025)
por: Zhong, Wei, et al.
Publicado: (2025)
Reinforcement Learning Improves Traversal of Hierarchical Knowledge in LLMs
por: Zhang, Renfei, et al.
Publicado: (2025)
por: Zhang, Renfei, et al.
Publicado: (2025)
Finding and Reactivating Post-Trained LLMs' Hidden Safety Mechanisms
por: Li, Mingjie, et al.
Publicado: (2026)
por: Li, Mingjie, et al.
Publicado: (2026)
Configurable Safety Tuning of Language Models with Synthetic Preference Data
por: Gallego, Victor
Publicado: (2024)
por: Gallego, Victor
Publicado: (2024)
The Language Barrier: Dissecting Safety Challenges of LLMs in Multilingual Contexts
por: Shen, Lingfeng, et al.
Publicado: (2024)
por: Shen, Lingfeng, et al.
Publicado: (2024)
MAPO: Advancing Multilingual Reasoning through Multilingual Alignment-as-Preference Optimization
por: She, Shuaijie, et al.
Publicado: (2024)
por: She, Shuaijie, et al.
Publicado: (2024)
Code-Switching Red-Teaming: LLM Evaluation for Safety and Multilingual Understanding
por: Yoo, Haneul, et al.
Publicado: (2024)
por: Yoo, Haneul, et al.
Publicado: (2024)
Flattery, Fluff, and Fog: Diagnosing and Mitigating Idiosyncratic Biases in Preference Models
por: Bharadwaj, Anirudh, et al.
Publicado: (2025)
por: Bharadwaj, Anirudh, et al.
Publicado: (2025)
Understanding Multilingualism in Mixture-of-Experts LLMs: Routing Mechanism, Expert Specialization, and Layerwise Steering
por: Chen, Yuxin, et al.
Publicado: (2026)
por: Chen, Yuxin, et al.
Publicado: (2026)
Investigating Language Preference of Multilingual RAG Systems
por: Park, Jeonghyun, et al.
Publicado: (2025)
por: Park, Jeonghyun, et al.
Publicado: (2025)
Mind the Pause: Disfluency-Aware Objective Tuning for Multilingual Speech Correction with LLMs
por: Kumar, Deepak, et al.
Publicado: (2026)
por: Kumar, Deepak, et al.
Publicado: (2026)
Debiasing Multilingual LLMs in Cross-lingual Latent Space
por: Peng, Qiwei, et al.
Publicado: (2025)
por: Peng, Qiwei, et al.
Publicado: (2025)
RLHF Can Speak Many Languages: Unlocking Multilingual Preference Optimization for LLMs
por: Dang, John, et al.
Publicado: (2024)
por: Dang, John, et al.
Publicado: (2024)
Towards Better Understanding of Cybercrime: The Role of Fine-Tuned LLMs in Translation
por: Valeros, Veronica, et al.
Publicado: (2024)
por: Valeros, Veronica, et al.
Publicado: (2024)
Cross-Task Defense: Instruction-Tuning LLMs for Content Safety
por: Fu, Yu, et al.
Publicado: (2024)
por: Fu, Yu, et al.
Publicado: (2024)
CAPO: Confidence Aware Preference Optimization Learning for Multilingual Preferences
por: Pokharel, Rhitabrat, et al.
Publicado: (2025)
por: Pokharel, Rhitabrat, et al.
Publicado: (2025)
CONGRAD:Conflicting Gradient Filtering for Multilingual Preference Alignment
por: Li, Jiangnan, et al.
Publicado: (2025)
por: Li, Jiangnan, et al.
Publicado: (2025)
CARE: Multilingual Human Preference Learning for Cultural Awareness
por: Guo, Geyang, et al.
Publicado: (2025)
por: Guo, Geyang, et al.
Publicado: (2025)
HIPPO: Enhancing the Table Understanding Capability of LLMs through Hybrid-Modal Preference Optimization
por: Wang, Haolan, et al.
Publicado: (2025)
por: Wang, Haolan, et al.
Publicado: (2025)
Exploring Representational Disparities Between Multilingual and Bilingual Translation Models
por: Verma, Neha, et al.
Publicado: (2023)
por: Verma, Neha, et al.
Publicado: (2023)
Investigating Multilingual Instruction-Tuning: Do Polyglot Models Demand for Multilingual Instructions?
por: Weber, Alexander Arno, et al.
Publicado: (2024)
por: Weber, Alexander Arno, et al.
Publicado: (2024)
Multilingual Amnesia: On the Transferability of Unlearning in Multilingual LLMs
por: Farashah, Alireza Dehghanpour, et al.
Publicado: (2026)
por: Farashah, Alireza Dehghanpour, et al.
Publicado: (2026)
Layer-wise Swapping for Generalizable Multilingual Safety
por: Shin, Hyunseo, et al.
Publicado: (2026)
por: Shin, Hyunseo, et al.
Publicado: (2026)
States Hidden in Hidden States: LLMs Emerge Discrete State Representations Implicitly
por: Chen, Junhao, et al.
Publicado: (2024)
por: Chen, Junhao, et al.
Publicado: (2024)
The Hidden Space of Transformer Language Adapters
por: Alabi, Jesujoba O., et al.
Publicado: (2024)
por: Alabi, Jesujoba O., et al.
Publicado: (2024)
Challenges in Adapting Multilingual LLMs to Low-Resource Languages using LoRA PEFT Tuning
por: Khade, Omkar, et al.
Publicado: (2024)
por: Khade, Omkar, et al.
Publicado: (2024)
Beyond QA Pairs: Assessing Parameter-Efficient Fine-Tuning for Fact Embedding in LLMs
por: Ratnakar, Shivam, et al.
Publicado: (2025)
por: Ratnakar, Shivam, et al.
Publicado: (2025)
Controlling Language Confusion in Multilingual LLMs
por: Lee, Nahyun, et al.
Publicado: (2025)
por: Lee, Nahyun, et al.
Publicado: (2025)
Fake It To Make It: Virtual Multiviews to Enhance Monocular Indoor Semantic Scene Completion
por: Selvakumar, Anith, et al.
Publicado: (2025)
por: Selvakumar, Anith, et al.
Publicado: (2025)
Ejemplares similares
-
GPT-DETOX: An In-Context Learning-Based Paraphraser for Text Detoxification
por: Pesaranghader, Ali, et al.
Publicado: (2024) -
S3D: A Simple and Cost-Effective Self-Speculative Decoding Scheme for Low-Memory GPUs
por: Zhong, Wei, et al.
Publicado: (2024) -
Understanding Hidden Computations in Chain-of-Thought Reasoning
por: Bharadwaj, Aryasomayajula Ram
Publicado: (2024) -
SMaRT: Select, Mix, and ReinvenT -- A Strategy Fusion Framework for LLM-Driven Reasoning and Planning
por: Verma, Nikhil, et al.
Publicado: (2025) -
Quality-Aware Translation Tagging in Multilingual RAG system
por: Moon, Hoyeon, et al.
Publicado: (2025)