Prompt Perturbation Consistency Learning for Robust Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Qiang, Yao, Nandi, Subhrangshu, Mehrabi, Ninareh, Steeg, Greg Ver, Kumar, Anoop, Rumshisky, Anna, Galstyan, Aram |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Making Sense Of Distributed Representations With Activation Spectroscopy
di: Reing, Kyle, et al.
Pubblicazione: (2025)
di: Reing, Kyle, et al.
Pubblicazione: (2025)
QuAILoRA: Quantization-Aware Initialization for LoRA
di: Lawton, Neal, et al.
Pubblicazione: (2024)
di: Lawton, Neal, et al.
Pubblicazione: (2024)
Learning Morphisms with Gauss-Newton Approximation for Growing Networks
di: Lawton, Neal, et al.
Pubblicazione: (2024)
di: Lawton, Neal, et al.
Pubblicazione: (2024)
Data Advisor: Dynamic Data Curation for Safety Alignment of Large Language Models
di: Wang, Fei, et al.
Pubblicazione: (2024)
di: Wang, Fei, et al.
Pubblicazione: (2024)
KG-LLM-Bench: A Scalable Benchmark for Evaluating LLM Reasoning on Textualized Knowledge Graphs
di: Markowitz, Elan, et al.
Pubblicazione: (2025)
di: Markowitz, Elan, et al.
Pubblicazione: (2025)
Tokenization Matters: Navigating Data-Scarce Tokenization for Gender Inclusive Language Technologies
di: Ovalle, Anaelia, et al.
Pubblicazione: (2023)
di: Ovalle, Anaelia, et al.
Pubblicazione: (2023)
Annealed Importance Sampling with q-Paths
di: Brekelmans, Rob, et al.
Pubblicazione: (2020)
di: Brekelmans, Rob, et al.
Pubblicazione: (2020)
K-Edit: Language Model Editing with Contextual Knowledge Awareness
di: Markowitz, Elan, et al.
Pubblicazione: (2025)
di: Markowitz, Elan, et al.
Pubblicazione: (2025)
Knowledge Enhanced Multi-Domain Recommendations in an AI Assistant Application
di: Markowitz, Elan, et al.
Pubblicazione: (2023)
di: Markowitz, Elan, et al.
Pubblicazione: (2023)
Exploring the Design Space of Diffusion Bridge Models
di: Zhang, Shaorong, et al.
Pubblicazione: (2024)
di: Zhang, Shaorong, et al.
Pubblicazione: (2024)
Is Your Diffusion Sampler Actually Correct? A Sampler-Centric Evaluation of Discrete Diffusion Language Models
di: Tang, Luhan, et al.
Pubblicazione: (2026)
di: Tang, Luhan, et al.
Pubblicazione: (2026)
Emergent Abilities in Reduced-Scale Generative Language Models
di: Muckatira, Sherin, et al.
Pubblicazione: (2024)
di: Muckatira, Sherin, et al.
Pubblicazione: (2024)
Robust Persona-Aware Toxicity Detection with Prompt Optimization and Learned Ensembling
di: Atil, Berk, et al.
Pubblicazione: (2026)
di: Atil, Berk, et al.
Pubblicazione: (2026)
Discrete Stochastic Localization for Non-autoregressive Generation
di: Wu, Yunshu, et al.
Pubblicazione: (2026)
di: Wu, Yunshu, et al.
Pubblicazione: (2026)
Flip-Flop Consistency: Unsupervised Training for Robustness to Prompt Perturbations in LLMs
di: Hejabi, Parsa, et al.
Pubblicazione: (2025)
di: Hejabi, Parsa, et al.
Pubblicazione: (2025)
Measurement-Aligned Sampling for Inverse Problem
di: Zhang, Shaorong, et al.
Pubblicazione: (2025)
di: Zhang, Shaorong, et al.
Pubblicazione: (2025)
DiCoRe: Enhancing Zero-shot Event Detection via Divergent-Convergent LLM Reasoning
di: Parekh, Tanmay, et al.
Pubblicazione: (2025)
di: Parekh, Tanmay, et al.
Pubblicazione: (2025)
Breaking the Benchmark: Revealing LLM Bias via Minimal Contextual Augmentation
di: Miandoab, Kaveh Eskandari, et al.
Pubblicazione: (2025)
di: Miandoab, Kaveh Eskandari, et al.
Pubblicazione: (2025)
Attribute Controlled Fine-tuning for Large Language Models: A Case Study on Detoxification
di: Meng, Tao, et al.
Pubblicazione: (2024)
di: Meng, Tao, et al.
Pubblicazione: (2024)
Your Diffusion Model is Secretly a Noise Classifier and Benefits from Contrastive Training
di: Wu, Yunshu, et al.
Pubblicazione: (2024)
di: Wu, Yunshu, et al.
Pubblicazione: (2024)
Asymmetric Bias in Text-to-Image Generation with Adversarial Attacks
di: Shahgir, Haz Sameen, et al.
Pubblicazione: (2023)
di: Shahgir, Haz Sameen, et al.
Pubblicazione: (2023)
Spectral Regularization for Diffusion Models
di: Chandran, Satish, et al.
Pubblicazione: (2026)
di: Chandran, Satish, et al.
Pubblicazione: (2026)
Towards Lightweight Reliability: Using Soft Prompts for Hallucination Mitigation in Large Language Models
di: Siddiqui, S M Tahmid, et al.
Pubblicazione: (2026)
di: Siddiqui, S M Tahmid, et al.
Pubblicazione: (2026)
Generation Order and Parallel Decoding in Masked Diffusion Models: An Information-Theoretic Perspective
di: Zhang, Shaorong, et al.
Pubblicazione: (2026)
di: Zhang, Shaorong, et al.
Pubblicazione: (2026)
Activation Steering for Synthetic Data Generation: The Role of Diversity in Downstream Safety Detection
di: Deshpande, Vijeta, et al.
Pubblicazione: (2026)
di: Deshpande, Vijeta, et al.
Pubblicazione: (2026)
Beyond Perplexity: A Geometric and Spectral Study of Low-Rank Pre-Training
di: Shivagunde, Namrata, et al.
Pubblicazione: (2026)
di: Shivagunde, Namrata, et al.
Pubblicazione: (2026)
FedDAPL: Toward Client-Private Generalization in Federated Learning
di: Loaliyan, Soroosh Safari, et al.
Pubblicazione: (2025)
di: Loaliyan, Soroosh Safari, et al.
Pubblicazione: (2025)
MMG: Mutual Information Estimation via the MMSE Gap in Diffusion
di: Yu, Longxuan, et al.
Pubblicazione: (2025)
di: Yu, Longxuan, et al.
Pubblicazione: (2025)
On the steerability of large language models toward data-driven personas
di: Li, Junyi, et al.
Pubblicazione: (2023)
di: Li, Junyi, et al.
Pubblicazione: (2023)
Toward More Accurate and Generalizable Evaluation Metrics for Task-Oriented Dialogs
di: Komma, Abishek, et al.
Pubblicazione: (2023)
di: Komma, Abishek, et al.
Pubblicazione: (2023)
Probing the Robustness of Large Language Models Safety to Latent Perturbations
di: Gu, Tianle, et al.
Pubblicazione: (2025)
di: Gu, Tianle, et al.
Pubblicazione: (2025)
Dissecting Language Models: Machine Unlearning via Selective Pruning
di: Pochinkov, Nicholas, et al.
Pubblicazione: (2024)
di: Pochinkov, Nicholas, et al.
Pubblicazione: (2024)
Improving Mutual Information Estimation with Annealed and Energy-Based Bounds
di: Brekelmans, Rob, et al.
Pubblicazione: (2023)
di: Brekelmans, Rob, et al.
Pubblicazione: (2023)
Interpretable Diffusion via Information Decomposition
di: Kong, Xianghao, et al.
Pubblicazione: (2023)
di: Kong, Xianghao, et al.
Pubblicazione: (2023)
MoreauPruner: Robust Pruning of Large Language Models against Weight Perturbations
di: Wang, Zixiao, et al.
Pubblicazione: (2024)
di: Wang, Zixiao, et al.
Pubblicazione: (2024)
PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts
di: Zhu, Kaijie, et al.
Pubblicazione: (2023)
di: Zhu, Kaijie, et al.
Pubblicazione: (2023)
Synthetic Data Generation and Joint Learning for Robust Code-Mixed Translation
di: Kartik, Kartik, et al.
Pubblicazione: (2024)
di: Kartik, Kartik, et al.
Pubblicazione: (2024)
Revise, Don't Freeze: Sampler-Matched Training for Self-Correcting Masked Diffusion Language Models
di: Yu, Longxuan, et al.
Pubblicazione: (2026)
di: Yu, Longxuan, et al.
Pubblicazione: (2026)
Wanda++: Pruning Large Language Models via Regional Gradients
di: Yang, Yifan, et al.
Pubblicazione: (2025)
di: Yang, Yifan, et al.
Pubblicazione: (2025)
Compress, Gather, and Recompute: REFORMing Long-Context Processing in Transformers
di: Song, Woomin, et al.
Pubblicazione: (2025)
di: Song, Woomin, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Making Sense Of Distributed Representations With Activation Spectroscopy
di: Reing, Kyle, et al.
Pubblicazione: (2025) -
QuAILoRA: Quantization-Aware Initialization for LoRA
di: Lawton, Neal, et al.
Pubblicazione: (2024) -
Learning Morphisms with Gauss-Newton Approximation for Growing Networks
di: Lawton, Neal, et al.
Pubblicazione: (2024) -
Data Advisor: Dynamic Data Curation for Safety Alignment of Large Language Models
di: Wang, Fei, et al.
Pubblicazione: (2024) -
KG-LLM-Bench: A Scalable Benchmark for Evaluating LLM Reasoning on Textualized Knowledge Graphs
di: Markowitz, Elan, et al.
Pubblicazione: (2025)