Prompts have evil twins
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Melamed, Rimon, McCabe, Lucas H., Wakhare, Tanay, Kim, Yejin, Huang, H. Howie, Boix-Adsera, Enric |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Demystifying optimized prompts in language models
von: Melamed, Rimon, et al.
Veröffentlicht: (2025)
von: Melamed, Rimon, et al.
Veröffentlicht: (2025)
Estimating Semantic Alphabet Size for LLM Uncertainty Quantification
von: McCabe, Lucas H., et al.
Veröffentlicht: (2025)
von: McCabe, Lucas H., et al.
Veröffentlicht: (2025)
Towards a theory of model distillation
von: Boix-Adsera, Enric
Veröffentlicht: (2024)
von: Boix-Adsera, Enric
Veröffentlicht: (2024)
SENECA: Small-Sample Discrete Entropy Estimation via Self-Consistent Missing Mass
von: McCabe, Lucas H., et al.
Veröffentlicht: (2026)
von: McCabe, Lucas H., et al.
Veröffentlicht: (2026)
Toward universal steering and monitoring of AI models
von: Beaglehole, Daniel, et al.
Veröffentlicht: (2025)
von: Beaglehole, Daniel, et al.
Veröffentlicht: (2025)
Improving Content Recommendation: Knowledge Graph-Based Semantic Contrastive Learning for Diversity and Cold-Start Users
von: Kim, Yejin, et al.
Veröffentlicht: (2024)
von: Kim, Yejin, et al.
Veröffentlicht: (2024)
Secret mixtures of experts inside your LLM
von: Boix-Adsera, Enric
Veröffentlicht: (2025)
von: Boix-Adsera, Enric
Veröffentlicht: (2025)
On the inductive bias of infinite-depth ResNets and the bottleneck rank
von: Boix-Adsera, Enric
Veröffentlicht: (2025)
von: Boix-Adsera, Enric
Veröffentlicht: (2025)
Causal Reasoning in Large Language Models: A Knowledge Graph Approach
von: Kim, Yejin, et al.
Veröffentlicht: (2024)
von: Kim, Yejin, et al.
Veröffentlicht: (2024)
Let Me Think! A Long Chain-of-Thought Can Be Worth Exponentially Many Short Ones
von: Mirtaheri, Parsa, et al.
Veröffentlicht: (2025)
von: Mirtaheri, Parsa, et al.
Veröffentlicht: (2025)
When can transformers reason with abstract symbols?
von: Boix-Adsera, Enric, et al.
Veröffentlicht: (2023)
von: Boix-Adsera, Enric, et al.
Veröffentlicht: (2023)
Evil twins are not that evil: Qualitative insights into machine-generated prompts
von: Rakotonirina, Nathanaël Carraz, et al.
Veröffentlicht: (2024)
von: Rakotonirina, Nathanaël Carraz, et al.
Veröffentlicht: (2024)
The power of fine-grained experts: Granularity boosts expressivity in Mixture of Experts
von: Boix-Adsera, Enric, et al.
Veröffentlicht: (2025)
von: Boix-Adsera, Enric, et al.
Veröffentlicht: (2025)
Iterated Entropy Derivatives and Binary Entropy Inequalities
von: Wakhare, Tanay
Veröffentlicht: (2023)
von: Wakhare, Tanay
Veröffentlicht: (2023)
Romik's Conjecture for the Jacobi Theta Function
von: Wakhare, Tanay
Veröffentlicht: (2019)
von: Wakhare, Tanay
Veröffentlicht: (2019)
Graph Eigenvalues and Projection Constants
von: Wakhare, Tanay
Veröffentlicht: (2026)
von: Wakhare, Tanay
Veröffentlicht: (2026)
Nonparametric Estimation and Comparison of Distance Distributions from Censored Data
von: McCabe, Lucas H.
Veröffentlicht: (2023)
von: McCabe, Lucas H.
Veröffentlicht: (2023)
Predicting Movie Hits Before They Happen with LLMs
von: Agah, Shaghayegh, et al.
Veröffentlicht: (2025)
von: Agah, Shaghayegh, et al.
Veröffentlicht: (2025)
The merged-staircase property: a necessary and nearly sufficient condition for SGD learning of sparse functions on two-layer neural networks
von: Abbe, Emmanuel, et al.
Veröffentlicht: (2022)
von: Abbe, Emmanuel, et al.
Veröffentlicht: (2022)
The Coming Generation of Computer Proficient Students: What It May Mean for Libraries.
von: McCabe, Gerard B., et al.
Veröffentlicht: (1994)
von: McCabe, Gerard B., et al.
Veröffentlicht: (1994)
Do Language Models Know When They'll Refuse? Probing Introspective Awareness of Safety Boundaries
von: Gondil, Tanay
Veröffentlicht: (2026)
von: Gondil, Tanay
Veröffentlicht: (2026)
Nerve Models of Subdivision Bifiltrations
von: Lesnick, Michael, et al.
Veröffentlicht: (2024)
von: Lesnick, Michael, et al.
Veröffentlicht: (2024)
Sparse Approximation of the Subdivision-Rips Bifiltration for Doubling Metrics
von: Lesnick, Michael, et al.
Veröffentlicht: (2024)
von: Lesnick, Michael, et al.
Veröffentlicht: (2024)
The Features at Convergence Theorem: a first-principles alternative to the Neural Feature Ansatz for how networks learn representations
von: Boix-Adsera, Enric, et al.
Veröffentlicht: (2025)
von: Boix-Adsera, Enric, et al.
Veröffentlicht: (2025)
School Board Censorship: Library Books and Curriculum Materials.
von: Cambron-McCabe, Nelda H.
Veröffentlicht: (1982)
von: Cambron-McCabe, Nelda H.
Veröffentlicht: (1982)
Personalized LLM Response Generation with Parameterized Memory Injection
von: Zhang, Kai, et al.
Veröffentlicht: (2024)
von: Zhang, Kai, et al.
Veröffentlicht: (2024)
Latent Preference Modeling for Cross-Session Personalized Tool Calling
von: Yoon, Yejin, et al.
Veröffentlicht: (2026)
von: Yoon, Yejin, et al.
Veröffentlicht: (2026)
On conjectural fermionic formulas for the Macdonald index in Argyres-Douglas theories
von: Chern, Shane, et al.
Veröffentlicht: (2026)
von: Chern, Shane, et al.
Veröffentlicht: (2026)
one-dimensional axial strain
von: McCabe, David
Veröffentlicht: (2025)
von: McCabe, David
Veröffentlicht: (2025)
A Square Matrix Classification Scheme
von: McCabe, David
Veröffentlicht: (2025)
von: McCabe, David
Veröffentlicht: (2025)
A Smoking-Gun Resolution of the Hubble Tension: Parameter-Free Local-Global Mapping
von: McCabe, Luke
Veröffentlicht: (2026)
von: McCabe, Luke
Veröffentlicht: (2026)
A Parameter-Free, First Principles Dual Derivation of the Cosmological Constant: The 25/12 Consilience in Chrono Singularity Unification
von: McCabe, Luke
Veröffentlicht: (2026)
von: McCabe, Luke
Veröffentlicht: (2026)
A Parameter-Free, First Principles Dual Derivation of the Cosmological Constant: The 25/12 Consilience in Chrono Singularity Unification
von: McCabe, Luke
Veröffentlicht: (2026)
von: McCabe, Luke
Veröffentlicht: (2026)
A Zero-Parameter Derivation of String Theory: Resolving the Landscape Problem.
von: McCabe, Luke
Veröffentlicht: (2026)
von: McCabe, Luke
Veröffentlicht: (2026)
correction: second difference matrix dimension reduction
von: McCabe, David
Veröffentlicht: (2025)
von: McCabe, David
Veröffentlicht: (2025)
The Zero-Parameter SymPy-Physics Complete Validation of String Theory From First Principles with CSU Framework
von: McCabe, Luke
Veröffentlicht: (2026)
von: McCabe, Luke
Veröffentlicht: (2026)
Race, Tea and Colonial Resettlement
von: McCabe, Jane
Veröffentlicht: (2019)
von: McCabe, Jane
Veröffentlicht: (2019)
Begging, Charity and Religion in Pre-Famine Ireland
von: McCabe, Ciarán
Veröffentlicht: (2019)
von: McCabe, Ciarán
Veröffentlicht: (2019)
Se extrae la energía humana de Venezuela : Creación de nuevas asociaciones para el desarrollo / Michael McCabe
von: McCabe, Michael
Veröffentlicht: (1995)
von: McCabe, Michael
Veröffentlicht: (1995)
If It Is To Be, It Is Up to Me To Help: Helping Anyone Overcome Reading/Spelling Problems. A Tutor's Text for Volunteer Tutors Such as Parents, Spouses, or Friends of Dyslexics.
von: McCabe, Don
Veröffentlicht: (1991)
von: McCabe, Don
Veröffentlicht: (1991)
Ähnliche Einträge
-
Demystifying optimized prompts in language models
von: Melamed, Rimon, et al.
Veröffentlicht: (2025) -
Estimating Semantic Alphabet Size for LLM Uncertainty Quantification
von: McCabe, Lucas H., et al.
Veröffentlicht: (2025) -
Towards a theory of model distillation
von: Boix-Adsera, Enric
Veröffentlicht: (2024) -
SENECA: Small-Sample Discrete Entropy Estimation via Self-Consistent Missing Mass
von: McCabe, Lucas H., et al.
Veröffentlicht: (2026) -
Toward universal steering and monitoring of AI models
von: Beaglehole, Daniel, et al.
Veröffentlicht: (2025)