Enregistré dans:
| Auteurs principaux: | Zhou, Xuhui, Kim, Hyunwoo, Brahman, Faeze, Jiang, Liwei, Zhu, Hao, Lu, Ximing, Xu, Frank, Lin, Bill Yuchen, Choi, Yejin, Mireshghallah, Niloofar, Bras, Ronan Le, Sap, Maarten |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | https://arxiv.org/abs/2409.16427 |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Leftover Lunch: Advantage-based Offline Reinforcement Learning for Language Models
par: Baheti, Ashutosh, et autres
Publié: (2023)
par: Baheti, Ashutosh, et autres
Publié: (2023)
Multi-Attribute Constraint Satisfaction via Language Model Rewriting
par: Baheti, Ashutosh, et autres
Publié: (2024)
par: Baheti, Ashutosh, et autres
Publié: (2024)
WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
par: Jiang, Liwei, et autres
Publié: (2024)
par: Jiang, Liwei, et autres
Publié: (2024)
Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory
par: Mireshghallah, Niloofar, et autres
Publié: (2023)
par: Mireshghallah, Niloofar, et autres
Publié: (2023)
From Dogwhistles to Bullhorns: Unveiling Coded Rhetoric with Language Models
par: Mendelsohn, Julia, et autres
Publié: (2023)
par: Mendelsohn, Julia, et autres
Publié: (2023)
WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
par: Lin, Bill Yuchen, et autres
Publié: (2024)
par: Lin, Bill Yuchen, et autres
Publié: (2024)
Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement
par: Jung, Jaehun, et autres
Publié: (2024)
par: Jung, Jaehun, et autres
Publié: (2024)
Information-Theoretic Distillation for Reference-less Summarization
par: Jung, Jaehun, et autres
Publié: (2024)
par: Jung, Jaehun, et autres
Publié: (2024)
AI-LieDar: Examine the Trade-off Between Utility and Truthfulness in LLM Agents
par: Su, Zhe, et autres
Publié: (2024)
par: Su, Zhe, et autres
Publié: (2024)
Impossible Distillation: from Low-Quality Model to High-Quality Dataset & Model for Summarization and Paraphrasing
par: Jung, Jaehun, et autres
Publié: (2023)
par: Jung, Jaehun, et autres
Publié: (2023)
MacGyver: Are Large Language Models Creative Problem Solvers?
par: Tian, Yufei, et autres
Publié: (2023)
par: Tian, Yufei, et autres
Publié: (2023)
Agent Lumos: Unified and Modular Training for Open-Source Language Agents
par: Yin, Da, et autres
Publié: (2023)
par: Yin, Da, et autres
Publié: (2023)
Is this the real life? Is this just fantasy? The Misleading Success of Simulating Social Interactions With LLMs
par: Zhou, Xuhui, et autres
Publié: (2024)
par: Zhou, Xuhui, et autres
Publié: (2024)
Trust No Bot: Discovering Personal Disclosures in Human-LLM Conversations in the Wild
par: Mireshghallah, Niloofar, et autres
Publié: (2024)
par: Mireshghallah, Niloofar, et autres
Publié: (2024)
ZebraLogic: On the Scaling Limits of LLMs for Logical Reasoning
par: Lin, Bill Yuchen, et autres
Publié: (2025)
par: Lin, Bill Yuchen, et autres
Publié: (2025)
Information-Guided Identification of Training Data Imprint in (Proprietary) Large Language Models
par: Ravichander, Abhilasha, et autres
Publié: (2025)
par: Ravichander, Abhilasha, et autres
Publié: (2025)
AI as Humanity's Salieri: Quantifying Linguistic Creativity of Language Models via Systematic Attribution of Machine Text against Web Text
par: Lu, Ximing, et autres
Publié: (2024)
par: Lu, Ximing, et autres
Publié: (2024)
How to Train Your Fact Verifier: Knowledge Transfer with Multimodal Open Models
par: Lee, Jaeyoung, et autres
Publié: (2024)
par: Lee, Jaeyoung, et autres
Publié: (2024)
ParaPO: Aligning Language Models to Reduce Verbatim Reproduction of Pre-training Data
par: Chen, Tong, et autres
Publié: (2025)
par: Chen, Tong, et autres
Publié: (2025)
SimpleToM: Exposing the Gap between Explicit ToM Inference and Implicit ToM Application in LLMs
par: Gu, Yuling, et autres
Publié: (2024)
par: Gu, Yuling, et autres
Publié: (2024)
GoodPoint: Learning Constructive Scientific Paper Feedback from Author Responses
par: Mun, Jimin, et autres
Publié: (2026)
par: Mun, Jimin, et autres
Publié: (2026)
Social World Models
par: Zhou, Xuhui, et autres
Publié: (2025)
par: Zhou, Xuhui, et autres
Publié: (2025)
What Makes it Ok to Set a Fire? Iterative Self-distillation of Contexts and Rationales for Disambiguating Defeasible Social and Moral Situations
par: Rao, Kavel, et autres
Publié: (2023)
par: Rao, Kavel, et autres
Publié: (2023)
RESTOR: Knowledge Recovery in Machine Unlearning
par: Rezaei, Keivan, et autres
Publié: (2024)
par: Rezaei, Keivan, et autres
Publié: (2024)
Spectrum Tuning: Post-Training for Distributional Coverage and In-Context Steerability
par: Sorensen, Taylor, et autres
Publié: (2025)
par: Sorensen, Taylor, et autres
Publié: (2025)
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
par: Han, Seungju, et autres
Publié: (2024)
par: Han, Seungju, et autres
Publié: (2024)
Alpaca against Vicuna: Using LLMs to Uncover Memorization of LLMs
par: Kassem, Aly M., et autres
Publié: (2024)
par: Kassem, Aly M., et autres
Publié: (2024)
PPMI: Privacy-Preserving LLM Interaction with Socratic Chain-of-Thought Reasoning and Homomorphically Encrypted Vector Databases
par: Bae, Yubeen, et autres
Publié: (2025)
par: Bae, Yubeen, et autres
Publié: (2025)
Let Them Down Easy! Contextual Effects of LLM Guardrails on User Perceptions and Preferences
par: Zheng, Mingqian, et autres
Publié: (2025)
par: Zheng, Mingqian, et autres
Publié: (2025)
Value Kaleidoscope: Engaging AI with Pluralistic Human Values, Rights, and Duties
par: Sorensen, Taylor, et autres
Publié: (2023)
par: Sorensen, Taylor, et autres
Publié: (2023)
In Search of the Long-Tail: Systematic Generation of Long-Tail Inferential Knowledge via Logical Rule Guided Search
par: Li, Huihan, et autres
Publié: (2023)
par: Li, Huihan, et autres
Publié: (2023)
A Roadmap to Pluralistic Alignment
par: Sorensen, Taylor, et autres
Publié: (2024)
par: Sorensen, Taylor, et autres
Publié: (2024)
Position: Privacy Is Not Just Memorization!
par: Mireshghallah, Niloofar, et autres
Publié: (2025)
par: Mireshghallah, Niloofar, et autres
Publié: (2025)
Synthetic Data Can Mislead Evaluations: Membership Inference as Machine Text Detection
par: Naseh, Ali, et autres
Publié: (2025)
par: Naseh, Ali, et autres
Publié: (2025)
Ambig-SWE: Interactive Agents to Overcome Underspecificity in Software Engineering
par: Vijayvargiya, Sanidhya, et autres
Publié: (2025)
par: Vijayvargiya, Sanidhya, et autres
Publié: (2025)
A False Sense of Privacy: Evaluating Textual Data Sanitization Beyond Surface-level Privacy Leakage
par: Xin, Rui, et autres
Publié: (2025)
par: Xin, Rui, et autres
Publié: (2025)
The Surprising Effectiveness of Membership Inference with Simple N-Gram Coverage
par: Hallinan, Skyler, et autres
Publié: (2025)
par: Hallinan, Skyler, et autres
Publié: (2025)
SafetyAnalyst: Interpretable, Transparent, and Steerable Safety Moderation for AI Behavior
par: Li, Jing-Jing, et autres
Publié: (2024)
par: Li, Jing-Jing, et autres
Publié: (2024)
A Call for Clarity in Beam Search: How It Works and When It Stops
par: Kasai, Jungo, et autres
Publié: (2022)
par: Kasai, Jungo, et autres
Publié: (2022)
Particip-AI: A Democratic Surveying Framework for Anticipating Future AI Use Cases, Harms and Benefits
par: Mun, Jimin, et autres
Publié: (2024)
par: Mun, Jimin, et autres
Publié: (2024)
Documents similaires
-
Leftover Lunch: Advantage-based Offline Reinforcement Learning for Language Models
par: Baheti, Ashutosh, et autres
Publié: (2023) -
Multi-Attribute Constraint Satisfaction via Language Model Rewriting
par: Baheti, Ashutosh, et autres
Publié: (2024) -
WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
par: Jiang, Liwei, et autres
Publié: (2024) -
Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory
par: Mireshghallah, Niloofar, et autres
Publié: (2023) -
From Dogwhistles to Bullhorns: Unveiling Coded Rhetoric with Language Models
par: Mendelsohn, Julia, et autres
Publié: (2023)