When Bad Data Leads to Good Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Li, Kenneth, Chen, Yida, Viégas, Fernanda, Wattenberg, Martin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Dialogue Action Tokens: Steering Language Models in Goal-Directed Dialogue with a Multi-Turn Planner
di: Li, Kenneth, et al.
Pubblicazione: (2024)
di: Li, Kenneth, et al.
Pubblicazione: (2024)
Inference-Time Intervention: Eliciting Truthful Answers from a Language Model
di: Li, Kenneth, et al.
Pubblicazione: (2023)
di: Li, Kenneth, et al.
Pubblicazione: (2023)
Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task
di: Li, Kenneth, et al.
Pubblicazione: (2022)
di: Li, Kenneth, et al.
Pubblicazione: (2022)
Measuring and Controlling Instruction (In)Stability in Language Model Dialogs
di: Li, Kenneth, et al.
Pubblicazione: (2024)
di: Li, Kenneth, et al.
Pubblicazione: (2024)
What Does it Mean for a Neural Network to Learn a "World Model"?
di: Li, Kenneth, et al.
Pubblicazione: (2025)
di: Li, Kenneth, et al.
Pubblicazione: (2025)
Relational Composition in Neural Networks: A Survey and Call to Action
di: Wattenberg, Martin, et al.
Pubblicazione: (2024)
di: Wattenberg, Martin, et al.
Pubblicazione: (2024)
AgentCollabBench: Diagnosing When Good Agents Make Bad Collaborators
di: Mazumder, Aritra, et al.
Pubblicazione: (2026)
di: Mazumder, Aritra, et al.
Pubblicazione: (2026)
Simulation, Modelling and Classification of Wiki Contributors: Spotting The Good, The Bad, and The Ugly
di: Méndez, Silvia García, et al.
Pubblicazione: (2024)
di: Méndez, Silvia García, et al.
Pubblicazione: (2024)
FADE: Why Bad Descriptions Happen to Good Features
di: Puri, Bruno, et al.
Pubblicazione: (2025)
di: Puri, Bruno, et al.
Pubblicazione: (2025)
Shared Global and Local Geometry of Language Model Embeddings
di: Lee, Andrew, et al.
Pubblicazione: (2025)
di: Lee, Andrew, et al.
Pubblicazione: (2025)
Can Interpretation Predict Behavior on Unseen Data?
di: Li, Victoria R., et al.
Pubblicazione: (2025)
di: Li, Victoria R., et al.
Pubblicazione: (2025)
Think Before You Lie: How Reasoning Leads to Honesty
di: Yuan, Ann, et al.
Pubblicazione: (2026)
di: Yuan, Ann, et al.
Pubblicazione: (2026)
Comparing Bad Apples to Good Oranges: Aligning Large Language Models via Joint Preference Optimization
di: Bansal, Hritik, et al.
Pubblicazione: (2024)
di: Bansal, Hritik, et al.
Pubblicazione: (2024)
The Geometry of Self-Verification in a Task-Specific Reasoning Model
di: Lee, Andrew, et al.
Pubblicazione: (2025)
di: Lee, Andrew, et al.
Pubblicazione: (2025)
How Bad is Training on Synthetic Data? A Statistical Analysis of Language Model Collapse
di: Seddik, Mohamed El Amine, et al.
Pubblicazione: (2024)
di: Seddik, Mohamed El Amine, et al.
Pubblicazione: (2024)
When Does Multimodality Lead to Better Time Series Forecasting?
di: Zhang, Xiyuan, et al.
Pubblicazione: (2025)
di: Zhang, Xiyuan, et al.
Pubblicazione: (2025)
When LLM Judge Scores Look Good but Best-of-N Decisions Fail
di: Landesberg, Eddie
Pubblicazione: (2026)
di: Landesberg, Eddie
Pubblicazione: (2026)
Model Merging and Safety Alignment: One Bad Model Spoils the Bunch
di: Hammoud, Hasan Abed Al Kader, et al.
Pubblicazione: (2024)
di: Hammoud, Hasan Abed Al Kader, et al.
Pubblicazione: (2024)
The Good, the Bad, and the Ugly: The Role of AI Quality Disclosure in Lie Detection
di: Bhattacharya, Haimanti, et al.
Pubblicazione: (2024)
di: Bhattacharya, Haimanti, et al.
Pubblicazione: (2024)
What Level of Automation is "Good Enough"? A Benchmark of Large Language Models for Meta-Analysis Data Extraction
di: Li, Lingbo, et al.
Pubblicazione: (2025)
di: Li, Lingbo, et al.
Pubblicazione: (2025)
ICLR: In-Context Learning of Representations
di: Park, Core Francisco, et al.
Pubblicazione: (2024)
di: Park, Core Francisco, et al.
Pubblicazione: (2024)
When Raw Data Prevails: Are Large Language Model Embeddings Effective in Numerical Data Representation for Medical Machine Learning Applications?
di: Gao, Yanjun, et al.
Pubblicazione: (2024)
di: Gao, Yanjun, et al.
Pubblicazione: (2024)
Large Language Models are Good Relational Learners
di: Wu, Fang, et al.
Pubblicazione: (2025)
di: Wu, Fang, et al.
Pubblicazione: (2025)
Chronotome: Real-Time Topic Modeling for Streaming Embedding Spaces
di: Lim, Matte, et al.
Pubblicazione: (2025)
di: Lim, Matte, et al.
Pubblicazione: (2025)
Large Language Models Badly Generalize across Option Length, Problem Types, and Irrelevant Noun Replacements
di: Zhao, Guangxiang, et al.
Pubblicazione: (2025)
di: Zhao, Guangxiang, et al.
Pubblicazione: (2025)
Communicating Activations Between Language Model Agents
di: Ramesh, Vignav, et al.
Pubblicazione: (2025)
di: Ramesh, Vignav, et al.
Pubblicazione: (2025)
Toxicity Detection Should Measure Contextual Harm, Not Text-Intrinsic Badness
di: Berezin, Sergei, et al.
Pubblicazione: (2025)
di: Berezin, Sergei, et al.
Pubblicazione: (2025)
What Makes Good Data for Alignment? A Comprehensive Study of Automatic Data Selection in Instruction Tuning
di: Liu, Wei, et al.
Pubblicazione: (2023)
di: Liu, Wei, et al.
Pubblicazione: (2023)
Large Language Models Are Bad Dice Players: LLMs Struggle to Generate Random Numbers from Statistical Distributions
di: Zhao, Minda, et al.
Pubblicazione: (2026)
di: Zhao, Minda, et al.
Pubblicazione: (2026)
Prompting in the Dark: Assessing Human Performance in Prompt Engineering for Data Labeling When Gold Labels Are Absent
di: He, Zeyu, et al.
Pubblicazione: (2025)
di: He, Zeyu, et al.
Pubblicazione: (2025)
Scaling Up Membership Inference: When and How Attacks Succeed on Large Language Models
di: Puerto, Haritz, et al.
Pubblicazione: (2024)
di: Puerto, Haritz, et al.
Pubblicazione: (2024)
AdaptThink: Reasoning Models Can Learn When to Think
di: Zhang, Jiajie, et al.
Pubblicazione: (2025)
di: Zhang, Jiajie, et al.
Pubblicazione: (2025)
AfroBench: How Good are Large Language Models on African Languages?
di: Ojo, Jessica, et al.
Pubblicazione: (2023)
di: Ojo, Jessica, et al.
Pubblicazione: (2023)
Designing a Dashboard for Transparency and Control of Conversational AI
di: Chen, Yida, et al.
Pubblicazione: (2024)
di: Chen, Yida, et al.
Pubblicazione: (2024)
Large Language Models Generate Harmful Content Using a Distinct, Unified Mechanism
di: Orgad, Hadas, et al.
Pubblicazione: (2026)
di: Orgad, Hadas, et al.
Pubblicazione: (2026)
ReAct Meets ActRe: When Language Agents Enjoy Training Data Autonomy
di: Yang, Zonghan, et al.
Pubblicazione: (2024)
di: Yang, Zonghan, et al.
Pubblicazione: (2024)
Large Language Models Are Latent Variable Models: Explaining and Finding Good Demonstrations for In-Context Learning
di: Wang, Xinyi, et al.
Pubblicazione: (2023)
di: Wang, Xinyi, et al.
Pubblicazione: (2023)
What Makes a Reward Model a Good Teacher? An Optimization Perspective
di: Razin, Noam, et al.
Pubblicazione: (2025)
di: Razin, Noam, et al.
Pubblicazione: (2025)
Stuffed Mamba: Oversized States Lead to the Inability to Forget
di: Chen, Yingfa, et al.
Pubblicazione: (2024)
di: Chen, Yingfa, et al.
Pubblicazione: (2024)
Are Hallucinations Bad Estimations?
di: Liu, Hude, et al.
Pubblicazione: (2025)
di: Liu, Hude, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Dialogue Action Tokens: Steering Language Models in Goal-Directed Dialogue with a Multi-Turn Planner
di: Li, Kenneth, et al.
Pubblicazione: (2024) -
Inference-Time Intervention: Eliciting Truthful Answers from a Language Model
di: Li, Kenneth, et al.
Pubblicazione: (2023) -
Emergent World Representations: Exploring a Sequence Model Trained on a Synthetic Task
di: Li, Kenneth, et al.
Pubblicazione: (2022) -
Measuring and Controlling Instruction (In)Stability in Language Model Dialogs
di: Li, Kenneth, et al.
Pubblicazione: (2024) -
What Does it Mean for a Neural Network to Learn a "World Model"?
di: Li, Kenneth, et al.
Pubblicazione: (2025)