Retracing the Past: LLMs Emit Training Data When They Get Lost
Fuente:
arXiv
Salvato in:
| Autori principali: | Ko, Myeongseob, Billa, Nikhil Reddy, Nguyen, Adam, Fleming, Charles, Jin, Ming, Jia, Ruoxi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Probing Knowledge Holes in Unlearned LLMs
di: Ko, Myeongseob, et al.
Pubblicazione: (2025)
di: Ko, Myeongseob, et al.
Pubblicazione: (2025)
Characterizing Model-Native Skills
di: Kang, Feiyang, et al.
Pubblicazione: (2026)
di: Kang, Feiyang, et al.
Pubblicazione: (2026)
The Signal is in the Steps: Local Scoring for Reasoning Data Selection
di: Just, Hoang Anh, et al.
Pubblicazione: (2025)
di: Just, Hoang Anh, et al.
Pubblicazione: (2025)
The Cascade Equivalence Hypothesis: When Do Speech LLMs Behave Like ASR$\rightarrow$LLM Pipelines?
di: Billa, Jayadev
Pubblicazione: (2026)
di: Billa, Jayadev
Pubblicazione: (2026)
LLMs Can Plan Only If We Tell Them
di: Sel, Bilgehan, et al.
Pubblicazione: (2025)
di: Sel, Bilgehan, et al.
Pubblicazione: (2025)
Modality Collapse as Mismatched Decoding: Information-Theoretic Limits of Multimodal LLMs
di: Billa, Jayadev
Pubblicazione: (2026)
di: Billa, Jayadev
Pubblicazione: (2026)
Intent Mismatch Causes LLMs to Get Lost in Multi-Turn Conversation
di: Liu, Geng, et al.
Pubblicazione: (2026)
di: Liu, Geng, et al.
Pubblicazione: (2026)
From Weak Cues to Real Identities: Evaluating Inference-Driven De-Anonymization in LLM Agents
di: Ko, Myeongseob, et al.
Pubblicazione: (2026)
di: Ko, Myeongseob, et al.
Pubblicazione: (2026)
Get more for less: Principled Data Selection for Warming Up Fine-Tuning in LLMs
di: Kang, Feiyang, et al.
Pubblicazione: (2024)
di: Kang, Feiyang, et al.
Pubblicazione: (2024)
Supervisory Prompt Training
di: Billa, Jean Ghislain, et al.
Pubblicazione: (2024)
di: Billa, Jean Ghislain, et al.
Pubblicazione: (2024)
AutoScale: Scale-Aware Data Mixing for Pre-Training LLMs
di: Kang, Feiyang, et al.
Pubblicazione: (2024)
di: Kang, Feiyang, et al.
Pubblicazione: (2024)
The Geometric Anatomy of Capability Acquisition in Transformers
di: Billa, Jayadev
Pubblicazione: (2026)
di: Billa, Jayadev
Pubblicazione: (2026)
When Weak LLMs Speak with Confidence, Preference Alignment Gets Stronger
di: Afzali, Amirabbas, et al.
Pubblicazione: (2026)
di: Afzali, Amirabbas, et al.
Pubblicazione: (2026)
Skin-in-the-Game: Decision Making via Multi-Stakeholder Alignment in LLMs
di: Sel, Bilgehan, et al.
Pubblicazione: (2024)
di: Sel, Bilgehan, et al.
Pubblicazione: (2024)
TravelBench : Exploring LLM Performance in Low-Resource Domains
di: Billa, Srinivas, et al.
Pubblicazione: (2025)
di: Billa, Srinivas, et al.
Pubblicazione: (2025)
FASTTRACK: Fast and Accurate Fact Tracing for LLMs
di: Chen, Si, et al.
Pubblicazione: (2024)
di: Chen, Si, et al.
Pubblicazione: (2024)
Lost in Transmission: When and Why LLMs Fail to Reason Globally
di: Schnabel, Tobias, et al.
Pubblicazione: (2025)
di: Schnabel, Tobias, et al.
Pubblicazione: (2025)
Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning
di: Dabas, Mahavir, et al.
Pubblicazione: (2025)
di: Dabas, Mahavir, et al.
Pubblicazione: (2025)
Boosting Alignment for Post-Unlearning Text-to-Image Generative Models
di: Ko, Myeongseob, et al.
Pubblicazione: (2024)
di: Ko, Myeongseob, et al.
Pubblicazione: (2024)
Algorithm of Thoughts: Enhancing Exploration of Ideas in Large Language Models
di: Sel, Bilgehan, et al.
Pubblicazione: (2023)
di: Sel, Bilgehan, et al.
Pubblicazione: (2023)
Remembering More, Risking More: Longitudinal Safety Risks in Memory-Equipped LLM Agents
di: Al-Tawaha, Ahmad, et al.
Pubblicazione: (2026)
di: Al-Tawaha, Ahmad, et al.
Pubblicazione: (2026)
Zero-Shot Confidence Estimation for Small LLMs: When Supervised Baselines Aren't Worth Training
di: Nguyen, Luong N.
Pubblicazione: (2026)
di: Nguyen, Luong N.
Pubblicazione: (2026)
Does Refusal Training in LLMs Generalize to the Past Tense?
di: Andriushchenko, Maksym, et al.
Pubblicazione: (2024)
di: Andriushchenko, Maksym, et al.
Pubblicazione: (2024)
LLM-FE: Automated Feature Engineering for Tabular Data with LLMs as Evolutionary Optimizers
di: Abhyankar, Nikhil, et al.
Pubblicazione: (2025)
di: Abhyankar, Nikhil, et al.
Pubblicazione: (2025)
Quagmires in SFT-RL Post-Training: When High SFT Scores Mislead and What to Use Instead
di: Kang, Feiyang, et al.
Pubblicazione: (2025)
di: Kang, Feiyang, et al.
Pubblicazione: (2025)
On A Scale From 1 to 5: Quantifying Hallucination in Faithfulness Evaluation
di: Jing, Xiaonan, et al.
Pubblicazione: (2024)
di: Jing, Xiaonan, et al.
Pubblicazione: (2024)
How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs
di: Zeng, Yi, et al.
Pubblicazione: (2024)
di: Zeng, Yi, et al.
Pubblicazione: (2024)
Lost in Stories: Consistency Bugs in Long Story Generation by LLMs
di: Li, Junjie, et al.
Pubblicazione: (2026)
di: Li, Junjie, et al.
Pubblicazione: (2026)
TeachLM: Post-Training LLMs for Education Using Authentic Learning Data
di: Perczel, Janos, et al.
Pubblicazione: (2025)
di: Perczel, Janos, et al.
Pubblicazione: (2025)
Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning
di: Javaji, Shashidhar Reddy, et al.
Pubblicazione: (2025)
di: Javaji, Shashidhar Reddy, et al.
Pubblicazione: (2025)
What Works for 'Lost-in-the-Middle' in LLMs? A Study on GM-Extract and Mitigations
di: Gupte, Mihir, et al.
Pubblicazione: (2025)
di: Gupte, Mihir, et al.
Pubblicazione: (2025)
Synthesizing Post-Training Data for LLMs through Multi-Agent Simulation
di: Tang, Shuo, et al.
Pubblicazione: (2024)
di: Tang, Shuo, et al.
Pubblicazione: (2024)
Easy Problems That LLMs Get Wrong
di: Williams, Sean, et al.
Pubblicazione: (2024)
di: Williams, Sean, et al.
Pubblicazione: (2024)
"Lost-in-the-Later": Framework for Quantifying Contextual Grounding in Large Language Models
di: Tao, Yufei, et al.
Pubblicazione: (2025)
di: Tao, Yufei, et al.
Pubblicazione: (2025)
When Personalization Misleads: Understanding and Mitigating Hallucinations in Personalized LLMs
di: Sun, Zhongxiang, et al.
Pubblicazione: (2026)
di: Sun, Zhongxiang, et al.
Pubblicazione: (2026)
AGR: Age Group fairness Reward for Bias Mitigation in LLMs
di: Cao, Shuirong, et al.
Pubblicazione: (2024)
di: Cao, Shuirong, et al.
Pubblicazione: (2024)
Lost in Sampling: Assessing Lexical Reachability in LLMs via the Word Coverage Score (WCS)
di: Awad, Samer, et al.
Pubblicazione: (2026)
di: Awad, Samer, et al.
Pubblicazione: (2026)
Benchmarks Saturate When The Model Gets Smarter Than The Judge
di: Ballon, Marthe, et al.
Pubblicazione: (2026)
di: Ballon, Marthe, et al.
Pubblicazione: (2026)
When Two LLMs Debate, Both Think They'll Win
di: Prasad, Pradyumna Shyama, et al.
Pubblicazione: (2025)
di: Prasad, Pradyumna Shyama, et al.
Pubblicazione: (2025)
From Fallback to Frontline: When Can LLMs be Superior Annotators of Human Perspectives?
di: Amin, Hasan, et al.
Pubblicazione: (2026)
di: Amin, Hasan, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Probing Knowledge Holes in Unlearned LLMs
di: Ko, Myeongseob, et al.
Pubblicazione: (2025) -
Characterizing Model-Native Skills
di: Kang, Feiyang, et al.
Pubblicazione: (2026) -
The Signal is in the Steps: Local Scoring for Reasoning Data Selection
di: Just, Hoang Anh, et al.
Pubblicazione: (2025) -
The Cascade Equivalence Hypothesis: When Do Speech LLMs Behave Like ASR$\rightarrow$LLM Pipelines?
di: Billa, Jayadev
Pubblicazione: (2026) -
LLMs Can Plan Only If We Tell Them
di: Sel, Bilgehan, et al.
Pubblicazione: (2025)