Saved in:
| Main Authors: | Ko, Myeongseob, Billa, Nikhil Reddy, Nguyen, Adam, Fleming, Charles, Jin, Ming, Jia, Ruoxi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2511.05518 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Probing Knowledge Holes in Unlearned LLMs
by: Ko, Myeongseob, et al.
Published: (2025)
by: Ko, Myeongseob, et al.
Published: (2025)
Characterizing Model-Native Skills
by: Kang, Feiyang, et al.
Published: (2026)
by: Kang, Feiyang, et al.
Published: (2026)
The Signal is in the Steps: Local Scoring for Reasoning Data Selection
by: Just, Hoang Anh, et al.
Published: (2025)
by: Just, Hoang Anh, et al.
Published: (2025)
The Cascade Equivalence Hypothesis: When Do Speech LLMs Behave Like ASR$\rightarrow$LLM Pipelines?
by: Billa, Jayadev
Published: (2026)
by: Billa, Jayadev
Published: (2026)
LLMs Can Plan Only If We Tell Them
by: Sel, Bilgehan, et al.
Published: (2025)
by: Sel, Bilgehan, et al.
Published: (2025)
Modality Collapse as Mismatched Decoding: Information-Theoretic Limits of Multimodal LLMs
by: Billa, Jayadev
Published: (2026)
by: Billa, Jayadev
Published: (2026)
From Weak Cues to Real Identities: Evaluating Inference-Driven De-Anonymization in LLM Agents
by: Ko, Myeongseob, et al.
Published: (2026)
by: Ko, Myeongseob, et al.
Published: (2026)
Intent Mismatch Causes LLMs to Get Lost in Multi-Turn Conversation
by: Liu, Geng, et al.
Published: (2026)
by: Liu, Geng, et al.
Published: (2026)
Get more for less: Principled Data Selection for Warming Up Fine-Tuning in LLMs
by: Kang, Feiyang, et al.
Published: (2024)
by: Kang, Feiyang, et al.
Published: (2024)
Supervisory Prompt Training
by: Billa, Jean Ghislain, et al.
Published: (2024)
by: Billa, Jean Ghislain, et al.
Published: (2024)
AutoScale: Scale-Aware Data Mixing for Pre-Training LLMs
by: Kang, Feiyang, et al.
Published: (2024)
by: Kang, Feiyang, et al.
Published: (2024)
The Geometric Anatomy of Capability Acquisition in Transformers
by: Billa, Jayadev
Published: (2026)
by: Billa, Jayadev
Published: (2026)
Skin-in-the-Game: Decision Making via Multi-Stakeholder Alignment in LLMs
by: Sel, Bilgehan, et al.
Published: (2024)
by: Sel, Bilgehan, et al.
Published: (2024)
Boosting Alignment for Post-Unlearning Text-to-Image Generative Models
by: Ko, Myeongseob, et al.
Published: (2024)
by: Ko, Myeongseob, et al.
Published: (2024)
TravelBench : Exploring LLM Performance in Low-Resource Domains
by: Billa, Srinivas, et al.
Published: (2025)
by: Billa, Srinivas, et al.
Published: (2025)
Just Enough Shifts: Mitigating Over-Refusal in Aligned Language Models with Targeted Representation Fine-Tuning
by: Dabas, Mahavir, et al.
Published: (2025)
by: Dabas, Mahavir, et al.
Published: (2025)
When Weak LLMs Speak with Confidence, Preference Alignment Gets Stronger
by: Afzali, Amirabbas, et al.
Published: (2026)
by: Afzali, Amirabbas, et al.
Published: (2026)
FASTTRACK: Fast and Accurate Fact Tracing for LLMs
by: Chen, Si, et al.
Published: (2024)
by: Chen, Si, et al.
Published: (2024)
Lost in Transmission: When and Why LLMs Fail to Reason Globally
by: Schnabel, Tobias, et al.
Published: (2025)
by: Schnabel, Tobias, et al.
Published: (2025)
LLM-FE: Automated Feature Engineering for Tabular Data with LLMs as Evolutionary Optimizers
by: Abhyankar, Nikhil, et al.
Published: (2025)
by: Abhyankar, Nikhil, et al.
Published: (2025)
Algorithm of Thoughts: Enhancing Exploration of Ideas in Large Language Models
by: Sel, Bilgehan, et al.
Published: (2023)
by: Sel, Bilgehan, et al.
Published: (2023)
Remembering More, Risking More: Longitudinal Safety Risks in Memory-Equipped LLM Agents
by: Al-Tawaha, Ahmad, et al.
Published: (2026)
by: Al-Tawaha, Ahmad, et al.
Published: (2026)
On A Scale From 1 to 5: Quantifying Hallucination in Faithfulness Evaluation
by: Jing, Xiaonan, et al.
Published: (2024)
by: Jing, Xiaonan, et al.
Published: (2024)
Quagmires in SFT-RL Post-Training: When High SFT Scores Mislead and What to Use Instead
by: Kang, Feiyang, et al.
Published: (2025)
by: Kang, Feiyang, et al.
Published: (2025)
Does Refusal Training in LLMs Generalize to the Past Tense?
by: Andriushchenko, Maksym, et al.
Published: (2024)
by: Andriushchenko, Maksym, et al.
Published: (2024)
Zero-Shot Confidence Estimation for Small LLMs: When Supervised Baselines Aren't Worth Training
by: Nguyen, Luong N.
Published: (2026)
by: Nguyen, Luong N.
Published: (2026)
How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs
by: Zeng, Yi, et al.
Published: (2024)
by: Zeng, Yi, et al.
Published: (2024)
Can AI Validate Science? Benchmarking LLMs for Accurate Scientific Claim $\rightarrow$ Evidence Reasoning
by: Javaji, Shashidhar Reddy, et al.
Published: (2025)
by: Javaji, Shashidhar Reddy, et al.
Published: (2025)
Lost in Stories: Consistency Bugs in Long Story Generation by LLMs
by: Li, Junjie, et al.
Published: (2026)
by: Li, Junjie, et al.
Published: (2026)
TeachLM: Post-Training LLMs for Education Using Authentic Learning Data
by: Perczel, Janos, et al.
Published: (2025)
by: Perczel, Janos, et al.
Published: (2025)
AGR: Age Group fairness Reward for Bias Mitigation in LLMs
by: Cao, Shuirong, et al.
Published: (2024)
by: Cao, Shuirong, et al.
Published: (2024)
Easy Problems That LLMs Get Wrong
by: Williams, Sean, et al.
Published: (2024)
by: Williams, Sean, et al.
Published: (2024)
What Works for 'Lost-in-the-Middle' in LLMs? A Study on GM-Extract and Mitigations
by: Gupte, Mihir, et al.
Published: (2025)
by: Gupte, Mihir, et al.
Published: (2025)
"Lost-in-the-Later": Framework for Quantifying Contextual Grounding in Large Language Models
by: Tao, Yufei, et al.
Published: (2025)
by: Tao, Yufei, et al.
Published: (2025)
Synthesizing Post-Training Data for LLMs through Multi-Agent Simulation
by: Tang, Shuo, et al.
Published: (2024)
by: Tang, Shuo, et al.
Published: (2024)
When Personalization Misleads: Understanding and Mitigating Hallucinations in Personalized LLMs
by: Sun, Zhongxiang, et al.
Published: (2026)
by: Sun, Zhongxiang, et al.
Published: (2026)
Benchmarks Saturate When The Model Gets Smarter Than The Judge
by: Ballon, Marthe, et al.
Published: (2026)
by: Ballon, Marthe, et al.
Published: (2026)
When Audio-LLMs Don't Listen: A Cross-Linguistic Study of Modality Arbitration
by: Billa, Jayadev
Published: (2026)
by: Billa, Jayadev
Published: (2026)
Lost in Sampling: Assessing Lexical Reachability in LLMs via the Word Coverage Score (WCS)
by: Awad, Samer, et al.
Published: (2026)
by: Awad, Samer, et al.
Published: (2026)
When Two LLMs Debate, Both Think They'll Win
by: Prasad, Pradyumna Shyama, et al.
Published: (2025)
by: Prasad, Pradyumna Shyama, et al.
Published: (2025)
Similar Items
-
Probing Knowledge Holes in Unlearned LLMs
by: Ko, Myeongseob, et al.
Published: (2025) -
Characterizing Model-Native Skills
by: Kang, Feiyang, et al.
Published: (2026) -
The Signal is in the Steps: Local Scoring for Reasoning Data Selection
by: Just, Hoang Anh, et al.
Published: (2025) -
The Cascade Equivalence Hypothesis: When Do Speech LLMs Behave Like ASR$\rightarrow$LLM Pipelines?
by: Billa, Jayadev
Published: (2026) -
LLMs Can Plan Only If We Tell Them
by: Sel, Bilgehan, et al.
Published: (2025)