Gespeichert in:
| Hauptverfasser: | Wu, Siyang, Bao, Honglin, Li, Sida, Holtzman, Ari, Evans, James A. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2509.23488 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LLM Probability Concentration: How Alignment Shrinks the Generative Horizon
von: Yang, Chenghao, et al.
Veröffentlicht: (2025)
von: Yang, Chenghao, et al.
Veröffentlicht: (2025)
Linearly Decoding Refused Knowledge in Aligned Language Models
von: Shrivastava, Aryan, et al.
Veröffentlicht: (2025)
von: Shrivastava, Aryan, et al.
Veröffentlicht: (2025)
Predicting vs. Acting: A Trade-off Between World Modeling & Agent Modeling
von: Li, Margaret, et al.
Veröffentlicht: (2024)
von: Li, Margaret, et al.
Veröffentlicht: (2024)
Forking Paths in Neural Text Generation
von: Bigelow, Eric, et al.
Veröffentlicht: (2024)
von: Bigelow, Eric, et al.
Veröffentlicht: (2024)
The Text Uncanny Valley: Non-Monotonic Performance Degradation in LLM Information Retrieval
von: Tong, Zekai, et al.
Veröffentlicht: (2026)
von: Tong, Zekai, et al.
Veröffentlicht: (2026)
Know Thyself? On the Incapability and Implications of AI Self-Recognition
von: Bai, Xiaoyan, et al.
Veröffentlicht: (2025)
von: Bai, Xiaoyan, et al.
Veröffentlicht: (2025)
Moral Mazes in the Era of LLMs
von: Nguyen, Dang, et al.
Veröffentlicht: (2026)
von: Nguyen, Dang, et al.
Veröffentlicht: (2026)
Beyond Perplexity: A Lightweight Benchmark for Knowledge Retention in Supervised Fine-Tuning
von: Shabgahi, Soheil Zibakhsh, et al.
Veröffentlicht: (2026)
von: Shabgahi, Soheil Zibakhsh, et al.
Veröffentlicht: (2026)
The Story is Not the Science: Execution-Grounded Evaluation of Mechanistic Interpretability Research
von: Bai, Xiaoyan, et al.
Veröffentlicht: (2026)
von: Bai, Xiaoyan, et al.
Veröffentlicht: (2026)
Perplexity Trap: PLM-Based Retrievers Overrate Low Perplexity Documents
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
von: Wang, Haoyu, et al.
Veröffentlicht: (2025)
Missing vs. Unused Knowledge Hypothesis for Language Model Bottlenecks in Patent Understanding
von: Wu, Siyang, et al.
Veröffentlicht: (2025)
von: Wu, Siyang, et al.
Veröffentlicht: (2025)
Narrative Flattening: How Post-Training Compresses Thematic, Affective, and Stylistic Variation in LLM Fiction
von: Li, Zehan, et al.
Veröffentlicht: (2026)
von: Li, Zehan, et al.
Veröffentlicht: (2026)
PLPP: Prompt Learning with Perplexity Is Self-Distillation for Vision-Language Models
von: Liu, Biao, et al.
Veröffentlicht: (2024)
von: Liu, Biao, et al.
Veröffentlicht: (2024)
AI as Entertainment
von: Kommers, Cody, et al.
Veröffentlicht: (2026)
von: Kommers, Cody, et al.
Veröffentlicht: (2026)
Rethinking GSPO: The Perplexity-Entropy Equivalence
von: Liu, Chi
Veröffentlicht: (2025)
von: Liu, Chi
Veröffentlicht: (2025)
WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
von: Lin, Bill Yuchen, et al.
Veröffentlicht: (2024)
von: Lin, Bill Yuchen, et al.
Veröffentlicht: (2024)
On the Fallacy of Global Token Perplexity in Spoken Language Model Evaluation
von: Hsu, Chan-Jan, et al.
Veröffentlicht: (2026)
von: Hsu, Chan-Jan, et al.
Veröffentlicht: (2026)
Measuring all the noises of LLM Evals
von: Wang, Sida
Veröffentlicht: (2025)
von: Wang, Sida
Veröffentlicht: (2025)
Language Models Should be Used to Surface the Unwritten Code of Science and Society
von: Bao, Honglin, et al.
Veröffentlicht: (2025)
von: Bao, Honglin, et al.
Veröffentlicht: (2025)
Benchmarking LLM Tool-Use in the Wild
von: Yu, Peijie, et al.
Veröffentlicht: (2026)
von: Yu, Peijie, et al.
Veröffentlicht: (2026)
HowkGPT: Investigating the Detection of ChatGPT-generated University Student Homework through Context-Aware Perplexity Analysis
von: Vasilatos, Christoforos, et al.
Veröffentlicht: (2023)
von: Vasilatos, Christoforos, et al.
Veröffentlicht: (2023)
Model Organisms Are Leaky: Perplexity Differencing Often Reveals Finetuning Objectives
von: Baker, Mohammed Abu, et al.
Veröffentlicht: (2026)
von: Baker, Mohammed Abu, et al.
Veröffentlicht: (2026)
When Perplexity Lies: Generation-Focused Distillation of Hybrid Sequence Models
von: Kostelec, Juan Gabriel, et al.
Veröffentlicht: (2026)
von: Kostelec, Juan Gabriel, et al.
Veröffentlicht: (2026)
MUSE: Machine Unlearning Six-Way Evaluation for Language Models
von: Shi, Weijia, et al.
Veröffentlicht: (2024)
von: Shi, Weijia, et al.
Veröffentlicht: (2024)
Contrastive Attribution in the Wild: An Interpretability Analysis of LLM Failures on Realistic Benchmarks
von: Tan, Rongyuan, et al.
Veröffentlicht: (2026)
von: Tan, Rongyuan, et al.
Veröffentlicht: (2026)
Cross-Session Threats in AI Agents: Benchmark, Evaluation, and Algorithms
von: Azarafrooz, Ari
Veröffentlicht: (2026)
von: Azarafrooz, Ari
Veröffentlicht: (2026)
Momentum Point-Perplexity Mechanics in Large Language Models
von: Tomaz, Lorenzo, et al.
Veröffentlicht: (2025)
von: Tomaz, Lorenzo, et al.
Veröffentlicht: (2025)
Perplexity Cannot Always Tell Right from Wrong
von: Veličković, Petar, et al.
Veröffentlicht: (2026)
von: Veličković, Petar, et al.
Veröffentlicht: (2026)
How do Language Models Generate Slang: A Systematic Comparison between Human and Machine-Generated Slang Usages
von: Wu, Siyang, et al.
Veröffentlicht: (2025)
von: Wu, Siyang, et al.
Veröffentlicht: (2025)
Benchmarking Contextual and Paralinguistic Reasoning in Speech-LLMs: A Case Study with In-the-Wild Data
von: Wang, Qiongqiong, et al.
Veröffentlicht: (2025)
von: Wang, Qiongqiong, et al.
Veröffentlicht: (2025)
Inconsistent Tokenizations Cause Language Models to be Perplexed by Japanese Grammar
von: Gambardella, Andrew, et al.
Veröffentlicht: (2025)
von: Gambardella, Andrew, et al.
Veröffentlicht: (2025)
Why Slop Matters
von: Kommers, Cody, et al.
Veröffentlicht: (2025)
von: Kommers, Cody, et al.
Veröffentlicht: (2025)
Not all tokens are created equal: Perplexity Attention Weighted Networks for AI generated text detection
von: Miralles-González, Pablo, et al.
Veröffentlicht: (2025)
von: Miralles-González, Pablo, et al.
Veröffentlicht: (2025)
MathFusion: Enhancing Mathematical Problem-solving of LLM through Instruction Fusion
von: Pei, Qizhi, et al.
Veröffentlicht: (2025)
von: Pei, Qizhi, et al.
Veröffentlicht: (2025)
ProAgent: Harnessing On-Demand Sensory Contexts for Proactive LLM Agent Systems in the Wild
von: Yang, Bufang, et al.
Veröffentlicht: (2025)
von: Yang, Bufang, et al.
Veröffentlicht: (2025)
Needle in the Web: A Benchmark for Retrieving Targeted Web Pages in the Wild
von: Wang, Yumeng, et al.
Veröffentlicht: (2025)
von: Wang, Yumeng, et al.
Veröffentlicht: (2025)
VibeSearchBench: Benchmarking Long-horizon Proactive Search in the Wild
von: Inc, Xiaohongshu
Veröffentlicht: (2026)
von: Inc, Xiaohongshu
Veröffentlicht: (2026)
StackingNet: Collective Inference Across Independent AI Foundation Models
von: Li, Siyang, et al.
Veröffentlicht: (2026)
von: Li, Siyang, et al.
Veröffentlicht: (2026)
RePPL: Recalibrating Perplexity by Uncertainty in Semantic Propagation and Language Generation for Explainable QA Hallucination Detection
von: Huang, Yiming, et al.
Veröffentlicht: (2025)
von: Huang, Yiming, et al.
Veröffentlicht: (2025)
Robust Guidance for Unsupervised Data Selection: Capturing Perplexing Named Entities for Domain-Specific Machine Translation
von: Ji, Seunghyun, et al.
Veröffentlicht: (2024)
von: Ji, Seunghyun, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
LLM Probability Concentration: How Alignment Shrinks the Generative Horizon
von: Yang, Chenghao, et al.
Veröffentlicht: (2025) -
Linearly Decoding Refused Knowledge in Aligned Language Models
von: Shrivastava, Aryan, et al.
Veröffentlicht: (2025) -
Predicting vs. Acting: A Trade-off Between World Modeling & Agent Modeling
von: Li, Margaret, et al.
Veröffentlicht: (2024) -
Forking Paths in Neural Text Generation
von: Bigelow, Eric, et al.
Veröffentlicht: (2024) -
The Text Uncanny Valley: Non-Monotonic Performance Degradation in LLM Information Retrieval
von: Tong, Zekai, et al.
Veröffentlicht: (2026)