Salvato in:
| Autori principali: | Goldstein, Daniel, Alcaide, Eric, Lu, Janna, Cheah, Eugene |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2505.03005 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Key-Value Means: Transformers with Expandable Block-Recurrent Compressed Memory
di: Goldstein, Daniel, et al.
Pubblicazione: (2026)
di: Goldstein, Daniel, et al.
Pubblicazione: (2026)
RWKV-7 "Goose" with Expressive Dynamic State Evolution
di: Peng, Bo, et al.
Pubblicazione: (2025)
di: Peng, Bo, et al.
Pubblicazione: (2025)
Forget Attention: Importance-Aware Attention Is All You Need
di: Shin, Soohyeong, et al.
Pubblicazione: (2026)
di: Shin, Soohyeong, et al.
Pubblicazione: (2026)
Adversarial Lens: Exploiting Attention Layers to Generate Adversarial Examples for Evaluation
di: Dhole, Kaustubh
Pubblicazione: (2025)
di: Dhole, Kaustubh
Pubblicazione: (2025)
AtManRL: Towards Faithful Reasoning via Differentiable Attention Saliency
di: Höth, Max Henning, et al.
Pubblicazione: (2026)
di: Höth, Max Henning, et al.
Pubblicazione: (2026)
Mixture of Attention Spans: Optimizing LLM Inference Efficiency with Heterogeneous Sliding-Window Lengths
di: Fu, Tianyu, et al.
Pubblicazione: (2024)
di: Fu, Tianyu, et al.
Pubblicazione: (2024)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
di: Oketunji, Abiodun Finbarrs
Pubblicazione: (2023)
di: Oketunji, Abiodun Finbarrs
Pubblicazione: (2023)
Encoder vs Decoder: Comparative Analysis of Encoder and Decoder Language Models on Multilingual NLU Tasks
di: Nielsen, Dan Saattrup, et al.
Pubblicazione: (2024)
di: Nielsen, Dan Saattrup, et al.
Pubblicazione: (2024)
SSSD: Simply-Scalable Speculative Decoding
di: Marzollo, Michele, et al.
Pubblicazione: (2024)
di: Marzollo, Michele, et al.
Pubblicazione: (2024)
Attention Drift: What Autoregressive Speculative Decoding Models Learn
di: Eldenk, Doğaç, et al.
Pubblicazione: (2026)
di: Eldenk, Doğaç, et al.
Pubblicazione: (2026)
Dodo: Dynamic Contextual Compression for Decoder-only LMs
di: Qin, Guanghui, et al.
Pubblicazione: (2023)
di: Qin, Guanghui, et al.
Pubblicazione: (2023)
Weakly Supervised Distillation of Hallucination Signals into Transformer Representations
di: Salehmohamed, Shoaib Sadiq, et al.
Pubblicazione: (2026)
di: Salehmohamed, Shoaib Sadiq, et al.
Pubblicazione: (2026)
Large Language Model (LLM) Bias Index -- LLMBI
di: Oketunji, Abiodun Finbarrs, et al.
Pubblicazione: (2023)
di: Oketunji, Abiodun Finbarrs, et al.
Pubblicazione: (2023)
Distilling Knowledge from Large Language Models: A Concept Bottleneck Model for Hate and Counter Speech Recognition
di: Labadie-Tamayo, Roberto, et al.
Pubblicazione: (2025)
di: Labadie-Tamayo, Roberto, et al.
Pubblicazione: (2025)
Turning the TIDE: Cross-Architecture Distillation for Diffusion Large Language Models
di: Zhang, Gongbo, et al.
Pubblicazione: (2026)
di: Zhang, Gongbo, et al.
Pubblicazione: (2026)
Harnessing Negative Signals: Reinforcement Distillation from Teacher Data for LLM Reasoning
di: Xu, Shuyao, et al.
Pubblicazione: (2025)
di: Xu, Shuyao, et al.
Pubblicazione: (2025)
Decodable but Not Corrected by Fixed Residual-Stream Linear Steering: Evidence from Medical LLM Failure Regimes
di: Liu, Ming
Pubblicazione: (2026)
di: Liu, Ming
Pubblicazione: (2026)
Revisiting Intermediate-Layer Matching in Knowledge Distillation: Layer-Selection Strategy Doesn't Matter (Much)
di: Yu, Zony, et al.
Pubblicazione: (2025)
di: Yu, Zony, et al.
Pubblicazione: (2025)
Generalizing Test-time Compute-optimal Scaling as an Optimizable Graph
di: Wang, Fali, et al.
Pubblicazione: (2025)
di: Wang, Fali, et al.
Pubblicazione: (2025)
Slim-SC: Thought Pruning for Efficient Scaling with Self-Consistency
di: Hong, Colin, et al.
Pubblicazione: (2025)
di: Hong, Colin, et al.
Pubblicazione: (2025)
Evaluating Explainable AI Attribution Methods in Neural Machine Translation via Attention-Guided Knowledge Distillation
di: Nourbakhsh, Aria, et al.
Pubblicazione: (2026)
di: Nourbakhsh, Aria, et al.
Pubblicazione: (2026)
QiMeng-Attention: SOTA Attention Operator is generated by SOTA Attention Algorithm
di: Zhou, Qirui, et al.
Pubblicazione: (2025)
di: Zhou, Qirui, et al.
Pubblicazione: (2025)
AgentTTS: Large Language Model Agent for Test-time Compute-optimal Scaling Strategy in Complex Tasks
di: Wang, Fali, et al.
Pubblicazione: (2025)
di: Wang, Fali, et al.
Pubblicazione: (2025)
Align-to-Distill: Trainable Attention Alignment for Knowledge Distillation in Neural Machine Translation
di: Jin, Heegon, et al.
Pubblicazione: (2024)
di: Jin, Heegon, et al.
Pubblicazione: (2024)
ALISON: Fast and Effective Stylometric Authorship Obfuscation
di: Xing, Eric, et al.
Pubblicazione: (2024)
di: Xing, Eric, et al.
Pubblicazione: (2024)
Pressure-Testing Deception Probes in LLMs: Scaling, Robustness, and the Geometry of Deceptive Representations
di: Kumar, Sachin
Pubblicazione: (2026)
di: Kumar, Sachin
Pubblicazione: (2026)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
di: Fadli, Samih
Pubblicazione: (2025)
di: Fadli, Samih
Pubblicazione: (2025)
Softmax Linear Attention: Reclaiming Global Competition
di: Xu, Mingwei, et al.
Pubblicazione: (2026)
di: Xu, Mingwei, et al.
Pubblicazione: (2026)
Sparse Logit Sampling: Accelerating Knowledge Distillation in LLMs
di: Anshumann, et al.
Pubblicazione: (2025)
di: Anshumann, et al.
Pubblicazione: (2025)
Alif: Advancing Urdu Large Language Models via Multilingual Synthetic Data Distillation
di: Shafique, Muhammad Ali, et al.
Pubblicazione: (2025)
di: Shafique, Muhammad Ali, et al.
Pubblicazione: (2025)
Sleepless Nights, Sugary Days: Creating Synthetic Users with Health Conditions for Realistic Coaching Agent Interactions
di: Yun, Taedong, et al.
Pubblicazione: (2025)
di: Yun, Taedong, et al.
Pubblicazione: (2025)
DIVERSED: Relaxed Speculative Decoding via Dynamic Ensemble Verification
di: Wang, Ziyi, et al.
Pubblicazione: (2026)
di: Wang, Ziyi, et al.
Pubblicazione: (2026)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
di: Peters, Sydney, et al.
Pubblicazione: (2025)
di: Peters, Sydney, et al.
Pubblicazione: (2025)
SpecExtend: A Drop-in Enhancement for Speculative Decoding of Long Sequences
di: Cha, Jungyoub, et al.
Pubblicazione: (2025)
di: Cha, Jungyoub, et al.
Pubblicazione: (2025)
ObfusQAte: A Proposed Framework to Evaluate LLM Robustness on Obfuscated Factual Question Answering
di: Ghosh, Shubhra, et al.
Pubblicazione: (2025)
di: Ghosh, Shubhra, et al.
Pubblicazione: (2025)
Dealing with Annotator Disagreement in Hate Speech Classification
di: Dehghan, Somaiyeh, et al.
Pubblicazione: (2025)
di: Dehghan, Somaiyeh, et al.
Pubblicazione: (2025)
Improving Discrete Diffusion Unmasking Policies Beyond Explicit Reference Policies
di: Hong, Chunsan, et al.
Pubblicazione: (2025)
di: Hong, Chunsan, et al.
Pubblicazione: (2025)
On Explaining with Attention Matrices
di: Naim, Omar, et al.
Pubblicazione: (2024)
di: Naim, Omar, et al.
Pubblicazione: (2024)
A Multi-Encoder Frozen-Decoder Approach for Fine-Tuning Large Language Models
di: Dhole, Kaustubh D.
Pubblicazione: (2025)
di: Dhole, Kaustubh D.
Pubblicazione: (2025)
AMALIA Technical Report: A Fully Open Source Large Language Model for European Portuguese
di: Simplício, Afonso, et al.
Pubblicazione: (2026)
di: Simplício, Afonso, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Key-Value Means: Transformers with Expandable Block-Recurrent Compressed Memory
di: Goldstein, Daniel, et al.
Pubblicazione: (2026) -
RWKV-7 "Goose" with Expressive Dynamic State Evolution
di: Peng, Bo, et al.
Pubblicazione: (2025) -
Forget Attention: Importance-Aware Attention Is All You Need
di: Shin, Soohyeong, et al.
Pubblicazione: (2026) -
Adversarial Lens: Exploiting Attention Layers to Generate Adversarial Examples for Evaluation
di: Dhole, Kaustubh
Pubblicazione: (2025) -
AtManRL: Towards Faithful Reasoning via Differentiable Attention Saliency
di: Höth, Max Henning, et al.
Pubblicazione: (2026)