Salvato in:
| Autori principali: | Zhao, Mingkuan, Hu, Wentao, Wang, Jiayin, Lai, Xin, Huang, Tianchen, Min, Yuheng, Yan, Rui, Zhu, Xiaoyan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2511.09596 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Fast Quiet-STaR: Thinking Without Thought Tokens
di: Huang, Wei, et al.
Pubblicazione: (2025)
di: Huang, Wei, et al.
Pubblicazione: (2025)
D-SMART: Enhancing LLM Dialogue Consistency via Dynamic Structured Memory And Reasoning Tree
di: Lei, Xiang, et al.
Pubblicazione: (2025)
di: Lei, Xiang, et al.
Pubblicazione: (2025)
Train-Attention: Meta-Learning Where to Focus in Continual Knowledge Learning
di: Seo, Yeongbin, et al.
Pubblicazione: (2024)
di: Seo, Yeongbin, et al.
Pubblicazione: (2024)
WeDLM: Reconciling Diffusion Language Models with Standard Causal Attention for Fast Inference
di: Liu, Aiwei, et al.
Pubblicazione: (2025)
di: Liu, Aiwei, et al.
Pubblicazione: (2025)
When Does Content-Based Routing Work? Representation Requirements for Selective Attention in Hybrid Sequence Models
di: Basu, Abhinaba
Pubblicazione: (2026)
di: Basu, Abhinaba
Pubblicazione: (2026)
From Brazilian Portuguese to European Portuguese
di: Sanches, João, et al.
Pubblicazione: (2024)
di: Sanches, João, et al.
Pubblicazione: (2024)
Towards Effective and Efficient Continual Pre-training of Large Language Models
di: Chen, Jie, et al.
Pubblicazione: (2024)
di: Chen, Jie, et al.
Pubblicazione: (2024)
Prompt Engineering and the Effectiveness of Large Language Models in Enhancing Human Productivity
di: Anam, Rizal Khoirul
Pubblicazione: (2025)
di: Anam, Rizal Khoirul
Pubblicazione: (2025)
Transactional Attention: Semantic Sponsorship for KV-Cache Retention
di: Basu, Abhinaba
Pubblicazione: (2026)
di: Basu, Abhinaba
Pubblicazione: (2026)
How Pruning Reshapes Features: Sparse Autoencoder Analysis of Weight-Pruned Language Models
di: Borobia, Hector, et al.
Pubblicazione: (2026)
di: Borobia, Hector, et al.
Pubblicazione: (2026)
CLMN: Concept based Language Models via Neural Symbolic Reasoning
di: Yang, Yibo
Pubblicazione: (2025)
di: Yang, Yibo
Pubblicazione: (2025)
Fact Grounded Attention: Eliminating Hallucination in Large Language Models Through Attention Level Knowledge Integration
di: Gupta, Aayush
Pubblicazione: (2025)
di: Gupta, Aayush
Pubblicazione: (2025)
Softmax Linear Attention: Reclaiming Global Competition
di: Xu, Mingwei, et al.
Pubblicazione: (2026)
di: Xu, Mingwei, et al.
Pubblicazione: (2026)
Co-NAML-LSTUR: A Combined Model with Attentive Multi-View Learning and Long- and Short-term User Representations for News Recommendation
di: Nguyen, Minh Hoang, et al.
Pubblicazione: (2025)
di: Nguyen, Minh Hoang, et al.
Pubblicazione: (2025)
Mitigating LLM Hallucinations through Domain-Grounded Tiered Retrieval
di: Haque, Md. Asraful, et al.
Pubblicazione: (2026)
di: Haque, Md. Asraful, et al.
Pubblicazione: (2026)
Multipole Semantic Attention: A Fast Approximation of Softmax Attention for Pretraining
di: Mitchell, Rupert, et al.
Pubblicazione: (2025)
di: Mitchell, Rupert, et al.
Pubblicazione: (2025)
Align-to-Distill: Trainable Attention Alignment for Knowledge Distillation in Neural Machine Translation
di: Jin, Heegon, et al.
Pubblicazione: (2024)
di: Jin, Heegon, et al.
Pubblicazione: (2024)
Investigating Neuron Ablation in Attention Heads: The Case for Peak Activation Centering
di: Pochinkov, Nicholas, et al.
Pubblicazione: (2024)
di: Pochinkov, Nicholas, et al.
Pubblicazione: (2024)
Enhancing OCR for Sino-Vietnamese Language Processing via Fine-tuned PaddleOCRv5
di: Nguyen, Minh Hoang, et al.
Pubblicazione: (2025)
di: Nguyen, Minh Hoang, et al.
Pubblicazione: (2025)
Towards Probabilistic Question Answering Over Tabular Data
di: Shen, Chen, et al.
Pubblicazione: (2025)
di: Shen, Chen, et al.
Pubblicazione: (2025)
RHealthTwin: Towards Responsible and Multimodal Digital Twins for Personalized Well-being
di: Ferdousi, Rahatara, et al.
Pubblicazione: (2025)
di: Ferdousi, Rahatara, et al.
Pubblicazione: (2025)
Advancing Explainability in Neural Machine Translation: Analytical Metrics for Attention and Alignment Consistency
di: Mishra, Anurag
Pubblicazione: (2024)
di: Mishra, Anurag
Pubblicazione: (2024)
$δ$-STEAL: LLM Stealing Attack with Local Differential Privacy
di: Dang, Kieu, et al.
Pubblicazione: (2025)
di: Dang, Kieu, et al.
Pubblicazione: (2025)
Multi-Model Synthetic Training for Mission-Critical Small Language Models
di: Platt, Nolan, et al.
Pubblicazione: (2025)
di: Platt, Nolan, et al.
Pubblicazione: (2025)
PaperAudit-Bench: Benchmarking Error Detection in Research Papers for Critical Automated Peer Review
di: Tu, Songjun, et al.
Pubblicazione: (2026)
di: Tu, Songjun, et al.
Pubblicazione: (2026)
PairCFR: Enhancing Model Training on Paired Counterfactually Augmented Data through Contrastive Learning
di: Qiu, Xiaoqi, et al.
Pubblicazione: (2024)
di: Qiu, Xiaoqi, et al.
Pubblicazione: (2024)
Bi-Attention HateXplain : Taking into account the sequential aspect of data during explainability in a multi-task context
di: Mondjo, Ghislain Dorian Tchuente
Pubblicazione: (2026)
di: Mondjo, Ghislain Dorian Tchuente
Pubblicazione: (2026)
How much do LLMs learn from negative examples?
di: Hamdan, Shadi, et al.
Pubblicazione: (2025)
di: Hamdan, Shadi, et al.
Pubblicazione: (2025)
Communicative Agents for Slideshow Storytelling Video Generation based on LLMs
di: Fan, Jingxing, et al.
Pubblicazione: (2025)
di: Fan, Jingxing, et al.
Pubblicazione: (2025)
An Epidemiological Knowledge Graph extracted from the World Health Organization's Disease Outbreak News
di: Consoli, Sergio, et al.
Pubblicazione: (2025)
di: Consoli, Sergio, et al.
Pubblicazione: (2025)
MedMemoryBench: Benchmarking Agent Memory in Personalized Healthcare
di: Wang, Yihao, et al.
Pubblicazione: (2026)
di: Wang, Yihao, et al.
Pubblicazione: (2026)
From Noise to Diversity: Random Embedding Injection in LLM Reasoning
di: Kim, Heejun, et al.
Pubblicazione: (2026)
di: Kim, Heejun, et al.
Pubblicazione: (2026)
NOTAI.AI: Explainable Detection of Machine-Generated Text via Curvature and Feature Attribution
di: Breneur, Oleksandr Marchenko, et al.
Pubblicazione: (2026)
di: Breneur, Oleksandr Marchenko, et al.
Pubblicazione: (2026)
Rethinking the Multilingual Reasoning Gap with Layer Swap
di: Lasbordes, Maxence, et al.
Pubblicazione: (2026)
di: Lasbordes, Maxence, et al.
Pubblicazione: (2026)
Cache-to-Cache: Direct Semantic Communication Between Large Language Models
di: Fu, Tianyu, et al.
Pubblicazione: (2025)
di: Fu, Tianyu, et al.
Pubblicazione: (2025)
An Iterative Optimizing Framework for Radiology Report Summarization with ChatGPT
di: Ma, Chong, et al.
Pubblicazione: (2023)
di: Ma, Chong, et al.
Pubblicazione: (2023)
The Knesset Corpus: An Annotated Corpus of Hebrew Parliamentary Proceedings
di: Goldin, Gili, et al.
Pubblicazione: (2024)
di: Goldin, Gili, et al.
Pubblicazione: (2024)
Math Natural Language Inference: this should be easy!
di: de Paiva, Valeria, et al.
Pubblicazione: (2025)
di: de Paiva, Valeria, et al.
Pubblicazione: (2025)
New Skills or Sharper Primitives? A Probabilistic Perspective on the Emergence of Reasoning in RLVR
di: Wang, Zhilin, et al.
Pubblicazione: (2026)
di: Wang, Zhilin, et al.
Pubblicazione: (2026)
Pitfalls in Evaluating Interpretability Agents
di: Haklay, Tal, et al.
Pubblicazione: (2026)
di: Haklay, Tal, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Fast Quiet-STaR: Thinking Without Thought Tokens
di: Huang, Wei, et al.
Pubblicazione: (2025) -
D-SMART: Enhancing LLM Dialogue Consistency via Dynamic Structured Memory And Reasoning Tree
di: Lei, Xiang, et al.
Pubblicazione: (2025) -
Train-Attention: Meta-Learning Where to Focus in Continual Knowledge Learning
di: Seo, Yeongbin, et al.
Pubblicazione: (2024) -
WeDLM: Reconciling Diffusion Language Models with Standard Causal Attention for Fast Inference
di: Liu, Aiwei, et al.
Pubblicazione: (2025) -
When Does Content-Based Routing Work? Representation Requirements for Selective Attention in Hybrid Sequence Models
di: Basu, Abhinaba
Pubblicazione: (2026)