Salvato in:
| Autori principali: | Choi, DongHyun, Spangher, Lucas, Hidey, Chris, Grabowski, Peter, Eskander, Ramy |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2504.02877 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Cross-Architecture Transfer Learning for Linear-Cost Inference Transformers
di: Choi, Sehyun
Pubblicazione: (2024)
di: Choi, Sehyun
Pubblicazione: (2024)
Factored Agents: Decoupling In-Context Learning and Memorization for Robust Tool Use
di: Roth, Nicholas, et al.
Pubblicazione: (2025)
di: Roth, Nicholas, et al.
Pubblicazione: (2025)
PatentEdits: Framing Patent Novelty as Textual Entailment
di: Lee, Ryan, et al.
Pubblicazione: (2024)
di: Lee, Ryan, et al.
Pubblicazione: (2024)
Residual Stream Duality in Modern Transformer Architectures
di: Zhang, Yifan
Pubblicazione: (2026)
di: Zhang, Yifan
Pubblicazione: (2026)
LLM Cache Bandit Revisited: Addressing Query Heterogeneity for Cost-Effective LLM Inference
di: Yang, Hantao, et al.
Pubblicazione: (2025)
di: Yang, Hantao, et al.
Pubblicazione: (2025)
Understanding LLMs: A Comprehensive Overview from Training to Inference
di: Liu, Yiheng, et al.
Pubblicazione: (2024)
di: Liu, Yiheng, et al.
Pubblicazione: (2024)
Peri-LN: Revisiting Normalization Layer in the Transformer Architecture
di: Kim, Jeonghoon, et al.
Pubblicazione: (2025)
di: Kim, Jeonghoon, et al.
Pubblicazione: (2025)
Learning Action Conditions from Instructional Manuals for Instruction Understanding
di: Wu, Te-Lin, et al.
Pubblicazione: (2022)
di: Wu, Te-Lin, et al.
Pubblicazione: (2022)
Revisiting Padded Transformer Expressivity: Which Architectural Choices Matter and Which Don't
di: Svete, Anej, et al.
Pubblicazione: (2026)
di: Svete, Anej, et al.
Pubblicazione: (2026)
PyramidKV: Dynamic KV Cache Compression based on Pyramidal Information Funneling
di: Cai, Zefan, et al.
Pubblicazione: (2024)
di: Cai, Zefan, et al.
Pubblicazione: (2024)
DiscoSum: Discourse-aware News Summarization
di: Spangher, Alexander, et al.
Pubblicazione: (2025)
di: Spangher, Alexander, et al.
Pubblicazione: (2025)
Improving LLM-as-a-Judge Inference with the Judgment Distribution
di: Wang, Victor, et al.
Pubblicazione: (2025)
di: Wang, Victor, et al.
Pubblicazione: (2025)
NewsEdits 2.0: Learning the Intentions Behind Updating News
di: Spangher, Alexander, et al.
Pubblicazione: (2024)
di: Spangher, Alexander, et al.
Pubblicazione: (2024)
Applications of the Transformer Architecture in AI-Assisted English Reading Comprehension
di: Li, Ping
Pubblicazione: (2026)
di: Li, Ping
Pubblicazione: (2026)
Transforming Slot Schema Induction with Generative Dialogue State Inference
di: Finch, James D., et al.
Pubblicazione: (2024)
di: Finch, James D., et al.
Pubblicazione: (2024)
Extra Global Attention Designation Using Keyword Detection in Sparse Transformer Architectures
di: Lucas, Evan, et al.
Pubblicazione: (2024)
di: Lucas, Evan, et al.
Pubblicazione: (2024)
Fast KVzip: Efficient and Accurate LLM Inference with Gated KV Eviction
di: Kim, Jang-Hyun, et al.
Pubblicazione: (2026)
di: Kim, Jang-Hyun, et al.
Pubblicazione: (2026)
Mirror Speculative Decoding: Breaking the Serial Barrier in LLM Inference
di: Bhendawade, Nikhil, et al.
Pubblicazione: (2025)
di: Bhendawade, Nikhil, et al.
Pubblicazione: (2025)
NewsHomepages: Homepage Layouts Capture Information Prioritization Decisions
di: Welsh, Ben, et al.
Pubblicazione: (2024)
di: Welsh, Ben, et al.
Pubblicazione: (2024)
DiffAdapt: Difficulty-Adaptive Reasoning for Token-Efficient LLM Inference
di: Liu, Xiang, et al.
Pubblicazione: (2025)
di: Liu, Xiang, et al.
Pubblicazione: (2025)
ModernBERT or DeBERTaV3? Examining Architecture and Data Influence on Transformer Encoder Models Performance
di: Antoun, Wissam, et al.
Pubblicazione: (2025)
di: Antoun, Wissam, et al.
Pubblicazione: (2025)
A Survey on LLM Inference-Time Self-Improvement
di: Dong, Xiangjue, et al.
Pubblicazione: (2024)
di: Dong, Xiangjue, et al.
Pubblicazione: (2024)
FunnelRAG: A Coarse-to-Fine Progressive Retrieval Paradigm for RAG
di: Zhao, Xinping, et al.
Pubblicazione: (2024)
di: Zhao, Xinping, et al.
Pubblicazione: (2024)
Transparent Screening for LLM Inference and Training Impacts
di: Pachot, Arnault, et al.
Pubblicazione: (2026)
di: Pachot, Arnault, et al.
Pubblicazione: (2026)
Understanding Multimodal Procedural Knowledge by Sequencing Multimodal Instructional Manuals
di: Wu, Te-Lin, et al.
Pubblicazione: (2021)
di: Wu, Te-Lin, et al.
Pubblicazione: (2021)
ArcLight: A Lightweight LLM Inference Architecture for Many-Core CPUs
di: Xu, Yuzhuang, et al.
Pubblicazione: (2026)
di: Xu, Yuzhuang, et al.
Pubblicazione: (2026)
Generalized Probabilistic Attention Mechanism in Transformers
di: Heo, DongNyeong, et al.
Pubblicazione: (2024)
di: Heo, DongNyeong, et al.
Pubblicazione: (2024)
RoBIn: A Transformer-Based Model For Risk Of Bias Inference With Machine Reading Comprehension
di: Dias, Abel Corrêa, et al.
Pubblicazione: (2024)
di: Dias, Abel Corrêa, et al.
Pubblicazione: (2024)
TransformLLM: Adapting Large Language Models via LLM-Transformed Reading Comprehension Text
di: Arbel, Iftach, et al.
Pubblicazione: (2024)
di: Arbel, Iftach, et al.
Pubblicazione: (2024)
Explaining Mixtures of Sources in News Articles
di: Spangher, Alexander, et al.
Pubblicazione: (2024)
di: Spangher, Alexander, et al.
Pubblicazione: (2024)
Folding Tensor and Sequence Parallelism for Memory-Efficient Transformer Training & Inference
di: Shyam, Vasu, et al.
Pubblicazione: (2026)
di: Shyam, Vasu, et al.
Pubblicazione: (2026)
Inferflow: an Efficient and Highly Configurable Inference Engine for Large Language Models
di: Shi, Shuming, et al.
Pubblicazione: (2024)
di: Shi, Shuming, et al.
Pubblicazione: (2024)
Horizon-LM: A RAM-Centric Architecture for LLM Training
di: Yuan, Zhengqing, et al.
Pubblicazione: (2026)
di: Yuan, Zhengqing, et al.
Pubblicazione: (2026)
Latent Recurrent Transformer: Architecture Exploration, Training Strategies, and Scaling Behavior
di: Huang, Zeyi, et al.
Pubblicazione: (2026)
di: Huang, Zeyi, et al.
Pubblicazione: (2026)
End-to-End Training for Back-Translation with Categorical Reparameterization Trick
di: Heo, DongNyeong, et al.
Pubblicazione: (2022)
di: Heo, DongNyeong, et al.
Pubblicazione: (2022)
Diagnosing Training Inference Mismatch in LLM Reinforcement Learning
di: Zhong, Tianle, et al.
Pubblicazione: (2026)
di: Zhong, Tianle, et al.
Pubblicazione: (2026)
Revisiting Word Embeddings in the LLM Era
di: Mahajan, Yash, et al.
Pubblicazione: (2025)
di: Mahajan, Yash, et al.
Pubblicazione: (2025)
Are Large Language Models Capable of Generating Human-Level Narratives?
di: Tian, Yufei, et al.
Pubblicazione: (2024)
di: Tian, Yufei, et al.
Pubblicazione: (2024)
Autoregressive Transformers for Disruption Prediction in Nuclear Fusion Plasmas
di: Spangher, Lucas, et al.
Pubblicazione: (2023)
di: Spangher, Lucas, et al.
Pubblicazione: (2023)
Revisiting Hierarchical Text Classification: Inference and Metrics
di: Plaud, Roman, et al.
Pubblicazione: (2024)
di: Plaud, Roman, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Cross-Architecture Transfer Learning for Linear-Cost Inference Transformers
di: Choi, Sehyun
Pubblicazione: (2024) -
Factored Agents: Decoupling In-Context Learning and Memorization for Robust Tool Use
di: Roth, Nicholas, et al.
Pubblicazione: (2025) -
PatentEdits: Framing Patent Novelty as Textual Entailment
di: Lee, Ryan, et al.
Pubblicazione: (2024) -
Residual Stream Duality in Modern Transformer Architectures
di: Zhang, Yifan
Pubblicazione: (2026) -
LLM Cache Bandit Revisited: Addressing Query Heterogeneity for Cost-Effective LLM Inference
di: Yang, Hantao, et al.
Pubblicazione: (2025)