How Language Models Process Out-of-Distribution Inputs: A Two-Pathway Framework
Fuente:
arXiv
Guardado en:
| Autor principal: | Saghir, Hamidreza |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
por: Fadli, Samih
Publicado: (2025)
por: Fadli, Samih
Publicado: (2025)
Aspect-Based Sentiment Analysis for Future Tourism Experiences: A BERT-MoE Framework for Persian User Reviews
por: Taskooh, Hamidreza Kazemi, et al.
Publicado: (2026)
por: Taskooh, Hamidreza Kazemi, et al.
Publicado: (2026)
Ouroboros: Dynamic Weight Generation for Recursive Transformers via Input-Conditioned LoRA Modulation
por: Jaber, Jaber, et al.
Publicado: (2026)
por: Jaber, Jaber, et al.
Publicado: (2026)
Language Models Are Implicitly Continuous
por: Marro, Samuele, et al.
Publicado: (2025)
por: Marro, Samuele, et al.
Publicado: (2025)
Pre-trained Models Perform the Best When Token Distributions Follow Zipf's Law
por: He, Yanjin, et al.
Publicado: (2025)
por: He, Yanjin, et al.
Publicado: (2025)
Combining Language and Topic Models for Hierarchical Text Classification
por: Toit, Jaco du, et al.
Publicado: (2025)
por: Toit, Jaco du, et al.
Publicado: (2025)
Memory Bank Compression for Continual Adaptation of Large Language Models
por: Katraouras, Thomas, et al.
Publicado: (2026)
por: Katraouras, Thomas, et al.
Publicado: (2026)
Neural Activation Patterns Across Language Model Architectures: A Comprehensive Analysis of Cognitive Task Performance
por: Naser-Moghadasi, Mahdi, et al.
Publicado: (2026)
por: Naser-Moghadasi, Mahdi, et al.
Publicado: (2026)
OCRR: A Benchmark for Online Correction Recovery under Distribution Shift
por: Grassi, Adrian
Publicado: (2026)
por: Grassi, Adrian
Publicado: (2026)
Rethinking Addressing in Language Models via Contexualized Equivariant Positional Encoding
por: Zhu, Jiajun, et al.
Publicado: (2025)
por: Zhu, Jiajun, et al.
Publicado: (2025)
Your Pretrained Model Tells the Difficulty Itself: A Self-Adaptive Curriculum Learning Paradigm for Natural Language Understanding
por: Feng, Qi, et al.
Publicado: (2025)
por: Feng, Qi, et al.
Publicado: (2025)
Kronecker Embeddings: Byte-Level Structured Token Representations for Parameter-Efficient Language Models
por: Shravan, Rohan
Publicado: (2026)
por: Shravan, Rohan
Publicado: (2026)
The Right Answer, the Wrong Direction: Why Transformers Fail at Counting and How to Fix It
por: Garcia, Gabriel
Publicado: (2026)
por: Garcia, Gabriel
Publicado: (2026)
Future Token Prediction -- Causal Language Modelling with Per-Token Semantic State Vector for Multi-Token Prediction
por: Walker, Nicholas
Publicado: (2024)
por: Walker, Nicholas
Publicado: (2024)
Bayesian Attention Mechanism: A Probabilistic Framework for Positional Encoding and Context Length Extrapolation
por: Bianchessi, Arthur S., et al.
Publicado: (2025)
por: Bianchessi, Arthur S., et al.
Publicado: (2025)
Discovering Transformer Circuits via a Hybrid Attribution and Pruning Framework
por: Gu, Hao, et al.
Publicado: (2025)
por: Gu, Hao, et al.
Publicado: (2025)
Sarcasm Detection in a Less-Resourced Language
por: Đoković, Lazar, et al.
Publicado: (2024)
por: Đoković, Lazar, et al.
Publicado: (2024)
When Models Can't Follow: Testing Instruction Adherence Across 256 LLMs
por: Young, Richard J., et al.
Publicado: (2025)
por: Young, Richard J., et al.
Publicado: (2025)
The Data Efficiency Frontier of Financial Foundation Models: Scaling Laws from Continued Pretraining
por: Ponnock, Jesse
Publicado: (2025)
por: Ponnock, Jesse
Publicado: (2025)
Stratified Hazard Sampling: Minimal-Variance Event Scheduling for CTMC/DTMC Discrete Diffusion and Flow Models
por: Jang, Seunghwan, et al.
Publicado: (2026)
por: Jang, Seunghwan, et al.
Publicado: (2026)
Advancing Transformer Architecture in Long-Context Large Language Models: A Comprehensive Survey
por: Huang, Yunpeng, et al.
Publicado: (2023)
por: Huang, Yunpeng, et al.
Publicado: (2023)
MCP: A Control-Theoretic Orchestration Framework for Synergistic Efficiency and Interpretability in Multimodal Large Language Models
por: Zhang, Luyan
Publicado: (2025)
por: Zhang, Luyan
Publicado: (2025)
Are LLM Uncertainty and Correctness Encoded by the Same Features? A Functional Dissociation via Sparse Autoencoders
por: Patel, Het, et al.
Publicado: (2026)
por: Patel, Het, et al.
Publicado: (2026)
HyDRA: Hybrid Dynamic Routing Architecture for Heterogeneous LLM Pools
por: Garg, Aashna, et al.
Publicado: (2026)
por: Garg, Aashna, et al.
Publicado: (2026)
Quantization-Robust LLM Unlearning via Low-Rank Adaptation
por: Abitante, João Vitor Boer, et al.
Publicado: (2026)
por: Abitante, João Vitor Boer, et al.
Publicado: (2026)
Annotation Entropy Predicts Per-Example Learning Dynamics in LoRA Fine-Tuning
por: Steele, Brady
Publicado: (2026)
por: Steele, Brady
Publicado: (2026)
Self-Pruned Key-Value Attention: Learning When to Write by Predicting Future Utility
por: Szilvasy, Gergely, et al.
Publicado: (2026)
por: Szilvasy, Gergely, et al.
Publicado: (2026)
Representation-Aware Unlearning via Activation Signatures: From Suppression to Entity-Signature Erasure
por: Mahmood, Syed Naveed, et al.
Publicado: (2026)
por: Mahmood, Syed Naveed, et al.
Publicado: (2026)
Continuous-Depth Transformers with Learned Control Dynamics
por: Jemley, Peter
Publicado: (2026)
por: Jemley, Peter
Publicado: (2026)
TIME: Temporally Intelligent Meta-reasoning Engine for Context-Triggered Explicit Reasoning
por: Das, Susmit
Publicado: (2026)
por: Das, Susmit
Publicado: (2026)
Mean-Pooled Cosine Similarity is Not Length-Invariant: Theory and Cross-Domain Evidence for a Length-Invariant Alternative
por: Mitra, Sibayan, et al.
Publicado: (2026)
por: Mitra, Sibayan, et al.
Publicado: (2026)
PersonalLLM: Tailoring LLMs to Individual Preferences
por: Zollo, Thomas P., et al.
Publicado: (2024)
por: Zollo, Thomas P., et al.
Publicado: (2024)
Thread Detection and Response Generation using Transformers with Prompt Optimisation
por: T, Kevin Joshua, et al.
Publicado: (2024)
por: T, Kevin Joshua, et al.
Publicado: (2024)
LLM Vocabulary Compression for Low-Compute Environments
por: Vennam, Sreeram, et al.
Publicado: (2024)
por: Vennam, Sreeram, et al.
Publicado: (2024)
Efficient Strategy for Improving Large Language Model (LLM) Capabilities
por: Gutiérrez, Julián Camilo Velandia
Publicado: (2025)
por: Gutiérrez, Julián Camilo Velandia
Publicado: (2025)
Merge-Bench: Resolve Merge Conflicts with Large Language Models
por: Schesch, Benedikt, et al.
Publicado: (2026)
por: Schesch, Benedikt, et al.
Publicado: (2026)
Towards Alignment-Centric Paradigm: A Survey of Instruction Tuning in Large Language Models
por: Han, Xudong, et al.
Publicado: (2025)
por: Han, Xudong, et al.
Publicado: (2025)
Domain-Specific Pretraining of Language Models: A Comparative Study in the Medical Field
por: Kerner, Tobias
Publicado: (2024)
por: Kerner, Tobias
Publicado: (2024)
How Human-Like Are Large Language Models? A Register-Aware Linguistic Evaluation Framework
por: Nieth, Björn, et al.
Publicado: (2026)
por: Nieth, Björn, et al.
Publicado: (2026)
Towards Resource-Efficient Multimodal Intelligence: Learned Routing among Specialized Expert Models
por: Saini, Mayank, et al.
Publicado: (2025)
por: Saini, Mayank, et al.
Publicado: (2025)
Ejemplares similares
-
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
por: Fadli, Samih
Publicado: (2025) -
Aspect-Based Sentiment Analysis for Future Tourism Experiences: A BERT-MoE Framework for Persian User Reviews
por: Taskooh, Hamidreza Kazemi, et al.
Publicado: (2026) -
Ouroboros: Dynamic Weight Generation for Recursive Transformers via Input-Conditioned LoRA Modulation
por: Jaber, Jaber, et al.
Publicado: (2026) -
Language Models Are Implicitly Continuous
por: Marro, Samuele, et al.
Publicado: (2025) -
Pre-trained Models Perform the Best When Token Distributions Follow Zipf's Law
por: He, Yanjin, et al.
Publicado: (2025)