Neural Activation Patterns Across Language Model Architectures: A Comprehensive Analysis of Cognitive Task Performance
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Naser-Moghadasi, Mahdi, Ghaderi, Faezeh |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Transformer Scalability Crisis: The First Comprehensive Empirical Analysis of Performance Walls in Modern Language Models
von: Moghadasi, Mahdi Naser, et al.
Veröffentlicht: (2026)
von: Moghadasi, Mahdi Naser, et al.
Veröffentlicht: (2026)
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
von: Fadli, Samih
Veröffentlicht: (2025)
von: Fadli, Samih
Veröffentlicht: (2025)
Advancing Transformer Architecture in Long-Context Large Language Models: A Comprehensive Survey
von: Huang, Yunpeng, et al.
Veröffentlicht: (2023)
von: Huang, Yunpeng, et al.
Veröffentlicht: (2023)
Variance Is Not Importance: Structural Analysis of Transformer Compressibility Across Model Scales
von: Salfati, Samuel
Veröffentlicht: (2026)
von: Salfati, Samuel
Veröffentlicht: (2026)
When Models Can't Follow: Testing Instruction Adherence Across 256 LLMs
von: Young, Richard J., et al.
Veröffentlicht: (2025)
von: Young, Richard J., et al.
Veröffentlicht: (2025)
Bridging the Gap: An Intermediate Language for Enhanced and Cost-Effective Grapheme-to-Phoneme Conversion with Homographs with Multiple Pronunciations Disambiguation
von: Bertina, Abbas, et al.
Veröffentlicht: (2025)
von: Bertina, Abbas, et al.
Veröffentlicht: (2025)
Emergent Lexical Semantics in Neural Language Models: Testing Martin's Law on LLM-Generated Text
von: Kugler, Kai
Veröffentlicht: (2025)
von: Kugler, Kai
Veröffentlicht: (2025)
Language Models Are Implicitly Continuous
von: Marro, Samuele, et al.
Veröffentlicht: (2025)
von: Marro, Samuele, et al.
Veröffentlicht: (2025)
Mixup Model Merge: Enhancing Model Merging Performance through Randomized Linear Interpolation
von: Zhou, Yue, et al.
Veröffentlicht: (2025)
von: Zhou, Yue, et al.
Veröffentlicht: (2025)
HyDRA: Hybrid Dynamic Routing Architecture for Heterogeneous LLM Pools
von: Garg, Aashna, et al.
Veröffentlicht: (2026)
von: Garg, Aashna, et al.
Veröffentlicht: (2026)
Pre-trained Models Perform the Best When Token Distributions Follow Zipf's Law
von: He, Yanjin, et al.
Veröffentlicht: (2025)
von: He, Yanjin, et al.
Veröffentlicht: (2025)
Combining Language and Topic Models for Hierarchical Text Classification
von: Toit, Jaco du, et al.
Veröffentlicht: (2025)
von: Toit, Jaco du, et al.
Veröffentlicht: (2025)
Representation-Aware Unlearning via Activation Signatures: From Suppression to Entity-Signature Erasure
von: Mahmood, Syed Naveed, et al.
Veröffentlicht: (2026)
von: Mahmood, Syed Naveed, et al.
Veröffentlicht: (2026)
Memory Bank Compression for Continual Adaptation of Large Language Models
von: Katraouras, Thomas, et al.
Veröffentlicht: (2026)
von: Katraouras, Thomas, et al.
Veröffentlicht: (2026)
Induce, Align, Predict: Zero-Shot Stance Detection via Cognitive Inductive Reasoning
von: Zhang, Bowen, et al.
Veröffentlicht: (2025)
von: Zhang, Bowen, et al.
Veröffentlicht: (2025)
Rethinking Addressing in Language Models via Contexualized Equivariant Positional Encoding
von: Zhu, Jiajun, et al.
Veröffentlicht: (2025)
von: Zhu, Jiajun, et al.
Veröffentlicht: (2025)
Kronecker Embeddings: Byte-Level Structured Token Representations for Parameter-Efficient Language Models
von: Shravan, Rohan
Veröffentlicht: (2026)
von: Shravan, Rohan
Veröffentlicht: (2026)
How Language Models Process Out-of-Distribution Inputs: A Two-Pathway Framework
von: Saghir, Hamidreza
Veröffentlicht: (2026)
von: Saghir, Hamidreza
Veröffentlicht: (2026)
Merge-Bench: Resolve Merge Conflicts with Large Language Models
von: Schesch, Benedikt, et al.
Veröffentlicht: (2026)
von: Schesch, Benedikt, et al.
Veröffentlicht: (2026)
Towards Alignment-Centric Paradigm: A Survey of Instruction Tuning in Large Language Models
von: Han, Xudong, et al.
Veröffentlicht: (2025)
von: Han, Xudong, et al.
Veröffentlicht: (2025)
D-COT: Disciplined Chain-of-Thought Learning for Efficient Reasoning in Small Language Models
von: Ubukata, Shunsuke
Veröffentlicht: (2026)
von: Ubukata, Shunsuke
Veröffentlicht: (2026)
Future Token Prediction -- Causal Language Modelling with Per-Token Semantic State Vector for Multi-Token Prediction
von: Walker, Nicholas
Veröffentlicht: (2024)
von: Walker, Nicholas
Veröffentlicht: (2024)
MCP: A Control-Theoretic Orchestration Framework for Synergistic Efficiency and Interpretability in Multimodal Large Language Models
von: Zhang, Luyan
Veröffentlicht: (2025)
von: Zhang, Luyan
Veröffentlicht: (2025)
Your Pretrained Model Tells the Difficulty Itself: A Self-Adaptive Curriculum Learning Paradigm for Natural Language Understanding
von: Feng, Qi, et al.
Veröffentlicht: (2025)
von: Feng, Qi, et al.
Veröffentlicht: (2025)
Prototype Transformer: Towards Language Model Architectures Interpretable by Design
von: Yordanov, Yordan, et al.
Veröffentlicht: (2026)
von: Yordanov, Yordan, et al.
Veröffentlicht: (2026)
Turning the TIDE: Cross-Architecture Distillation for Diffusion Large Language Models
von: Zhang, Gongbo, et al.
Veröffentlicht: (2026)
von: Zhang, Gongbo, et al.
Veröffentlicht: (2026)
Cognitive Load Limits in Large Language Models: Benchmarking Multi-Hop Reasoning
von: Adapala, Sai Teja Reddy
Veröffentlicht: (2025)
von: Adapala, Sai Teja Reddy
Veröffentlicht: (2025)
Sarcasm Detection in a Less-Resourced Language
von: Đoković, Lazar, et al.
Veröffentlicht: (2024)
von: Đoković, Lazar, et al.
Veröffentlicht: (2024)
Painless Activation Steering: An Automated, Lightweight Approach for Post-Training Large Language Models
von: Cui, Sasha, et al.
Veröffentlicht: (2025)
von: Cui, Sasha, et al.
Veröffentlicht: (2025)
ACE: Exploring Activation Cosine Similarity and Variance for Accurate and Calibration-Efficient LLM Pruning
von: Mi, Zhendong, et al.
Veröffentlicht: (2025)
von: Mi, Zhendong, et al.
Veröffentlicht: (2025)
Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring
von: Heyman, Alex, et al.
Veröffentlicht: (2025)
von: Heyman, Alex, et al.
Veröffentlicht: (2025)
The Data Efficiency Frontier of Financial Foundation Models: Scaling Laws from Continued Pretraining
von: Ponnock, Jesse
Veröffentlicht: (2025)
von: Ponnock, Jesse
Veröffentlicht: (2025)
Stratified Hazard Sampling: Minimal-Variance Event Scheduling for CTMC/DTMC Discrete Diffusion and Flow Models
von: Jang, Seunghwan, et al.
Veröffentlicht: (2026)
von: Jang, Seunghwan, et al.
Veröffentlicht: (2026)
KSHSeek: Data-Driven Approaches to Mitigating and Detecting Knowledge-Shortcut Hallucinations in Generative Models
von: Liu, Zhongxin, et al.
Veröffentlicht: (2025)
von: Liu, Zhongxin, et al.
Veröffentlicht: (2025)
EmoLoom-2B: Fast Base-Model Screening for Emotion Classification and VAD with Lexicon-Weak Supervision and KV-Off Evaluation
von: Li, Zilin, et al.
Veröffentlicht: (2026)
von: Li, Zilin, et al.
Veröffentlicht: (2026)
Transformers Boost the Performance of Decision Trees on Tabular Data across Sample Sizes
von: Jayawardhana, Mayuka, et al.
Veröffentlicht: (2025)
von: Jayawardhana, Mayuka, et al.
Veröffentlicht: (2025)
Adaptive Activation Cancellation for Hallucination Mitigation in Large Language Models
von: Yocam, Eric, et al.
Veröffentlicht: (2026)
von: Yocam, Eric, et al.
Veröffentlicht: (2026)
Efficient Strategy for Improving Large Language Model (LLM) Capabilities
von: Gutiérrez, Julián Camilo Velandia
Veröffentlicht: (2025)
von: Gutiérrez, Julián Camilo Velandia
Veröffentlicht: (2025)
TensorLens: End-to-End Transformer Analysis via High-Order Attention Tensors
von: Atad, Ido Andrew, et al.
Veröffentlicht: (2026)
von: Atad, Ido Andrew, et al.
Veröffentlicht: (2026)
Discovering Transformer Circuits via a Hybrid Attribution and Pruning Framework
von: Gu, Hao, et al.
Veröffentlicht: (2025)
von: Gu, Hao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Transformer Scalability Crisis: The First Comprehensive Empirical Analysis of Performance Walls in Modern Language Models
von: Moghadasi, Mahdi Naser, et al.
Veröffentlicht: (2026) -
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
von: Fadli, Samih
Veröffentlicht: (2025) -
Advancing Transformer Architecture in Long-Context Large Language Models: A Comprehensive Survey
von: Huang, Yunpeng, et al.
Veröffentlicht: (2023) -
Variance Is Not Importance: Structural Analysis of Transformer Compressibility Across Model Scales
von: Salfati, Samuel
Veröffentlicht: (2026) -
When Models Can't Follow: Testing Instruction Adherence Across 256 LLMs
von: Young, Richard J., et al.
Veröffentlicht: (2025)