From Attention to Activation: Unravelling the Enigmas of Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kaul, Prannay, Ma, Chengcheng, Elezi, Ismail, Deng, Jiankang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Test-Time Scaling with Diffusion Language Models via Reward-Guided Stitching
von: Miles, Roy, et al.
Veröffentlicht: (2026)
von: Miles, Roy, et al.
Veröffentlicht: (2026)
Three Heads Are Better Than One: Complementary Experts for Long-Tailed Semi-supervised Learning
von: Ma, Chengcheng, et al.
Veröffentlicht: (2023)
von: Ma, Chengcheng, et al.
Veröffentlicht: (2023)
G3DR: Generative 3D Reconstruction in ImageNet
von: Reddy, Pradyumna, et al.
Veröffentlicht: (2024)
von: Reddy, Pradyumna, et al.
Veröffentlicht: (2024)
$V_kD:$ Improving Knowledge Distillation using Orthogonal Projections
von: Miles, Roy, et al.
Veröffentlicht: (2024)
von: Miles, Roy, et al.
Veröffentlicht: (2024)
Deep Active Learning: A Reality Check
von: Gashi, Edrina, et al.
Veröffentlicht: (2024)
von: Gashi, Edrina, et al.
Veröffentlicht: (2024)
"Principal Components" Enable A New Language of Images
von: Wen, Xin, et al.
Veröffentlicht: (2025)
von: Wen, Xin, et al.
Veröffentlicht: (2025)
VeLoRA: Memory Efficient Training using Rank-1 Sub-Token Projections
von: Miles, Roy, et al.
Veröffentlicht: (2024)
von: Miles, Roy, et al.
Veröffentlicht: (2024)
RetouchLLM: Training-free Code-based Image Retouching with Vision Language Models
von: Ye-Bin, Moon, et al.
Veröffentlicht: (2025)
von: Ye-Bin, Moon, et al.
Veröffentlicht: (2025)
CASteer: Cross-Attention Steering for Controllable Concept Erasure
von: Gaintseva, Tatiana, et al.
Veröffentlicht: (2025)
von: Gaintseva, Tatiana, et al.
Veröffentlicht: (2025)
Fractal Calibration for long-tailed object detection
von: Alexandridis, Konstantinos Panagiotis, et al.
Veröffentlicht: (2024)
von: Alexandridis, Konstantinos Panagiotis, et al.
Veröffentlicht: (2024)
SATGround: A Spatially-Aware Approach for Visual Grounding in Remote Sensing
von: Toker, Aysim, et al.
Veröffentlicht: (2025)
von: Toker, Aysim, et al.
Veröffentlicht: (2025)
Stable and Explainable Personality Trait Evaluation in Large Language Models with Internal Activations
von: Ma, Xiaoxu, et al.
Veröffentlicht: (2026)
von: Ma, Xiaoxu, et al.
Veröffentlicht: (2026)
Prefixing Attention Sinks can Mitigate Activation Outliers for Large Language Model Quantization
von: Son, Seungwoo, et al.
Veröffentlicht: (2024)
von: Son, Seungwoo, et al.
Veröffentlicht: (2024)
Do You See What I Am Pointing At? Gesture-Based Egocentric Video Question Answering
von: Choi, Yura, et al.
Veröffentlicht: (2026)
von: Choi, Yura, et al.
Veröffentlicht: (2026)
Hybrid EEG--Driven Brain--Computer Interface: A Large Language Model Framework for Personalized Language Rehabilitation
von: Hossain, Ismail, et al.
Veröffentlicht: (2025)
von: Hossain, Ismail, et al.
Veröffentlicht: (2025)
An Enigma of Artificial Reason: Investigating the Production-Evaluation Gap in Large Reasoning Models
von: Sun, Mingzhong, et al.
Veröffentlicht: (2026)
von: Sun, Mingzhong, et al.
Veröffentlicht: (2026)
First Activations Matter: Training-Free Methods for Dynamic Activation in Large Language Models
von: Ma, Chi, et al.
Veröffentlicht: (2024)
von: Ma, Chi, et al.
Veröffentlicht: (2024)
Unraveling Arithmetic in Large Language Models: The Role of Algebraic Structures
von: Chang, Fu-Chieh, et al.
Veröffentlicht: (2024)
von: Chang, Fu-Chieh, et al.
Veröffentlicht: (2024)
Massive Activations in Large Language Models
von: Sun, Mingjie, et al.
Veröffentlicht: (2024)
von: Sun, Mingjie, et al.
Veröffentlicht: (2024)
Measuring Maximum Activations in Open Large Language Models
von: Chen, Luxuan, et al.
Veröffentlicht: (2026)
von: Chen, Luxuan, et al.
Veröffentlicht: (2026)
Activation-Guided Consensus Merging for Large Language Models
von: Yao, Yuxuan, et al.
Veröffentlicht: (2025)
von: Yao, Yuxuan, et al.
Veröffentlicht: (2025)
Why do Large Language Models Fail in Low-resource Translation? Unraveling the Token Dynamics of Large Language Models for Machine Translation
von: Qian, Shenbin, et al.
Veröffentlicht: (2026)
von: Qian, Shenbin, et al.
Veröffentlicht: (2026)
Learning to Check: Unleashing Potentials for Self-Correction in Large Language Models
von: Zhang, Che, et al.
Veröffentlicht: (2024)
von: Zhang, Che, et al.
Veröffentlicht: (2024)
Formality is Favored: Unraveling the Learning Preferences of Large Language Models on Data with Conflicting Knowledge
von: Li, Jiahuan, et al.
Veröffentlicht: (2024)
von: Li, Jiahuan, et al.
Veröffentlicht: (2024)
Unraveling Interwoven Roles of Large Language Models in Authorship Privacy: Obfuscation, Mimicking, and Verification
von: Nguyen, Tuc, et al.
Veröffentlicht: (2025)
von: Nguyen, Tuc, et al.
Veröffentlicht: (2025)
Q-Sparse: All Large Language Models can be Fully Sparsely-Activated
von: Wang, Hongyu, et al.
Veröffentlicht: (2024)
von: Wang, Hongyu, et al.
Veröffentlicht: (2024)
Unraveling Cross-Modality Knowledge Conflicts in Large Vision-Language Models
von: Zhu, Tinghui, et al.
Veröffentlicht: (2024)
von: Zhu, Tinghui, et al.
Veröffentlicht: (2024)
Head-wise Shareable Attention for Large Language Models
von: Cao, Zouying, et al.
Veröffentlicht: (2024)
von: Cao, Zouying, et al.
Veröffentlicht: (2024)
Attention Heads of Large Language Models: A Survey
von: Zheng, Zifan, et al.
Veröffentlicht: (2024)
von: Zheng, Zifan, et al.
Veröffentlicht: (2024)
PocketLLM: Ultimate Compression of Large Language Models via Meta Networks
von: Tian, Ye, et al.
Veröffentlicht: (2025)
von: Tian, Ye, et al.
Veröffentlicht: (2025)
Unraveling the Dominance of Large Language Models Over Transformer Models for Bangla Natural Language Inference: A Comprehensive Study
von: Faria, Fatema Tuj Johora, et al.
Veröffentlicht: (2024)
von: Faria, Fatema Tuj Johora, et al.
Veröffentlicht: (2024)
Unraveling and Mitigating Retriever Inconsistencies in Retrieval-Augmented Large Language Models
von: Li, Mingda, et al.
Veröffentlicht: (2024)
von: Li, Mingda, et al.
Veröffentlicht: (2024)
Controlling Large Language Model Agents with Entropic Activation Steering
von: Rahn, Nate, et al.
Veröffentlicht: (2024)
von: Rahn, Nate, et al.
Veröffentlicht: (2024)
Controlling Large Language Models Through Concept Activation Vectors
von: Zhang, Hanyu, et al.
Veröffentlicht: (2025)
von: Zhang, Hanyu, et al.
Veröffentlicht: (2025)
Riddle Quest : The Enigma of Words
von: Parasa, Niharika Sri, et al.
Veröffentlicht: (2026)
von: Parasa, Niharika Sri, et al.
Veröffentlicht: (2026)
Activation-Informed Merging of Large Language Models
von: Nobari, Amin Heyrani, et al.
Veröffentlicht: (2025)
von: Nobari, Amin Heyrani, et al.
Veröffentlicht: (2025)
Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
von: Qiu, Zihan, et al.
Veröffentlicht: (2025)
von: Qiu, Zihan, et al.
Veröffentlicht: (2025)
PoLLMgraph: Unraveling Hallucinations in Large Language Models via State Transition Dynamics
von: Zhu, Derui, et al.
Veröffentlicht: (2024)
von: Zhu, Derui, et al.
Veröffentlicht: (2024)
Signs as Tokens: A Retrieval-Enhanced Multilingual Sign Language Generator
von: Zuo, Ronglai, et al.
Veröffentlicht: (2024)
von: Zuo, Ronglai, et al.
Veröffentlicht: (2024)
Polynomial Composition Activations: Unleashing the Dynamics of Large Language Models
von: Zhuo, Zhijian, et al.
Veröffentlicht: (2024)
von: Zhuo, Zhijian, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Test-Time Scaling with Diffusion Language Models via Reward-Guided Stitching
von: Miles, Roy, et al.
Veröffentlicht: (2026) -
Three Heads Are Better Than One: Complementary Experts for Long-Tailed Semi-supervised Learning
von: Ma, Chengcheng, et al.
Veröffentlicht: (2023) -
G3DR: Generative 3D Reconstruction in ImageNet
von: Reddy, Pradyumna, et al.
Veröffentlicht: (2024) -
$V_kD:$ Improving Knowledge Distillation using Orthogonal Projections
von: Miles, Roy, et al.
Veröffentlicht: (2024) -
Deep Active Learning: A Reality Check
von: Gashi, Edrina, et al.
Veröffentlicht: (2024)