Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
Fuente:
arXiv
Salvato in:
| Autore principale: | Fadli, Samih |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Beyond Hallucinations: A Composite Score for Measuring Reliability in Open-Source Large Language Models
di: Salla, Rohit Kumar, et al.
Pubblicazione: (2025)
di: Salla, Rohit Kumar, et al.
Pubblicazione: (2025)
Towards Alignment-Centric Paradigm: A Survey of Instruction Tuning in Large Language Models
di: Han, Xudong, et al.
Pubblicazione: (2025)
di: Han, Xudong, et al.
Pubblicazione: (2025)
Merge-Bench: Resolve Merge Conflicts with Large Language Models
di: Schesch, Benedikt, et al.
Pubblicazione: (2026)
di: Schesch, Benedikt, et al.
Pubblicazione: (2026)
Automated CAD Modeling Sequence Generation from Text Descriptions via Transformer-Based Large Language Models
di: Liao, Jianxing, et al.
Pubblicazione: (2025)
di: Liao, Jianxing, et al.
Pubblicazione: (2025)
Turning the TIDE: Cross-Architecture Distillation for Diffusion Large Language Models
di: Zhang, Gongbo, et al.
Pubblicazione: (2026)
di: Zhang, Gongbo, et al.
Pubblicazione: (2026)
Cognitive Load Limits in Large Language Models: Benchmarking Multi-Hop Reasoning
di: Adapala, Sai Teja Reddy
Pubblicazione: (2025)
di: Adapala, Sai Teja Reddy
Pubblicazione: (2025)
Painless Activation Steering: An Automated, Lightweight Approach for Post-Training Large Language Models
di: Cui, Sasha, et al.
Pubblicazione: (2025)
di: Cui, Sasha, et al.
Pubblicazione: (2025)
The Geometry of Thought: How Scale Restructures Reasoning In Large Language Models
di: Anderson, Samuel Cyrenius
Pubblicazione: (2026)
di: Anderson, Samuel Cyrenius
Pubblicazione: (2026)
Memory Bank Compression for Continual Adaptation of Large Language Models
di: Katraouras, Thomas, et al.
Pubblicazione: (2026)
di: Katraouras, Thomas, et al.
Pubblicazione: (2026)
Assessing Large Language Models on Islamic Legal Reasoning: Evidence from Inheritance Law Evaluation
di: Bouchekif, Abdessalam, et al.
Pubblicazione: (2025)
di: Bouchekif, Abdessalam, et al.
Pubblicazione: (2025)
SECURA: Sigmoid-Enhanced CUR Decomposition with Uninterrupted Retention and Low-Rank Adaptation in Large Language Models
di: Zhang, Yuxuan
Pubblicazione: (2025)
di: Zhang, Yuxuan
Pubblicazione: (2025)
Reasoning Large Language Model Errors Arise from Hallucinating Critical Problem Features
di: Heyman, Alex, et al.
Pubblicazione: (2025)
di: Heyman, Alex, et al.
Pubblicazione: (2025)
The Anti-Ouroboros Effect: Emergent Resilience in Large Language Models from Recursive Selective Feedback
di: Adapala, Sai Teja Reddy
Pubblicazione: (2025)
di: Adapala, Sai Teja Reddy
Pubblicazione: (2025)
Quantization Undoes Alignment: Bias Emergence in Compressed LLMs Across Models and Precision Levels
di: Rath, Plawan Kumar, et al.
Pubblicazione: (2026)
di: Rath, Plawan Kumar, et al.
Pubblicazione: (2026)
Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring
di: Heyman, Alex, et al.
Pubblicazione: (2025)
di: Heyman, Alex, et al.
Pubblicazione: (2025)
MCP: A Control-Theoretic Orchestration Framework for Synergistic Efficiency and Interpretability in Multimodal Large Language Models
di: Zhang, Luyan
Pubblicazione: (2025)
di: Zhang, Luyan
Pubblicazione: (2025)
Machine Unlearning for Masked Diffusion Language Models
di: Lee, Georu, et al.
Pubblicazione: (2026)
di: Lee, Georu, et al.
Pubblicazione: (2026)
Truth as a Compression Artifact in Language Model Training
di: Krestnikov, Konstantin
Pubblicazione: (2026)
di: Krestnikov, Konstantin
Pubblicazione: (2026)
Prototype Transformer: Towards Language Model Architectures Interpretable by Design
di: Yordanov, Yordan, et al.
Pubblicazione: (2026)
di: Yordanov, Yordan, et al.
Pubblicazione: (2026)
CoE: Collaborative Entropy for Uncertainty Quantification in Agentic Multi-LLM Systems
di: Sun, Kangkang, et al.
Pubblicazione: (2026)
di: Sun, Kangkang, et al.
Pubblicazione: (2026)
Intention Collapse: Intention-Level Metrics for Reasoning in Language Models
di: Vera, Patricio
Pubblicazione: (2026)
di: Vera, Patricio
Pubblicazione: (2026)
Language Models Are Implicitly Continuous
di: Marro, Samuele, et al.
Pubblicazione: (2025)
di: Marro, Samuele, et al.
Pubblicazione: (2025)
Domain-Specific Pretraining of Language Models: A Comparative Study in the Medical Field
di: Kerner, Tobias
Pubblicazione: (2024)
di: Kerner, Tobias
Pubblicazione: (2024)
Annotation Entropy Predicts Per-Example Learning Dynamics in LoRA Fine-Tuning
di: Steele, Brady
Pubblicazione: (2026)
di: Steele, Brady
Pubblicazione: (2026)
The Drill-Down and Fabricate Test (DDFT): A Protocol for Measuring Epistemic Robustness in Language Models
di: Baxi, Rahul
Pubblicazione: (2025)
di: Baxi, Rahul
Pubblicazione: (2025)
The Readout Shortcut: Positional Number Copying Dominates Arithmetic CoT Readout in Small Language Models
di: Liu, Ming
Pubblicazione: (2026)
di: Liu, Ming
Pubblicazione: (2026)
Combining Language and Topic Models for Hierarchical Text Classification
di: Toit, Jaco du, et al.
Pubblicazione: (2025)
di: Toit, Jaco du, et al.
Pubblicazione: (2025)
Self-Pruned Key-Value Attention: Learning When to Write by Predicting Future Utility
di: Szilvasy, Gergely, et al.
Pubblicazione: (2026)
di: Szilvasy, Gergely, et al.
Pubblicazione: (2026)
Rethinking Addressing in Language Models via Contexualized Equivariant Positional Encoding
di: Zhu, Jiajun, et al.
Pubblicazione: (2025)
di: Zhu, Jiajun, et al.
Pubblicazione: (2025)
Alif: Advancing Urdu Large Language Models via Multilingual Synthetic Data Distillation
di: Shafique, Muhammad Ali, et al.
Pubblicazione: (2025)
di: Shafique, Muhammad Ali, et al.
Pubblicazione: (2025)
Aligning LLMs on a Budget: Inference-Time Alignment with Heuristic Reward Models
di: Nakamura, Mason, et al.
Pubblicazione: (2025)
di: Nakamura, Mason, et al.
Pubblicazione: (2025)
Kronecker Embeddings: Byte-Level Structured Token Representations for Parameter-Efficient Language Models
di: Shravan, Rohan
Pubblicazione: (2026)
di: Shravan, Rohan
Pubblicazione: (2026)
How Language Models Process Out-of-Distribution Inputs: A Two-Pathway Framework
di: Saghir, Hamidreza
Pubblicazione: (2026)
di: Saghir, Hamidreza
Pubblicazione: (2026)
Latent Instruction Representation Alignment: defending against jailbreaks, backdoors and undesired knowledge in LLMs
di: Easley, Eric, et al.
Pubblicazione: (2026)
di: Easley, Eric, et al.
Pubblicazione: (2026)
Neural Activation Patterns Across Language Model Architectures: A Comprehensive Analysis of Cognitive Task Performance
di: Naser-Moghadasi, Mahdi, et al.
Pubblicazione: (2026)
di: Naser-Moghadasi, Mahdi, et al.
Pubblicazione: (2026)
D-COT: Disciplined Chain-of-Thought Learning for Efficient Reasoning in Small Language Models
di: Ubukata, Shunsuke
Pubblicazione: (2026)
di: Ubukata, Shunsuke
Pubblicazione: (2026)
Future Token Prediction -- Causal Language Modelling with Per-Token Semantic State Vector for Multi-Token Prediction
di: Walker, Nicholas
Pubblicazione: (2024)
di: Walker, Nicholas
Pubblicazione: (2024)
Language as a Wave Phenomenon: Semantic Phase Locking and Interference in Neural Networks
di: Yıldırım, Alper, et al.
Pubblicazione: (2025)
di: Yıldırım, Alper, et al.
Pubblicazione: (2025)
Emergent Lexical Semantics in Neural Language Models: Testing Martin's Law on LLM-Generated Text
di: Kugler, Kai
Pubblicazione: (2025)
di: Kugler, Kai
Pubblicazione: (2025)
Your Pretrained Model Tells the Difficulty Itself: A Self-Adaptive Curriculum Learning Paradigm for Natural Language Understanding
di: Feng, Qi, et al.
Pubblicazione: (2025)
di: Feng, Qi, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Beyond Hallucinations: A Composite Score for Measuring Reliability in Open-Source Large Language Models
di: Salla, Rohit Kumar, et al.
Pubblicazione: (2025) -
Towards Alignment-Centric Paradigm: A Survey of Instruction Tuning in Large Language Models
di: Han, Xudong, et al.
Pubblicazione: (2025) -
Merge-Bench: Resolve Merge Conflicts with Large Language Models
di: Schesch, Benedikt, et al.
Pubblicazione: (2026) -
Automated CAD Modeling Sequence Generation from Text Descriptions via Transformer-Based Large Language Models
di: Liao, Jianxing, et al.
Pubblicazione: (2025) -
Turning the TIDE: Cross-Architecture Distillation for Diffusion Large Language Models
di: Zhang, Gongbo, et al.
Pubblicazione: (2026)