The Remarkable Robustness of LLMs: Stages of Inference?
Fuente:
arXiv
Salvato in:
| Autori principali: | Lad, Vedang, Lee, Jin Hwa, Gurnee, Wes, Tegmark, Max |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Language Models Represent Space and Time
di: Gurnee, Wes, et al.
Pubblicazione: (2023)
di: Gurnee, Wes, et al.
Pubblicazione: (2023)
Language Models Use Trigonometry to Do Addition
di: Kantamneni, Subhash, et al.
Pubblicazione: (2025)
di: Kantamneni, Subhash, et al.
Pubblicazione: (2025)
Confidence Regulation Neurons in Language Models
di: Stolfo, Alessandro, et al.
Pubblicazione: (2024)
di: Stolfo, Alessandro, et al.
Pubblicazione: (2024)
Refusal in Language Models Is Mediated by a Single Direction
di: Arditi, Andy, et al.
Pubblicazione: (2024)
di: Arditi, Andy, et al.
Pubblicazione: (2024)
Universal Neurons in GPT2 Language Models
di: Gurnee, Wes, et al.
Pubblicazione: (2024)
di: Gurnee, Wes, et al.
Pubblicazione: (2024)
Dense SAE Latents Are Features, Not Bugs
di: Sun, Xiaoqing, et al.
Pubblicazione: (2025)
di: Sun, Xiaoqing, et al.
Pubblicazione: (2025)
Two-Stage Regularization-Based Structured Pruning for LLMs
di: Feng, Mingkuan, et al.
Pubblicazione: (2025)
di: Feng, Mingkuan, et al.
Pubblicazione: (2025)
Distinct Computations Emerge From Compositional Curricula in In-Context Learning
di: Lee, Jin Hwa, et al.
Pubblicazione: (2025)
di: Lee, Jin Hwa, et al.
Pubblicazione: (2025)
The Impact of Inference Acceleration on Bias of LLMs
di: Kirsten, Elisabeth, et al.
Pubblicazione: (2024)
di: Kirsten, Elisabeth, et al.
Pubblicazione: (2024)
Not All Language Model Features Are One-Dimensionally Linear
di: Engels, Joshua, et al.
Pubblicazione: (2024)
di: Engels, Joshua, et al.
Pubblicazione: (2024)
Not All Layers of LLMs Are Necessary During Inference
di: Fan, Siqi, et al.
Pubblicazione: (2024)
di: Fan, Siqi, et al.
Pubblicazione: (2024)
Reward-Robust RLHF in LLMs
di: Yan, Yuzi, et al.
Pubblicazione: (2024)
di: Yan, Yuzi, et al.
Pubblicazione: (2024)
2SSP: A Two-Stage Framework for Structured Pruning of LLMs
di: Sandri, Fabrizio, et al.
Pubblicazione: (2025)
di: Sandri, Fabrizio, et al.
Pubblicazione: (2025)
Self-Training Meets Consistency: Improving LLMs' Reasoning with Consistency-Driven Rationale Evaluation
di: Lee, Jaehyeok, et al.
Pubblicazione: (2024)
di: Lee, Jaehyeok, et al.
Pubblicazione: (2024)
Nudging: Inference-time Alignment of LLMs via Guided Decoding
di: Fei, Yu, et al.
Pubblicazione: (2024)
di: Fei, Yu, et al.
Pubblicazione: (2024)
GLASS: Global-Local Aggregation for Inference-time Sparsification of LLMs
di: Sattarifard, Amirmohsen, et al.
Pubblicazione: (2025)
di: Sattarifard, Amirmohsen, et al.
Pubblicazione: (2025)
Agile-Quant: Activation-Guided Quantization for Faster Inference of LLMs on the Edge
di: Shen, Xuan, et al.
Pubblicazione: (2023)
di: Shen, Xuan, et al.
Pubblicazione: (2023)
Detecting Hallucinations in SpeechLLMs at Inference Time Using Attention Maps
di: Waldendorf, Jonas, et al.
Pubblicazione: (2026)
di: Waldendorf, Jonas, et al.
Pubblicazione: (2026)
ProxySPEX: Inference-Efficient Interpretability via Sparse Feature Interactions in LLMs
di: Butler, Landon, et al.
Pubblicazione: (2025)
di: Butler, Landon, et al.
Pubblicazione: (2025)
Semantic Entropy Probes: Robust and Cheap Hallucination Detection in LLMs
di: Kossen, Jannik, et al.
Pubblicazione: (2024)
di: Kossen, Jannik, et al.
Pubblicazione: (2024)
Shifting Perspectives: Steering Vectors for Robust Bias Mitigation in LLMs
di: Siddique, Zara, et al.
Pubblicazione: (2025)
di: Siddique, Zara, et al.
Pubblicazione: (2025)
GraphDancer: Training LLMs to Explore and Reason over Graphs via Two-Stage Curriculum Post-Training
di: Bai, Yuyang, et al.
Pubblicazione: (2026)
di: Bai, Yuyang, et al.
Pubblicazione: (2026)
Geometric Signatures of Compositionality Across a Language Model's Lifetime
di: Lee, Jin Hwa, et al.
Pubblicazione: (2024)
di: Lee, Jin Hwa, et al.
Pubblicazione: (2024)
From Associations to Activations: Comparing Behavioral and Hidden-State Semantic Geometry in LLMs
di: Schiekiera, Louis, et al.
Pubblicazione: (2026)
di: Schiekiera, Louis, et al.
Pubblicazione: (2026)
AQUA: Attention via QUery mAgnitudes for Memory and Compute Efficient Inference in LLMs
di: S, Santhosh G, et al.
Pubblicazione: (2025)
di: S, Santhosh G, et al.
Pubblicazione: (2025)
Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
di: Sheshadri, Abhay, et al.
Pubblicazione: (2024)
di: Sheshadri, Abhay, et al.
Pubblicazione: (2024)
SPARC: Subspace-Aware Prompt Adaptation for Robust Continual Learning in LLMs
di: Jayasuriya, Dinithi, et al.
Pubblicazione: (2025)
di: Jayasuriya, Dinithi, et al.
Pubblicazione: (2025)
Enhancing Effectiveness and Robustness in a Low-Resource Regime via Decision-Boundary-aware Data Augmentation
di: Jin, Kyohoon, et al.
Pubblicazione: (2024)
di: Jin, Kyohoon, et al.
Pubblicazione: (2024)
Steering Safely or Off a Cliff? Rethinking Specificity and Robustness in Inference-Time Interventions
di: Goyal, Navita, et al.
Pubblicazione: (2026)
di: Goyal, Navita, et al.
Pubblicazione: (2026)
Is Conformal Factuality for RAG-based LLMs Robust? Novel Metrics and Systematic Insights
di: Chen, Yi, et al.
Pubblicazione: (2026)
di: Chen, Yi, et al.
Pubblicazione: (2026)
Adaptive Inference-Time Compute: LLMs Can Predict if They Can Do Better, Even Mid-Generation
di: Manvi, Rohin, et al.
Pubblicazione: (2024)
di: Manvi, Rohin, et al.
Pubblicazione: (2024)
LLMs Should Express Uncertainty Explicitly
di: Guo, Junyu, et al.
Pubblicazione: (2026)
di: Guo, Junyu, et al.
Pubblicazione: (2026)
Sustainable LLM Inference for Edge AI: Evaluating Quantized LLMs for Energy Efficiency, Output Accuracy, and Inference Latency
di: Husom, Erik Johannes, et al.
Pubblicazione: (2025)
di: Husom, Erik Johannes, et al.
Pubblicazione: (2025)
Towards Self-Robust LLMs: Intrinsic Prompt Noise Resistance via CoIPO
di: Yang, Xin, et al.
Pubblicazione: (2026)
di: Yang, Xin, et al.
Pubblicazione: (2026)
Robust Checkpoint Selection for Multimodal LLMs via Agentic Evaluation and Stability-Aware Ranking
di: Xu, Qinwu, et al.
Pubblicazione: (2026)
di: Xu, Qinwu, et al.
Pubblicazione: (2026)
The Reversal Curse: LLMs trained on "A is B" fail to learn "B is A"
di: Berglund, Lukas, et al.
Pubblicazione: (2023)
di: Berglund, Lukas, et al.
Pubblicazione: (2023)
Sample-Efficient Alignment for LLMs
di: Liu, Zichen, et al.
Pubblicazione: (2024)
di: Liu, Zichen, et al.
Pubblicazione: (2024)
LLMs Know What to Drop: Self-Attention Guided KV Cache Eviction for Efficient Long-Context Inference
di: Wang, Guangtao, et al.
Pubblicazione: (2025)
di: Wang, Guangtao, et al.
Pubblicazione: (2025)
SAGE: Shaping Anchors for Guided Exploration in RLVR of LLMs
di: Lee, Chanuk, et al.
Pubblicazione: (2026)
di: Lee, Chanuk, et al.
Pubblicazione: (2026)
An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
di: Hao, Yuren, et al.
Pubblicazione: (2025)
di: Hao, Yuren, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Language Models Represent Space and Time
di: Gurnee, Wes, et al.
Pubblicazione: (2023) -
Language Models Use Trigonometry to Do Addition
di: Kantamneni, Subhash, et al.
Pubblicazione: (2025) -
Confidence Regulation Neurons in Language Models
di: Stolfo, Alessandro, et al.
Pubblicazione: (2024) -
Refusal in Language Models Is Mediated by a Single Direction
di: Arditi, Andy, et al.
Pubblicazione: (2024) -
Universal Neurons in GPT2 Language Models
di: Gurnee, Wes, et al.
Pubblicazione: (2024)