Watch the Weights: Unsupervised monitoring and control of fine-tuned LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhong, Ziqian, Raghunathan, Aditi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ImpossibleBench: Measuring LLMs' Propensity of Exploiting Test Cases
von: Zhong, Ziqian, et al.
Veröffentlicht: (2025)
von: Zhong, Ziqian, et al.
Veröffentlicht: (2025)
Base Models Look Human To AI Detectors
von: Xu, Yixuan Even, et al.
Veröffentlicht: (2026)
von: Xu, Yixuan Even, et al.
Veröffentlicht: (2026)
Understanding Finetuning for Factual Knowledge Extraction
von: Ghosal, Gaurav, et al.
Veröffentlicht: (2024)
von: Ghosal, Gaurav, et al.
Veröffentlicht: (2024)
Self-Trained Verification for Training- and Test-Time Self-Improvement
von: Wu, Chen Henry, et al.
Veröffentlicht: (2026)
von: Wu, Chen Henry, et al.
Veröffentlicht: (2026)
Understanding Catastrophic Forgetting in Language Models via Implicit Inference
von: Kotha, Suhas, et al.
Veröffentlicht: (2023)
von: Kotha, Suhas, et al.
Veröffentlicht: (2023)
Jailbreaking in the Haystack
von: Shah, Rishi Rajesh, et al.
Veröffentlicht: (2025)
von: Shah, Rishi Rajesh, et al.
Veröffentlicht: (2025)
Mitigating Bias in RAG: Controlling the Embedder
von: Kim, Taeyoun, et al.
Veröffentlicht: (2025)
von: Kim, Taeyoun, et al.
Veröffentlicht: (2025)
Testing the Limits of Jailbreaking Defenses with the Purple Problem
von: Kim, Taeyoun, et al.
Veröffentlicht: (2024)
von: Kim, Taeyoun, et al.
Veröffentlicht: (2024)
Complexity-aware fine-tuning
von: Goncharov, Andrey, et al.
Veröffentlicht: (2025)
von: Goncharov, Andrey, et al.
Veröffentlicht: (2025)
Context-Parametric Inversion: Why Instruction Finetuning Can Worsen Context Reliance
von: Goyal, Sachin, et al.
Veröffentlicht: (2024)
von: Goyal, Sachin, et al.
Veröffentlicht: (2024)
Mode-Conditioning Unlocks Superior Test-Time Scaling
von: Wu, Chen Henry, et al.
Veröffentlicht: (2025)
von: Wu, Chen Henry, et al.
Veröffentlicht: (2025)
Uncertainty quantification in fine-tuned LLMs using LoRA ensembles
von: Balabanov, Oleksandr, et al.
Veröffentlicht: (2024)
von: Balabanov, Oleksandr, et al.
Veröffentlicht: (2024)
Sharpness-Aware Pretraining Mitigates Catastrophic Forgetting
von: Watts, Ishaan, et al.
Veröffentlicht: (2026)
von: Watts, Ishaan, et al.
Veröffentlicht: (2026)
Repetition Improves Language Model Embeddings
von: Springer, Jacob Mitchell, et al.
Veröffentlicht: (2024)
von: Springer, Jacob Mitchell, et al.
Veröffentlicht: (2024)
Replaying pre-training data improves fine-tuning
von: Kotha, Suhas, et al.
Veröffentlicht: (2026)
von: Kotha, Suhas, et al.
Veröffentlicht: (2026)
Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction
von: Nagarajan, Vaishnavh, et al.
Veröffentlicht: (2025)
von: Nagarajan, Vaishnavh, et al.
Veröffentlicht: (2025)
Memento: Fine-tuning LLM Agents without Fine-tuning LLMs
von: Zhou, Huichi, et al.
Veröffentlicht: (2025)
von: Zhou, Huichi, et al.
Veröffentlicht: (2025)
You can remove GPT2's LayerNorm by fine-tuning
von: Heimersheim, Stefan
Veröffentlicht: (2024)
von: Heimersheim, Stefan
Veröffentlicht: (2024)
LayerNorm: A key component in parameter-efficient fine-tuning
von: ValizadehAslani, Taha, et al.
Veröffentlicht: (2024)
von: ValizadehAslani, Taha, et al.
Veröffentlicht: (2024)
Algorithmic Capabilities of Random Transformers
von: Zhong, Ziqian, et al.
Veröffentlicht: (2024)
von: Zhong, Ziqian, et al.
Veröffentlicht: (2024)
The representation landscape of few-shot learning and fine-tuning in large language models
von: Doimo, Diego, et al.
Veröffentlicht: (2024)
von: Doimo, Diego, et al.
Veröffentlicht: (2024)
The more polypersonal the better -- a short look on space geometry of fine-tuned layers
von: Kudriashov, Sergei, et al.
Veröffentlicht: (2025)
von: Kudriashov, Sergei, et al.
Veröffentlicht: (2025)
Topic Modeling with Fine-tuning LLMs and Bag of Sentences
von: Schneider, Johannes
Veröffentlicht: (2024)
von: Schneider, Johannes
Veröffentlicht: (2024)
RoLoRA: Fine-tuning Rotated Outlier-free LLMs for Effective Weight-Activation Quantization
von: Huang, Xijie, et al.
Veröffentlicht: (2024)
von: Huang, Xijie, et al.
Veröffentlicht: (2024)
Deep literature reviews: an application of fine-tuned language models to migration research
von: Iacus, Stefano M., et al.
Veröffentlicht: (2025)
von: Iacus, Stefano M., et al.
Veröffentlicht: (2025)
Does fine-tuning GPT-3 with the OpenAI API leak personally-identifiable information?
von: Sun, Albert Yu, et al.
Veröffentlicht: (2023)
von: Sun, Albert Yu, et al.
Veröffentlicht: (2023)
Rethinking harmless refusals when fine-tuning foundation models
von: Pop, Florin, et al.
Veröffentlicht: (2024)
von: Pop, Florin, et al.
Veröffentlicht: (2024)
T-MARS: Improving Visual Representations by Circumventing Text Feature Learning
von: Maini, Pratyush, et al.
Veröffentlicht: (2023)
von: Maini, Pratyush, et al.
Veröffentlicht: (2023)
Multi-task retriever fine-tuning for domain-specific and efficient RAG
von: Béchard, Patrice, et al.
Veröffentlicht: (2025)
von: Béchard, Patrice, et al.
Veröffentlicht: (2025)
Scaling Laws for Precision
von: Kumar, Tanishq, et al.
Veröffentlicht: (2024)
von: Kumar, Tanishq, et al.
Veröffentlicht: (2024)
Multilingual Amnesia: On the Transferability of Unlearning in Multilingual LLMs
von: Farashah, Alireza Dehghanpour, et al.
Veröffentlicht: (2026)
von: Farashah, Alireza Dehghanpour, et al.
Veröffentlicht: (2026)
Improving embedding with contrastive fine-tuning on small datasets with expert-augmented scores
von: Lu, Jun, et al.
Veröffentlicht: (2024)
von: Lu, Jun, et al.
Veröffentlicht: (2024)
Robust and Efficient Fine-tuning of LLMs with Bayesian Reparameterization of Low-Rank Adaptation
von: Sengupta, Ayan, et al.
Veröffentlicht: (2024)
von: Sengupta, Ayan, et al.
Veröffentlicht: (2024)
Flip-Flop Consistency: Unsupervised Training for Robustness to Prompt Perturbations in LLMs
von: Hejabi, Parsa, et al.
Veröffentlicht: (2025)
von: Hejabi, Parsa, et al.
Veröffentlicht: (2025)
Towards Safer Pretraining: Analyzing and Filtering Harmful Content in Webscale datasets for Responsible LLMs
von: Mendu, Sai Krishna, et al.
Veröffentlicht: (2025)
von: Mendu, Sai Krishna, et al.
Veröffentlicht: (2025)
Pando: Do Interpretability Methods Work When Models Won't Explain Themselves?
von: Zhong, Ziqian, et al.
Veröffentlicht: (2026)
von: Zhong, Ziqian, et al.
Veröffentlicht: (2026)
Confidence-Credibility Aware Weighted Ensembles of Small LLMs Outperform Large LLMs in Emotion Detection
von: Elgabry, Menna, et al.
Veröffentlicht: (2025)
von: Elgabry, Menna, et al.
Veröffentlicht: (2025)
Guiding Giants: Lightweight Controllers for Weighted Activation Steering in LLMs
von: Hegazy, Amr, et al.
Veröffentlicht: (2025)
von: Hegazy, Amr, et al.
Veröffentlicht: (2025)
Pruning Weights but Not Truth: Safeguarding Truthfulness While Pruning LLMs
von: Fu, Yao, et al.
Veröffentlicht: (2025)
von: Fu, Yao, et al.
Veröffentlicht: (2025)
Simultaneous Masking, Not Prompting Optimization: A Paradigm Shift in Fine-tuning LLMs for Simultaneous Translation
von: Raffel, Matthew, et al.
Veröffentlicht: (2024)
von: Raffel, Matthew, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
ImpossibleBench: Measuring LLMs' Propensity of Exploiting Test Cases
von: Zhong, Ziqian, et al.
Veröffentlicht: (2025) -
Base Models Look Human To AI Detectors
von: Xu, Yixuan Even, et al.
Veröffentlicht: (2026) -
Understanding Finetuning for Factual Knowledge Extraction
von: Ghosal, Gaurav, et al.
Veröffentlicht: (2024) -
Self-Trained Verification for Training- and Test-Time Self-Improvement
von: Wu, Chen Henry, et al.
Veröffentlicht: (2026) -
Understanding Catastrophic Forgetting in Language Models via Implicit Inference
von: Kotha, Suhas, et al.
Veröffentlicht: (2023)