Relating Misfit to Gain in Weak-to-Strong Generalization Beyond the Squared Loss
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mulgund, Abhijeet, Pabbaraju, Chirag |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
von: Fadli, Samih
Veröffentlicht: (2025)
von: Fadli, Samih
Veröffentlicht: (2025)
Merge-Bench: Resolve Merge Conflicts with Large Language Models
von: Schesch, Benedikt, et al.
Veröffentlicht: (2026)
von: Schesch, Benedikt, et al.
Veröffentlicht: (2026)
ACE: Exploring Activation Cosine Similarity and Variance for Accurate and Calibration-Efficient LLM Pruning
von: Mi, Zhendong, et al.
Veröffentlicht: (2025)
von: Mi, Zhendong, et al.
Veröffentlicht: (2025)
Scalable GPU-Accelerated Euler Characteristic Curves: Optimization and Differentiable Learning for PyTorch
von: Saxena, Udit
Veröffentlicht: (2025)
von: Saxena, Udit
Veröffentlicht: (2025)
Synergy over Discrepancy: A Partition-Based Approach to Multi-Domain LLM Fine-Tuning
von: Ye, Hua, et al.
Veröffentlicht: (2025)
von: Ye, Hua, et al.
Veröffentlicht: (2025)
KerZOO: Kernel Function Informed Zeroth-Order Optimization for Accurate and Accelerated LLM Fine-Tuning
von: Mi, Zhendong, et al.
Veröffentlicht: (2025)
von: Mi, Zhendong, et al.
Veröffentlicht: (2025)
Revisiting LRP: Positional Attribution as the Missing Ingredient for Transformer Explainability
von: Bakish, Yarden, et al.
Veröffentlicht: (2025)
von: Bakish, Yarden, et al.
Veröffentlicht: (2025)
QuAnTS: Question Answering on Time Series
von: Divo, Felix, et al.
Veröffentlicht: (2025)
von: Divo, Felix, et al.
Veröffentlicht: (2025)
Evaluating the Systematic Reasoning Abilities of Large Language Models through Graph Coloring
von: Heyman, Alex, et al.
Veröffentlicht: (2025)
von: Heyman, Alex, et al.
Veröffentlicht: (2025)
Latent Instruction Representation Alignment: defending against jailbreaks, backdoors and undesired knowledge in LLMs
von: Easley, Eric, et al.
Veröffentlicht: (2026)
von: Easley, Eric, et al.
Veröffentlicht: (2026)
Descriptive Collision in Sparse Autoencoder Auto-Interpretability: When One Explanation Describes Many Features
von: McCann, Jordan F.
Veröffentlicht: (2026)
von: McCann, Jordan F.
Veröffentlicht: (2026)
Super Apriel: One Checkpoint, Many Speeds
von: Labs, SLAM, et al.
Veröffentlicht: (2026)
von: Labs, SLAM, et al.
Veröffentlicht: (2026)
Learned Relay Representations for Forward-Thinking Discrete Diffusion Models
von: Rozonoyer, Benjamin, et al.
Veröffentlicht: (2026)
von: Rozonoyer, Benjamin, et al.
Veröffentlicht: (2026)
Why LoRA Resists Label Noise: A Theoretical Framework for Noise-Robust Parameter-Efficient Fine-Tuning
von: Steele, Brady
Veröffentlicht: (2026)
von: Steele, Brady
Veröffentlicht: (2026)
Variance Is Not Importance: Structural Analysis of Transformer Compressibility Across Model Scales
von: Salfati, Samuel
Veröffentlicht: (2026)
von: Salfati, Samuel
Veröffentlicht: (2026)
TensorLens: End-to-End Transformer Analysis via High-Order Attention Tensors
von: Atad, Ido Andrew, et al.
Veröffentlicht: (2026)
von: Atad, Ido Andrew, et al.
Veröffentlicht: (2026)
Transformer Scalability Crisis: The First Comprehensive Empirical Analysis of Performance Walls in Modern Language Models
von: Moghadasi, Mahdi Naser, et al.
Veröffentlicht: (2026)
von: Moghadasi, Mahdi Naser, et al.
Veröffentlicht: (2026)
On Semantic Loss Fine-Tuning Approach for Preventing Model Collapse in Causal Reasoning
von: Deshmukh, Pratik, et al.
Veröffentlicht: (2026)
von: Deshmukh, Pratik, et al.
Veröffentlicht: (2026)
Latent Cache Flow: Model-to-Model Communication Without Text
von: Rossi, Maximillian, et al.
Veröffentlicht: (2026)
von: Rossi, Maximillian, et al.
Veröffentlicht: (2026)
Node-Level Uncertainty Estimation in LLM-Generated SQL
von: Hasson, Hilaf, et al.
Veröffentlicht: (2025)
von: Hasson, Hilaf, et al.
Veröffentlicht: (2025)
Thread Detection and Response Generation using Transformers with Prompt Optimisation
von: T, Kevin Joshua, et al.
Veröffentlicht: (2024)
von: T, Kevin Joshua, et al.
Veröffentlicht: (2024)
Ouroboros: Dynamic Weight Generation for Recursive Transformers via Input-Conditioned LoRA Modulation
von: Jaber, Jaber, et al.
Veröffentlicht: (2026)
von: Jaber, Jaber, et al.
Veröffentlicht: (2026)
From Syntax to Semantics: Unveiling the Emergence of Chirality in SMILES Translation Models
von: Li, Zehao, et al.
Veröffentlicht: (2026)
von: Li, Zehao, et al.
Veröffentlicht: (2026)
Listwise Direct Preference Optimization with Multi-Dimensional Preference Mixing
von: Sun, Yuhui, et al.
Veröffentlicht: (2025)
von: Sun, Yuhui, et al.
Veröffentlicht: (2025)
Weakly Supervised Distillation of Hallucination Signals into Transformer Representations
von: Salehmohamed, Shoaib Sadiq, et al.
Veröffentlicht: (2026)
von: Salehmohamed, Shoaib Sadiq, et al.
Veröffentlicht: (2026)
Before the Last Token: Diagnosing Final-Token Safety Probe Failures
von: Doda, Shravan
Veröffentlicht: (2026)
von: Doda, Shravan
Veröffentlicht: (2026)
Social Cooperation in Conversational AI Agents
von: Çelikok, Mustafa Mert, et al.
Veröffentlicht: (2025)
von: Çelikok, Mustafa Mert, et al.
Veröffentlicht: (2025)
Discovering Transformer Circuits via a Hybrid Attribution and Pruning Framework
von: Gu, Hao, et al.
Veröffentlicht: (2025)
von: Gu, Hao, et al.
Veröffentlicht: (2025)
PoTS: Proof-of-Training-Steps for Backdoor Detection in Large Language Models
von: Seddik, Issam, et al.
Veröffentlicht: (2025)
von: Seddik, Issam, et al.
Veröffentlicht: (2025)
When Models Can't Follow: Testing Instruction Adherence Across 256 LLMs
von: Young, Richard J., et al.
Veröffentlicht: (2025)
von: Young, Richard J., et al.
Veröffentlicht: (2025)
FastForward Pruning: Efficient LLM Pruning via Single-Step Reinforcement Learning
von: Yuan, Xin, et al.
Veröffentlicht: (2025)
von: Yuan, Xin, et al.
Veröffentlicht: (2025)
Bayesian Attention Mechanism: A Probabilistic Framework for Positional Encoding and Context Length Extrapolation
von: Bianchessi, Arthur S., et al.
Veröffentlicht: (2025)
von: Bianchessi, Arthur S., et al.
Veröffentlicht: (2025)
Rethinking Addressing in Language Models via Contexualized Equivariant Positional Encoding
von: Zhu, Jiajun, et al.
Veröffentlicht: (2025)
von: Zhu, Jiajun, et al.
Veröffentlicht: (2025)
OFMU: Optimization-Driven Framework for Machine Unlearning
von: Asif, Sadia, et al.
Veröffentlicht: (2025)
von: Asif, Sadia, et al.
Veröffentlicht: (2025)
Persona Features Control Emergent Misalignment
von: Wang, Miles, et al.
Veröffentlicht: (2025)
von: Wang, Miles, et al.
Veröffentlicht: (2025)
Autonomous Deep Agent
von: Yu, Amy, et al.
Veröffentlicht: (2025)
von: Yu, Amy, et al.
Veröffentlicht: (2025)
The Anti-Ouroboros Effect: Emergent Resilience in Large Language Models from Recursive Selective Feedback
von: Adapala, Sai Teja Reddy
Veröffentlicht: (2025)
von: Adapala, Sai Teja Reddy
Veröffentlicht: (2025)
FastGRPO: Accelerating Policy Optimization via Concurrency-aware Speculative Decoding and Online Draft Learning
von: Zhang, Yizhou, et al.
Veröffentlicht: (2025)
von: Zhang, Yizhou, et al.
Veröffentlicht: (2025)
The Data Efficiency Frontier of Financial Foundation Models: Scaling Laws from Continued Pretraining
von: Ponnock, Jesse
Veröffentlicht: (2025)
von: Ponnock, Jesse
Veröffentlicht: (2025)
Overclocking LLM Reasoning: Monitoring and Controlling Thinking Path Lengths in LLMs
von: Eisenstadt, Roy, et al.
Veröffentlicht: (2025)
von: Eisenstadt, Roy, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Entropy-Based Measurement of Value Drift and Alignment Work in Large Language Models
von: Fadli, Samih
Veröffentlicht: (2025) -
Merge-Bench: Resolve Merge Conflicts with Large Language Models
von: Schesch, Benedikt, et al.
Veröffentlicht: (2026) -
ACE: Exploring Activation Cosine Similarity and Variance for Accurate and Calibration-Efficient LLM Pruning
von: Mi, Zhendong, et al.
Veröffentlicht: (2025) -
Scalable GPU-Accelerated Euler Characteristic Curves: Optimization and Differentiable Learning for PyTorch
von: Saxena, Udit
Veröffentlicht: (2025) -
Synergy over Discrepancy: A Partition-Based Approach to Multi-Domain LLM Fine-Tuning
von: Ye, Hua, et al.
Veröffentlicht: (2025)