Saved in:
| Main Authors: | Singh, Sidak Pal, He, Bobby, Hofmann, Thomas, Schölkopf, Bernhard |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2403.07379 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Rethinking Attention: Exploring Shallow Feed-Forward Neural Networks as an Alternative to Attention Layers in Transformers
by: Bozic, Vukasin, et al.
Published: (2023)
by: Bozic, Vukasin, et al.
Published: (2023)
Accelerating Neural Network Training Along Sharp and Flat Directions
by: Zakarin, Daniyar, et al.
Published: (2025)
by: Zakarin, Daniyar, et al.
Published: (2025)
Avoiding spurious sharpness minimization broadens applicability of SAM
by: Singh, Sidak Pal, et al.
Published: (2025)
by: Singh, Sidak Pal, et al.
Published: (2025)
Landscaping Linear Mode Connectivity
by: Singh, Sidak Pal, et al.
Published: (2024)
by: Singh, Sidak Pal, et al.
Published: (2024)
Some Fundamental Aspects about Lipschitz Continuity of Neural Networks
by: Khromov, Grigory, et al.
Published: (2023)
by: Khromov, Grigory, et al.
Published: (2023)
Theoretical characterisation of the Gauss-Newton conditioning in Neural Networks
by: Zhao, Jim, et al.
Published: (2024)
by: Zhao, Jim, et al.
Published: (2024)
Identifying Intervenable and Interpretable Features via Orthogonality Regularization
by: Miller, Moritz, et al.
Published: (2026)
by: Miller, Moritz, et al.
Published: (2026)
On the Emergence and Test-Time Use of Structural Information in Large Language Models
by: Chen, Michelle Chao, et al.
Published: (2026)
by: Chen, Michelle Chao, et al.
Published: (2026)
Generalized Interpolating Discrete Diffusion
by: von Rütte, Dimitri, et al.
Published: (2025)
by: von Rütte, Dimitri, et al.
Published: (2025)
Counterfactual reasoning: an analysis of in-context emergence
by: Miller, Moritz, et al.
Published: (2025)
by: Miller, Moritz, et al.
Published: (2025)
Improving Large Language Model Safety with Contrastive Representation Learning
by: Simko, Samuel, et al.
Published: (2025)
by: Simko, Samuel, et al.
Published: (2025)
Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents
by: Song, Yifan, et al.
Published: (2024)
by: Song, Yifan, et al.
Published: (2024)
Verbalized Machine Learning: Revisiting Machine Learning with Language Models
by: Xiao, Tim Z., et al.
Published: (2024)
by: Xiao, Tim Z., et al.
Published: (2024)
MathGAP: Out-of-Distribution Evaluation on Problems with Arbitrarily Complex Proofs
by: Opedal, Andreas, et al.
Published: (2024)
by: Opedal, Andreas, et al.
Published: (2024)
Towards Meta-Pruning via Optimal Transport
by: Theus, Alexander, et al.
Published: (2024)
by: Theus, Alexander, et al.
Published: (2024)
Flipping Against All Odds: Reducing LLM Coin Flip Bias via Verbalized Rejection Sampling
by: Xiao, Tim Z., et al.
Published: (2025)
by: Xiao, Tim Z., et al.
Published: (2025)
Hyperbolic Busemann Neural Networks
by: Chen, Ziheng, et al.
Published: (2026)
by: Chen, Ziheng, et al.
Published: (2026)
Orthogonal Finetuning Made Scalable
by: Qiu, Zeju, et al.
Published: (2025)
by: Qiu, Zeju, et al.
Published: (2025)
PEFT-Arena: Understanding Parameter-Efficient Finetuning from a Stability-Plasticity Perspective
by: Huang, Yangyi, et al.
Published: (2026)
by: Huang, Yangyi, et al.
Published: (2026)
Local vs Global continual learning
by: Lanzillotta, Giulia, et al.
Published: (2024)
by: Lanzillotta, Giulia, et al.
Published: (2024)
What Does It Mean to Be a Transformer? Insights from a Theoretical Hessian Analysis
by: Ormaniec, Weronika, et al.
Published: (2024)
by: Ormaniec, Weronika, et al.
Published: (2024)
Limits of Transformer Language Models on Learning to Compose Algorithms
by: Thomm, Jonathan, et al.
Published: (2024)
by: Thomm, Jonathan, et al.
Published: (2024)
Learning Beyond Pattern Matching? Assaying Mathematical Understanding in LLMs
by: Guo, Siyuan, et al.
Published: (2024)
by: Guo, Siyuan, et al.
Published: (2024)
Reparameterized LLM Training via Orthogonal Equivalence Transformation
by: Qiu, Zeju, et al.
Published: (2025)
by: Qiu, Zeju, et al.
Published: (2025)
Learning to Reason Efficiently with A* Post-Training
by: Opedal, Andreas, et al.
Published: (2026)
by: Opedal, Andreas, et al.
Published: (2026)
Language Models Can Reduce Asymmetry in Information Markets
by: Rahaman, Nasim, et al.
Published: (2024)
by: Rahaman, Nasim, et al.
Published: (2024)
Analyzing the Role of Semantic Representations in the Era of Large Language Models
by: Jin, Zhijing, et al.
Published: (2024)
by: Jin, Zhijing, et al.
Published: (2024)
Do Language Models Exhibit the Same Cognitive Biases in Problem Solving as Human Learners?
by: Opedal, Andreas, et al.
Published: (2024)
by: Opedal, Andreas, et al.
Published: (2024)
Can Large Language Models Infer Causation from Correlation?
by: Jin, Zhijing, et al.
Published: (2023)
by: Jin, Zhijing, et al.
Published: (2023)
Transformer Fusion with Optimal Transport
by: Imfeld, Moritz, et al.
Published: (2023)
by: Imfeld, Moritz, et al.
Published: (2023)
SLEB: Streamlining LLMs through Redundancy Verification and Elimination of Transformer Blocks
by: Song, Jiwon, et al.
Published: (2024)
by: Song, Jiwon, et al.
Published: (2024)
CLT-Forge: A Scalable Library for Cross-Layer Transcoders and Attribution Graphs
by: Draye, Florent, et al.
Published: (2026)
by: Draye, Florent, et al.
Published: (2026)
Length Desensitization in Direct Preference Optimization
by: Liu, Wei, et al.
Published: (2024)
by: Liu, Wei, et al.
Published: (2024)
CausalCite: A Causal Formulation of Paper Citations
by: Kumar, Ishan, et al.
Published: (2023)
by: Kumar, Ishan, et al.
Published: (2023)
On Affine Homotopy between Language Encoders
by: Chan, Robin SM, et al.
Published: (2024)
by: Chan, Robin SM, et al.
Published: (2024)
GRASP: Replace Redundant Layers with Adaptive Singular Parameters for Efficient Model Compression
by: Liu, Kainan, et al.
Published: (2024)
by: Liu, Kainan, et al.
Published: (2024)
RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding
by: Hoshino, Yuichiro, et al.
Published: (2025)
by: Hoshino, Yuichiro, et al.
Published: (2025)
Physics of Learning: A Lagrangian perspective to different learning paradigms
by: Guo, Siyuan, et al.
Published: (2025)
by: Guo, Siyuan, et al.
Published: (2025)
Understanding Reference Policies in Direct Preference Optimization
by: Liu, Yixin, et al.
Published: (2024)
by: Liu, Yixin, et al.
Published: (2024)
Accelerating Direct Preference Optimization with Prefix Sharing
by: Wang, Franklin, et al.
Published: (2024)
by: Wang, Franklin, et al.
Published: (2024)
Similar Items
-
Rethinking Attention: Exploring Shallow Feed-Forward Neural Networks as an Alternative to Attention Layers in Transformers
by: Bozic, Vukasin, et al.
Published: (2023) -
Accelerating Neural Network Training Along Sharp and Flat Directions
by: Zakarin, Daniyar, et al.
Published: (2025) -
Avoiding spurious sharpness minimization broadens applicability of SAM
by: Singh, Sidak Pal, et al.
Published: (2025) -
Landscaping Linear Mode Connectivity
by: Singh, Sidak Pal, et al.
Published: (2024) -
Some Fundamental Aspects about Lipschitz Continuity of Neural Networks
by: Khromov, Grigory, et al.
Published: (2023)