Gespeichert in:
| Hauptverfasser: | Qiu, Zeju, Buchholz, Simon, Xiao, Tim Z., Dax, Maximilian, Schölkopf, Bernhard, Liu, Weiyang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2506.08001 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Orthogonal Finetuning Made Scalable
von: Qiu, Zeju, et al.
Veröffentlicht: (2025)
von: Qiu, Zeju, et al.
Veröffentlicht: (2025)
POET-X: Memory-efficient LLM Training by Scaling Orthogonal Transformation
von: Qiu, Zeju, et al.
Veröffentlicht: (2026)
von: Qiu, Zeju, et al.
Veröffentlicht: (2026)
Pion: A Spectrum-Preserving Optimizer via Orthogonal Equivalence Transformation
von: Shi, Kexuan, et al.
Veröffentlicht: (2026)
von: Shi, Kexuan, et al.
Veröffentlicht: (2026)
Parameter-Efficient Orthogonal Finetuning via Butterfly Factorization
von: Liu, Weiyang, et al.
Veröffentlicht: (2023)
von: Liu, Weiyang, et al.
Veröffentlicht: (2023)
Can Large Language Models Understand Symbolic Graphics Programs?
von: Qiu, Zeju, et al.
Veröffentlicht: (2024)
von: Qiu, Zeju, et al.
Veröffentlicht: (2024)
Identifying Intervenable and Interpretable Features via Orthogonality Regularization
von: Miller, Moritz, et al.
Veröffentlicht: (2026)
von: Miller, Moritz, et al.
Veröffentlicht: (2026)
Rigidity-Aware Geometric Pretraining for Protein Design and Conformational Ensembles
von: Ni, Zhanghan, et al.
Veröffentlicht: (2026)
von: Ni, Zhanghan, et al.
Veröffentlicht: (2026)
Controlling Text-to-Image Diffusion by Orthogonal Finetuning
von: Qiu, Zeju, et al.
Veröffentlicht: (2023)
von: Qiu, Zeju, et al.
Veröffentlicht: (2023)
Flipping Against All Odds: Reducing LLM Coin Flip Bias via Verbalized Rejection Sampling
von: Xiao, Tim Z., et al.
Veröffentlicht: (2025)
von: Xiao, Tim Z., et al.
Veröffentlicht: (2025)
Algorithmic causal structure emerging through compression
von: Wendong, Liang, et al.
Veröffentlicht: (2025)
von: Wendong, Liang, et al.
Veröffentlicht: (2025)
PEFT-Arena: Understanding Parameter-Efficient Finetuning from a Stability-Plasticity Perspective
von: Huang, Yangyi, et al.
Veröffentlicht: (2026)
von: Huang, Yangyi, et al.
Veröffentlicht: (2026)
ShiftAddLLM: Accelerating Pretrained LLMs via Post-Training Multiplication-Less Reparameterization
von: You, Haoran, et al.
Veröffentlicht: (2024)
von: You, Haoran, et al.
Veröffentlicht: (2024)
Verbalized Machine Learning: Revisiting Machine Learning with Language Models
von: Xiao, Tim Z., et al.
Veröffentlicht: (2024)
von: Xiao, Tim Z., et al.
Veröffentlicht: (2024)
Your Finetuned Large Language Model is Already a Powerful Out-of-distribution Detector
von: Zhang, Andi, et al.
Veröffentlicht: (2024)
von: Zhang, Andi, et al.
Veröffentlicht: (2024)
Limits of Transformer Language Models on Learning to Compose Algorithms
von: Thomm, Jonathan, et al.
Veröffentlicht: (2024)
von: Thomm, Jonathan, et al.
Veröffentlicht: (2024)
Learning to Reason Efficiently with A* Post-Training
von: Opedal, Andreas, et al.
Veröffentlicht: (2026)
von: Opedal, Andreas, et al.
Veröffentlicht: (2026)
Learning Interpretable Concepts: Unifying Causal Representation Learning and Foundation Models
von: Rajendran, Goutham, et al.
Veröffentlicht: (2024)
von: Rajendran, Goutham, et al.
Veröffentlicht: (2024)
A Compact Representation for Bayesian Neural Networks By Removing Permutation Symmetry
von: Xiao, Tim Z., et al.
Veröffentlicht: (2023)
von: Xiao, Tim Z., et al.
Veröffentlicht: (2023)
Improving Large Language Model Safety with Contrastive Representation Learning
von: Simko, Samuel, et al.
Veröffentlicht: (2025)
von: Simko, Samuel, et al.
Veröffentlicht: (2025)
Counterfactual reasoning: an analysis of in-context emergence
von: Miller, Moritz, et al.
Veröffentlicht: (2025)
von: Miller, Moritz, et al.
Veröffentlicht: (2025)
Latent Context Compilation: Distilling Long Context into Compact Portable Memory
von: Li, Zeju, et al.
Veröffentlicht: (2026)
von: Li, Zeju, et al.
Veröffentlicht: (2026)
MathGAP: Out-of-Distribution Evaluation on Problems with Arbitrarily Complex Proofs
von: Opedal, Andreas, et al.
Veröffentlicht: (2024)
von: Opedal, Andreas, et al.
Veröffentlicht: (2024)
Training-Free Tokenizer Transplantation via Orthogonal Matching Pursuit
von: Goddard, Charles, et al.
Veröffentlicht: (2025)
von: Goddard, Charles, et al.
Veröffentlicht: (2025)
Causal Component Analysis
von: Wendong, Liang, et al.
Veröffentlicht: (2023)
von: Wendong, Liang, et al.
Veröffentlicht: (2023)
ButterflyQuant: Ultra-low-bit LLM Quantization through Learnable Orthogonal Butterfly Transforms
von: Xu, Bingxin, et al.
Veröffentlicht: (2025)
von: Xu, Bingxin, et al.
Veröffentlicht: (2025)
Generalized Interpolating Discrete Diffusion
von: von Rütte, Dimitri, et al.
Veröffentlicht: (2025)
von: von Rütte, Dimitri, et al.
Veröffentlicht: (2025)
Learning Beyond Pattern Matching? Assaying Mathematical Understanding in LLMs
von: Guo, Siyuan, et al.
Veröffentlicht: (2024)
von: Guo, Siyuan, et al.
Veröffentlicht: (2024)
Improving LLM Code Reasoning via Semantic Equivalence Self-Play with Formal Verification
von: Barone, Antonio Valerio Miceli, et al.
Veröffentlicht: (2026)
von: Barone, Antonio Valerio Miceli, et al.
Veröffentlicht: (2026)
Robustness of Nonlinear Representation Learning
von: Buchholz, Simon, et al.
Veröffentlicht: (2025)
von: Buchholz, Simon, et al.
Veröffentlicht: (2025)
OrthoRank: Token Selection via Sink Token Orthogonality for Efficient LLM inference
von: Shin, Seungjun, et al.
Veröffentlicht: (2025)
von: Shin, Seungjun, et al.
Veröffentlicht: (2025)
Can Large Language Models Infer Causation from Correlation?
von: Jin, Zhijing, et al.
Veröffentlicht: (2023)
von: Jin, Zhijing, et al.
Veröffentlicht: (2023)
Analyzing the Role of Semantic Representations in the Era of Large Language Models
von: Jin, Zhijing, et al.
Veröffentlicht: (2024)
von: Jin, Zhijing, et al.
Veröffentlicht: (2024)
Intrinsically Interpretable Attention via Sparse Post-Training
von: Draye, Florent, et al.
Veröffentlicht: (2025)
von: Draye, Florent, et al.
Veröffentlicht: (2025)
MatryoshkaKV: Adaptive KV Compression via Trainable Orthogonal Projection
von: Lin, Bokai, et al.
Veröffentlicht: (2024)
von: Lin, Bokai, et al.
Veröffentlicht: (2024)
Selection of LLM Fine-Tuning Data based on Orthogonal Rules
von: Li, Xiaomin, et al.
Veröffentlicht: (2024)
von: Li, Xiaomin, et al.
Veröffentlicht: (2024)
Enhancing LLM Reasoning via Critique Models with Test-Time and Training-Time Supervision
von: Xi, Zhiheng, et al.
Veröffentlicht: (2024)
von: Xi, Zhiheng, et al.
Veröffentlicht: (2024)
Do Language Models Exhibit the Same Cognitive Biases in Problem Solving as Human Learners?
von: Opedal, Andreas, et al.
Veröffentlicht: (2024)
von: Opedal, Andreas, et al.
Veröffentlicht: (2024)
InfLLM: Training-Free Long-Context Extrapolation for LLMs with an Efficient Context Memory
von: Xiao, Chaojun, et al.
Veröffentlicht: (2024)
von: Xiao, Chaojun, et al.
Veröffentlicht: (2024)
TransformLLM: Adapting Large Language Models via LLM-Transformed Reading Comprehension Text
von: Arbel, Iftach, et al.
Veröffentlicht: (2024)
von: Arbel, Iftach, et al.
Veröffentlicht: (2024)
Rethinking GSPO: The Perplexity-Entropy Equivalence
von: Liu, Chi
Veröffentlicht: (2025)
von: Liu, Chi
Veröffentlicht: (2025)
Ähnliche Einträge
-
Orthogonal Finetuning Made Scalable
von: Qiu, Zeju, et al.
Veröffentlicht: (2025) -
POET-X: Memory-efficient LLM Training by Scaling Orthogonal Transformation
von: Qiu, Zeju, et al.
Veröffentlicht: (2026) -
Pion: A Spectrum-Preserving Optimizer via Orthogonal Equivalence Transformation
von: Shi, Kexuan, et al.
Veröffentlicht: (2026) -
Parameter-Efficient Orthogonal Finetuning via Butterfly Factorization
von: Liu, Weiyang, et al.
Veröffentlicht: (2023) -
Can Large Language Models Understand Symbolic Graphics Programs?
von: Qiu, Zeju, et al.
Veröffentlicht: (2024)