Saved in:
| Main Authors: | Fan, Ying, Du, Yilun, Ramchandran, Kannan, Lee, Kangwook |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2409.15647 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Optimizing DDPM Sampling with Shortcut Fine-Tuning
by: Fan, Ying, et al.
Published: (2023)
by: Fan, Ying, et al.
Published: (2023)
Self-Improving Transformers Overcome Easy-to-Hard and Length Generalization Challenges
by: Lee, Nayoung, et al.
Published: (2025)
by: Lee, Nayoung, et al.
Published: (2025)
Looped Transformers are Better at Learning Learning Algorithms
by: Yang, Liu, et al.
Published: (2023)
by: Yang, Liu, et al.
Published: (2023)
Adaptive Sparse Möbius Transforms for Learning Polynomials
by: Erginbas, Yigit Efe, et al.
Published: (2026)
by: Erginbas, Yigit Efe, et al.
Published: (2026)
Toward a Theory of Tokenization in LLMs
by: Rajaraman, Nived, et al.
Published: (2024)
by: Rajaraman, Nived, et al.
Published: (2024)
Learning to Understand: Identifying Interactions via the Möbius Transform
by: Kang, Justin S., et al.
Published: (2024)
by: Kang, Justin S., et al.
Published: (2024)
The Fair Value of Data Under Heterogeneous Privacy Constraints in Federated Learning
by: Kang, Justin, et al.
Published: (2023)
by: Kang, Justin, et al.
Published: (2023)
Transformers on Markov Data: Constant Depth Suffices
by: Rajaraman, Nived, et al.
Published: (2024)
by: Rajaraman, Nived, et al.
Published: (2024)
Statistical Complexity and Optimal Algorithms for Non-linear Ridge Bandits
by: Rajaraman, Nived, et al.
Published: (2023)
by: Rajaraman, Nived, et al.
Published: (2023)
VersaPRM: Multi-Domain Process Reward Model via Synthetic Reasoning Data
by: Zeng, Thomas, et al.
Published: (2025)
by: Zeng, Thomas, et al.
Published: (2025)
Online Assortment and Price Optimization Under Contextual Choice Models
by: Erginbas, Yigit Efe, et al.
Published: (2025)
by: Erginbas, Yigit Efe, et al.
Published: (2025)
Dual Operating Modes of In-Context Learning
by: Lin, Ziqian, et al.
Published: (2024)
by: Lin, Ziqian, et al.
Published: (2024)
Towards Anytime-Valid Statistical Watermarking
by: Huang, Baihe, et al.
Published: (2026)
by: Huang, Baihe, et al.
Published: (2026)
Human-in-the-Loop Systems for Adaptive Learning Using Generative AI
by: Tarun, Bhavishya, et al.
Published: (2025)
by: Tarun, Bhavishya, et al.
Published: (2025)
An Odd Estimator for Shapley Values
by: Fumagalli, Fabian, et al.
Published: (2026)
by: Fumagalli, Fabian, et al.
Published: (2026)
Quantitative Bounds for Length Generalization in Transformers
by: Izzo, Zachary, et al.
Published: (2025)
by: Izzo, Zachary, et al.
Published: (2025)
The Expressive Power of Low-Rank Adaptation
by: Zeng, Yuchen, et al.
Published: (2023)
by: Zeng, Yuchen, et al.
Published: (2023)
EmbedLLM: Learning Compact Representations of Large Language Models
by: Zhuang, Richard, et al.
Published: (2024)
by: Zhuang, Richard, et al.
Published: (2024)
Any-Order Flexible Length Masked Diffusion
by: Kim, Jaeyeon, et al.
Published: (2025)
by: Kim, Jaeyeon, et al.
Published: (2025)
A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behavior
by: Mayne, Harry, et al.
Published: (2026)
by: Mayne, Harry, et al.
Published: (2026)
Fine-Tuning Without Forgetting In-Context Learning: A Theoretical Analysis of Linear Attention Models
by: Lee, Chungpa, et al.
Published: (2026)
by: Lee, Chungpa, et al.
Published: (2026)
In-Context Learning with Hypothesis-Class Guidance
by: Lin, Ziqian, et al.
Published: (2025)
by: Lin, Ziqian, et al.
Published: (2025)
Cross-Cluster Weighted Forests
by: Ramchandran, Maya, et al.
Published: (2021)
by: Ramchandran, Maya, et al.
Published: (2021)
ProxySPEX: Inference-Efficient Interpretability via Sparse Feature Interactions in LLMs
by: Butler, Landon, et al.
Published: (2025)
by: Butler, Landon, et al.
Published: (2025)
Towards Optimal Statistical Watermarking
by: Huang, Baihe, et al.
Published: (2023)
by: Huang, Baihe, et al.
Published: (2023)
Memorization Capacity for Additive Fine-Tuning with Small ReLU Networks
by: Sohn, Jy-yong, et al.
Published: (2024)
by: Sohn, Jy-yong, et al.
Published: (2024)
Transformers in the Dark: Navigating Unknown Search Spaces via Bandit Feedback
by: Kim, Jungtaek, et al.
Published: (2026)
by: Kim, Jungtaek, et al.
Published: (2026)
Sample Complexity and Representation Ability of Test-time Scaling Paradigms
by: Huang, Baihe, et al.
Published: (2025)
by: Huang, Baihe, et al.
Published: (2025)
Variation Spaces for Multi-Output Neural Networks: Insights on Multi-Task Learning and Network Compression
by: Shenouda, Joseph, et al.
Published: (2023)
by: Shenouda, Joseph, et al.
Published: (2023)
Contextual Feedback Loops: Amplifying Deep Reasoning with Iterative Top-Down Feedback
by: Fein-Ashley, Jacob, et al.
Published: (2024)
by: Fein-Ashley, Jacob, et al.
Published: (2024)
Equilibrium Matching: Generative Modeling with Implicit Energy-Based Models
by: Wang, Runqian, et al.
Published: (2025)
by: Wang, Runqian, et al.
Published: (2025)
Stability and Generalization in Looped Transformers
by: Labovich, Asher
Published: (2026)
by: Labovich, Asher
Published: (2026)
On Vanishing Variance in Transformer Length Generalization
by: Li, Ruining, et al.
Published: (2025)
by: Li, Ruining, et al.
Published: (2025)
Predictive Pipelined Decoding: A Compute-Latency Trade-off for Exact LLM Decoding
by: Yang, Seongjun, et al.
Published: (2023)
by: Yang, Seongjun, et al.
Published: (2023)
Task Vectors in In-Context Learning: Emergence, Formation, and Benefit
by: Yang, Liu, et al.
Published: (2025)
by: Yang, Liu, et al.
Published: (2025)
Arithmetic Transformers Can Length-Generalize in Both Operand Length and Count
by: Cho, Hanseul, et al.
Published: (2024)
by: Cho, Hanseul, et al.
Published: (2024)
From Markov to Laplace: How Mamba In-Context Learns Markov Chains
by: Bondaschi, Marco, et al.
Published: (2025)
by: Bondaschi, Marco, et al.
Published: (2025)
Compositional Generative Modeling: A Single Model is Not All You Need
by: Du, Yilun, et al.
Published: (2024)
by: Du, Yilun, et al.
Published: (2024)
Infected Smallville: How Disease Threat Shapes Sociality in LLM Agents
by: Choi, Soyeon, et al.
Published: (2025)
by: Choi, Soyeon, et al.
Published: (2025)
SPEX: Scaling Feature Interaction Explanations for LLMs
by: Kang, Justin Singh, et al.
Published: (2025)
by: Kang, Justin Singh, et al.
Published: (2025)
Similar Items
-
Optimizing DDPM Sampling with Shortcut Fine-Tuning
by: Fan, Ying, et al.
Published: (2023) -
Self-Improving Transformers Overcome Easy-to-Hard and Length Generalization Challenges
by: Lee, Nayoung, et al.
Published: (2025) -
Looped Transformers are Better at Learning Learning Algorithms
by: Yang, Liu, et al.
Published: (2023) -
Adaptive Sparse Möbius Transforms for Learning Polynomials
by: Erginbas, Yigit Efe, et al.
Published: (2026) -
Toward a Theory of Tokenization in LLMs
by: Rajaraman, Nived, et al.
Published: (2024)