Gespeichert in:
| Hauptverfasser: | Lin, Ziqian, Bharti, Shubham Kumar, Lee, Kangwook |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2502.19787 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Dual Operating Modes of In-Context Learning
von: Lin, Ziqian, et al.
Veröffentlicht: (2024)
von: Lin, Ziqian, et al.
Veröffentlicht: (2024)
Task Vectors in In-Context Learning: Emergence, Formation, and Benefit
von: Yang, Liu, et al.
Veröffentlicht: (2025)
von: Yang, Liu, et al.
Veröffentlicht: (2025)
Fine-Tuning Without Forgetting In-Context Learning: A Theoretical Analysis of Linear Attention Models
von: Lee, Chungpa, et al.
Veröffentlicht: (2026)
von: Lee, Chungpa, et al.
Veröffentlicht: (2026)
Can MLLMs Perform Text-to-Image In-Context Learning?
von: Zeng, Yuchen, et al.
Veröffentlicht: (2024)
von: Zeng, Yuchen, et al.
Veröffentlicht: (2024)
Transformers in the Dark: Navigating Unknown Search Spaces via Bandit Feedback
von: Kim, Jungtaek, et al.
Veröffentlicht: (2026)
von: Kim, Jungtaek, et al.
Veröffentlicht: (2026)
Optimizing DDPM Sampling with Shortcut Fine-Tuning
von: Fan, Ying, et al.
Veröffentlicht: (2023)
von: Fan, Ying, et al.
Veröffentlicht: (2023)
Can Mamba Learn How to Learn? A Comparative Study on In-Context Learning Tasks
von: Park, Jongho, et al.
Veröffentlicht: (2024)
von: Park, Jongho, et al.
Veröffentlicht: (2024)
The Expressive Power of Low-Rank Adaptation
von: Zeng, Yuchen, et al.
Veröffentlicht: (2023)
von: Zeng, Yuchen, et al.
Veröffentlicht: (2023)
Everything Everywhere All at Once: LLMs can In-Context Learn Multiple Tasks in Superposition
von: Xiong, Zheyang, et al.
Veröffentlicht: (2024)
von: Xiong, Zheyang, et al.
Veröffentlicht: (2024)
Looped Transformers are Better at Learning Learning Algorithms
von: Yang, Liu, et al.
Veröffentlicht: (2023)
von: Yang, Liu, et al.
Veröffentlicht: (2023)
Variation Spaces for Multi-Output Neural Networks: Insights on Multi-Task Learning and Network Compression
von: Shenouda, Joseph, et al.
Veröffentlicht: (2023)
von: Shenouda, Joseph, et al.
Veröffentlicht: (2023)
Rare-to-Frequent: Unlocking Compositional Generation Power of Diffusion Models on Rare Concepts with LLM Guidance
von: Park, Dongmin, et al.
Veröffentlicht: (2024)
von: Park, Dongmin, et al.
Veröffentlicht: (2024)
Looped Transformers for Length Generalization
von: Fan, Ying, et al.
Veröffentlicht: (2024)
von: Fan, Ying, et al.
Veröffentlicht: (2024)
Memorization Capacity for Additive Fine-Tuning with Small ReLU Networks
von: Sohn, Jy-yong, et al.
Veröffentlicht: (2024)
von: Sohn, Jy-yong, et al.
Veröffentlicht: (2024)
A Novel Data-Dependent Learning Paradigm for Large Hypothesis Classes
von: Pour, Alireza F., et al.
Veröffentlicht: (2025)
von: Pour, Alireza F., et al.
Veröffentlicht: (2025)
DS-AL: A Dual-Stream Analytic Learning for Exemplar-Free Class-Incremental Learning
von: Zhuang, Huiping, et al.
Veröffentlicht: (2024)
von: Zhuang, Huiping, et al.
Veröffentlicht: (2024)
ReJump: A Tree-Jump Representation for Analyzing and Improving LLM Reasoning
von: Zeng, Yuchen, et al.
Veröffentlicht: (2025)
von: Zeng, Yuchen, et al.
Veröffentlicht: (2025)
Leave-One-Out Prediction for General Hypothesis Classes
von: Qian, Jian, et al.
Veröffentlicht: (2026)
von: Qian, Jian, et al.
Veröffentlicht: (2026)
Self-Improving Transformers Overcome Easy-to-Hard and Length Generalization Challenges
von: Lee, Nayoung, et al.
Veröffentlicht: (2025)
von: Lee, Nayoung, et al.
Veröffentlicht: (2025)
Predictive Pipelined Decoding: A Compute-Latency Trade-off for Exact LLM Decoding
von: Yang, Seongjun, et al.
Veröffentlicht: (2023)
von: Yang, Seongjun, et al.
Veröffentlicht: (2023)
Smoothness Adaptive Hypothesis Transfer Learning
von: Lin, Haotian, et al.
Veröffentlicht: (2024)
von: Lin, Haotian, et al.
Veröffentlicht: (2024)
On Hypothesis Transfer Learning of Functional Linear Models
von: Lin, Haotian, et al.
Veröffentlicht: (2022)
von: Lin, Haotian, et al.
Veröffentlicht: (2022)
Infected Smallville: How Disease Threat Shapes Sociality in LLM Agents
von: Choi, Soyeon, et al.
Veröffentlicht: (2025)
von: Choi, Soyeon, et al.
Veröffentlicht: (2025)
From Artificial Needles to Real Haystacks: Improving Retrieval Capabilities in LLMs by Finetuning on Synthetic Data
von: Xiong, Zheyang, et al.
Veröffentlicht: (2024)
von: Xiong, Zheyang, et al.
Veröffentlicht: (2024)
What Do Language Models Learn in Context? The Structured Task Hypothesis
von: Li, Jiaoda, et al.
Veröffentlicht: (2024)
von: Li, Jiaoda, et al.
Veröffentlicht: (2024)
ENTP: Encoder-only Next Token Prediction
von: Ewer, Ethan, et al.
Veröffentlicht: (2024)
von: Ewer, Ethan, et al.
Veröffentlicht: (2024)
How to Correctly Report LLM-as-a-Judge Evaluations
von: Lee, Chungpa, et al.
Veröffentlicht: (2025)
von: Lee, Chungpa, et al.
Veröffentlicht: (2025)
Parameter-Efficient Fine-Tuning of State Space Models
von: Galim, Kevin, et al.
Veröffentlicht: (2024)
von: Galim, Kevin, et al.
Veröffentlicht: (2024)
Multi-Bin Batching for Increasing LLM Inference Throughput
von: Guldogan, Ozgur, et al.
Veröffentlicht: (2024)
von: Guldogan, Ozgur, et al.
Veröffentlicht: (2024)
Automated Type Annotation in Python Using Large Language Models
von: Bharti, Varun, et al.
Veröffentlicht: (2025)
von: Bharti, Varun, et al.
Veröffentlicht: (2025)
Hypothesis Class Determines Explanation: Why Accurate Models Disagree on Feature Attribution
von: B, Thackshanaramana
Veröffentlicht: (2026)
von: B, Thackshanaramana
Veröffentlicht: (2026)
Bayesian Active Learning in the Presence of Nuisance Parameters
von: Sloman, Sabina J., et al.
Veröffentlicht: (2023)
von: Sloman, Sabina J., et al.
Veröffentlicht: (2023)
Muon with Spectral Guidance: Efficient Optimization for Scientific Machine Learning
von: Lu, Binghang, et al.
Veröffentlicht: (2026)
von: Lu, Binghang, et al.
Veröffentlicht: (2026)
HTM-EAR: Importance-Preserving Tiered Memory with Hybrid Routing under Saturation
von: Singh, Shubham Kumar
Veröffentlicht: (2026)
von: Singh, Shubham Kumar
Veröffentlicht: (2026)
Rethinking Attention Output Projection: Structured Hadamard Transforms for Efficient Transformers
von: Aggarwal, Shubham, et al.
Veröffentlicht: (2026)
von: Aggarwal, Shubham, et al.
Veröffentlicht: (2026)
MixKVQ: Query-Aware Mixed-Precision KV Cache Quantization for Long-Context Reasoning
von: Zhang, Tao, et al.
Veröffentlicht: (2025)
von: Zhang, Tao, et al.
Veröffentlicht: (2025)
Quantization vs Pruning: Insights from the Strong Lottery Ticket Hypothesis
von: Kumar, Aakash, et al.
Veröffentlicht: (2025)
von: Kumar, Aakash, et al.
Veröffentlicht: (2025)
Hypothesis Spaces for Deep Learning
von: Wang, Rui, et al.
Veröffentlicht: (2024)
von: Wang, Rui, et al.
Veröffentlicht: (2024)
Multi-Task Corrupted Prediction for Learning Robust Audio-Visual Speech Representation
von: Kim, Sungnyun, et al.
Veröffentlicht: (2025)
von: Kim, Sungnyun, et al.
Veröffentlicht: (2025)
LookaheadKV: Fast and Accurate KV Cache Eviction by Glimpsing into the Future without Generation
von: Ahn, Jinwoo, et al.
Veröffentlicht: (2026)
von: Ahn, Jinwoo, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Dual Operating Modes of In-Context Learning
von: Lin, Ziqian, et al.
Veröffentlicht: (2024) -
Task Vectors in In-Context Learning: Emergence, Formation, and Benefit
von: Yang, Liu, et al.
Veröffentlicht: (2025) -
Fine-Tuning Without Forgetting In-Context Learning: A Theoretical Analysis of Linear Attention Models
von: Lee, Chungpa, et al.
Veröffentlicht: (2026) -
Can MLLMs Perform Text-to-Image In-Context Learning?
von: Zeng, Yuchen, et al.
Veröffentlicht: (2024) -
Transformers in the Dark: Navigating Unknown Search Spaces via Bandit Feedback
von: Kim, Jungtaek, et al.
Veröffentlicht: (2026)