Saved in:
| Main Authors: | Park, Jongho, Park, Jaeseung, Xiong, Zheyang, Lee, Nayoung, Cho, Jaewoong, Oymak, Samet, Lee, Kangwook, Papailiopoulos, Dimitris |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2402.04248 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries
by: Kim, Junhyuck, et al.
Published: (2024)
by: Kim, Junhyuck, et al.
Published: (2024)
From Artificial Needles to Real Haystacks: Improving Retrieval Capabilities in LLMs by Finetuning on Synthetic Data
by: Xiong, Zheyang, et al.
Published: (2024)
by: Xiong, Zheyang, et al.
Published: (2024)
Predictive Pipelined Decoding: A Compute-Latency Trade-off for Exact LLM Decoding
by: Yang, Seongjun, et al.
Published: (2023)
by: Yang, Seongjun, et al.
Published: (2023)
Task Vectors in In-Context Learning: Emergence, Formation, and Benefit
by: Yang, Liu, et al.
Published: (2025)
by: Yang, Liu, et al.
Published: (2025)
Everything Everywhere All at Once: LLMs can In-Context Learn Multiple Tasks in Superposition
by: Xiong, Zheyang, et al.
Published: (2024)
by: Xiong, Zheyang, et al.
Published: (2024)
Self-Improving Transformers Overcome Easy-to-Hard and Length Generalization Challenges
by: Lee, Nayoung, et al.
Published: (2025)
by: Lee, Nayoung, et al.
Published: (2025)
Looped Transformers are Better at Learning Learning Algorithms
by: Yang, Liu, et al.
Published: (2023)
by: Yang, Liu, et al.
Published: (2023)
Extrapolation by Association: Length Generalization Transfer in Transformers
by: Cai, Ziyang, et al.
Published: (2025)
by: Cai, Ziyang, et al.
Published: (2025)
Not All Bits Are Equal: Scale-Dependent Memory Optimization Strategies for Reasoning Models
by: Kim, Junhyuck, et al.
Published: (2025)
by: Kim, Junhyuck, et al.
Published: (2025)
In-Context Learning Under Regime Change
by: Dudley, Carson, et al.
Published: (2026)
by: Dudley, Carson, et al.
Published: (2026)
Dual Operating Modes of In-Context Learning
by: Lin, Ziqian, et al.
Published: (2024)
by: Lin, Ziqian, et al.
Published: (2024)
Can Transformers Learn Optimal Filtering for Unknown Systems?
by: Balim, Haldun, et al.
Published: (2023)
by: Balim, Haldun, et al.
Published: (2023)
Can Large Language Models Develop Strategic Reasoning? Post-training Insights from Learning Chess
by: Hwang, Dongyoon, et al.
Published: (2025)
by: Hwang, Dongyoon, et al.
Published: (2025)
Task Diversity Shortens the ICL Plateau
by: Kim, Jaeyeon, et al.
Published: (2024)
by: Kim, Jaeyeon, et al.
Published: (2024)
Rare-to-Frequent: Unlocking Compositional Generation Power of Diffusion Models on Rare Concepts with LLM Guidance
by: Park, Dongmin, et al.
Published: (2024)
by: Park, Dongmin, et al.
Published: (2024)
When and How Unlabeled Data Provably Improve In-Context Learning
by: Li, Yingcong, et al.
Published: (2025)
by: Li, Yingcong, et al.
Published: (2025)
In-Context Learning with Hypothesis-Class Guidance
by: Lin, Ziqian, et al.
Published: (2025)
by: Lin, Ziqian, et al.
Published: (2025)
Can MLLMs Perform Text-to-Image In-Context Learning?
by: Zeng, Yuchen, et al.
Published: (2024)
by: Zeng, Yuchen, et al.
Published: (2024)
How Well Can Transformers Emulate In-context Newton's Method?
by: Giannou, Angeliki, et al.
Published: (2024)
by: Giannou, Angeliki, et al.
Published: (2024)
Learning to Bet for Horizon-Aware Anytime-Valid Testing
by: Taga, Ege Onur, et al.
Published: (2026)
by: Taga, Ege Onur, et al.
Published: (2026)
Fine-Tuning Without Forgetting In-Context Learning: A Theoretical Analysis of Linear Attention Models
by: Lee, Chungpa, et al.
Published: (2026)
by: Lee, Chungpa, et al.
Published: (2026)
Latent Chain-of-Thought Improves Structured-Data Transformers
by: Dudley, Carson, et al.
Published: (2026)
by: Dudley, Carson, et al.
Published: (2026)
Image Clustering Conditioned on Text Criteria
by: Kwon, Sehyun, et al.
Published: (2023)
by: Kwon, Sehyun, et al.
Published: (2023)
Generative Quantum Data Embeddings for Supervised Learning
by: Heo, Jaewoong, et al.
Published: (2026)
by: Heo, Jaewoong, et al.
Published: (2026)
Learning to Correct: Calibrated Reinforcement Learning for Multi-Attempt Chain-of-Thought
by: Ildiz, Muhammed Emrullah, et al.
Published: (2026)
by: Ildiz, Muhammed Emrullah, et al.
Published: (2026)
A Lightweight CNN-Transformer Model for Learning Traveling Salesman Problems
by: Jung, Minseop, et al.
Published: (2023)
by: Jung, Minseop, et al.
Published: (2023)
SAiD: Speech-driven Blendshape Facial Animation with Diffusion
by: Park, Inkyu, et al.
Published: (2023)
by: Park, Inkyu, et al.
Published: (2023)
How Data Mixing Shapes In-Context Learning: Asymptotic Equivalence for Transformers with MLPs
by: Demir, Samet, et al.
Published: (2025)
by: Demir, Samet, et al.
Published: (2025)
Variation Spaces for Multi-Output Neural Networks: Insights on Multi-Task Learning and Network Compression
by: Shenouda, Joseph, et al.
Published: (2023)
by: Shenouda, Joseph, et al.
Published: (2023)
Beyond RLHF: A Unified Theoretical Framework of Alignment
by: Yun, Jihun, et al.
Published: (2025)
by: Yun, Jihun, et al.
Published: (2025)
MX+: Pushing the Limits of Microscaling Formats for Efficient Large Language Model Serving
by: Lee, Jungi, et al.
Published: (2025)
by: Lee, Jungi, et al.
Published: (2025)
Parareal Neural Networks Emulating a Parallel-in-time Algorithm
by: Lee, Chang-Ock, et al.
Published: (2021)
by: Lee, Chang-Ock, et al.
Published: (2021)
Is Prompt Selection Necessary for Task-Free Online Continual Learning?
by: Park, Seoyoung, et al.
Published: (2026)
by: Park, Seoyoung, et al.
Published: (2026)
Lookahead Unmasking Elicits Accurate Decoding in Diffusion Language Models
by: Lee, Sanghyun, et al.
Published: (2025)
by: Lee, Sanghyun, et al.
Published: (2025)
Transductive Generalization via Optimal Transport and Its Application to Graph Node Classification
by: Park, MoonJeong, et al.
Published: (2026)
by: Park, MoonJeong, et al.
Published: (2026)
DualFL: A Duality-based Federated Learning Algorithm with Communication Acceleration in the General Convex Regime
by: Park, Jongho, et al.
Published: (2023)
by: Park, Jongho, et al.
Published: (2023)
Learning a Patent-Informed Biomedical Knowledge Graph Reveals Technological Potential of Drug Repositioning Candidates
by: Jegal, Yongseung, et al.
Published: (2023)
by: Jegal, Yongseung, et al.
Published: (2023)
Can Mamba Learn In Context with Outliers? A Theoretical Generalization Analysis
by: Li, Hongkang, et al.
Published: (2025)
by: Li, Hongkang, et al.
Published: (2025)
Neural Functions for Learning Periodic Signal
by: Cho, Woojin, et al.
Published: (2025)
by: Cho, Woojin, et al.
Published: (2025)
Multi-Task Corrupted Prediction for Learning Robust Audio-Visual Speech Representation
by: Kim, Sungnyun, et al.
Published: (2025)
by: Kim, Sungnyun, et al.
Published: (2025)
Similar Items
-
Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries
by: Kim, Junhyuck, et al.
Published: (2024) -
From Artificial Needles to Real Haystacks: Improving Retrieval Capabilities in LLMs by Finetuning on Synthetic Data
by: Xiong, Zheyang, et al.
Published: (2024) -
Predictive Pipelined Decoding: A Compute-Latency Trade-off for Exact LLM Decoding
by: Yang, Seongjun, et al.
Published: (2023) -
Task Vectors in In-Context Learning: Emergence, Formation, and Benefit
by: Yang, Liu, et al.
Published: (2025) -
Everything Everywhere All at Once: LLMs can In-Context Learn Multiple Tasks in Superposition
by: Xiong, Zheyang, et al.
Published: (2024)