Task Diversity Shortens the ICL Plateau
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kim, Jaeyeon, Kwon, Sehyun, Choi, Joo Young, Park, Jongho, Cho, Jaewoong, Lee, Jason D., Ryu, Ernest K. |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Simple Drop-in LoRA Conditioning on Attention Layers Will Improve Your Diffusion Model
von: Choi, Joo Young, et al.
Veröffentlicht: (2024)
von: Choi, Joo Young, et al.
Veröffentlicht: (2024)
Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries
von: Kim, Junhyuck, et al.
Veröffentlicht: (2024)
von: Kim, Junhyuck, et al.
Veröffentlicht: (2024)
LoRA Training Provably Converges to a Low-Rank Global Minimum or It Fails Loudly (But it Probably Won't Fail)
von: Kim, Junsu, et al.
Veröffentlicht: (2025)
von: Kim, Junsu, et al.
Veröffentlicht: (2025)
Cross-Architecture Transfer Learning for Linear-Cost Inference Transformers
von: Choi, Sehyun
Veröffentlicht: (2024)
von: Choi, Sehyun
Veröffentlicht: (2024)
Clip-Low Increases Entropy and Clip-High Decreases Entropy in Reinforcement Learning of Large Language Models
von: Park, Jaesung R., et al.
Veröffentlicht: (2025)
von: Park, Jaesung R., et al.
Veröffentlicht: (2025)
Break the Breakout: Reinventing LM Defense Against Jailbreak Attacks with Self-Refinement
von: Kim, Heegyu, et al.
Veröffentlicht: (2024)
von: Kim, Heegyu, et al.
Veröffentlicht: (2024)
Can Mamba Learn How to Learn? A Comparative Study on In-Context Learning Tasks
von: Park, Jongho, et al.
Veröffentlicht: (2024)
von: Park, Jongho, et al.
Veröffentlicht: (2024)
Image Clustering Conditioned on Text Criteria
von: Kwon, Sehyun, et al.
Veröffentlicht: (2023)
von: Kwon, Sehyun, et al.
Veröffentlicht: (2023)
Beyond RLHF: A Unified Theoretical Framework of Alignment
von: Yun, Jihun, et al.
Veröffentlicht: (2025)
von: Yun, Jihun, et al.
Veröffentlicht: (2025)
Shortened LLaMA: Depth Pruning for Large Language Models with Comparison of Retraining Methods
von: Kim, Bo-Kyeong, et al.
Veröffentlicht: (2024)
von: Kim, Bo-Kyeong, et al.
Veröffentlicht: (2024)
CHILL at SemEval-2025 Task 2: You Can't Just Throw Entities and Hope -- Make Your LLM to Get Them Right
von: Lee, Jaebok, et al.
Veröffentlicht: (2025)
von: Lee, Jaebok, et al.
Veröffentlicht: (2025)
Predictive Pipelined Decoding: A Compute-Latency Trade-off for Exact LLM Decoding
von: Yang, Seongjun, et al.
Veröffentlicht: (2023)
von: Yang, Seongjun, et al.
Veröffentlicht: (2023)
Rare-to-Frequent: Unlocking Compositional Generation Power of Diffusion Models on Rare Concepts with LLM Guidance
von: Park, Dongmin, et al.
Veröffentlicht: (2024)
von: Park, Dongmin, et al.
Veröffentlicht: (2024)
Overcoming Spurious Solutions in Semi-Dual Neural Optimal Transport: A Smoothing Approach for Learning the Optimal Transport Plan
von: Choi, Jaemoo, et al.
Veröffentlicht: (2025)
von: Choi, Jaemoo, et al.
Veröffentlicht: (2025)
Stress-Testing Long-Context Language Models with Lifelong ICL and Task Haystack
von: Xu, Xiaoyue, et al.
Veröffentlicht: (2024)
von: Xu, Xiaoyue, et al.
Veröffentlicht: (2024)
Fast and Accurate Neural Rendering Using Semi-Gradients
von: Cho, In-Young, et al.
Veröffentlicht: (2024)
von: Cho, In-Young, et al.
Veröffentlicht: (2024)
Pruning and Distilling Mixture-of-Experts into Dense Language Models
von: Kim, Junhyuck, et al.
Veröffentlicht: (2026)
von: Kim, Junhyuck, et al.
Veröffentlicht: (2026)
Bridging the Gap between Expert and Language Models: Concept-guided Chess Commentary Generation and Evaluation
von: Kim, Jaechang, et al.
Veröffentlicht: (2024)
von: Kim, Jaechang, et al.
Veröffentlicht: (2024)
QUICK: Quantization-aware Interleaving and Conflict-free Kernel for efficient LLM inference
von: Kim, Taesu, et al.
Veröffentlicht: (2024)
von: Kim, Taesu, et al.
Veröffentlicht: (2024)
DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors
von: Lee, Keon, et al.
Veröffentlicht: (2024)
von: Lee, Keon, et al.
Veröffentlicht: (2024)
Learning a Patent-Informed Biomedical Knowledge Graph Reveals Technological Potential of Drug Repositioning Candidates
von: Jegal, Yongseung, et al.
Veröffentlicht: (2023)
von: Jegal, Yongseung, et al.
Veröffentlicht: (2023)
Adaptive Task Vectors for Large Language Models
von: Kang, Joonseong, et al.
Veröffentlicht: (2025)
von: Kang, Joonseong, et al.
Veröffentlicht: (2025)
Rate-Optimal Noise Annealing in Semi-Dual Neural Optimal Transport: Tangential Identifiability, Off-Manifold Ambiguity, and Guaranteed Recovery
von: Chu, Raymond, et al.
Veröffentlicht: (2026)
von: Chu, Raymond, et al.
Veröffentlicht: (2026)
Soft Head Selection for Injecting ICL-Derived Task Embeddings
von: Park, Jungwon, et al.
Veröffentlicht: (2025)
von: Park, Jungwon, et al.
Veröffentlicht: (2025)
Batch-ICL: Effective, Efficient, and Order-Agnostic In-Context Learning
von: Zhang, Kaiyi, et al.
Veröffentlicht: (2024)
von: Zhang, Kaiyi, et al.
Veröffentlicht: (2024)
Merging Methods for Multilingual Knowledge Editing for Large Language Models: An Empirical Odyssey
von: Lee, Kunil, et al.
Veröffentlicht: (2026)
von: Lee, Kunil, et al.
Veröffentlicht: (2026)
Rethinking State Tracking in Recurrent Models Through Error Control Dynamics
von: Chung, Jiwan, et al.
Veröffentlicht: (2026)
von: Chung, Jiwan, et al.
Veröffentlicht: (2026)
LoRA Training in the NTK Regime has No Spurious Local Minima
von: Jang, Uijeong, et al.
Veröffentlicht: (2024)
von: Jang, Uijeong, et al.
Veröffentlicht: (2024)
Sequence Shortening for Context-Aware Machine Translation
von: Mąka, Paweł, et al.
Veröffentlicht: (2024)
von: Mąka, Paweł, et al.
Veröffentlicht: (2024)
MAGE: All-[MASK] Block Already Knows Where to Look in Diffusion LLM
von: Kwon, Omin, et al.
Veröffentlicht: (2026)
von: Kwon, Omin, et al.
Veröffentlicht: (2026)
Auto-ICL: In-Context Learning without Human Supervision
von: Yang, Jinghan, et al.
Veröffentlicht: (2023)
von: Yang, Jinghan, et al.
Veröffentlicht: (2023)
Generative AI Meets Semantic Communication: Evolution and Revolution of Communication Tasks
von: Grassucci, Eleonora, et al.
Veröffentlicht: (2024)
von: Grassucci, Eleonora, et al.
Veröffentlicht: (2024)
Legal2LogicICL: Improving Generalization in Transforming Legal Cases to Logical Formulas via Diverse Few-Shot Learning
von: Xue, Jieying, et al.
Veröffentlicht: (2026)
von: Xue, Jieying, et al.
Veröffentlicht: (2026)
Finite-Time Bounds for Average-Reward Fitted Q-Iteration
von: Lee, Jongmin, et al.
Veröffentlicht: (2025)
von: Lee, Jongmin, et al.
Veröffentlicht: (2025)
Why Policy Gradient Algorithms Work for Undiscounted Total-Reward MDPs
von: Lee, Jongmin, et al.
Veröffentlicht: (2025)
von: Lee, Jongmin, et al.
Veröffentlicht: (2025)
RetICL: Sequential Retrieval of In-Context Examples with Reinforcement Learning
von: Scarlatos, Alexander, et al.
Veröffentlicht: (2023)
von: Scarlatos, Alexander, et al.
Veröffentlicht: (2023)
Expanding Foundational Language Capabilities in Open-Source LLMs through a Korean Case Study
von: Lim, Junghwan, et al.
Veröffentlicht: (2025)
von: Lim, Junghwan, et al.
Veröffentlicht: (2025)
Neural Optimal Transport in Hilbert Spaces: Characterizing Spurious Solutions and Gaussian Smoothing
von: Choi, Jae-Hwan, et al.
Veröffentlicht: (2026)
von: Choi, Jae-Hwan, et al.
Veröffentlicht: (2026)
Efficient Generative Modeling with Residual Vector Quantization-Based Tokens
von: Kim, Jaehyeon, et al.
Veröffentlicht: (2024)
von: Kim, Jaehyeon, et al.
Veröffentlicht: (2024)
Culinary Class Wars: Evaluating LLMs using ASH in Cuisine Transfer Task
von: Lee, Hoonick, et al.
Veröffentlicht: (2024)
von: Lee, Hoonick, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Simple Drop-in LoRA Conditioning on Attention Layers Will Improve Your Diffusion Model
von: Choi, Joo Young, et al.
Veröffentlicht: (2024) -
Lexico: Extreme KV Cache Compression via Sparse Coding over Universal Dictionaries
von: Kim, Junhyuck, et al.
Veröffentlicht: (2024) -
LoRA Training Provably Converges to a Low-Rank Global Minimum or It Fails Loudly (But it Probably Won't Fail)
von: Kim, Junsu, et al.
Veröffentlicht: (2025) -
Cross-Architecture Transfer Learning for Linear-Cost Inference Transformers
von: Choi, Sehyun
Veröffentlicht: (2024) -
Clip-Low Increases Entropy and Clip-High Decreases Entropy in Reinforcement Learning of Large Language Models
von: Park, Jaesung R., et al.
Veröffentlicht: (2025)