MarginSel : Max-Margin Demonstration Selection for LLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Ambati, Rajeev Bhatt, Lester, James, Srivastava, Shashank, Chaturvedi, Snigdha |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Optimizing Attention with Mirror Descent: Generalized Max-Margin Token Selection
by: Julistiono, Addison Kristanto, et al.
Published: (2024)
by: Julistiono, Addison Kristanto, et al.
Published: (2024)
Meta-Sel: Efficient Demonstration Selection for In-Context Learning via Supervised Meta-Learning
by: Wang, Xubin, et al.
Published: (2026)
by: Wang, Xubin, et al.
Published: (2026)
Socratic Students: Teaching Language Models to Learn by Asking Questions
by: Ambati, Rajeev Bhatt, et al.
Published: (2025)
by: Ambati, Rajeev Bhatt, et al.
Published: (2025)
Larger or Smaller Reward Margins to Select Preferences for Alignment?
by: Huang, Kexin, et al.
Published: (2025)
by: Huang, Kexin, et al.
Published: (2025)
Fuse to Forget: Bias Reduction and Selective Memorization through Model Fusion
by: Zaman, Kerem, et al.
Published: (2023)
by: Zaman, Kerem, et al.
Published: (2023)
Is Chain-of-Thought Really Not Explainability? Chain-of-Thought Can Be Faithful without Hint Verbalization
by: Zaman, Kerem, et al.
Published: (2025)
by: Zaman, Kerem, et al.
Published: (2025)
Adaptive Margin RLHF via Preference over Preferences
by: Chittepu, Yaswanth, et al.
Published: (2025)
by: Chittepu, Yaswanth, et al.
Published: (2025)
A Causal Lens for Evaluating Faithfulness Metrics
by: Zaman, Kerem, et al.
Published: (2025)
by: Zaman, Kerem, et al.
Published: (2025)
AlphaDPO: Adaptive Reward Margin for Direct Preference Optimization
by: Wu, Junkang, et al.
Published: (2024)
by: Wu, Junkang, et al.
Published: (2024)
Steering LLMs for Formal Theorem Proving
by: Kirtania, Shashank, et al.
Published: (2025)
by: Kirtania, Shashank, et al.
Published: (2025)
Affective and Dynamic Beam Search for Story Generation
by: Huang, Tenghao, et al.
Published: (2023)
by: Huang, Tenghao, et al.
Published: (2023)
A Common Pitfall of Margin-based Language Model Alignment: Gradient Entanglement
by: Yuan, Hui, et al.
Published: (2024)
by: Yuan, Hui, et al.
Published: (2024)
Geometry-Lite: Interpretable Safety Probing via Layer-Wise Margin Geometry
by: Sim, Woo Seob, et al.
Published: (2026)
by: Sim, Woo Seob, et al.
Published: (2026)
Unraveling the Mechanics of Learning-Based Demonstration Selection for In-Context Learning
by: Liu, Hui, et al.
Published: (2024)
by: Liu, Hui, et al.
Published: (2024)
Comparative Analysis of Demonstration Selection Algorithms for LLM In-Context Learning
by: Shu, Dong, et al.
Published: (2024)
by: Shu, Dong, et al.
Published: (2024)
Linear-Time Demonstration Selection for In-Context Learning via Gradient Estimation
by: Zhang, Ziniu, et al.
Published: (2025)
by: Zhang, Ziniu, et al.
Published: (2025)
LLMem: Estimating GPU Memory Usage for Fine-Tuning Pre-Trained LLMs
by: Kim, Taeho, et al.
Published: (2024)
by: Kim, Taeho, et al.
Published: (2024)
Affinity and Diversity: A Unified Metric for Demonstration Selection via Internal Representations
by: Kato, Mariko, et al.
Published: (2025)
by: Kato, Mariko, et al.
Published: (2025)
Optimizing LLMs for Resource-Constrained Environments: A Survey of Model Compression Techniques
by: Girija, Sanjay Surendranath, et al.
Published: (2025)
by: Girija, Sanjay Surendranath, et al.
Published: (2025)
Dynamic Context Evolution for Scalable Synthetic Data Generation
by: Lingo, Ryan, et al.
Published: (2026)
by: Lingo, Ryan, et al.
Published: (2026)
Selective Prompting Tuning for Personalized Conversations with LLMs
by: Huang, Qiushi, et al.
Published: (2024)
by: Huang, Qiushi, et al.
Published: (2024)
SelectIT: Selective Instruction Tuning for LLMs via Uncertainty-Aware Self-Reflection
by: Liu, Liangxin, et al.
Published: (2024)
by: Liu, Liangxin, et al.
Published: (2024)
The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence
by: MiniMax, et al.
Published: (2026)
by: MiniMax, et al.
Published: (2026)
Forgetting-MarI: LLM Unlearning via Marginal Information Regularization
by: Xu, Shizhou, et al.
Published: (2025)
by: Xu, Shizhou, et al.
Published: (2025)
Generation and De-Identification of Indian Clinical Discharge Summaries using LLMs
by: Singh, Sanjeet, et al.
Published: (2024)
by: Singh, Sanjeet, et al.
Published: (2024)
Enhancing Domain-Specific Retrieval-Augmented Generation: Synthetic Data Generation and Evaluation using Reasoning Models
by: Jadon, Aryan, et al.
Published: (2025)
by: Jadon, Aryan, et al.
Published: (2025)
Point of Order: Action-Aware LLM Persona Modeling for Realistic Civic Simulation
by: Merrill, Scott, et al.
Published: (2025)
by: Merrill, Scott, et al.
Published: (2025)
Recover-LoRA: Data-Free Accuracy Recovery of Degraded Language Models via Low-Rank Adaptation
by: Das, Devleena, et al.
Published: (2025)
by: Das, Devleena, et al.
Published: (2025)
Enhancing LLM Problem Solving with REAP: Reflection, Explicit Problem Deconstruction, and Advanced Prompting
by: Lingo, Ryan, et al.
Published: (2024)
by: Lingo, Ryan, et al.
Published: (2024)
Geometry of Decision Making in Language Models
by: Joshi, Abhinav, et al.
Published: (2025)
by: Joshi, Abhinav, et al.
Published: (2025)
Hypertokens: Holographic Associative Memory in Tokenized LLMs
by: Augeri, Christopher James
Published: (2025)
by: Augeri, Christopher James
Published: (2025)
Easy Problems That LLMs Get Wrong
by: Williams, Sean, et al.
Published: (2024)
by: Williams, Sean, et al.
Published: (2024)
Max It or Miss It: Benchmarking LLM On Solving Extremal Problems
by: Gao, Binxin, et al.
Published: (2025)
by: Gao, Binxin, et al.
Published: (2025)
Morpheme Boundary Detection & Grammatical Feature Prediction for Gujarati : Dataset & Model
by: Baxi, Jatayu, et al.
Published: (2021)
by: Baxi, Jatayu, et al.
Published: (2021)
DemoShapley: Valuation of Demonstrations for In-Context Learning
by: Xie, Shan, et al.
Published: (2024)
by: Xie, Shan, et al.
Published: (2024)
Reinforcement Learning for Diffusion LLMs with Entropy-Guided Step Selection and Stepwise Advantages
by: Kunde, Vishnu Teja, et al.
Published: (2026)
by: Kunde, Vishnu Teja, et al.
Published: (2026)
A Case Study of Selected PTQ Baselines for Reasoning LLMs on Ascend NPU
by: Luo, Yuchen, et al.
Published: (2026)
by: Luo, Yuchen, et al.
Published: (2026)
TokenSelect: Efficient Long-Context Inference and Length Extrapolation for LLMs via Dynamic Token-Level KV Cache Selection
by: Wu, Wei, et al.
Published: (2024)
by: Wu, Wei, et al.
Published: (2024)
Evolving Demonstration Optimization for Chain-of-Thought Feature Transformation
by: Wang, Xinyuan, et al.
Published: (2026)
by: Wang, Xinyuan, et al.
Published: (2026)
ClusterUCB: Efficient Gradient-Based Data Selection for Targeted Fine-Tuning of LLMs
by: Wang, Zige, et al.
Published: (2025)
by: Wang, Zige, et al.
Published: (2025)
Similar Items
-
Optimizing Attention with Mirror Descent: Generalized Max-Margin Token Selection
by: Julistiono, Addison Kristanto, et al.
Published: (2024) -
Meta-Sel: Efficient Demonstration Selection for In-Context Learning via Supervised Meta-Learning
by: Wang, Xubin, et al.
Published: (2026) -
Socratic Students: Teaching Language Models to Learn by Asking Questions
by: Ambati, Rajeev Bhatt, et al.
Published: (2025) -
Larger or Smaller Reward Margins to Select Preferences for Alignment?
by: Huang, Kexin, et al.
Published: (2025) -
Fuse to Forget: Bias Reduction and Selective Memorization through Model Fusion
by: Zaman, Kerem, et al.
Published: (2023)