OptiSeq: Ordering Examples On-The-Fly for In-Context Learning
Fuente:
arXiv
Salvato in:
| Autori principali: | Bhope, Rahul Atul, Venkateswaran, Praveen, Jayaram, K. R., Isahagian, Vatche, Muthusamy, Vinod, Venkatasubramanian, Nalini |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
VOILA: Value-of-Information Guided Fidelity Selection for Cost-Aware Multimodal Question Answering
di: Bhope, Rahul Atul, et al.
Pubblicazione: (2026)
di: Bhope, Rahul Atul, et al.
Pubblicazione: (2026)
Shift Happens: Mixture of Experts based Continual Adaptation in Federated Learning
di: Bhope, Rahul Atul, et al.
Pubblicazione: (2025)
di: Bhope, Rahul Atul, et al.
Pubblicazione: (2025)
FLOW-BENCH: Towards Conversational Generation of Enterprise Workflows
di: Duesterwald, Evelyn, et al.
Pubblicazione: (2025)
di: Duesterwald, Evelyn, et al.
Pubblicazione: (2025)
Trajectory-Informed Memory Generation for Self-Improving Agent Systems
di: Fang, Gaodan, et al.
Pubblicazione: (2026)
di: Fang, Gaodan, et al.
Pubblicazione: (2026)
On Automating Security Policies with Contemporary LLMs
di: Saura, Pablo Fernández, et al.
Pubblicazione: (2025)
di: Saura, Pablo Fernández, et al.
Pubblicazione: (2025)
BitDecoding: Unlocking Tensor Cores for Long-Context LLMs with Low-Bit KV Cache
di: Du, Dayou, et al.
Pubblicazione: (2025)
di: Du, Dayou, et al.
Pubblicazione: (2025)
Embedding Retrofitting: Data Engineering for better RAG
di: Sharma, Anantha
Pubblicazione: (2026)
di: Sharma, Anantha
Pubblicazione: (2026)
DyLLM: Efficient Diffusion LLM Inference via Saliency-based Token Selection and Partial Attention
di: Lee, Younjoo, et al.
Pubblicazione: (2026)
di: Lee, Younjoo, et al.
Pubblicazione: (2026)
AutoLALA: Automatic Loop Algebraic Locality Analysis for AI and HPC Kernels
di: Zhu, Yifan, et al.
Pubblicazione: (2026)
di: Zhu, Yifan, et al.
Pubblicazione: (2026)
Performance Characterization of Expert Router for Scalable LLM Inference
di: Pichlmeier, Josef, et al.
Pubblicazione: (2024)
di: Pichlmeier, Josef, et al.
Pubblicazione: (2024)
Harnessing Chain-of-Thought Metadata for Task Routing and Adversarial Prompt Detection
di: Marinelli, Ryan, et al.
Pubblicazione: (2025)
di: Marinelli, Ryan, et al.
Pubblicazione: (2025)
DeepSeek-R1 Outperforms Gemini 2.0 Pro, OpenAI o1, and o3-mini in Bilingual Complex Ophthalmology Reasoning
di: Xu, Pusheng, et al.
Pubblicazione: (2025)
di: Xu, Pusheng, et al.
Pubblicazione: (2025)
TeleEval-OS: Performance evaluations of large language models for operations scheduling
di: Wang, Yanyan, et al.
Pubblicazione: (2025)
di: Wang, Yanyan, et al.
Pubblicazione: (2025)
Hold Onto That Thought: Assessing KV Cache Compression On Reasoning
di: Liu, Minghui, et al.
Pubblicazione: (2025)
di: Liu, Minghui, et al.
Pubblicazione: (2025)
A Data-driven ML Approach for Maximizing Performance in LLM-Adapter Serving
di: Agullo, Ferran, et al.
Pubblicazione: (2025)
di: Agullo, Ferran, et al.
Pubblicazione: (2025)
DeepContext: A Context-aware, Cross-platform, and Cross-framework Tool for Performance Profiling and Analysis of Deep Learning Workloads
di: Zhao, Qidong, et al.
Pubblicazione: (2024)
di: Zhao, Qidong, et al.
Pubblicazione: (2024)
DIM-SUM: Dynamic IMputation for Smart Utility Management
di: Hildebrant, Ryan, et al.
Pubblicazione: (2025)
di: Hildebrant, Ryan, et al.
Pubblicazione: (2025)
SuperCoder: Assembly Program Superoptimization with Large Language Models
di: Wei, Anjiang, et al.
Pubblicazione: (2025)
di: Wei, Anjiang, et al.
Pubblicazione: (2025)
Protean Compiler: An Agile Framework to Drive Fine-grain Phase Ordering
di: Ashouri, Amir H., et al.
Pubblicazione: (2026)
di: Ashouri, Amir H., et al.
Pubblicazione: (2026)
Optimization of Armv9 architecture general large language model inference performance based on Llama.cpp
di: Chen, Longhao, et al.
Pubblicazione: (2024)
di: Chen, Longhao, et al.
Pubblicazione: (2024)
Quamba2: A Robust and Scalable Post-training Quantization Framework for Selective State Space Models
di: Chiang, Hung-Yueh, et al.
Pubblicazione: (2025)
di: Chiang, Hung-Yueh, et al.
Pubblicazione: (2025)
Energy-Aware LLMs: A step towards sustainable AI for downstream applications
di: Tran, Nguyen Phuc, et al.
Pubblicazione: (2025)
di: Tran, Nguyen Phuc, et al.
Pubblicazione: (2025)
Investigating Execution-Aware Language Models for Code Optimization
di: Di Menna, Federico, et al.
Pubblicazione: (2025)
di: Di Menna, Federico, et al.
Pubblicazione: (2025)
QServe: W4A8KV4 Quantization and System Co-design for Efficient LLM Serving
di: Lin, Yujun, et al.
Pubblicazione: (2024)
di: Lin, Yujun, et al.
Pubblicazione: (2024)
REAM: Merging Improves Pruning of Experts in LLMs
di: Jha, Saurav, et al.
Pubblicazione: (2026)
di: Jha, Saurav, et al.
Pubblicazione: (2026)
QJL: 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead
di: Zandieh, Amir, et al.
Pubblicazione: (2024)
di: Zandieh, Amir, et al.
Pubblicazione: (2024)
Evaluating the Efficacy of Foundational Models: Advancing Benchmarking Practices to Enhance Fine-Tuning Decision-Making
di: Amujo, Oluyemi Enoch, et al.
Pubblicazione: (2024)
di: Amujo, Oluyemi Enoch, et al.
Pubblicazione: (2024)
Data Efficacy for Language Model Training
di: Dai, Yalun, et al.
Pubblicazione: (2025)
di: Dai, Yalun, et al.
Pubblicazione: (2025)
Bench360: Benchmarking Local LLM Inference from 360 Degrees
di: Stuhlmann, Linus, et al.
Pubblicazione: (2025)
di: Stuhlmann, Linus, et al.
Pubblicazione: (2025)
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing
di: Zheng, Wenhao, et al.
Pubblicazione: (2025)
di: Zheng, Wenhao, et al.
Pubblicazione: (2025)
Morpheme Boundary Detection & Grammatical Feature Prediction for Gujarati : Dataset & Model
di: Baxi, Jatayu, et al.
Pubblicazione: (2021)
di: Baxi, Jatayu, et al.
Pubblicazione: (2021)
An energy-based comparative analysis of common approaches to text classification in the Legal domain
di: Gultekin, Sinan, et al.
Pubblicazione: (2023)
di: Gultekin, Sinan, et al.
Pubblicazione: (2023)
ECCO: Evidence-Driven Causal Reasoning for Compiler Optimization
di: Pan, Haolin, et al.
Pubblicazione: (2026)
di: Pan, Haolin, et al.
Pubblicazione: (2026)
MMLU-CF: A Contamination-free Multi-task Language Understanding Benchmark
di: Zhao, Qihao, et al.
Pubblicazione: (2024)
di: Zhao, Qihao, et al.
Pubblicazione: (2024)
Accelerating Diffusion LLMs via Adaptive Parallel Decoding
di: Israel, Daniel, et al.
Pubblicazione: (2025)
di: Israel, Daniel, et al.
Pubblicazione: (2025)
Model Compression and Efficient Inference for Large Language Models: A Survey
di: Wang, Wenxiao, et al.
Pubblicazione: (2024)
di: Wang, Wenxiao, et al.
Pubblicazione: (2024)
Enhancing Inference Efficiency of Large Language Models: Investigating Optimization Strategies and Architectural Innovations
di: Tyukin, Georgy
Pubblicazione: (2024)
di: Tyukin, Georgy
Pubblicazione: (2024)
Systematic Evaluation of Optimization Techniques for Long-Context Language Models
di: Ahmed, Ammar, et al.
Pubblicazione: (2025)
di: Ahmed, Ammar, et al.
Pubblicazione: (2025)
Learning, Potential, and Retention: An Approach for Evaluating Adaptive AI-Enabled Medical Devices
di: Burgon, Alexis, et al.
Pubblicazione: (2026)
di: Burgon, Alexis, et al.
Pubblicazione: (2026)
Performance of Confidential Computing GPUs
di: Ibarra, Antonio Martínez, et al.
Pubblicazione: (2025)
di: Ibarra, Antonio Martínez, et al.
Pubblicazione: (2025)
Documenti analoghi
-
VOILA: Value-of-Information Guided Fidelity Selection for Cost-Aware Multimodal Question Answering
di: Bhope, Rahul Atul, et al.
Pubblicazione: (2026) -
Shift Happens: Mixture of Experts based Continual Adaptation in Federated Learning
di: Bhope, Rahul Atul, et al.
Pubblicazione: (2025) -
FLOW-BENCH: Towards Conversational Generation of Enterprise Workflows
di: Duesterwald, Evelyn, et al.
Pubblicazione: (2025) -
Trajectory-Informed Memory Generation for Self-Improving Agent Systems
di: Fang, Gaodan, et al.
Pubblicazione: (2026) -
On Automating Security Policies with Contemporary LLMs
di: Saura, Pablo Fernández, et al.
Pubblicazione: (2025)