BASS: Batched Attention-optimized Speculative Sampling
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Qian, Haifeng, Gonugondla, Sujan Kumar, Ha, Sungsoo, Shang, Mingyue, Gouda, Sanjay Krishna, Nallapati, Ramesh, Sengupta, Sudipta, Ma, Xiaofei, Deoras, Anoop |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Bifurcated Attention: Accelerating Massively Parallel Decoding with Shared Prefixes in LLMs
par: Athiwaratkun, Ben, et autres
Publié: (2024)
par: Athiwaratkun, Ben, et autres
Publié: (2024)
Token Alignment via Character Matching for Subword Completion
par: Athiwaratkun, Ben, et autres
Publié: (2024)
par: Athiwaratkun, Ben, et autres
Publié: (2024)
The N-Grammys: Accelerating Autoregressive Inference with Learning-Free Batched Speculation
par: Stewart, Lawrence, et autres
Publié: (2024)
par: Stewart, Lawrence, et autres
Publié: (2024)
CodeFort: Robust Training for Code Generation Models
par: Zhang, Yuhao, et autres
Publié: (2024)
par: Zhang, Yuhao, et autres
Publié: (2024)
Approximately Aligned Decoding
par: Melcer, Daniel, et autres
Publié: (2024)
par: Melcer, Daniel, et autres
Publié: (2024)
Constrained Decoding for Fill-in-the-Middle Code Language Models via Efficient Left and Right Quotienting of Context-Sensitive Grammars
par: Melcer, Daniel, et autres
Publié: (2024)
par: Melcer, Daniel, et autres
Publié: (2024)
Lightweight reranking for language model generations
par: Jain, Siddhartha, et autres
Publié: (2023)
par: Jain, Siddhartha, et autres
Publié: (2023)
LibEvolutionEval: A Benchmark and Study for Version-Specific Code Generation
par: Kuhar, Sachit, et autres
Publié: (2024)
par: Kuhar, Sachit, et autres
Publié: (2024)
Structural Code Search using Natural Language Queries
par: Limpanukorn, Ben, et autres
Publié: (2025)
par: Limpanukorn, Ben, et autres
Publié: (2025)
CodeAssistBench (CAB): Dataset & Benchmarking for Multi-turn Chat-Based Code Assistance
par: Kim, Myeongsoo, et autres
Publié: (2025)
par: Kim, Myeongsoo, et autres
Publié: (2025)
Automated Evaluation of Retrieval-Augmented Language Models with Task-Specific Exam Generation
par: Guinet, Gauthier, et autres
Publié: (2024)
par: Guinet, Gauthier, et autres
Publié: (2024)
UTFix: Change Aware Unit Test Repairing using LLM
par: Rahman, Shanto, et autres
Publié: (2025)
par: Rahman, Shanto, et autres
Publié: (2025)
Code Representation Learning At Scale
par: Zhang, Dejiao, et autres
Publié: (2024)
par: Zhang, Dejiao, et autres
Publié: (2024)
Multi-IaC-Eval: Benchmarking Cloud Infrastructure as Code Across Multiple Formats
par: Davidson, Sam, et autres
Publié: (2025)
par: Davidson, Sam, et autres
Publié: (2025)
LeDex: Training LLMs to Better Self-Debug and Explain Code
par: Jiang, Nan, et autres
Publié: (2024)
par: Jiang, Nan, et autres
Publié: (2024)
Batch Speculative Decoding Done Right
par: Zhang, Ranran Haoran, et autres
Publié: (2025)
par: Zhang, Ranran Haoran, et autres
Publié: (2025)
Breaking the Code: Security Assessment of AI Code Agents Through Systematic Jailbreaking Attacks
par: Saha, Shoumik, et autres
Publié: (2025)
par: Saha, Shoumik, et autres
Publié: (2025)
Dynamic Graph Attention Networks for Travel Time Distribution Prediction in Urban Arterial Roads
par: Yousefzadeh, Nooshin, et autres
Publié: (2024)
par: Yousefzadeh, Nooshin, et autres
Publié: (2024)
Co‐Pyrolysis of Calophyllum inophyllum Seeds and Polypropylene: Thermokinetics and Batch Studies
par: Subhashree Padhy, et autres
Publié: (2025)
par: Subhashree Padhy, et autres
Publié: (2025)
Fewer Truncations Improve Language Modeling
par: Ding, Hantian, et autres
Publié: (2024)
par: Ding, Hantian, et autres
Publié: (2024)
Iterated Energy-based Flow Matching for Sampling from Boltzmann Densities
par: Woo, Dongyeop, et autres
Publié: (2024)
par: Woo, Dongyeop, et autres
Publié: (2024)
Sampling Decisions
par: Chertkov, Michael, et autres
Publié: (2025)
par: Chertkov, Michael, et autres
Publié: (2025)
Global Resolution: Optimal Multi-Draft Speculative Sampling via Convex Minimization
par: Thomas, Rahul Krishna, et autres
Publié: (2025)
par: Thomas, Rahul Krishna, et autres
Publié: (2025)
TerraFormer: Automated Infrastructure-as-Code with LLMs Fine-Tuned via Policy-Guided Verifier Feedback
par: Jana, Prithwish, et autres
Publié: (2026)
par: Jana, Prithwish, et autres
Publié: (2026)
SpecAgent: A Speculative Retrieval and Forecasting Agent for Code Completion
par: Ma, George, et autres
Publié: (2025)
par: Ma, George, et autres
Publié: (2025)
Batch Transformer: Look for Attention in Batch
par: Her, Myung Beom, et autres
Publié: (2024)
par: Her, Myung Beom, et autres
Publié: (2024)
Code-Aware Prompting: A study of Coverage Guided Test Generation in Regression Setting using LLM
par: Ryan, Gabriel, et autres
Publié: (2024)
par: Ryan, Gabriel, et autres
Publié: (2024)
Equation-of-State Independent Relations in Rapidly Rotating Hybrid Stars
par: Roy, Sujan Kumar
Publié: (2025)
par: Roy, Sujan Kumar
Publié: (2025)
Named Entity Recognition in Hindi using Maximum Entropy and Transliteration
par: Sujan Kumar Saha
Publié: (2008)
par: Sujan Kumar Saha
Publié: (2008)
TETRIS: Optimal Draft Token Selection for Batch Speculative Decoding
par: Wu, Zhaoxuan, et autres
Publié: (2025)
par: Wu, Zhaoxuan, et autres
Publié: (2025)
MineDraft: A Framework for Batch Parallel Speculative Decoding
par: Tang, Zhenwei, et autres
Publié: (2026)
par: Tang, Zhenwei, et autres
Publié: (2026)
Pre-trained Recommender Systems: A Causal Debiasing Perspective
par: Lin, Ziqian, et autres
Publié: (2023)
par: Lin, Ziqian, et autres
Publié: (2023)
Empowering Multi-Turn Tool-Integrated Agentic Reasoning with Group Turn Policy Optimization
par: Ding, Yifeng, et autres
Publié: (2025)
par: Ding, Yifeng, et autres
Publié: (2025)
BASS XLV: Quantifying AGN Selection Effects in the Chandra COSMOS-Legacy Survey with BASS
par: Tokayer, Yarone M., et autres
Publié: (2025)
par: Tokayer, Yarone M., et autres
Publié: (2025)
A Novel Granule‐Based Formulation as a Health Supplement With an Extended Shelf Life, Derived From the Pulp of Bael ( Aegle marmelos ) and Various Herbs
par: Vipin Kumar, et autres
Publié: (2025)
par: Vipin Kumar, et autres
Publié: (2025)
An existence and uniqueness result using bounded variation estimates in Galerkin approximations
par: Mondal, Ramesh, et autres
Publié: (2023)
par: Mondal, Ramesh, et autres
Publié: (2023)
Logic-Scaffolding: Personalized Aspect-Instructed Recommendation Explanation Generation using LLMs
par: Rahdari, Behnam, et autres
Publié: (2023)
par: Rahdari, Behnam, et autres
Publié: (2023)
ContextWeaver: Selective and Dependency-Structured Memory Construction for LLM Agents
par: Wu, Yating, et autres
Publié: (2026)
par: Wu, Yating, et autres
Publié: (2026)
Lossless Token Sequence Compression via Meta-Tokens
par: Harvill, John, et autres
Publié: (2025)
par: Harvill, John, et autres
Publié: (2025)
Speculative Speculative Decoding
par: Kumar, Tanishq, et autres
Publié: (2026)
par: Kumar, Tanishq, et autres
Publié: (2026)
Documents similaires
-
Bifurcated Attention: Accelerating Massively Parallel Decoding with Shared Prefixes in LLMs
par: Athiwaratkun, Ben, et autres
Publié: (2024) -
Token Alignment via Character Matching for Subword Completion
par: Athiwaratkun, Ben, et autres
Publié: (2024) -
The N-Grammys: Accelerating Autoregressive Inference with Learning-Free Batched Speculation
par: Stewart, Lawrence, et autres
Publié: (2024) -
CodeFort: Robust Training for Code Generation Models
par: Zhang, Yuhao, et autres
Publié: (2024) -
Approximately Aligned Decoding
par: Melcer, Daniel, et autres
Publié: (2024)