SALSA: Single-pass Autoregressive LLM Structured Classification
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Berdichevsky, Ruslan, Nahum-Gefen, Shai, Zaken, Elad Ben |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
BitFit: Simple Parameter-efficient Fine-tuning for Transformer-based Masked Language-models
par: Ben-Zaken, Elad, et autres
Publié: (2021)
par: Ben-Zaken, Elad, et autres
Publié: (2021)
SALSA: Speedy ASR-LLM Synchronous Aggregation
par: Mittal, Ashish, et autres
Publié: (2024)
par: Mittal, Ashish, et autres
Publié: (2024)
Single-pass Detection of Jailbreaking Input in Large Language Models
par: Candogan, Leyla Naz, et autres
Publié: (2025)
par: Candogan, Leyla Naz, et autres
Publié: (2025)
Superposed Decoding: Multiple Generations from a Single Autoregressive Inference Pass
par: Shen, Ethan, et autres
Publié: (2024)
par: Shen, Ethan, et autres
Publié: (2024)
Automatic Replication of LLM Mistakes in Medical Conversations
par: Proniakin, Oleksii, et autres
Publié: (2025)
par: Proniakin, Oleksii, et autres
Publié: (2025)
Learning a Continue-Thinking Token for Enhanced Test-Time Scaling
par: Ringel, Liran, et autres
Publié: (2025)
par: Ringel, Liran, et autres
Publié: (2025)
Chain of LoRA: Efficient Fine-tuning of Language Models via Residual Learning
par: Xia, Wenhan, et autres
Publié: (2024)
par: Xia, Wenhan, et autres
Publié: (2024)
FFN-SkipLLM: A Hidden Gem for Autoregressive Decoding with Adaptive Feed Forward Skipping
par: Jaiswal, Ajay, et autres
Publié: (2024)
par: Jaiswal, Ajay, et autres
Publié: (2024)
IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems
par: Levi, Elad, et autres
Publié: (2025)
par: Levi, Elad, et autres
Publié: (2025)
BARRED: Synthetic Training of Custom Policy Guardrails via Asymmetric Debate
par: Mazza, Arnon, et autres
Publié: (2026)
par: Mazza, Arnon, et autres
Publié: (2026)
Why Any-Order Autoregressive Models Need Two-Stream Attention: A Structural-Semantic Tradeoff
par: Pynadath, Patrick, et autres
Publié: (2026)
par: Pynadath, Patrick, et autres
Publié: (2026)
The Branch Not Taken: Predicting Branching in Online Conversations
par: Meital, Shai, et autres
Publié: (2024)
par: Meital, Shai, et autres
Publié: (2024)
Think Big, Generate Quick: LLM-to-SLM for Fast Autoregressive Decoding
par: Bergner, Benjamin, et autres
Publié: (2024)
par: Bergner, Benjamin, et autres
Publié: (2024)
Single Character Perturbations Break LLM Alignment
par: Lin, Leon, et autres
Publié: (2024)
par: Lin, Leon, et autres
Publié: (2024)
Cognitive Fatigue in Autoregressive Transformers: Formalization and Measurement
par: Marwah, Riju, et autres
Publié: (2026)
par: Marwah, Riju, et autres
Publié: (2026)
Evaluating Memory Structure in LLM Agents
par: Shutova, Alina, et autres
Publié: (2026)
par: Shutova, Alina, et autres
Publié: (2026)
ERD: A Framework for Improving LLM Reasoning for Cognitive Distortion Classification
par: Lim, Sehee, et autres
Publié: (2024)
par: Lim, Sehee, et autres
Publié: (2024)
Methods of improving LLM training stability
par: Rybakov, Oleg, et autres
Publié: (2024)
par: Rybakov, Oleg, et autres
Publié: (2024)
Beyond Autoregression: Discrete Diffusion for Complex Reasoning and Planning
par: Ye, Jiacheng, et autres
Publié: (2024)
par: Ye, Jiacheng, et autres
Publié: (2024)
Dynamic Context Pruning for Efficient and Interpretable Autoregressive Transformers
par: Anagnostidis, Sotiris, et autres
Publié: (2023)
par: Anagnostidis, Sotiris, et autres
Publié: (2023)
On Mesa-Optimization in Autoregressively Trained Transformers: Emergence and Capability
par: Zheng, Chenyu, et autres
Publié: (2024)
par: Zheng, Chenyu, et autres
Publié: (2024)
Projected Autoregression: Autoregressive Language Generation in Continuous State Space
par: Naparstek, Oshri
Publié: (2026)
par: Naparstek, Oshri
Publié: (2026)
Intent-based Prompt Calibration: Enhancing prompt optimization with synthetic boundary cases
par: Levi, Elad, et autres
Publié: (2024)
par: Levi, Elad, et autres
Publié: (2024)
FBI-LLM: Scaling Up Fully Binarized LLMs from Scratch via Autoregressive Distillation
par: Ma, Liqun, et autres
Publié: (2024)
par: Ma, Liqun, et autres
Publié: (2024)
Not All LLM-Generated Data Are Equal: Rethinking Data Weighting in Text Classification
par: Kuo, Hsun-Yu, et autres
Publié: (2024)
par: Kuo, Hsun-Yu, et autres
Publié: (2024)
TELEClass: Taxonomy Enrichment and LLM-Enhanced Hierarchical Text Classification with Minimal Supervision
par: Zhang, Yunyi, et autres
Publié: (2024)
par: Zhang, Yunyi, et autres
Publié: (2024)
Knowledge Distillation in Automated Annotation: Supervised Text Classification with LLM-Generated Training Labels
par: Pangakis, Nicholas, et autres
Publié: (2024)
par: Pangakis, Nicholas, et autres
Publié: (2024)
Odysseys: Benchmarking Web Agents on Realistic Long Horizon Tasks
par: Jang, Lawrence Keunho, et autres
Publié: (2026)
par: Jang, Lawrence Keunho, et autres
Publié: (2026)
DORE: A Dataset For Portuguese Definition Generation
par: Furtado, Anna Beatriz Dimas, et autres
Publié: (2024)
par: Furtado, Anna Beatriz Dimas, et autres
Publié: (2024)
Flexible Language Modeling in Continuous Space with Transformer-based Autoregressive Flows
par: Zhang, Ruixiang, et autres
Publié: (2025)
par: Zhang, Ruixiang, et autres
Publié: (2025)
Mask-Enhanced Autoregressive Prediction: Pay Less Attention to Learn More
par: Zhuang, Xialie, et autres
Publié: (2025)
par: Zhuang, Xialie, et autres
Publié: (2025)
Beyond Autoregression: Fast LLMs via Self-Distillation Through Time
par: Deschenaux, Justin, et autres
Publié: (2024)
par: Deschenaux, Justin, et autres
Publié: (2024)
Beyond a Single Extractor: Re-thinking HTML-to-Text Extraction for LLM Pretraining
par: Li, Jeffrey, et autres
Publié: (2026)
par: Li, Jeffrey, et autres
Publié: (2026)
Memory-Efficient Structured Backpropagation for On-Device LLM Fine-Tuning
par: Park, Juneyoung, et autres
Publié: (2026)
par: Park, Juneyoung, et autres
Publié: (2026)
Transformers represent belief state geometry in their residual stream
par: Shai, Adam S., et autres
Publié: (2024)
par: Shai, Adam S., et autres
Publié: (2024)
Reviving Any-Subset Autoregressive Models with Principled Parallel Sampling and Speculative Decoding
par: Guo, Gabe, et autres
Publié: (2025)
par: Guo, Gabe, et autres
Publié: (2025)
Why Are Positional Encodings Nonessential for Deep Autoregressive Transformers? Revisiting a Petroglyph
par: Irie, Kazuki
Publié: (2024)
par: Irie, Kazuki
Publié: (2024)
Shared Latent Space by Both Languages in Non-Autoregressive Neural Machine Translation
par: Heo, DongNyeong, et autres
Publié: (2023)
par: Heo, DongNyeong, et autres
Publié: (2023)
AutoTimes: Autoregressive Time Series Forecasters via Large Language Models
par: Liu, Yong, et autres
Publié: (2024)
par: Liu, Yong, et autres
Publié: (2024)
Lory: Fully Differentiable Mixture-of-Experts for Autoregressive Language Model Pre-training
par: Zhong, Zexuan, et autres
Publié: (2024)
par: Zhong, Zexuan, et autres
Publié: (2024)
Documents similaires
-
BitFit: Simple Parameter-efficient Fine-tuning for Transformer-based Masked Language-models
par: Ben-Zaken, Elad, et autres
Publié: (2021) -
SALSA: Speedy ASR-LLM Synchronous Aggregation
par: Mittal, Ashish, et autres
Publié: (2024) -
Single-pass Detection of Jailbreaking Input in Large Language Models
par: Candogan, Leyla Naz, et autres
Publié: (2025) -
Superposed Decoding: Multiple Generations from a Single Autoregressive Inference Pass
par: Shen, Ethan, et autres
Publié: (2024) -
Automatic Replication of LLM Mistakes in Medical Conversations
par: Proniakin, Oleksii, et autres
Publié: (2025)