HARP: Hesitation-Aware Reframing in Transformer Inference Pass
Fuente:
arXiv
Guardado en:
| Autores principales: | Storaï, Romain, Hwang, Seung-won |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Intended Target Identification for Anomia Patients with Gradient-based Selective Augmentation
por: Kim, Jongho, et al.
Publicado: (2025)
por: Kim, Jongho, et al.
Publicado: (2025)
AcuRank: Uncertainty-Aware Adaptive Computation for Listwise Reranking
por: Yoon, Soyoung, et al.
Publicado: (2025)
por: Yoon, Soyoung, et al.
Publicado: (2025)
Counterfactual-Consistency Prompting for Relative Temporal Understanding in Large Language Models
por: Kim, Jongho, et al.
Publicado: (2025)
por: Kim, Jongho, et al.
Publicado: (2025)
CoEx -- Co-evolving World-model and Exploration
por: Kim, Minsoo, et al.
Publicado: (2025)
por: Kim, Minsoo, et al.
Publicado: (2025)
LEAP: Layer-wise Exit-Aware Pretraining for Efficient Transformer Inference
por: Kapadia, Shashank, et al.
Publicado: (2026)
por: Kapadia, Shashank, et al.
Publicado: (2026)
SAFE: Stepwise Atomic Feedback for Error correction in Multi-hop Reasoning
por: Kwon, Daeyong, et al.
Publicado: (2026)
por: Kwon, Daeyong, et al.
Publicado: (2026)
DEMO: Reframing Dialogue Interaction with Fine-grained Element Modeling
por: Wang, Minzheng, et al.
Publicado: (2024)
por: Wang, Minzheng, et al.
Publicado: (2024)
Reframing Data Value for Large Language Models Through the Lens of Plausibility
por: Rammal, Mohamad Rida, et al.
Publicado: (2024)
por: Rammal, Mohamad Rida, et al.
Publicado: (2024)
A Comment On "The Illusion of Thinking": Reframing the Reasoning Cliff as an Agentic Gap
por: Khan, Sheraz, et al.
Publicado: (2025)
por: Khan, Sheraz, et al.
Publicado: (2025)
Chain of Grounded Objectives: Bridging Process and Goal-oriented Prompting for Code Generation
por: Yeo, Sangyeop, et al.
Publicado: (2025)
por: Yeo, Sangyeop, et al.
Publicado: (2025)
Relation-based Counterfactual Data Augmentation and Contrastive Learning for Robustifying Natural Language Inference Models
por: Yang, Heerin, et al.
Publicado: (2024)
por: Yang, Heerin, et al.
Publicado: (2024)
CREFT: Sequential Multi-Agent LLM for Character Relation Extraction
por: Chun, Ye Eun, et al.
Publicado: (2025)
por: Chun, Ye Eun, et al.
Publicado: (2025)
Accelerating Transformer Inference for Translation via Parallel Decoding
por: Santilli, Andrea, et al.
Publicado: (2023)
por: Santilli, Andrea, et al.
Publicado: (2023)
Think Before You Act: Decision Transformers with Working Memory
por: Kang, Jikun, et al.
Publicado: (2023)
por: Kang, Jikun, et al.
Publicado: (2023)
Capacity-Aware Inference: Mitigating the Straggler Effect in Mixture of Experts
por: He, Shwai, et al.
Publicado: (2025)
por: He, Shwai, et al.
Publicado: (2025)
Block Transformer: Global-to-Local Language Modeling for Fast Inference
por: Ho, Namgyu, et al.
Publicado: (2024)
por: Ho, Namgyu, et al.
Publicado: (2024)
Cross-Architecture Transfer Learning for Linear-Cost Inference Transformers
por: Choi, Sehyun
Publicado: (2024)
por: Choi, Sehyun
Publicado: (2024)
KV Cache Transform Coding for Compact Storage in LLM Inference
por: Staniszewski, Konrad, et al.
Publicado: (2025)
por: Staniszewski, Konrad, et al.
Publicado: (2025)
TransformerFAM: Feedback attention is working memory
por: Hwang, Dongseong, et al.
Publicado: (2024)
por: Hwang, Dongseong, et al.
Publicado: (2024)
Leveraging LLM Inconsistency to Boost Pass@k Performance
por: Dalal, Uri, et al.
Publicado: (2025)
por: Dalal, Uri, et al.
Publicado: (2025)
Value-Aware Numerical Representations for Transformer Language Models
por: Dutulescu, Andreea, et al.
Publicado: (2026)
por: Dutulescu, Andreea, et al.
Publicado: (2026)
Inference-Aware Fine-Tuning for Best-of-N Sampling in Large Language Models
por: Chow, Yinlam, et al.
Publicado: (2024)
por: Chow, Yinlam, et al.
Publicado: (2024)
SwiftKV: Fast Prefill-Optimized Inference with Knowledge-Preserving Model Transformation
por: Qiao, Aurick, et al.
Publicado: (2024)
por: Qiao, Aurick, et al.
Publicado: (2024)
Integrating Locality-Aware Attention with Transformers for General Geometry PDEs
por: Koh, Minsu, et al.
Publicado: (2025)
por: Koh, Minsu, et al.
Publicado: (2025)
Pass@K Policy Optimization: Solving Harder Reinforcement Learning Problems
por: Walder, Christian, et al.
Publicado: (2025)
por: Walder, Christian, et al.
Publicado: (2025)
M2R2: Mixture of Multi-Rate Residuals for Efficient Transformer Inference
por: Bhendawade, Nikhil, et al.
Publicado: (2025)
por: Bhendawade, Nikhil, et al.
Publicado: (2025)
Transformer-Based Multimodal Knowledge Graph Completion with Link-Aware Contexts
por: Ma, Haodi, et al.
Publicado: (2025)
por: Ma, Haodi, et al.
Publicado: (2025)
ChameleonLLM: Batch-Aware Dynamic Low-Rank Adaptation via Inference-Time Clusters
por: Yuksel, Kamer Ali, et al.
Publicado: (2025)
por: Yuksel, Kamer Ali, et al.
Publicado: (2025)
Alloc-MoE: Budget-Aware Expert Activation Allocation for Efficient Mixture-of-Experts Inference
por: Liu, Baihui, et al.
Publicado: (2026)
por: Liu, Baihui, et al.
Publicado: (2026)
LycheeCluster: Efficient Long-Context Inference with Structure-Aware Chunking and Hierarchical KV Indexing
por: Li, Dongfang, et al.
Publicado: (2026)
por: Li, Dongfang, et al.
Publicado: (2026)
Generative Adapter: Contextualizing Language Models in Parameters with A Single Forward Pass
por: Chen, Tong, et al.
Publicado: (2024)
por: Chen, Tong, et al.
Publicado: (2024)
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
por: Chen, Zhipeng, et al.
Publicado: (2025)
por: Chen, Zhipeng, et al.
Publicado: (2025)
Improving Quantized Model Performance in Qualitative Analysis with Multi-Pass Prompt Verification
por: Adeseye, Aisvarya, et al.
Publicado: (2026)
por: Adeseye, Aisvarya, et al.
Publicado: (2026)
IntPro: A Proxy Agent for Context-Aware Intent Understanding via Retrieval-conditioned Inference
por: Liu, Guanming, et al.
Publicado: (2026)
por: Liu, Guanming, et al.
Publicado: (2026)
Memory Injections: Correcting Multi-Hop Reasoning Failures during Inference in Transformer-Based Language Models
por: Sakarvadia, Mansi, et al.
Publicado: (2023)
por: Sakarvadia, Mansi, et al.
Publicado: (2023)
Relevance to Utility: Process-Supervised Rewrite for RAG
por: Kim, Jaeyoung, et al.
Publicado: (2025)
por: Kim, Jaeyoung, et al.
Publicado: (2025)
SPARE: Single-Pass Annotation with Reference-Guided Evaluation for Automatic Process Supervision and Reward Modelling
por: Rizvi, Md Imbesat Hassan, et al.
Publicado: (2025)
por: Rizvi, Md Imbesat Hassan, et al.
Publicado: (2025)
Dist2ill: Distributional Distillation for One-Pass Uncertainty Estimation in Large Language Models
por: Zhao, Yicong, et al.
Publicado: (2025)
por: Zhao, Yicong, et al.
Publicado: (2025)
KVTuner: Sensitivity-Aware Layer-Wise Mixed-Precision KV Cache Quantization for Efficient and Nearly Lossless LLM Inference
por: Li, Xing, et al.
Publicado: (2025)
por: Li, Xing, et al.
Publicado: (2025)
Characterizing Large Language Model Geometry Helps Solve Toxicity Detection and Generation
por: Balestriero, Randall, et al.
Publicado: (2023)
por: Balestriero, Randall, et al.
Publicado: (2023)
Ejemplares similares
-
Intended Target Identification for Anomia Patients with Gradient-based Selective Augmentation
por: Kim, Jongho, et al.
Publicado: (2025) -
AcuRank: Uncertainty-Aware Adaptive Computation for Listwise Reranking
por: Yoon, Soyoung, et al.
Publicado: (2025) -
Counterfactual-Consistency Prompting for Relative Temporal Understanding in Large Language Models
por: Kim, Jongho, et al.
Publicado: (2025) -
CoEx -- Co-evolving World-model and Exploration
por: Kim, Minsoo, et al.
Publicado: (2025) -
LEAP: Layer-wise Exit-Aware Pretraining for Efficient Transformer Inference
por: Kapadia, Shashank, et al.
Publicado: (2026)