Top-H Decoding: Adapting the Creativity and Coherence with Bounded Entropy in Text Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Potraghloo, Erfan Baghaei, Azizi, Seyedarmin, Kundu, Souvik, Pedram, Massoud |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
One Token Away from Collapse: The Fragility of Instruction-Tuned Helpfulness
by: Potraghloo, Erfan Baghaei, et al.
Published: (2026)
by: Potraghloo, Erfan Baghaei, et al.
Published: (2026)
Activation Steering for Chain-of-Thought Compression
by: Azizi, Seyedarmin, et al.
Published: (2025)
by: Azizi, Seyedarmin, et al.
Published: (2025)
LaMDA: Large Model Fine-Tuning via Spectrally Decomposed Low-Dimensional Adaptation
by: Azizi, Seyedarmin, et al.
Published: (2024)
by: Azizi, Seyedarmin, et al.
Published: (2024)
Power-SMC: Low-Latency Sequence-Level Power Sampling for Training-Free LLM Reasoning
by: Azizi, Seyedarmin, et al.
Published: (2026)
by: Azizi, Seyedarmin, et al.
Published: (2026)
Efficient Noise Mitigation for Enhancing Inference Accuracy in DNNs on Mixed-Signal Accelerators
by: Azizi, Seyedarmin, et al.
Published: (2024)
by: Azizi, Seyedarmin, et al.
Published: (2024)
Memory-Efficient Vision Transformers: An Activation-Aware Mixed-Rank Compression Strategy
by: Azizi, Seyedarmin, et al.
Published: (2024)
by: Azizi, Seyedarmin, et al.
Published: (2024)
Sensitivity-Aware Mixed-Precision Quantization and Width Optimization of Deep Neural Networks Through Cluster-Based Tree-Structured Parzen Estimation
by: Azizi, Seyedarmin, et al.
Published: (2023)
by: Azizi, Seyedarmin, et al.
Published: (2023)
Training-Free Acceleration of ViTs with Delayed Spatial Merging
by: Heo, Jung Hwan, et al.
Published: (2023)
by: Heo, Jung Hwan, et al.
Published: (2023)
SkipKV: Selective Skipping of KV Generation and Storage for Efficient Inference with Large Reasoning Models
by: Tian, Jiayi, et al.
Published: (2025)
by: Tian, Jiayi, et al.
Published: (2025)
VISTA: Vision-Language Inference for Training-Free Stock Time-Series Analysis
by: Khezresmaeilzadeh, Tina, et al.
Published: (2025)
by: Khezresmaeilzadeh, Tina, et al.
Published: (2025)
PEANO-ViT: Power-Efficient Approximations of Non-Linearities in Vision Transformers
by: Sadeghi, Mohammad Erfan, et al.
Published: (2024)
by: Sadeghi, Mohammad Erfan, et al.
Published: (2024)
COFT: Counterfactual-Conformal Decoding for Fair Chain-of-Thought Reasoning in Large Language Models
by: Fayyazi, Arya, et al.
Published: (2026)
by: Fayyazi, Arya, et al.
Published: (2026)
Hakim: Farsi Text Embedding Model
by: Sarmadi, Mehran, et al.
Published: (2025)
by: Sarmadi, Mehran, et al.
Published: (2025)
Reflection-Window Decoding: Text Generation with Selective Refinement
by: Tang, Zeyu, et al.
Published: (2025)
by: Tang, Zeyu, et al.
Published: (2025)
GEAR: An Efficient KV Cache Compression Recipe for Near-Lossless Generative Inference of LLM
by: Kang, Hao, et al.
Published: (2024)
by: Kang, Hao, et al.
Published: (2024)
CITER: Collaborative Inference for Efficient Large Language Model Decoding with Token-Level Routing
by: Zheng, Wenhao, et al.
Published: (2025)
by: Zheng, Wenhao, et al.
Published: (2025)
Locally Coherent Parallel Decoding in Diffusion Language Models
by: Hersche, Michael, et al.
Published: (2026)
by: Hersche, Michael, et al.
Published: (2026)
Decoding the AI Pen: Techniques and Challenges in Detecting AI-Generated Text
by: Abdali, Sara, et al.
Published: (2024)
by: Abdali, Sara, et al.
Published: (2024)
Entropy Adaptive Decoding: Dynamic Model Switching for Efficient Inference
by: Simonds, Toby
Published: (2025)
by: Simonds, Toby
Published: (2025)
AFLoRA: Adaptive Freezing of Low Rank Adaptation in Parameter Efficient Fine-Tuning of Large Models
by: Liu, Zeyu, et al.
Published: (2024)
by: Liu, Zeyu, et al.
Published: (2024)
Unveiling Intrinsic Dimension of Texts: from Academic Abstract to Creative Story
by: Pedashenko, Vladislav, et al.
Published: (2025)
by: Pedashenko, Vladislav, et al.
Published: (2025)
When Models Examine Themselves: Vocabulary-Activation Correspondence in Self-Referential Processing
by: Dadfar, Zachary Pedram
Published: (2026)
by: Dadfar, Zachary Pedram
Published: (2026)
Understanding the Performance and Estimating the Cost of LLM Fine-Tuning
by: Xia, Yuchen, et al.
Published: (2024)
by: Xia, Yuchen, et al.
Published: (2024)
Rep2Text: Decoding Full Text from a Single LLM Token Representation
by: Zhao, Haiyan, et al.
Published: (2025)
by: Zhao, Haiyan, et al.
Published: (2025)
TransformLLM: Adapting Large Language Models via LLM-Transformed Reading Comprehension Text
by: Arbel, Iftach, et al.
Published: (2024)
by: Arbel, Iftach, et al.
Published: (2024)
Cooking Up Creativity: Enhancing LLM Creativity through Structured Recombination
by: Mizrahi, Moran, et al.
Published: (2025)
by: Mizrahi, Moran, et al.
Published: (2025)
FACTER: Fairness-Aware Conformal Thresholding and Prompt Engineering for Enabling Fair LLM-Based Recommender Systems
by: Fayyazi, Arya, et al.
Published: (2025)
by: Fayyazi, Arya, et al.
Published: (2025)
ShiftAddLLM: Accelerating Pretrained LLMs via Post-Training Multiplication-Less Reparameterization
by: You, Haoran, et al.
Published: (2024)
by: You, Haoran, et al.
Published: (2024)
EnvGen: Generating and Adapting Environments via LLMs for Training Embodied Agents
by: Zala, Abhay, et al.
Published: (2024)
by: Zala, Abhay, et al.
Published: (2024)
CreativityBench: Evaluating Agent Creative Reasoning via Affordance-Based Tool Repurposing
by: Qian, Cheng, et al.
Published: (2026)
by: Qian, Cheng, et al.
Published: (2026)
AdaptAgent: Adapting Multimodal Web Agents with Few-Shot Learning from Human Demonstrations
by: Verma, Gaurav, et al.
Published: (2024)
by: Verma, Gaurav, et al.
Published: (2024)
Draft-Conditioned Constrained Decoding for Structured Generation in LLMs
by: Reddy, Avinash, et al.
Published: (2026)
by: Reddy, Avinash, et al.
Published: (2026)
QUIET: A Multi-Blank Cascaded Story Cloze Benchmark for LLM Creative Generation Capability
by: Zou, Bo, et al.
Published: (2026)
by: Zou, Bo, et al.
Published: (2026)
Weaver: Foundation Models for Creative Writing
by: Wang, Tiannan, et al.
Published: (2024)
by: Wang, Tiannan, et al.
Published: (2024)
Towards Adapting Open-Source Large Language Models for Expert-Level Clinical Note Generation
by: Wang, Hanyin, et al.
Published: (2024)
by: Wang, Hanyin, et al.
Published: (2024)
A Comprehensive Study of Decoder-Only LLMs for Text-to-Image Generation
by: Wang, Andrew Z., et al.
Published: (2025)
by: Wang, Andrew Z., et al.
Published: (2025)
POSS: Position Specialist Generates Better Draft for Speculative Decoding
by: Huang, Langlin, et al.
Published: (2025)
by: Huang, Langlin, et al.
Published: (2025)
Adapting Language Models via Token Translation
by: Feng, Zhili, et al.
Published: (2024)
by: Feng, Zhili, et al.
Published: (2024)
Think Big, Generate Quick: LLM-to-SLM for Fast Autoregressive Decoding
by: Bergner, Benjamin, et al.
Published: (2024)
by: Bergner, Benjamin, et al.
Published: (2024)
Contrastive Decoding for Synthetic Data Generation in Low-Resource Language Modeling
by: Ulm, Jannek, et al.
Published: (2025)
by: Ulm, Jannek, et al.
Published: (2025)
Similar Items
-
One Token Away from Collapse: The Fragility of Instruction-Tuned Helpfulness
by: Potraghloo, Erfan Baghaei, et al.
Published: (2026) -
Activation Steering for Chain-of-Thought Compression
by: Azizi, Seyedarmin, et al.
Published: (2025) -
LaMDA: Large Model Fine-Tuning via Spectrally Decomposed Low-Dimensional Adaptation
by: Azizi, Seyedarmin, et al.
Published: (2024) -
Power-SMC: Low-Latency Sequence-Level Power Sampling for Training-Free LLM Reasoning
by: Azizi, Seyedarmin, et al.
Published: (2026) -
Efficient Noise Mitigation for Enhancing Inference Accuracy in DNNs on Mixed-Signal Accelerators
by: Azizi, Seyedarmin, et al.
Published: (2024)