Power-SMC: Low-Latency Sequence-Level Power Sampling for Training-Free LLM Reasoning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Azizi, Seyedarmin, Potraghloo, Erfan Baghaei, Ahmadi, Minoo, Kundu, Souvik, Pedram, Massoud |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Top-H Decoding: Adapting the Creativity and Coherence with Bounded Entropy in Text Generation
von: Potraghloo, Erfan Baghaei, et al.
Veröffentlicht: (2025)
von: Potraghloo, Erfan Baghaei, et al.
Veröffentlicht: (2025)
Activation Steering for Chain-of-Thought Compression
von: Azizi, Seyedarmin, et al.
Veröffentlicht: (2025)
von: Azizi, Seyedarmin, et al.
Veröffentlicht: (2025)
One Token Away from Collapse: The Fragility of Instruction-Tuned Helpfulness
von: Potraghloo, Erfan Baghaei, et al.
Veröffentlicht: (2026)
von: Potraghloo, Erfan Baghaei, et al.
Veröffentlicht: (2026)
LaMDA: Large Model Fine-Tuning via Spectrally Decomposed Low-Dimensional Adaptation
von: Azizi, Seyedarmin, et al.
Veröffentlicht: (2024)
von: Azizi, Seyedarmin, et al.
Veröffentlicht: (2024)
VISTA: Vision-Language Inference for Training-Free Stock Time-Series Analysis
von: Khezresmaeilzadeh, Tina, et al.
Veröffentlicht: (2025)
von: Khezresmaeilzadeh, Tina, et al.
Veröffentlicht: (2025)
Training-Free Acceleration of ViTs with Delayed Spatial Merging
von: Heo, Jung Hwan, et al.
Veröffentlicht: (2023)
von: Heo, Jung Hwan, et al.
Veröffentlicht: (2023)
Efficient Noise Mitigation for Enhancing Inference Accuracy in DNNs on Mixed-Signal Accelerators
von: Azizi, Seyedarmin, et al.
Veröffentlicht: (2024)
von: Azizi, Seyedarmin, et al.
Veröffentlicht: (2024)
Memory-Efficient Vision Transformers: An Activation-Aware Mixed-Rank Compression Strategy
von: Azizi, Seyedarmin, et al.
Veröffentlicht: (2024)
von: Azizi, Seyedarmin, et al.
Veröffentlicht: (2024)
PEANO-ViT: Power-Efficient Approximations of Non-Linearities in Vision Transformers
von: Sadeghi, Mohammad Erfan, et al.
Veröffentlicht: (2024)
von: Sadeghi, Mohammad Erfan, et al.
Veröffentlicht: (2024)
SkipKV: Selective Skipping of KV Generation and Storage for Efficient Inference with Large Reasoning Models
von: Tian, Jiayi, et al.
Veröffentlicht: (2025)
von: Tian, Jiayi, et al.
Veröffentlicht: (2025)
Sensitivity-Aware Mixed-Precision Quantization and Width Optimization of Deep Neural Networks Through Cluster-Based Tree-Structured Parzen Estimation
von: Azizi, Seyedarmin, et al.
Veröffentlicht: (2023)
von: Azizi, Seyedarmin, et al.
Veröffentlicht: (2023)
RocketPPA: Code-Level Power, Performance, and Area Prediction via LLM and Mixture of Experts
von: Abdollahi, Armin, et al.
Veröffentlicht: (2025)
von: Abdollahi, Armin, et al.
Veröffentlicht: (2025)
Neuro-Symbolic Financial Reasoning via Deterministic Fact Ledgers and Adversarial Low-Latency Hallucination Detector
von: Agand, Pedram
Veröffentlicht: (2026)
von: Agand, Pedram
Veröffentlicht: (2026)
FACTER: Fairness-Aware Conformal Thresholding and Prompt Engineering for Enabling Fair LLM-Based Recommender Systems
von: Fayyazi, Arya, et al.
Veröffentlicht: (2025)
von: Fayyazi, Arya, et al.
Veröffentlicht: (2025)
MARCO: Hardware-Aware Neural Architecture Search for Edge Devices with Multi-Agent Reinforcement Learning and Conformal Prediction Filtering
von: Fayyazi, Arya, et al.
Veröffentlicht: (2025)
von: Fayyazi, Arya, et al.
Veröffentlicht: (2025)
Scalable Power Sampling: Unlocking Efficient, Training-Free Reasoning for LLMs via Distribution Sharpening
von: Ji, Xiaotong, et al.
Veröffentlicht: (2026)
von: Ji, Xiaotong, et al.
Veröffentlicht: (2026)
TimingLLM: A Two-Stage Retrieval-Augmented Framework for Pre-Synthesis Timing Prediction from Verilog
von: Abdollahi, Armin, et al.
Veröffentlicht: (2026)
von: Abdollahi, Armin, et al.
Veröffentlicht: (2026)
DILEMMA: Joint LLM Quantization and Distributed LLM Inference Over Edge Computing Systems
von: Hosseinzadeh, Minoo, et al.
Veröffentlicht: (2025)
von: Hosseinzadeh, Minoo, et al.
Veröffentlicht: (2025)
PRISM: Enhancing Protein Inverse Folding through Fine-Grained Retrieval on Structure-Sequence Multimodal Representations
von: Mahbub, Sazan, et al.
Veröffentlicht: (2025)
von: Mahbub, Sazan, et al.
Veröffentlicht: (2025)
Dynamic Co-Optimization Compiler: Leveraging Multi-Agent Reinforcement Learning for Enhanced DNN Accelerator Performance
von: Fayyazi, Arya, et al.
Veröffentlicht: (2024)
von: Fayyazi, Arya, et al.
Veröffentlicht: (2024)
Enhancing Multi-Modal Video Sentiment Classification Through Semi-Supervised Clustering
von: Saadatinia, Mehrshad, et al.
Veröffentlicht: (2025)
von: Saadatinia, Mehrshad, et al.
Veröffentlicht: (2025)
Improving the Throughput of Diffusion-based Large Language Models via a Training-Free Confidence-Aware Calibration
von: Shen, Jucheng, et al.
Veröffentlicht: (2025)
von: Shen, Jucheng, et al.
Veröffentlicht: (2025)
Accelerated Training on Low-Power Edge Devices
von: Ahmed, Mohamed Aboelenien, et al.
Veröffentlicht: (2025)
von: Ahmed, Mohamed Aboelenien, et al.
Veröffentlicht: (2025)
Simultaneous Learning and Optimization via Misspecified Saddle Point Problems
von: Ahmadi, Mohammad Mahdi, et al.
Veröffentlicht: (2025)
von: Ahmadi, Mohammad Mahdi, et al.
Veröffentlicht: (2025)
Block Selective Reprogramming for On-device Training of Vision Transformers
von: Sarkar, Sreetama, et al.
Veröffentlicht: (2024)
von: Sarkar, Sreetama, et al.
Veröffentlicht: (2024)
Probabilistic Numeric SMC Sampling for Bayesian Nonlinear System Identification in Continuous Time
von: Longbottom, Joe D., et al.
Veröffentlicht: (2024)
von: Longbottom, Joe D., et al.
Veröffentlicht: (2024)
The Power of Second Order Methods for Sequence Preconditioning
von: Marsden, Annie, et al.
Veröffentlicht: (2026)
von: Marsden, Annie, et al.
Veröffentlicht: (2026)
DPQ-HD: Post-Training Compression for Ultra-Low Power Hyperdimensional Computing
von: Pandey, Nilesh Prasad, et al.
Veröffentlicht: (2025)
von: Pandey, Nilesh Prasad, et al.
Veröffentlicht: (2025)
Junk DNA Hypothesis: Pruning Small Pre-Trained Weights Irreversibly and Monotonically Impairs "Difficult" Downstream Tasks in LLMs
von: Yin, Lu, et al.
Veröffentlicht: (2023)
von: Yin, Lu, et al.
Veröffentlicht: (2023)
Spiking Decision Transformers: Local Plasticity, Phase-Coding, and Dendritic Routing for Low-Power Sequence Control
von: Pandey, Vishal, et al.
Veröffentlicht: (2025)
von: Pandey, Vishal, et al.
Veröffentlicht: (2025)
Think Just Enough: Sequence-Level Entropy as a Confidence Signal for LLM Reasoning
von: Sharma, Aman, et al.
Veröffentlicht: (2025)
von: Sharma, Aman, et al.
Veröffentlicht: (2025)
Understanding the Expressive Power and Mechanisms of Transformer for Sequence Modeling
von: Wang, Mingze, et al.
Veröffentlicht: (2024)
von: Wang, Mingze, et al.
Veröffentlicht: (2024)
Accelerating LLM Inference with Flexible N:M Sparsity via A Fully Digital Compute-in-Memory Accelerator
von: Ramachandran, Akshat, et al.
Veröffentlicht: (2025)
von: Ramachandran, Akshat, et al.
Veröffentlicht: (2025)
InsightBuild: LLM-Powered Causal Reasoning in Smart Building Systems
von: Neogi, Pinaki Prasad Guha, et al.
Veröffentlicht: (2025)
von: Neogi, Pinaki Prasad Guha, et al.
Veröffentlicht: (2025)
Video Killed the Energy Budget: Characterizing the Latency and Power Regimes of Open Text-to-Video Models
von: Delavande, Julien, et al.
Veröffentlicht: (2025)
von: Delavande, Julien, et al.
Veröffentlicht: (2025)
VESPO: Variational Sequence-Level Soft Policy Optimization for Stable Off-Policy LLM Training
von: Shen, Guobin, et al.
Veröffentlicht: (2026)
von: Shen, Guobin, et al.
Veröffentlicht: (2026)
NES: An Instruction-Free, Low-Latency Next Edit Suggestion Framework Powered by Learned Historical Editing Trajectories
von: Chen, Xinfang, et al.
Veröffentlicht: (2025)
von: Chen, Xinfang, et al.
Veröffentlicht: (2025)
Unleashing The Power of Pre-Trained Language Models for Irregularly Sampled Time Series
von: Zhang, Weijia, et al.
Veröffentlicht: (2024)
von: Zhang, Weijia, et al.
Veröffentlicht: (2024)
MicroScopiQ: Accelerating Foundational Models through Outlier-Aware Microscaling Quantization
von: Ramachandran, Akshat, et al.
Veröffentlicht: (2024)
von: Ramachandran, Akshat, et al.
Veröffentlicht: (2024)
Is Temperature Sample Efficient for Softmax Gaussian Mixture of Experts?
von: Nguyen, Huy, et al.
Veröffentlicht: (2024)
von: Nguyen, Huy, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Top-H Decoding: Adapting the Creativity and Coherence with Bounded Entropy in Text Generation
von: Potraghloo, Erfan Baghaei, et al.
Veröffentlicht: (2025) -
Activation Steering for Chain-of-Thought Compression
von: Azizi, Seyedarmin, et al.
Veröffentlicht: (2025) -
One Token Away from Collapse: The Fragility of Instruction-Tuned Helpfulness
von: Potraghloo, Erfan Baghaei, et al.
Veröffentlicht: (2026) -
LaMDA: Large Model Fine-Tuning via Spectrally Decomposed Low-Dimensional Adaptation
von: Azizi, Seyedarmin, et al.
Veröffentlicht: (2024) -
VISTA: Vision-Language Inference for Training-Free Stock Time-Series Analysis
von: Khezresmaeilzadeh, Tina, et al.
Veröffentlicht: (2025)