TrimR: Verifier-based Training-Free Thinking Compression for Efficient Test-Time Scaling
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lin, Weizhe, Li, Xing, Yang, Zhiyuan, Fu, Xiaojin, Zhen, Hui-Ling, Wang, Yaoyuan, Yu, Xianzhi, Liu, Wulong, Li, Xiaosong, Yuan, Mingxuan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SVDq: 1.25-bit and 410x Key Cache Compression for LLM Attention
von: Yankun, Hong, et al.
Veröffentlicht: (2025)
von: Yankun, Hong, et al.
Veröffentlicht: (2025)
Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling
von: Sun, Shengyin, et al.
Veröffentlicht: (2025)
von: Sun, Shengyin, et al.
Veröffentlicht: (2025)
Revisiting Judge Decoding from First Principles via Training-Free Distributional Divergence
von: Sun, Shengyin, et al.
Veröffentlicht: (2026)
von: Sun, Shengyin, et al.
Veröffentlicht: (2026)
KVTuner: Sensitivity-Aware Layer-Wise Mixed-Precision KV Cache Quantization for Efficient and Nearly Lossless LLM Inference
von: Li, Xing, et al.
Veröffentlicht: (2025)
von: Li, Xing, et al.
Veröffentlicht: (2025)
MOSS: Efficient and Accurate FP8 LLM Training with Microscaling and Automatic Scaling
von: Zhang, Yu, et al.
Veröffentlicht: (2025)
von: Zhang, Yu, et al.
Veröffentlicht: (2025)
Towards Efficient Agents: A Co-Design of Inference Architecture and System
von: Lin, Weizhe, et al.
Veröffentlicht: (2025)
von: Lin, Weizhe, et al.
Veröffentlicht: (2025)
AttentionPredictor: Temporal Patterns Matter for KV Cache Compression
von: Yang, Qingyue, et al.
Veröffentlicht: (2025)
von: Yang, Qingyue, et al.
Veröffentlicht: (2025)
Analytical FFN-to-MoE Restructuring via Activation Pattern Analysis
von: Pei, Zehua, et al.
Veröffentlicht: (2025)
von: Pei, Zehua, et al.
Veröffentlicht: (2025)
What Matters For Safety Alignment?
von: Li, Xing, et al.
Veröffentlicht: (2026)
von: Li, Xing, et al.
Veröffentlicht: (2026)
Behavioral Fingerprinting of Large Language Models
von: Pei, Zehua, et al.
Veröffentlicht: (2025)
von: Pei, Zehua, et al.
Veröffentlicht: (2025)
AgentCollab: A Self-Evaluation-Driven Collaboration Paradigm for Efficient LLM Agents
von: Gao, Wenbo, et al.
Veröffentlicht: (2026)
von: Gao, Wenbo, et al.
Veröffentlicht: (2026)
MixPE: Quantization and Hardware Co-design for Efficient LLM Inference
von: Zhang, Yu, et al.
Veröffentlicht: (2024)
von: Zhang, Yu, et al.
Veröffentlicht: (2024)
VisionTrim: Unified Vision Token Compression for Training-Free MLLM Acceleration
von: Yu, Hanxun, et al.
Veröffentlicht: (2026)
von: Yu, Hanxun, et al.
Veröffentlicht: (2026)
Unlocking Efficient Long-to-Short LLM Reasoning with Model Merging
von: Wu, Han, et al.
Veröffentlicht: (2025)
von: Wu, Han, et al.
Veröffentlicht: (2025)
PSD: Pushing the Pareto Frontier of Diffusion LLMs via Parallel Speculative Decoding
von: Sun, Shengyin, et al.
Veröffentlicht: (2026)
von: Sun, Shengyin, et al.
Veröffentlicht: (2026)
Unleashing Low-Bit Inference on Ascend NPUs: A Comprehensive Evaluation of HiFloat Formats
von: Zhao, Pengxiang, et al.
Veröffentlicht: (2026)
von: Zhao, Pengxiang, et al.
Veröffentlicht: (2026)
SwiftMem: Fast Agentic Memory via Query-aware Indexing
von: Tian, Anxin, et al.
Veröffentlicht: (2026)
von: Tian, Anxin, et al.
Veröffentlicht: (2026)
Faster and Better LLMs via Latency-Aware Test-Time Scaling
von: Wang, Zili, et al.
Veröffentlicht: (2025)
von: Wang, Zili, et al.
Veröffentlicht: (2025)
MemDLM: Memory-Enhanced DLM Training
von: Pei, Zehua, et al.
Veröffentlicht: (2026)
von: Pei, Zehua, et al.
Veröffentlicht: (2026)
ViCrop-Det: Spatial Attention Entropy Guided Cropping for Training-Free Small-Object Detection
von: Wang, Hui, et al.
Veröffentlicht: (2026)
von: Wang, Hui, et al.
Veröffentlicht: (2026)
Benchmarking Post-Training Quantization of Large Language Models under Microscaling Floating Point Formats
von: Zhang, Manyi, et al.
Veröffentlicht: (2026)
von: Zhang, Manyi, et al.
Veröffentlicht: (2026)
Iterative Deepening Sampling as Efficient Test-Time Scaling
von: Chen, Weizhe, et al.
Veröffentlicht: (2025)
von: Chen, Weizhe, et al.
Veröffentlicht: (2025)
Beyond Speedup -- Utilizing KV Cache for Sampling and Reasoning
von: Xing, Zeyu, et al.
Veröffentlicht: (2026)
von: Xing, Zeyu, et al.
Veröffentlicht: (2026)
FocuSFT: Bilevel Optimization for Dilution-Aware Long-Context Fine-Tuning
von: Pei, Zehua, et al.
Veröffentlicht: (2026)
von: Pei, Zehua, et al.
Veröffentlicht: (2026)
From Pruning to Grafting: Dynamic Knowledge Redistribution via Learnable Layer Fusion
von: Pei, Zehua, et al.
Veröffentlicht: (2024)
von: Pei, Zehua, et al.
Veröffentlicht: (2024)
EAQuant: Enhancing Post-Training Quantization for MoE Models via Expert-Aware Optimization
von: Fu, Zhongqian, et al.
Veröffentlicht: (2025)
von: Fu, Zhongqian, et al.
Veröffentlicht: (2025)
PASER: Post-Training Data Selection for Efficient Pruned Large Language Model Recovery
von: He, Bowei, et al.
Veröffentlicht: (2025)
von: He, Bowei, et al.
Veröffentlicht: (2025)
Test-Time Adaptation by Causal Trimming
von: Liu, Yingnan, et al.
Veröffentlicht: (2025)
von: Liu, Yingnan, et al.
Veröffentlicht: (2025)
PreMoE: Proactive Inference for Efficient Mixture-of-Experts
von: Pei, Zehua, et al.
Veröffentlicht: (2025)
von: Pei, Zehua, et al.
Veröffentlicht: (2025)
Critique to Verify: Accurate and Honest Test-Time Scaling with RL-Trained Verifiers
von: Yang, Zhicheng, et al.
Veröffentlicht: (2025)
von: Yang, Zhicheng, et al.
Veröffentlicht: (2025)
Entropy-Gated Branching for Efficient Test-Time Reasoning
von: Li, Xianzhi, et al.
Veröffentlicht: (2025)
von: Li, Xianzhi, et al.
Veröffentlicht: (2025)
Think2Drive: Efficient Reinforcement Learning by Thinking in Latent World Model for Quasi-Realistic Autonomous Driving (in CARLA-v2)
von: Li, Qifeng, et al.
Veröffentlicht: (2024)
von: Li, Qifeng, et al.
Veröffentlicht: (2024)
REG: A Regularization Optimizer for Robust Training Dynamics
von: Liu, Zehua, et al.
Veröffentlicht: (2025)
von: Liu, Zehua, et al.
Veröffentlicht: (2025)
CAT: A Causally Graph Attention Network for Trimming Heterophilic Graph
von: He, Silu, et al.
Veröffentlicht: (2023)
von: He, Silu, et al.
Veröffentlicht: (2023)
ThinkLess: A Training-Free Inference-Efficient Method for Reducing Reasoning Redundancy
von: Li, Gengyang, et al.
Veröffentlicht: (2025)
von: Li, Gengyang, et al.
Veröffentlicht: (2025)
EPD-Serve: A Flexible Multimodal EPD Disaggregation Inference Serving System On Ascend
von: Bai, Fan, et al.
Veröffentlicht: (2026)
von: Bai, Fan, et al.
Veröffentlicht: (2026)
Trimming the Fat: Efficient Compression of 3D Gaussian Splats through Pruning
von: Ali, Muhammad Salman, et al.
Veröffentlicht: (2024)
von: Ali, Muhammad Salman, et al.
Veröffentlicht: (2024)
Improved Algorithm for Permutation Testing
von: Zhang, Xiaojin
Veröffentlicht: (2020)
von: Zhang, Xiaojin
Veröffentlicht: (2020)
Unleashing the Potential of Pre-Trained Diffusion Models for Generalizable Person Re-Identification
von: Li, Jiachen, et al.
Veröffentlicht: (2025)
von: Li, Jiachen, et al.
Veröffentlicht: (2025)
HAP: Hybrid Adaptive Parallelism for Efficient Mixture-of-Experts Inference
von: Lin, Haoran, et al.
Veröffentlicht: (2025)
von: Lin, Haoran, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SVDq: 1.25-bit and 410x Key Cache Compression for LLM Attention
von: Yankun, Hong, et al.
Veröffentlicht: (2025) -
Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling
von: Sun, Shengyin, et al.
Veröffentlicht: (2025) -
Revisiting Judge Decoding from First Principles via Training-Free Distributional Divergence
von: Sun, Shengyin, et al.
Veröffentlicht: (2026) -
KVTuner: Sensitivity-Aware Layer-Wise Mixed-Precision KV Cache Quantization for Efficient and Nearly Lossless LLM Inference
von: Li, Xing, et al.
Veröffentlicht: (2025) -
MOSS: Efficient and Accurate FP8 LLM Training with Microscaling and Automatic Scaling
von: Zhang, Yu, et al.
Veröffentlicht: (2025)