Dual-Dimensional Consistency: Balancing Budget and Quality in Adaptive Inference-Time Scaling
Fuente:
arXiv
Saved in:
| Main Authors: | Xu, Rongman, Li, Yifei, Zhao, Tianzhe, Wu, Yanrui, Li, Bo, Yan, Hang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
$ϕ$-Decoding: Adaptive Foresight Sampling for Balanced Inference-Time Exploration and Exploitation
by: Xu, Fangzhi, et al.
Published: (2025)
by: Xu, Fangzhi, et al.
Published: (2025)
SAGE: Scale-Aware Gradual Evolution for Continual Knowledge Graph Embedding
by: Li, Yifei, et al.
Published: (2025)
by: Li, Yifei, et al.
Published: (2025)
Locomo-Plus: Beyond-Factual Cognitive Memory Evaluation Framework for LLM Agents
by: Li, Yifei, et al.
Published: (2026)
by: Li, Yifei, et al.
Published: (2026)
Hierarchical Budget Policy Optimization for Adaptive Reasoning
by: Lyu, Shangke, et al.
Published: (2025)
by: Lyu, Shangke, et al.
Published: (2025)
Solve-Detect-Verify: Inference-Time Scaling with Flexible Generative Verifier
by: Zhong, Jianyuan, et al.
Published: (2025)
by: Zhong, Jianyuan, et al.
Published: (2025)
Inference-Time Budget Control for LLM Search Agents
by: Fang, Zhengru, et al.
Published: (2026)
by: Fang, Zhengru, et al.
Published: (2026)
Dual Consistent Constraint via Disentangled Consistency and Complementarity for Multi-view Clustering
by: Li, Bo, et al.
Published: (2025)
by: Li, Bo, et al.
Published: (2025)
Dimensional Balance Improves Large Scale Spatiotemporal Prediction Performance
by: Chen, Jing, et al.
Published: (2026)
by: Chen, Jing, et al.
Published: (2026)
ContextBudget: Budget-Aware Context Management for Long-Horizon Search Agents
by: Wu, Yong, et al.
Published: (2026)
by: Wu, Yong, et al.
Published: (2026)
MACS: Modality-Aware Capacity Scaling for Efficient Multimodal MoE Inference
by: Li, Bo, et al.
Published: (2026)
by: Li, Bo, et al.
Published: (2026)
OccamToken: Efficient VLM Inference with Training-Free and Budget-Adaptive Token Pruning
by: Li, Geng, et al.
Published: (2026)
by: Li, Geng, et al.
Published: (2026)
MeetBench-XL: Calibrated Multi-Dimensional Evaluation and Learned Dual-Policy Agents for Real-Time Meetings
by: Hu, Yuelin, et al.
Published: (2026)
by: Hu, Yuelin, et al.
Published: (2026)
Time Matters: Scaling Laws for Any Budget
by: Inbar, Itay, et al.
Published: (2024)
by: Inbar, Itay, et al.
Published: (2024)
Inference-Time Scaling for Generalist Reward Modeling
by: Liu, Zijun, et al.
Published: (2025)
by: Liu, Zijun, et al.
Published: (2025)
Adaptive Inference-Time Scaling via Cyclic Diffusion Search
by: Lee, Gyubin, et al.
Published: (2025)
by: Lee, Gyubin, et al.
Published: (2025)
SubQuad: Near-Quadratic-Free Structure Inference with Distribution-Balanced Objectives in Adaptive Receptor framework
by: Fu, Rong, et al.
Published: (2026)
by: Fu, Rong, et al.
Published: (2026)
PersonaDual: Balancing Personalization and Objectivity via Adaptive Reasoning
by: Liu, Xiaoyou, et al.
Published: (2026)
by: Liu, Xiaoyou, et al.
Published: (2026)
LoopServe: An Adaptive Dual-phase LLM Inference Acceleration System for Multi-Turn Dialogues
by: Li, Haoyang, et al.
Published: (2025)
by: Li, Haoyang, et al.
Published: (2025)
UniScale: Adaptive Unified Inference Scaling via Online Joint Optimization of Model Routing and Test-Time Scaling
by: Huang, Kaiyu, et al.
Published: (2026)
by: Huang, Kaiyu, et al.
Published: (2026)
Scale-Adaptive Balancing of Exploration and Exploitation in Classical Planning
by: Wissow, Stephen, et al.
Published: (2023)
by: Wissow, Stephen, et al.
Published: (2023)
Predict-then-Diffuse: Adaptive Response Length for Compute-Budgeted Inference in Diffusion LLMs
by: Rottoli, Michael, et al.
Published: (2026)
by: Rottoli, Michael, et al.
Published: (2026)
Wider or Deeper? Scaling LLM Inference-Time Compute with Adaptive Branching Tree Search
by: Inoue, Yuichi, et al.
Published: (2025)
by: Inoue, Yuichi, et al.
Published: (2025)
Physics-Informed Inference Time Scaling for Solving High-Dimensional PDE via Defect Correction
by: Fan, Zexi, et al.
Published: (2025)
by: Fan, Zexi, et al.
Published: (2025)
Uncertainty-aware Prototype Learning with Variational Inference for Few-shot Point Cloud Segmentation
by: Zhao, Yifei, et al.
Published: (2026)
by: Zhao, Yifei, et al.
Published: (2026)
QPART: Adaptive Model Quantization and Dynamic Workload Balancing for Accuracy-aware Edge Inference
by: Li, Xiangchen, et al.
Published: (2025)
by: Li, Xiangchen, et al.
Published: (2025)
FreqPolicy: Efficient Flow-based Visuomotor Policy via Frequency Consistency
by: Su, Yifei, et al.
Published: (2025)
by: Su, Yifei, et al.
Published: (2025)
DISCO Balances the Scales: Adaptive Domain- and Difficulty-Aware Reinforcement Learning on Imbalanced Data
by: Zhou, Yuhang, et al.
Published: (2025)
by: Zhou, Yuhang, et al.
Published: (2025)
Budget-Aware Tool-Use Enables Effective Agent Scaling
by: Liu, Tengxiao, et al.
Published: (2025)
by: Liu, Tengxiao, et al.
Published: (2025)
Sample, Scrutinize and Scale: Effective Inference-Time Search by Scaling Verification
by: Zhao, Eric, et al.
Published: (2025)
by: Zhao, Eric, et al.
Published: (2025)
A Multi-Dimensional Quality Scoring Framework for Decentralized LLM Inference with Proof of Quality
by: Tian, Arther, et al.
Published: (2026)
by: Tian, Arther, et al.
Published: (2026)
Adaptive Orchestration for Large-Scale Inference on Heterogeneous Accelerator Systems Balancing Cost, Performance, and Resilience
by: Biran, Yahav, et al.
Published: (2025)
by: Biran, Yahav, et al.
Published: (2025)
Highly Efficient Test-Time Scaling for T2I Diffusion Models with Text Embedding Perturbation
by: Xu, Hang, et al.
Published: (2025)
by: Xu, Hang, et al.
Published: (2025)
Consist-Retinex: One-Step Noise-Emphasized Consistency Training Accelerates High-Quality Retinex Enhancement
by: Xu, Jian, et al.
Published: (2025)
by: Xu, Jian, et al.
Published: (2025)
Adaptive KV-Cache Compression without Manually Setting Budget
by: Tang, Chenxia, et al.
Published: (2025)
by: Tang, Chenxia, et al.
Published: (2025)
Reasoning on a Budget: A Survey of Adaptive and Controllable Test-Time Compute in LLMs
by: Alomrani, Mohammad Ali, et al.
Published: (2025)
by: Alomrani, Mohammad Ali, et al.
Published: (2025)
Ada-KV: Optimizing KV Cache Eviction by Adaptive Budget Allocation for Efficient LLM Inference
by: Feng, Yuan, et al.
Published: (2024)
by: Feng, Yuan, et al.
Published: (2024)
Budget-Constrained Tool Learning with Planning
by: Zheng, Yuanhang, et al.
Published: (2024)
by: Zheng, Yuanhang, et al.
Published: (2024)
BEAT: Balanced Frequency Adaptive Tuning for Long-Term Time-Series Forecasting
by: Li, Zhixuan, et al.
Published: (2025)
by: Li, Zhixuan, et al.
Published: (2025)
Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models
by: Wu, Yangzhen, et al.
Published: (2024)
by: Wu, Yangzhen, et al.
Published: (2024)
TheoremForge: Scaling up Formal Data Synthesis with Low-Budget Agentic Workflow
by: Tao, Yicheng, et al.
Published: (2026)
by: Tao, Yicheng, et al.
Published: (2026)
Similar Items
-
$ϕ$-Decoding: Adaptive Foresight Sampling for Balanced Inference-Time Exploration and Exploitation
by: Xu, Fangzhi, et al.
Published: (2025) -
SAGE: Scale-Aware Gradual Evolution for Continual Knowledge Graph Embedding
by: Li, Yifei, et al.
Published: (2025) -
Locomo-Plus: Beyond-Factual Cognitive Memory Evaluation Framework for LLM Agents
by: Li, Yifei, et al.
Published: (2026) -
Hierarchical Budget Policy Optimization for Adaptive Reasoning
by: Lyu, Shangke, et al.
Published: (2025) -
Solve-Detect-Verify: Inference-Time Scaling with Flexible Generative Verifier
by: Zhong, Jianyuan, et al.
Published: (2025)