AutoL2S: Auto Long-Short Reasoning for Efficient Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Luo, Feng, Chuang, Yu-Neng, Wang, Guanchu, Le, Hoang Anh Duy, Zhong, Shaochen, Liu, Hongyi, Yuan, Jiayi, Sui, Yang, Braverman, Vladimir, Chaudhary, Vipin, Hu, Xia |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
KV Cache Compression, But What Must We Give in Return? A Comprehensive Benchmark of Long Context Capable Approaches
by: Yuan, Jiayi, et al.
Published: (2024)
by: Yuan, Jiayi, et al.
Published: (2024)
Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models
by: Sui, Yang, et al.
Published: (2025)
by: Sui, Yang, et al.
Published: (2025)
Taylor Unswift: Secured Weight Release for Large Language Models via Taylor Expansion
by: Wang, Guanchu, et al.
Published: (2024)
by: Wang, Guanchu, et al.
Published: (2024)
FaithLM: Towards Faithful Explanations for Large Language Models
by: Chuang, Yu-Neng, et al.
Published: (2024)
by: Chuang, Yu-Neng, et al.
Published: (2024)
70% Size, 100% Accuracy: Lossless LLM Compression for Efficient GPU Inference via Dynamic-Length Float (DFloat11)
by: Zhang, Tianyi, et al.
Published: (2025)
by: Zhang, Tianyi, et al.
Published: (2025)
Self-ensemble: Mitigating Confidence Mis-calibration for Large Language Models
by: Xu, Zicheng, et al.
Published: (2025)
by: Xu, Zicheng, et al.
Published: (2025)
Demystifying OPD: Length Inflation and Stabilization Strategies for Large Language Models
by: Luo, Feng, et al.
Published: (2026)
by: Luo, Feng, et al.
Published: (2026)
Confident or Seek Stronger: Exploring Uncertainty-Based On-device LLM Routing From Benchmarking to Generalization
by: Chuang, Yu-Neng, et al.
Published: (2025)
by: Chuang, Yu-Neng, et al.
Published: (2025)
DTS: Enhancing Large Reasoning Models via Decoding Tree Sketching
by: Xu, Zicheng, et al.
Published: (2025)
by: Xu, Zicheng, et al.
Published: (2025)
Word Salad Chopper: Reasoning Models Waste A Ton Of Decoding Budget On Useless Repetitions, Self-Knowingly
by: Xie, Wenya, et al.
Published: (2025)
by: Xie, Wenya, et al.
Published: (2025)
Winner-Take-All Column Row Sampling for Memory Efficient Adaptation of Language Model
by: Liu, Zirui, et al.
Published: (2023)
by: Liu, Zirui, et al.
Published: (2023)
100-LongBench: Are de facto Long-Context Benchmarks Literally Evaluating Long-Context Ability?
by: Yang, Wang, et al.
Published: (2025)
by: Yang, Wang, et al.
Published: (2025)
LoRATK: LoRA Once, Backdoor Everywhere in the Share-and-Play Ecosystem
by: Liu, Hongyi, et al.
Published: (2024)
by: Liu, Hongyi, et al.
Published: (2024)
KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache
by: Liu, Zirui, et al.
Published: (2024)
by: Liu, Zirui, et al.
Published: (2024)
Assessing and Enhancing Large Language Models in Rare Disease Question-answering
by: Wang, Guanchu, et al.
Published: (2024)
by: Wang, Guanchu, et al.
Published: (2024)
TVE: Learning Meta-attribution for Transferable Vision Explainer
by: Wang, Guanchu, et al.
Published: (2023)
by: Wang, Guanchu, et al.
Published: (2023)
Quantize What Counts: More for Keys, Less for Values
by: Hariri, Mohsen, et al.
Published: (2025)
by: Hariri, Mohsen, et al.
Published: (2025)
Statistical Test for Anomaly Detections by Variational Auto-Encoders
by: Miwa, Daiki, et al.
Published: (2024)
by: Miwa, Daiki, et al.
Published: (2024)
TP-GMOT: Tracking Generic Multiple Object by Textual Prompt with Motion-Appearance Cost (MAC) SORT
by: Anh, Duy Le Dinh, et al.
Published: (2024)
by: Anh, Duy Le Dinh, et al.
Published: (2024)
The LLM Data Auditor: A Metric-oriented Survey on Quality and Trustworthiness in Evaluating Synthetic Data
by: Zhang, Kaituo, et al.
Published: (2026)
by: Zhang, Kaituo, et al.
Published: (2026)
Auto-Relational Reasoning
by: Konstantoulas, Ioannis, et al.
Published: (2026)
by: Konstantoulas, Ioannis, et al.
Published: (2026)
UniAutoML: A Human-Centered Framework for Unified Discriminative and Generative AutoML with Large Language Models
by: Guo, Jiayi, et al.
Published: (2024)
by: Guo, Jiayi, et al.
Published: (2024)
AutoGluon-Multimodal (AutoMM): Supercharging Multimodal AutoML with Foundation Models
by: Tang, Zhiqiang, et al.
Published: (2024)
by: Tang, Zhiqiang, et al.
Published: (2024)
Hydra: A Modular Architecture for Efficient Long-Context Reasoning
by: Chaudhary, Siddharth, et al.
Published: (2025)
by: Chaudhary, Siddharth, et al.
Published: (2025)
Personalized Query Auto-Completion for Long and Short-Term Interests with Adaptive Detoxification Generation
by: Wang, Zhibo, et al.
Published: (2025)
by: Wang, Zhibo, et al.
Published: (2025)
AutoViVQA: A Large-Scale Automatically Constructed Dataset for Vietnamese Visual Question Answering
by: Tuong, Nguyen Anh, et al.
Published: (2026)
by: Tuong, Nguyen Anh, et al.
Published: (2026)
Longer Context, Deeper Thinking: Uncovering the Role of Long-Context Ability in Reasoning
by: Yang, Wang, et al.
Published: (2025)
by: Yang, Wang, et al.
Published: (2025)
Robust Asynchronous Planning via Auto-Formalization
by: Zhang, Jiayi, et al.
Published: (2026)
by: Zhang, Jiayi, et al.
Published: (2026)
AutoOdom: Learning Auto-regressive Proprioceptive Odometry for Legged Locomotion
by: Luo, Changsheng, et al.
Published: (2025)
by: Luo, Changsheng, et al.
Published: (2025)
LRD-Net: A Lightweight Real-Centered Detection Network for Cross-Domain Face Forgery Detection
by: Zhang, Xuecen, et al.
Published: (2026)
by: Zhang, Xuecen, et al.
Published: (2026)
Scout Before You Attend: Sketch-and-Walk Sparse Attention for Efficient LLM Inference
by: Le, Hoang Anh Duy, et al.
Published: (2026)
by: Le, Hoang Anh Duy, et al.
Published: (2026)
Ranking Reasoning LLMs under Test-Time Scaling
by: Hariri, Mohsen, et al.
Published: (2026)
by: Hariri, Mohsen, et al.
Published: (2026)
Enhanced Kalman with Adaptive Appearance Motion SORT for Grounded Generic Multiple Object Tracking
by: Anh, Duy Le Dinh, et al.
Published: (2024)
by: Anh, Duy Le Dinh, et al.
Published: (2024)
Proof of Thought : Neurosymbolic Program Synthesis allows Robust and Interpretable Reasoning
by: Ganguly, Debargha, et al.
Published: (2024)
by: Ganguly, Debargha, et al.
Published: (2024)
AutoReason: Automatic Few-Shot Reasoning Decomposition
by: Sevinc, Arda, et al.
Published: (2024)
by: Sevinc, Arda, et al.
Published: (2024)
VideoAuto-R1: Video Auto Reasoning via Thinking Once, Answering Twice
by: Liu, Shuming, et al.
Published: (2026)
by: Liu, Shuming, et al.
Published: (2026)
Resurrect Mask AutoRegressive Modeling for Efficient and Scalable Image Generation
by: Xin, Yi, et al.
Published: (2025)
by: Xin, Yi, et al.
Published: (2025)
Speculative Thinking: Enhancing Small-Model Reasoning with Large Model Guidance at Inference Time
by: Yang, Wang, et al.
Published: (2025)
by: Yang, Wang, et al.
Published: (2025)
AutoSSVH: Exploring Automated Frame Sampling for Efficient Self-Supervised Video Hashing
by: Lian, Niu, et al.
Published: (2025)
by: Lian, Niu, et al.
Published: (2025)
HFGN: A Python code of Hartree-Fock calculations for Graphene Nanoribbons
by: Hoang Anh Le
Published: (2025)
by: Hoang Anh Le
Published: (2025)
Similar Items
-
KV Cache Compression, But What Must We Give in Return? A Comprehensive Benchmark of Long Context Capable Approaches
by: Yuan, Jiayi, et al.
Published: (2024) -
Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models
by: Sui, Yang, et al.
Published: (2025) -
Taylor Unswift: Secured Weight Release for Large Language Models via Taylor Expansion
by: Wang, Guanchu, et al.
Published: (2024) -
FaithLM: Towards Faithful Explanations for Large Language Models
by: Chuang, Yu-Neng, et al.
Published: (2024) -
70% Size, 100% Accuracy: Lossless LLM Compression for Efficient GPU Inference via Dynamic-Length Float (DFloat11)
by: Zhang, Tianyi, et al.
Published: (2025)