Preservation Is Not Enough for Width Growth: Regime-Sensitive Selection of Dense LM Warm Starts
Fuente:
arXiv
Salvato in:
| Autore principale: | Unlu, Eren |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Don't Start What You Can't Finish: A Counterfactual Audit of Support-State Triage in LLM Agents
di: Unlu, Eren
Pubblicazione: (2026)
di: Unlu, Eren
Pubblicazione: (2026)
Know When to Trust the Skill: Delayed Appraisal and Epistemic Vigilance for Single-Agent LLMs
di: Unlu, Eren
Pubblicazione: (2026)
di: Unlu, Eren
Pubblicazione: (2026)
Geotokens and Geotransformers
di: Unlu, Eren
Pubblicazione: (2024)
di: Unlu, Eren
Pubblicazione: (2024)
STEP: Warm-Started Visuomotor Policies with Spatiotemporal Consistency Prediction
di: Li, Jinhao, et al.
Pubblicazione: (2026)
di: Li, Jinhao, et al.
Pubblicazione: (2026)
Asking Is Not Enough: Protocol Sensitivity in LLM Confidence Calibration
di: Kim, Hankyeol, et al.
Pubblicazione: (2026)
di: Kim, Hankyeol, et al.
Pubblicazione: (2026)
Policy Optimization for PDE Control with a Warm Start
di: Zhang, Xiangyuan, et al.
Pubblicazione: (2024)
di: Zhang, Xiangyuan, et al.
Pubblicazione: (2024)
Deep Learning Warm Starts for Trajectory Optimization on the International Space Station
di: Banerjee, Somrita, et al.
Pubblicazione: (2025)
di: Banerjee, Somrita, et al.
Pubblicazione: (2025)
A Multi-Stage Warm-Start Deep Learning Framework for Unit Commitment
di: Za'ter, Muhy Eddin, et al.
Pubblicazione: (2026)
di: Za'ter, Muhy Eddin, et al.
Pubblicazione: (2026)
DASH: Warm-Starting Neural Network Training in Stationary Settings without Loss of Plasticity
di: Shin, Baekrok, et al.
Pubblicazione: (2024)
di: Shin, Baekrok, et al.
Pubblicazione: (2024)
Selecting Language Models for Social Science: Start Small, Start Open, and Validate
di: Stoltz, Dustin S., et al.
Pubblicazione: (2026)
di: Stoltz, Dustin S., et al.
Pubblicazione: (2026)
WARP: A Benchmark for Primal-Dual Warm-Starting of Interior-Point Solvers
di: Suri, Dhruv, et al.
Pubblicazione: (2026)
di: Suri, Dhruv, et al.
Pubblicazione: (2026)
Newton's Lantern: A Reinforcement Learning Framework for Finetuning AC Power Flow Warm Start Models
di: Bose, Shourya, et al.
Pubblicazione: (2026)
di: Bose, Shourya, et al.
Pubblicazione: (2026)
Learning to Solve the Quadratic Assignment Problem with Warm-Started MCMC Finetuning
di: Pan, Yicheng, et al.
Pubblicazione: (2026)
di: Pan, Yicheng, et al.
Pubblicazione: (2026)
When Few Steps Are Enough: Training-Free Acceleration of Identity-Preserved Generation
di: Zheng, Dongqi
Pubblicazione: (2026)
di: Zheng, Dongqi
Pubblicazione: (2026)
When Is Enough Not Enough? Illusory Completion in Search Agents
di: Ko, Dayoon, et al.
Pubblicazione: (2026)
di: Ko, Dayoon, et al.
Pubblicazione: (2026)
From the Pursuit of Universal AGI Architecture to Systematic Approach to Heterogenous AGI: Addressing Alignment, Energy, & AGI Grand Challenges
di: Kurshan, Eren
Pubblicazione: (2023)
di: Kurshan, Eren
Pubblicazione: (2023)
Virtual Width Networks
di: Seed, et al.
Pubblicazione: (2025)
di: Seed, et al.
Pubblicazione: (2025)
Is Monitoring Enough? Strategic Agent Selection For Stealthy Attack in Multi-Agent Discussions
di: Xiang, Qiuchi, et al.
Pubblicazione: (2026)
di: Xiang, Qiuchi, et al.
Pubblicazione: (2026)
Loop-Extrusion Linkage: Spectral Ordering and Interval-Based Structure Discovery for Continuous Optimization
di: Unlu, Eren
Pubblicazione: (2026)
di: Unlu, Eren
Pubblicazione: (2026)
A Regime Theory of Controller Class Selection for LLM Action Decisions
di: Jiang, Zhaoyang, et al.
Pubblicazione: (2026)
di: Jiang, Zhaoyang, et al.
Pubblicazione: (2026)
Learning to Select: Query-Aware Adaptive Dimension Selection for Dense Retrieval
di: Wu, Zhanyu, et al.
Pubblicazione: (2026)
di: Wu, Zhanyu, et al.
Pubblicazione: (2026)
Training for Compositional Sensitivity Reduces Dense Retrieval Generalization
di: Ralev, Radoslav, et al.
Pubblicazione: (2026)
di: Ralev, Radoslav, et al.
Pubblicazione: (2026)
Adaptive Width Neural Networks
di: Errica, Federico, et al.
Pubblicazione: (2025)
di: Errica, Federico, et al.
Pubblicazione: (2025)
The Gatekeeper Knows Enough
di: Abebayew, Fikresilase Wondmeneh
Pubblicazione: (2025)
di: Abebayew, Fikresilase Wondmeneh
Pubblicazione: (2025)
Identifying the Risks of LM Agents with an LM-Emulated Sandbox
di: Ruan, Yangjun, et al.
Pubblicazione: (2023)
di: Ruan, Yangjun, et al.
Pubblicazione: (2023)
Consolidating LAMA with Best-First Width Search
di: Corrêa, Augusto B., et al.
Pubblicazione: (2024)
di: Corrêa, Augusto B., et al.
Pubblicazione: (2024)
Screening Is Enough
di: Nakanishi, Ken M.
Pubblicazione: (2026)
di: Nakanishi, Ken M.
Pubblicazione: (2026)
On the Diminishing Returns of Width for Continual Learning
di: Guha, Etash, et al.
Pubblicazione: (2024)
di: Guha, Etash, et al.
Pubblicazione: (2024)
Mixture of Universal Experts: Scaling Virtual Width via Depth-Width Transformation
di: Chen, Yilong, et al.
Pubblicazione: (2026)
di: Chen, Yilong, et al.
Pubblicazione: (2026)
Dynamic-Width Speculative Beam Decoding for Efficient LLM Inference
di: Qin, Zongyue, et al.
Pubblicazione: (2024)
di: Qin, Zongyue, et al.
Pubblicazione: (2024)
Improving Epidemic Analyses with Privacy-Preserving Integration of Sensitive Data
di: Guan, Zihan, et al.
Pubblicazione: (2025)
di: Guan, Zihan, et al.
Pubblicazione: (2025)
QV May Be Enough: Toward the Essence of Attention in LLMs
di: Edward, Zhang
Pubblicazione: (2026)
di: Edward, Zhang
Pubblicazione: (2026)
Greedy Is Enough: Sparse Action Discovery in Agentic LLMs
di: Majumdar, Angshul
Pubblicazione: (2026)
di: Majumdar, Angshul
Pubblicazione: (2026)
Xmodel-LM Technical Report
di: Wang, Yichuan, et al.
Pubblicazione: (2024)
di: Wang, Yichuan, et al.
Pubblicazione: (2024)
SleepLM: Natural-Language Intelligence for Human Sleep
di: Xu, Zongzhe, et al.
Pubblicazione: (2026)
di: Xu, Zongzhe, et al.
Pubblicazione: (2026)
Agents Are Not Enough
di: Shah, Chirag, et al.
Pubblicazione: (2024)
di: Shah, Chirag, et al.
Pubblicazione: (2024)
Rethinking Science in the Age of Artificial Intelligence
di: Eren, Maksim E., et al.
Pubblicazione: (2025)
di: Eren, Maksim E., et al.
Pubblicazione: (2025)
Softmax is not Enough (for Adaptive Conformal Classification)
di: Attar, Navid Akhavan, et al.
Pubblicazione: (2026)
di: Attar, Navid Akhavan, et al.
Pubblicazione: (2026)
When Can Model-Free Reinforcement Learning be Enough for Thinking?
di: Hanna, Josiah P., et al.
Pubblicazione: (2025)
di: Hanna, Josiah P., et al.
Pubblicazione: (2025)
Selective Steering: Norm-Preserving Control Through Discriminative Layer Selection
di: Dang, Quy-Anh, et al.
Pubblicazione: (2026)
di: Dang, Quy-Anh, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Don't Start What You Can't Finish: A Counterfactual Audit of Support-State Triage in LLM Agents
di: Unlu, Eren
Pubblicazione: (2026) -
Know When to Trust the Skill: Delayed Appraisal and Epistemic Vigilance for Single-Agent LLMs
di: Unlu, Eren
Pubblicazione: (2026) -
Geotokens and Geotransformers
di: Unlu, Eren
Pubblicazione: (2024) -
STEP: Warm-Started Visuomotor Policies with Spatiotemporal Consistency Prediction
di: Li, Jinhao, et al.
Pubblicazione: (2026) -
Asking Is Not Enough: Protocol Sensitivity in LLM Confidence Calibration
di: Kim, Hankyeol, et al.
Pubblicazione: (2026)