More Compute Is What You Need
Fuente:
arXiv
Salvato in:
| Autore principale: | Guo, Zhen |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
More Agents Is All You Need
di: Li, Junyou, et al.
Pubblicazione: (2024)
di: Li, Junyou, et al.
Pubblicazione: (2024)
Synthetic Data RL: Task Definition Is All You Need
di: Guo, Yiduo, et al.
Pubblicazione: (2025)
di: Guo, Yiduo, et al.
Pubblicazione: (2025)
Tensor Product Attention Is All You Need
di: Zhang, Yifan, et al.
Pubblicazione: (2025)
di: Zhang, Yifan, et al.
Pubblicazione: (2025)
Attention Smoothing Is All You Need For Unlearning
di: Zade, Saleh Zare, et al.
Pubblicazione: (2026)
di: Zade, Saleh Zare, et al.
Pubblicazione: (2026)
Are More LLM Calls All You Need? Towards Scaling Laws of Compound Inference Systems
di: Chen, Lingjiao, et al.
Pubblicazione: (2024)
di: Chen, Lingjiao, et al.
Pubblicazione: (2024)
You Don't Need Prompt Engineering Anymore: The Prompting Inversion
di: Khan, Imran
Pubblicazione: (2025)
di: Khan, Imran
Pubblicazione: (2025)
Attention Is All You Need for KV Cache in Diffusion LLMs
di: Nguyen-Tri, Quan, et al.
Pubblicazione: (2025)
di: Nguyen-Tri, Quan, et al.
Pubblicazione: (2025)
What Matters in Transformers? Not All Attention is Needed
di: He, Shwai, et al.
Pubblicazione: (2024)
di: He, Shwai, et al.
Pubblicazione: (2024)
Think When You Need: Self-Adaptive Chain-of-Thought Learning
di: Yang, Junjie, et al.
Pubblicazione: (2025)
di: Yang, Junjie, et al.
Pubblicazione: (2025)
The Factorization Curse: Which Tokens You Predict Underlie the Reversal Curse and More
di: Kitouni, Ouail, et al.
Pubblicazione: (2024)
di: Kitouni, Ouail, et al.
Pubblicazione: (2024)
You Are What You Train: Effects of Data Composition on Training Context-aware Machine Translation Models
di: Mąka, Paweł, et al.
Pubblicazione: (2025)
di: Mąka, Paweł, et al.
Pubblicazione: (2025)
Guidance is All You Need: Temperature-Guided Reasoning in Large Language Models
di: Gomaa, Eyad, et al.
Pubblicazione: (2024)
di: Gomaa, Eyad, et al.
Pubblicazione: (2024)
An Extra RMSNorm is All You Need for Fine Tuning to 1.58 Bits
di: Steinmetz, Cody, et al.
Pubblicazione: (2025)
di: Steinmetz, Cody, et al.
Pubblicazione: (2025)
All You Need is One: Capsule Prompt Tuning with a Single Vector
di: Liu, Yiyang, et al.
Pubblicazione: (2025)
di: Liu, Yiyang, et al.
Pubblicazione: (2025)
Beyond the Last Answer: Your Reasoning Trace Uncovers More than You Think
di: Hammoud, Hasan Abed Al Kader, et al.
Pubblicazione: (2025)
di: Hammoud, Hasan Abed Al Kader, et al.
Pubblicazione: (2025)
CoDAR: Continuous Diffusion Language Models are More Powerful Than You Think
di: Shen, Junzhe, et al.
Pubblicazione: (2026)
di: Shen, Junzhe, et al.
Pubblicazione: (2026)
Similarity is Not All You Need: Endowing Retrieval Augmented Generation with Multi Layered Thoughts
di: Gan, Chunjing, et al.
Pubblicazione: (2024)
di: Gan, Chunjing, et al.
Pubblicazione: (2024)
SynthDST: Synthetic Data is All You Need for Few-Shot Dialog State Tracking
di: Kulkarni, Atharva, et al.
Pubblicazione: (2024)
di: Kulkarni, Atharva, et al.
Pubblicazione: (2024)
Evaluation is All You Need: Strategic Overclaiming of LLM Reasoning Capabilities Through Evaluation Design
di: Sun, Lin, et al.
Pubblicazione: (2025)
di: Sun, Lin, et al.
Pubblicazione: (2025)
Are Retrials All You Need? Enhancing Large Language Model Reasoning Without Verbalized Feedback
di: Potamitis, Nearchos, et al.
Pubblicazione: (2025)
di: Potamitis, Nearchos, et al.
Pubblicazione: (2025)
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective
di: He, Shenghua, et al.
Pubblicazione: (2025)
di: He, Shenghua, et al.
Pubblicazione: (2025)
Understand What LLM Needs: Dual Preference Alignment for Retrieval-Augmented Generation
di: Dong, Guanting, et al.
Pubblicazione: (2024)
di: Dong, Guanting, et al.
Pubblicazione: (2024)
Forget What You Know about LLMs Evaluations -- LLMs are Like a Chameleon
di: Cohen-Inger, Nurit, et al.
Pubblicazione: (2025)
di: Cohen-Inger, Nurit, et al.
Pubblicazione: (2025)
SecEncoder: Logs are All You Need in Security
di: Bulut, Muhammed Fatih, et al.
Pubblicazione: (2024)
di: Bulut, Muhammed Fatih, et al.
Pubblicazione: (2024)
What Would You Ask When You First Saw $a^2+b^2=c^2$? Evaluating LLM on Curiosity-Driven Questioning
di: Javaji, Shashidhar Reddy, et al.
Pubblicazione: (2024)
di: Javaji, Shashidhar Reddy, et al.
Pubblicazione: (2024)
Did You Forget What I Asked? Prospective Memory Failures in Large Language Models
di: Mittal, Avni
Pubblicazione: (2026)
di: Mittal, Avni
Pubblicazione: (2026)
One Jump Is All You Need: Short-Cutting Transformers for Early Exit Prediction with One Jump to Fit All Exit Levels
di: Seshadri, Amrit Diggavi
Pubblicazione: (2025)
di: Seshadri, Amrit Diggavi
Pubblicazione: (2025)
Input-Time Scaling: Adding Noise and Irrelevance into Less-Is-More Drastically Improves Reasoning Performance and Efficiency
di: Huang, Rapheal, et al.
Pubblicazione: (2025)
di: Huang, Rapheal, et al.
Pubblicazione: (2025)
Less Noise, More Voice: Reinforcement Learning for Reasoning via Instruction Purification
di: Guo, Yiju, et al.
Pubblicazione: (2026)
di: Guo, Yiju, et al.
Pubblicazione: (2026)
Can't Remember Details in Long Documents? You Need Some R&R
di: Agrawal, Devanshu, et al.
Pubblicazione: (2024)
di: Agrawal, Devanshu, et al.
Pubblicazione: (2024)
You only need 4 extra tokens: Synergistic Test-time Adaptation for LLMs
di: Xu, Yijie, et al.
Pubblicazione: (2025)
di: Xu, Yijie, et al.
Pubblicazione: (2025)
Explain Less, Understand More: Jargon Detection via Personalized Parameter-Efficient Fine-tuning
di: Wu, Bohao, et al.
Pubblicazione: (2025)
di: Wu, Bohao, et al.
Pubblicazione: (2025)
Planning and Editing What You Retrieve for Enhanced Tool Learning
di: Huang, Tenghao, et al.
Pubblicazione: (2024)
di: Huang, Tenghao, et al.
Pubblicazione: (2024)
RecurFormer: Not All Transformer Heads Need Self-Attention
di: Yan, Ruiqing, et al.
Pubblicazione: (2024)
di: Yan, Ruiqing, et al.
Pubblicazione: (2024)
More Thinking, More Bias: Length-Driven Position Bias in Reasoning Models
di: Wang, Xiao
Pubblicazione: (2026)
di: Wang, Xiao
Pubblicazione: (2026)
What Differentiates Educational Literature? A Multimodal Fusion Approach of Transformers and Computational Linguistics
di: Bird, Jordan J.
Pubblicazione: (2024)
di: Bird, Jordan J.
Pubblicazione: (2024)
Packing Analysis: Packing Is More Appropriate for Large Models or Datasets in Supervised Fine-tuning
di: Wang, Shuhe, et al.
Pubblicazione: (2024)
di: Wang, Shuhe, et al.
Pubblicazione: (2024)
Augmenting Human Evaluation with LLM Judges: How Many Human Reviews Do You Need?
di: Kim, Jane Paik
Pubblicazione: (2026)
di: Kim, Jane Paik
Pubblicazione: (2026)
Optimisation Is Not What You Need
di: Ibias, Alfredo
Pubblicazione: (2025)
di: Ibias, Alfredo
Pubblicazione: (2025)
More Expressive Attention with Negative Weights
di: Lv, Ang, et al.
Pubblicazione: (2024)
di: Lv, Ang, et al.
Pubblicazione: (2024)
Documenti analoghi
-
More Agents Is All You Need
di: Li, Junyou, et al.
Pubblicazione: (2024) -
Synthetic Data RL: Task Definition Is All You Need
di: Guo, Yiduo, et al.
Pubblicazione: (2025) -
Tensor Product Attention Is All You Need
di: Zhang, Yifan, et al.
Pubblicazione: (2025) -
Attention Smoothing Is All You Need For Unlearning
di: Zade, Saleh Zare, et al.
Pubblicazione: (2026) -
Are More LLM Calls All You Need? Towards Scaling Laws of Compound Inference Systems
di: Chen, Lingjiao, et al.
Pubblicazione: (2024)