Inverse Scaling in Test-Time Compute
Fuente:
arXiv
Saved in:
| Main Authors: | Gema, Aryo Pradipta, Hägele, Alexander, Chen, Runjin, Arditi, Andy, Goldman-Wetzler, Jacob, Fraser-Taliente, Kit, Sleight, Henry, Petrini, Linda, Michael, Julian, Alex, Beatrice, Minervini, Pasquale, Chen, Yanda, Benton, Joe, Perez, Ethan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Lost in Time: Clock and Calendar Understanding Challenges in Multimodal LLMs
by: Saxena, Rohit, et al.
Published: (2025)
by: Saxena, Rohit, et al.
Published: (2025)
The Hot Mess of AI: How Does Misalignment Scale With Model Intelligence and Task Complexity?
by: Hägele, Alexander, et al.
Published: (2026)
by: Hägele, Alexander, et al.
Published: (2026)
Edinburgh Clinical NLP at SemEval-2024 Task 2: Fine-tune your model unless you have access to GPT-4
by: Gema, Aryo Pradipta, et al.
Published: (2024)
by: Gema, Aryo Pradipta, et al.
Published: (2024)
Parameter-Efficient Fine-Tuning of LLaMA for the Clinical Domain
by: Gema, Aryo Pradipta, et al.
Published: (2023)
by: Gema, Aryo Pradipta, et al.
Published: (2023)
Beyond Data Filtering: Knowledge Localization for Capability Removal in LLMs
by: Shilov, Igor, et al.
Published: (2025)
by: Shilov, Igor, et al.
Published: (2025)
Persona Vectors: Monitoring and Controlling Character Traits in Language Models
by: Chen, Runjin, et al.
Published: (2025)
by: Chen, Runjin, et al.
Published: (2025)
Edinburgh Clinical NLP at MEDIQA-CORR 2024: Guiding Large Language Models with Hints
by: Gema, Aryo Pradipta, et al.
Published: (2024)
by: Gema, Aryo Pradipta, et al.
Published: (2024)
SCOPE: Self-Play via Co-Evolving Policies for Open-Ended Tasks
by: Kwan, Wai-Chung, et al.
Published: (2026)
by: Kwan, Wai-Chung, et al.
Published: (2026)
The $T^{μν}$ of the conformal scalars
by: Fraser-Taliente, Kit, et al.
Published: (2026)
by: Fraser-Taliente, Kit, et al.
Published: (2026)
DeCoRe: Decoding by Contrasting Retrieval Heads to Mitigate Hallucinations
by: Gema, Aryo Pradipta, et al.
Published: (2024)
by: Gema, Aryo Pradipta, et al.
Published: (2024)
Self-Training Large Language Models for Tool-Use Without Demonstrations
by: Luo, Ne, et al.
Published: (2025)
by: Luo, Ne, et al.
Published: (2025)
GRADA: Graph-based Reranking against Adversarial Documents Attack
by: Zheng, Jingjie, et al.
Published: (2025)
by: Zheng, Jingjie, et al.
Published: (2025)
Noiser: Bounded Input Perturbations for Attributing Large Language Models
by: Madani, Mohammad Reza Ghasemi, et al.
Published: (2025)
by: Madani, Mohammad Reza Ghasemi, et al.
Published: (2025)
An Analysis of Decoding Methods for LLM-based Agents for Faithful Multi-Hop Question Answering
by: Murphy, Alexander, et al.
Published: (2025)
by: Murphy, Alexander, et al.
Published: (2025)
Diffusion Models for Cayley Graphs
by: Douglas, Michael R., et al.
Published: (2025)
by: Douglas, Michael R., et al.
Published: (2025)
Teaching Models to Verbalize Reward Hacking in Chain-of-Thought Reasoning
by: Turpin, Miles, et al.
Published: (2025)
by: Turpin, Miles, et al.
Published: (2025)
Unsupervised Elicitation of Language Models
by: Wen, Jiaxin, et al.
Published: (2025)
by: Wen, Jiaxin, et al.
Published: (2025)
Analysing the Residual Stream of Language Models Under Knowledge Conflicts
by: Zhao, Yu, et al.
Published: (2024)
by: Zhao, Yu, et al.
Published: (2024)
Steering Knowledge Selection Behaviours in LLMs via SAE-Based Representation Engineering
by: Zhao, Yu, et al.
Published: (2024)
by: Zhao, Yu, et al.
Published: (2024)
The sphere free energy of the vector models to order $1/N$
by: Fraser-Taliente, Ludo
Published: (2025)
by: Fraser-Taliente, Ludo
Published: (2025)
Local CFTs extremise $F$
by: Fraser-Taliente, Ludo
Published: (2026)
by: Fraser-Taliente, Ludo
Published: (2026)
Quantum field theories with many fields
by: Fraser-Taliente, Ludo
Published: (2026)
by: Fraser-Taliente, Ludo
Published: (2026)
Not So Flat Metrics
by: Fraser-Taliente, Kit, et al.
Published: (2024)
by: Fraser-Taliente, Kit, et al.
Published: (2024)
CoMAT: Chain of Mathematically Annotated Thought Improves Mathematical Reasoning
by: Leang, Joshua Ong Jun, et al.
Published: (2024)
by: Leang, Joshua Ong Jun, et al.
Published: (2024)
Scalpel vs. Hammer: GRPO Amplifies Existing Capabilities, SFT Replaces Them
by: Rajani, Neel, et al.
Published: (2025)
by: Rajani, Neel, et al.
Published: (2025)
Same Answer, Different Representations: Hidden instability in VLMs
by: Wani, Farooq Ahmad, et al.
Published: (2026)
by: Wani, Farooq Ahmad, et al.
Published: (2026)
The Hallucinations Leaderboard -- An Open Effort to Measure Hallucinations in Large Language Models
by: Hong, Giwon, et al.
Published: (2024)
by: Hong, Giwon, et al.
Published: (2024)
Symbolic Regression with Multimodal Large Language Models and Kolmogorov Arnold Networks
by: Harvey, Thomas R., et al.
Published: (2025)
by: Harvey, Thomas R., et al.
Published: (2025)
PiCSAR: Probabilistic Confidence Selection And Ranking for Reasoning Chains
by: Leang, Joshua Ong Jun, et al.
Published: (2025)
by: Leang, Joshua Ong Jun, et al.
Published: (2025)
$F$-extremization determines certain large-$N$ CFTs
by: Fraser-Taliente, Ludo, et al.
Published: (2024)
by: Fraser-Taliente, Ludo, et al.
Published: (2024)
Melonic limits of the quartic Yukawa model and general features of melonic CFTs
by: Fraser-Taliente, Ludo, et al.
Published: (2024)
by: Fraser-Taliente, Ludo, et al.
Published: (2024)
Do Composed Image Retrieval Benchmarks Require Multimodal Composition?
by: Attimonelli, Matteo, et al.
Published: (2026)
by: Attimonelli, Matteo, et al.
Published: (2026)
Speeding up and reducing memory usage for scientific machine learning via mixed precision
by: Hayford, Joel, et al.
Published: (2024)
by: Hayford, Joel, et al.
Published: (2024)
Can GPT-3.5 Generate and Code Discharge Summaries?
by: Falis, Matúš, et al.
Published: (2024)
by: Falis, Matúš, et al.
Published: (2024)
A Comparative Study on Patient Language across Therapeutic Domains for Effective Patient Voice Classification in Online Health Discussions
by: Lysandrou, Giorgos, et al.
Published: (2024)
by: Lysandrou, Giorgos, et al.
Published: (2024)
Enumerating Calabi-Yau Manifolds: Placing bounds on the number of diffeomorphism classes in the Kreuzer-Skarke list
by: Chandra, Aditi, et al.
Published: (2023)
by: Chandra, Aditi, et al.
Published: (2023)
Computation of Quark Masses from String Theory
by: Constantin, Andrei, et al.
Published: (2024)
by: Constantin, Andrei, et al.
Published: (2024)
SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
by: Kutasov, Jonathan, et al.
Published: (2025)
by: Kutasov, Jonathan, et al.
Published: (2025)
Fermion Masses and Mixing in String-Inspired Models
by: Constantin, Andrei, et al.
Published: (2024)
by: Constantin, Andrei, et al.
Published: (2024)
A Nonlocal Schwinger Model
by: Fraser-Taliente, Ludovic, et al.
Published: (2024)
by: Fraser-Taliente, Ludovic, et al.
Published: (2024)
Similar Items
-
Lost in Time: Clock and Calendar Understanding Challenges in Multimodal LLMs
by: Saxena, Rohit, et al.
Published: (2025) -
The Hot Mess of AI: How Does Misalignment Scale With Model Intelligence and Task Complexity?
by: Hägele, Alexander, et al.
Published: (2026) -
Edinburgh Clinical NLP at SemEval-2024 Task 2: Fine-tune your model unless you have access to GPT-4
by: Gema, Aryo Pradipta, et al.
Published: (2024) -
Parameter-Efficient Fine-Tuning of LLaMA for the Clinical Domain
by: Gema, Aryo Pradipta, et al.
Published: (2023) -
Beyond Data Filtering: Knowledge Localization for Capability Removal in LLMs
by: Shilov, Igor, et al.
Published: (2025)