Think Again! The Effect of Test-Time Compute on Preferences, Opinions, and Beliefs of Large Language Models
Fuente:
arXiv
Salvato in:
| Autori principali: | Kour, George, Nakash, Itay, Anaby-Tavor, Ateret, Shmueli-Scheuer, Michal |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Efficient Agent Evaluation via Diversity-Guided User Simulation
di: Nakash, Itay, et al.
Pubblicazione: (2026)
di: Nakash, Itay, et al.
Pubblicazione: (2026)
Breaking ReAct Agents: Foot-in-the-Door Attack Will Get You In
di: Nakash, Itay, et al.
Pubblicazione: (2024)
di: Nakash, Itay, et al.
Pubblicazione: (2024)
Effective Red-Teaming of Policy-Adherent Agents
di: Nakash, Itay, et al.
Pubblicazione: (2025)
di: Nakash, Itay, et al.
Pubblicazione: (2025)
Exploring Straightforward Conversational Red-Teaming
di: Kour, George, et al.
Pubblicazione: (2024)
di: Kour, George, et al.
Pubblicazione: (2024)
On the Robustness of Agentic Function Calling
di: Rabinovich, Ella, et al.
Pubblicazione: (2025)
di: Rabinovich, Ella, et al.
Pubblicazione: (2025)
From Zero to Hero: Cold-Start Anomaly Detection
di: Reiss, Tal, et al.
Pubblicazione: (2024)
di: Reiss, Tal, et al.
Pubblicazione: (2024)
What's the Plan? Evaluating and Developing Planning-Aware Techniques for Language Models
di: Hirsch, Eran, et al.
Pubblicazione: (2024)
di: Hirsch, Eran, et al.
Pubblicazione: (2024)
CRISP: Complex Reasoning with Interpretable Step-based Plans
di: Vetzler, Matan, et al.
Pubblicazione: (2025)
di: Vetzler, Matan, et al.
Pubblicazione: (2025)
A Novel Metric for Measuring the Robustness of Large Language Models in Non-adversarial Scenarios
di: Ackerman, Samuel, et al.
Pubblicazione: (2024)
di: Ackerman, Samuel, et al.
Pubblicazione: (2024)
ErrorMap and ErrorAtlas: Charting the Failure Landscape of Large Language Models
di: Ashury-Tahan, Shir, et al.
Pubblicazione: (2026)
di: Ashury-Tahan, Shir, et al.
Pubblicazione: (2026)
Agentic CLEAR: Automating Multi-Level Evaluation of LLM Agents
di: Yehudai, Asaf, et al.
Pubblicazione: (2026)
di: Yehudai, Asaf, et al.
Pubblicazione: (2026)
Near-Miss: Latent Policy Failure Detection in Agentic Workflows
di: Rabinovich, Ella, et al.
Pubblicazione: (2026)
di: Rabinovich, Ella, et al.
Pubblicazione: (2026)
SpeCrawler: Generating OpenAPI Specifications from API Documentation Using Large Language Models
di: Lazar, Koren, et al.
Pubblicazione: (2024)
di: Lazar, Koren, et al.
Pubblicazione: (2024)
CLEAR: Error Analysis via LLM-as-a-Judge Made Easy
di: Yehudai, Asaf, et al.
Pubblicazione: (2025)
di: Yehudai, Asaf, et al.
Pubblicazione: (2025)
Robustness as an Emergent Property of Task Performance
di: Ashury-Tahan, Shir, et al.
Pubblicazione: (2026)
di: Ashury-Tahan, Shir, et al.
Pubblicazione: (2026)
Cognitive-Mental-LLM: Evaluating Reasoning in Large Language Models for Mental Health Prediction via Online Text
di: Patil, Avinash, et al.
Pubblicazione: (2025)
di: Patil, Avinash, et al.
Pubblicazione: (2025)
Towards Enforcing Company Policy Adherence in Agentic Workflows
di: Zwerdling, Naama, et al.
Pubblicazione: (2025)
di: Zwerdling, Naama, et al.
Pubblicazione: (2025)
Efficient Benchmarking of Language Models
di: Perlitz, Yotam, et al.
Pubblicazione: (2023)
di: Perlitz, Yotam, et al.
Pubblicazione: (2023)
AdapThink: Adaptive Thinking Preferences for Reasoning Language Model
di: Wan, Xu, et al.
Pubblicazione: (2025)
di: Wan, Xu, et al.
Pubblicazione: (2025)
Ask Again, Then Fail: Large Language Models' Vacillations in Judgment
di: Xie, Qiming, et al.
Pubblicazione: (2023)
di: Xie, Qiming, et al.
Pubblicazione: (2023)
Towards Thinking-Optimal Scaling of Test-Time Compute for LLM Reasoning
di: Yang, Wenkai, et al.
Pubblicazione: (2025)
di: Yang, Wenkai, et al.
Pubblicazione: (2025)
OASBuilder: Generating OpenAPI Specifications from Online API Documentation with Large Language Models
di: Lazar, Koren, et al.
Pubblicazione: (2025)
di: Lazar, Koren, et al.
Pubblicazione: (2025)
Thinking Long, but Short: Stable Sequential Test-Time Scaling for Large Reasoning Models
di: Metel, Michael R., et al.
Pubblicazione: (2026)
di: Metel, Michael R., et al.
Pubblicazione: (2026)
Survey on Evaluation of LLM-based Agents
di: Yehudai, Asaf, et al.
Pubblicazione: (2025)
di: Yehudai, Asaf, et al.
Pubblicazione: (2025)
Provable Scaling Laws for the Test-Time Compute of Large Language Models
di: Chen, Yanxi, et al.
Pubblicazione: (2024)
di: Chen, Yanxi, et al.
Pubblicazione: (2024)
Test-Time Fairness and Robustness in Large Language Models
di: Cotta, Leonardo, et al.
Pubblicazione: (2024)
di: Cotta, Leonardo, et al.
Pubblicazione: (2024)
Preference Packing: Efficient Preference Optimization for Large Language Models
di: Cho, Jaekyung
Pubblicazione: (2026)
di: Cho, Jaekyung
Pubblicazione: (2026)
FairBelief -- Assessing Harmful Beliefs in Language Models
di: Setzu, Mattia, et al.
Pubblicazione: (2024)
di: Setzu, Mattia, et al.
Pubblicazione: (2024)
Think Carefully and Check Again! Meta-Generation Unlocking LLMs for Low-Resource Cross-Lingual Summarization
di: Li, Zhecheng, et al.
Pubblicazione: (2024)
di: Li, Zhecheng, et al.
Pubblicazione: (2024)
Exploring the Potential for Large Language Models to Demonstrate Rational Probabilistic Beliefs
di: Freedman, Gabriel, et al.
Pubblicazione: (2025)
di: Freedman, Gabriel, et al.
Pubblicazione: (2025)
Aligning Large Language Models with Searcher Preferences
di: Wu, Wei, et al.
Pubblicazione: (2026)
di: Wu, Wei, et al.
Pubblicazione: (2026)
Stay Tuned: An Empirical Study of the Impact of Hyperparameters on LLM Tuning in Real-World Applications
di: Halfon, Alon, et al.
Pubblicazione: (2024)
di: Halfon, Alon, et al.
Pubblicazione: (2024)
Test-Time Learning for Large Language Models
di: Hu, Jinwu, et al.
Pubblicazione: (2025)
di: Hu, Jinwu, et al.
Pubblicazione: (2025)
Training-Free Test-Time Contrastive Learning for Large Language Models
di: Zheng, Kaiwen, et al.
Pubblicazione: (2026)
di: Zheng, Kaiwen, et al.
Pubblicazione: (2026)
Using LLMs to Model the Beliefs and Preferences of Targeted Populations
di: Namikoshi, Keiichi, et al.
Pubblicazione: (2024)
di: Namikoshi, Keiichi, et al.
Pubblicazione: (2024)
Are You Sure? Rank Them Again: Repeated Ranking For Better Preference Datasets
di: Devine, Peter
Pubblicazione: (2024)
di: Devine, Peter
Pubblicazione: (2024)
Cognitive Decision Routing in Large Language Models: When to Think Fast, When to Think Slow
di: Du, Y., et al.
Pubblicazione: (2025)
di: Du, Y., et al.
Pubblicazione: (2025)
Weights-Rotated Preference Optimization for Large Language Models
di: Yang, Chenxu, et al.
Pubblicazione: (2025)
di: Yang, Chenxu, et al.
Pubblicazione: (2025)
Do Language Models Think Consistently? A Study of Value Preferences Across Varying Response Lengths
di: Nair, Inderjeet, et al.
Pubblicazione: (2025)
di: Nair, Inderjeet, et al.
Pubblicazione: (2025)
The Cost of Thinking: Increased Jailbreak Risk in Large Language Models
di: Yang, Fan
Pubblicazione: (2025)
di: Yang, Fan
Pubblicazione: (2025)
Documenti analoghi
-
Efficient Agent Evaluation via Diversity-Guided User Simulation
di: Nakash, Itay, et al.
Pubblicazione: (2026) -
Breaking ReAct Agents: Foot-in-the-Door Attack Will Get You In
di: Nakash, Itay, et al.
Pubblicazione: (2024) -
Effective Red-Teaming of Policy-Adherent Agents
di: Nakash, Itay, et al.
Pubblicazione: (2025) -
Exploring Straightforward Conversational Red-Teaming
di: Kour, George, et al.
Pubblicazione: (2024) -
On the Robustness of Agentic Function Calling
di: Rabinovich, Ella, et al.
Pubblicazione: (2025)