GuidedSampling: Steering LLMs Towards Diverse Candidate Solutions at Inference-Time
Fuente:
arXiv
Guardado en:
| Autores principales: | Handa, Divij, Parmar, Mihir, RRV, Aswin, Uddin, Md Nayem, Palangi, Hamid, Baral, Chitta |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
ThinkTuning: Instilling Cognitive Reflections without Distillation
por: RRV, Aswin, et al.
Publicado: (2025)
por: RRV, Aswin, et al.
Publicado: (2025)
When "Competency" in Reasoning Opens the Door to Vulnerability: Jailbreaking LLMs via Novel Complex Ciphers
por: Handa, Divij, et al.
Publicado: (2024)
por: Handa, Divij, et al.
Publicado: (2024)
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models
por: RRV, Aswin, et al.
Publicado: (2026)
por: RRV, Aswin, et al.
Publicado: (2026)
Chaos with Keywords: Exposing Large Language Models Sycophantic Hallucination to Misleading Keywords and Evaluating Defense Strategies
por: RRV, Aswin, et al.
Publicado: (2024)
por: RRV, Aswin, et al.
Publicado: (2024)
Step-by-Step Reasoning to Solve Grid Puzzles: Where do LLMs Falter?
por: Tyagi, Nemika, et al.
Publicado: (2024)
por: Tyagi, Nemika, et al.
Publicado: (2024)
PHANTOM RECALL: When Familiar Puzzles Fool Smart Models
por: Mukhopadhyay, Souradeep, et al.
Publicado: (2025)
por: Mukhopadhyay, Souradeep, et al.
Publicado: (2025)
UnSeenTimeQA: Time-Sensitive Question-Answering Beyond LLMs' Memorization
por: Uddin, Md Nayem, et al.
Publicado: (2024)
por: Uddin, Md Nayem, et al.
Publicado: (2024)
PLAN-TUNING: Post-Training Language Models to Learn Step-by-Step Planning for Complex Problem Solving
por: Parmar, Mihir, et al.
Publicado: (2025)
por: Parmar, Mihir, et al.
Publicado: (2025)
ActionReasoningBench: Reasoning about Actions with and without Ramification Constraints
por: Handa, Divij, et al.
Publicado: (2024)
por: Handa, Divij, et al.
Publicado: (2024)
Don't Blame the Annotator: Bias Already Starts in the Annotation Instructions
por: Parmar, Mihir, et al.
Publicado: (2022)
por: Parmar, Mihir, et al.
Publicado: (2022)
Insights into Alignment: Evaluating DPO and its Variants Across Multiple Tasks
por: Saeidi, Amir, et al.
Publicado: (2024)
por: Saeidi, Amir, et al.
Publicado: (2024)
Triple Preference Optimization: Achieving Better Alignment using a Single Step Optimization
por: Saeidi, Amir, et al.
Publicado: (2024)
por: Saeidi, Amir, et al.
Publicado: (2024)
Multi-LogiEval: Towards Evaluating Multi-Step Logical Reasoning Ability of Large Language Models
por: Patel, Nisarg, et al.
Publicado: (2024)
por: Patel, Nisarg, et al.
Publicado: (2024)
Hypothesis Generation for Materials Discovery and Design Using Goal-Driven and Constraint-Guided LLM Agents
por: Kumbhar, Shrinidhi, et al.
Publicado: (2025)
por: Kumbhar, Shrinidhi, et al.
Publicado: (2025)
Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution
por: Dineen, Jacob, et al.
Publicado: (2026)
por: Dineen, Jacob, et al.
Publicado: (2026)
From Recall to Forgetting: Benchmarking Long-Term Memory for Personalized Agents
por: Uddin, Md Nayem, et al.
Publicado: (2026)
por: Uddin, Md Nayem, et al.
Publicado: (2026)
LogicBench: Towards Systematic Evaluation of Logical Reasoning Ability of Large Language Models
por: Parmar, Mihir, et al.
Publicado: (2024)
por: Parmar, Mihir, et al.
Publicado: (2024)
ScholarPeer: A Context-Aware Multi-Agent Framework for Automated Peer Review
por: Goyal, Palash, et al.
Publicado: (2026)
por: Goyal, Palash, et al.
Publicado: (2026)
Towards LogiGLUE: A Brief Survey and A Benchmark for Analyzing Logical Reasoning Capabilities of Language Models
por: Luo, Man, et al.
Publicado: (2023)
por: Luo, Man, et al.
Publicado: (2023)
Polymath: A Challenging Multi-modal Mathematical Reasoning Benchmark
por: Gupta, Himanshu, et al.
Publicado: (2024)
por: Gupta, Himanshu, et al.
Publicado: (2024)
Diversity of Thought Improves Reasoning Abilities of LLMs
por: Naik, Ranjita, et al.
Publicado: (2023)
por: Naik, Ranjita, et al.
Publicado: (2023)
PlanGEN: A Multi-Agent Framework for Generating Planning and Reasoning Trajectories for Complex Problem Solving
por: Parmar, Mihir, et al.
Publicado: (2025)
por: Parmar, Mihir, et al.
Publicado: (2025)
VeriGuard: Enhancing LLM Agent Safety via Verified Code Generation
por: Miculicich, Lesly, et al.
Publicado: (2025)
por: Miculicich, Lesly, et al.
Publicado: (2025)
OptAgent: Optimizing Query Rewriting for E-commerce via Multi-Agent Simulation
por: Handa, Divij, et al.
Publicado: (2025)
por: Handa, Divij, et al.
Publicado: (2025)
LLM-Based Multi-Agent Blackboard System for Information Discovery in Data Science
por: Salemi, Alireza, et al.
Publicado: (2025)
por: Salemi, Alireza, et al.
Publicado: (2025)
TFRBench: A Reasoning Benchmark for Evaluating Forecasting Systems
por: Ahamed, Md Atik, et al.
Publicado: (2026)
por: Ahamed, Md Atik, et al.
Publicado: (2026)
Investigating the Shortcomings of LLMs in Step-by-Step Legal Reasoning
por: Mishra, Venkatesh, et al.
Publicado: (2025)
por: Mishra, Venkatesh, et al.
Publicado: (2025)
QA-LIGN: Aligning LLMs through Constitutionally Decomposed QA
por: Dineen, Jacob, et al.
Publicado: (2025)
por: Dineen, Jacob, et al.
Publicado: (2025)
Towards Enhancing Coherence in Extractive Summarization: Dataset and Experiments with LLMs
por: Parmar, Mihir, et al.
Publicado: (2024)
por: Parmar, Mihir, et al.
Publicado: (2024)
ToW: Thoughts of Words Improve Reasoning in Large Language Models
por: Xu, Zhikun, et al.
Publicado: (2024)
por: Xu, Zhikun, et al.
Publicado: (2024)
When Can LLMs Learn to Reason with Weak Supervision?
por: Rahman, Salman, et al.
Publicado: (2026)
por: Rahman, Salman, et al.
Publicado: (2026)
InfSplign: Inference-Time Spatial Alignment of Text-to-Image Diffusion Models
por: Rastegar, Sarah, et al.
Publicado: (2025)
por: Rastegar, Sarah, et al.
Publicado: (2025)
MMTABREAL: Real-World Benchmark for Multimodal Table Understanding
por: Titiya, Prasham, et al.
Publicado: (2025)
por: Titiya, Prasham, et al.
Publicado: (2025)
Fints: Efficient Inference-Time Personalization for LLMs with Fine-Grained Instance-Tailored Steering
por: Du, Kounianhua, et al.
Publicado: (2025)
por: Du, Kounianhua, et al.
Publicado: (2025)
Low-Rank Adaptation of Time Series Foundational Models for Out-of-Domain Modality Forecasting
por: Gupta, Divij, et al.
Publicado: (2024)
por: Gupta, Divij, et al.
Publicado: (2024)
Reasoning-Aware Training for Time Series Forecasting
por: Ahamed, Md Atik, et al.
Publicado: (2026)
por: Ahamed, Md Atik, et al.
Publicado: (2026)
Beyond LoRA: Exploring Efficient Fine-Tuning Techniques for Time Series Foundational Models
por: Gupta, Divij, et al.
Publicado: (2024)
por: Gupta, Divij, et al.
Publicado: (2024)
HEART: Emotionally-Driven Test-Time Scaling of Language Models
por: Pinto, Gabriela, et al.
Publicado: (2025)
por: Pinto, Gabriela, et al.
Publicado: (2025)
GrAInS: Gradient-based Attribution for Inference-Time Steering of LLMs and VLMs
por: Nguyen, Duy, et al.
Publicado: (2025)
por: Nguyen, Duy, et al.
Publicado: (2025)
Latent Reward Steering: An Adaptive Inference-Time Framework that Implicitly Promotes Cognitive Behaviors in Reasoning LLMs
por: Li, Jiakang, et al.
Publicado: (2026)
por: Li, Jiakang, et al.
Publicado: (2026)
Ejemplares similares
-
ThinkTuning: Instilling Cognitive Reflections without Distillation
por: RRV, Aswin, et al.
Publicado: (2025) -
When "Competency" in Reasoning Opens the Door to Vulnerability: Jailbreaking LLMs via Novel Complex Ciphers
por: Handa, Divij, et al.
Publicado: (2024) -
Mid-Training with Self-Generated Data Improves Reinforcement Learning in Language Models
por: RRV, Aswin, et al.
Publicado: (2026) -
Chaos with Keywords: Exposing Large Language Models Sycophantic Hallucination to Misleading Keywords and Evaluating Defense Strategies
por: RRV, Aswin, et al.
Publicado: (2024) -
Step-by-Step Reasoning to Solve Grid Puzzles: Where do LLMs Falter?
por: Tyagi, Nemika, et al.
Publicado: (2024)