TOOLVERIFIER: Generalization to New Tools via Self-Verification
Fuente:
arXiv
Guardado en:
| Autores principales: | Mekala, Dheeraj, Weston, Jason, Lanchantin, Jack, Raileanu, Roberta, Lomeli, Maria, Shang, Jingbo, Dwivedi-Yu, Jane |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Source2Synth: Synthetic Data Generation and Curation Grounded in Real Data Sources
por: Lupidi, Alisia, et al.
Publicado: (2024)
por: Lupidi, Alisia, et al.
Publicado: (2024)
Smaller Language Models are capable of selecting Instruction-Tuning Training Data for Larger Language Models
por: Mekala, Dheeraj, et al.
Publicado: (2024)
por: Mekala, Dheeraj, et al.
Publicado: (2024)
DOCMASTER: A Unified Platform for Annotation, Training, & Inference in Document Question-Answering
por: Nguyen, Alex, et al.
Publicado: (2024)
por: Nguyen, Alex, et al.
Publicado: (2024)
When is the consistent prediction likely to be a correct prediction?
por: Nguyen, Alex, et al.
Publicado: (2024)
por: Nguyen, Alex, et al.
Publicado: (2024)
TIER: Trajectory-Invariant Execution Rewards for Multi-Step Tool Composition
por: Kulkarni, Anay, et al.
Publicado: (2026)
por: Kulkarni, Anay, et al.
Publicado: (2026)
Adaptive Decoding via Latent Preference Optimization
por: Dhuliawala, Shehzaad, et al.
Publicado: (2024)
por: Dhuliawala, Shehzaad, et al.
Publicado: (2024)
Diverse Preference Optimization
por: Lanchantin, Jack, et al.
Publicado: (2025)
por: Lanchantin, Jack, et al.
Publicado: (2025)
Jointly Reinforcing Diversity and Quality in Language Model Generations
por: Li, Tianjian, et al.
Publicado: (2025)
por: Li, Tianjian, et al.
Publicado: (2025)
GLoRe: When, Where, and How to Improve LLM Reasoning via Global and Local Refinements
por: Havrilla, Alex, et al.
Publicado: (2024)
por: Havrilla, Alex, et al.
Publicado: (2024)
CoT-Self-Instruct: Building high-quality synthetic prompts for reasoning and non-reasoning tasks
por: Yu, Ping, et al.
Publicado: (2025)
por: Yu, Ping, et al.
Publicado: (2025)
SPICE: Self-Play In Corpus Environments Improves Reasoning
por: Liu, Bo, et al.
Publicado: (2025)
por: Liu, Bo, et al.
Publicado: (2025)
LLM-First Search: Self-Guided Exploration of the Solution Space
por: Herr, Nathan, et al.
Publicado: (2025)
por: Herr, Nathan, et al.
Publicado: (2025)
OptimalThinkingBench: Evaluating Over and Underthinking in LLMs
por: Aggarwal, Pranjal, et al.
Publicado: (2025)
por: Aggarwal, Pranjal, et al.
Publicado: (2025)
NaturalThoughts: Selecting and Distilling Reasoning Traces for General Reasoning Tasks
por: Li, Yang, et al.
Publicado: (2025)
por: Li, Yang, et al.
Publicado: (2025)
LLM Pretraining with Continuous Concepts
por: Tack, Jihoon, et al.
Publicado: (2025)
por: Tack, Jihoon, et al.
Publicado: (2025)
MORL-Prompt: An Empirical Analysis of Multi-Objective Reinforcement Learning for Discrete Prompt Optimization
por: Jafari, Yasaman, et al.
Publicado: (2024)
por: Jafari, Yasaman, et al.
Publicado: (2024)
Self-Taught Evaluators
por: Wang, Tianlu, et al.
Publicado: (2024)
por: Wang, Tianlu, et al.
Publicado: (2024)
Self-Improving Pretraining: using post-trained models to pretrain better models
por: Tan, Ellen Xiaoqing, et al.
Publicado: (2026)
por: Tan, Ellen Xiaoqing, et al.
Publicado: (2026)
Bridging Offline and Online Reinforcement Learning for LLMs
por: Lanchantin, Jack, et al.
Publicado: (2025)
por: Lanchantin, Jack, et al.
Publicado: (2025)
FairPair: A Robust Evaluation of Biases in Language Models through Paired Perturbations
por: Dwivedi-Yu, Jane, et al.
Publicado: (2024)
por: Dwivedi-Yu, Jane, et al.
Publicado: (2024)
UltraGen: Extremely Fine-grained Controllable Generation via Attribute Reconstruction and Global Preference Optimization
por: Yun, Longfei, et al.
Publicado: (2025)
por: Yun, Longfei, et al.
Publicado: (2025)
DreamCraft: Text-Guided Generation of Functional 3D Environments in Minecraft
por: Earle, Sam, et al.
Publicado: (2024)
por: Earle, Sam, et al.
Publicado: (2024)
Watermarks for Language Models via Probabilistic Automata
por: Wang, Yangkun, et al.
Publicado: (2025)
por: Wang, Yangkun, et al.
Publicado: (2025)
Sparks of Science: Hypothesis Generation Using Structured Paper Data
por: O'Neill, Charles, et al.
Publicado: (2025)
por: O'Neill, Charles, et al.
Publicado: (2025)
Are Large Language Models Strategic Decision Makers? A Study of Performance and Bias in Two-Player Non-Zero-Sum Games
por: Herr, Nathan, et al.
Publicado: (2024)
por: Herr, Nathan, et al.
Publicado: (2024)
Incubating Text Classifiers Following User Instruction with Nothing but LLM
por: Peng, Letian, et al.
Publicado: (2024)
por: Peng, Letian, et al.
Publicado: (2024)
Quantifying and Optimizing Global Faithfulness in Persona-driven Role-playing
por: Peng, Letian, et al.
Publicado: (2024)
por: Peng, Letian, et al.
Publicado: (2024)
Codifying Character Logic in Role-Playing
por: Peng, Letian, et al.
Publicado: (2025)
por: Peng, Letian, et al.
Publicado: (2025)
Evaluating the Smooth Control of Attribute Intensity in Text Generation with LLMs
por: Zhou, Shang, et al.
Publicado: (2024)
por: Zhou, Shang, et al.
Publicado: (2024)
Efficient Tool Use with Chain-of-Abstraction Reasoning
por: Gao, Silin, et al.
Publicado: (2024)
por: Gao, Silin, et al.
Publicado: (2024)
Self-Alignment with Instruction Backtranslation
por: Li, Xian, et al.
Publicado: (2023)
por: Li, Xian, et al.
Publicado: (2023)
Improving Language Plasticity via Pretraining with Active Forgetting
por: Chen, Yihong, et al.
Publicado: (2023)
por: Chen, Yihong, et al.
Publicado: (2023)
Self-Challenging Language Model Agents
por: Zhou, Yifei, et al.
Publicado: (2025)
por: Zhou, Yifei, et al.
Publicado: (2025)
Task-Dependent Evaluation of LLM Output Homogenization: A Taxonomy-Guided Framework
por: Jain, Shomik, et al.
Publicado: (2025)
por: Jain, Shomik, et al.
Publicado: (2025)
Correlation and Navigation in the Vocabulary Key Representation Space of Language Models
por: Peng, Letian, et al.
Publicado: (2024)
por: Peng, Letian, et al.
Publicado: (2024)
Self-Consistency Preference Optimization
por: Prasad, Archiki, et al.
Publicado: (2024)
por: Prasad, Archiki, et al.
Publicado: (2024)
Linear Correlation in LM's Compositional Generalization and Hallucination
por: Peng, Letian, et al.
Publicado: (2025)
por: Peng, Letian, et al.
Publicado: (2025)
RESTRAIN: From Spurious Votes to Signals -- Self-Driven RL with Self-Penalization
por: Yu, Zhaoning, et al.
Publicado: (2025)
por: Yu, Zhaoning, et al.
Publicado: (2025)
Distilling System 2 into System 1
por: Yu, Ping, et al.
Publicado: (2024)
por: Yu, Ping, et al.
Publicado: (2024)
Codified Foreshadowing-Payoff Text Generation
por: Yun, Longfei, et al.
Publicado: (2026)
por: Yun, Longfei, et al.
Publicado: (2026)
Ejemplares similares
-
Source2Synth: Synthetic Data Generation and Curation Grounded in Real Data Sources
por: Lupidi, Alisia, et al.
Publicado: (2024) -
Smaller Language Models are capable of selecting Instruction-Tuning Training Data for Larger Language Models
por: Mekala, Dheeraj, et al.
Publicado: (2024) -
DOCMASTER: A Unified Platform for Annotation, Training, & Inference in Document Question-Answering
por: Nguyen, Alex, et al.
Publicado: (2024) -
When is the consistent prediction likely to be a correct prediction?
por: Nguyen, Alex, et al.
Publicado: (2024) -
TIER: Trajectory-Invariant Execution Rewards for Multi-Step Tool Composition
por: Kulkarni, Anay, et al.
Publicado: (2026)