Towards Practical Tool Usage for Continually Learning LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Huang, Jerry, Parthasarathi, Prasanna, Rezagholizadeh, Mehdi, Chandar, Sarath |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Do Robot Snakes Dream like Electric Sheep? Investigating the Effects of Architectural Inductive Biases on Hallucination
por: Huang, Jerry, et al.
Publicado: (2024)
por: Huang, Jerry, et al.
Publicado: (2024)
Context-Aware Assistant Selection for Improved Inference Acceleration with Large Language Models
por: Huang, Jerry, et al.
Publicado: (2024)
por: Huang, Jerry, et al.
Publicado: (2024)
Do Large Language Models Know How Much They Know?
por: Prato, Gabriele, et al.
Publicado: (2025)
por: Prato, Gabriele, et al.
Publicado: (2025)
EpiK-Eval: Evaluation for Language Models as Epistemic Models
por: Prato, Gabriele, et al.
Publicado: (2023)
por: Prato, Gabriele, et al.
Publicado: (2023)
GRPO-$λ$: Credit Assignment improves LLM Reasoning
por: Parthasarathi, Prasanna, et al.
Publicado: (2025)
por: Parthasarathi, Prasanna, et al.
Publicado: (2025)
Are self-explanations from Large Language Models faithful?
por: Madsen, Andreas, et al.
Publicado: (2024)
por: Madsen, Andreas, et al.
Publicado: (2024)
Effect of Document Packing on the Latent Multi-Hop Reasoning Capabilities of Large Language Models
por: Prato, Gabriele, et al.
Publicado: (2025)
por: Prato, Gabriele, et al.
Publicado: (2025)
Nested-ReFT: Efficient Reinforcement Learning for Large Language Model Fine-Tuning via Off-Policy Rollouts
por: Heuillet, Maxime, et al.
Publicado: (2025)
por: Heuillet, Maxime, et al.
Publicado: (2025)
Why Don't Prompt-Based Fairness Metrics Correlate?
por: Zayed, Abdelrahman, et al.
Publicado: (2024)
por: Zayed, Abdelrahman, et al.
Publicado: (2024)
Should We Attend More or Less? Modulating Attention for Fairness
por: Zayed, Abdelrahman, et al.
Publicado: (2023)
por: Zayed, Abdelrahman, et al.
Publicado: (2023)
Revisiting Replay and Gradient Alignment for Continual Pre-Training of Large Language Models
por: Abbes, Istabrak, et al.
Publicado: (2025)
por: Abbes, Istabrak, et al.
Publicado: (2025)
Too Big to Fool: Resisting Deception in Language Models
por: Samsami, Mohammad Reza, et al.
Publicado: (2024)
por: Samsami, Mohammad Reza, et al.
Publicado: (2024)
The Markovian Thinker: Architecture-Agnostic Linear Scaling of Reasoning
por: Aghajohari, Milad, et al.
Publicado: (2025)
por: Aghajohari, Milad, et al.
Publicado: (2025)
Towards Optimizing the Costs of LLM Usage
por: Shekhar, Shivanshu, et al.
Publicado: (2024)
por: Shekhar, Shivanshu, et al.
Publicado: (2024)
Balcony: A Lightweight Approach to Dynamic Inference of Generative Language Models
por: Jamialahmadi, Benyamin, et al.
Publicado: (2025)
por: Jamialahmadi, Benyamin, et al.
Publicado: (2025)
Small Encoders Can Rival Large Decoders in Detecting Groundedness
por: Abbes, Istabrak, et al.
Publicado: (2025)
por: Abbes, Istabrak, et al.
Publicado: (2025)
Tool Unlearning for Tool-Augmented LLMs
por: Cheng, Jiali, et al.
Publicado: (2025)
por: Cheng, Jiali, et al.
Publicado: (2025)
How Well Can a Long Sequence Model Model Long Sequences? Comparing Architechtural Inductive Biases on Long-Context Abilities
por: Huang, Jerry
Publicado: (2024)
por: Huang, Jerry
Publicado: (2024)
Stabilizing Reinforcement Learning with LLMs: Formulation and Practices
por: Zheng, Chujie, et al.
Publicado: (2025)
por: Zheng, Chujie, et al.
Publicado: (2025)
LLMem: Estimating GPU Memory Usage for Fine-Tuning Pre-Trained LLMs
por: Kim, Taeho, et al.
Publicado: (2024)
por: Kim, Taeho, et al.
Publicado: (2024)
LLMs in the Imaginarium: Tool Learning through Simulated Trial and Error
por: Wang, Boshi, et al.
Publicado: (2024)
por: Wang, Boshi, et al.
Publicado: (2024)
LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws
por: Mayilvahanan, Prasanna, et al.
Publicado: (2025)
por: Mayilvahanan, Prasanna, et al.
Publicado: (2025)
Adapting While Learning: Grounding LLMs for Scientific Problems with Intelligent Tool Usage Adaptation
por: Lyu, Bohan, et al.
Publicado: (2024)
por: Lyu, Bohan, et al.
Publicado: (2024)
Characterizing Large Language Model Geometry Helps Solve Toxicity Detection and Generation
por: Balestriero, Randall, et al.
Publicado: (2023)
por: Balestriero, Randall, et al.
Publicado: (2023)
SPARC: Subspace-Aware Prompt Adaptation for Robust Continual Learning in LLMs
por: Jayasuriya, Dinithi, et al.
Publicado: (2025)
por: Jayasuriya, Dinithi, et al.
Publicado: (2025)
PoTPTQ: A Two-step Power-of-Two Post-training for LLMs
por: Wang, Xinyu, et al.
Publicado: (2025)
por: Wang, Xinyu, et al.
Publicado: (2025)
Tool Learning with Foundation Models
por: Qin, Yujia, et al.
Publicado: (2023)
por: Qin, Yujia, et al.
Publicado: (2023)
Faithfulness Measurable Masked Language Models
por: Madsen, Andreas, et al.
Publicado: (2023)
por: Madsen, Andreas, et al.
Publicado: (2023)
ToolACE-R: Model-aware Iterative Training and Adaptive Refinement for Tool Learning
por: Zeng, Xingshan, et al.
Publicado: (2025)
por: Zeng, Xingshan, et al.
Publicado: (2025)
True Knowledge Comes from Practice: Aligning LLMs with Embodied Environments via Reinforcement Learning
por: Tan, Weihao, et al.
Publicado: (2024)
por: Tan, Weihao, et al.
Publicado: (2024)
ToolRL: Reward is All Tool Learning Needs
por: Qian, Cheng, et al.
Publicado: (2025)
por: Qian, Cheng, et al.
Publicado: (2025)
Squeezing More from the Stream : Learning Representation Online for Streaming Reinforcement Learning
por: Nilaksh, et al.
Publicado: (2026)
por: Nilaksh, et al.
Publicado: (2026)
Are LLMs Ready for Neural-integrated Mechanistic Modeling? A Benchmark and Agentic Framework
por: Guan, Zihan, et al.
Publicado: (2026)
por: Guan, Zihan, et al.
Publicado: (2026)
MINT: Evaluating LLMs in Multi-turn Interaction with Tools and Language Feedback
por: Wang, Xingyao, et al.
Publicado: (2023)
por: Wang, Xingyao, et al.
Publicado: (2023)
CACTUS: Chemistry Agent Connecting Tool-Usage to Science
por: McNaughton, Andrew D., et al.
Publicado: (2024)
por: McNaughton, Andrew D., et al.
Publicado: (2024)
Beyond Mode-Seeking RL: Trajectory-Balance Post-Training for Diffusion Language Models
por: Ahmadi, Saba, et al.
Publicado: (2026)
por: Ahmadi, Saba, et al.
Publicado: (2026)
AgentKernelArena: Generalization-Aware Benchmarking of GPU Kernel Optimization Agents
por: Younesian, Sharareh, et al.
Publicado: (2026)
por: Younesian, Sharareh, et al.
Publicado: (2026)
LABO: Towards Learning Optimal Label Regularization via Bi-level Optimization
por: Lu, Peng, et al.
Publicado: (2023)
por: Lu, Peng, et al.
Publicado: (2023)
Unified Tool Integration for LLMs: A Protocol-Agnostic Approach to Function Calling
por: Ding, Peng, et al.
Publicado: (2025)
por: Ding, Peng, et al.
Publicado: (2025)
WikiContradict: A Benchmark for Evaluating LLMs on Real-World Knowledge Conflicts from Wikipedia
por: Hou, Yufang, et al.
Publicado: (2024)
por: Hou, Yufang, et al.
Publicado: (2024)
Ejemplares similares
-
Do Robot Snakes Dream like Electric Sheep? Investigating the Effects of Architectural Inductive Biases on Hallucination
por: Huang, Jerry, et al.
Publicado: (2024) -
Context-Aware Assistant Selection for Improved Inference Acceleration with Large Language Models
por: Huang, Jerry, et al.
Publicado: (2024) -
Do Large Language Models Know How Much They Know?
por: Prato, Gabriele, et al.
Publicado: (2025) -
EpiK-Eval: Evaluation for Language Models as Epistemic Models
por: Prato, Gabriele, et al.
Publicado: (2023) -
GRPO-$λ$: Credit Assignment improves LLM Reasoning
por: Parthasarathi, Prasanna, et al.
Publicado: (2025)