Agent Lumos: Unified and Modular Training for Open-Source Language Agents
Fuente:
arXiv
Salvato in:
| Autori principali: | Yin, Da, Brahman, Faeze, Ravichander, Abhilasha, Chandu, Khyathi, Chang, Kai-Wei, Choi, Yejin, Lin, Bill Yuchen |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
RESTOR: Knowledge Recovery in Machine Unlearning
di: Rezaei, Keivan, et al.
Pubblicazione: (2024)
di: Rezaei, Keivan, et al.
Pubblicazione: (2024)
WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
di: Lin, Bill Yuchen, et al.
Pubblicazione: (2024)
di: Lin, Bill Yuchen, et al.
Pubblicazione: (2024)
The Art of Saying No: Contextual Noncompliance in Language Models
di: Brahman, Faeze, et al.
Pubblicazione: (2024)
di: Brahman, Faeze, et al.
Pubblicazione: (2024)
Selective "Selective Prediction": Reducing Unnecessary Abstention in Vision-Language Reasoning
di: Srinivasan, Tejas, et al.
Pubblicazione: (2024)
di: Srinivasan, Tejas, et al.
Pubblicazione: (2024)
Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement
di: Jung, Jaehun, et al.
Pubblicazione: (2024)
di: Jung, Jaehun, et al.
Pubblicazione: (2024)
MacGyver: Are Large Language Models Creative Problem Solvers?
di: Tian, Yufei, et al.
Pubblicazione: (2023)
di: Tian, Yufei, et al.
Pubblicazione: (2023)
HALoGEN: Fantastic LLM Hallucinations and Where to Find Them
di: Ravichander, Abhilasha, et al.
Pubblicazione: (2025)
di: Ravichander, Abhilasha, et al.
Pubblicazione: (2025)
Information-Guided Identification of Training Data Imprint in (Proprietary) Large Language Models
di: Ravichander, Abhilasha, et al.
Pubblicazione: (2025)
di: Ravichander, Abhilasha, et al.
Pubblicazione: (2025)
L3GO: Language Agents with Chain-of-3D-Thoughts for Generating Unconventional Objects
di: Yamada, Yutaro, et al.
Pubblicazione: (2024)
di: Yamada, Yutaro, et al.
Pubblicazione: (2024)
WildHallucinations: Evaluating Long-form Factuality in LLMs with Real-World Entity Queries
di: Zhao, Wenting, et al.
Pubblicazione: (2024)
di: Zhao, Wenting, et al.
Pubblicazione: (2024)
How to Train Your Fact Verifier: Knowledge Transfer with Multimodal Open Models
di: Lee, Jaeyoung, et al.
Pubblicazione: (2024)
di: Lee, Jaeyoung, et al.
Pubblicazione: (2024)
Why and How LLMs Hallucinate: Connecting the Dots with Subsequence Associations
di: Sun, Yiyou, et al.
Pubblicazione: (2025)
di: Sun, Yiyou, et al.
Pubblicazione: (2025)
In Agents We Trust, but Who Do Agents Trust? Latent Source Preferences Steer LLM Generations
di: Khan, Mohammad Aflah, et al.
Pubblicazione: (2026)
di: Khan, Mohammad Aflah, et al.
Pubblicazione: (2026)
Tailoring with Targeted Precision: Edit-Based Agents for Open-Domain Procedure Customization
di: Lal, Yash Kumar, et al.
Pubblicazione: (2023)
di: Lal, Yash Kumar, et al.
Pubblicazione: (2023)
Artifacts or Abduction: How Do LLMs Answer Multiple-Choice Questions Without the Question?
di: Balepur, Nishant, et al.
Pubblicazione: (2024)
di: Balepur, Nishant, et al.
Pubblicazione: (2024)
What Has Been Lost with Synthetic Evaluation?
di: Gill, Alexander, et al.
Pubblicazione: (2025)
di: Gill, Alexander, et al.
Pubblicazione: (2025)
Scale Can't Overcome Pragmatics: The Impact of Reporting Bias on Vision-Language Reasoning
di: Kamath, Amita, et al.
Pubblicazione: (2026)
di: Kamath, Amita, et al.
Pubblicazione: (2026)
The Curious Case of Factuality Finetuning: Models' Internal Beliefs Can Improve Factuality
di: Newman, Benjamin, et al.
Pubblicazione: (2025)
di: Newman, Benjamin, et al.
Pubblicazione: (2025)
Information-Theoretic Distillation for Reference-less Summarization
di: Jung, Jaehun, et al.
Pubblicazione: (2024)
di: Jung, Jaehun, et al.
Pubblicazione: (2024)
OfficeBench: Benchmarking Language Agents across Multiple Applications for Office Automation
di: Wang, Zilong, et al.
Pubblicazione: (2024)
di: Wang, Zilong, et al.
Pubblicazione: (2024)
What Makes it Ok to Set a Fire? Iterative Self-distillation of Contexts and Rationales for Disambiguating Defeasible Social and Moral Situations
di: Rao, Kavel, et al.
Pubblicazione: (2023)
di: Rao, Kavel, et al.
Pubblicazione: (2023)
Train for Truth, Keep the Skills: Binary Retrieval-Augmented Reward Mitigates Hallucinations
di: Chen, Tong, et al.
Pubblicazione: (2025)
di: Chen, Tong, et al.
Pubblicazione: (2025)
Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents
di: Song, Yifan, et al.
Pubblicazione: (2024)
di: Song, Yifan, et al.
Pubblicazione: (2024)
AI-LieDar: Examine the Trade-off Between Utility and Truthfulness in LLM Agents
di: Su, Zhe, et al.
Pubblicazione: (2024)
di: Su, Zhe, et al.
Pubblicazione: (2024)
AI as Humanity's Salieri: Quantifying Linguistic Creativity of Language Models via Systematic Attribution of Machine Text against Web Text
di: Lu, Ximing, et al.
Pubblicazione: (2024)
di: Lu, Ximing, et al.
Pubblicazione: (2024)
Creativity Support in the Age of Large Language Models: An Empirical Study Involving Emerging Writers
di: Chakrabarty, Tuhin, et al.
Pubblicazione: (2023)
di: Chakrabarty, Tuhin, et al.
Pubblicazione: (2023)
Impossible Distillation: from Low-Quality Model to High-Quality Dataset & Model for Summarization and Paraphrasing
di: Jung, Jaehun, et al.
Pubblicazione: (2023)
di: Jung, Jaehun, et al.
Pubblicazione: (2023)
WildTeaming at Scale: From In-the-Wild Jailbreaks to (Adversarially) Safer Language Models
di: Jiang, Liwei, et al.
Pubblicazione: (2024)
di: Jiang, Liwei, et al.
Pubblicazione: (2024)
Deal, or no deal (or who knows)? Forecasting Uncertainty in Conversations using Large Language Models
di: Sicilia, Anthony, et al.
Pubblicazione: (2024)
di: Sicilia, Anthony, et al.
Pubblicazione: (2024)
Certainly Uncertain: A Benchmark and Metric for Multimodal Epistemic and Aleatoric Awareness
di: Chandu, Khyathi Raghavi, et al.
Pubblicazione: (2024)
di: Chandu, Khyathi Raghavi, et al.
Pubblicazione: (2024)
Reasoning Up the Instruction Ladder for Controllable Language Models
di: Zheng, Zishuo, et al.
Pubblicazione: (2025)
di: Zheng, Zishuo, et al.
Pubblicazione: (2025)
WildGuard: Open One-Stop Moderation Tools for Safety Risks, Jailbreaks, and Refusals of LLMs
di: Han, Seungju, et al.
Pubblicazione: (2024)
di: Han, Seungju, et al.
Pubblicazione: (2024)
The Surprising Effectiveness of Membership Inference with Simple N-Gram Coverage
di: Hallinan, Skyler, et al.
Pubblicazione: (2025)
di: Hallinan, Skyler, et al.
Pubblicazione: (2025)
Large-Scale Data Selection for Instruction Tuning
di: Ivison, Hamish, et al.
Pubblicazione: (2025)
di: Ivison, Hamish, et al.
Pubblicazione: (2025)
Fractional Rotation, Full Potential? Investigating Performance and Convergence of Partial RoPE
di: Khan, Mohammad Aflah, et al.
Pubblicazione: (2026)
di: Khan, Mohammad Aflah, et al.
Pubblicazione: (2026)
UNcommonsense Reasoning: Abductive Reasoning about Uncommon Situations
di: Zhao, Wenting, et al.
Pubblicazione: (2023)
di: Zhao, Wenting, et al.
Pubblicazione: (2023)
Husky: A Unified, Open-Source Language Agent for Multi-Step Reasoning
di: Kim, Joongwon, et al.
Pubblicazione: (2024)
di: Kim, Joongwon, et al.
Pubblicazione: (2024)
Leftover Lunch: Advantage-based Offline Reinforcement Learning for Language Models
di: Baheti, Ashutosh, et al.
Pubblicazione: (2023)
di: Baheti, Ashutosh, et al.
Pubblicazione: (2023)
WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences
di: Lu, Yujie, et al.
Pubblicazione: (2024)
di: Lu, Yujie, et al.
Pubblicazione: (2024)
PlaSma: Making Small Language Models Better Procedural Knowledge Models for (Counterfactual) Planning
di: Brahman, Faeze, et al.
Pubblicazione: (2023)
di: Brahman, Faeze, et al.
Pubblicazione: (2023)
Documenti analoghi
-
RESTOR: Knowledge Recovery in Machine Unlearning
di: Rezaei, Keivan, et al.
Pubblicazione: (2024) -
WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
di: Lin, Bill Yuchen, et al.
Pubblicazione: (2024) -
The Art of Saying No: Contextual Noncompliance in Language Models
di: Brahman, Faeze, et al.
Pubblicazione: (2024) -
Selective "Selective Prediction": Reducing Unnecessary Abstention in Vision-Language Reasoning
di: Srinivasan, Tejas, et al.
Pubblicazione: (2024) -
Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement
di: Jung, Jaehun, et al.
Pubblicazione: (2024)