SwissNYF: Tool Grounded LLM Agents for Black Box Setting
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Kumar, Somnath Sendhil, Jain, Dhruv, Agarwal, Eshaan, Pandey, Raunak |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SoK: Measuring What Matters for Closed-Loop Security Agents
von: Khurana, Mudita, et al.
Veröffentlicht: (2025)
von: Khurana, Mudita, et al.
Veröffentlicht: (2025)
From Sycophancy to Sensemaking: Premise Governance for Human-AI Decision Making
von: Jain, Raunak
Veröffentlicht: (2026)
von: Jain, Raunak
Veröffentlicht: (2026)
Collaborative Causal Sensemaking: Closing the Complementarity Gap in Human-AI Decision Support
von: Jain, Raunak
Veröffentlicht: (2025)
von: Jain, Raunak
Veröffentlicht: (2025)
VidModEx: Interpretable and Efficient Black Box Model Extraction for High-Dimensional Spaces
von: Kumar, Somnath Sendhil, et al.
Veröffentlicht: (2024)
von: Kumar, Somnath Sendhil, et al.
Veröffentlicht: (2024)
Diagnosing LLM Judge Reliability: Conformal Prediction Sets and Transitivity Violations
von: Gupta, Manan, et al.
Veröffentlicht: (2026)
von: Gupta, Manan, et al.
Veröffentlicht: (2026)
Large Language Models as Pokémon Battle Agents: Strategic Play and Content Generation
von: Jain, Daksh, et al.
Veröffentlicht: (2025)
von: Jain, Daksh, et al.
Veröffentlicht: (2025)
LUDOBENCH: Evaluating LLM Behavioural Decision-Making Through Spot-Based Board Game Scenarios in Ludo
von: Jain, Ojas, et al.
Veröffentlicht: (2026)
von: Jain, Ojas, et al.
Veröffentlicht: (2026)
VoiceAgentBench: Are Voice Assistants ready for agentic tasks?
von: Jain, Dhruv, et al.
Veröffentlicht: (2025)
von: Jain, Dhruv, et al.
Veröffentlicht: (2025)
Beyond Sentiment: A Multi-Agent Pipeline for Actionable Business Advice from Reviews
von: Bhandari, Kartikey Singh, et al.
Veröffentlicht: (2026)
von: Bhandari, Kartikey Singh, et al.
Veröffentlicht: (2026)
Conformal Sets in Multiple-Choice Question Answering under Black-Box Settings with Provable Coverage Guarantees
von: Yang, Guang, et al.
Veröffentlicht: (2025)
von: Yang, Guang, et al.
Veröffentlicht: (2025)
Do Large Language Models Perform the Way People Expect? Measuring the Human Generalization Function
von: Vafa, Keyon, et al.
Veröffentlicht: (2024)
von: Vafa, Keyon, et al.
Veröffentlicht: (2024)
From Black Boxes to Conversations: Incorporating XAI in a Conversational Agent
von: Nguyen, Van Bach, et al.
Veröffentlicht: (2022)
von: Nguyen, Van Bach, et al.
Veröffentlicht: (2022)
In-Context Explainers: Harnessing LLMs for Explaining Black Box Models
von: Kroeger, Nicholas, et al.
Veröffentlicht: (2023)
von: Kroeger, Nicholas, et al.
Veröffentlicht: (2023)
Evaluating Novelty in AI-Generated Research Plans Using Multi-Workflow LLM Pipelines
von: Saraogi, Devesh, et al.
Veröffentlicht: (2025)
von: Saraogi, Devesh, et al.
Veröffentlicht: (2025)
GroundAct: Can LLM Agents Ground Actions in Environmental States?
von: Wang, Zixuan, et al.
Veröffentlicht: (2025)
von: Wang, Zixuan, et al.
Veröffentlicht: (2025)
LLM-as-a-Judge for Time Series Explanations
von: Sivalingam, Preetham, et al.
Veröffentlicht: (2026)
von: Sivalingam, Preetham, et al.
Veröffentlicht: (2026)
Generation Z's Ability to Discriminate Between AI-generated and Human-Authored Text on Discord
von: Ramu, Dhruv, et al.
Veröffentlicht: (2023)
von: Ramu, Dhruv, et al.
Veröffentlicht: (2023)
PromptWizard: Task-Aware Prompt Optimization Framework
von: Agarwal, Eshaan, et al.
Veröffentlicht: (2024)
von: Agarwal, Eshaan, et al.
Veröffentlicht: (2024)
Investigating Implicit Bias in Large Language Models: A Large-Scale Study of Over 50 LLMs
von: Kumar, Divyanshu, et al.
Veröffentlicht: (2024)
von: Kumar, Divyanshu, et al.
Veröffentlicht: (2024)
SLIDE: Reference-free Evaluation for Machine Translation using a Sliding Document Window
von: Raunak, Vikas, et al.
Veröffentlicht: (2023)
von: Raunak, Vikas, et al.
Veröffentlicht: (2023)
Confidence is Not Competence
von: Sanyal, Debdeep, et al.
Veröffentlicht: (2025)
von: Sanyal, Debdeep, et al.
Veröffentlicht: (2025)
A Hybrid Supervised-LLM Pipeline for Actionable Suggestion Mining in Unstructured Customer Reviews
von: Trivedi, Aakash, et al.
Veröffentlicht: (2026)
von: Trivedi, Aakash, et al.
Veröffentlicht: (2026)
HYDRA: Model Factorization Framework for Black-Box LLM Personalization
von: Zhuang, Yuchen, et al.
Veröffentlicht: (2024)
von: Zhuang, Yuchen, et al.
Veröffentlicht: (2024)
Potemkin Understanding in Large Language Models
von: Mancoridis, Marina, et al.
Veröffentlicht: (2025)
von: Mancoridis, Marina, et al.
Veröffentlicht: (2025)
Swiss-Bench 003: Evaluating LLM Reliability and Adversarial Security for Swiss Regulatory Contexts
von: Uenal, Fatih
Veröffentlicht: (2026)
von: Uenal, Fatih
Veröffentlicht: (2026)
On Instruction-Finetuning Neural Machine Translation Models
von: Raunak, Vikas, et al.
Veröffentlicht: (2024)
von: Raunak, Vikas, et al.
Veröffentlicht: (2024)
Distilling LLM Agent into Small Models with Retrieval and Code Tools
von: Kang, Minki, et al.
Veröffentlicht: (2025)
von: Kang, Minki, et al.
Veröffentlicht: (2025)
Uncertainty Quantification for Language Models: A Suite of Black-Box, White-Box, LLM Judge, and Ensemble Scorers
von: Bouchard, Dylan, et al.
Veröffentlicht: (2025)
von: Bouchard, Dylan, et al.
Veröffentlicht: (2025)
Krutrim LLM: Multilingual Foundational Model for over a Billion People
von: Kallappa, Aditya, et al.
Veröffentlicht: (2025)
von: Kallappa, Aditya, et al.
Veröffentlicht: (2025)
neuralFOMO: Can LLMs Handle Being Second Best? Measuring Envy-Like Preferences in Multi-Agent Settings
von: Ramamoorthy, Arnav, et al.
Veröffentlicht: (2025)
von: Ramamoorthy, Arnav, et al.
Veröffentlicht: (2025)
Small LLMs Are Weak Tool Learners: A Multi-LLM Agent
von: Shen, Weizhou, et al.
Veröffentlicht: (2024)
von: Shen, Weizhou, et al.
Veröffentlicht: (2024)
ALAS: Autonomous Learning Agent for Self-Updating Language Models
von: Atreja, Dhruv
Veröffentlicht: (2025)
von: Atreja, Dhruv
Veröffentlicht: (2025)
AgentSHAP: Interpreting LLM Agent Tool Importance with Monte Carlo Shapley Value Estimation
von: Horovicz, Miriam
Veröffentlicht: (2025)
von: Horovicz, Miriam
Veröffentlicht: (2025)
A Survey of Calibration Process for Black-Box LLMs
von: Xie, Liangru, et al.
Veröffentlicht: (2024)
von: Xie, Liangru, et al.
Veröffentlicht: (2024)
Black-Box On-Policy Distillation of Large Language Models
von: Ye, Tianzhu, et al.
Veröffentlicht: (2025)
von: Ye, Tianzhu, et al.
Veröffentlicht: (2025)
Characterizing the Robustness of Black-Box LLM Planners Under Perturbed Observations with Adaptive Stress Testing
von: Chakraborty, Neeloy, et al.
Veröffentlicht: (2025)
von: Chakraborty, Neeloy, et al.
Veröffentlicht: (2025)
Linguistic Complexity and Socio-cultural Patterns in Hip-Hop Lyrics
von: Bansal, Aayam, et al.
Veröffentlicht: (2025)
von: Bansal, Aayam, et al.
Veröffentlicht: (2025)
Auditing Black-Box LLM APIs with a Rank-Based Uniformity Test
von: Zhu, Xiaoyuan, et al.
Veröffentlicht: (2025)
von: Zhu, Xiaoyuan, et al.
Veröffentlicht: (2025)
TrailBlazer: History-Guided Reinforcement Learning for Black-Box LLM Jailbreaking
von: Yoon, Sung-Hoon, et al.
Veröffentlicht: (2026)
von: Yoon, Sung-Hoon, et al.
Veröffentlicht: (2026)
Language Generation in the Limit
von: Kleinberg, Jon, et al.
Veröffentlicht: (2024)
von: Kleinberg, Jon, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
SoK: Measuring What Matters for Closed-Loop Security Agents
von: Khurana, Mudita, et al.
Veröffentlicht: (2025) -
From Sycophancy to Sensemaking: Premise Governance for Human-AI Decision Making
von: Jain, Raunak
Veröffentlicht: (2026) -
Collaborative Causal Sensemaking: Closing the Complementarity Gap in Human-AI Decision Support
von: Jain, Raunak
Veröffentlicht: (2025) -
VidModEx: Interpretable and Efficient Black Box Model Extraction for High-Dimensional Spaces
von: Kumar, Somnath Sendhil, et al.
Veröffentlicht: (2024) -
Diagnosing LLM Judge Reliability: Conformal Prediction Sets and Transitivity Violations
von: Gupta, Manan, et al.
Veröffentlicht: (2026)