WorldAPIs: The World Is Worth How Many APIs? A Thought Experiment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ou, Jiefu, Uzunoglu, Arda, Van Durme, Benjamin, Khashabi, Daniel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
arXiv2Table: Toward Realistic Benchmarking and Evaluation for LLM-Based Literature-Review Table Generation
von: Wang, Weiqi, et al.
Veröffentlicht: (2025)
von: Wang, Weiqi, et al.
Veröffentlicht: (2025)
Trust Functions: Near-Lossless Weak-to-Strong Generalization by Learning When to Trust the Weak Teacher
von: Uzunoglu, Arda, et al.
Veröffentlicht: (2026)
von: Uzunoglu, Arda, et al.
Veröffentlicht: (2026)
The Flaw of Averages: Quantifying Uniformity of Performance on Benchmarks
von: Uzunoglu, Arda, et al.
Veröffentlicht: (2025)
von: Uzunoglu, Arda, et al.
Veröffentlicht: (2025)
Many-Tier Instruction Hierarchy in LLM Agents
von: Zhang, Jingyu, et al.
Veröffentlicht: (2026)
von: Zhang, Jingyu, et al.
Veröffentlicht: (2026)
Crystal: Characterizing Relative Impact of Scholarly Publications
von: Collison, Hannah, et al.
Veröffentlicht: (2026)
von: Collison, Hannah, et al.
Veröffentlicht: (2026)
Are Finer Citations Always Better? Rethinking Granularity for Attributed Generation
von: Wang, Hexuan, et al.
Veröffentlicht: (2026)
von: Wang, Hexuan, et al.
Veröffentlicht: (2026)
Benchmarking Procedural Language Understanding for Low-Resource Languages: A Case Study on Turkish
von: Uzunoglu, Arda, et al.
Veröffentlicht: (2023)
von: Uzunoglu, Arda, et al.
Veröffentlicht: (2023)
Beyond Perfect APIs: A Comprehensive Evaluation of LLM Agents Under Real-World API Complexity
von: Kim, Doyoung, et al.
Veröffentlicht: (2026)
von: Kim, Doyoung, et al.
Veröffentlicht: (2026)
Compressed Chain of Thought: Efficient Reasoning Through Dense Representations
von: Cheng, Jeffrey, et al.
Veröffentlicht: (2024)
von: Cheng, Jeffrey, et al.
Veröffentlicht: (2024)
Verifiable by Design: Aligning Language Models to Quote from Pre-Training Data
von: Zhang, Jingyu, et al.
Veröffentlicht: (2024)
von: Zhang, Jingyu, et al.
Veröffentlicht: (2024)
Certified Mitigation of Worst-Case LLM Copyright Infringement
von: Zhang, Jingyu, et al.
Veröffentlicht: (2025)
von: Zhang, Jingyu, et al.
Veröffentlicht: (2025)
From Summary to Action: Enhancing Large Language Models for Complex Tasks with Open World APIs
von: Liu, Yulong, et al.
Veröffentlicht: (2024)
von: Liu, Yulong, et al.
Veröffentlicht: (2024)
Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirements
von: Zhang, Jingyu, et al.
Veröffentlicht: (2024)
von: Zhang, Jingyu, et al.
Veröffentlicht: (2024)
StableToolBench-MirrorAPI: Modeling Tool Environments as Mirrors of 7,000+ Real-World APIs
von: Guo, Zhicheng, et al.
Veröffentlicht: (2025)
von: Guo, Zhicheng, et al.
Veröffentlicht: (2025)
CLAIMCHECK: How Grounded are LLM Critiques of Scientific Papers?
von: Ou, Jiefu, et al.
Veröffentlicht: (2025)
von: Ou, Jiefu, et al.
Veröffentlicht: (2025)
RORA: Robust Free-Text Rationale Evaluation
von: Jiang, Zhengping, et al.
Veröffentlicht: (2024)
von: Jiang, Zhengping, et al.
Veröffentlicht: (2024)
Dated Data: Tracing Knowledge Cutoffs in Large Language Models
von: Cheng, Jeffrey, et al.
Veröffentlicht: (2024)
von: Cheng, Jeffrey, et al.
Veröffentlicht: (2024)
Auditing Prompt Caching in Language Model APIs
von: Gu, Chenchen, et al.
Veröffentlicht: (2025)
von: Gu, Chenchen, et al.
Veröffentlicht: (2025)
Attacks on Third-Party APIs of Large Language Models
von: Zhao, Wanru, et al.
Veröffentlicht: (2024)
von: Zhao, Wanru, et al.
Veröffentlicht: (2024)
Exploiting Novel GPT-4 APIs
von: Pelrine, Kellin, et al.
Veröffentlicht: (2023)
von: Pelrine, Kellin, et al.
Veröffentlicht: (2023)
An Investigation into Misuse of Java Security APIs by Large Language Models
von: Mousavi, Zahra, et al.
Veröffentlicht: (2024)
von: Mousavi, Zahra, et al.
Veröffentlicht: (2024)
PARADISE: Evaluating Implicit Planning Skills of Language Models with Procedural Warnings and Tips Dataset
von: Uzunoglu, Arda, et al.
Veröffentlicht: (2024)
von: Uzunoglu, Arda, et al.
Veröffentlicht: (2024)
Octopus: On-device language model for function calling of software APIs
von: Chen, Wei, et al.
Veröffentlicht: (2024)
von: Chen, Wei, et al.
Veröffentlicht: (2024)
Beyond Indistinguishability: Measuring Extraction Risk in LLM APIs
von: Liu, Ruixuan, et al.
Veröffentlicht: (2026)
von: Liu, Ruixuan, et al.
Veröffentlicht: (2026)
"According to ...": Prompting Language Models Improves Quoting from Pre-Training Data
von: Weller, Orion, et al.
Veröffentlicht: (2023)
von: Weller, Orion, et al.
Veröffentlicht: (2023)
Instructional Text Across Disciplines: A Survey of Representations, Downstream Tasks, and Open Challenges Toward Capable AI Agents
von: Safa, Abdulfattah, et al.
Veröffentlicht: (2024)
von: Safa, Abdulfattah, et al.
Veröffentlicht: (2024)
Differentially Private Synthetic Data via Foundation Model APIs 2: Text
von: Xie, Chulin, et al.
Veröffentlicht: (2024)
von: Xie, Chulin, et al.
Veröffentlicht: (2024)
SELF-[IN]CORRECT: LLMs Struggle with Discriminating Self-Generated Responses
von: Jiang, Dongwei, et al.
Veröffentlicht: (2024)
von: Jiang, Dongwei, et al.
Veröffentlicht: (2024)
Beyond Text-to-SQL: An Agentic LLM System for Governed Enterprise Analytics APIs
von: Singh, Gundeep, et al.
Veröffentlicht: (2026)
von: Singh, Gundeep, et al.
Veröffentlicht: (2026)
FineTuneBench: How well do commercial fine-tuning APIs infuse knowledge into LLMs?
von: Wu, Eric, et al.
Veröffentlicht: (2024)
von: Wu, Eric, et al.
Veröffentlicht: (2024)
Lost in Moderation: How Commercial Content Moderation APIs Over- and Under-Moderate Group-Targeted Hate Speech and Linguistic Variations
von: Hartmann, David, et al.
Veröffentlicht: (2025)
von: Hartmann, David, et al.
Veröffentlicht: (2025)
(Why) Is My Prompt Getting Worse? Rethinking Regression Testing for Evolving LLM APIs
von: Ma, Wanqin, et al.
Veröffentlicht: (2023)
von: Ma, Wanqin, et al.
Veröffentlicht: (2023)
Auditing Black-Box LLM APIs with a Rank-Based Uniformity Test
von: Zhu, Xiaoyuan, et al.
Veröffentlicht: (2025)
von: Zhu, Xiaoyuan, et al.
Veröffentlicht: (2025)
Are You Getting What You Pay For? Auditing Model Substitution in LLM APIs
von: Cai, Will, et al.
Veröffentlicht: (2025)
von: Cai, Will, et al.
Veröffentlicht: (2025)
Enabling Communication via APIs for Mainframe Applications
von: Kanvar, Vini, et al.
Veröffentlicht: (2024)
von: Kanvar, Vini, et al.
Veröffentlicht: (2024)
Firefly: Illuminating Large-Scale Verified Tool-Call Data Generation from Real APIs
von: Lu, Yuxuan, et al.
Veröffentlicht: (2026)
von: Lu, Yuxuan, et al.
Veröffentlicht: (2026)
RATIONALYST: Mining Implicit Rationales for Process Supervision of Reasoning
von: Jiang, Dongwei, et al.
Veröffentlicht: (2024)
von: Jiang, Dongwei, et al.
Veröffentlicht: (2024)
Core: Robust Factual Precision with Informative Sub-Claim Identification
von: Jiang, Zhengping, et al.
Veröffentlicht: (2024)
von: Jiang, Zhengping, et al.
Veröffentlicht: (2024)
Compactor: Calibrated Query-Agnostic KV Cache Compression with Approximate Leverage Scores
von: Chari, Vivek, et al.
Veröffentlicht: (2025)
von: Chari, Vivek, et al.
Veröffentlicht: (2025)
LLMs Provide Unstable Answers to Legal Questions
von: Blair-Stanek, Andrew, et al.
Veröffentlicht: (2025)
von: Blair-Stanek, Andrew, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
arXiv2Table: Toward Realistic Benchmarking and Evaluation for LLM-Based Literature-Review Table Generation
von: Wang, Weiqi, et al.
Veröffentlicht: (2025) -
Trust Functions: Near-Lossless Weak-to-Strong Generalization by Learning When to Trust the Weak Teacher
von: Uzunoglu, Arda, et al.
Veröffentlicht: (2026) -
The Flaw of Averages: Quantifying Uniformity of Performance on Benchmarks
von: Uzunoglu, Arda, et al.
Veröffentlicht: (2025) -
Many-Tier Instruction Hierarchy in LLM Agents
von: Zhang, Jingyu, et al.
Veröffentlicht: (2026) -
Crystal: Characterizing Relative Impact of Scholarly Publications
von: Collison, Hannah, et al.
Veröffentlicht: (2026)