Peering Inside the Black Box: Uncovering LLM Errors in Optimization Modelling through Component-Level Evaluation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Refai, Dania, Ahmed, Moataz |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HYDRA: Model Factorization Framework for Black-Box LLM Personalization
von: Zhuang, Yuchen, et al.
Veröffentlicht: (2024)
von: Zhuang, Yuchen, et al.
Veröffentlicht: (2024)
Topic Modelling Black Box Optimization
von: Akramov, Roman, et al.
Veröffentlicht: (2025)
von: Akramov, Roman, et al.
Veröffentlicht: (2025)
Uncertainty Quantification for Language Models: A Suite of Black-Box, White-Box, LLM Judge, and Ensemble Scorers
von: Bouchard, Dylan, et al.
Veröffentlicht: (2025)
von: Bouchard, Dylan, et al.
Veröffentlicht: (2025)
FactSelfCheck: Fact-Level Black-Box Hallucination Detection for LLMs
von: Sawczyn, Albert, et al.
Veröffentlicht: (2025)
von: Sawczyn, Albert, et al.
Veröffentlicht: (2025)
In-Context Explainers: Harnessing LLMs for Explaining Black Box Models
von: Kroeger, Nicholas, et al.
Veröffentlicht: (2023)
von: Kroeger, Nicholas, et al.
Veröffentlicht: (2023)
From Human Annotation to Automation: LLM-in-the-Loop Active Learning for Arabic Sentiment Analysis
von: Refai, Dania, et al.
Veröffentlicht: (2025)
von: Refai, Dania, et al.
Veröffentlicht: (2025)
BBox-Adapter: Lightweight Adapting for Black-Box Large Language Models
von: Sun, Haotian, et al.
Veröffentlicht: (2024)
von: Sun, Haotian, et al.
Veröffentlicht: (2024)
Large Language Model Confidence Estimation via Black-Box Access
von: Pedapati, Tejaswini, et al.
Veröffentlicht: (2024)
von: Pedapati, Tejaswini, et al.
Veröffentlicht: (2024)
Mafin: Enhancing Black-Box Embeddings with Model Augmented Fine-Tuning
von: Zhang, Mingtian, et al.
Veröffentlicht: (2024)
von: Zhang, Mingtian, et al.
Veröffentlicht: (2024)
Trial and Error: Exploration-Based Trajectory Optimization for LLM Agents
von: Song, Yifan, et al.
Veröffentlicht: (2024)
von: Song, Yifan, et al.
Veröffentlicht: (2024)
How to Train Your Advisor: Steering Black-Box LLMs with Advisor Models
von: Asawa, Parth, et al.
Veröffentlicht: (2025)
von: Asawa, Parth, et al.
Veröffentlicht: (2025)
Training Deliberative Monitors for Black-Box Scheming Detection
von: Sinha, Aditya, et al.
Veröffentlicht: (2026)
von: Sinha, Aditya, et al.
Veröffentlicht: (2026)
Evaluating the Sensitivity of BiLSTM Forecasting Models to Sequence Length and Input Noise
von: Albelali, Salma, et al.
Veröffentlicht: (2025)
von: Albelali, Salma, et al.
Veröffentlicht: (2025)
RKadiyala at SemEval-2024 Task 8: Black-Box Word-Level Text Boundary Detection in Partially Machine Generated Texts
von: Kadiyala, Ram Mohan Rao
Veröffentlicht: (2024)
von: Kadiyala, Ram Mohan Rao
Veröffentlicht: (2024)
Matryoshka Pilot: Learning to Drive Black-Box LLMs with LLMs
von: Li, Changhao, et al.
Veröffentlicht: (2024)
von: Li, Changhao, et al.
Veröffentlicht: (2024)
Are Large Language Models Reliable AI Scientists? Assessing Reverse-Engineering of Black-Box Systems
von: Geng, Jiayi, et al.
Veröffentlicht: (2025)
von: Geng, Jiayi, et al.
Veröffentlicht: (2025)
Efficient Non-Parametric Uncertainty Quantification for Black-Box Large Language Models and Decision Planning
von: Tsai, Yao-Hung Hubert, et al.
Veröffentlicht: (2024)
von: Tsai, Yao-Hung Hubert, et al.
Veröffentlicht: (2024)
Think Inside the JSON: Reinforcement Strategy for Strict LLM Schema Adherence
von: Agarwal, Bhavik, et al.
Veröffentlicht: (2025)
von: Agarwal, Bhavik, et al.
Veröffentlicht: (2025)
Uncovering Customer Issues through Topological Natural Language Analysis
von: Pi, Shu-Ting, et al.
Veröffentlicht: (2024)
von: Pi, Shu-Ting, et al.
Veröffentlicht: (2024)
Collab-RAG: Boosting Retrieval-Augmented Generation for Complex Question Answering via White-Box and Black-Box LLM Collaboration
von: Xu, Ran, et al.
Veröffentlicht: (2025)
von: Xu, Ran, et al.
Veröffentlicht: (2025)
PiCO: Peer Review in LLMs based on the Consistency Optimization
von: Ning, Kun-Peng, et al.
Veröffentlicht: (2024)
von: Ning, Kun-Peng, et al.
Veröffentlicht: (2024)
Bias Similarity Measurement: A Black-Box Audit of Fairness Across LLMs
von: Jeong, Hyejun, et al.
Veröffentlicht: (2024)
von: Jeong, Hyejun, et al.
Veröffentlicht: (2024)
TICKing All the Boxes: Generated Checklists Improve LLM Evaluation and Generation
von: Cook, Jonathan, et al.
Veröffentlicht: (2024)
von: Cook, Jonathan, et al.
Veröffentlicht: (2024)
Evaluating Cooperation in LLM Social Groups through Elected Leadership
von: Faulkner, Ryan, et al.
Veröffentlicht: (2026)
von: Faulkner, Ryan, et al.
Veröffentlicht: (2026)
From Signal Degradation to Computation Collapse: Uncovering the Two Failure Modes of LLM Quantization
von: Zhou, Chenxi, et al.
Veröffentlicht: (2026)
von: Zhou, Chenxi, et al.
Veröffentlicht: (2026)
Automated Optimization Modeling via a Localizable Error-Driven Perspective
von: Liu, Weiting, et al.
Veröffentlicht: (2026)
von: Liu, Weiting, et al.
Veröffentlicht: (2026)
Error Taxonomy-Guided Prompt Optimization
von: Singh, Mayank, et al.
Veröffentlicht: (2026)
von: Singh, Mayank, et al.
Veröffentlicht: (2026)
PAL: Proxy-Guided Black-Box Attack on Large Language Models
von: Sitawarin, Chawin, et al.
Veröffentlicht: (2024)
von: Sitawarin, Chawin, et al.
Veröffentlicht: (2024)
GTBench: Uncovering the Strategic Reasoning Limitations of LLMs via Game-Theoretic Evaluations
von: Duan, Jinhao, et al.
Veröffentlicht: (2024)
von: Duan, Jinhao, et al.
Veröffentlicht: (2024)
Kov: Transferable and Naturalistic Black-Box LLM Attacks using Markov Decision Processes and Tree Search
von: Moss, Robert J.
Veröffentlicht: (2024)
von: Moss, Robert J.
Veröffentlicht: (2024)
Black-Box Reliability Certification for AI Agents via Self-Consistency Sampling and Conformal Calibration
von: Mouzouni, Charafeddine
Veröffentlicht: (2026)
von: Mouzouni, Charafeddine
Veröffentlicht: (2026)
Matrix: Peer-to-Peer Multi-Agent Synthetic Data Generation Framework
von: Wang, Dong, et al.
Veröffentlicht: (2025)
von: Wang, Dong, et al.
Veröffentlicht: (2025)
Layer by Layer: Uncovering Hidden Representations in Language Models
von: Skean, Oscar, et al.
Veröffentlicht: (2025)
von: Skean, Oscar, et al.
Veröffentlicht: (2025)
Uncovering Competency Gaps in Large Language Models and Their Benchmarks
von: Bohacek, Maty, et al.
Veröffentlicht: (2025)
von: Bohacek, Maty, et al.
Veröffentlicht: (2025)
Hidden Leaks in Time Series Forecasting: How Data Leakage Affects LSTM Evaluation Across Configurations and Validation Strategies
von: Albelali, Salma, et al.
Veröffentlicht: (2025)
von: Albelali, Salma, et al.
Veröffentlicht: (2025)
Temporal Consistency for LLM Reasoning Process Error Identification
von: Guo, Jiacheng, et al.
Veröffentlicht: (2025)
von: Guo, Jiacheng, et al.
Veröffentlicht: (2025)
LLM Assertiveness can be Mechanistically Decomposed into Emotional and Logical Components
von: Tsujimura, Hikaru, et al.
Veröffentlicht: (2025)
von: Tsujimura, Hikaru, et al.
Veröffentlicht: (2025)
Policies Permitting LLM Use for Polishing Peer Reviews Are Currently Not Enforceable
von: Saha, Rounak, et al.
Veröffentlicht: (2026)
von: Saha, Rounak, et al.
Veröffentlicht: (2026)
Contextual Drag: How Errors in the Context Affect LLM Reasoning
von: Cheng, Yun, et al.
Veröffentlicht: (2026)
von: Cheng, Yun, et al.
Veröffentlicht: (2026)
Evaluating Mathematical Reasoning of Large Language Models: A Focus on Error Identification and Correction
von: Li, Xiaoyuan, et al.
Veröffentlicht: (2024)
von: Li, Xiaoyuan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
HYDRA: Model Factorization Framework for Black-Box LLM Personalization
von: Zhuang, Yuchen, et al.
Veröffentlicht: (2024) -
Topic Modelling Black Box Optimization
von: Akramov, Roman, et al.
Veröffentlicht: (2025) -
Uncertainty Quantification for Language Models: A Suite of Black-Box, White-Box, LLM Judge, and Ensemble Scorers
von: Bouchard, Dylan, et al.
Veröffentlicht: (2025) -
FactSelfCheck: Fact-Level Black-Box Hallucination Detection for LLMs
von: Sawczyn, Albert, et al.
Veröffentlicht: (2025) -
In-Context Explainers: Harnessing LLMs for Explaining Black Box Models
von: Kroeger, Nicholas, et al.
Veröffentlicht: (2023)