Adaptively profiling models with task elicitation
Fuente:
arXiv
Guardado en:
| Autores principales: | Brown, Davis, Balehannina, Prithvi, Jin, Helen, Havaldar, Shreya, Hassani, Hamed, Wong, Eric |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Evaluating the Performance of Large Language Models via Debates
por: Moniri, Behrad, et al.
Publicado: (2024)
por: Moniri, Behrad, et al.
Publicado: (2024)
Detecting Safety Violations Across Many Agent Traces
por: Stein, Adam, et al.
Publicado: (2026)
por: Stein, Adam, et al.
Publicado: (2026)
Probabilistic Soundness Guarantees in LLM Reasoning Chains
por: You, Weiqiu, et al.
Publicado: (2025)
por: You, Weiqiu, et al.
Publicado: (2025)
BrowserArena: Evaluating LLM Agents on Real-World Web Navigation Tasks
por: Anupam, Sagnik, et al.
Publicado: (2025)
por: Anupam, Sagnik, et al.
Publicado: (2025)
Uncertainty in Language Models: Assessment through Rank-Calibration
por: Huang, Xinmeng, et al.
Publicado: (2024)
por: Huang, Xinmeng, et al.
Publicado: (2024)
Language models show human-like content effects on reasoning tasks
por: Dasgupta, Ishita, et al.
Publicado: (2022)
por: Dasgupta, Ishita, et al.
Publicado: (2022)
One-Shot Safety Alignment for Large Language Models via Optimal Dualization
por: Huang, Xinmeng, et al.
Publicado: (2024)
por: Huang, Xinmeng, et al.
Publicado: (2024)
SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks
por: Robey, Alexander, et al.
Publicado: (2023)
por: Robey, Alexander, et al.
Publicado: (2023)
A Communication-Theoretic Framework for LLM Agents: Cost-Aware Adaptive Reliability
por: Omidvar, Hamed, et al.
Publicado: (2026)
por: Omidvar, Hamed, et al.
Publicado: (2026)
Beyond Next Word Prediction: Developing Comprehensive Evaluation Frameworks for measuring LLM performance on real world applications
por: Agrawal, Vishakha, et al.
Publicado: (2025)
por: Agrawal, Vishakha, et al.
Publicado: (2025)
Once Upon an Input: Reasoning via Per-Instance Program Synthesis
por: Stein, Adam, et al.
Publicado: (2025)
por: Stein, Adam, et al.
Publicado: (2025)
Instruction Following by Principled Boosting Attention of Large Language Models
por: Guardieiro, Vitoria, et al.
Publicado: (2025)
por: Guardieiro, Vitoria, et al.
Publicado: (2025)
Discourse vs emissions: Analysis of corporate narratives, symbolic practices, and mimicry through LLMs
por: Hassani, Bertrand Kian, et al.
Publicado: (2025)
por: Hassani, Bertrand Kian, et al.
Publicado: (2025)
Training-free LLM Merging for Multi-task Learning
por: Fu, Zichuan, et al.
Publicado: (2025)
por: Fu, Zichuan, et al.
Publicado: (2025)
Aviary: training language agents on challenging scientific tasks
por: Narayanan, Siddharth, et al.
Publicado: (2024)
por: Narayanan, Siddharth, et al.
Publicado: (2024)
Open Knowledge Base Canonicalization with Multi-task Learning
por: Liu, Bingchen, et al.
Publicado: (2024)
por: Liu, Bingchen, et al.
Publicado: (2024)
Calibrating Language Models with Adaptive Temperature Scaling
por: Xie, Johnathan, et al.
Publicado: (2024)
por: Xie, Johnathan, et al.
Publicado: (2024)
On-device System of Compositional Multi-tasking in Large Language Models
por: Bohdal, Ondrej, et al.
Publicado: (2025)
por: Bohdal, Ondrej, et al.
Publicado: (2025)
Efficient Compositional Multi-tasking for On-device Large Language Models
por: Bohdal, Ondrej, et al.
Publicado: (2025)
por: Bohdal, Ondrej, et al.
Publicado: (2025)
VoiceAgentBench: Are Voice Assistants ready for agentic tasks?
por: Jain, Dhruv, et al.
Publicado: (2025)
por: Jain, Dhruv, et al.
Publicado: (2025)
MAP's not dead yet: Uncovering true language model modes by conditioning away degeneracy
por: Yoshida, Davis, et al.
Publicado: (2023)
por: Yoshida, Davis, et al.
Publicado: (2023)
The FIX Benchmark: Extracting Features Interpretable to eXperts
por: Jin, Helen, et al.
Publicado: (2024)
por: Jin, Helen, et al.
Publicado: (2024)
XLGoBench: Detecting cross-lingual skill gaps with algorithmic tasks
por: Jain, Purvam, et al.
Publicado: (2026)
por: Jain, Purvam, et al.
Publicado: (2026)
CoT-ICL Lab: A Synthetic Framework for Studying Chain-of-Thought Learning from In-Context Demonstrations
por: Kothapalli, Vignesh, et al.
Publicado: (2025)
por: Kothapalli, Vignesh, et al.
Publicado: (2025)
Sample-aware Adaptive Structured Pruning for Large Language Models
por: Kong, Jun, et al.
Publicado: (2025)
por: Kong, Jun, et al.
Publicado: (2025)
CALM : A Multi-task Benchmark for Comprehensive Assessment of Language Model Bias
por: Gupta, Vipul, et al.
Publicado: (2023)
por: Gupta, Vipul, et al.
Publicado: (2023)
Adaptive Selection of LoRA Components in Privacy-Preserving Federated Learning
por: Kim, Myoungjun, et al.
Publicado: (2026)
por: Kim, Myoungjun, et al.
Publicado: (2026)
Enhancing Chemical Reaction and Retrosynthesis Prediction with Large Language Model and Dual-task Learning
por: Lin, Xuan, et al.
Publicado: (2025)
por: Lin, Xuan, et al.
Publicado: (2025)
QuestBench: Can LLMs ask the right question to acquire information in reasoning tasks?
por: Li, Belinda Z., et al.
Publicado: (2025)
por: Li, Belinda Z., et al.
Publicado: (2025)
Flaw or Artifact? Rethinking Prompt Sensitivity in Evaluating LLMs
por: Hua, Andong, et al.
Publicado: (2025)
por: Hua, Andong, et al.
Publicado: (2025)
MetaTool: Facilitating Large Language Models to Master Tools with Meta-task Augmentation
por: Wang, Xiaohan, et al.
Publicado: (2024)
por: Wang, Xiaohan, et al.
Publicado: (2024)
Multi-Target Cross-Lingual Summarization: a novel task and a language-neutral approach
por: Pernes, Diogo, et al.
Publicado: (2024)
por: Pernes, Diogo, et al.
Publicado: (2024)
To Think or Not to Think: The Hidden Cost of Meta-Training with Excessive CoT Examples
por: Kothapalli, Vignesh, et al.
Publicado: (2025)
por: Kothapalli, Vignesh, et al.
Publicado: (2025)
Sysformer: Safeguarding Frozen Large Language Models with Adaptive System Prompts
por: Sharma, Kartik, et al.
Publicado: (2025)
por: Sharma, Kartik, et al.
Publicado: (2025)
REALISTA: Realistic Latent Adversarial Attacks that Elicit LLM Hallucinations
por: Liang, Buyun, et al.
Publicado: (2026)
por: Liang, Buyun, et al.
Publicado: (2026)
Spectral Logit Sculpting: Adaptive Low-Rank Logit Transformation for Controlled Text Generation
por: Li, Jin, et al.
Publicado: (2025)
por: Li, Jin, et al.
Publicado: (2025)
Assessing the Impact of Prompting Methods on ChatGPT's Mathematical Capabilities
por: Chen, Yuhao, et al.
Publicado: (2023)
por: Chen, Yuhao, et al.
Publicado: (2023)
Granite-Function Calling Model: Introducing Function Calling Abilities via Multi-task Learning of Granular Tasks
por: Abdelaziz, Ibrahim, et al.
Publicado: (2024)
por: Abdelaziz, Ibrahim, et al.
Publicado: (2024)
Generation and De-Identification of Indian Clinical Discharge Summaries using LLMs
por: Singh, Sanjeet, et al.
Publicado: (2024)
por: Singh, Sanjeet, et al.
Publicado: (2024)
A Framework to Implement 1+N Multi-task Fine-tuning Pattern in LLMs Using the CGC-LORA Algorithm
por: Song, Chao, et al.
Publicado: (2024)
por: Song, Chao, et al.
Publicado: (2024)
Ejemplares similares
-
Evaluating the Performance of Large Language Models via Debates
por: Moniri, Behrad, et al.
Publicado: (2024) -
Detecting Safety Violations Across Many Agent Traces
por: Stein, Adam, et al.
Publicado: (2026) -
Probabilistic Soundness Guarantees in LLM Reasoning Chains
por: You, Weiqiu, et al.
Publicado: (2025) -
BrowserArena: Evaluating LLM Agents on Real-World Web Navigation Tasks
por: Anupam, Sagnik, et al.
Publicado: (2025) -
Uncertainty in Language Models: Assessment through Rank-Calibration
por: Huang, Xinmeng, et al.
Publicado: (2024)