Gespeichert in:
| Hauptverfasser: | Kadasi, Pritam, Upperwal, Abhishek, Singh, Mayank |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2602.03103 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ADAPT: Learning Task Mixtures for Budget-Constrained Instruction Tuning
von: Kadasi, Pritam, et al.
Veröffentlicht: (2025)
von: Kadasi, Pritam, et al.
Veröffentlicht: (2025)
When LLMs Stop Following Steps: A Diagnostic Study of Procedural Execution in Language Models
von: Panda, Sailesh, et al.
Veröffentlicht: (2026)
von: Panda, Sailesh, et al.
Veröffentlicht: (2026)
One Instruction Does Not Fit All: How Well Do Embeddings Align Personas and Instructions in Low-Resource Indian Languages?
von: Shah, Arya, et al.
Veröffentlicht: (2026)
von: Shah, Arya, et al.
Veröffentlicht: (2026)
Eka-Eval: An Evaluation Framework for Low-Resource Multilingual Large Language Models
von: Sinha, Samridhi Raj, et al.
Veröffentlicht: (2025)
von: Sinha, Samridhi Raj, et al.
Veröffentlicht: (2025)
How Many Parameters Does Your Task Really Need? Task Specific Pruning with LLM-Sieve
von: Reda, Waleed, et al.
Veröffentlicht: (2025)
von: Reda, Waleed, et al.
Veröffentlicht: (2025)
Short-Context Dominance: How Much Local Context Natural Language Actually Needs?
von: Vakilian, Vala, et al.
Veröffentlicht: (2025)
von: Vakilian, Vala, et al.
Veröffentlicht: (2025)
How Much Can RAG Help the Reasoning of LLM?
von: Liu, Jingyu, et al.
Veröffentlicht: (2024)
von: Liu, Jingyu, et al.
Veröffentlicht: (2024)
How Much Heavy Lifting Can an Agent Harness Do?: Measuring the LLM's Residual Role in a Planning Agent
von: Jung, Sungwoo, et al.
Veröffentlicht: (2026)
von: Jung, Sungwoo, et al.
Veröffentlicht: (2026)
Evaluating the Evaluator: Measuring LLMs' Adherence to Task Evaluation Instructions
von: Murugadoss, Bhuvanashree, et al.
Veröffentlicht: (2024)
von: Murugadoss, Bhuvanashree, et al.
Veröffentlicht: (2024)
SELF-GUIDE: Better Task-Specific Instruction Following via Self-Synthetic Finetuning
von: Zhao, Chenyang, et al.
Veröffentlicht: (2024)
von: Zhao, Chenyang, et al.
Veröffentlicht: (2024)
How Much is Too Much? Exploring LoRA Rank Trade-offs for Retaining Knowledge and Domain Robustness
von: Rathore, Darshita, et al.
Veröffentlicht: (2025)
von: Rathore, Darshita, et al.
Veröffentlicht: (2025)
DETAIL Matters: Measuring the Impact of Prompt Specificity on Reasoning in Large Language Models
von: Kim, Olivia
Veröffentlicht: (2025)
von: Kim, Olivia
Veröffentlicht: (2025)
Do Large Language Models Know How Much They Know?
von: Prato, Gabriele, et al.
Veröffentlicht: (2025)
von: Prato, Gabriele, et al.
Veröffentlicht: (2025)
How Much Can We Forget about Data Contamination?
von: Bordt, Sebastian, et al.
Veröffentlicht: (2024)
von: Bordt, Sebastian, et al.
Veröffentlicht: (2024)
COMI-LINGUA: Expert Annotated Large-Scale Dataset for Multitask NLP in Hindi-English Code-Mixing
von: Sheth, Rajvee, et al.
Veröffentlicht: (2025)
von: Sheth, Rajvee, et al.
Veröffentlicht: (2025)
PythonSaga: Redefining the Benchmark to Evaluate Code Generating LLMs
von: Yadav, Ankit, et al.
Veröffentlicht: (2024)
von: Yadav, Ankit, et al.
Veröffentlicht: (2024)
Where Does Toxicity Live? Mechanistic Localization and Targeted Suppression in Language Models
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2026)
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2026)
Quality-Conditioned Agreement in Automated Short Answer Scoring: Mid-Range Degradation and the Impact of Task-Specific Adaptation
von: Schleifer, Abigail Victoria Gurin, et al.
Veröffentlicht: (2026)
von: Schleifer, Abigail Victoria Gurin, et al.
Veröffentlicht: (2026)
How Much Do LLMs Hallucinate across Languages? On Realistic Multilingual Estimation of LLM Hallucination
von: Islam, Saad Obaid ul, et al.
Veröffentlicht: (2025)
von: Islam, Saad Obaid ul, et al.
Veröffentlicht: (2025)
Cross-lingual Editing in Multilingual Language Models
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2024)
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2024)
How Good Are LLMs for Literary Translation, Really? Literary Translation Evaluation with Humans and LLMs
von: Zhang, Ran, et al.
Veröffentlicht: (2024)
von: Zhang, Ran, et al.
Veröffentlicht: (2024)
Where It Really Matters: Few-Shot Environmental Conservation Media Monitoring for Low-Resource Languages
von: Jain, Sameer, et al.
Veröffentlicht: (2024)
von: Jain, Sameer, et al.
Veröffentlicht: (2024)
Measuring What Matters -- or What's Convenient?: Robustness of LLM-Based Scoring Systems to Construct-Irrelevant Factors
von: Walsh, Cole, et al.
Veröffentlicht: (2026)
von: Walsh, Cole, et al.
Veröffentlicht: (2026)
Every Answer Matters: Evaluating Commonsense with Probabilistic Measures
von: Cheng, Qi, et al.
Veröffentlicht: (2024)
von: Cheng, Qi, et al.
Veröffentlicht: (2024)
Instruction Embedding: Latent Representations of Instructions Towards Task Identification
von: Li, Yiwei, et al.
Veröffentlicht: (2024)
von: Li, Yiwei, et al.
Veröffentlicht: (2024)
How Much Would a Clinician Edit This Draft? Evaluating LLM Alignment for Patient Message Response Drafting
von: Seegmiller, Parker, et al.
Veröffentlicht: (2026)
von: Seegmiller, Parker, et al.
Veröffentlicht: (2026)
From No to Know: Taxonomy, Challenges, and Opportunities for Negation Understanding in Multimodal Foundation Models
von: Vatsa, Mayank, et al.
Veröffentlicht: (2025)
von: Vatsa, Mayank, et al.
Veröffentlicht: (2025)
Char-mander Use mBackdoor! A Study of Cross-lingual Backdoor Attacks in Multilingual LLMs
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2025)
von: Beniwal, Himanshu, et al.
Veröffentlicht: (2025)
No Universal Prompt: Unifying Reasoning through Adaptive Prompting for Temporal Table Reasoning
von: Rajgaria, Abhishek, et al.
Veröffentlicht: (2025)
von: Rajgaria, Abhishek, et al.
Veröffentlicht: (2025)
Fragile Thoughts: How Large Language Models Handle Chain-of-Thought Perturbations
von: Aravindan, Ashwath Vaithinathan, et al.
Veröffentlicht: (2026)
von: Aravindan, Ashwath Vaithinathan, et al.
Veröffentlicht: (2026)
ROSE: A Reward-Oriented Data Selection Framework for LLM Task-Specific Instruction Tuning
von: Wu, Yang, et al.
Veröffentlicht: (2024)
von: Wu, Yang, et al.
Veröffentlicht: (2024)
How effective are VLMs in assisting humans in inferring the quality of mental models from Multimodal short answers?
von: Sil, Pritam, et al.
Veröffentlicht: (2026)
von: Sil, Pritam, et al.
Veröffentlicht: (2026)
What Really is Commonsense Knowledge?
von: Do, Quyet V., et al.
Veröffentlicht: (2024)
von: Do, Quyet V., et al.
Veröffentlicht: (2024)
Intent-conditioned and Non-toxic Counterspeech Generation using Multi-Task Instruction Tuning with RLAIF
von: Hengle, Amey, et al.
Veröffentlicht: (2024)
von: Hengle, Amey, et al.
Veröffentlicht: (2024)
BertaQA: How Much Do Language Models Know About Local Culture?
von: Etxaniz, Julen, et al.
Veröffentlicht: (2024)
von: Etxaniz, Julen, et al.
Veröffentlicht: (2024)
R-Horizon: How Far Can Your Large Reasoning Model Really Go in Breadth and Depth?
von: Lu, Yi, et al.
Veröffentlicht: (2025)
von: Lu, Yi, et al.
Veröffentlicht: (2025)
Error Taxonomy-Guided Prompt Optimization
von: Singh, Mayank, et al.
Veröffentlicht: (2026)
von: Singh, Mayank, et al.
Veröffentlicht: (2026)
Theory of Mind in Action: The Instruction Inference Task in Dynamic Human-Agent Collaboration
von: Saad, Fardin, et al.
Veröffentlicht: (2025)
von: Saad, Fardin, et al.
Veröffentlicht: (2025)
Outlier Dimensions Encode Task-Specific Knowledge
von: Rudman, William, et al.
Veröffentlicht: (2023)
von: Rudman, William, et al.
Veröffentlicht: (2023)
Do Audio LLMs Really LISTEN, or Just Transcribe? Measuring Lexical vs. Acoustic Emotion Cues Reliance
von: Chen, Jingyi, et al.
Veröffentlicht: (2025)
von: Chen, Jingyi, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
ADAPT: Learning Task Mixtures for Budget-Constrained Instruction Tuning
von: Kadasi, Pritam, et al.
Veröffentlicht: (2025) -
When LLMs Stop Following Steps: A Diagnostic Study of Procedural Execution in Language Models
von: Panda, Sailesh, et al.
Veröffentlicht: (2026) -
One Instruction Does Not Fit All: How Well Do Embeddings Align Personas and Instructions in Low-Resource Indian Languages?
von: Shah, Arya, et al.
Veröffentlicht: (2026) -
Eka-Eval: An Evaluation Framework for Low-Resource Multilingual Large Language Models
von: Sinha, Samridhi Raj, et al.
Veröffentlicht: (2025) -
How Many Parameters Does Your Task Really Need? Task Specific Pruning with LLM-Sieve
von: Reda, Waleed, et al.
Veröffentlicht: (2025)