Separating Constraint Compliance from Semantic Accuracy: A Novel Benchmark for Evaluating Instruction-Following Under Compression
Fuente:
arXiv
Salvato in:
| Autore principale: | Baxi, Rahul |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
The Drill-Down and Fabricate Test (DDFT): A Protocol for Measuring Epistemic Robustness in Language Models
di: Baxi, Rahul
Pubblicazione: (2025)
di: Baxi, Rahul
Pubblicazione: (2025)
Attentive Reasoning Queries: A Systematic Method for Optimizing Instruction-Following in Large Language Models
di: Karov, Bar, et al.
Pubblicazione: (2025)
di: Karov, Bar, et al.
Pubblicazione: (2025)
Towards Safer Chatbots: Automated Policy Compliance Evaluation of Custom GPTs
di: Rodriguez, David, et al.
Pubblicazione: (2025)
di: Rodriguez, David, et al.
Pubblicazione: (2025)
InstructCMP: Length Control in Sentence Compression through Instruction-based Large Language Models
di: Juseon-Do, et al.
Pubblicazione: (2024)
di: Juseon-Do, et al.
Pubblicazione: (2024)
LLMs and Memorization: On Quality and Specificity of Copyright Compliance
di: Mueller, Felix B, et al.
Pubblicazione: (2024)
di: Mueller, Felix B, et al.
Pubblicazione: (2024)
Is this Idea Novel? An Automated Benchmark for Judgment of Research Ideas
di: Schopf, Tim, et al.
Pubblicazione: (2026)
di: Schopf, Tim, et al.
Pubblicazione: (2026)
No Dataset Needed for Downstream Knowledge Benchmarking: Response Dispersion Inversely Correlates with Accuracy on Domain-specific QA
di: Simione II, Robert L
Pubblicazione: (2024)
di: Simione II, Robert L
Pubblicazione: (2024)
Instruction Complexity Induces Positional Collapse in Adversarial LLM Evaluation
di: Cacioli, Jon-Paul
Pubblicazione: (2026)
di: Cacioli, Jon-Paul
Pubblicazione: (2026)
PARAPHRASUS : A Comprehensive Benchmark for Evaluating Paraphrase Detection Models
di: Michail, Andrianos, et al.
Pubblicazione: (2024)
di: Michail, Andrianos, et al.
Pubblicazione: (2024)
Integrating Emotional and Linguistic Models for Ethical Compliance in Large Language Models
di: Chang, Edward Y.
Pubblicazione: (2024)
di: Chang, Edward Y.
Pubblicazione: (2024)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
di: Peters, Sydney, et al.
Pubblicazione: (2025)
di: Peters, Sydney, et al.
Pubblicazione: (2025)
PrivacyBench: A Conversational Benchmark for Evaluating Privacy in Personalized AI
di: Mukhopadhyay, Srija, et al.
Pubblicazione: (2025)
di: Mukhopadhyay, Srija, et al.
Pubblicazione: (2025)
UA-Legal-Bench: A Benchmark for Evaluating Large Language Models on Ukrainian Legal Reasoning
di: Ovcharov, Volodymyr
Pubblicazione: (2026)
di: Ovcharov, Volodymyr
Pubblicazione: (2026)
PSST: A Benchmark for Evaluation-driven Text Public-Speaking Style Transfer
di: Sun, Huashan, et al.
Pubblicazione: (2023)
di: Sun, Huashan, et al.
Pubblicazione: (2023)
LeCoDe: A Benchmark Dataset for Interactive Legal Consultation Dialogue Evaluation
di: Yuan, Weikang, et al.
Pubblicazione: (2025)
di: Yuan, Weikang, et al.
Pubblicazione: (2025)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
di: Saji, Alan, et al.
Pubblicazione: (2025)
di: Saji, Alan, et al.
Pubblicazione: (2025)
EduAdapt: A Question Answer Benchmark Dataset for Evaluating Grade-Level Adaptability in LLMs
di: Naeem, Numaan, et al.
Pubblicazione: (2025)
di: Naeem, Numaan, et al.
Pubblicazione: (2025)
Cetvel: A Unified Benchmark for Evaluating Language Understanding, Generation and Cultural Capacity of LLMs for Turkish
di: Er, Yakup Abrek, et al.
Pubblicazione: (2025)
di: Er, Yakup Abrek, et al.
Pubblicazione: (2025)
From Reading to Compressing: Exploring the Multi-document Reader for Prompt Compression
di: Choi, Eunseong, et al.
Pubblicazione: (2024)
di: Choi, Eunseong, et al.
Pubblicazione: (2024)
Instructional Agents: Reducing Teaching Faculty Workload through Multi-Agent Instructional Design
di: Yao, Huaiyuan, et al.
Pubblicazione: (2025)
di: Yao, Huaiyuan, et al.
Pubblicazione: (2025)
FIN-bench-v2: A Unified and Robust Benchmark Suite for Evaluating Finnish Large Language Models
di: Kytöniemi, Joona, et al.
Pubblicazione: (2025)
di: Kytöniemi, Joona, et al.
Pubblicazione: (2025)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
di: Oketunji, Abiodun Finbarrs
Pubblicazione: (2023)
di: Oketunji, Abiodun Finbarrs
Pubblicazione: (2023)
Controllable Text Summarization: Unraveling Challenges, Approaches, and Prospects -- A Survey
di: Urlana, Ashok, et al.
Pubblicazione: (2023)
di: Urlana, Ashok, et al.
Pubblicazione: (2023)
Exploring News Summarization and Enrichment in a Highly Resource-Scarce Indian Language: A Case Study of Mizo
di: Bala, Abhinaba, et al.
Pubblicazione: (2024)
di: Bala, Abhinaba, et al.
Pubblicazione: (2024)
Streamlining Redundant Layers to Compress Large Language Models
di: Chen, Xiaodong, et al.
Pubblicazione: (2024)
di: Chen, Xiaodong, et al.
Pubblicazione: (2024)
How Instruction-Tuning Imparts Length Control: A Cross-Lingual Mechanistic Analysis
di: Rocchetti, Elisabetta, et al.
Pubblicazione: (2025)
di: Rocchetti, Elisabetta, et al.
Pubblicazione: (2025)
mEdIT: Multilingual Text Editing via Instruction Tuning
di: Raheja, Vipul, et al.
Pubblicazione: (2024)
di: Raheja, Vipul, et al.
Pubblicazione: (2024)
Demystifying Instruction Mixing for Fine-tuning Large Language Models
di: Wang, Renxi, et al.
Pubblicazione: (2023)
di: Wang, Renxi, et al.
Pubblicazione: (2023)
From Benchmarking to Reasoning: A Dual-Aspect, Large-Scale Evaluation of LLMs on Vietnamese Legal Text
di: Le, Van-Truong
Pubblicazione: (2026)
di: Le, Van-Truong
Pubblicazione: (2026)
Residual Drift Dominates Contradiction in Multi-Turn Constraint Reasoning
di: Kawada, Sebastien
Pubblicazione: (2026)
di: Kawada, Sebastien
Pubblicazione: (2026)
The Perplexity Paradox: Why Code Compresses Better Than Math in LLM Prompts
di: Johnson, Warren
Pubblicazione: (2026)
di: Johnson, Warren
Pubblicazione: (2026)
SUBLLM: A Novel Efficient Architecture with Token Sequence Subsampling for LLM
di: Wang, Quandong, et al.
Pubblicazione: (2024)
di: Wang, Quandong, et al.
Pubblicazione: (2024)
Exploring RWKV for Sentence Embeddings: Layer-wise Analysis and Baseline Comparison for Semantic Similarity
di: Pan, Xinghan
Pubblicazione: (2025)
di: Pan, Xinghan
Pubblicazione: (2025)
Automated MCQA Benchmarking at Scale: Evaluating Reasoning Traces as Retrieval Sources for Domain Adaptation of Small Language Models
di: Gokdemir, Ozan, et al.
Pubblicazione: (2025)
di: Gokdemir, Ozan, et al.
Pubblicazione: (2025)
Cultural Benchmarking of LLMs in Standard and Dialectal Arabic Dialogues
di: Kautsar, Muhammad Dehan Al, et al.
Pubblicazione: (2026)
di: Kautsar, Muhammad Dehan Al, et al.
Pubblicazione: (2026)
Tool-Schema Compression Enables Agentic RAG Under Constrained Context Budgets
di: Sakizli, Furkan
Pubblicazione: (2026)
di: Sakizli, Furkan
Pubblicazione: (2026)
Historical Ink: Semantic Shift Detection for 19th Century Spanish
di: Montes, Tony, et al.
Pubblicazione: (2024)
di: Montes, Tony, et al.
Pubblicazione: (2024)
DeEscalWild: A Real-World Benchmark for Automated De-Escalation Training with SLMs
di: Hasan, Md Hasebul, et al.
Pubblicazione: (2026)
di: Hasan, Md Hasebul, et al.
Pubblicazione: (2026)
Evaluating the efficacy of LLM Safety Solutions : The Palit Benchmark Dataset
di: Palit, Sayon, et al.
Pubblicazione: (2025)
di: Palit, Sayon, et al.
Pubblicazione: (2025)
CausalGym: Benchmarking causal interpretability methods on linguistic tasks
di: Arora, Aryaman, et al.
Pubblicazione: (2024)
di: Arora, Aryaman, et al.
Pubblicazione: (2024)
Documenti analoghi
-
The Drill-Down and Fabricate Test (DDFT): A Protocol for Measuring Epistemic Robustness in Language Models
di: Baxi, Rahul
Pubblicazione: (2025) -
Attentive Reasoning Queries: A Systematic Method for Optimizing Instruction-Following in Large Language Models
di: Karov, Bar, et al.
Pubblicazione: (2025) -
Towards Safer Chatbots: Automated Policy Compliance Evaluation of Custom GPTs
di: Rodriguez, David, et al.
Pubblicazione: (2025) -
InstructCMP: Length Control in Sentence Compression through Instruction-based Large Language Models
di: Juseon-Do, et al.
Pubblicazione: (2024) -
LLMs and Memorization: On Quality and Specificity of Copyright Compliance
di: Mueller, Felix B, et al.
Pubblicazione: (2024)