Separating Constraint Compliance from Semantic Accuracy: A Novel Benchmark for Evaluating Instruction-Following Under Compression
Fuente:
arXiv
Guardado en:
| Autor principal: | Baxi, Rahul |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
The Drill-Down and Fabricate Test (DDFT): A Protocol for Measuring Epistemic Robustness in Language Models
por: Baxi, Rahul
Publicado: (2025)
por: Baxi, Rahul
Publicado: (2025)
Attentive Reasoning Queries: A Systematic Method for Optimizing Instruction-Following in Large Language Models
por: Karov, Bar, et al.
Publicado: (2025)
por: Karov, Bar, et al.
Publicado: (2025)
Towards Safer Chatbots: Automated Policy Compliance Evaluation of Custom GPTs
por: Rodriguez, David, et al.
Publicado: (2025)
por: Rodriguez, David, et al.
Publicado: (2025)
InstructCMP: Length Control in Sentence Compression through Instruction-based Large Language Models
por: Juseon-Do, et al.
Publicado: (2024)
por: Juseon-Do, et al.
Publicado: (2024)
LLMs and Memorization: On Quality and Specificity of Copyright Compliance
por: Mueller, Felix B, et al.
Publicado: (2024)
por: Mueller, Felix B, et al.
Publicado: (2024)
Is this Idea Novel? An Automated Benchmark for Judgment of Research Ideas
por: Schopf, Tim, et al.
Publicado: (2026)
por: Schopf, Tim, et al.
Publicado: (2026)
No Dataset Needed for Downstream Knowledge Benchmarking: Response Dispersion Inversely Correlates with Accuracy on Domain-specific QA
por: Simione II, Robert L
Publicado: (2024)
por: Simione II, Robert L
Publicado: (2024)
Instruction Complexity Induces Positional Collapse in Adversarial LLM Evaluation
por: Cacioli, Jon-Paul
Publicado: (2026)
por: Cacioli, Jon-Paul
Publicado: (2026)
PARAPHRASUS : A Comprehensive Benchmark for Evaluating Paraphrase Detection Models
por: Michail, Andrianos, et al.
Publicado: (2024)
por: Michail, Andrianos, et al.
Publicado: (2024)
Integrating Emotional and Linguistic Models for Ethical Compliance in Large Language Models
por: Chang, Edward Y.
Publicado: (2024)
por: Chang, Edward Y.
Publicado: (2024)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
por: Peters, Sydney, et al.
Publicado: (2025)
por: Peters, Sydney, et al.
Publicado: (2025)
PrivacyBench: A Conversational Benchmark for Evaluating Privacy in Personalized AI
por: Mukhopadhyay, Srija, et al.
Publicado: (2025)
por: Mukhopadhyay, Srija, et al.
Publicado: (2025)
UA-Legal-Bench: A Benchmark for Evaluating Large Language Models on Ukrainian Legal Reasoning
por: Ovcharov, Volodymyr
Publicado: (2026)
por: Ovcharov, Volodymyr
Publicado: (2026)
PSST: A Benchmark for Evaluation-driven Text Public-Speaking Style Transfer
por: Sun, Huashan, et al.
Publicado: (2023)
por: Sun, Huashan, et al.
Publicado: (2023)
LeCoDe: A Benchmark Dataset for Interactive Legal Consultation Dialogue Evaluation
por: Yuan, Weikang, et al.
Publicado: (2025)
por: Yuan, Weikang, et al.
Publicado: (2025)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
por: Saji, Alan, et al.
Publicado: (2025)
por: Saji, Alan, et al.
Publicado: (2025)
EduAdapt: A Question Answer Benchmark Dataset for Evaluating Grade-Level Adaptability in LLMs
por: Naeem, Numaan, et al.
Publicado: (2025)
por: Naeem, Numaan, et al.
Publicado: (2025)
Cetvel: A Unified Benchmark for Evaluating Language Understanding, Generation and Cultural Capacity of LLMs for Turkish
por: Er, Yakup Abrek, et al.
Publicado: (2025)
por: Er, Yakup Abrek, et al.
Publicado: (2025)
From Reading to Compressing: Exploring the Multi-document Reader for Prompt Compression
por: Choi, Eunseong, et al.
Publicado: (2024)
por: Choi, Eunseong, et al.
Publicado: (2024)
Instructional Agents: Reducing Teaching Faculty Workload through Multi-Agent Instructional Design
por: Yao, Huaiyuan, et al.
Publicado: (2025)
por: Yao, Huaiyuan, et al.
Publicado: (2025)
FIN-bench-v2: A Unified and Robust Benchmark Suite for Evaluating Finnish Large Language Models
por: Kytöniemi, Joona, et al.
Publicado: (2025)
por: Kytöniemi, Joona, et al.
Publicado: (2025)
Evaluating the Efficacy of Hybrid Deep Learning Models in Distinguishing AI-Generated Text
por: Oketunji, Abiodun Finbarrs
Publicado: (2023)
por: Oketunji, Abiodun Finbarrs
Publicado: (2023)
Controllable Text Summarization: Unraveling Challenges, Approaches, and Prospects -- A Survey
por: Urlana, Ashok, et al.
Publicado: (2023)
por: Urlana, Ashok, et al.
Publicado: (2023)
Exploring News Summarization and Enrichment in a Highly Resource-Scarce Indian Language: A Case Study of Mizo
por: Bala, Abhinaba, et al.
Publicado: (2024)
por: Bala, Abhinaba, et al.
Publicado: (2024)
Streamlining Redundant Layers to Compress Large Language Models
por: Chen, Xiaodong, et al.
Publicado: (2024)
por: Chen, Xiaodong, et al.
Publicado: (2024)
How Instruction-Tuning Imparts Length Control: A Cross-Lingual Mechanistic Analysis
por: Rocchetti, Elisabetta, et al.
Publicado: (2025)
por: Rocchetti, Elisabetta, et al.
Publicado: (2025)
mEdIT: Multilingual Text Editing via Instruction Tuning
por: Raheja, Vipul, et al.
Publicado: (2024)
por: Raheja, Vipul, et al.
Publicado: (2024)
Demystifying Instruction Mixing for Fine-tuning Large Language Models
por: Wang, Renxi, et al.
Publicado: (2023)
por: Wang, Renxi, et al.
Publicado: (2023)
From Benchmarking to Reasoning: A Dual-Aspect, Large-Scale Evaluation of LLMs on Vietnamese Legal Text
por: Le, Van-Truong
Publicado: (2026)
por: Le, Van-Truong
Publicado: (2026)
Residual Drift Dominates Contradiction in Multi-Turn Constraint Reasoning
por: Kawada, Sebastien
Publicado: (2026)
por: Kawada, Sebastien
Publicado: (2026)
The Perplexity Paradox: Why Code Compresses Better Than Math in LLM Prompts
por: Johnson, Warren
Publicado: (2026)
por: Johnson, Warren
Publicado: (2026)
SUBLLM: A Novel Efficient Architecture with Token Sequence Subsampling for LLM
por: Wang, Quandong, et al.
Publicado: (2024)
por: Wang, Quandong, et al.
Publicado: (2024)
Exploring RWKV for Sentence Embeddings: Layer-wise Analysis and Baseline Comparison for Semantic Similarity
por: Pan, Xinghan
Publicado: (2025)
por: Pan, Xinghan
Publicado: (2025)
Automated MCQA Benchmarking at Scale: Evaluating Reasoning Traces as Retrieval Sources for Domain Adaptation of Small Language Models
por: Gokdemir, Ozan, et al.
Publicado: (2025)
por: Gokdemir, Ozan, et al.
Publicado: (2025)
Cultural Benchmarking of LLMs in Standard and Dialectal Arabic Dialogues
por: Kautsar, Muhammad Dehan Al, et al.
Publicado: (2026)
por: Kautsar, Muhammad Dehan Al, et al.
Publicado: (2026)
Tool-Schema Compression Enables Agentic RAG Under Constrained Context Budgets
por: Sakizli, Furkan
Publicado: (2026)
por: Sakizli, Furkan
Publicado: (2026)
Historical Ink: Semantic Shift Detection for 19th Century Spanish
por: Montes, Tony, et al.
Publicado: (2024)
por: Montes, Tony, et al.
Publicado: (2024)
DeEscalWild: A Real-World Benchmark for Automated De-Escalation Training with SLMs
por: Hasan, Md Hasebul, et al.
Publicado: (2026)
por: Hasan, Md Hasebul, et al.
Publicado: (2026)
Evaluating the efficacy of LLM Safety Solutions : The Palit Benchmark Dataset
por: Palit, Sayon, et al.
Publicado: (2025)
por: Palit, Sayon, et al.
Publicado: (2025)
CausalGym: Benchmarking causal interpretability methods on linguistic tasks
por: Arora, Aryaman, et al.
Publicado: (2024)
por: Arora, Aryaman, et al.
Publicado: (2024)
Ejemplares similares
-
The Drill-Down and Fabricate Test (DDFT): A Protocol for Measuring Epistemic Robustness in Language Models
por: Baxi, Rahul
Publicado: (2025) -
Attentive Reasoning Queries: A Systematic Method for Optimizing Instruction-Following in Large Language Models
por: Karov, Bar, et al.
Publicado: (2025) -
Towards Safer Chatbots: Automated Policy Compliance Evaluation of Custom GPTs
por: Rodriguez, David, et al.
Publicado: (2025) -
InstructCMP: Length Control in Sentence Compression through Instruction-based Large Language Models
por: Juseon-Do, et al.
Publicado: (2024) -
LLMs and Memorization: On Quality and Specificity of Copyright Compliance
por: Mueller, Felix B, et al.
Publicado: (2024)