Instructions Shape Production of Language, not Processing
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Waldis, Andreas, Choshen, Leshem, Hou, Yufang, Perlitz, Yotam |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Holmes: A Benchmark to Assess the Linguistic Competence of Language Models
von: Waldis, Andreas, et al.
Veröffentlicht: (2024)
von: Waldis, Andreas, et al.
Veröffentlicht: (2024)
How to Handle Different Types of Out-of-Distribution Scenarios in Computational Argumentation? A Comprehensive and Fine-Grained Field Study
von: Waldis, Andreas, et al.
Veröffentlicht: (2023)
von: Waldis, Andreas, et al.
Veröffentlicht: (2023)
Dive into the Chasm: Probing the Gap between In- and Cross-Topic Generalization
von: Waldis, Andreas, et al.
Veröffentlicht: (2024)
von: Waldis, Andreas, et al.
Veröffentlicht: (2024)
Growing Pains: Extensible and Efficient LLM Benchmarking Via Fixed Parameter Calibration
von: Habba, Eliya, et al.
Veröffentlicht: (2026)
von: Habba, Eliya, et al.
Veröffentlicht: (2026)
DOVE: A Large-Scale Multi-Dimensional Predictions Dataset Towards Meaningful LLM Evaluation
von: Habba, Eliya, et al.
Veröffentlicht: (2025)
von: Habba, Eliya, et al.
Veröffentlicht: (2025)
Do These LLM Benchmarks Agree? Fixing Benchmark Evaluation with BenchBench
von: Perlitz, Yotam, et al.
Veröffentlicht: (2024)
von: Perlitz, Yotam, et al.
Veröffentlicht: (2024)
Efficient Benchmarking of Language Models
von: Perlitz, Yotam, et al.
Veröffentlicht: (2023)
von: Perlitz, Yotam, et al.
Veröffentlicht: (2023)
Can Gradient Descent Simulate Prompting?
von: Zhang, Eric, et al.
Veröffentlicht: (2025)
von: Zhang, Eric, et al.
Veröffentlicht: (2025)
A Hitchhiker's Guide to Scaling Law Estimation
von: Choshen, Leshem, et al.
Veröffentlicht: (2024)
von: Choshen, Leshem, et al.
Veröffentlicht: (2024)
Deductive Closure Training of Language Models for Coherence, Accuracy, and Updatability
von: Akyürek, Afra Feyza, et al.
Veröffentlicht: (2024)
von: Akyürek, Afra Feyza, et al.
Veröffentlicht: (2024)
The Mighty ToRR: A Benchmark for Table Reasoning and Robustness
von: Ashury-Tahan, Shir, et al.
Veröffentlicht: (2025)
von: Ashury-Tahan, Shir, et al.
Veröffentlicht: (2025)
The ShareLM Collection and Plugin: Contributing Human-Model Chats for the Benefit of the Community
von: Don-Yehiya, Shachar, et al.
Veröffentlicht: (2024)
von: Don-Yehiya, Shachar, et al.
Veröffentlicht: (2024)
Naturally Occurring Feedback is Common, Extractable and Useful
von: Don-Yehiya, Shachar, et al.
Veröffentlicht: (2024)
von: Don-Yehiya, Shachar, et al.
Veröffentlicht: (2024)
Pretraining Language Models for Diachronic Linguistic Change Discovery
von: Fittschen, Elisabeth, et al.
Veröffentlicht: (2025)
von: Fittschen, Elisabeth, et al.
Veröffentlicht: (2025)
Fuse to Forget: Bias Reduction and Selective Memorization through Model Fusion
von: Zaman, Kerem, et al.
Veröffentlicht: (2023)
von: Zaman, Kerem, et al.
Veröffentlicht: (2023)
Unitxt: Flexible, Shareable and Reusable Data Preparation and Evaluation for Generative AI
von: Bandel, Elron, et al.
Veröffentlicht: (2024)
von: Bandel, Elron, et al.
Veröffentlicht: (2024)
Mediocrity is the key for LLM as a Judge Anchor Selection
von: Don-Yehiya, Shachar, et al.
Veröffentlicht: (2026)
von: Don-Yehiya, Shachar, et al.
Veröffentlicht: (2026)
A Pipeline to Assess Merging Methods via Behavior and Internals
von: Sigrist, Yutaro, et al.
Veröffentlicht: (2025)
von: Sigrist, Yutaro, et al.
Veröffentlicht: (2025)
Jump to Conclusions: Short-Cutting Transformers With Linear Transformations
von: Din, Alexander Yom, et al.
Veröffentlicht: (2023)
von: Din, Alexander Yom, et al.
Veröffentlicht: (2023)
PolySQL: Scaling Text-to-SQL Evaluation Across SQL Dialects via Automated Backend Isomorphism
von: Perlitz, Yotam, et al.
Veröffentlicht: (2026)
von: Perlitz, Yotam, et al.
Veröffentlicht: (2026)
Do LLMs Benefit From Their Own Words?
von: Huang, Jenny Y., et al.
Veröffentlicht: (2026)
von: Huang, Jenny Y., et al.
Veröffentlicht: (2026)
LLM Hypnosis: Exploiting User Feedback for Unauthorized Knowledge Injection to All Users
von: Hilel, Almog, et al.
Veröffentlicht: (2025)
von: Hilel, Almog, et al.
Veröffentlicht: (2025)
ComPEFT: Compression for Communicating Parameter Efficient Updates via Sparsification and Quantization
von: Yadav, Prateek, et al.
Veröffentlicht: (2023)
von: Yadav, Prateek, et al.
Veröffentlicht: (2023)
ErrorMap and ErrorAtlas: Charting the Failure Landscape of Large Language Models
von: Ashury-Tahan, Shir, et al.
Veröffentlicht: (2026)
von: Ashury-Tahan, Shir, et al.
Veröffentlicht: (2026)
NeurIPS 2023 LLM Efficiency Fine-tuning Competition
von: Saroufim, Mark, et al.
Veröffentlicht: (2025)
von: Saroufim, Mark, et al.
Veröffentlicht: (2025)
How Safe is Your Safety Metric? Automatic Concatenation Tests for Metric Reliability
von: Fandina, Ora Nova, et al.
Veröffentlicht: (2024)
von: Fandina, Ora Nova, et al.
Veröffentlicht: (2024)
The Lou Dataset -- Exploring the Impact of Gender-Fair Language in German Text Classification
von: Waldis, Andreas, et al.
Veröffentlicht: (2024)
von: Waldis, Andreas, et al.
Veröffentlicht: (2024)
Beneath the Surface of Consistency: Exploring Cross-lingual Knowledge Representation Sharing in LLMs
von: Ifergan, Maxim, et al.
Veröffentlicht: (2024)
von: Ifergan, Maxim, et al.
Veröffentlicht: (2024)
Overview of PerpectiveArg2024: The First Shared Task on Perspective Argument Retrieval
von: Falk, Neele, et al.
Veröffentlicht: (2024)
von: Falk, Neele, et al.
Veröffentlicht: (2024)
Will it Merge? On The Causes of Model Mergeability
von: Rahamim, Adir, et al.
Veröffentlicht: (2026)
von: Rahamim, Adir, et al.
Veröffentlicht: (2026)
NumeroLogic: Number Encoding for Enhanced LLMs' Numerical Reasoning
von: Schwartz, Eli, et al.
Veröffentlicht: (2024)
von: Schwartz, Eli, et al.
Veröffentlicht: (2024)
Resolving Interference (RI): Disentangling Models for Improved Model Merging
von: Ramesh, Pratik, et al.
Veröffentlicht: (2026)
von: Ramesh, Pratik, et al.
Veröffentlicht: (2026)
Unforgettable Generalization in Language Models
von: Zhang, Eric, et al.
Veröffentlicht: (2024)
von: Zhang, Eric, et al.
Veröffentlicht: (2024)
Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty
von: Damani, Mehul, et al.
Veröffentlicht: (2025)
von: Damani, Mehul, et al.
Veröffentlicht: (2025)
CLEAR: Error Analysis via LLM-as-a-Judge Made Easy
von: Yehudai, Asaf, et al.
Veröffentlicht: (2025)
von: Yehudai, Asaf, et al.
Veröffentlicht: (2025)
Diversity Over Size: On the Effect of Sample and Topic Sizes for Topic-Dependent Argument Mining Datasets
von: Schiller, Benjamin, et al.
Veröffentlicht: (2022)
von: Schiller, Benjamin, et al.
Veröffentlicht: (2022)
Beyond Outcome Verification: Verifiable Process Reward Models for Structured Reasoning
von: Pronesti, Massimiliano, et al.
Veröffentlicht: (2026)
von: Pronesti, Massimiliano, et al.
Veröffentlicht: (2026)
CRISP: Complex Reasoning with Interpretable Step-based Plans
von: Vetzler, Matan, et al.
Veröffentlicht: (2025)
von: Vetzler, Matan, et al.
Veröffentlicht: (2025)
Robustness as an Emergent Property of Task Performance
von: Ashury-Tahan, Shir, et al.
Veröffentlicht: (2026)
von: Ashury-Tahan, Shir, et al.
Veröffentlicht: (2026)
TaxoAlign: Scholarly Taxonomy Generation Using Language Models
von: Lahiri, Avishek, et al.
Veröffentlicht: (2025)
von: Lahiri, Avishek, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Holmes: A Benchmark to Assess the Linguistic Competence of Language Models
von: Waldis, Andreas, et al.
Veröffentlicht: (2024) -
How to Handle Different Types of Out-of-Distribution Scenarios in Computational Argumentation? A Comprehensive and Fine-Grained Field Study
von: Waldis, Andreas, et al.
Veröffentlicht: (2023) -
Dive into the Chasm: Probing the Gap between In- and Cross-Topic Generalization
von: Waldis, Andreas, et al.
Veröffentlicht: (2024) -
Growing Pains: Extensible and Efficient LLM Benchmarking Via Fixed Parameter Calibration
von: Habba, Eliya, et al.
Veröffentlicht: (2026) -
DOVE: A Large-Scale Multi-Dimensional Predictions Dataset Towards Meaningful LLM Evaluation
von: Habba, Eliya, et al.
Veröffentlicht: (2025)