PACIFIC: a framework for generating benchmarks to check Precise Automatically Checked Instruction Following In Code
Fuente:
arXiv
Saved in:
| Main Authors: | Dreyfuss, Itay, Nassar, Antonio Abu, Ackerman, Samuel, David, Axel Ben, Farchi, Eitan, Katan, Rami, Raz, Orna, Zalmanovici, Marcel |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Using Combinatorial Optimization to Design a High quality LLM Solution
by: Ackerman, Samuel, et al.
Published: (2024)
by: Ackerman, Samuel, et al.
Published: (2024)
Automatic Generation of Benchmarks and Reliable LLM Judgment for Code Tasks
by: Farchi, Eitan, et al.
Published: (2024)
by: Farchi, Eitan, et al.
Published: (2024)
Generating Unseen Code Tests In Infinitum
by: Zalmanovici, Marcel, et al.
Published: (2024)
by: Zalmanovici, Marcel, et al.
Published: (2024)
Evaluating perturbation robustness of generative systems that use COBOL code inputs
by: Ackerman, Samuel, et al.
Published: (2025)
by: Ackerman, Samuel, et al.
Published: (2025)
Statistical multi-metric evaluation and visualization of LLM system predictive performance
by: Ackerman, Samuel, et al.
Published: (2025)
by: Ackerman, Samuel, et al.
Published: (2025)
Automated Validation of LLM-based Evaluators for Software Engineering Artifacts
by: Fandina, Ora Nova, et al.
Published: (2025)
by: Fandina, Ora Nova, et al.
Published: (2025)
Enhancing Formal Software Specification with Artificial Intelligence
by: Nassar, Antonio Abu, et al.
Published: (2026)
by: Nassar, Antonio Abu, et al.
Published: (2026)
Vintage Code, Modern Judges: Meta-Validation in Low Data Regimes
by: Fandina, Ora Nova, et al.
Published: (2025)
by: Fandina, Ora Nova, et al.
Published: (2025)
How Safe is Your Safety Metric? Automatic Concatenation Tests for Metric Reliability
by: Fandina, Ora Nova, et al.
Published: (2024)
by: Fandina, Ora Nova, et al.
Published: (2024)
Exploring Straightforward Conversational Red-Teaming
by: Kour, George, et al.
Published: (2024)
by: Kour, George, et al.
Published: (2024)
Beyond Blind Spots: Analytic Hints for Mitigating LLM-Based Evaluation Pitfalls
by: Fandina, Ora Nova, et al.
Published: (2025)
by: Fandina, Ora Nova, et al.
Published: (2025)
An Agent-Based Framework for the Automatic Validation of Mathematical Optimization Models
by: Zadorojniy, Alexander, et al.
Published: (2025)
by: Zadorojniy, Alexander, et al.
Published: (2025)
A Novel Metric for Measuring the Robustness of Large Language Models in Non-adversarial Scenarios
by: Ackerman, Samuel, et al.
Published: (2024)
by: Ackerman, Samuel, et al.
Published: (2024)
Effective Technical Reviews
by: Ballentine, Scott, et al.
Published: (2024)
by: Ballentine, Scott, et al.
Published: (2024)
Quality Engineering for Agile and DevOps on the Cloud and Edge
by: Farchi, Eitan, et al.
Published: (2023)
by: Farchi, Eitan, et al.
Published: (2023)
A Practical Approach to Combinatorial Test Design
by: Farchi, Eitan, et al.
Published: (2024)
by: Farchi, Eitan, et al.
Published: (2024)
Uncovering Code Insights: Leveraging GitHub Artifacts for Deeper Code Understanding
by: Nevo, Ziv, et al.
Published: (2025)
by: Nevo, Ziv, et al.
Published: (2025)
Planted-solution SAT and Ising benchmarks from integer factorization
by: Hen, Itay
Published: (2026)
by: Hen, Itay
Published: (2026)
LaajMeter: A Framework for LaaJ Evaluation
by: Ackerman, Samuel, et al.
Published: (2025)
by: Ackerman, Samuel, et al.
Published: (2025)
Technique to Baseline QE Artefact Generation Aligned to Quality Metrics
by: Farchi, Eitan, et al.
Published: (2025)
by: Farchi, Eitan, et al.
Published: (2025)
Check out the check-in: airport work hazards
by: Ellen Rosskam (Author)
Published: (2002)
by: Ellen Rosskam (Author)
Published: (2002)
Generalized Coverage Criteria for Combinatorial Sequence Testing
by: Elyasaf, Achiya, et al.
Published: (2022)
by: Elyasaf, Achiya, et al.
Published: (2022)
Check-up pour les professionnels du check-in
by: Ellen Rosskam (Author)
Published: (2002)
by: Ellen Rosskam (Author)
Published: (2002)
A Hierarchy of Nondeterminism
by: Radi, Bader Abu, et al.
Published: (2022)
by: Radi, Bader Abu, et al.
Published: (2022)
Automatic Fact-checking in English and Telugu
by: Chikkala, Ravi Kiran, et al.
Published: (2025)
by: Chikkala, Ravi Kiran, et al.
Published: (2025)
Black-Box Bug-Amplification for Multithreaded Software
by: Weiss, Yeshayahu, et al.
Published: (2025)
by: Weiss, Yeshayahu, et al.
Published: (2025)
Simulation by Rounds of Letter-to-Letter Transducers
by: Nassar, Antonio Abu, et al.
Published: (2021)
by: Nassar, Antonio Abu, et al.
Published: (2021)
Boosting Instruction Following at Scale
by: Elder, Ben, et al.
Published: (2025)
by: Elder, Ben, et al.
Published: (2025)
A Note About Majority Colorings of Countable DAGs
by: Bosek, Bartłomiej, et al.
Published: (2024)
by: Bosek, Bartłomiej, et al.
Published: (2024)
CastillonMiguel/CastillonMiguel-Check_Complementary_data: Release check
by: Miguel Castillón
Published: (2025)
by: Miguel Castillón
Published: (2025)
Precision over Diversity: High-Precision Reward Generalizes to Robust Instruction Following
by: Zeng, Yirong, et al.
Published: (2026)
by: Zeng, Yirong, et al.
Published: (2026)
Human Learning about AI
by: Dreyfuss, Bnaya, et al.
Published: (2024)
by: Dreyfuss, Bnaya, et al.
Published: (2024)
Diseño y construcción de una secadora automática para cacao a base de aire caliente tipo rotatorio para una capacidad de 500 kg
by: Javier Orna
Published: (2018)
by: Javier Orna
Published: (2018)
Bridging vision language model (VLM) evaluation gaps with a framework for scalable and cost-effective benchmark generation
by: Rädsch, Tim, et al.
Published: (2025)
by: Rädsch, Tim, et al.
Published: (2025)
FactCheck Editor: Multilingual Text Editor with End-to-End fact-checking
by: Setty, Vinay
Published: (2024)
by: Setty, Vinay
Published: (2024)
THE PACIFIC: COMMUNIST ON THE DOCKS
Published: (1951)
Published: (1951)
Is this correct? Let's check!
by: Ben-Eliezer, Omri, et al.
Published: (2022)
by: Ben-Eliezer, Omri, et al.
Published: (2022)
Light-IF: Endowing LLMs with Generalizable Reasoning via Preview and Self-Checking for Complex Instruction Following
by: Wang, Chenyang, et al.
Published: (2025)
by: Wang, Chenyang, et al.
Published: (2025)
Automatic Layout Planning for Visually-Rich Documents with Instruction-Following Models
by: Zhu, Wanrong, et al.
Published: (2024)
by: Zhu, Wanrong, et al.
Published: (2024)
Real-Time Weather Image Classification with SVM
by: Ship, Eden, et al.
Published: (2024)
by: Ship, Eden, et al.
Published: (2024)
Similar Items
-
Using Combinatorial Optimization to Design a High quality LLM Solution
by: Ackerman, Samuel, et al.
Published: (2024) -
Automatic Generation of Benchmarks and Reliable LLM Judgment for Code Tasks
by: Farchi, Eitan, et al.
Published: (2024) -
Generating Unseen Code Tests In Infinitum
by: Zalmanovici, Marcel, et al.
Published: (2024) -
Evaluating perturbation robustness of generative systems that use COBOL code inputs
by: Ackerman, Samuel, et al.
Published: (2025) -
Statistical multi-metric evaluation and visualization of LLM system predictive performance
by: Ackerman, Samuel, et al.
Published: (2025)