CRISP: Complex Reasoning with Interpretable Step-based Plans
Fuente:
arXiv
Saved in:
| Main Authors: | Vetzler, Matan, Lazar, Koren, Uziel, Guy, Hirsch, Eran, Anaby-Tavor, Ateret, Choshen, Leshem |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Effective Red-Teaming of Policy-Adherent Agents
by: Nakash, Itay, et al.
Published: (2025)
by: Nakash, Itay, et al.
Published: (2025)
What's the Plan? Evaluating and Developing Planning-Aware Techniques for Language Models
by: Hirsch, Eran, et al.
Published: (2024)
by: Hirsch, Eran, et al.
Published: (2024)
SpeCrawler: Generating OpenAPI Specifications from API Documentation Using Large Language Models
by: Lazar, Koren, et al.
Published: (2024)
by: Lazar, Koren, et al.
Published: (2024)
Breaking ReAct Agents: Foot-in-the-Door Attack Will Get You In
by: Nakash, Itay, et al.
Published: (2024)
by: Nakash, Itay, et al.
Published: (2024)
Think Again! The Effect of Test-Time Compute on Preferences, Opinions, and Beliefs of Large Language Models
by: Kour, George, et al.
Published: (2025)
by: Kour, George, et al.
Published: (2025)
On the Robustness of Agentic Function Calling
by: Rabinovich, Ella, et al.
Published: (2025)
by: Rabinovich, Ella, et al.
Published: (2025)
OASBuilder: Generating OpenAPI Specifications from Online API Documentation with Large Language Models
by: Lazar, Koren, et al.
Published: (2025)
by: Lazar, Koren, et al.
Published: (2025)
Towards Enforcing Company Policy Adherence in Agentic Workflows
by: Zwerdling, Naama, et al.
Published: (2025)
by: Zwerdling, Naama, et al.
Published: (2025)
Exploring Straightforward Conversational Red-Teaming
by: Kour, George, et al.
Published: (2024)
by: Kour, George, et al.
Published: (2024)
Efficient Agent Evaluation via Diversity-Guided User Simulation
by: Nakash, Itay, et al.
Published: (2026)
by: Nakash, Itay, et al.
Published: (2026)
Near-Miss: Latent Policy Failure Detection in Agentic Workflows
by: Rabinovich, Ella, et al.
Published: (2026)
by: Rabinovich, Ella, et al.
Published: (2026)
A Novel Metric for Measuring the Robustness of Large Language Models in Non-adversarial Scenarios
by: Ackerman, Samuel, et al.
Published: (2024)
by: Ackerman, Samuel, et al.
Published: (2024)
A Hitchhiker's Guide to Scaling Law Estimation
by: Choshen, Leshem, et al.
Published: (2024)
by: Choshen, Leshem, et al.
Published: (2024)
Fuse to Forget: Bias Reduction and Selective Memorization through Model Fusion
by: Zaman, Kerem, et al.
Published: (2023)
by: Zaman, Kerem, et al.
Published: (2023)
ComPEFT: Compression for Communicating Parameter Efficient Updates via Sparsification and Quantization
by: Yadav, Prateek, et al.
Published: (2023)
by: Yadav, Prateek, et al.
Published: (2023)
Beneath the Surface of Consistency: Exploring Cross-lingual Knowledge Representation Sharing in LLMs
by: Ifergan, Maxim, et al.
Published: (2024)
by: Ifergan, Maxim, et al.
Published: (2024)
From Zero to Hero: Cold-Start Anomaly Detection
by: Reiss, Tal, et al.
Published: (2024)
by: Reiss, Tal, et al.
Published: (2024)
Do LLMs Benefit From Their Own Words?
by: Huang, Jenny Y., et al.
Published: (2026)
by: Huang, Jenny Y., et al.
Published: (2026)
Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty
by: Damani, Mehul, et al.
Published: (2025)
by: Damani, Mehul, et al.
Published: (2025)
ErrorMap and ErrorAtlas: Charting the Failure Landscape of Large Language Models
by: Ashury-Tahan, Shir, et al.
Published: (2026)
by: Ashury-Tahan, Shir, et al.
Published: (2026)
Robustness as an Emergent Property of Task Performance
by: Ashury-Tahan, Shir, et al.
Published: (2026)
by: Ashury-Tahan, Shir, et al.
Published: (2026)
Optimizing the Interface Between Knowledge Graphs and LLMs for Complex Reasoning
by: Markovic, Vasilije, et al.
Published: (2025)
by: Markovic, Vasilije, et al.
Published: (2025)
Resolving Interference (RI): Disentangling Models for Improved Model Merging
by: Ramesh, Pratik, et al.
Published: (2026)
by: Ramesh, Pratik, et al.
Published: (2026)
tinyBenchmarks: evaluating LLMs with fewer examples
by: Polo, Felipe Maia, et al.
Published: (2024)
by: Polo, Felipe Maia, et al.
Published: (2024)
Unitxt: Flexible, Shareable and Reusable Data Preparation and Evaluation for Generative AI
by: Bandel, Elron, et al.
Published: (2024)
by: Bandel, Elron, et al.
Published: (2024)
TextArena
by: Guertler, Leon, et al.
Published: (2025)
by: Guertler, Leon, et al.
Published: (2025)
Planning Beyond Text: Graph-based Reasoning for Complex Narrative Generation
by: Gu, Hanwen, et al.
Published: (2026)
by: Gu, Hanwen, et al.
Published: (2026)
PLAN-TUNING: Post-Training Language Models to Learn Step-by-Step Planning for Complex Problem Solving
by: Parmar, Mihir, et al.
Published: (2025)
by: Parmar, Mihir, et al.
Published: (2025)
Genie: Achieving Human Parity in Content-Grounded Datasets Generation
by: Yehudai, Asaf, et al.
Published: (2024)
by: Yehudai, Asaf, et al.
Published: (2024)
GenerationPrograms: Fine-grained Attribution with Executable Programs
by: Wan, David, et al.
Published: (2025)
by: Wan, David, et al.
Published: (2025)
From Building Blocks to Planning: Multi-Step Spatial Reasoning in LLMs with Reinforcement Learning
by: Tahmasbi, Amir, et al.
Published: (2025)
by: Tahmasbi, Amir, et al.
Published: (2025)
PrefixNLI: Detecting Factual Inconsistencies as Soon as They Arise
by: Harary, Sapir, et al.
Published: (2025)
by: Harary, Sapir, et al.
Published: (2025)
A Survey on Model MoErging: Recycling and Routing Among Specialized Experts for Collaborative Learning
by: Yadav, Prateek, et al.
Published: (2024)
by: Yadav, Prateek, et al.
Published: (2024)
Creating a digital poet
by: Tohar, Vered, et al.
Published: (2026)
by: Tohar, Vered, et al.
Published: (2026)
Prompt Repetition Improves Non-Reasoning LLMs
by: Leviathan, Yaniv, et al.
Published: (2025)
by: Leviathan, Yaniv, et al.
Published: (2025)
Compress then Serve: Serving Thousands of LoRA Adapters with Little Overhead
by: Brüel-Gabrielsson, Rickard, et al.
Published: (2024)
by: Brüel-Gabrielsson, Rickard, et al.
Published: (2024)
Survey on Evaluation of LLM-based Agents
by: Yehudai, Asaf, et al.
Published: (2025)
by: Yehudai, Asaf, et al.
Published: (2025)
OraPlan-SQL: A Planning-Centric Framework for Complex Bilingual NL2SQL Reasoning
by: Liu, Marianne Menglin, et al.
Published: (2025)
by: Liu, Marianne Menglin, et al.
Published: (2025)
PlanGEN: A Multi-Agent Framework for Generating Planning and Reasoning Trajectories for Complex Problem Solving
by: Parmar, Mihir, et al.
Published: (2025)
by: Parmar, Mihir, et al.
Published: (2025)
LLM Reasoners: New Evaluation, Library, and Analysis of Step-by-Step Reasoning with Large Language Models
by: Hao, Shibo, et al.
Published: (2024)
by: Hao, Shibo, et al.
Published: (2024)
Similar Items
-
Effective Red-Teaming of Policy-Adherent Agents
by: Nakash, Itay, et al.
Published: (2025) -
What's the Plan? Evaluating and Developing Planning-Aware Techniques for Language Models
by: Hirsch, Eran, et al.
Published: (2024) -
SpeCrawler: Generating OpenAPI Specifications from API Documentation Using Large Language Models
by: Lazar, Koren, et al.
Published: (2024) -
Breaking ReAct Agents: Foot-in-the-Door Attack Will Get You In
by: Nakash, Itay, et al.
Published: (2024) -
Think Again! The Effect of Test-Time Compute on Preferences, Opinions, and Beliefs of Large Language Models
by: Kour, George, et al.
Published: (2025)