Think Inside the JSON: Reinforcement Strategy for Strict LLM Schema Adherence
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Agarwal, Bhavik, Joshi, Ishan, Rojkova, Viktoria |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
From Paper to Structured JSON: An Agentic Workflow for Compliant BMR Digital Transformation
par: Agarwal, Bhavik, et autres
Publié: (2025)
par: Agarwal, Bhavik, et autres
Publié: (2025)
J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement Learning
par: Whitehouse, Chenxi, et autres
Publié: (2025)
par: Whitehouse, Chenxi, et autres
Publié: (2025)
RAGulating Compliance: A Multi-Agent Knowledge Graph for Regulatory QA
par: Agarwal, Bhavik, et autres
Publié: (2025)
par: Agarwal, Bhavik, et autres
Publié: (2025)
Process Reward Models That Think
par: Khalifa, Muhammad, et autres
Publié: (2025)
par: Khalifa, Muhammad, et autres
Publié: (2025)
EvoSchema: Towards Text-to-SQL Robustness Against Schema Evolution
par: Zhang, Tianshu, et autres
Publié: (2026)
par: Zhang, Tianshu, et autres
Publié: (2026)
Beyond Introspection: Reinforcing Thinking via Externalist Behavioral Feedback
par: Yang, Diji, et autres
Publié: (2024)
par: Yang, Diji, et autres
Publié: (2024)
MDKeyChunker: Single-Call LLM Enrichment with Rolling Keys and Key-Based Restructuring for High-Accuracy RAG
par: Mangla, Bhavik
Publié: (2026)
par: Mangla, Bhavik
Publié: (2026)
Peering Inside the Black Box: Uncovering LLM Errors in Optimization Modelling through Component-Level Evaluation
par: Refai, Dania, et autres
Publié: (2025)
par: Refai, Dania, et autres
Publié: (2025)
Flow-of-Options: Diversified and Improved LLM Reasoning by Thinking Through Options
par: Nair, Lakshmi, et autres
Publié: (2025)
par: Nair, Lakshmi, et autres
Publié: (2025)
Think Big, Generate Quick: LLM-to-SLM for Fast Autoregressive Decoding
par: Bergner, Benjamin, et autres
Publié: (2024)
par: Bergner, Benjamin, et autres
Publié: (2024)
OmniStruct: Universal Text-to-Structure Generation across Diverse Schemas
par: Huang, James Y., et autres
Publié: (2025)
par: Huang, James Y., et autres
Publié: (2025)
RASL: Retrieval Augmented Schema Linking for Massive Database Text-to-SQL
par: Eben, Jeffrey, et autres
Publié: (2025)
par: Eben, Jeffrey, et autres
Publié: (2025)
BRiTE: Bootstrapping Reinforced Thinking Process to Enhance Language Model Reasoning
par: Zhong, Han, et autres
Publié: (2025)
par: Zhong, Han, et autres
Publié: (2025)
Automatic Labelling with Open-source LLMs using Dynamic Label Schema Integration
par: Walshe, Thomas, et autres
Publié: (2025)
par: Walshe, Thomas, et autres
Publié: (2025)
L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
par: Aggarwal, Pranjal, et autres
Publié: (2025)
par: Aggarwal, Pranjal, et autres
Publié: (2025)
From Isolated Conversations to Hierarchical Schemas: Dynamic Tree Memory Representation for LLMs
par: Rezazadeh, Alireza, et autres
Publié: (2024)
par: Rezazadeh, Alireza, et autres
Publié: (2024)
Demystifying Hybrid Thinking: Can LLMs Truly Switch Between Think and No-Think?
par: Wang, Shouren, et autres
Publié: (2025)
par: Wang, Shouren, et autres
Publié: (2025)
Think-Augmented Function Calling: Improving LLM Parameter Accuracy Through Embedded Reasoning
par: Wei, Lei, et autres
Publié: (2026)
par: Wei, Lei, et autres
Publié: (2026)
LLM-based Schema-Guided Extraction and Validation of Missing-Person Intelligence from Heterogeneous Data Sources
par: Castillo, Joshua, et autres
Publié: (2026)
par: Castillo, Joshua, et autres
Publié: (2026)
The Hidden Signal of Verifier Strictness: Controlling and Improving Step-Wise Verification via Selective Latent Steering
par: Zhou, Yefan, et autres
Publié: (2026)
par: Zhou, Yefan, et autres
Publié: (2026)
PolyKV: A Shared Asymmetrically-Compressed KV Cache Pool for Multi-Agent LLM Inference
par: Patel, Ishan, et autres
Publié: (2026)
par: Patel, Ishan, et autres
Publié: (2026)
To Think or Not to Think: Exploring the Unthinking Vulnerability in Large Reasoning Models
par: Zhu, Zihao, et autres
Publié: (2025)
par: Zhu, Zihao, et autres
Publié: (2025)
AdapThink: Adaptive Thinking Preferences for Reasoning Language Model
par: Wan, Xu, et autres
Publié: (2025)
par: Wan, Xu, et autres
Publié: (2025)
AdaptThink: Reasoning Models Can Learn When to Think
par: Zhang, Jiajie, et autres
Publié: (2025)
par: Zhang, Jiajie, et autres
Publié: (2025)
The Evolving Landscape of LLM- and VLM-Integrated Reinforcement Learning
par: Schoepp, Sheila, et autres
Publié: (2025)
par: Schoepp, Sheila, et autres
Publié: (2025)
Reinforce LLM Reasoning through Multi-Agent Reflection
par: Yuan, Yurun, et autres
Publié: (2025)
par: Yuan, Yurun, et autres
Publié: (2025)
Diagnosing Training Inference Mismatch in LLM Reinforcement Learning
par: Zhong, Tianle, et autres
Publié: (2026)
par: Zhong, Tianle, et autres
Publié: (2026)
Offline Reinforcement Learning for LLM Multi-Step Reasoning
par: Wang, Huaijie, et autres
Publié: (2024)
par: Wang, Huaijie, et autres
Publié: (2024)
To Think or Not to Think: The Hidden Cost of Meta-Training with Excessive CoT Examples
par: Kothapalli, Vignesh, et autres
Publié: (2025)
par: Kothapalli, Vignesh, et autres
Publié: (2025)
BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation
par: Kim, Eunsu, et autres
Publié: (2025)
par: Kim, Eunsu, et autres
Publié: (2025)
Pluralistic Behavior Suite: Stress-Testing Multi-Turn Adherence to Custom Behavioral Policies
par: Varshney, Prasoon, et autres
Publié: (2025)
par: Varshney, Prasoon, et autres
Publié: (2025)
Cascading Adaptors to Leverage English Data to Improve Performance of Question Answering for Low-Resource Languages
par: Pandya, Hariom A., et autres
Publié: (2021)
par: Pandya, Hariom A., et autres
Publié: (2021)
LLM as a Neural Architect: Controlled Generation of Image Captioning Models Under Strict API Contracts
par: Jesani, Krunal, et autres
Publié: (2025)
par: Jesani, Krunal, et autres
Publié: (2025)
Inverse Reinforcement Learning with Dynamic Reward Scaling for LLM Alignment
par: Cheng, Ruoxi, et autres
Publié: (2025)
par: Cheng, Ruoxi, et autres
Publié: (2025)
OneKE: A Dockerized Schema-Guided LLM Agent-based Knowledge Extraction System
par: Luo, Yujie, et autres
Publié: (2024)
par: Luo, Yujie, et autres
Publié: (2024)
Think Right: Learning to Mitigate Under-Over Thinking via Adaptive, Attentive Compression
par: Singh, Joykirat, et autres
Publié: (2025)
par: Singh, Joykirat, et autres
Publié: (2025)
Thinking About Thinking: SAGE-nano's Inverse Reasoning for Self-Aware Language Models
par: Jha, Basab, et autres
Publié: (2025)
par: Jha, Basab, et autres
Publié: (2025)
ThinkRouter: Efficient Reasoning via Routing Thinking between Latent and Discrete Spaces
par: Xu, Xin, et autres
Publié: (2026)
par: Xu, Xin, et autres
Publié: (2026)
TASER: Table Agents for Schema-guided Extraction and Recommendation
par: Cho, Nicole, et autres
Publié: (2025)
par: Cho, Nicole, et autres
Publié: (2025)
Efficient Reasoning with Hidden Thinking
par: Shen, Xuan, et autres
Publié: (2025)
par: Shen, Xuan, et autres
Publié: (2025)
Documents similaires
-
From Paper to Structured JSON: An Agentic Workflow for Compliant BMR Digital Transformation
par: Agarwal, Bhavik, et autres
Publié: (2025) -
J1: Incentivizing Thinking in LLM-as-a-Judge via Reinforcement Learning
par: Whitehouse, Chenxi, et autres
Publié: (2025) -
RAGulating Compliance: A Multi-Agent Knowledge Graph for Regulatory QA
par: Agarwal, Bhavik, et autres
Publié: (2025) -
Process Reward Models That Think
par: Khalifa, Muhammad, et autres
Publié: (2025) -
EvoSchema: Towards Text-to-SQL Robustness Against Schema Evolution
par: Zhang, Tianshu, et autres
Publié: (2026)