A Pragmatic Way to Measure Chain-of-Thought Monitorability
Fuente:
arXiv
Salvato in:
| Autori principali: | Emmons, Scott, Zimmermann, Roland S., Elson, David K., Shah, Rohin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
FGDM: Reasoning Aware Multi-Agentic Framework for Software Bug Detection using Chain of Thought and Tree of Thought Prompting
di: Padmanabhuni, Srita, et al.
Pubblicazione: (2026)
di: Padmanabhuni, Srita, et al.
Pubblicazione: (2026)
Aligned, Orthogonal or In-conflict: When can we safely optimize Chain-of-Thought?
di: Kaufmann, Max, et al.
Pubblicazione: (2026)
di: Kaufmann, Max, et al.
Pubblicazione: (2026)
Monitoring Machine Learning Systems: A Multivocal Literature Review
di: Naveed, Hira, et al.
Pubblicazione: (2025)
di: Naveed, Hira, et al.
Pubblicazione: (2025)
TeleResilienceBench: Quantifying Resilience for LLM Reasoning in Telecommunications
di: Gajjar, Pranshav, et al.
Pubblicazione: (2026)
di: Gajjar, Pranshav, et al.
Pubblicazione: (2026)
StackSight: Unveiling WebAssembly through Large Language Models and Neurosymbolic Chain-of-Thought Decompilation
di: Fang, Weike, et al.
Pubblicazione: (2024)
di: Fang, Weike, et al.
Pubblicazione: (2024)
Can LLMs Reason Like Automated Theorem Provers for Rust Verification? VCoT-Bench: Evaluating via Verification Chain of Thought
di: Xie, Zichen, et al.
Pubblicazione: (2026)
di: Xie, Zichen, et al.
Pubblicazione: (2026)
Expert-Driven Monitoring of Operational ML Models
di: Leest, Joran, et al.
Pubblicazione: (2024)
di: Leest, Joran, et al.
Pubblicazione: (2024)
Beyond the Comfort Zone: Emerging Solutions to Overcome Challenges in Integrating LLMs into Software Products
di: Nahar, Nadia, et al.
Pubblicazione: (2024)
di: Nahar, Nadia, et al.
Pubblicazione: (2024)
Understanding Practitioners Perspectives on Monitoring Machine Learning Systems
di: Naveed, Hira, et al.
Pubblicazione: (2025)
di: Naveed, Hira, et al.
Pubblicazione: (2025)
Uncertainty-Guided Label Rebalancing for CPS Safety Monitoring
di: Ayotunde, John, et al.
Pubblicazione: (2026)
di: Ayotunde, John, et al.
Pubblicazione: (2026)
Blockchain-Enabled Accountability in Data Supply Chain: A Data Bill of Materials Approach
di: Liu, Yue, et al.
Pubblicazione: (2024)
di: Liu, Yue, et al.
Pubblicazione: (2024)
OSS-Bench: Benchmark Generator for Coding LLMs
di: Jiang, Yuancheng, et al.
Pubblicazione: (2025)
di: Jiang, Yuancheng, et al.
Pubblicazione: (2025)
Towards MLOps: A DevOps Tools Recommender System for Machine Learning System
di: Shah, Pir Sami Ullah, et al.
Pubblicazione: (2024)
di: Shah, Pir Sami Ullah, et al.
Pubblicazione: (2024)
Leveraging Code Cohesion Analysis to Identify Source Code Supply Chain Attacks
di: Reuben, Maor, et al.
Pubblicazione: (2025)
di: Reuben, Maor, et al.
Pubblicazione: (2025)
From Data Lifting to Continuous Risk Estimation: A Process-Aware Pipeline for Predictive Monitoring of Clinical Pathways
di: Ardimento, Pasquale, et al.
Pubblicazione: (2026)
di: Ardimento, Pasquale, et al.
Pubblicazione: (2026)
OS-Harm: A Benchmark for Measuring Safety of Computer Use Agents
di: Kuntz, Thomas, et al.
Pubblicazione: (2025)
di: Kuntz, Thomas, et al.
Pubblicazione: (2025)
Towards Enhancing the Reproducibility of Deep Learning Bugs: An Empirical Study
di: Shah, Mehil B., et al.
Pubblicazione: (2024)
di: Shah, Mehil B., et al.
Pubblicazione: (2024)
An Empirical Analysis of Machine Learning Model and Dataset Documentation, Supply Chain, and Licensing Challenges on Hugging Face
di: Stalnaker, Trevor, et al.
Pubblicazione: (2025)
di: Stalnaker, Trevor, et al.
Pubblicazione: (2025)
How to Sustainably Monitor ML-Enabled Systems? Accuracy and Energy Efficiency Tradeoffs in Concept Drift Detection
di: Omar, Rafiullah, et al.
Pubblicazione: (2024)
di: Omar, Rafiullah, et al.
Pubblicazione: (2024)
Clotho: Measuring Task-Specific Pre-Generation Test Adequacy for LLM Inputs
di: Yoon, Juyeon, et al.
Pubblicazione: (2025)
di: Yoon, Juyeon, et al.
Pubblicazione: (2025)
Architecting Digital Twins for Intelligent Transportation Systems
di: Bhatt, Hiya, et al.
Pubblicazione: (2025)
di: Bhatt, Hiya, et al.
Pubblicazione: (2025)
The Constraint Tax: Measuring Validity-Correctness Tradeoffs in Structured Outputs for Small Language Models
di: Ray, Jaideep
Pubblicazione: (2026)
di: Ray, Jaideep
Pubblicazione: (2026)
Push Your Agent: Measuring and Enforcing Quantitative Goal Persistence in Long-Horizon LLM Agents
di: Cai, Yuandao, et al.
Pubblicazione: (2026)
di: Cai, Yuandao, et al.
Pubblicazione: (2026)
Agile Story-Point Estimation: Is RAG a Better Way to Go?
di: Maha, Lamyea, et al.
Pubblicazione: (2026)
di: Maha, Lamyea, et al.
Pubblicazione: (2026)
MoTCoder: Elevating Large Language Models with Modular of Thought for Challenging Programming Tasks
di: Li, Jingyao, et al.
Pubblicazione: (2023)
di: Li, Jingyao, et al.
Pubblicazione: (2023)
Towards Predicting Multi-Vulnerability Attack Chains in Software Supply Chains from Software Bill of Materials Graphs
di: Baird, Laura, et al.
Pubblicazione: (2026)
di: Baird, Laura, et al.
Pubblicazione: (2026)
SMARLA: A Safety Monitoring Approach for Deep Reinforcement Learning Agents
di: Zolfagharian, Amirhossein, et al.
Pubblicazione: (2023)
di: Zolfagharian, Amirhossein, et al.
Pubblicazione: (2023)
A Metamorphic Testing Perspective on Knowledge Distillation for Language Models of Code: Does the Student Deeply Mimic the Teacher?
di: Awal, Md. Abdul, et al.
Pubblicazione: (2025)
di: Awal, Md. Abdul, et al.
Pubblicazione: (2025)
Basic Legibility Protocols Improve Trusted Monitoring
di: Sreevatsa, Ashwin, et al.
Pubblicazione: (2026)
di: Sreevatsa, Ashwin, et al.
Pubblicazione: (2026)
Leveraging LLMs for Legacy Code Modernization: Challenges and Opportunities for LLM-Generated Documentation
di: Diggs, Colin, et al.
Pubblicazione: (2024)
di: Diggs, Colin, et al.
Pubblicazione: (2024)
Monitizer: Automating Design and Evaluation of Neural Network Monitors
di: Azeem, Muqsit, et al.
Pubblicazione: (2024)
di: Azeem, Muqsit, et al.
Pubblicazione: (2024)
Carbon-Taxed Transformers: A Green Compression Pipeline for Overgrown Language Models
di: Alam, Ajmain Inqiad, et al.
Pubblicazione: (2026)
di: Alam, Ajmain Inqiad, et al.
Pubblicazione: (2026)
Automated Machine Learning: A Case Study on Non-Intrusive Appliance Load Monitoring
di: Moin, Armin, et al.
Pubblicazione: (2022)
di: Moin, Armin, et al.
Pubblicazione: (2022)
Enhancing Business Process Simulation Models with Extraneous Activity Delays
di: Chapela-Campa, David, et al.
Pubblicazione: (2022)
di: Chapela-Campa, David, et al.
Pubblicazione: (2022)
What do we know about Hugging Face? A systematic literature review and quantitative validation of qualitative claims
di: Jones, Jason, et al.
Pubblicazione: (2024)
di: Jones, Jason, et al.
Pubblicazione: (2024)
David vs. Goliath: A comparative study of different-sized LLMs for code generation in the domain of automotive scenario generation
di: Bauerfeind, Philipp, et al.
Pubblicazione: (2025)
di: Bauerfeind, Philipp, et al.
Pubblicazione: (2025)
StackEval: Benchmarking LLMs in Coding Assistance
di: Shah, Nidhish, et al.
Pubblicazione: (2024)
di: Shah, Nidhish, et al.
Pubblicazione: (2024)
BugGen: A Self-Correcting Multi-Agent LLM Pipeline for Realistic RTL Bug Synthesis
di: Jasper, Surya, et al.
Pubblicazione: (2025)
di: Jasper, Surya, et al.
Pubblicazione: (2025)
StructCoder: Structure-Aware Transformer for Code Generation
di: Tipirneni, Sindhu, et al.
Pubblicazione: (2022)
di: Tipirneni, Sindhu, et al.
Pubblicazione: (2022)
Bridging Expert Knowledge with Deep Learning Techniques for Just-In-Time Defect Prediction
di: Zhou, Xin, et al.
Pubblicazione: (2024)
di: Zhou, Xin, et al.
Pubblicazione: (2024)
Documenti analoghi
-
FGDM: Reasoning Aware Multi-Agentic Framework for Software Bug Detection using Chain of Thought and Tree of Thought Prompting
di: Padmanabhuni, Srita, et al.
Pubblicazione: (2026) -
Aligned, Orthogonal or In-conflict: When can we safely optimize Chain-of-Thought?
di: Kaufmann, Max, et al.
Pubblicazione: (2026) -
Monitoring Machine Learning Systems: A Multivocal Literature Review
di: Naveed, Hira, et al.
Pubblicazione: (2025) -
TeleResilienceBench: Quantifying Resilience for LLM Reasoning in Telecommunications
di: Gajjar, Pranshav, et al.
Pubblicazione: (2026) -
StackSight: Unveiling WebAssembly through Large Language Models and Neurosymbolic Chain-of-Thought Decompilation
di: Fang, Weike, et al.
Pubblicazione: (2024)