Salvato in:
| Autore principale: | Sahoo, Subramanyam |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2512.13821 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
The Good, The Bad, and The Hybrid: A Reward Structure Showdown in Reasoning Models Training
di: Sahoo, Subramanyam
Pubblicazione: (2025)
di: Sahoo, Subramanyam
Pubblicazione: (2025)
The Deepfake Detective: Interpreting Neural Forensics Through Sparse Features and Manifolds
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2025)
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2025)
Calibration Collapse Under Sycophancy Fine-Tuning: How Reward Hacking Breaks Uncertainty Quantification in LLMs
di: Sahoo, Subramanyam
Pubblicazione: (2026)
di: Sahoo, Subramanyam
Pubblicazione: (2026)
The Horcrux: Mechanistically Interpretable Task Decomposition for Detecting and Mitigating Reward Hacking in Embodied AI Systems
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2025)
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2025)
Multi-Agent Systems Execute Arbitrary Malicious Code
di: Triedman, Harold, et al.
Pubblicazione: (2025)
di: Triedman, Harold, et al.
Pubblicazione: (2025)
Boardwalk Empire: How Generative AI is Revolutionizing Economic Paradigms
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2024)
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2024)
Position: The Complexity of Perfect AI Alignment -- Formalizing the RLHF Trilemma
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2025)
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2025)
I Can't Believe It's Not Robust: Catastrophic Collapse of Safety Classifiers under Embedding Drift
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2026)
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2026)
When Shallow Wins: Silent Failures and the Depth-Accuracy Paradox in Latent Reasoning
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2026)
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2026)
SAHOO: Safeguarded Alignment for High-Order Optimization Objectives in Recursive Self-Improvement
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2026)
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2026)
The Reasoning Trap -- Logical Reasoning as a Mechanistic Pathway to Situational Awareness
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2026)
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2026)
Executable Counterfactuals: Improving LLMs' Causal Reasoning Through Code
di: Vashishtha, Aniket, et al.
Pubblicazione: (2025)
di: Vashishtha, Aniket, et al.
Pubblicazione: (2025)
Towards Effectively Leveraging Execution Traces for Program Repair with Code LLMs
di: Haque, Mirazul, et al.
Pubblicazione: (2025)
di: Haque, Mirazul, et al.
Pubblicazione: (2025)
GitChameleon: Unmasking the Version-Switching Capabilities of Code Generation Models
di: Islah, Nizar, et al.
Pubblicazione: (2024)
di: Islah, Nizar, et al.
Pubblicazione: (2024)
GasTrace: Detecting Sandwich Attack Malicious Accounts in Ethereum
di: Liu, Zekai, et al.
Pubblicazione: (2024)
di: Liu, Zekai, et al.
Pubblicazione: (2024)
VoxelCodeBench: Benchmarking 3D World Modeling Through Code Generation
di: Zheng, Yan, et al.
Pubblicazione: (2026)
di: Zheng, Yan, et al.
Pubblicazione: (2026)
Scaling Laws Revisited: Modeling the Role of Data Quality in Language Model Pretraining
di: Subramanyam, Anirudh, et al.
Pubblicazione: (2025)
di: Subramanyam, Anirudh, et al.
Pubblicazione: (2025)
Towards Quantum Machine Learning for Malicious Code Analysis
di: Lopez, Jesus, et al.
Pubblicazione: (2025)
di: Lopez, Jesus, et al.
Pubblicazione: (2025)
Localizing Malicious Outputs from CodeLLM
di: Borana, Mayukh, et al.
Pubblicazione: (2025)
di: Borana, Mayukh, et al.
Pubblicazione: (2025)
Keypoint Aware Masked Image Modelling
di: Krishna, Madhava, et al.
Pubblicazione: (2024)
di: Krishna, Madhava, et al.
Pubblicazione: (2024)
Malicious and Unintentional Disclosure Risks in Large Language Models for Code Generation
di: Rabin, Rafiqul, et al.
Pubblicazione: (2025)
di: Rabin, Rafiqul, et al.
Pubblicazione: (2025)
Learning Unmasking Policies for Diffusion Language Models
di: Jazbec, Metod, et al.
Pubblicazione: (2025)
di: Jazbec, Metod, et al.
Pubblicazione: (2025)
MOCHA: Are Code Language Models Robust Against Multi-Turn Malicious Coding Prompts?
di: Wahed, Muntasir, et al.
Pubblicazione: (2025)
di: Wahed, Muntasir, et al.
Pubblicazione: (2025)
Self-Execution Simulation Improves Coding Models
di: Maimon, Gallil, et al.
Pubblicazione: (2026)
di: Maimon, Gallil, et al.
Pubblicazione: (2026)
Detecting Malicious AI Agents Through Simulated Interactions
di: Pi, Yulu, et al.
Pubblicazione: (2025)
di: Pi, Yulu, et al.
Pubblicazione: (2025)
Unmasking Trees for Tabular Data
di: McCarter, Calvin
Pubblicazione: (2024)
di: McCarter, Calvin
Pubblicazione: (2024)
Interactive and Hybrid Imitation Learning: Provably Beating Behavior Cloning
di: Li, Yichen, et al.
Pubblicazione: (2024)
di: Li, Yichen, et al.
Pubblicazione: (2024)
Autonomous Vehicle Decision-Making Framework for Considering Malicious Behavior at Unsignalized Intersections
di: Li, Qing, et al.
Pubblicazione: (2024)
di: Li, Qing, et al.
Pubblicazione: (2024)
FreeMOCA: Memory-Free Continual Learning for Malicious Code Analysis
di: Asadi, Zahra, et al.
Pubblicazione: (2026)
di: Asadi, Zahra, et al.
Pubblicazione: (2026)
Byzantine Outside, Curious Inside: Reconstructing Data Through Malicious Updates
di: Yue, Kai, et al.
Pubblicazione: (2025)
di: Yue, Kai, et al.
Pubblicazione: (2025)
Reasoning Through Execution: Unifying Process and Outcome Rewards for Code Generation
di: Yu, Zhuohao, et al.
Pubblicazione: (2024)
di: Yu, Zhuohao, et al.
Pubblicazione: (2024)
Understanding Diffusion Models via Code Execution
di: Yu, Cheng
Pubblicazione: (2025)
di: Yu, Cheng
Pubblicazione: (2025)
Trace is the Next AutoDiff: Generative Optimization with Rich Feedback, Execution Traces, and LLMs
di: Cheng, Ching-An, et al.
Pubblicazione: (2024)
di: Cheng, Ching-An, et al.
Pubblicazione: (2024)
Execution Guided Line-by-Line Code Generation
di: Lavon, Boaz, et al.
Pubblicazione: (2025)
di: Lavon, Boaz, et al.
Pubblicazione: (2025)
Provably Safe Model Updates
di: Elmecker-Plakolm, Leo, et al.
Pubblicazione: (2025)
di: Elmecker-Plakolm, Leo, et al.
Pubblicazione: (2025)
Beyond Masked and Unmasked: Discrete Diffusion Models via Partial Masking
di: Chao, Chen-Hao, et al.
Pubblicazione: (2025)
di: Chao, Chen-Hao, et al.
Pubblicazione: (2025)
TaxBreak: Unmasking the Hidden Costs of LLM Inference Through Overhead Decomposition
di: Vellaisamy, Prabhu, et al.
Pubblicazione: (2026)
di: Vellaisamy, Prabhu, et al.
Pubblicazione: (2026)
Execution-Grounded Credit Assignment for GRPO in Code Generation
di: Kumar, Abhijit, et al.
Pubblicazione: (2026)
di: Kumar, Abhijit, et al.
Pubblicazione: (2026)
Lookahead Unmasking Elicits Accurate Decoding in Diffusion Language Models
di: Lee, Sanghyun, et al.
Pubblicazione: (2025)
di: Lee, Sanghyun, et al.
Pubblicazione: (2025)
Accelerated Sampling from Masked Diffusion Models via Entropy Bounded Unmasking
di: Ben-Hamu, Heli, et al.
Pubblicazione: (2025)
di: Ben-Hamu, Heli, et al.
Pubblicazione: (2025)
Documenti analoghi
-
The Good, The Bad, and The Hybrid: A Reward Structure Showdown in Reasoning Models Training
di: Sahoo, Subramanyam
Pubblicazione: (2025) -
The Deepfake Detective: Interpreting Neural Forensics Through Sparse Features and Manifolds
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2025) -
Calibration Collapse Under Sycophancy Fine-Tuning: How Reward Hacking Breaks Uncertainty Quantification in LLMs
di: Sahoo, Subramanyam
Pubblicazione: (2026) -
The Horcrux: Mechanistically Interpretable Task Decomposition for Detecting and Mitigating Reward Hacking in Embodied AI Systems
di: Sahoo, Subramanyam, et al.
Pubblicazione: (2025) -
Multi-Agent Systems Execute Arbitrary Malicious Code
di: Triedman, Harold, et al.
Pubblicazione: (2025)