Probing AI Safety with Source Code
Fuente:
arXiv
Saved in:
| Main Authors: | Narayan, Ujwal, Chaudhari, Shreyas, Kalyan, Ashwin, Rajpurohit, Tanmay, Narasimhan, Karthik, Deshpande, Ameet, Murahari, Vishvak |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Agent Context Protocols Enhance Collective Inference
by: Bhardwaj, Devansh, et al.
Published: (2025)
by: Bhardwaj, Devansh, et al.
Published: (2025)
QualEval: Qualitative Evaluation for Model Improvement
by: Murahari, Vishvak, et al.
Published: (2023)
by: Murahari, Vishvak, et al.
Published: (2023)
RLHF Deciphered: A Critical Analysis of Reinforcement Learning from Human Feedback for LLMs
by: Chaudhari, Shreyas, et al.
Published: (2024)
by: Chaudhari, Shreyas, et al.
Published: (2024)
PersonaGym: Evaluating Persona Agents and LLMs
by: Samuel, Vinay, et al.
Published: (2024)
by: Samuel, Vinay, et al.
Published: (2024)
GEO: Generative Engine Optimization
by: Aggarwal, Pranjal, et al.
Published: (2023)
by: Aggarwal, Pranjal, et al.
Published: (2023)
Language Models can Subtly Deceive Without Lying: A Case Study on Strategic Phrasing in Legislation
by: Dogra, Atharvan, et al.
Published: (2024)
by: Dogra, Atharvan, et al.
Published: (2024)
Engagement Undermines Safety: How Stereotypes and Toxicity Shape Humor in Language Models
by: Dogra, Atharvan, et al.
Published: (2025)
by: Dogra, Atharvan, et al.
Published: (2025)
Bias Runs Deep: Implicit Reasoning Biases in Persona-Assigned LLMs
by: Gupta, Shashank, et al.
Published: (2023)
by: Gupta, Shashank, et al.
Published: (2023)
MARCUS: An Event-Centric NLP Pipeline that generates Character Arcs from Narratives
by: Bhyravajjula, Sriharsh, et al.
Published: (2025)
by: Bhyravajjula, Sriharsh, et al.
Published: (2025)
Can Language Models Solve Olympiad Programming?
by: Shi, Quan, et al.
Published: (2024)
by: Shi, Quan, et al.
Published: (2024)
Language-Guided World Models: A Model-Based Approach to AI Control
by: Zhang, Alex, et al.
Published: (2024)
by: Zhang, Alex, et al.
Published: (2024)
Abstract Reward Processes: Leveraging State Abstraction for Consistent Off-Policy Evaluation
by: Chaudhari, Shreyas, et al.
Published: (2024)
by: Chaudhari, Shreyas, et al.
Published: (2024)
Retaining by Doing: The Role of On-Policy Data in Mitigating Forgetting
by: Chen, Howard, et al.
Published: (2025)
by: Chen, Howard, et al.
Published: (2025)
$τ$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
by: Yao, Shunyu, et al.
Published: (2024)
by: Yao, Shunyu, et al.
Published: (2024)
LLMs are Superior Feedback Providers: Bootstrapping Reasoning for Lie Detection with Self-Generated Feedback
by: Banerjee, Tanushree, et al.
Published: (2024)
by: Banerjee, Tanushree, et al.
Published: (2024)
When Models Know More Than They Can Explain: Quantifying Knowledge Transfer in Human-AI Collaboration
by: Shi, Quan, et al.
Published: (2025)
by: Shi, Quan, et al.
Published: (2025)
From RAGs to rich parameters: Probing how language models utilize external knowledge over parametric information for factual queries
by: Wadhwa, Hitesh, et al.
Published: (2024)
by: Wadhwa, Hitesh, et al.
Published: (2024)
Contextual Experience Replay for Self-Improvement of Language Agents
by: Liu, Yitao, et al.
Published: (2025)
by: Liu, Yitao, et al.
Published: (2025)
$τ^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
by: Barres, Victor, et al.
Published: (2025)
by: Barres, Victor, et al.
Published: (2025)
Dimension-Free Parameterized Approximation Schemes for Hybrid Clustering
by: Gadekar, Ameet, et al.
Published: (2025)
by: Gadekar, Ameet, et al.
Published: (2025)
Plan-Then-Execute: An Empirical Study of User Trust and Team Performance When Using LLM Agents As A Daily Assistant
by: He, Gaole, et al.
Published: (2025)
by: He, Gaole, et al.
Published: (2025)
Cognitive Architectures for Language Agents
by: Sumers, Theodore R., et al.
Published: (2023)
by: Sumers, Theodore R., et al.
Published: (2023)
On the Quantization Robustness of Diffusion Language Models in Coding Benchmarks
by: Gupta, Aarav, et al.
Published: (2026)
by: Gupta, Aarav, et al.
Published: (2026)
$τ$-Knowledge: Evaluating Conversational Agents over Unstructured Knowledge
by: Shi, Quan, et al.
Published: (2026)
by: Shi, Quan, et al.
Published: (2026)
Build, Borrow, or Just Fine-Tune? A Political Scientist's Guide to Choosing NLP Models
by: Meher, Shreyas
Published: (2026)
by: Meher, Shreyas
Published: (2026)
LoRA Soups: Merging LoRAs for Practical Skill Composition Tasks
by: Prabhakar, Akshara, et al.
Published: (2024)
by: Prabhakar, Akshara, et al.
Published: (2024)
Probing Language Models on Their Knowledge Source
by: Tighidet, Zineddine, et al.
Published: (2024)
by: Tighidet, Zineddine, et al.
Published: (2024)
VideoGameBench: Can Vision-Language Models complete popular video games?
by: Zhang, Alex L., et al.
Published: (2025)
by: Zhang, Alex L., et al.
Published: (2025)
Mind the Gap! Choice Independence in Using Multilingual LLMs for Persuasive Co-Writing Tasks in Different Languages
by: Biswas, Shreyan, et al.
Published: (2025)
by: Biswas, Shreyan, et al.
Published: (2025)
A Critical Study of What Code-LLMs (Do Not) Learn
by: Anand, Abhinav, et al.
Published: (2024)
by: Anand, Abhinav, et al.
Published: (2024)
Global Beats, Local Tongue: Studying Code Switching in K-pop Hits on Billboard Charts
by: Sankaran, Aditya Narayan, et al.
Published: (2025)
by: Sankaran, Aditya Narayan, et al.
Published: (2025)
Graphing the Truth: Structured Visualizations for Automated Hallucination Detection in LLMs
by: Agrawal, Tanmay
Published: (2025)
by: Agrawal, Tanmay
Published: (2025)
CodeNav: Beyond tool-use to using real-world codebases with LLM agents
by: Gupta, Tanmay, et al.
Published: (2024)
by: Gupta, Tanmay, et al.
Published: (2024)
Reinforcement Learning for Long-Horizon Multi-Turn Search Agents
by: Kalyan, Vivek, et al.
Published: (2025)
by: Kalyan, Vivek, et al.
Published: (2025)
Comparing Developer and LLM Biases in Code Evaluation
by: Mittal, Aditya, et al.
Published: (2026)
by: Mittal, Aditya, et al.
Published: (2026)
Quantifying reliance on external information over parametric knowledge during Retrieval Augmented Generation (RAG) using mechanistic analysis
by: Ghosh, Reshmi, et al.
Published: (2024)
by: Ghosh, Reshmi, et al.
Published: (2024)
CodeCipher: Learning to Obfuscate Source Code Against LLMs
by: Lin, Yalan, et al.
Published: (2024)
by: Lin, Yalan, et al.
Published: (2024)
Text Is Not All You Need: Multimodal Prompting Helps LLMs Understand Humor
by: Baluja, Ashwin
Published: (2024)
by: Baluja, Ashwin
Published: (2024)
ShieldGemma: Generative AI Content Moderation Based on Gemma
by: Zeng, Wenjun, et al.
Published: (2024)
by: Zeng, Wenjun, et al.
Published: (2024)
Tool-Aware Planning in Contact Center AI: Evaluating LLMs through Lineage-Guided Query Decomposition
by: Nathan, Varun, et al.
Published: (2026)
by: Nathan, Varun, et al.
Published: (2026)
Similar Items
-
Agent Context Protocols Enhance Collective Inference
by: Bhardwaj, Devansh, et al.
Published: (2025) -
QualEval: Qualitative Evaluation for Model Improvement
by: Murahari, Vishvak, et al.
Published: (2023) -
RLHF Deciphered: A Critical Analysis of Reinforcement Learning from Human Feedback for LLMs
by: Chaudhari, Shreyas, et al.
Published: (2024) -
PersonaGym: Evaluating Persona Agents and LLMs
by: Samuel, Vinay, et al.
Published: (2024) -
GEO: Generative Engine Optimization
by: Aggarwal, Pranjal, et al.
Published: (2023)