Saved in:
| Main Authors: | Chhabra, Anshuman, Datta, Shrestha, Nahin, Shahriar Kabir, Mohapatra, Prasant |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2510.23883 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Revisiting the Effectiveness of LLM Pruning for Test-Time Scaling
by: Monjur, Ocean, et al.
Published: (2026)
by: Monjur, Ocean, et al.
Published: (2026)
Less Diverse, Less Safe: The Indirect But Pervasive Risk of Test-Time Scaling in Large Language Models
by: Nahin, Shahriar Kabir, et al.
Published: (2025)
by: Nahin, Shahriar Kabir, et al.
Published: (2025)
Revisiting Zero-Shot Abstractive Summarization in the Era of Large Language Models from the Perspective of Position Bias
by: Chhabra, Anshuman, et al.
Published: (2024)
by: Chhabra, Anshuman, et al.
Published: (2024)
Outlier Gradient Analysis: Efficiently Identifying Detrimental Training Samples for Deep Learning Models
by: Chhabra, Anshuman, et al.
Published: (2024)
by: Chhabra, Anshuman, et al.
Published: (2024)
Golden Layers and Where to Find Them: Improved Knowledge Editing for Large Language Models Via Layer Gradient Analysis
by: Datta, Shrestha, et al.
Published: (2026)
by: Datta, Shrestha, et al.
Published: (2026)
SafeLens: Deliberate and Efficient Video Guardrails with Fast-and-Slow Screening
by: Nahin, Shahriar Kabir, et al.
Published: (2026)
by: Nahin, Shahriar Kabir, et al.
Published: (2026)
A Survey on Agentic Security: Applications, Threats and Defenses
by: Shahriar, Asif, et al.
Published: (2025)
by: Shahriar, Asif, et al.
Published: (2025)
Watching the AI Watchdogs: A Fairness and Robustness Analysis of AI Safety Moderation Classifiers
by: Achara, Akshit, et al.
Published: (2025)
by: Achara, Akshit, et al.
Published: (2025)
What Is The Performance Ceiling of My Classifier? Utilizing Category-Wise Influence Functions for Pareto Frontier Analysis
by: Nahin, Shahriar Kabir, et al.
Published: (2025)
by: Nahin, Shahriar Kabir, et al.
Published: (2025)
Grounding Natural Language for Multi-agent Decision-Making with Multi-agentic LLMs
by: Huh, Dom, et al.
Published: (2025)
by: Huh, Dom, et al.
Published: (2025)
Maximize Your Diffusion: A Study into Reward Maximization and Alignment for Diffusion-based Control
by: Huh, Dom, et al.
Published: (2025)
by: Huh, Dom, et al.
Published: (2025)
First is Not Really Better Than Last: Evaluating Layer Choice and Aggregation Strategies in Language Model Data Influence Estimation
by: Vitel, Dmytro, et al.
Published: (2025)
by: Vitel, Dmytro, et al.
Published: (2025)
Multi-agent Auto-Bidding with Latent Graph Diffusion Models
by: Huh, Dom, et al.
Published: (2025)
by: Huh, Dom, et al.
Published: (2025)
Multi-agent Reinforcement Learning: A Comprehensive Survey
by: Huh, Dom, et al.
Published: (2023)
by: Huh, Dom, et al.
Published: (2023)
Representation Learning For Efficient Deep Multi-Agent Reinforcement Learning
by: Huh, Dom, et al.
Published: (2024)
by: Huh, Dom, et al.
Published: (2024)
Assessing LLMs for Zero-shot Abstractive Summarization Through the Lens of Relevance Paraphrasing
by: Askari, Hadi, et al.
Published: (2024)
by: Askari, Hadi, et al.
Published: (2024)
Security Threats in Agentic AI System
by: Khan, Raihan, et al.
Published: (2024)
by: Khan, Raihan, et al.
Published: (2024)
SAGA: A Security Architecture for Governing AI Agentic Systems
by: Syros, Georgios, et al.
Published: (2025)
by: Syros, Georgios, et al.
Published: (2025)
Privacy and Security Threat for OpenAI GPTs
by: Wenying, Wei, et al.
Published: (2025)
by: Wenying, Wei, et al.
Published: (2025)
Securing Agentic AI: Threat Modeling and Risk Analysis for Network Monitoring Agentic AI System
by: Zambare, Pallavi, et al.
Published: (2025)
by: Zambare, Pallavi, et al.
Published: (2025)
Rethinking Reasoning in LLMs: Neuro-Symbolic Local RetoMaton Beyond ICL and CoT
by: Mamidala, Rushitha Santhoshi, et al.
Published: (2025)
by: Mamidala, Rushitha Santhoshi, et al.
Published: (2025)
A2AS: Agentic AI Runtime Security and Self-Defense
by: Neelou, Eugene, et al.
Published: (2025)
by: Neelou, Eugene, et al.
Published: (2025)
ASTRIDE: A Security Threat Modeling Platform for Agentic-AI Applications
by: Bandara, Eranga, et al.
Published: (2025)
by: Bandara, Eranga, et al.
Published: (2025)
Towards Understanding and Improving Refusal in Compressed Models via Mechanistic Interpretability
by: Chhabra, Vishnu Kabir, et al.
Published: (2025)
by: Chhabra, Vishnu Kabir, et al.
Published: (2025)
Securing Agentic AI: A Comprehensive Threat Model and Mitigation Framework for Generative AI Agents
by: Narajala, Vineeth Sai, et al.
Published: (2025)
by: Narajala, Vineeth Sai, et al.
Published: (2025)
Cyber Defense Benchmark: Agentic Threat Hunting Evaluation for LLMs in SecOps
by: Chona, Alankrit, et al.
Published: (2026)
by: Chona, Alankrit, et al.
Published: (2026)
Curvature-Weighted Capacity Allocation: A Minimum Description Length Framework for Layer-Adaptive Large Language Model Optimization
by: Amaefuna, Theophilus, et al.
Published: (2026)
by: Amaefuna, Theophilus, et al.
Published: (2026)
Security in the Fine-Tuning Lifecycle of Large Language Models: Threats, Defenses,Evaluation, and Future Directions
by: Li, Wenjuan, et al.
Published: (2026)
by: Li, Wenjuan, et al.
Published: (2026)
Agentic AI for Autonomous Defense in Software Supply Chain Security: Beyond Provenance to Vulnerability Mitigation
by: Syed, Toqeer Ali, et al.
Published: (2025)
by: Syed, Toqeer Ali, et al.
Published: (2025)
A Formal Security Framework for MCP-Based AI Agents: Threat Taxonomy, Verification Models, and Defense Mechanisms
by: Acharya, Nirajan, et al.
Published: (2026)
by: Acharya, Nirajan, et al.
Published: (2026)
Secure Multiparty Generative AI
by: Shrestha, Manil, et al.
Published: (2024)
by: Shrestha, Manil, et al.
Published: (2024)
Neuroplasticity and Corruption in Model Mechanisms: A Case Study Of Indirect Object Identification
by: Chhabra, Vishnu Kabir, et al.
Published: (2025)
by: Chhabra, Vishnu Kabir, et al.
Published: (2025)
On the transferability of Sparse Autoencoders for interpreting compressed models
by: Gupte, Suchit, et al.
Published: (2025)
by: Gupte, Suchit, et al.
Published: (2025)
Self-Supervision in Time for Satellite Images(S3-TSS): A novel method of SSL technique in Satellite images
by: Maurya, Akansh, et al.
Published: (2024)
by: Maurya, Akansh, et al.
Published: (2024)
AI Agents Under Threat: A Survey of Key Security Challenges and Future Pathways
by: Deng, Zehang, et al.
Published: (2024)
by: Deng, Zehang, et al.
Published: (2024)
LayerIF: Estimating Layer Quality for Large Language Models using Influence Functions
by: Askari, Hadi, et al.
Published: (2025)
by: Askari, Hadi, et al.
Published: (2025)
ThreatGPT: An Agentic AI Framework for Enhancing Public Safety through Threat Modeling
by: Zisad, Sharif Noor, et al.
Published: (2025)
by: Zisad, Sharif Noor, et al.
Published: (2025)
Towards Secure Retrieval-Augmented Generation: A Comprehensive Review of Threats, Defenses and Benchmarks
by: Mu, Yanming, et al.
Published: (2026)
by: Mu, Yanming, et al.
Published: (2026)
MCP-Guard: A Multi-Stage Defense-in-Depth Framework for Securing Model Context Protocol in Agentic AI
by: Xing, Wenpeng, et al.
Published: (2025)
by: Xing, Wenpeng, et al.
Published: (2025)
Unraveling Indirect In-Context Learning Using Influence Functions
by: Askari, Hadi, et al.
Published: (2025)
by: Askari, Hadi, et al.
Published: (2025)
Similar Items
-
Revisiting the Effectiveness of LLM Pruning for Test-Time Scaling
by: Monjur, Ocean, et al.
Published: (2026) -
Less Diverse, Less Safe: The Indirect But Pervasive Risk of Test-Time Scaling in Large Language Models
by: Nahin, Shahriar Kabir, et al.
Published: (2025) -
Revisiting Zero-Shot Abstractive Summarization in the Era of Large Language Models from the Perspective of Position Bias
by: Chhabra, Anshuman, et al.
Published: (2024) -
Outlier Gradient Analysis: Efficiently Identifying Detrimental Training Samples for Deep Learning Models
by: Chhabra, Anshuman, et al.
Published: (2024) -
Golden Layers and Where to Find Them: Improved Knowledge Editing for Large Language Models Via Layer Gradient Analysis
by: Datta, Shrestha, et al.
Published: (2026)