Agentic AI Needs a Systems Theory
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Miehling, Erik, Ramamurthy, Karthikeyan Natesan, Varshney, Kush R., Riemer, Matthew, Bouneffouf, Djallel, Richards, John T., Dhurandhar, Amit, Daly, Elizabeth M., Hind, Michael, Sattigeri, Prasanna, Wei, Dennis, Rawat, Ambrish, Gajcin, Jasmina, Geyer, Werner |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
The Ultimate Test of Superintelligent AI Agents: Can an AI Balance Care and Control in Asymmetric Relationships?
von: Bouneffouf, Djallel, et al.
Veröffentlicht: (2025)
von: Bouneffouf, Djallel, et al.
Veröffentlicht: (2025)
Evaluating the Prompt Steerability of Large Language Models
von: Miehling, Erik, et al.
Veröffentlicht: (2024)
von: Miehling, Erik, et al.
Veröffentlicht: (2024)
Value Alignment from Unstructured Text
von: Padhi, Inkit, et al.
Veröffentlicht: (2024)
von: Padhi, Inkit, et al.
Veröffentlicht: (2024)
Who Sees the Risk? Stakeholder Conflicts and Explanatory Policies in LLM-based Risk Assessment
von: Yadav, Srishti, et al.
Veröffentlicht: (2025)
von: Yadav, Srishti, et al.
Veröffentlicht: (2025)
The Effectiveness of Approximate Regularized Replay for Efficient Supervised Fine-Tuning of Large Language Models
von: Riemer, Matthew, et al.
Veröffentlicht: (2025)
von: Riemer, Matthew, et al.
Veröffentlicht: (2025)
Ranking Large Language Models without Ground Truth
von: Dhurandhar, Amit, et al.
Veröffentlicht: (2024)
von: Dhurandhar, Amit, et al.
Veröffentlicht: (2024)
Scopes of Alignment
von: Varshney, Kush R., et al.
Veröffentlicht: (2025)
von: Varshney, Kush R., et al.
Veröffentlicht: (2025)
Programming Refusal with Conditional Activation Steering
von: Lee, Bruce W., et al.
Veröffentlicht: (2024)
von: Lee, Bruce W., et al.
Veröffentlicht: (2024)
Trust Regions for Explanations via Black-Box Probabilistic Certification
von: Dhurandhar, Amit, et al.
Veröffentlicht: (2024)
von: Dhurandhar, Amit, et al.
Veröffentlicht: (2024)
Identifying Sub-networks in Neural Networks via Functionally Similar Representations
von: Gao, Tian, et al.
Veröffentlicht: (2024)
von: Gao, Tian, et al.
Veröffentlicht: (2024)
Contextual Moral Value Alignment Through Context-Based Aggregation
von: Dognin, Pierre, et al.
Veröffentlicht: (2024)
von: Dognin, Pierre, et al.
Veröffentlicht: (2024)
Interpreting LLM-as-a-Judge Policies via Verifiable Global Explanations
von: Gajcin, Jasmina, et al.
Veröffentlicht: (2025)
von: Gajcin, Jasmina, et al.
Veröffentlicht: (2025)
Multi-Level Explanations for Generative Language Models
von: Paes, Lucas Monteiro, et al.
Veröffentlicht: (2024)
von: Paes, Lucas Monteiro, et al.
Veröffentlicht: (2024)
LCGuard: Latent Communication Guard for Safe KV Sharing in Multi-Agent Systems
von: Asif, Sadia, et al.
Veröffentlicht: (2026)
von: Asif, Sadia, et al.
Veröffentlicht: (2026)
Language Models in Dialogue: Conversational Maxims for Human-AI Interactions
von: Miehling, Erik, et al.
Veröffentlicht: (2024)
von: Miehling, Erik, et al.
Veröffentlicht: (2024)
AgentSCOPE: Evaluating Contextual Privacy Across Agentic Workflows
von: Ngong, Ivoline C., et al.
Veröffentlicht: (2026)
von: Ngong, Ivoline C., et al.
Veröffentlicht: (2026)
Generate, Evaluate, Iterate: Synthetic Data for Human-in-the-Loop Refinement of LLM Judges
von: Do, Hyo Jin, et al.
Veröffentlicht: (2025)
von: Do, Hyo Jin, et al.
Veröffentlicht: (2025)
AI Steerability 360: A Toolkit for Steering Large Language Models
von: Miehling, Erik, et al.
Veröffentlicht: (2026)
von: Miehling, Erik, et al.
Veröffentlicht: (2026)
CELL your Model: Contrastive Explanations for Large Language Models
von: Luss, Ronny, et al.
Veröffentlicht: (2024)
von: Luss, Ronny, et al.
Veröffentlicht: (2024)
Attack Atlas: A Practitioner's Perspective on Challenges and Pitfalls in Red Teaming GenAI
von: Rawat, Ambrish, et al.
Veröffentlicht: (2024)
von: Rawat, Ambrish, et al.
Veröffentlicht: (2024)
Mitigating Misalignment Contagion by Steering with Implicit Traits
von: Chang, Maria, et al.
Veröffentlicht: (2026)
von: Chang, Maria, et al.
Veröffentlicht: (2026)
Final-Model-Only Data Attribution with a Unifying View of Gradient-Based Methods
von: Wei, Dennis, et al.
Veröffentlicht: (2024)
von: Wei, Dennis, et al.
Veröffentlicht: (2024)
Language Models Coupled with Metacognition Can Outperform Reasoning Models
von: Khandelwal, Vedant, et al.
Veröffentlicht: (2025)
von: Khandelwal, Vedant, et al.
Veröffentlicht: (2025)
Survey: Multi-Armed Bandits Meet Large Language Models
von: Bouneffouf, Djallel, et al.
Veröffentlicht: (2025)
von: Bouneffouf, Djallel, et al.
Veröffentlicht: (2025)
Alignment Studio: Aligning Large Language Models to Particular Contextual Regulations
von: Achintalwar, Swapnaja, et al.
Veröffentlicht: (2024)
von: Achintalwar, Swapnaja, et al.
Veröffentlicht: (2024)
Protecting Users From Themselves: Safeguarding Contextual Privacy in Interactions with Conversational Agents
von: Ngong, Ivoline, et al.
Veröffentlicht: (2025)
von: Ngong, Ivoline, et al.
Veröffentlicht: (2025)
Adversarial Prompt Evaluation: Systematic Benchmarking of Guardrails Against Prompt Input Attacks on LLMs
von: Zizzo, Giulio, et al.
Veröffentlicht: (2025)
von: Zizzo, Giulio, et al.
Veröffentlicht: (2025)
Large Language Model Confidence Estimation via Black-Box Access
von: Pedapati, Tejaswini, et al.
Veröffentlicht: (2024)
von: Pedapati, Tejaswini, et al.
Veröffentlicht: (2024)
Granite Guardian
von: Padhi, Inkit, et al.
Veröffentlicht: (2024)
von: Padhi, Inkit, et al.
Veröffentlicht: (2024)
Redefining Counterfactual Explanations for Reinforcement Learning: Overview, Challenges and Opportunities
von: Gajcin, Jasmina, et al.
Veröffentlicht: (2022)
von: Gajcin, Jasmina, et al.
Veröffentlicht: (2022)
ACTER: Diverse and Actionable Counterfactual Sequences for Explaining and Diagnosing RL Policies
von: Gajcin, Jasmina, et al.
Veröffentlicht: (2024)
von: Gajcin, Jasmina, et al.
Veröffentlicht: (2024)
Assessing AI Utility: The Random Guesser Test for Sequential Decision-Making Systems
von: Ide, Shun, et al.
Veröffentlicht: (2024)
von: Ide, Shun, et al.
Veröffentlicht: (2024)
Conversational Topic Recommendation in Counseling and Psychotherapy with Decision Transformer and Large Language Models
von: Gunal, Aylin, et al.
Veröffentlicht: (2024)
von: Gunal, Aylin, et al.
Veröffentlicht: (2024)
Enhancing Value Alignment of LLMs with Multi-agent system and Combinatorial Fusion
von: Wu, Yuanhong, et al.
Veröffentlicht: (2026)
von: Wu, Yuanhong, et al.
Veröffentlicht: (2026)
ICX360: In-Context eXplainability 360 Toolkit
von: Wei, Dennis, et al.
Veröffentlicht: (2025)
von: Wei, Dennis, et al.
Veröffentlicht: (2025)
Semifactual Explanations for Reinforcement Learning
von: Gajcin, Jasmina, et al.
Veröffentlicht: (2024)
von: Gajcin, Jasmina, et al.
Veröffentlicht: (2024)
Position: Theory of Mind Benchmarks are Broken for Large Language Models
von: Riemer, Matthew, et al.
Veröffentlicht: (2024)
von: Riemer, Matthew, et al.
Veröffentlicht: (2024)
Reasoning about concepts with LLMs: Inconsistencies abound
von: Uceda-Sosa, Rosario, et al.
Veröffentlicht: (2024)
von: Uceda-Sosa, Rosario, et al.
Veröffentlicht: (2024)
CoFrNets: Interpretable Neural Architecture Inspired by Continued Fractions
von: Puri, Isha, et al.
Veröffentlicht: (2025)
von: Puri, Isha, et al.
Veröffentlicht: (2025)
Targeted Advertising on Social Networks Using Online Variational Tensor Regression
von: Idé, Tsuyoshi, et al.
Veröffentlicht: (2022)
von: Idé, Tsuyoshi, et al.
Veröffentlicht: (2022)
Ähnliche Einträge
-
The Ultimate Test of Superintelligent AI Agents: Can an AI Balance Care and Control in Asymmetric Relationships?
von: Bouneffouf, Djallel, et al.
Veröffentlicht: (2025) -
Evaluating the Prompt Steerability of Large Language Models
von: Miehling, Erik, et al.
Veröffentlicht: (2024) -
Value Alignment from Unstructured Text
von: Padhi, Inkit, et al.
Veröffentlicht: (2024) -
Who Sees the Risk? Stakeholder Conflicts and Explanatory Policies in LLM-based Risk Assessment
von: Yadav, Srishti, et al.
Veröffentlicht: (2025) -
The Effectiveness of Approximate Regularized Replay for Efficient Supervised Fine-Tuning of Large Language Models
von: Riemer, Matthew, et al.
Veröffentlicht: (2025)