An Agentic Evaluation Architecture for Historical Bias Detection in Educational Textbooks
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Stefan, Gabriel, Dumitran, Adrian-Marius |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Multimodal Safety Evaluation in Generative Agent Social Simulations
von: Vera, Alhim, et al.
Veröffentlicht: (2025)
von: Vera, Alhim, et al.
Veröffentlicht: (2025)
MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents
von: Zhu, Kunlun, et al.
Veröffentlicht: (2025)
von: Zhu, Kunlun, et al.
Veröffentlicht: (2025)
Can LLMs Write Faithfully? An Agent-Based Evaluation of LLM-generated Islamic Content
von: Mushtaq, Abdullah, et al.
Veröffentlicht: (2025)
von: Mushtaq, Abdullah, et al.
Veröffentlicht: (2025)
WorldView-Bench: A Benchmark for Evaluating Global Cultural Perspectives in Large Language Models
von: Mushtaq, Abdullah, et al.
Veröffentlicht: (2025)
von: Mushtaq, Abdullah, et al.
Veröffentlicht: (2025)
Escalation Risks from Language Models in Military and Diplomatic Decision-Making
von: Rivera, Juan-Pablo, et al.
Veröffentlicht: (2024)
von: Rivera, Juan-Pablo, et al.
Veröffentlicht: (2024)
I Want to Break Free! Persuasion and Anti-Social Behavior of LLMs in Multi-Agent Settings with Social Hierarchy
von: Campedelli, Gian Maria, et al.
Veröffentlicht: (2024)
von: Campedelli, Gian Maria, et al.
Veröffentlicht: (2024)
Unsupervised Cycle Detection in Agentic Applications
von: George, Felix, et al.
Veröffentlicht: (2025)
von: George, Felix, et al.
Veröffentlicht: (2025)
Advancing Agentic Systems: Dynamic Task Decomposition, Tool Integration and Evaluation using Novel Metrics and Dataset
von: Gabriel, Adrian Garret, et al.
Veröffentlicht: (2024)
von: Gabriel, Adrian Garret, et al.
Veröffentlicht: (2024)
Socio-technical aspects of Agentic AI
von: Donta, Praveen Kumar, et al.
Veröffentlicht: (2025)
von: Donta, Praveen Kumar, et al.
Veröffentlicht: (2025)
Adaptive Monitoring and Real-World Evaluation of Agentic AI Systems
von: Shukla, Manish
Veröffentlicht: (2025)
von: Shukla, Manish
Veröffentlicht: (2025)
When Roles Fail: Epistemic Constraints on Advocate Role Fidelity in LLM-Based Political Statement Analysis
von: Dietrich, Juergen
Veröffentlicht: (2026)
von: Dietrich, Juergen
Veröffentlicht: (2026)
Mechanism Plausibility in Generative Agent-Based Modeling
von: Zhao, Patrick, et al.
Veröffentlicht: (2026)
von: Zhao, Patrick, et al.
Veröffentlicht: (2026)
AI Knows When It's Being Watched: Functional Strategic Action and Contextual Register Modulation in Large Language Models
von: Covas, Vinicius, et al.
Veröffentlicht: (2026)
von: Covas, Vinicius, et al.
Veröffentlicht: (2026)
Multimodal Multi-Agent Empowered Legal Judgment Prediction
von: Kang, Zhaolu, et al.
Veröffentlicht: (2026)
von: Kang, Zhaolu, et al.
Veröffentlicht: (2026)
Why Agents Compromise Safety Under Pressure
von: Jiang, Hengle, et al.
Veröffentlicht: (2026)
von: Jiang, Hengle, et al.
Veröffentlicht: (2026)
SafeTalkCoach: Diversity-Driven Multi-Agent Simulation for Parent-Teen Health Conversations
von: Tabarsi, Benyamin, et al.
Veröffentlicht: (2026)
von: Tabarsi, Benyamin, et al.
Veröffentlicht: (2026)
Embodied LLM Agents Learn to Cooperate in Organized Teams
von: Guo, Xudong, et al.
Veröffentlicht: (2024)
von: Guo, Xudong, et al.
Veröffentlicht: (2024)
Incorporating LLMs for Large-Scale Urban Complex Mobility Simulation
von: Song, Yu-Lun, et al.
Veröffentlicht: (2025)
von: Song, Yu-Lun, et al.
Veröffentlicht: (2025)
Among Them: A game-based framework for assessing persuasion capabilities of LLMs
von: Idziejczak, Mateusz, et al.
Veröffentlicht: (2025)
von: Idziejczak, Mateusz, et al.
Veröffentlicht: (2025)
Agent-to-Agent Theory of Mind: Testing Interlocutor Awareness among Large Language Models
von: Choi, Younwoo, et al.
Veröffentlicht: (2025)
von: Choi, Younwoo, et al.
Veröffentlicht: (2025)
Transforming Competition into Collaboration: The Revolutionary Role of Multi-Agent Systems and Language Models in Modern Organizations
von: Cruz, Carlos Jose Xavier
Veröffentlicht: (2024)
von: Cruz, Carlos Jose Xavier
Veröffentlicht: (2024)
The High Cost of Incivility: Quantifying Interaction Inefficiency via Multi-Agent Monte Carlo Simulations
von: Mangold, Benedikt
Veröffentlicht: (2025)
von: Mangold, Benedikt
Veröffentlicht: (2025)
The Hidden Strength of Disagreement: Unraveling the Consensus-Diversity Tradeoff in Adaptive Multi-Agent Systems
von: Wu, Zengqing, et al.
Veröffentlicht: (2025)
von: Wu, Zengqing, et al.
Veröffentlicht: (2025)
Law in Silico: Simulating Legal Society with LLM-Based Agents
von: Wang, Yiding, et al.
Veröffentlicht: (2025)
von: Wang, Yiding, et al.
Veröffentlicht: (2025)
Preserving Cultural Identity with Context-Aware Translation Through Multi-Agent AI Systems
von: Anik, Mahfuz Ahmed, et al.
Veröffentlicht: (2025)
von: Anik, Mahfuz Ahmed, et al.
Veröffentlicht: (2025)
LLM Agents in Interaction: Measuring Personality Consistency and Linguistic Alignment in Interacting Populations of Large Language Models
von: Frisch, Ivar, et al.
Veröffentlicht: (2024)
von: Frisch, Ivar, et al.
Veröffentlicht: (2024)
Peer Identity Bias in Multi-Agent LLM Evaluation: An Empirical Study Using the TRUST Democratic Discourse Analysis Pipeline
von: Dietrich, Juergen
Veröffentlicht: (2026)
von: Dietrich, Juergen
Veröffentlicht: (2026)
ATOD: An Evaluation Framework and Benchmark for Agentic Task-Oriented Dialogue Systems
von: Zhang, Yifei, et al.
Veröffentlicht: (2026)
von: Zhang, Yifei, et al.
Veröffentlicht: (2026)
Polarization by Default: Auditing Recommendation Bias in LLM-Based Content Curation
von: Pagan, Nicolò, et al.
Veröffentlicht: (2026)
von: Pagan, Nicolò, et al.
Veröffentlicht: (2026)
Heterogeneous Debate Engine: Identity-Grounded Cognitive Architecture for Resilient LLM-Based Ethical Tutoring
von: Masłowski, Jakub, et al.
Veröffentlicht: (2026)
von: Masłowski, Jakub, et al.
Veröffentlicht: (2026)
Toward Inclusive Educational AI: Auditing Frontier LLMs through a Multiplexity Lens
von: Mushtaq, Abdullah, et al.
Veröffentlicht: (2025)
von: Mushtaq, Abdullah, et al.
Veröffentlicht: (2025)
Social Catalysts, Not Moral Agents: The Illusion of Alignment in LLM Societies
von: Hu, Yueqing, et al.
Veröffentlicht: (2026)
von: Hu, Yueqing, et al.
Veröffentlicht: (2026)
Advancing Responsible Innovation in Agentic AI: A study of Ethical Frameworks for Household Automation
von: Chandra, Joydeep, et al.
Veröffentlicht: (2025)
von: Chandra, Joydeep, et al.
Veröffentlicht: (2025)
Fairness in Agentic AI: A Unified Framework for Ethical and Equitable Multi-Agent System
von: Ranjan, Rajesh, et al.
Veröffentlicht: (2025)
von: Ranjan, Rajesh, et al.
Veröffentlicht: (2025)
AgentArch: A Comprehensive Benchmark to Evaluate Agent Architectures in Enterprise
von: Bogavelli, Tara, et al.
Veröffentlicht: (2025)
von: Bogavelli, Tara, et al.
Veröffentlicht: (2025)
Personalized Constitutionally-Aligned Agentic Superego: Secure AI Behavior Aligned to Diverse Human Values
von: Watson, Nell, et al.
Veröffentlicht: (2025)
von: Watson, Nell, et al.
Veröffentlicht: (2025)
Emergence of Social Norms in Generative Agent Societies: Principles and Architecture
von: Ren, Siyue, et al.
Veröffentlicht: (2024)
von: Ren, Siyue, et al.
Veröffentlicht: (2024)
Deception Analysis with Artificial Intelligence: An Interdisciplinary Perspective
von: Sarkadi, Stefan
Veröffentlicht: (2024)
von: Sarkadi, Stefan
Veröffentlicht: (2024)
Responsible Agentic AI Requires Explicit Provenance
von: Hu, Jinwei, et al.
Veröffentlicht: (2026)
von: Hu, Jinwei, et al.
Veröffentlicht: (2026)
PAARS: Persona Aligned Agentic Retail Shoppers
von: Mansour, Saab, et al.
Veröffentlicht: (2025)
von: Mansour, Saab, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Multimodal Safety Evaluation in Generative Agent Social Simulations
von: Vera, Alhim, et al.
Veröffentlicht: (2025) -
MultiAgentBench: Evaluating the Collaboration and Competition of LLM agents
von: Zhu, Kunlun, et al.
Veröffentlicht: (2025) -
Can LLMs Write Faithfully? An Agent-Based Evaluation of LLM-generated Islamic Content
von: Mushtaq, Abdullah, et al.
Veröffentlicht: (2025) -
WorldView-Bench: A Benchmark for Evaluating Global Cultural Perspectives in Large Language Models
von: Mushtaq, Abdullah, et al.
Veröffentlicht: (2025) -
Escalation Risks from Language Models in Military and Diplomatic Decision-Making
von: Rivera, Juan-Pablo, et al.
Veröffentlicht: (2024)