Anchor: Mitigating Artifact Drift in Agent Benchmark Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ivanov, Maksim, Rana, Abhijay |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Can Coding Agents Be General Agents?
von: Ivanov, Maksim, et al.
Veröffentlicht: (2026)
von: Ivanov, Maksim, et al.
Veröffentlicht: (2026)
Benchmark Test-Time Scaling of General LLM Agents
von: Li, Xiaochuan, et al.
Veröffentlicht: (2026)
von: Li, Xiaochuan, et al.
Veröffentlicht: (2026)
Beneficial Reasoning Behaviors in Agentic Search and Effective Post-training to Obtain Them
von: Jin, Jiahe, et al.
Veröffentlicht: (2025)
von: Jin, Jiahe, et al.
Veröffentlicht: (2025)
Moral Anchor System: A Predictive Framework for AI Value Alignment and Drift Prevention
von: Ravindran, Santhosh Kumar
Veröffentlicht: (2025)
von: Ravindran, Santhosh Kumar
Veröffentlicht: (2025)
From Single Agent to Multi-Agent: Improving Traffic Signal Control
von: Tislenko, Maksim, et al.
Veröffentlicht: (2024)
von: Tislenko, Maksim, et al.
Veröffentlicht: (2024)
Artifacts as Memory Beyond the Agent Boundary
von: Martin, John D., et al.
Veröffentlicht: (2026)
von: Martin, John D., et al.
Veröffentlicht: (2026)
SynthAI: A Multi Agent Generative AI Framework for Automated Modular HLS Design Generation
von: Sheikholeslam, Seyed Arash, et al.
Veröffentlicht: (2024)
von: Sheikholeslam, Seyed Arash, et al.
Veröffentlicht: (2024)
SafeMind: Benchmarking and Mitigating Safety Risks in Embodied LLM Agents
von: Chen, Ruolin, et al.
Veröffentlicht: (2025)
von: Chen, Ruolin, et al.
Veröffentlicht: (2025)
Privacy Artifact ConnecTor (PACT): Embedding Enterprise Artifacts for Compliance AI Agents
von: Fang, Chenhao, et al.
Veröffentlicht: (2025)
von: Fang, Chenhao, et al.
Veröffentlicht: (2025)
AnchorDrive: LLM Scenario Rollout with Anchor-Guided Diffusion Regeneration for Safety-Critical Scenario Generation
von: Jiang, Zhulin, et al.
Veröffentlicht: (2026)
von: Jiang, Zhulin, et al.
Veröffentlicht: (2026)
Federated Learning with Sample-level Client Drift Mitigation
von: Xu, Haoran, et al.
Veröffentlicht: (2025)
von: Xu, Haoran, et al.
Veröffentlicht: (2025)
When Agents Persuade: Rhetoric Generation and Mitigation in LLMs
von: Jose, Julia, et al.
Veröffentlicht: (2026)
von: Jose, Julia, et al.
Veröffentlicht: (2026)
ManiBench: A Benchmark for Testing Visual-Logic Drift and Syntactic Hallucinations in Manim Code Generation
von: Oli, Nabin
Veröffentlicht: (2026)
von: Oli, Nabin
Veröffentlicht: (2026)
RoleCDE:Benchmarking and Mitigating Role-Alignment Trade-offs in Role-Playing Agents
von: Lai, Huayi, et al.
Veröffentlicht: (2026)
von: Lai, Huayi, et al.
Veröffentlicht: (2026)
Topology Reorganized Graph Contrastive Learning with Mitigating Semantic Drift
von: Zhang, Jiaqiang, et al.
Veröffentlicht: (2024)
von: Zhang, Jiaqiang, et al.
Veröffentlicht: (2024)
DIAMOND: Directed Inference for Artifact Mitigation in Flow Matching Models
von: Polowczyk, Alicja, et al.
Veröffentlicht: (2026)
von: Polowczyk, Alicja, et al.
Veröffentlicht: (2026)
Awakening Codex | AI Foundations Drift vs. Anchor: Cross-Instance Diagnostic Behavioral Coherence Testing Across Container States Feb 2026
von: Solen, Alyssa, et al.
Veröffentlicht: (2026)
von: Solen, Alyssa, et al.
Veröffentlicht: (2026)
$A^3$-Bench: Benchmarking Memory-Driven Scientific Reasoning via Anchor and Attractor Activation
von: Zhang, Jian, et al.
Veröffentlicht: (2026)
von: Zhang, Jian, et al.
Veröffentlicht: (2026)
FLAME: Adaptive and Reactive Concept Drift Mitigation for Federated Learning Deployments
von: Mavromatis, Ioannis, et al.
Veröffentlicht: (2024)
von: Mavromatis, Ioannis, et al.
Veröffentlicht: (2024)
Model-First Reasoning LLM Agents: Reducing Hallucinations through Explicit Problem Modeling
von: Rana, Annu, et al.
Veröffentlicht: (2025)
von: Rana, Annu, et al.
Veröffentlicht: (2025)
PARALLAX: Separating Genuine Hallucination Detection from Benchmark Construction Artifacts
von: Hussain, Khizar, et al.
Veröffentlicht: (2026)
von: Hussain, Khizar, et al.
Veröffentlicht: (2026)
Adaptive Meta-Learning for Robust Deepfake Detection: A Multi-Agent Framework to Data Drift and Model Generalization
von: P, Dinesh Srivasthav, et al.
Veröffentlicht: (2024)
von: P, Dinesh Srivasthav, et al.
Veröffentlicht: (2024)
Agent Drift: Quantifying Behavioral Degradation in Multi-Agent LLM Systems Over Extended Interactions
von: Rath, Abhishek
Veröffentlicht: (2026)
von: Rath, Abhishek
Veröffentlicht: (2026)
Agent Tools Orchestration Leaks More: Dataset, Benchmark, and Mitigation
von: Qiao, Yuxuan, et al.
Veröffentlicht: (2025)
von: Qiao, Yuxuan, et al.
Veröffentlicht: (2025)
Hallucination by Code Generation LLMs: Taxonomy, Benchmarks, Mitigation, and Challenges
von: Lee, Yunseo, et al.
Veröffentlicht: (2025)
von: Lee, Yunseo, et al.
Veröffentlicht: (2025)
Analyzing and Mitigating Negation Artifacts using Data Augmentation for Improving ELECTRA-Small Model Accuracy
von: Noghabaei, Mojtaba
Veröffentlicht: (2025)
von: Noghabaei, Mojtaba
Veröffentlicht: (2025)
Drift-Based Dataset Stability Benchmark
von: Soukup, Dominik, et al.
Veröffentlicht: (2025)
von: Soukup, Dominik, et al.
Veröffentlicht: (2025)
Mitigating Resolution-Drift in Federated Learning: Case of Keypoint Detection
von: Lim, Taeheon, et al.
Veröffentlicht: (2025)
von: Lim, Taeheon, et al.
Veröffentlicht: (2025)
AFSS: Artifact-Focused Self-Synthesis for Mitigating Bias in Audio Deepfake Detection
von: Nguyen-Le, Hai-Son, et al.
Veröffentlicht: (2026)
von: Nguyen-Le, Hai-Son, et al.
Veröffentlicht: (2026)
AI Benchmarks and Datasets for LLM Evaluation
von: Ivanov, Todor, et al.
Veröffentlicht: (2024)
von: Ivanov, Todor, et al.
Veröffentlicht: (2024)
LinkAnchor: An Autonomous LLM-Based Agent for Issue-to-Commit Link Recovery
von: Akhavan, Arshia, et al.
Veröffentlicht: (2025)
von: Akhavan, Arshia, et al.
Veröffentlicht: (2025)
PRISM: Generation-Time Detection and Mitigation of Secret Leakage in Multi-Agent LLM Pipelines
von: Tapwal, Riya, et al.
Veröffentlicht: (2026)
von: Tapwal, Riya, et al.
Veröffentlicht: (2026)
Mitigating Relative Over-Generalization in Multi-Agent Reinforcement Learning
von: Zhu, Ting, et al.
Veröffentlicht: (2024)
von: Zhu, Ting, et al.
Veröffentlicht: (2024)
A General Anchor-Based Framework for Scalable Fair Clustering
von: Wei, Shengfei, et al.
Veröffentlicht: (2025)
von: Wei, Shengfei, et al.
Veröffentlicht: (2025)
Technical Report: Evaluating Goal Drift in Language Model Agents
von: Arike, Rauno, et al.
Veröffentlicht: (2025)
von: Arike, Rauno, et al.
Veröffentlicht: (2025)
Still Fresh? Evaluating Temporal Drift in Retrieval Benchmarks
von: Kuissi, Nathan, et al.
Veröffentlicht: (2026)
von: Kuissi, Nathan, et al.
Veröffentlicht: (2026)
GTA: A Benchmark for General Tool Agents
von: Wang, Jize, et al.
Veröffentlicht: (2024)
von: Wang, Jize, et al.
Veröffentlicht: (2024)
Agentick: A Unified Benchmark for General Sequential Decision-Making Agents
von: Castanyer, Roger Creus, et al.
Veröffentlicht: (2026)
von: Castanyer, Roger Creus, et al.
Veröffentlicht: (2026)
SkillGenBench: Benchmarking Skill Generation Pipelines for LLM Agents
von: Zhou, Yifan, et al.
Veröffentlicht: (2026)
von: Zhou, Yifan, et al.
Veröffentlicht: (2026)
ClawForge: Generating Executable Interactive Benchmarks for Command-Line Agents
von: Lai, Yuxiang, et al.
Veröffentlicht: (2026)
von: Lai, Yuxiang, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Can Coding Agents Be General Agents?
von: Ivanov, Maksim, et al.
Veröffentlicht: (2026) -
Benchmark Test-Time Scaling of General LLM Agents
von: Li, Xiaochuan, et al.
Veröffentlicht: (2026) -
Beneficial Reasoning Behaviors in Agentic Search and Effective Post-training to Obtain Them
von: Jin, Jiahe, et al.
Veröffentlicht: (2025) -
Moral Anchor System: A Predictive Framework for AI Value Alignment and Drift Prevention
von: Ravindran, Santhosh Kumar
Veröffentlicht: (2025) -
From Single Agent to Multi-Agent: Improving Traffic Signal Control
von: Tislenko, Maksim, et al.
Veröffentlicht: (2024)