ArGen: Auto-Regulation of Generative AI via GRPO and Policy-as-Code
Fuente:
arXiv
Gespeichert in:
| 1. Verfasser: | Madan, Kapil |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
NOTAI.AI: Explainable Detection of Machine-Generated Text via Curvature and Feature Attribution
von: Breneur, Oleksandr Marchenko, et al.
Veröffentlicht: (2026)
von: Breneur, Oleksandr Marchenko, et al.
Veröffentlicht: (2026)
Rethinking the Multilingual Reasoning Gap with Layer Swap
von: Lasbordes, Maxence, et al.
Veröffentlicht: (2026)
von: Lasbordes, Maxence, et al.
Veröffentlicht: (2026)
The AI Fiction Paradox
von: Elkins, Katherine
Veröffentlicht: (2026)
von: Elkins, Katherine
Veröffentlicht: (2026)
Who's Asking? Investigating Bias Through the Lens of Disability Framed Queries in LLMs
von: Hari, Vishnu, et al.
Veröffentlicht: (2025)
von: Hari, Vishnu, et al.
Veröffentlicht: (2025)
$δ$-STEAL: LLM Stealing Attack with Local Differential Privacy
von: Dang, Kieu, et al.
Veröffentlicht: (2025)
von: Dang, Kieu, et al.
Veröffentlicht: (2025)
CLMN: Concept based Language Models via Neural Symbolic Reasoning
von: Yang, Yibo
Veröffentlicht: (2025)
von: Yang, Yibo
Veröffentlicht: (2025)
AIPsy-Affect: A Keyword-Free Clinical Stimulus Battery for Mechanistic Interpretability of Emotion in Language Models
von: Keeman, Michael
Veröffentlicht: (2026)
von: Keeman, Michael
Veröffentlicht: (2026)
Product-of-Experts Training Reduces Dataset Artifacts in Natural Language Inference
von: Mathew, Aby Mammen
Veröffentlicht: (2026)
von: Mathew, Aby Mammen
Veröffentlicht: (2026)
When Does Content-Based Routing Work? Representation Requirements for Selective Attention in Hybrid Sequence Models
von: Basu, Abhinaba
Veröffentlicht: (2026)
von: Basu, Abhinaba
Veröffentlicht: (2026)
Automated but Atrophied? Student Over-Reliance vs Expert Augmentation of AI in Learning and Cybersecurity
von: Khan, Koffka
Veröffentlicht: (2025)
von: Khan, Koffka
Veröffentlicht: (2025)
Mechanistic Analysis of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning
von: Imanov, Olaf Yunus Laitinen
Veröffentlicht: (2026)
von: Imanov, Olaf Yunus Laitinen
Veröffentlicht: (2026)
Latent Object Permanence: Topological Phase Transitions, Free-Energy Principles, and Renormalization Group Flows in Deep Transformer Manifolds
von: Alpay, Faruk, et al.
Veröffentlicht: (2026)
von: Alpay, Faruk, et al.
Veröffentlicht: (2026)
Inference acceleration for large language models using "stairs" assisted greedy generation
von: Grigaliūnas, Domas, et al.
Veröffentlicht: (2024)
von: Grigaliūnas, Domas, et al.
Veröffentlicht: (2024)
Strategic Doctrine Language Models (sdLM): A Learning-System Framework for Doctrinal Consistency and Geopolitical Forecasting
von: Imanov, Olaf Yunus Laitinen, et al.
Veröffentlicht: (2026)
von: Imanov, Olaf Yunus Laitinen, et al.
Veröffentlicht: (2026)
Automated Feedback Generation for Undergraduate Mathematics: Development and Evaluation of an AI Teaching Assistant
von: Gohr, Aron, et al.
Veröffentlicht: (2026)
von: Gohr, Aron, et al.
Veröffentlicht: (2026)
Breaking to Build: A Threat Model of Prompt-Based Attacks for Securing LLMs
von: Hill, Brennen, et al.
Veröffentlicht: (2025)
von: Hill, Brennen, et al.
Veröffentlicht: (2025)
Harnessing non-adversarial robustness in large language models
von: Zhou, Qinghua, et al.
Veröffentlicht: (2026)
von: Zhou, Qinghua, et al.
Veröffentlicht: (2026)
Navigational Thinking as an Emerging Paradigm of Computer Science in the Age of Generative AI
von: Levin, Ilya
Veröffentlicht: (2026)
von: Levin, Ilya
Veröffentlicht: (2026)
Embedding Explainable AI in NHS Clinical Safety: The Explainability-Enabled Clinical Safety Framework (ECSF)
von: Gigiu, Robert
Veröffentlicht: (2025)
von: Gigiu, Robert
Veröffentlicht: (2025)
Unpacking Hateful Memes: Presupposed Context and False Claims
von: Cai, Weibin, et al.
Veröffentlicht: (2025)
von: Cai, Weibin, et al.
Veröffentlicht: (2025)
XAutoLM: Efficient Fine-Tuning of Language Models via Meta-Learning and AutoML
von: Estevanell-Valladares, Ernesto L., et al.
Veröffentlicht: (2025)
von: Estevanell-Valladares, Ernesto L., et al.
Veröffentlicht: (2025)
Vis-CoT: A Human-in-the-Loop Framework for Interactive Visualization and Intervention in LLM Chain-of-Thought Reasoning
von: Pather, Kaviraj, et al.
Veröffentlicht: (2025)
von: Pather, Kaviraj, et al.
Veröffentlicht: (2025)
BreakFun: Jailbreaking LLMs via Schema Exploitation
von: Oskooei, Amirkia Rafiei, et al.
Veröffentlicht: (2025)
von: Oskooei, Amirkia Rafiei, et al.
Veröffentlicht: (2025)
Do Reasoning Models Enhance Embedding Models?
von: Chan, Wun Yu, et al.
Veröffentlicht: (2026)
von: Chan, Wun Yu, et al.
Veröffentlicht: (2026)
ProbeScale: Probing Analysis to Optimize Neural Scaling Laws for Efficient Small Language Model Inference
von: Das, Sourav
Veröffentlicht: (2026)
von: Das, Sourav
Veröffentlicht: (2026)
Unsolvability Ceiling in Multi-LLM Routing: An Empirical Study of Evaluation Artifacts
von: Garg, Saloni, et al.
Veröffentlicht: (2026)
von: Garg, Saloni, et al.
Veröffentlicht: (2026)
Research on a hybrid LSTM-CNN-Attention model for text-based web content classification
von: Kuz, Mykola, et al.
Veröffentlicht: (2025)
von: Kuz, Mykola, et al.
Veröffentlicht: (2025)
Linguistic Collapse: Neural Collapse in (Large) Language Models
von: Wu, Robert, et al.
Veröffentlicht: (2024)
von: Wu, Robert, et al.
Veröffentlicht: (2024)
When Names Change Verdicts: Intervention Consistency Reveals Systematic Bias in LLM Decision-Making
von: Basu, Abhinaba, et al.
Veröffentlicht: (2026)
von: Basu, Abhinaba, et al.
Veröffentlicht: (2026)
How much do LLMs learn from negative examples?
von: Hamdan, Shadi, et al.
Veröffentlicht: (2025)
von: Hamdan, Shadi, et al.
Veröffentlicht: (2025)
How Pruning Reshapes Features: Sparse Autoencoder Analysis of Weight-Pruned Language Models
von: Borobia, Hector, et al.
Veröffentlicht: (2026)
von: Borobia, Hector, et al.
Veröffentlicht: (2026)
The Concept Allocation Zone: Tracking How Concepts Form Across Transformer Depth
von: Henry, James
Veröffentlicht: (2026)
von: Henry, James
Veröffentlicht: (2026)
Seeing Hate Differently: Hate Subspace Modeling for Culture-Aware Hate Speech Detection
von: Cai, Weibin, et al.
Veröffentlicht: (2025)
von: Cai, Weibin, et al.
Veröffentlicht: (2025)
mHC-SSM: Manifold-Constrained Hyper-Connections for State Space Language Models with Stream-Specialized Adapters
von: Mutlu, Abdulvahap, et al.
Veröffentlicht: (2026)
von: Mutlu, Abdulvahap, et al.
Veröffentlicht: (2026)
Extracting Sentence Embeddings from Pretrained Transformer Models
von: Stankevičius, Lukas, et al.
Veröffentlicht: (2024)
von: Stankevičius, Lukas, et al.
Veröffentlicht: (2024)
Sentiment Analysis of Lithuanian Online Reviews Using Large Language Models
von: Vileikytė, Brigita, et al.
Veröffentlicht: (2024)
von: Vileikytė, Brigita, et al.
Veröffentlicht: (2024)
ReFactor GNNs: Revisiting Factorisation-based Models from a Message-Passing Perspective
von: Chen, Yihong, et al.
Veröffentlicht: (2022)
von: Chen, Yihong, et al.
Veröffentlicht: (2022)
IFMTBench: A Comprehensive Benchmark for Multilingual Translation Instruction Following
von: Sun, Mingrui, et al.
Veröffentlicht: (2026)
von: Sun, Mingrui, et al.
Veröffentlicht: (2026)
Fine-tuning of Large Language Models for Constituency Parsing Using a Sequence to Sequence Approach
von: Delgado, Francisco Jose Cortes, et al.
Veröffentlicht: (2025)
von: Delgado, Francisco Jose Cortes, et al.
Veröffentlicht: (2025)
Doğal Dil İşlemede Tokenizasyon Standartları ve Ölçümü: Türkçe Üzerinden Büyük Dil Modellerinin Karşılaştırmalı Analizi
von: Bayram, M. Ali, et al.
Veröffentlicht: (2025)
von: Bayram, M. Ali, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
NOTAI.AI: Explainable Detection of Machine-Generated Text via Curvature and Feature Attribution
von: Breneur, Oleksandr Marchenko, et al.
Veröffentlicht: (2026) -
Rethinking the Multilingual Reasoning Gap with Layer Swap
von: Lasbordes, Maxence, et al.
Veröffentlicht: (2026) -
The AI Fiction Paradox
von: Elkins, Katherine
Veröffentlicht: (2026) -
Who's Asking? Investigating Bias Through the Lens of Disability Framed Queries in LLMs
von: Hari, Vishnu, et al.
Veröffentlicht: (2025) -
$δ$-STEAL: LLM Stealing Attack with Local Differential Privacy
von: Dang, Kieu, et al.
Veröffentlicht: (2025)