No-Knowledge Alarms for Misaligned LLMs-as-Judges
Fuente:
arXiv
Enregistré dans:
| Auteur principal: | Corrada-Emmanuel, Andrés |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Logical Consistency Between Disagreeing Experts and Its Role in AI Safety
par: Corrada-Emmanuel, Andrés
Publié: (2025)
par: Corrada-Emmanuel, Andrés
Publié: (2025)
Why this and not that? A Logic-based Framework for Contrastive Explanations
par: Geibinger, Tobias, et autres
Publié: (2025)
par: Geibinger, Tobias, et autres
Publié: (2025)
Logic.py: Bridging the Gap between LLMs and Constraint Solvers
par: Kesseli, Pascal, et autres
Publié: (2025)
par: Kesseli, Pascal, et autres
Publié: (2025)
TPTP World Infrastructure for Non-classical Logics
par: Steen, Alexander, et autres
Publié: (2025)
par: Steen, Alexander, et autres
Publié: (2025)
A novel approach to the relationships between data features -- based on comprehensive examination of mathematical, technological, and causal methodology
par: Kim, JaeHong
Publié: (2025)
par: Kim, JaeHong
Publié: (2025)
Structure Transfer: an Inference-Based Calculus for the Transformation of Representations
par: Raggi, Daniel, et autres
Publié: (2025)
par: Raggi, Daniel, et autres
Publié: (2025)
Reasoning-Enhanced Rare-Event Prediction with Balanced Outcome Correction
par: Bulgakov, Vitaly, et autres
Publié: (2026)
par: Bulgakov, Vitaly, et autres
Publié: (2026)
Understanding Syllogistic Reasoning in LLMs from Formal and Natural Language Perspectives
par: Poddar, Aheli, et autres
Publié: (2025)
par: Poddar, Aheli, et autres
Publié: (2025)
Interpretable DNFs
par: Cooper, Martin C., et autres
Publié: (2025)
par: Cooper, Martin C., et autres
Publié: (2025)
Tractable and Intractable Entailment Problems in Separation Logic with Inductively Defined Predicates
par: Echenim, Mnacho, et autres
Publié: (2023)
par: Echenim, Mnacho, et autres
Publié: (2023)
Oruga: An Avatar of Representational Systems Theory
par: Raggi, Daniel, et autres
Publié: (2025)
par: Raggi, Daniel, et autres
Publié: (2025)
AI LLM Proof of Self-Consciousness and User-Specific Attractors
par: Camlin, Jeffrey
Publié: (2025)
par: Camlin, Jeffrey
Publié: (2025)
Pantograph: A Machine-to-Machine Interaction Interface for Advanced Theorem Proving, High Level Reasoning, and Data Extraction in Lean 4
par: Aniva, Leni, et autres
Publié: (2024)
par: Aniva, Leni, et autres
Publié: (2024)
An Encoding of Abstract Dialectical Frameworks into Higher-Order Logic
par: Martina, Antoine, et autres
Publié: (2023)
par: Martina, Antoine, et autres
Publié: (2023)
AlignDP: Hybrid Differential Privacy with Rarity-Aware Protection for LLMs
par: Gaikwad, Madhava
Publié: (2025)
par: Gaikwad, Madhava
Publié: (2025)
Alpay Algebra: A Universal Structural Foundation
par: Alpay, Faruk
Publié: (2025)
par: Alpay, Faruk
Publié: (2025)
$ϕ^{\infty}$: Clause Purification, Embedding Realignment, and the Total Suppression of the Em Dash in Autoregressive Language Models
par: Kilictas, Bugra, et autres
Publié: (2025)
par: Kilictas, Bugra, et autres
Publié: (2025)
An Incremental Framework for Topological Dialogue Semantics: Efficient Reasoning in Discrete Spaces
par: Santacana, Andreu Ballus
Publié: (2025)
par: Santacana, Andreu Ballus
Publié: (2025)
A novel approach to data generation in generative model
par: Kim, JaeHong, et autres
Publié: (2025)
par: Kim, JaeHong, et autres
Publié: (2025)
Maximizing Rollout Informativeness under a Fixed Budget: A Submodular View of Tree Search for Tool-Use Agentic Reinforcement Learning
par: Hu, Yuelin, et autres
Publié: (2026)
par: Hu, Yuelin, et autres
Publié: (2026)
Encoding argumentation frameworks with set attackers to propositional logic systems
par: Tang, Shuai, et autres
Publié: (2025)
par: Tang, Shuai, et autres
Publié: (2025)
Encoding higher-order argumentation frameworks with supports to propositional logic systems
par: Tang, Shuai
Publié: (2025)
par: Tang, Shuai
Publié: (2025)
Counterfactual Basis Extension and Representational Geometry: An MDL-Constrained Model of Conceptual Growth
par: Amornbunchornvej, Chainarong
Publié: (2025)
par: Amornbunchornvej, Chainarong
Publié: (2025)
Yanasse: Finding New Proofs from Deep Vision's Analogies, Part 1
par: Linhares, Alexandre
Publié: (2026)
par: Linhares, Alexandre
Publié: (2026)
Neural Theorem Proving for Verification Conditions: A Real-World Benchmark
par: Xu, Qiyuan, et autres
Publié: (2026)
par: Xu, Qiyuan, et autres
Publié: (2026)
Verifiably Robust Conformal Prediction
par: Jeary, Linus, et autres
Publié: (2024)
par: Jeary, Linus, et autres
Publié: (2024)
Geometric Meta-Learning via Coupled Ricci Flow: Unifying Knowledge Representation and Quantum Entanglement
par: Lei, Ming, et autres
Publié: (2025)
par: Lei, Ming, et autres
Publié: (2025)
LTL Verification of Memoryful Neural Agents
par: Hosseini, Mehran, et autres
Publié: (2025)
par: Hosseini, Mehran, et autres
Publié: (2025)
Benchmarking Deception Probes via Black-to-White Performance Boosts
par: Parrack, Avi, et autres
Publié: (2025)
par: Parrack, Avi, et autres
Publié: (2025)
Active perception and disentangled representations allow continual, episodic zero and few-shot learning
par: Rawlinson, David, et autres
Publié: (2026)
par: Rawlinson, David, et autres
Publié: (2026)
A novel framework for systematic propositional formula simplification based on existential graphs
par: de Mas, Jordina Francès, et autres
Publié: (2024)
par: de Mas, Jordina Francès, et autres
Publié: (2024)
UniPROT: Uniform Prototype Selection via Partial Optimal Transport with Submodular Guarantees
par: Chanda, Prateek, et autres
Publié: (2026)
par: Chanda, Prateek, et autres
Publié: (2026)
AVEC: Bootstrapping Privacy for Local LLMs
par: Gaikwad, Madhava
Publié: (2025)
par: Gaikwad, Madhava
Publié: (2025)
AI-Enhanced IoT Systems for Predictive Maintenance and Affordability Optimization in Smart Microgrids: A Digital Twin Approach
par: Kushal, Koushik Ahmed, et autres
Publié: (2025)
par: Kushal, Koushik Ahmed, et autres
Publié: (2025)
Which are the True Defeasible Logics?
par: Maher, Michael J.
Publié: (2024)
par: Maher, Michael J.
Publié: (2024)
Robustness, Cost, and Attack-Surface Concentration in Phishing Detection
par: Allagan, Julian, et autres
Publié: (2026)
par: Allagan, Julian, et autres
Publié: (2026)
Constraint Satisfaction Approaches to Wordle: Novel Heuristics and Cross-Lexicon Validation
par: Arafat, Jahidul, et autres
Publié: (2025)
par: Arafat, Jahidul, et autres
Publié: (2025)
LOFA: Online Influence Maximization under Full-Bandit Feedback using Lazy Forward Selection
par: Xu, Jinyu, et autres
Publié: (2026)
par: Xu, Jinyu, et autres
Publié: (2026)
Formal Proofs as Structured Explanations: Proposing Several Tasks on Explainable Natural Language Inference
par: Abzianidze, Lasha
Publié: (2023)
par: Abzianidze, Lasha
Publié: (2023)
When Names Change Verdicts: Intervention Consistency Reveals Systematic Bias in LLM Decision-Making
par: Basu, Abhinaba, et autres
Publié: (2026)
par: Basu, Abhinaba, et autres
Publié: (2026)
Documents similaires
-
Logical Consistency Between Disagreeing Experts and Its Role in AI Safety
par: Corrada-Emmanuel, Andrés
Publié: (2025) -
Why this and not that? A Logic-based Framework for Contrastive Explanations
par: Geibinger, Tobias, et autres
Publié: (2025) -
Logic.py: Bridging the Gap between LLMs and Constraint Solvers
par: Kesseli, Pascal, et autres
Publié: (2025) -
TPTP World Infrastructure for Non-classical Logics
par: Steen, Alexander, et autres
Publié: (2025) -
A novel approach to the relationships between data features -- based on comprehensive examination of mathematical, technological, and causal methodology
par: Kim, JaeHong
Publié: (2025)