Great Models Think Alike and this Undermines AI Oversight
Fuente:
arXiv
Saved in:
| Main Authors: | Goel, Shashwat, Struber, Joschka, Auzina, Ilze Amanda, Chandra, Karuna K, Kumaraguru, Ponnurangam, Kiela, Douwe, Prabhu, Ameya, Bethge, Matthias, Geiping, Jonas |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Intrinsic Credit Assignment for Long Horizon Interaction
by: Auzina, Ilze Amanda, et al.
Published: (2026)
by: Auzina, Ilze Amanda, et al.
Published: (2026)
Can Language Models Falsify? Evaluating Algorithmic Reasoning with Counterexample Creation
by: Sinha, Shiven, et al.
Published: (2025)
by: Sinha, Shiven, et al.
Published: (2025)
Wu's Method can Boost Symbolic AI to Rival Silver Medalists and AlphaGeometry to Outperform Gold Medalists at IMO Geometry
by: Sinha, Shiven, et al.
Published: (2024)
by: Sinha, Shiven, et al.
Published: (2024)
Corrective Machine Unlearning
by: Goel, Shashwat, et al.
Published: (2024)
by: Goel, Shashwat, et al.
Published: (2024)
Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLMs
by: Panfilov, Alexander, et al.
Published: (2025)
by: Panfilov, Alexander, et al.
Published: (2025)
Answer Matching Outperforms Multiple Choice for Language Model Evaluation
by: Chandak, Nikhil, et al.
Published: (2025)
by: Chandak, Nikhil, et al.
Published: (2025)
Scaling Open-Ended Reasoning to Predict the Future
by: Chandak, Nikhil, et al.
Published: (2025)
by: Chandak, Nikhil, et al.
Published: (2025)
Ketto and the Science of Giving: A Data-Driven Investigation of Crowdfunding for India
by: Chandra, Karuna, et al.
Published: (2025)
by: Chandra, Karuna, et al.
Published: (2025)
Single-Agent LLMs Outperform Multi-Agent Systems on Multi-Hop Reasoning Under Equal Thinking Token Budgets
by: Tran, Dat, et al.
Published: (2026)
by: Tran, Dat, et al.
Published: (2026)
Enhancing AI Safety Through the Fusion of Low Rank Adapters
by: Gudipudi, Satya Swaroop, et al.
Published: (2024)
by: Gudipudi, Satya Swaroop, et al.
Published: (2024)
What if I ask in \textit{alia lingua}? Measuring Functional Similarity Across Languages
by: Mishra, Debangan, et al.
Published: (2025)
by: Mishra, Debangan, et al.
Published: (2025)
Round-Trip Translation Reveals What Frontier Multilingual Benchmarks Miss
by: Skorobogat, Ronald, et al.
Published: (2026)
by: Skorobogat, Ronald, et al.
Published: (2026)
FutureSim: Replaying World Events to Evaluate Adaptive Agents
by: Goel, Shashwat, et al.
Published: (2026)
by: Goel, Shashwat, et al.
Published: (2026)
Pitfalls in Evaluating Language Model Forecasters
by: Paleka, Daniel, et al.
Published: (2025)
by: Paleka, Daniel, et al.
Published: (2025)
Mapping Post-Training Forgetting in Language Models at Scale
by: Harmon, Jackson, et al.
Published: (2025)
by: Harmon, Jackson, et al.
Published: (2025)
Are We Done with Object-Centric Learning?
by: Rubinstein, Alexander, et al.
Published: (2025)
by: Rubinstein, Alexander, et al.
Published: (2025)
Random Representations Outperform Online Continually Learned Representations
by: Prabhu, Ameya, et al.
Published: (2024)
by: Prabhu, Ameya, et al.
Published: (2024)
The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs
by: Sinha, Akshit, et al.
Published: (2025)
by: Sinha, Akshit, et al.
Published: (2025)
A Cognac Shot To Forget Bad Memories: Corrective Unlearning for Graph Neural Networks
by: Kolipaka, Varshita, et al.
Published: (2024)
by: Kolipaka, Varshita, et al.
Published: (2024)
Rethinking Thinking Tokens: Understanding Why They Underperform in Practice
by: Vennam, Sreeram, et al.
Published: (2024)
by: Vennam, Sreeram, et al.
Published: (2024)
TAMAS: Benchmarking Adversarial Risks in Multi-Agent LLM Systems
by: Kavathekar, Ishan, et al.
Published: (2025)
by: Kavathekar, Ishan, et al.
Published: (2025)
I Can't Believe It's Corrupt: Evaluating Corruption in Multi-Agent Governance Systems
by: P, Vedanta S, et al.
Published: (2026)
by: P, Vedanta S, et al.
Published: (2026)
Pretraining Frequency Predicts Compositional Generalization of CLIP on Real-World Tasks
by: Wiedemer, Thaddäus, et al.
Published: (2025)
by: Wiedemer, Thaddäus, et al.
Published: (2025)
Adaptive Attacks on Trusted Monitors Subvert AI Control Protocols
by: Terekhov, Mikhail, et al.
Published: (2025)
by: Terekhov, Mikhail, et al.
Published: (2025)
Topo Goes Political: TDA-Based Controversy Detection in Imbalanced Reddit Political Data
by: Arun, Arvindh, et al.
Published: (2025)
by: Arun, Arvindh, et al.
Published: (2025)
Sample Complexity of Causal Identification with Temporal Heterogeneity
by: Rathod, Ameya, et al.
Published: (2026)
by: Rathod, Ameya, et al.
Published: (2026)
LLM generation novelty through the lens of semantic similarity
by: Davydov, Philipp, et al.
Published: (2025)
by: Davydov, Philipp, et al.
Published: (2025)
BlitzRank: Principled Zero-shot Ranking Agents with Tournament Graphs
by: Agrawal, Sheshansh, et al.
Published: (2026)
by: Agrawal, Sheshansh, et al.
Published: (2026)
Representation Surgery: Theory and Practice of Affine Steering
by: Singh, Shashwat, et al.
Published: (2024)
by: Singh, Shashwat, et al.
Published: (2024)
ONEBench to Test Them All: Sample-Level Benchmarking Over Open-Ended Capabilities
by: Ghosh, Adhiraj, et al.
Published: (2024)
by: Ghosh, Adhiraj, et al.
Published: (2024)
A Sober Look at Progress in Language Model Reasoning: Pitfalls and Paths to Reproducibility
by: Hochlehnert, Andreas, et al.
Published: (2025)
by: Hochlehnert, Andreas, et al.
Published: (2025)
Solving Spatial Supersensing Without Spatial Supersensing
by: Udandarao, Vishaal, et al.
Published: (2025)
by: Udandarao, Vishaal, et al.
Published: (2025)
Efficient Lifelong Model Evaluation in an Era of Rapid Progress
by: Prabhu, Ameya, et al.
Published: (2024)
by: Prabhu, Ameya, et al.
Published: (2024)
CiteME: Can Language Models Accurately Cite Scientific Claims?
by: Press, Ori, et al.
Published: (2024)
by: Press, Ori, et al.
Published: (2024)
Intrinsic Guardrails: How Semantic Geometry of Personality Interacts with Emergent Misalignment in LLMs
by: Aneja, Krishak, et al.
Published: (2026)
by: Aneja, Krishak, et al.
Published: (2026)
Anchor Points: Benchmarking Models with Much Fewer Examples
by: Vivek, Rajan, et al.
Published: (2023)
by: Vivek, Rajan, et al.
Published: (2023)
Document Optimization for Black-Box Retrieval via Reinforcement Learning
by: Uzan, Omri, et al.
Published: (2026)
by: Uzan, Omri, et al.
Published: (2026)
Personal Narratives Empower Politically Disinclined Individuals to Engage in Political Discussions
by: Chebrolu, Tejasvi, et al.
Published: (2025)
by: Chebrolu, Tejasvi, et al.
Published: (2025)
LLM Vocabulary Compression for Low-Compute Environments
by: Vennam, Sreeram, et al.
Published: (2024)
by: Vennam, Sreeram, et al.
Published: (2024)
VGGSounder: Audio-Visual Evaluations for Foundation Models
by: Zverev, Daniil, et al.
Published: (2025)
by: Zverev, Daniil, et al.
Published: (2025)
Similar Items
-
Intrinsic Credit Assignment for Long Horizon Interaction
by: Auzina, Ilze Amanda, et al.
Published: (2026) -
Can Language Models Falsify? Evaluating Algorithmic Reasoning with Counterexample Creation
by: Sinha, Shiven, et al.
Published: (2025) -
Wu's Method can Boost Symbolic AI to Rival Silver Medalists and AlphaGeometry to Outperform Gold Medalists at IMO Geometry
by: Sinha, Shiven, et al.
Published: (2024) -
Corrective Machine Unlearning
by: Goel, Shashwat, et al.
Published: (2024) -
Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLMs
by: Panfilov, Alexander, et al.
Published: (2025)