Saved in:
| Main Authors: | Ryd, Emil, Bartsch, Henning, Stastny, Julian, Benton, Joe, Hebbar, Vivek |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2604.22082 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Sandbagging in a Simple Survival Bandit Problem
by: Dyer, Joel, et al.
Published: (2025)
by: Dyer, Joel, et al.
Published: (2025)
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases
by: Gan, Eric, et al.
Published: (2026)
by: Gan, Eric, et al.
Published: (2026)
Efficiently Aligning Language Models with Online Natural Language Feedback
by: Ye, Christine, et al.
Published: (2026)
by: Ye, Christine, et al.
Published: (2026)
Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
by: Sheshadri, Abhay, et al.
Published: (2024)
by: Sheshadri, Abhay, et al.
Published: (2024)
Open Character Training: Shaping the Persona of AI Assistants through Constitutional AI
by: Maiya, Sharan, et al.
Published: (2025)
by: Maiya, Sharan, et al.
Published: (2025)
AI Sandbagging: Language Models can Strategically Underperform on Evaluations
by: van der Weij, Teun, et al.
Published: (2024)
by: van der Weij, Teun, et al.
Published: (2024)
When Can LLMs Learn to Reason with Weak Supervision?
by: Rahman, Salman, et al.
Published: (2026)
by: Rahman, Salman, et al.
Published: (2026)
Fine Flood Forecasts: Incorporating local data into global models through fine-tuning
by: Ryd, Emil, et al.
Published: (2025)
by: Ryd, Emil, et al.
Published: (2025)
Technical Report: Evaluating Goal Drift in Language Model Agents
by: Arike, Rauno, et al.
Published: (2025)
by: Arike, Rauno, et al.
Published: (2025)
Truthfulness Despite Weak Supervision: Evaluating and Training LLMs Using Peer Prediction
by: Qiu, Tianyi Alex, et al.
Published: (2026)
by: Qiu, Tianyi Alex, et al.
Published: (2026)
Towards eliciting latent knowledge from LLMs with mechanistic interpretability
by: Cywiński, Bartosz, et al.
Published: (2025)
by: Cywiński, Bartosz, et al.
Published: (2025)
Weakly Supervised Learning on Large Graphs
by: Prakash, Aditya
Published: (2025)
by: Prakash, Aditya
Published: (2025)
Learning and Generating Diverse Residential Load Patterns Using GAN with Weakly-Supervised Training and Weight Selection
by: Liang, Xinyu, et al.
Published: (2025)
by: Liang, Xinyu, et al.
Published: (2025)
Neuro-symbolic Weak Supervision: Theory and Semantics
by: Upreti, Nijesh, et al.
Published: (2025)
by: Upreti, Nijesh, et al.
Published: (2025)
AdapterSwap: Continuous Training of LLMs with Data Removal and Access-Control Guarantees
by: Fleshman, William, et al.
Published: (2024)
by: Fleshman, William, et al.
Published: (2024)
Polysemanticity and Capacity in Neural Networks
by: Scherlis, Adam, et al.
Published: (2022)
by: Scherlis, Adam, et al.
Published: (2022)
A General Framework for Learning from Weak Supervision
by: Chen, Hao, et al.
Published: (2024)
by: Chen, Hao, et al.
Published: (2024)
Revisiting Training Scale: An Empirical Study of Token Count, Power Consumption, and Parameter Efficiency
by: Dwyer, Joe
Published: (2026)
by: Dwyer, Joe
Published: (2026)
Rethinking Weak Supervision in Anomaly Detection: A Comprehensive Benchmark
by: Yao, Xu, et al.
Published: (2026)
by: Yao, Xu, et al.
Published: (2026)
Auditing Games for Sandbagging
by: Taylor, Jordan, et al.
Published: (2025)
by: Taylor, Jordan, et al.
Published: (2025)
Weak Supervision Performance Evaluation via Partial Identification
by: Polo, Felipe Maia, et al.
Published: (2023)
by: Polo, Felipe Maia, et al.
Published: (2023)
Learning Stable Predictors from Weak Supervision under Distribution Shift
by: Shoeibi, Mehrdad, et al.
Published: (2026)
by: Shoeibi, Mehrdad, et al.
Published: (2026)
Advanced Weakly-Supervised Formula Exploration for Neuro-Symbolic Mathematical Reasoning
by: Wu, Yuxuan, et al.
Published: (2025)
by: Wu, Yuxuan, et al.
Published: (2025)
Weakly Supervised AUC Optimization: A Unified Partial AUC Approach
by: Xie, Zheng, et al.
Published: (2023)
by: Xie, Zheng, et al.
Published: (2023)
Meta-Learn Unimodal Signals with Weak Supervision for Multimodal Sentiment Analysis
by: Mai, Sijie, et al.
Published: (2024)
by: Mai, Sijie, et al.
Published: (2024)
The Elicitation Game: Evaluating Capability Elicitation Techniques
by: Hofstätter, Felix, et al.
Published: (2025)
by: Hofstätter, Felix, et al.
Published: (2025)
Automated Consistency Analysis of LLMs
by: Patwardhan, Aditya, et al.
Published: (2025)
by: Patwardhan, Aditya, et al.
Published: (2025)
Weak Supervision for Improved Precision in Search Systems
by: Vasudevan, Sriram
Published: (2025)
by: Vasudevan, Sriram
Published: (2025)
FastBUS: A Fast Bayesian Framework for Unified Weakly-Supervised Learning
by: Wang, Ziquan, et al.
Published: (2026)
by: Wang, Ziquan, et al.
Published: (2026)
CleverCatch: A Knowledge-Guided Weak Supervision Model for Fraud Detection
by: Mozafari, Amirhossein, et al.
Published: (2025)
by: Mozafari, Amirhossein, et al.
Published: (2025)
How to Train Data-Efficient LLMs
by: Sachdeva, Noveen, et al.
Published: (2024)
by: Sachdeva, Noveen, et al.
Published: (2024)
Test Time Training for Supervised Causal Learning
by: Deng, Zizhen, et al.
Published: (2026)
by: Deng, Zizhen, et al.
Published: (2026)
Training LLM Agents to Empower Humans
by: Ellis, Evan, et al.
Published: (2025)
by: Ellis, Evan, et al.
Published: (2025)
A Unified and Stable Risk Minimization Framework for Weakly Supervised Learning with Theoretical Guarantees
by: Zhang, Miao, et al.
Published: (2025)
by: Zhang, Miao, et al.
Published: (2025)
GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents
by: Cai, Shaofei, et al.
Published: (2024)
by: Cai, Shaofei, et al.
Published: (2024)
WindSeer: Real-time volumetric wind prediction over complex terrain aboard a small UAV
by: Achermann, Florian, et al.
Published: (2024)
by: Achermann, Florian, et al.
Published: (2024)
NEAT: Concept driven Neuron Attribution in LLMs
by: Kavuri, Vivek Hruday, et al.
Published: (2025)
by: Kavuri, Vivek Hruday, et al.
Published: (2025)
Self-Supervised Pre-Training for Precipitation Post-Processor
by: An, Sojung, et al.
Published: (2023)
by: An, Sojung, et al.
Published: (2023)
Can LLMs Learn to Reason Robustly under Noisy Supervision?
by: Yang, Shenzhi, et al.
Published: (2026)
by: Yang, Shenzhi, et al.
Published: (2026)
Weakly Supervised Veracity Classification with LLM-Predicted Credibility Signals
by: Leite, João A., et al.
Published: (2023)
by: Leite, João A., et al.
Published: (2023)
Similar Items
-
Sandbagging in a Simple Survival Bandit Problem
by: Dyer, Joel, et al.
Published: (2025) -
Auditing Sabotage Bench: A Benchmark for Detecting and Fixing Research Sabotage in ML Codebases
by: Gan, Eric, et al.
Published: (2026) -
Efficiently Aligning Language Models with Online Natural Language Feedback
by: Ye, Christine, et al.
Published: (2026) -
Latent Adversarial Training Improves Robustness to Persistent Harmful Behaviors in LLMs
by: Sheshadri, Abhay, et al.
Published: (2024) -
Open Character Training: Shaping the Persona of AI Assistants through Constitutional AI
by: Maiya, Sharan, et al.
Published: (2025)