LLM Novice Uplift on Dual-Use, In Silico Biology Tasks
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Chen Bo Calvin, Knight, Christina Q., Kruus, Nicholas, Hausenloy, Jason, Medeiros, Pedro, Li, Nathaniel, Kim, Aiden, Orlovskiy, Yury, Breen, Coleman, Cai, Bryce, Götting, Jasper, Liu, Andrew Bo, Nedungadi, Samira, Rodriguez, Paula, He, Yannis Yiming, Shaaban, Mohamed, Wang, Zifan, Donoughe, Seth, Michael, Julian |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Virology Capabilities Test (VCT): A Multimodal Virology Q&A Benchmark
di: Götting, Jasper, et al.
Pubblicazione: (2025)
di: Götting, Jasper, et al.
Pubblicazione: (2025)
STREAM (ChemBio): A Standard for Transparently Reporting Evaluations in AI Model Reports
di: McCaslin, Tegan, et al.
Pubblicazione: (2025)
di: McCaslin, Tegan, et al.
Pubblicazione: (2025)
Learners Teaching Novices: An Uplifting Alternative Assessment
di: Malik, Ali, et al.
Pubblicazione: (2024)
di: Malik, Ali, et al.
Pubblicazione: (2024)
Best Practices for Biorisk Evaluations on Open-Weight Bio-Foundation Models
di: Wei, Boyi, et al.
Pubblicazione: (2025)
di: Wei, Boyi, et al.
Pubblicazione: (2025)
Opportunities and Challenges of Frontier Data Governance With Synthetic Data
di: Thakur, Madhavendra, et al.
Pubblicazione: (2025)
di: Thakur, Madhavendra, et al.
Pubblicazione: (2025)
Invisible Saboteurs: Sycophantic LLMs Mislead Novices in Problem-Solving Tasks
di: Bo, Jessica Y., et al.
Pubblicazione: (2025)
di: Bo, Jessica Y., et al.
Pubblicazione: (2025)
SimpleStrat: Diversifying Language Model Generation with Stratification
di: Wong, Justin, et al.
Pubblicazione: (2024)
di: Wong, Justin, et al.
Pubblicazione: (2024)
The Humanist Programming Novice as Novice
di: Elior, Ofer
Pubblicazione: (2025)
di: Elior, Ofer
Pubblicazione: (2025)
Uncertainty Resolution in Misinformation Detection
di: Orlovskiy, Yury, et al.
Pubblicazione: (2024)
di: Orlovskiy, Yury, et al.
Pubblicazione: (2024)
Who's the Leader? Analyzing Novice Workflows in LLM-Assisted Debugging of Machine Learning Code
di: Bo, Jessica Y., et al.
Pubblicazione: (2025)
di: Bo, Jessica Y., et al.
Pubblicazione: (2025)
Continuous Sign Language Recognition with Adapted Conformer via Unsupervised Pretraining
di: Aloysius, Neena, et al.
Pubblicazione: (2024)
di: Aloysius, Neena, et al.
Pubblicazione: (2024)
Disability Representations: Finding Biases in Automatic Image Generation
di: Tevissen, Yannis
Pubblicazione: (2024)
di: Tevissen, Yannis
Pubblicazione: (2024)
Constrained Spectral Uplifting for HDR Environment Maps
di: L. Tódová, et al.
Pubblicazione: (2025)
di: L. Tódová, et al.
Pubblicazione: (2025)
Evaluating Human-AI Safety: A Framework for Measuring Harmful Capability Uplift
di: Vaccaro, Michelle, et al.
Pubblicazione: (2026)
di: Vaccaro, Michelle, et al.
Pubblicazione: (2026)
Critiquing Antipatterns in Novice Code
di: Leo C. Ureel II
Pubblicazione: (2020)
di: Leo C. Ureel II
Pubblicazione: (2020)
Howzat? Appealing to Expert Judgement for Evaluating Human and AI Next-Step Hints for Novice Programmers
di: Brown, Neil C. C., et al.
Pubblicazione: (2024)
di: Brown, Neil C. C., et al.
Pubblicazione: (2024)
SecureGate: Learning When to Reveal PII Safely via Token-Gated Dual-Adapters for Federated LLMs
di: Shaaban, Mohamed, et al.
Pubblicazione: (2026)
di: Shaaban, Mohamed, et al.
Pubblicazione: (2026)
Escaping Neal's Funnel: a multi-stage sampling method for hierarchical models
di: Gundersen, Aiden, et al.
Pubblicazione: (2025)
di: Gundersen, Aiden, et al.
Pubblicazione: (2025)
HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
di: Mazeika, Mantas, et al.
Pubblicazione: (2024)
di: Mazeika, Mantas, et al.
Pubblicazione: (2024)
xFinder: Large Language Models as Automated Evaluators for Reliable Evaluation
di: Yu, Qingchen, et al.
Pubblicazione: (2024)
di: Yu, Qingchen, et al.
Pubblicazione: (2024)
GuessArena: Guess Who I Am? A Self-Adaptive Framework for Evaluating LLMs in Domain-Specific Knowledge and Reasoning
di: Yu, Qingchen, et al.
Pubblicazione: (2025)
di: Yu, Qingchen, et al.
Pubblicazione: (2025)
QCSE: A Pretrained Quantum Context-Sensitive Word Embedding for Natural Language Processing
di: Varmantchaonala, Charles M., et al.
Pubblicazione: (2025)
di: Varmantchaonala, Charles M., et al.
Pubblicazione: (2025)
Exploring How Multiple Levels of GPT-Generated Programming Hints Support or Disappoint Novices
di: Xiao, Ruiwei, et al.
Pubblicazione: (2024)
di: Xiao, Ruiwei, et al.
Pubblicazione: (2024)
Multi-Level Feedback Generation with Large Language Models for Empowering Novice Peer Counselors
di: Chaszczewicz, Alicja, et al.
Pubblicazione: (2024)
di: Chaszczewicz, Alicja, et al.
Pubblicazione: (2024)
Sparse Code Uplifting for Efficient 3D Language Gaussian Splatting
di: Budimir, Lovre Antonio, et al.
Pubblicazione: (2026)
di: Budimir, Lovre Antonio, et al.
Pubblicazione: (2026)
Fairness Evaluation for Uplift Modeling in the Absence of Ground Truth
di: Kadioglu, Serdar, et al.
Pubblicazione: (2024)
di: Kadioglu, Serdar, et al.
Pubblicazione: (2024)
Driving Education Advancements of Novice Drivers: A Systematic Literature Review
di: Tusti, Anannya Ghosh, et al.
Pubblicazione: (2025)
di: Tusti, Anannya Ghosh, et al.
Pubblicazione: (2025)
NoviCode: Generating Programs from Natural Language Utterances by Novices
di: Mordechai, Asaf Achi, et al.
Pubblicazione: (2024)
di: Mordechai, Asaf Achi, et al.
Pubblicazione: (2024)
The Alignment Trap: Complexity Barriers
di: Yao, Jasper
Pubblicazione: (2025)
di: Yao, Jasper
Pubblicazione: (2025)
Uplifting Lower-Income Data: Strategies for Socioeconomic Perspective Shifts in Large Multi-modal Models
di: Nwatu, Joan, et al.
Pubblicazione: (2024)
di: Nwatu, Joan, et al.
Pubblicazione: (2024)
Pose-Based Sign Language Appearance Transfer
di: Moryossef, Amit, et al.
Pubblicazione: (2024)
di: Moryossef, Amit, et al.
Pubblicazione: (2024)
SignBank+: Preparing a Multilingual Sign Language Dataset for Machine Translation Using Large Language Models
di: Moryossef, Amit, et al.
Pubblicazione: (2023)
di: Moryossef, Amit, et al.
Pubblicazione: (2023)
NOVI : Chatbot System for University Novice with BERT and LLMs
di: Nam, Yoonji, et al.
Pubblicazione: (2024)
di: Nam, Yoonji, et al.
Pubblicazione: (2024)
First, Do No Harm: AI Supervisor Scaffolds Novice Growth in Counselor Education
di: Xu, Chen, et al.
Pubblicazione: (2025)
di: Xu, Chen, et al.
Pubblicazione: (2025)
NLP for The Greek Language: A Longer Survey
di: Papantoniou, Katerina, et al.
Pubblicazione: (2024)
di: Papantoniou, Katerina, et al.
Pubblicazione: (2024)
Attention Heads of Large Language Models: A Survey
di: Zheng, Zifan, et al.
Pubblicazione: (2024)
di: Zheng, Zifan, et al.
Pubblicazione: (2024)
MCP4IFC: IFC-Based Building Design Using Large Language Models
di: Nithyanantham, Bharathi Kannan, et al.
Pubblicazione: (2025)
di: Nithyanantham, Bharathi Kannan, et al.
Pubblicazione: (2025)
3DPFIX: Improving Remote Novices' 3D Printing Troubleshooting through Human-AI Collaboration
di: Kwon, Nahyun, et al.
Pubblicazione: (2024)
di: Kwon, Nahyun, et al.
Pubblicazione: (2024)
From Novice to Expert: LLM Agent Policy Optimization via Step-wise Reinforcement Learning
di: Deng, Zhirui, et al.
Pubblicazione: (2024)
di: Deng, Zhirui, et al.
Pubblicazione: (2024)
From Reality to Recognition: Evaluating Visualization Analogies for Novice Chart Comprehension
di: Huang, Oliver, et al.
Pubblicazione: (2025)
di: Huang, Oliver, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Virology Capabilities Test (VCT): A Multimodal Virology Q&A Benchmark
di: Götting, Jasper, et al.
Pubblicazione: (2025) -
STREAM (ChemBio): A Standard for Transparently Reporting Evaluations in AI Model Reports
di: McCaslin, Tegan, et al.
Pubblicazione: (2025) -
Learners Teaching Novices: An Uplifting Alternative Assessment
di: Malik, Ali, et al.
Pubblicazione: (2024) -
Best Practices for Biorisk Evaluations on Open-Weight Bio-Foundation Models
di: Wei, Boyi, et al.
Pubblicazione: (2025) -
Opportunities and Challenges of Frontier Data Governance With Synthetic Data
di: Thakur, Madhavendra, et al.
Pubblicazione: (2025)