Risks and Opportunities of Open-Source Generative AI
Fuente:
arXiv
Saved in:
| Main Authors: | Eiras, Francisco, Petrov, Aleksandar, Vidgen, Bertie, Schroeder, Christian, Pizzati, Fabio, Elkins, Katherine, Mukhopadhyay, Supratik, Bibi, Adel, Purewal, Aaron, Botos, Csaba, Steibel, Fabro, Keshtkar, Fazel, Barez, Fazl, Smith, Genevieve, Guadagni, Gianluca, Chun, Jon, Cabot, Jordi, Imperial, Joseph, Nolazco, Juan Arturo, Landay, Lori, Jackson, Matthew, Torr, Phillip H. S., Darrell, Trevor, Lee, Yong, Foerster, Jakob |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Near to Mid-term Risks and Opportunities of Open-Source Generative AI
by: Eiras, Francisco, et al.
Published: (2024)
by: Eiras, Francisco, et al.
Published: (2024)
A Outra Face do Horário Gratuito: Partidos Políticos e Eleições Proporcionais na Televisão
by: Fabro Steibel
Published: (2008)
by: Fabro Steibel
Published: (2008)
Cinema de ação em dia de eleição: Queima de arquivo (1996), Arnold Schwarzenegger e outras mídias da política
by: Fabro Boaz Steibel
Published: (2007)
by: Fabro Boaz Steibel
Published: (2007)
Beyond Linear Probes: Dynamic Safety Monitoring for Language Models
by: Oldfield, James, et al.
Published: (2025)
by: Oldfield, James, et al.
Published: (2025)
Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models
by: Lan, Michael, et al.
Published: (2023)
by: Lan, Michael, et al.
Published: (2023)
Do Sparse Autoencoders Generalize? A Case Study of Answerability
by: Heindrich, Lovis, et al.
Published: (2025)
by: Heindrich, Lovis, et al.
Published: (2025)
SafetyNet: Detecting Harmful Outputs in LLMs by Modeling and Monitoring Deceptive Behaviors
by: Chaudhary, Maheep, et al.
Published: (2025)
by: Chaudhary, Maheep, et al.
Published: (2025)
Understanding Addition in Transformers
by: Quirke, Philip, et al.
Published: (2023)
by: Quirke, Philip, et al.
Published: (2023)
Same Question, Different Words: A Latent Adversarial Framework for Prompt Robustness
by: Fu, Tingchen, et al.
Published: (2025)
by: Fu, Tingchen, et al.
Published: (2025)
Label Delay in Online Continual Learning
by: Csaba, Botos, et al.
Published: (2023)
by: Csaba, Botos, et al.
Published: (2023)
Query Circuits: Explaining How Language Models Answer User Prompts
by: Wu, Tung-Yu, et al.
Published: (2025)
by: Wu, Tung-Yu, et al.
Published: (2025)
Rethinking Safety in LLM Fine-tuning: An Optimization Perspective
by: Kim, Minseon, et al.
Published: (2025)
by: Kim, Minseon, et al.
Published: (2025)
SafetyPrompts: a Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety
by: Röttger, Paul, et al.
Published: (2024)
by: Röttger, Paul, et al.
Published: (2024)
Classification is a RAG problem: A case study on hate speech detection
by: Willats, Richard, et al.
Published: (2025)
by: Willats, Richard, et al.
Published: (2025)
Quantifying Feature Space Universality Across Large Language Models via Sparse Autoencoders
by: Lan, Michael, et al.
Published: (2024)
by: Lan, Michael, et al.
Published: (2024)
Towards Interpreting Visual Information Processing in Vision-Language Models
by: Neo, Clement, et al.
Published: (2024)
by: Neo, Clement, et al.
Published: (2024)
Rethinking AI Cultural Alignment
by: Bravansky, Michal, et al.
Published: (2025)
by: Bravansky, Michal, et al.
Published: (2025)
Understanding Addition and Subtraction in Transformers
by: Quirke, Philip, et al.
Published: (2024)
by: Quirke, Philip, et al.
Published: (2024)
PoisonBench: Assessing Large Language Model Vulnerability to Data Poisoning
by: Fu, Tingchen, et al.
Published: (2024)
by: Fu, Tingchen, et al.
Published: (2024)
Token Taxes: mitigating AGI's economic risks
by: Irwin, Lucas, et al.
Published: (2026)
by: Irwin, Lucas, et al.
Published: (2026)
Large Language Models Relearn Removed Concepts
by: Lo, Michelle, et al.
Published: (2024)
by: Lo, Michelle, et al.
Published: (2024)
VAL-Bench: Belief Consistency as a measure for Value Alignment in Language Models
by: Gupta, Aman, et al.
Published: (2025)
by: Gupta, Aman, et al.
Published: (2025)
Interpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions
by: Neo, Clement, et al.
Published: (2024)
by: Neo, Clement, et al.
Published: (2024)
Interpreting Learned Feedback Patterns in Large Language Models
by: Marks, Luke, et al.
Published: (2023)
by: Marks, Luke, et al.
Published: (2023)
Enhancing Neural Network Interpretability with Feature-Aligned Sparse Autoencoders
by: Marks, Luke, et al.
Published: (2024)
by: Marks, Luke, et al.
Published: (2024)
Towards Understanding Subliminal Learning: When and How Hidden Biases Transfer
by: Schrodi, Simon, et al.
Published: (2025)
by: Schrodi, Simon, et al.
Published: (2025)
Visualizing Neural Network Imagination
by: Wichers, Nevan, et al.
Published: (2024)
by: Wichers, Nevan, et al.
Published: (2024)
Do as I do (Safely): Mitigating Task-Specific Fine-tuning Risks in Large Language Models
by: Eiras, Francisco, et al.
Published: (2024)
by: Eiras, Francisco, et al.
Published: (2024)
Segment, Select, Correct: A Framework for Weakly-Supervised Referring Segmentation
by: Eiras, Francisco, et al.
Published: (2023)
by: Eiras, Francisco, et al.
Published: (2023)
A Narrative-Driven Computational Framework for Clinician Burnout Surveillance
by: Bukhari, Syed Ahmad Chan, et al.
Published: (2025)
by: Bukhari, Syed Ahmad Chan, et al.
Published: (2025)
Why human-AI relationships need socioaffective alignment
by: Kirk, Hannah Rose, et al.
Published: (2025)
by: Kirk, Hannah Rose, et al.
Published: (2025)
De la bonne monnaie. Documents originaux de « l’Union catholique des études sociales et économiques » de Fribourg (1884–1903)
by: Máté, Botos
Published: (2026)
by: Máté, Botos
Published: (2026)
De la bonne monnaie. Documents originaux de « l’Union catholique des études sociales et économiques » de Fribourg (1884–1903)
by: Máté, Botos
Published: (2026)
by: Máté, Botos
Published: (2026)
On Pretraining Data Diversity for Self-Supervised Learning
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2024)
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2024)
SynthCLIP: Are We Ready for a Fully Synthetic CLIP Training?
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2024)
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2024)
Efficient Error Certification for Physics-Informed Neural Networks
by: Eiras, Francisco, et al.
Published: (2023)
by: Eiras, Francisco, et al.
Published: (2023)
Embodied AI: Emerging Risks and Opportunities for Policy Action
by: Perlo, Jared, et al.
Published: (2025)
by: Perlo, Jared, et al.
Published: (2025)
Beyond Linear Steering: Unified Multi-Attribute Control for Language Models
by: Oozeer, Narmeen, et al.
Published: (2025)
by: Oozeer, Narmeen, et al.
Published: (2025)
Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer
by: Simhi, Adi, et al.
Published: (2025)
by: Simhi, Adi, et al.
Published: (2025)
Chain-of-Thought Hijacking
by: Zhao, Jianli, et al.
Published: (2025)
by: Zhao, Jianli, et al.
Published: (2025)
Similar Items
-
Near to Mid-term Risks and Opportunities of Open-Source Generative AI
by: Eiras, Francisco, et al.
Published: (2024) -
A Outra Face do Horário Gratuito: Partidos Políticos e Eleições Proporcionais na Televisão
by: Fabro Steibel
Published: (2008) -
Cinema de ação em dia de eleição: Queima de arquivo (1996), Arnold Schwarzenegger e outras mídias da política
by: Fabro Boaz Steibel
Published: (2007) -
Beyond Linear Probes: Dynamic Safety Monitoring for Language Models
by: Oldfield, James, et al.
Published: (2025) -
Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models
by: Lan, Michael, et al.
Published: (2023)