Near to Mid-term Risks and Opportunities of Open-Source Generative AI
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Eiras, Francisco, Petrov, Aleksandar, Vidgen, Bertie, de Witt, Christian Schroeder, Pizzati, Fabio, Elkins, Katherine, Mukhopadhyay, Supratik, Bibi, Adel, Csaba, Botos, Steibel, Fabro, Barez, Fazl, Smith, Genevieve, Guadagni, Gianluca, Chun, Jon, Cabot, Jordi, Imperial, Joseph Marvin, Nolazco-Flores, Juan A., Landay, Lori, Jackson, Matthew, Röttger, Paul, Torr, Philip H. S., Darrell, Trevor, Lee, Yong Suk, Foerster, Jakob |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Risks and Opportunities of Open-Source Generative AI
von: Eiras, Francisco, et al.
Veröffentlicht: (2024)
von: Eiras, Francisco, et al.
Veröffentlicht: (2024)
A Outra Face do Horário Gratuito: Partidos Políticos e Eleições Proporcionais na Televisão
von: Fabro Steibel
Veröffentlicht: (2008)
von: Fabro Steibel
Veröffentlicht: (2008)
SafetyPrompts: a Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety
von: Röttger, Paul, et al.
Veröffentlicht: (2024)
von: Röttger, Paul, et al.
Veröffentlicht: (2024)
Cinema de ação em dia de eleição: Queima de arquivo (1996), Arnold Schwarzenegger e outras mídias da política
von: Fabro Boaz Steibel
Veröffentlicht: (2007)
von: Fabro Boaz Steibel
Veröffentlicht: (2007)
Beyond Linear Probes: Dynamic Safety Monitoring for Language Models
von: Oldfield, James, et al.
Veröffentlicht: (2025)
von: Oldfield, James, et al.
Veröffentlicht: (2025)
Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models
von: Lan, Michael, et al.
Veröffentlicht: (2023)
von: Lan, Michael, et al.
Veröffentlicht: (2023)
Do Sparse Autoencoders Generalize? A Case Study of Answerability
von: Heindrich, Lovis, et al.
Veröffentlicht: (2025)
von: Heindrich, Lovis, et al.
Veröffentlicht: (2025)
XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
von: Röttger, Paul, et al.
Veröffentlicht: (2023)
von: Röttger, Paul, et al.
Veröffentlicht: (2023)
SafetyNet: Detecting Harmful Outputs in LLMs by Modeling and Monitoring Deceptive Behaviors
von: Chaudhary, Maheep, et al.
Veröffentlicht: (2025)
von: Chaudhary, Maheep, et al.
Veröffentlicht: (2025)
Understanding Addition in Transformers
von: Quirke, Philip, et al.
Veröffentlicht: (2023)
von: Quirke, Philip, et al.
Veröffentlicht: (2023)
Same Question, Different Words: A Latent Adversarial Framework for Prompt Robustness
von: Fu, Tingchen, et al.
Veröffentlicht: (2025)
von: Fu, Tingchen, et al.
Veröffentlicht: (2025)
Do as I do (Safely): Mitigating Task-Specific Fine-tuning Risks in Large Language Models
von: Eiras, Francisco, et al.
Veröffentlicht: (2024)
von: Eiras, Francisco, et al.
Veröffentlicht: (2024)
Label Delay in Online Continual Learning
von: Csaba, Botos, et al.
Veröffentlicht: (2023)
von: Csaba, Botos, et al.
Veröffentlicht: (2023)
Query Circuits: Explaining How Language Models Answer User Prompts
von: Wu, Tung-Yu, et al.
Veröffentlicht: (2025)
von: Wu, Tung-Yu, et al.
Veröffentlicht: (2025)
Rethinking Safety in LLM Fine-tuning: An Optimization Perspective
von: Kim, Minseon, et al.
Veröffentlicht: (2025)
von: Kim, Minseon, et al.
Veröffentlicht: (2025)
Classification is a RAG problem: A case study on hate speech detection
von: Willats, Richard, et al.
Veröffentlicht: (2025)
von: Willats, Richard, et al.
Veröffentlicht: (2025)
SimpleSafetyTests: a Test Suite for Identifying Critical Safety Risks in Large Language Models
von: Vidgen, Bertie, et al.
Veröffentlicht: (2023)
von: Vidgen, Bertie, et al.
Veröffentlicht: (2023)
Quantifying Feature Space Universality Across Large Language Models via Sparse Autoencoders
von: Lan, Michael, et al.
Veröffentlicht: (2024)
von: Lan, Michael, et al.
Veröffentlicht: (2024)
Towards Interpreting Visual Information Processing in Vision-Language Models
von: Neo, Clement, et al.
Veröffentlicht: (2024)
von: Neo, Clement, et al.
Veröffentlicht: (2024)
Rethinking AI Cultural Alignment
von: Bravansky, Michal, et al.
Veröffentlicht: (2025)
von: Bravansky, Michal, et al.
Veröffentlicht: (2025)
Understanding Addition and Subtraction in Transformers
von: Quirke, Philip, et al.
Veröffentlicht: (2024)
von: Quirke, Philip, et al.
Veröffentlicht: (2024)
Prompting a Pretrained Transformer Can Be a Universal Approximator
von: Petrov, Aleksandar, et al.
Veröffentlicht: (2024)
von: Petrov, Aleksandar, et al.
Veröffentlicht: (2024)
When Do Prompting and Prefix-Tuning Work? A Theory of Capabilities and Limitations
von: Petrov, Aleksandar, et al.
Veröffentlicht: (2023)
von: Petrov, Aleksandar, et al.
Veröffentlicht: (2023)
PoisonBench: Assessing Large Language Model Vulnerability to Data Poisoning
von: Fu, Tingchen, et al.
Veröffentlicht: (2024)
von: Fu, Tingchen, et al.
Veröffentlicht: (2024)
Token Taxes: mitigating AGI's economic risks
von: Irwin, Lucas, et al.
Veröffentlicht: (2026)
von: Irwin, Lucas, et al.
Veröffentlicht: (2026)
Large Language Models Relearn Removed Concepts
von: Lo, Michelle, et al.
Veröffentlicht: (2024)
von: Lo, Michelle, et al.
Veröffentlicht: (2024)
VAL-Bench: Belief Consistency as a measure for Value Alignment in Language Models
von: Gupta, Aman, et al.
Veröffentlicht: (2025)
von: Gupta, Aman, et al.
Veröffentlicht: (2025)
Interpreting Context Look-ups in Transformers: Investigating Attention-MLP Interactions
von: Neo, Clement, et al.
Veröffentlicht: (2024)
von: Neo, Clement, et al.
Veröffentlicht: (2024)
Interpreting Learned Feedback Patterns in Large Language Models
von: Marks, Luke, et al.
Veröffentlicht: (2023)
von: Marks, Luke, et al.
Veröffentlicht: (2023)
Enhancing Neural Network Interpretability with Feature-Aligned Sparse Autoencoders
von: Marks, Luke, et al.
Veröffentlicht: (2024)
von: Marks, Luke, et al.
Veröffentlicht: (2024)
Towards Understanding Subliminal Learning: When and How Hidden Biases Transfer
von: Schrodi, Simon, et al.
Veröffentlicht: (2025)
von: Schrodi, Simon, et al.
Veröffentlicht: (2025)
Visualizing Neural Network Imagination
von: Wichers, Nevan, et al.
Veröffentlicht: (2024)
von: Wichers, Nevan, et al.
Veröffentlicht: (2024)
Segment, Select, Correct: A Framework for Weakly-Supervised Referring Segmentation
von: Eiras, Francisco, et al.
Veröffentlicht: (2023)
von: Eiras, Francisco, et al.
Veröffentlicht: (2023)
Why human-AI relationships need socioaffective alignment
von: Kirk, Hannah Rose, et al.
Veröffentlicht: (2025)
von: Kirk, Hannah Rose, et al.
Veröffentlicht: (2025)
De la bonne monnaie. Documents originaux de « l’Union catholique des études sociales et économiques » de Fribourg (1884–1903)
von: Máté, Botos
Veröffentlicht: (2026)
von: Máté, Botos
Veröffentlicht: (2026)
De la bonne monnaie. Documents originaux de « l’Union catholique des études sociales et économiques » de Fribourg (1884–1903)
von: Máté, Botos
Veröffentlicht: (2026)
von: Máté, Botos
Veröffentlicht: (2026)
On Pretraining Data Diversity for Self-Supervised Learning
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2024)
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2024)
SynthCLIP: Are We Ready for a Fully Synthetic CLIP Training?
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2024)
von: Hammoud, Hasan Abed Al Kader, et al.
Veröffentlicht: (2024)
Efficient Error Certification for Physics-Informed Neural Networks
von: Eiras, Francisco, et al.
Veröffentlicht: (2023)
von: Eiras, Francisco, et al.
Veröffentlicht: (2023)
On the Coexistence and Ensembling of Watermarks
von: Petrov, Aleksandar, et al.
Veröffentlicht: (2025)
von: Petrov, Aleksandar, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Risks and Opportunities of Open-Source Generative AI
von: Eiras, Francisco, et al.
Veröffentlicht: (2024) -
A Outra Face do Horário Gratuito: Partidos Políticos e Eleições Proporcionais na Televisão
von: Fabro Steibel
Veröffentlicht: (2008) -
SafetyPrompts: a Systematic Review of Open Datasets for Evaluating and Improving Large Language Model Safety
von: Röttger, Paul, et al.
Veröffentlicht: (2024) -
Cinema de ação em dia de eleição: Queima de arquivo (1996), Arnold Schwarzenegger e outras mídias da política
von: Fabro Boaz Steibel
Veröffentlicht: (2007) -
Beyond Linear Probes: Dynamic Safety Monitoring for Language Models
von: Oldfield, James, et al.
Veröffentlicht: (2025)