Bridging the Data Provenance Gap Across Text, Speech and Video
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Longpre, Shayne, Singh, Nikhil, Cherep, Manuel, Tiwary, Kushagra, Materzynska, Joanna, Brannon, William, Mahari, Robert, Obeng-Marnu, Naana, Dey, Manan, Hamdy, Mohammed, Saxena, Nayan, Anis, Ahmad Mustafa, Alghamdi, Emad A., Chien, Vu Minh, Yin, Da, Qian, Kun, Li, Yizhi, Liang, Minnie, Dinh, An, Mohanty, Shrestha, Mataciunas, Deividas, South, Tobin, Zhang, Jianguo, Lee, Ariel N., Lund, Campbell S., Klamm, Christopher, Sileo, Damien, Misra, Diganta, Shippole, Enrico, Klyman, Kevin, Miranda, Lester JV, Muennighoff, Niklas, Ye, Seonghyeon, Kim, Seungone, Gupta, Vipul, Sharma, Vivek, Zhou, Xuhui, Xiong, Caiming, Villa, Luis, Biderman, Stella, Pentland, Alex, Hooker, Sara, Kabbara, Jad |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Data Authenticity, Consent, & Provenance for AI are all broken: what will it take to fix them?
von: Longpre, Shayne, et al.
Veröffentlicht: (2024)
von: Longpre, Shayne, et al.
Veröffentlicht: (2024)
Consent in Crisis: The Rapid Decline of the AI Data Commons
von: Longpre, Shayne, et al.
Veröffentlicht: (2024)
von: Longpre, Shayne, et al.
Veröffentlicht: (2024)
Future and AI-Ready Data Strategies: Response to DOC RFI on AI and Open Government Data Assets
von: Oderinwale, Hamidah, et al.
Veröffentlicht: (2024)
von: Oderinwale, Hamidah, et al.
Veröffentlicht: (2024)
ATLAS: Adaptive Transfer Scaling Laws for Multilingual Pretraining, Finetuning, and Decoding the Curse of Multilinguality
von: Longpre, Shayne, et al.
Veröffentlicht: (2025)
von: Longpre, Shayne, et al.
Veröffentlicht: (2025)
The 2024 Foundation Model Transparency Index
von: Bommasani, Rishi, et al.
Veröffentlicht: (2024)
von: Bommasani, Rishi, et al.
Veröffentlicht: (2024)
Foundation Model Transparency Reports
von: Bommasani, Rishi, et al.
Veröffentlicht: (2024)
von: Bommasani, Rishi, et al.
Veröffentlicht: (2024)
The 2025 Foundation Model Transparency Index
von: Wan, Alexander, et al.
Veröffentlicht: (2025)
von: Wan, Alexander, et al.
Veröffentlicht: (2025)
LePaRD: A Large-Scale Dataset of Judges Citing Precedents
von: Mahari, Robert, et al.
Veröffentlicht: (2023)
von: Mahari, Robert, et al.
Veröffentlicht: (2023)
A Systematic Review of NeurIPS Dataset Management Practices
von: Wu, Yiwei, et al.
Veröffentlicht: (2024)
von: Wu, Yiwei, et al.
Veröffentlicht: (2024)
Insights from an experiment crowdsourcing data from thousands of US Amazon users: The importance of transparency, money, and data use
von: Berke, Alex, et al.
Veröffentlicht: (2024)
von: Berke, Alex, et al.
Veröffentlicht: (2024)
zkTax: A pragmatic way to support zero-knowledge tax disclosures
von: Berke, Alex, et al.
Veröffentlicht: (2023)
von: Berke, Alex, et al.
Veröffentlicht: (2023)
KMMLU: Measuring Massive Multitask Language Understanding in Korean
von: Son, Guijin, et al.
Veröffentlicht: (2024)
von: Son, Guijin, et al.
Veröffentlicht: (2024)
Building trust in luxury brands through behavioral analytics of customer experience
von: Cherep, Nataliia
Veröffentlicht: (2025)
von: Cherep, Nataliia
Veröffentlicht: (2025)
AI-Powered Autonomous Weapons Risk Geopolitical Instability and Threaten AI Research
von: Simmons-Edler, Riley, et al.
Veröffentlicht: (2024)
von: Simmons-Edler, Riley, et al.
Veröffentlicht: (2024)
Leveraging Large Language Models for Learning Complex Legal Concepts through Storytelling
von: Jiang, Hang, et al.
Veröffentlicht: (2024)
von: Jiang, Hang, et al.
Veröffentlicht: (2024)
Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models
von: Kim, Seungone, et al.
Veröffentlicht: (2024)
von: Kim, Seungone, et al.
Veröffentlicht: (2024)
Self-Explore: Enhancing Mathematical Reasoning in Language Models with Fine-grained Rewards
von: Hwang, Hyeonbin, et al.
Veröffentlicht: (2024)
von: Hwang, Hyeonbin, et al.
Veröffentlicht: (2024)
The Other, the New, and the Old. Three Images of Leith in Irvine Welsh’s Porno
von: Deividas Zibalas
Veröffentlicht: (2021)
von: Deividas Zibalas
Veröffentlicht: (2021)
Los impulsos en la concepción materialista de la razón de Max Horkheimer
von: Paula García Cherep
Veröffentlicht: (2021)
von: Paula García Cherep
Veröffentlicht: (2021)
Prometheus: Inducing Fine-grained Evaluation Capability in Language Models
von: Kim, Seungone, et al.
Veröffentlicht: (2023)
von: Kim, Seungone, et al.
Veröffentlicht: (2023)
OctoPack: Instruction Tuning Code Large Language Models
von: Muennighoff, Niklas, et al.
Veröffentlicht: (2023)
von: Muennighoff, Niklas, et al.
Veröffentlicht: (2023)
LLM-as-an-Interviewer: Beyond Static Testing Through Dynamic LLM Evaluation
von: Kim, Eunsu, et al.
Veröffentlicht: (2024)
von: Kim, Eunsu, et al.
Veröffentlicht: (2024)
Community-centric modeling of citation dynamics explains collective citation patterns in science, law, and patents
von: Kojaku, Sadamori, et al.
Veröffentlicht: (2025)
von: Kojaku, Sadamori, et al.
Veröffentlicht: (2025)
ConGraT: Self-Supervised Contrastive Pretraining for Joint Graph and Text Embeddings
von: Brannon, William, et al.
Veröffentlicht: (2023)
von: Brannon, William, et al.
Veröffentlicht: (2023)
Contrastive Learning from Synthetic Audio Doppelgängers
von: Cherep, Manuel, et al.
Veröffentlicht: (2024)
von: Cherep, Manuel, et al.
Veröffentlicht: (2024)
Bibliographic Instruction for Freshman Students at Florida International University.
von: Dunbar, H. Minnie
Veröffentlicht: (1986)
von: Dunbar, H. Minnie
Veröffentlicht: (1986)
Acceptable Use Policies for Foundation Models
von: Klyman, Kevin
Veröffentlicht: (2024)
von: Klyman, Kevin
Veröffentlicht: (2024)
On the Relationship between Truth and Political Bias in Language Models
von: Fulay, Suyash, et al.
Veröffentlicht: (2024)
von: Fulay, Suyash, et al.
Veröffentlicht: (2024)
The Service Implications of a Rhetorical Approach to Information Literacy
von: Brannon, Brittany
Veröffentlicht: (2017)
von: Brannon, Brittany
Veröffentlicht: (2017)
Examining the Fieldwork Experience from the Site Supervisor Perspective: A Mixed-Methods Study Using Vygotsky's Zone of Proximal Development Theory
von: Brannon, Sian
Veröffentlicht: (2013)
von: Brannon, Sian
Veröffentlicht: (2013)
Assessment in Fieldwork Courses: What Are We Rating?
von: Brannon, Sian
Veröffentlicht: (2014)
von: Brannon, Sian
Veröffentlicht: (2014)
Say No to Speed Bumps!
von: Brannon, Sian
Veröffentlicht: (2010)
von: Brannon, Sian
Veröffentlicht: (2010)
Intentionality is a Design Decision: Measuring Functional Intentionality for Accountable AI Systems
von: Chiappetta, Allessia, et al.
Veröffentlicht: (2026)
von: Chiappetta, Allessia, et al.
Veröffentlicht: (2026)
Scaling Synthetic Logical Reasoning Datasets with Context-Sensitive Declarative Grammars
von: Sileo, Damien
Veröffentlicht: (2024)
von: Sileo, Damien
Veröffentlicht: (2024)
Logic Haystacks: Probing LLMs Long-Context Logical Reasoning (Without Easily Identifiable Unrelated Padding)
von: Sileo, Damien
Veröffentlicht: (2025)
von: Sileo, Damien
Veröffentlicht: (2025)
Attention Overflow: Language Model Input Blur during Long-Context Missing Items Recommendation
von: Sileo, Damien
Veröffentlicht: (2024)
von: Sileo, Damien
Veröffentlicht: (2024)
Discurso de aceptación de la orden Gustavo Machado
von: Enriqueta Sileo
Veröffentlicht: (2015)
von: Enriqueta Sileo
Veröffentlicht: (2015)
Handbook for Louisiana Library Trustees.
von: Lynch, Minnie-Lou, Ed.
Veröffentlicht: (1980)
von: Lynch, Minnie-Lou, Ed.
Veröffentlicht: (1980)
Na ante-sala da discriminação: o preço dos atributos de sexo ecor no Brasil (19891999)
von: Ciro Biderman
Veröffentlicht: (2004)
von: Ciro Biderman
Veröffentlicht: (2004)
Verifiable evaluations of machine learning models using zkSNARKs
von: South, Tobin, et al.
Veröffentlicht: (2024)
von: South, Tobin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Data Authenticity, Consent, & Provenance for AI are all broken: what will it take to fix them?
von: Longpre, Shayne, et al.
Veröffentlicht: (2024) -
Consent in Crisis: The Rapid Decline of the AI Data Commons
von: Longpre, Shayne, et al.
Veröffentlicht: (2024) -
Future and AI-Ready Data Strategies: Response to DOC RFI on AI and Open Government Data Assets
von: Oderinwale, Hamidah, et al.
Veröffentlicht: (2024) -
ATLAS: Adaptive Transfer Scaling Laws for Multilingual Pretraining, Finetuning, and Decoding the Curse of Multilinguality
von: Longpre, Shayne, et al.
Veröffentlicht: (2025) -
The 2024 Foundation Model Transparency Index
von: Bommasani, Rishi, et al.
Veröffentlicht: (2024)