The DeepSpeak Dataset
Fuente:
arXiv
Guardado en:
| Autores principales: | Barrington, Sarah, Bohacek, Maty, Farid, Hany |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Nepotistically Trained Generative-AI Models Collapse
por: Bohacek, Matyas, et al.
Publicado: (2023)
por: Bohacek, Matyas, et al.
Publicado: (2023)
Human Action CLIPs: Detecting AI-generated Human Motion
por: Bohacek, Matyas, et al.
Publicado: (2024)
por: Bohacek, Matyas, et al.
Publicado: (2024)
GenAI Confessions: Black-box Membership Inference for Generative Image Models
por: Bohacek, Matyas, et al.
Publicado: (2025)
por: Bohacek, Matyas, et al.
Publicado: (2025)
Does Head Pose Correction Improve Biometric Facial Recognition?
por: Norman, Justin, et al.
Publicado: (2025)
por: Norman, Justin, et al.
Publicado: (2025)
Detecting Deepfake Talking Heads from Facial Biometric Anomalies
por: Norman, Justin D., et al.
Publicado: (2025)
por: Norman, Justin D., et al.
Publicado: (2025)
Dataset of News Articles with Provenance Metadata for Media Relevance Assessment
por: Peterka, Tomas, et al.
Publicado: (2025)
por: Peterka, Tomas, et al.
Publicado: (2025)
AI-Powered Facial Mask Removal Is Not Suitable For Identification
por: Cooper, Emily A, et al.
Publicado: (2026)
por: Cooper, Emily A, et al.
Publicado: (2026)
Synthetic Human Action Video Data Generation with Pose Transfer
por: Knapp, Vaclav, et al.
Publicado: (2025)
por: Knapp, Vaclav, et al.
Publicado: (2025)
Large Language Models and Provenance Metadata for Determining the Relevance of Images and Videos in News Stories
por: Peterka, Tomas, et al.
Publicado: (2025)
por: Peterka, Tomas, et al.
Publicado: (2025)
MCTED: A Machine-Learning-Ready Dataset for Digital Elevation Model Generation From Mars Imagery
por: Osadnik, Rafał, et al.
Publicado: (2025)
por: Osadnik, Rafał, et al.
Publicado: (2025)
Can Pose Transfer Models Generate Realistic Human Motion?
por: Knapp, Vaclav, et al.
Publicado: (2025)
por: Knapp, Vaclav, et al.
Publicado: (2025)
Finding AI-Generated Faces in the Wild
por: Porcile, Gonzalo J. Aniano, et al.
Publicado: (2023)
por: Porcile, Gonzalo J. Aniano, et al.
Publicado: (2023)
Uncovering Conceptual Blindspots in Generative Image Models Using Sparse Autoencoders
por: Bohacek, Matyas, et al.
Publicado: (2025)
por: Bohacek, Matyas, et al.
Publicado: (2025)
Advancing Automated Deception Detection: A Multimodal Approach to Feature Extraction and Analysis
por: Bahaa, Mohamed, et al.
Publicado: (2024)
por: Bahaa, Mohamed, et al.
Publicado: (2024)
Open Stamped Parts Dataset
por: Antiles, Sarah, et al.
Publicado: (2024)
por: Antiles, Sarah, et al.
Publicado: (2024)
You Only Speak Once to See
por: Yang, Wenhao, et al.
Publicado: (2024)
por: Yang, Wenhao, et al.
Publicado: (2024)
iKUN: Speak to Trackers without Retraining
por: Du, Yunhao, et al.
Publicado: (2023)
por: Du, Yunhao, et al.
Publicado: (2023)
From Press to Pixels: Evolving Urdu Text Recognition
por: Arif, Samee, et al.
Publicado: (2025)
por: Arif, Samee, et al.
Publicado: (2025)
When Vision Speaks for Sound
por: Wen, Xiaofei, et al.
Publicado: (2026)
por: Wen, Xiaofei, et al.
Publicado: (2026)
ViSpeak: Visual Instruction Feedback in Streaming Videos
por: Fu, Shenghao, et al.
Publicado: (2025)
por: Fu, Shenghao, et al.
Publicado: (2025)
EgoSpeak: Learning When to Speak for Egocentric Conversational Agents in the Wild
por: Kim, Junhyeok, et al.
Publicado: (2025)
por: Kim, Junhyeok, et al.
Publicado: (2025)
Data-Efficient Generation for Dataset Distillation
por: Li, Zhe, et al.
Publicado: (2024)
por: Li, Zhe, et al.
Publicado: (2024)
Let ViT Speak: Generative Language-Image Pre-training
por: Fang, Yan, et al.
Publicado: (2026)
por: Fang, Yan, et al.
Publicado: (2026)
LouvreSAE: Sparse Autoencoders for Interpretable and Controllable Style Transfer
por: Panda, Raina, et al.
Publicado: (2025)
por: Panda, Raina, et al.
Publicado: (2025)
BenchSeg: A Large-Scale Dataset and Benchmark for Multi-View Food Video Segmentation
por: AlMughrabi, Ahmad, et al.
Publicado: (2026)
por: AlMughrabi, Ahmad, et al.
Publicado: (2026)
ReferEverything: Towards Segmenting Everything We Can Speak of in Videos
por: Bagchi, Anurag, et al.
Publicado: (2024)
por: Bagchi, Anurag, et al.
Publicado: (2024)
EmbodiedHead: Real-Time Listening and Speaking Avatar for Conversational Agents
por: Zhang, Yu, et al.
Publicado: (2026)
por: Zhang, Yu, et al.
Publicado: (2026)
Let Samples Speak: Mitigating Spurious Correlation by Exploiting the Clusterness of Samples
por: Li, Weiwei, et al.
Publicado: (2025)
por: Li, Weiwei, et al.
Publicado: (2025)
Dataset Distillation with Probabilistic Latent Features
por: Li, Zhe, et al.
Publicado: (2025)
por: Li, Zhe, et al.
Publicado: (2025)
Autonomous Crack Detection using Deep Learning on Synthetic Thermogram Datasets
por: Pimpalkhare, Chinmay Makarand, et al.
Publicado: (2024)
por: Pimpalkhare, Chinmay Makarand, et al.
Publicado: (2024)
Comparison Of Deep Object Detectors On A New Vulnerable Pedestrian Dataset
por: Sharma, Devansh, et al.
Publicado: (2022)
por: Sharma, Devansh, et al.
Publicado: (2022)
StyleTalk++: A Unified Framework for Controlling the Speaking Styles of Talking Heads
por: Wang, Suzhen, et al.
Publicado: (2024)
por: Wang, Suzhen, et al.
Publicado: (2024)
PolySLGen: Online Multimodal Speaking-Listening Reaction Generation in Polyadic Interaction
por: Lin, Zhi-Yi, et al.
Publicado: (2026)
por: Lin, Zhi-Yi, et al.
Publicado: (2026)
Disentangled Concepts Speak Louder Than Words: Explainable Video Action Recognition
por: Lee, Jongseo, et al.
Publicado: (2025)
por: Lee, Jongseo, et al.
Publicado: (2025)
TalkCLIP: Talking Head Generation with Text-Guided Expressive Speaking Styles
por: Ma, Yifeng, et al.
Publicado: (2023)
por: Ma, Yifeng, et al.
Publicado: (2023)
Video Dataset Condensation with Diffusion Models
por: Li, Zhe, et al.
Publicado: (2025)
por: Li, Zhe, et al.
Publicado: (2025)
People are poorly equipped to detect AI-powered voice clones
por: Barrington, Sarah, et al.
Publicado: (2024)
por: Barrington, Sarah, et al.
Publicado: (2024)
Cross-Dataset Semantic Segmentation Performance Analysis: Unifying NIST Point Cloud City Datasets for 3D Deep Learning
por: Dimopoulos, Alexander Nikitas, et al.
Publicado: (2025)
por: Dimopoulos, Alexander Nikitas, et al.
Publicado: (2025)
Relighting from a Single Image: Datasets and Deep Intrinsic-based Architecture
por: Yang, Yixiong, et al.
Publicado: (2024)
por: Yang, Yixiong, et al.
Publicado: (2024)
ANNA: A Deep Learning Based Dataset in Heterogeneous Traffic for Autonomous Vehicles
por: Kamal, Mahedi, et al.
Publicado: (2024)
por: Kamal, Mahedi, et al.
Publicado: (2024)
Ejemplares similares
-
Nepotistically Trained Generative-AI Models Collapse
por: Bohacek, Matyas, et al.
Publicado: (2023) -
Human Action CLIPs: Detecting AI-generated Human Motion
por: Bohacek, Matyas, et al.
Publicado: (2024) -
GenAI Confessions: Black-box Membership Inference for Generative Image Models
por: Bohacek, Matyas, et al.
Publicado: (2025) -
Does Head Pose Correction Improve Biometric Facial Recognition?
por: Norman, Justin, et al.
Publicado: (2025) -
Detecting Deepfake Talking Heads from Facial Biometric Anomalies
por: Norman, Justin D., et al.
Publicado: (2025)