Saved in:
| Main Authors: | Della Santa, Francesco, Lalli, Morgana |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2501.16100 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Renal Cell Carcinoma subtyping: learning from multi-resolution localization
by: Mohamad, Mohamad, et al.
Published: (2024)
by: Mohamad, Mohamad, et al.
Published: (2024)
Video Editing for Audio-Visual Dubbing
by: Manela, Binyamin, et al.
Published: (2025)
by: Manela, Binyamin, et al.
Published: (2025)
MAVOS-DD: Multilingual Audio-Video Open-Set Deepfake Detection Benchmark
by: Croitoru, Florinel-Alin, et al.
Published: (2025)
by: Croitoru, Florinel-Alin, et al.
Published: (2025)
Seeing Voices: Generating A-Roll Video from Audio with Mirage
by: Sundararaman, Aditi, et al.
Published: (2025)
by: Sundararaman, Aditi, et al.
Published: (2025)
Prisma: An Open Source Toolkit for Mechanistic Interpretability in Vision and Video
by: Joseph, Sonia, et al.
Published: (2025)
by: Joseph, Sonia, et al.
Published: (2025)
VidTok: A Versatile and Open-Source Video Tokenizer
by: Tang, Anni, et al.
Published: (2024)
by: Tang, Anni, et al.
Published: (2024)
X-AVDT: Audio-Visual Cross-Attention for Robust Deepfake Detection
by: Kim, Youngseo, et al.
Published: (2026)
by: Kim, Youngseo, et al.
Published: (2026)
Source-Free Domain Adaptation for YOLO Object Detection
by: Varailhon, Simon, et al.
Published: (2024)
by: Varailhon, Simon, et al.
Published: (2024)
Detection-Fusion for Knowledge Graph Extraction from Videos
by: Das, Taniya, et al.
Published: (2024)
by: Das, Taniya, et al.
Published: (2024)
Sounding Highlights: Dual-Pathway Audio Encoders for Audio-Visual Video Highlight Detection
by: Joo, Seohyun, et al.
Published: (2026)
by: Joo, Seohyun, et al.
Published: (2026)
Video Anomaly Detection with Structured Keywords
by: Foltz, Thomas
Published: (2025)
by: Foltz, Thomas
Published: (2025)
Automated Model Evaluation for Object Detection via Prediction Consistency and Reliability
by: Yoo, Seungju, et al.
Published: (2025)
by: Yoo, Seungju, et al.
Published: (2025)
PaSTe: Improving the Efficiency of Visual Anomaly Detection at the Edge
by: Barusco, Manuel, et al.
Published: (2024)
by: Barusco, Manuel, et al.
Published: (2024)
Adversarial Attacks on Audio Deepfake Detection: A Benchmark and Comparative Study
by: Uddin, Kutub, et al.
Published: (2025)
by: Uddin, Kutub, et al.
Published: (2025)
AI-Generated Video Detection via Perceptual Straightening
by: Internò, Christian, et al.
Published: (2025)
by: Internò, Christian, et al.
Published: (2025)
Advanced Gesture Recognition for Autism Spectrum Disorder Detection: Integrating YOLOv7, Video Augmentation, and VideoMAE for Naturalistic Video Analysis
by: Singh, Amit Kumar, et al.
Published: (2024)
by: Singh, Amit Kumar, et al.
Published: (2024)
Cataract-LMM Large-Scale Multi-Source Multi-Task Benchmark for Deep Learning in Surgical Video Analysis
by: Ahmadi, Mohammad Javad, et al.
Published: (2025)
by: Ahmadi, Mohammad Javad, et al.
Published: (2025)
UniRepLKNet: A Universal Perception Large-Kernel ConvNet for Audio, Video, Point Cloud, Time-Series and Image Recognition
by: Ding, Xiaohan, et al.
Published: (2023)
by: Ding, Xiaohan, et al.
Published: (2023)
TrackNetV4: Enhancing Fast Sports Object Tracking with Motion Attention Maps
by: Raj, Arjun, et al.
Published: (2024)
by: Raj, Arjun, et al.
Published: (2024)
LAPT: Label-driven Automated Prompt Tuning for OOD Detection with Vision-Language Models
by: Zhang, Yabin, et al.
Published: (2024)
by: Zhang, Yabin, et al.
Published: (2024)
SAVe: Self-Supervised Audio-visual Deepfake Detection Exploiting Visual Artifacts and Audio-visual Misalignment
by: Shahzad, Sahibzada Adil, et al.
Published: (2026)
by: Shahzad, Sahibzada Adil, et al.
Published: (2026)
ALARM: Automated MLLM-Based Anomaly Detection in Complex-EnviRonment Monitoring with Uncertainty Quantification
by: Zhang, Congjing, et al.
Published: (2025)
by: Zhang, Congjing, et al.
Published: (2025)
Automated Bi-Fold Weighted Ensemble Algorithms and its Application to Brain Tumor Detection and Classification
by: Huang, PoTsang B., et al.
Published: (2024)
by: Huang, PoTsang B., et al.
Published: (2024)
DAVE: Diagnostic benchmark for Audio Visual Evaluation
by: Radevski, Gorjan, et al.
Published: (2025)
by: Radevski, Gorjan, et al.
Published: (2025)
Intelligent Video Recording Optimization using Activity Detection for Surveillance Systems
by: Elmir, Youssef, et al.
Published: (2024)
by: Elmir, Youssef, et al.
Published: (2024)
Privacy-Aware Video Anomaly Detection through Orthogonal Subspace Projection
by: Wang, Lei, et al.
Published: (2026)
by: Wang, Lei, et al.
Published: (2026)
Anomaly-Aware Vision-Language Adapters for Zero-Shot Anomaly Detection
by: Aqeel, Muhammad, et al.
Published: (2026)
by: Aqeel, Muhammad, et al.
Published: (2026)
Dual Branch VideoMamba with Gated Class Token Fusion for Violence Detection
by: Senadeera, Damith Chamalke, et al.
Published: (2025)
by: Senadeera, Damith Chamalke, et al.
Published: (2025)
A Scalable and Generalized Deep Learning Framework for Anomaly Detection in Surveillance Videos
by: Jebur, Sabah Abdulazeez, et al.
Published: (2024)
by: Jebur, Sabah Abdulazeez, et al.
Published: (2024)
Source-Free Domain Adaptation with Diffusion-Guided Source Data Generation
by: Chopra, Shivang, et al.
Published: (2024)
by: Chopra, Shivang, et al.
Published: (2024)
Unleash the Potential of CLIP for Video Highlight Detection
by: Han, Donghoon, et al.
Published: (2024)
by: Han, Donghoon, et al.
Published: (2024)
EgoSurgery-Tool: A Dataset of Surgical Tool and Hand Detection from Egocentric Open Surgery Videos
by: Fujii, Ryo, et al.
Published: (2024)
by: Fujii, Ryo, et al.
Published: (2024)
VERA: Explainable Video Anomaly Detection via Verbalized Learning of Vision-Language Models
by: Ye, Muchao, et al.
Published: (2024)
by: Ye, Muchao, et al.
Published: (2024)
Just Dance with $π$! A Poly-modal Inductor for Weakly-supervised Video Anomaly Detection
by: Majhi, Snehashis, et al.
Published: (2025)
by: Majhi, Snehashis, et al.
Published: (2025)
Reasoning-Guided Grounding: Elevating Video Anomaly Detection through Multimodal Large Language Models
by: Agarwal, Sakshi, et al.
Published: (2026)
by: Agarwal, Sakshi, et al.
Published: (2026)
Video Anomaly Detection via Spatio-Temporal Pseudo-Anomaly Generation : A Unified Approach
by: Rai, Ayush K., et al.
Published: (2023)
by: Rai, Ayush K., et al.
Published: (2023)
Audio-3DVG: Unified Audio -- Point Cloud Fusion for 3D Visual Grounding
by: Cao-Dinh, Duc, et al.
Published: (2025)
by: Cao-Dinh, Duc, et al.
Published: (2025)
PortraitTalk: Towards Customizable One-Shot Audio-to-Talking Face Generation
by: Nazarieh, Fatemeh, et al.
Published: (2024)
by: Nazarieh, Fatemeh, et al.
Published: (2024)
Semantic Communication based on Generative AI: A New Approach to Image Compression and Edge Optimization
by: Pezone, Francesco
Published: (2025)
by: Pezone, Francesco
Published: (2025)
Towards Automated Machine Learning Research
by: Ardeshir, Shervin
Published: (2024)
by: Ardeshir, Shervin
Published: (2024)
Similar Items
-
Renal Cell Carcinoma subtyping: learning from multi-resolution localization
by: Mohamad, Mohamad, et al.
Published: (2024) -
Video Editing for Audio-Visual Dubbing
by: Manela, Binyamin, et al.
Published: (2025) -
MAVOS-DD: Multilingual Audio-Video Open-Set Deepfake Detection Benchmark
by: Croitoru, Florinel-Alin, et al.
Published: (2025) -
Seeing Voices: Generating A-Roll Video from Audio with Mirage
by: Sundararaman, Aditi, et al.
Published: (2025) -
Prisma: An Open Source Toolkit for Mechanistic Interpretability in Vision and Video
by: Joseph, Sonia, et al.
Published: (2025)