Comparative Analysis of Image, Video, and Audio Classifiers for Automated News Video Segmentation
Fuente:
arXiv
Saved in:
| Main Authors: | Attard, Jonathan, Seychell, Dylan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Hybrid Deterministic Framework for Named Entity Extraction in Broadcast News Video
by: Lucas, Andrea Filiberto, et al.
Published: (2026)
by: Lucas, Andrea Filiberto, et al.
Published: (2026)
Correlation of Object Detection Performance with Visual Saliency and Depth Estimation
by: Bartolo, Matthias, et al.
Published: (2024)
by: Bartolo, Matthias, et al.
Published: (2024)
Integrating Saliency Ranking and Reinforcement Learning for Enhanced Object Detection
by: Bartolo, Matthias, et al.
Published: (2024)
by: Bartolo, Matthias, et al.
Published: (2024)
X2SAM: Any Segmentation in Images and Videos
by: Wang, Hao, et al.
Published: (2026)
by: Wang, Hao, et al.
Published: (2026)
Automated Detection of Sport Highlights from Audio and Video Sources
by: Della Santa, Francesco, et al.
Published: (2025)
by: Della Santa, Francesco, et al.
Published: (2025)
SAM2 for Image and Video Segmentation: A Comprehensive Survey
by: Jiaxing, Zhang, et al.
Published: (2025)
by: Jiaxing, Zhang, et al.
Published: (2025)
Exploring Audio Hallucination in Egocentric Video Understanding
by: Seth, Ashish, et al.
Published: (2026)
by: Seth, Ashish, et al.
Published: (2026)
SALOVA: Segment-Augmented Long Video Assistant for Targeted Retrieval and Routing in Long-Form Video Analysis
by: Kim, Junho, et al.
Published: (2024)
by: Kim, Junho, et al.
Published: (2024)
VideoGUI: A Benchmark for GUI Automation from Instructional Videos
by: Lin, Kevin Qinghong, et al.
Published: (2024)
by: Lin, Kevin Qinghong, et al.
Published: (2024)
Automated Wicket-Taking Delivery Segmentation and Trajectory-Based Dismissal-Zone Analysis in Cricket Videos Using OCR-Guided YOLOv8
by: Karmoker, Joy, et al.
Published: (2025)
by: Karmoker, Joy, et al.
Published: (2025)
AVBench: Human-Aligned and Automated Evaluation Benchmark for Audio-Video Generative Models
by: Yang, Jialiang, et al.
Published: (2026)
by: Yang, Jialiang, et al.
Published: (2026)
General and Task-Oriented Video Segmentation
by: Chen, Mu, et al.
Published: (2024)
by: Chen, Mu, et al.
Published: (2024)
Semi-Supervised Audio-Visual Video Action Recognition with Audio Source Localization Guided Mixup
by: Kang, Seokun, et al.
Published: (2025)
by: Kang, Seokun, et al.
Published: (2025)
Audio-Sync Video Generation with Multi-Stream Temporal Control
by: Weng, Shuchen, et al.
Published: (2025)
by: Weng, Shuchen, et al.
Published: (2025)
Audio-centric Video Understanding Benchmark without Text Shortcut
by: Yang, Yudong, et al.
Published: (2025)
by: Yang, Yudong, et al.
Published: (2025)
SAM 2: Segment Anything in Images and Videos
by: Ravi, Nikhila, et al.
Published: (2024)
by: Ravi, Nikhila, et al.
Published: (2024)
Attend-Fusion: Efficient Audio-Visual Fusion for Video Classification
by: Awan, Mahrukh, et al.
Published: (2024)
by: Awan, Mahrukh, et al.
Published: (2024)
Affective Video Content Analysis: Decade Review and New Perspectives
by: Xue, Junxiao, et al.
Published: (2023)
by: Xue, Junxiao, et al.
Published: (2023)
Semantic Segmentation of Video Sequences with Convolutional LSTMs
by: Pfeuffer, Andreas, et al.
Published: (2019)
by: Pfeuffer, Andreas, et al.
Published: (2019)
Segment Anything for Videos: A Systematic Survey
by: Zhang, Chunhui, et al.
Published: (2024)
by: Zhang, Chunhui, et al.
Published: (2024)
Consistent Video Editing as Flow-Driven Image-to-Video Generation
by: Wang, Ge, et al.
Published: (2025)
by: Wang, Ge, et al.
Published: (2025)
Audio-Driven Talking Face Video Generation with Joint Uncertainty Learning
by: Xie, Yifan, et al.
Published: (2025)
by: Xie, Yifan, et al.
Published: (2025)
Empowering LLMs with Pseudo-Untrimmed Videos for Audio-Visual Temporal Understanding
by: Tang, Yolo Yunlong, et al.
Published: (2024)
by: Tang, Yolo Yunlong, et al.
Published: (2024)
Comparing Learning Paradigms for Egocentric Video Summarization
by: Wen, Daniel
Published: (2025)
by: Wen, Daniel
Published: (2025)
Causally Steered Diffusion for Automated Video Counterfactual Generation
by: Spyrou, Nikos, et al.
Published: (2025)
by: Spyrou, Nikos, et al.
Published: (2025)
FreeSliders: Training-Free, Modality-Agnostic Concept Sliders for Fine-Grained Diffusion Control in Images, Audio, and Video
by: Ezra, Rotem, et al.
Published: (2025)
by: Ezra, Rotem, et al.
Published: (2025)
Cross-Modal Transferable Image-to-Video Attack on Video Quality Metrics
by: Gotin, Georgii, et al.
Published: (2025)
by: Gotin, Georgii, et al.
Published: (2025)
Video Editing for Audio-Visual Dubbing
by: Manela, Binyamin, et al.
Published: (2025)
by: Manela, Binyamin, et al.
Published: (2025)
On Moving Object Segmentation from Monocular Video with Transformers
by: Homeyer, Christian, et al.
Published: (2024)
by: Homeyer, Christian, et al.
Published: (2024)
Space-time Reinforcement Network for Video Object Segmentation
by: Chen, Yadang, et al.
Published: (2024)
by: Chen, Yadang, et al.
Published: (2024)
SMART: Shot-Aware Multimodal Video Moment Retrieval with Audio-Enhanced MLLM
by: Yu, An, et al.
Published: (2025)
by: Yu, An, et al.
Published: (2025)
ViSAudio: End-to-End Video-Driven Binaural Spatial Audio Generation
by: Zhang, Mengchen, et al.
Published: (2025)
by: Zhang, Mengchen, et al.
Published: (2025)
Echoes Over Time: Unlocking Length Generalization in Video-to-Audio Generation Models
by: Simon, Christian, et al.
Published: (2026)
by: Simon, Christian, et al.
Published: (2026)
Leveraging the Video-level Semantic Consistency of Event for Audio-visual Event Localization
by: Jiang, Yuanyuan, et al.
Published: (2022)
by: Jiang, Yuanyuan, et al.
Published: (2022)
AI as a Tool for Fair Journalism: Case Studies from Malta
by: Seychell, Dylan, et al.
Published: (2024)
by: Seychell, Dylan, et al.
Published: (2024)
Spatiotemporal Learning with Context-aware Video Tubelets for Ultrasound Video Analysis
by: Li, Gary Y., et al.
Published: (2025)
by: Li, Gary Y., et al.
Published: (2025)
Towards Video Anomaly Retrieval from Video Anomaly Detection: New Benchmarks and Model
by: Wu, Peng, et al.
Published: (2023)
by: Wu, Peng, et al.
Published: (2023)
RefSAM: Efficiently Adapting Segmenting Anything Model for Referring Video Object Segmentation
by: Li, Yonglin, et al.
Published: (2023)
by: Li, Yonglin, et al.
Published: (2023)
Audio-visual Event Localization on Portrait Mode Short Videos
by: Liu, Wuyang, et al.
Published: (2025)
by: Liu, Wuyang, et al.
Published: (2025)
Label-anticipated Event Disentanglement for Audio-Visual Video Parsing
by: Zhou, Jinxing, et al.
Published: (2024)
by: Zhou, Jinxing, et al.
Published: (2024)
Similar Items
-
A Hybrid Deterministic Framework for Named Entity Extraction in Broadcast News Video
by: Lucas, Andrea Filiberto, et al.
Published: (2026) -
Correlation of Object Detection Performance with Visual Saliency and Depth Estimation
by: Bartolo, Matthias, et al.
Published: (2024) -
Integrating Saliency Ranking and Reinforcement Learning for Enhanced Object Detection
by: Bartolo, Matthias, et al.
Published: (2024) -
X2SAM: Any Segmentation in Images and Videos
by: Wang, Hao, et al.
Published: (2026) -
Automated Detection of Sport Highlights from Audio and Video Sources
by: Della Santa, Francesco, et al.
Published: (2025)