Gespeichert in:
| Hauptverfasser: | Xie, Wen, Zhu, Yanjun, Overgoor, Gijs, Bart, Yakov, Garcia, Agata Lapedriza, Ostadabbas, Sarah |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2510.26569 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Bridging Knowledge Gap Between Image Inpainting and Large-Area Visible Watermark Removal
von: Leng, Yicheng, et al.
Veröffentlicht: (2025)
von: Leng, Yicheng, et al.
Veröffentlicht: (2025)
A Roadmap for Multilingual, Multimodal Domain Independent Deception Detection
von: Boumber, Dainis, et al.
Veröffentlicht: (2024)
von: Boumber, Dainis, et al.
Veröffentlicht: (2024)
U-Net-Like Spiking Neural Networks for Single Image Dehazing
von: Li, Huibin, et al.
Veröffentlicht: (2025)
von: Li, Huibin, et al.
Veröffentlicht: (2025)
AVControl: Efficient Framework for Training Audio-Visual Controls
von: Ben-Yosef, Matan, et al.
Veröffentlicht: (2026)
von: Ben-Yosef, Matan, et al.
Veröffentlicht: (2026)
Visual Style Prompt Learning Using Diffusion Models for Blind Face Restoration
von: Lu, Wanglong, et al.
Veröffentlicht: (2024)
von: Lu, Wanglong, et al.
Veröffentlicht: (2024)
FACEMUG: A Multimodal Generative and Fusion Framework for Local Facial Editing
von: Lu, Wanglong, et al.
Veröffentlicht: (2024)
von: Lu, Wanglong, et al.
Veröffentlicht: (2024)
Leum-VL Technical Report
von: He, Yuxuan, et al.
Veröffentlicht: (2026)
von: He, Yuxuan, et al.
Veröffentlicht: (2026)
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
von: Raoufi, Behnam, et al.
Veröffentlicht: (2025)
von: Raoufi, Behnam, et al.
Veröffentlicht: (2025)
FLD+: Data-efficient Evaluation Metric for Generative Models
von: Jeevan, Pranav, et al.
Veröffentlicht: (2024)
von: Jeevan, Pranav, et al.
Veröffentlicht: (2024)
WaveMixSR-V2: Enhancing Super-resolution with Higher Efficiency
von: Jeevan, Pranav, et al.
Veröffentlicht: (2024)
von: Jeevan, Pranav, et al.
Veröffentlicht: (2024)
Normalizing Flow-Based Metric for Image Generation
von: Jeevan, Pranav, et al.
Veröffentlicht: (2024)
von: Jeevan, Pranav, et al.
Veröffentlicht: (2024)
Geo2Sound: A Scalable Geo-Aligned Framework for Soundscape Generation from Satellite Imagery
von: Wu, Kunlin, et al.
Veröffentlicht: (2026)
von: Wu, Kunlin, et al.
Veröffentlicht: (2026)
A Hybrid Deterministic Framework for Named Entity Extraction in Broadcast News Video
von: Lucas, Andrea Filiberto, et al.
Veröffentlicht: (2026)
von: Lucas, Andrea Filiberto, et al.
Veröffentlicht: (2026)
Learning Joint Denoising, Demosaicing, and Compression from the Raw Natural Image Noise Dataset
von: Brummer, Benoit, et al.
Veröffentlicht: (2025)
von: Brummer, Benoit, et al.
Veröffentlicht: (2025)
Light Future: Multimodal Action Frame Prediction via InstructPix2Pix
von: Zhong, Zesen, et al.
Veröffentlicht: (2025)
von: Zhong, Zesen, et al.
Veröffentlicht: (2025)
AIM 2024 Challenge on Video Saliency Prediction: Methods and Results
von: Moskalenko, Andrey, et al.
Veröffentlicht: (2024)
von: Moskalenko, Andrey, et al.
Veröffentlicht: (2024)
NTIRE 2026 Challenge on Video Saliency Prediction: Methods and Results
von: Moskalenko, Andrey, et al.
Veröffentlicht: (2026)
von: Moskalenko, Andrey, et al.
Veröffentlicht: (2026)
Phase-Aware Wavelet-Based-Scattering Encoder-Decoder for Dense Predictions
von: Marrakchi, Ghassen, et al.
Veröffentlicht: (2026)
von: Marrakchi, Ghassen, et al.
Veröffentlicht: (2026)
Scene Detection Policies and Keyframe Extraction Strategies for Large-Scale Video Analysis
von: Korolkov, Vasilii
Veröffentlicht: (2025)
von: Korolkov, Vasilii
Veröffentlicht: (2025)
EDSNet: Efficient-DSNet for Video Summarization
von: Prasad, Ashish, et al.
Veröffentlicht: (2024)
von: Prasad, Ashish, et al.
Veröffentlicht: (2024)
A Real-Time Diminished Reality Approach to Privacy in MR Collaboration
von: Fane, Christian
Veröffentlicht: (2025)
von: Fane, Christian
Veröffentlicht: (2025)
PCRI: Measuring Context Robustness in Multimodal Models for Enterprise Applications
von: Patel, Hitesh Laxmichand, et al.
Veröffentlicht: (2025)
von: Patel, Hitesh Laxmichand, et al.
Veröffentlicht: (2025)
Graph-PiT: Enhancing Structural Coherence in Part-Based Image Synthesis via Graph Priors
von: Zhang, Junbin, et al.
Veröffentlicht: (2026)
von: Zhang, Junbin, et al.
Veröffentlicht: (2026)
MetaErr: Towards Predicting Error Patterns in Deep Neural Networks
von: Totakura, Varun, et al.
Veröffentlicht: (2026)
von: Totakura, Varun, et al.
Veröffentlicht: (2026)
CCVA-FL: Cross-Client Variations Adaptive Federated Learning for Medical Imaging
von: Gupta, Sunny, et al.
Veröffentlicht: (2024)
von: Gupta, Sunny, et al.
Veröffentlicht: (2024)
Taming the Tail: Leveraging Asymmetric Loss and Pade Approximation to Overcome Medical Image Long-Tailed Class Imbalance
von: Kashyap, Pankhi, et al.
Veröffentlicht: (2024)
von: Kashyap, Pankhi, et al.
Veröffentlicht: (2024)
HY-Himmel Technical Report: Hierarchical Interleaved Multi-stream Motion Encoding for Long Video Understanding
von: Jin, Haopeng, et al.
Veröffentlicht: (2026)
von: Jin, Haopeng, et al.
Veröffentlicht: (2026)
Digital analysis of early color photographs taken using regular color screen processes
von: Hubička, Jan, et al.
Veröffentlicht: (2023)
von: Hubička, Jan, et al.
Veröffentlicht: (2023)
Semantic2Graph: Graph-based Multi-modal Feature Fusion for Action Segmentation in Videos
von: Zhang, Junbin, et al.
Veröffentlicht: (2022)
von: Zhang, Junbin, et al.
Veröffentlicht: (2022)
RCI: A Score for Evaluating Global and Local Reasoning in Multimodal Benchmarks
von: Agarwal, Amit, et al.
Veröffentlicht: (2025)
von: Agarwal, Amit, et al.
Veröffentlicht: (2025)
Efficient and Privacy-Protecting Background Removal for 2D Video Streaming using iPhone 15 Pro Max LiDAR
von: Kinnevan, Jessica, et al.
Veröffentlicht: (2025)
von: Kinnevan, Jessica, et al.
Veröffentlicht: (2025)
ForensicFormer: Hierarchical Multi-Scale Reasoning for Cross-Domain Image Forgery Detection
von: Samson, Hema Hariharan
Veröffentlicht: (2026)
von: Samson, Hema Hariharan
Veröffentlicht: (2026)
Evaluation Metric for Quality Control and Generative Models in Histopathology Images
von: Jeevan, Pranav, et al.
Veröffentlicht: (2024)
von: Jeevan, Pranav, et al.
Veröffentlicht: (2024)
Development of ultra-high efficiency soft X-ray angle-resolved photoemission spectroscopy equipped with deep prior-based denoising method
von: Yamagami, Kohei, et al.
Veröffentlicht: (2025)
von: Yamagami, Kohei, et al.
Veröffentlicht: (2025)
Automatic Detection of Intro and Credits in Video using CLIP and Multihead Attention
von: Korolkov, Vasilii, et al.
Veröffentlicht: (2025)
von: Korolkov, Vasilii, et al.
Veröffentlicht: (2025)
DSCSNet: A Dynamic Sparse Compression Sensing Network for Closely-Spaced Infrared Small Target Unmixing
von: Tang, Zhiyang, et al.
Veröffentlicht: (2026)
von: Tang, Zhiyang, et al.
Veröffentlicht: (2026)
WaveMix: A Resource-efficient Neural Network for Image Analysis
von: Jeevan, Pranav, et al.
Veröffentlicht: (2022)
von: Jeevan, Pranav, et al.
Veröffentlicht: (2022)
Which Backbone to Use: A Resource-efficient Domain Specific Comparison for Computer Vision
von: Jeevan, Pranav, et al.
Veröffentlicht: (2024)
von: Jeevan, Pranav, et al.
Veröffentlicht: (2024)
Image and Video Compression using Generative Sparse Representation with Fidelity Controls
von: Jiang, Wei, et al.
Veröffentlicht: (2024)
von: Jiang, Wei, et al.
Veröffentlicht: (2024)
Image-Based Leopard Seal Recognition: Approaches and Challenges in Current Automated Systems
von: Salazar, Jorge Yero, et al.
Veröffentlicht: (2024)
von: Salazar, Jorge Yero, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Bridging Knowledge Gap Between Image Inpainting and Large-Area Visible Watermark Removal
von: Leng, Yicheng, et al.
Veröffentlicht: (2025) -
A Roadmap for Multilingual, Multimodal Domain Independent Deception Detection
von: Boumber, Dainis, et al.
Veröffentlicht: (2024) -
U-Net-Like Spiking Neural Networks for Single Image Dehazing
von: Li, Huibin, et al.
Veröffentlicht: (2025) -
AVControl: Efficient Framework for Training Audio-Visual Controls
von: Ben-Yosef, Matan, et al.
Veröffentlicht: (2026) -
Visual Style Prompt Learning Using Diffusion Models for Blind Face Restoration
von: Lu, Wanglong, et al.
Veröffentlicht: (2024)