Gespeichert in:
| Hauptverfasser: | Mukherjee, Kushin, Ren, Donghao, Moritz, Dominik, Assogba, Yannick |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2508.04650 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
High-Entropy Tokens as Multimodal Failure Points in Vision-Language Models
von: He, Mengqi, et al.
Veröffentlicht: (2025)
von: He, Mengqi, et al.
Veröffentlicht: (2025)
Efficient Diffusion Training through Parallelization with Truncated Karhunen-Loève Expansion
von: Ren, Yumeng, et al.
Veröffentlicht: (2025)
von: Ren, Yumeng, et al.
Veröffentlicht: (2025)
Fine-Grained Open-Vocabulary Object Detection with Fined-Grained Prompts: Task, Dataset and Benchmark
von: Liu, Ying, et al.
Veröffentlicht: (2025)
von: Liu, Ying, et al.
Veröffentlicht: (2025)
Universal Adversarial Perturbations for Vision-Language Pre-trained Models
von: Zhang, Peng-Fei, et al.
Veröffentlicht: (2024)
von: Zhang, Peng-Fei, et al.
Veröffentlicht: (2024)
Visual Language Models show widespread visual deficits on neuropsychological tests
von: Tangtartharakul, Gene, et al.
Veröffentlicht: (2025)
von: Tangtartharakul, Gene, et al.
Veröffentlicht: (2025)
A Review of Pseudo-Labeling for Computer Vision
von: Kage, Patrick, et al.
Veröffentlicht: (2024)
von: Kage, Patrick, et al.
Veröffentlicht: (2024)
Beyond Routing: Characterising Expert Tuning and Representation in Vision Mixture-of-Experts
von: Tangtartharakul, Gene, et al.
Veröffentlicht: (2026)
von: Tangtartharakul, Gene, et al.
Veröffentlicht: (2026)
Context-Aware Full Body Anonymization using Text-to-Image Diffusion Models
von: Zwick, Pascal, et al.
Veröffentlicht: (2024)
von: Zwick, Pascal, et al.
Veröffentlicht: (2024)
Mechanisms of Prompt-Induced Hallucination in Vision-Language Models
von: Rudman, William, et al.
Veröffentlicht: (2026)
von: Rudman, William, et al.
Veröffentlicht: (2026)
Synthetic Photography Detection: A Visual Guidance for Identifying Synthetic Images Created by AI
von: Mathys, Melanie, et al.
Veröffentlicht: (2024)
von: Mathys, Melanie, et al.
Veröffentlicht: (2024)
MAR-MAER: Metric-Aware and Ambiguity-Adaptive Autoregressive Image Generation
von: Dong, Kai, et al.
Veröffentlicht: (2026)
von: Dong, Kai, et al.
Veröffentlicht: (2026)
Domain Generalized Stereo Matching with Uncertainty-guided Data Augmentation
von: Du, Shuangli, et al.
Veröffentlicht: (2025)
von: Du, Shuangli, et al.
Veröffentlicht: (2025)
Pose Matters: Evaluating Vision Transformers and CNNs for Human Action Recognition on Small COCO Subsets
von: Tang, MingZe, et al.
Veröffentlicht: (2025)
von: Tang, MingZe, et al.
Veröffentlicht: (2025)
CerberusDet: Unified Multi-Dataset Object Detection
von: Tolstykh, Irina, et al.
Veröffentlicht: (2024)
von: Tolstykh, Irina, et al.
Veröffentlicht: (2024)
Towards Interpretable Visual Decoding with Attention to Brain Representations
von: Feng, Pinyuan, et al.
Veröffentlicht: (2025)
von: Feng, Pinyuan, et al.
Veröffentlicht: (2025)
Measuring proximity to standard planes during fetal brain ultrasound scanning
von: Di Vece, Chiara, et al.
Veröffentlicht: (2024)
von: Di Vece, Chiara, et al.
Veröffentlicht: (2024)
Generating Image Adversarial Examples by Embedding Digital Watermarks
von: Xiang, Yuexin, et al.
Veröffentlicht: (2020)
von: Xiang, Yuexin, et al.
Veröffentlicht: (2020)
Quaternion Convolutional Neural Networks: Current Advances and Future Directions
von: Altamirano-Gomez, Gerardo, et al.
Veröffentlicht: (2023)
von: Altamirano-Gomez, Gerardo, et al.
Veröffentlicht: (2023)
MultiHateClip: A Multilingual Benchmark Dataset for Hateful Video Detection on YouTube and Bilibili
von: Wang, Han, et al.
Veröffentlicht: (2024)
von: Wang, Han, et al.
Veröffentlicht: (2024)
Sparse vs Contiguous Adversarial Pixel Perturbations in Multimodal Models: An Empirical Analysis
von: Botocan, Cristian-Alexandru, et al.
Veröffentlicht: (2024)
von: Botocan, Cristian-Alexandru, et al.
Veröffentlicht: (2024)
Generation of Complex 3D Human Motion by Temporal and Spatial Composition of Diffusion Models
von: Mandelli, Lorenzo, et al.
Veröffentlicht: (2024)
von: Mandelli, Lorenzo, et al.
Veröffentlicht: (2024)
Deep EM with Hierarchical Latent Label Modelling for Multi-Site Prostate Lesion Segmentation
von: Yan, Wen, et al.
Veröffentlicht: (2026)
von: Yan, Wen, et al.
Veröffentlicht: (2026)
VACoDe: Visual Augmented Contrastive Decoding
von: Kim, Sihyeon, et al.
Veröffentlicht: (2024)
von: Kim, Sihyeon, et al.
Veröffentlicht: (2024)
Surrealistic-like Image Generation with Vision-Language Models
von: Ayten, Elif, et al.
Veröffentlicht: (2024)
von: Ayten, Elif, et al.
Veröffentlicht: (2024)
Skeleton-based sign language recognition using a dual-stream spatio-temporal dynamic graph convolutional network
von: Liu, Liangjin, et al.
Veröffentlicht: (2025)
von: Liu, Liangjin, et al.
Veröffentlicht: (2025)
Vision Transformer-based Model for Severity Quantification of Lung Pneumonia Using Chest X-ray Images
von: Slika, Bouthaina, et al.
Veröffentlicht: (2023)
von: Slika, Bouthaina, et al.
Veröffentlicht: (2023)
Mechanistically Interpretable Neural Encoding Reveals Fine-Grained Functional Selectivity in Human Visual Cortex
von: Grosbard, Idan Daniel, et al.
Veröffentlicht: (2026)
von: Grosbard, Idan Daniel, et al.
Veröffentlicht: (2026)
Enhancing Eye Feature Estimation from Event Data Streams through Adaptive Inference State Space Modeling
von: Nguyen, Viet Dung, et al.
Veröffentlicht: (2026)
von: Nguyen, Viet Dung, et al.
Veröffentlicht: (2026)
Generative inpainting of incomplete Euclidean distance matrices of trajectories generated by a fractional Brownian motion
von: Lobashev, Alexander, et al.
Veröffentlicht: (2024)
von: Lobashev, Alexander, et al.
Veröffentlicht: (2024)
Exploring Transfer Learning for Deep Learning Polyp Detection in Colonoscopy Images Using YOLOv8
von: Vazquez, Fabian, et al.
Veröffentlicht: (2025)
von: Vazquez, Fabian, et al.
Veröffentlicht: (2025)
CoMA: Complementary Masking and Hierarchical Dynamic Multi-Window Self-Attention in a Unified Pre-training Framework
von: Li, Jiaxuan, et al.
Veröffentlicht: (2025)
von: Li, Jiaxuan, et al.
Veröffentlicht: (2025)
SeNeDiF-OOD: Semantic Nested Dichotomy Fusion for Out-of-Distribution Detection Methodology in Open-World Classification. A Case Study on Monument Style Classification
von: Antequera-Sánchez, Ignacio, et al.
Veröffentlicht: (2026)
von: Antequera-Sánchez, Ignacio, et al.
Veröffentlicht: (2026)
Efficient Neural Network Encoding for 3D Color Lookup Tables
von: Zehtab, Vahid, et al.
Veröffentlicht: (2024)
von: Zehtab, Vahid, et al.
Veröffentlicht: (2024)
Deep Domain Adaptation: A Sim2Real Neural Approach for Improving Eye-Tracking Systems
von: Nguyen, Viet Dung, et al.
Veröffentlicht: (2024)
von: Nguyen, Viet Dung, et al.
Veröffentlicht: (2024)
Parameterizing Dataset Distillation via Gaussian Splatting
von: Jiang, Chenyang, et al.
Veröffentlicht: (2025)
von: Jiang, Chenyang, et al.
Veröffentlicht: (2025)
Eye-gaze Guided Multi-modal Alignment for Medical Representation Learning
von: Ma, Chong, et al.
Veröffentlicht: (2024)
von: Ma, Chong, et al.
Veröffentlicht: (2024)
Beyond Specialization: Assessing the Capabilities of MLLMs in Age and Gender Estimation
von: Kuprashevich, Maksim, et al.
Veröffentlicht: (2024)
von: Kuprashevich, Maksim, et al.
Veröffentlicht: (2024)
Synthetic Image Generation in Cyber Influence Operations: An Emergent Threat?
von: Mathys, Melanie, et al.
Veröffentlicht: (2024)
von: Mathys, Melanie, et al.
Veröffentlicht: (2024)
FUTURE-AI: International consensus guideline for trustworthy and deployable artificial intelligence in healthcare
von: Lekadir, Karim, et al.
Veröffentlicht: (2023)
von: Lekadir, Karim, et al.
Veröffentlicht: (2023)
Lookism: The overlooked bias in computer vision
von: Gulati, Aditya, et al.
Veröffentlicht: (2024)
von: Gulati, Aditya, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
High-Entropy Tokens as Multimodal Failure Points in Vision-Language Models
von: He, Mengqi, et al.
Veröffentlicht: (2025) -
Efficient Diffusion Training through Parallelization with Truncated Karhunen-Loève Expansion
von: Ren, Yumeng, et al.
Veröffentlicht: (2025) -
Fine-Grained Open-Vocabulary Object Detection with Fined-Grained Prompts: Task, Dataset and Benchmark
von: Liu, Ying, et al.
Veröffentlicht: (2025) -
Universal Adversarial Perturbations for Vision-Language Pre-trained Models
von: Zhang, Peng-Fei, et al.
Veröffentlicht: (2024) -
Visual Language Models show widespread visual deficits on neuropsychological tests
von: Tangtartharakul, Gene, et al.
Veröffentlicht: (2025)