Saved in:
| Main Authors: | Krotov, Aleksei, Tebo, Alison, Picart, Dylan K., Algave, Aaron Dean |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2409.09560 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Object-oriented backdoor attack against image captioning
by: Li, Meiling, et al.
Published: (2024)
by: Li, Meiling, et al.
Published: (2024)
Semantic search for 100M+ galaxy images using AI-generated captions
by: Koblischke, Nolan, et al.
Published: (2025)
by: Koblischke, Nolan, et al.
Published: (2025)
Cognitive resilience: Unraveling the proficiency of image-captioning models to interpret masked visual content
by: Du, Zhicheng, et al.
Published: (2024)
by: Du, Zhicheng, et al.
Published: (2024)
COIN: Counterfactual inpainting for weakly supervised semantic segmentation for medical images
by: Shvetsov, Dmytro, et al.
Published: (2024)
by: Shvetsov, Dmytro, et al.
Published: (2024)
Multi-Modal interpretable automatic video captioning
by: Hanna-Asaad, Antoine, et al.
Published: (2024)
by: Hanna-Asaad, Antoine, et al.
Published: (2024)
Attention-based transformer models for image captioning across languages: An in-depth survey and evaluation
by: Albadarneh, Israa A., et al.
Published: (2025)
by: Albadarneh, Israa A., et al.
Published: (2025)
Knowledge distillation to effectively attain both region-of-interest and global semantics from an image where multiple objects appear
by: Jin, Seonwhee
Published: (2024)
by: Jin, Seonwhee
Published: (2024)
Image captioning for Brazilian Portuguese using GRIT model
by: de Alencar, Rafael Silva, et al.
Published: (2024)
by: de Alencar, Rafael Silva, et al.
Published: (2024)
Of-SemWat: High-payload text embedding for semantic watermarking of AI-generated images with arbitrary size
by: Tondi, Benedetta, et al.
Published: (2025)
by: Tondi, Benedetta, et al.
Published: (2025)
VQArt-Bench: A semantically rich VQA Benchmark for Art and Cultural Heritage
by: Alfarano, A., et al.
Published: (2025)
by: Alfarano, A., et al.
Published: (2025)
Correlation of Object Detection Performance with Visual Saliency and Depth Estimation
by: Bartolo, Matthias, et al.
Published: (2024)
by: Bartolo, Matthias, et al.
Published: (2024)
Multi-style conversion for semantic segmentation of lesions in fundus images by adversarial attacks
by: Playout, Clément, et al.
Published: (2024)
by: Playout, Clément, et al.
Published: (2024)
Criteria-first, semantics-later: reproducible structure discovery in image-based sciences
by: Bumberger, Jan
Published: (2026)
by: Bumberger, Jan
Published: (2026)
Integrating Saliency Ranking and Reinforcement Learning for Enhanced Object Detection
by: Bartolo, Matthias, et al.
Published: (2024)
by: Bartolo, Matthias, et al.
Published: (2024)
Good at captioning, bad at counting: Benchmarking GPT-4V on Earth observation data
by: Zhang, Chenhui, et al.
Published: (2024)
by: Zhang, Chenhui, et al.
Published: (2024)
VeCLIP: Improving CLIP Training via Visual-enriched Captions
by: Lai, Zhengfeng, et al.
Published: (2023)
by: Lai, Zhengfeng, et al.
Published: (2023)
MGHanD: Multi-modal Guidance for authentic Hand Diffusion
by: Eum, Taehyeon, et al.
Published: (2025)
by: Eum, Taehyeon, et al.
Published: (2025)
Comprehensive framework for evaluation of deep neural networks in detection and quantification of lymphoma from PET/CT images: clinical insights, pitfalls, and observer agreement analyses
by: Ahamed, Shadab, et al.
Published: (2023)
by: Ahamed, Shadab, et al.
Published: (2023)
Improving face generation quality and prompt following with synthetic captions
by: Tarasiou, Michail, et al.
Published: (2024)
by: Tarasiou, Michail, et al.
Published: (2024)
Evaluating alignment between humans and neural network representations in image-based learning tasks
by: Demircan, Can, et al.
Published: (2023)
by: Demircan, Can, et al.
Published: (2023)
Automated detection of underdiagnosed medical conditions via opportunistic imaging
by: Aali, Asad, et al.
Published: (2024)
by: Aali, Asad, et al.
Published: (2024)
Training-free Detection of AI-generated images via Cropping Robustness
by: Choi, Sungik, et al.
Published: (2025)
by: Choi, Sungik, et al.
Published: (2025)
Found in Translation: semantic approaches for enhancing AI interpretability in face verification
by: Doh, Miriam, et al.
Published: (2025)
by: Doh, Miriam, et al.
Published: (2025)
F$^3$OCUS -- Federated Finetuning of Vision-Language Foundation Models with Optimal Client Layer Updating Strategy via Multi-objective Meta-Heuristics
by: Saha, Pramit, et al.
Published: (2024)
by: Saha, Pramit, et al.
Published: (2024)
Experience-Guided Self-Adaptive Cascaded Agents for Breast Cancer Screening and Diagnosis with Reduced Biopsy Referrals
by: Saha, Pramit, et al.
Published: (2026)
by: Saha, Pramit, et al.
Published: (2026)
Better Understanding Differences in Attribution Methods via Systematic Evaluations
by: Rao, Sukrut, et al.
Published: (2023)
by: Rao, Sukrut, et al.
Published: (2023)
SCENES: Subpixel Correspondence Estimation With Epipolar Supervision
by: Kloepfer, Dominik A., et al.
Published: (2024)
by: Kloepfer, Dominik A., et al.
Published: (2024)
Towards a multimodal framework for remote sensing image change retrieval and captioning
by: Ferrod, Roger, et al.
Published: (2024)
by: Ferrod, Roger, et al.
Published: (2024)
Automated Model Evaluation for Object Detection via Prediction Consistency and Reliability
by: Yoo, Seungju, et al.
Published: (2025)
by: Yoo, Seungju, et al.
Published: (2025)
MoLE: Enhancing Human-centric Text-to-image Diffusion via Mixture of Low-rank Experts
by: Zhu, Jie, et al.
Published: (2024)
by: Zhu, Jie, et al.
Published: (2024)
CTARR: A fast and robust method for identifying anatomical regions on CT images via atlas registration
by: Buddenkotte, Thomas, et al.
Published: (2024)
by: Buddenkotte, Thomas, et al.
Published: (2024)
Hearing Touch: Audio-Visual Pretraining for Contact-Rich Manipulation
by: Mejia, Jared, et al.
Published: (2024)
by: Mejia, Jared, et al.
Published: (2024)
FedPIA -- Permuting and Integrating Adapters leveraging Wasserstein Barycenters for Finetuning Foundation Models in Multi-Modal Federated Learning
by: Saha, Pramit, et al.
Published: (2024)
by: Saha, Pramit, et al.
Published: (2024)
Orion: A Unified Visual Agent for Multimodal Perception, Advanced Visual Reasoning and Execution
by: Reddy, N Dinesh, et al.
Published: (2025)
by: Reddy, N Dinesh, et al.
Published: (2025)
Compositional Discrete Latent Code for High Fidelity, Productive Diffusion Models
by: Lavoie, Samuel, et al.
Published: (2025)
by: Lavoie, Samuel, et al.
Published: (2025)
SoftPQ: Robust Instance Segmentation Evaluation via Soft Matching and Tunable Thresholds
by: Karmakar, Ranit, et al.
Published: (2025)
by: Karmakar, Ranit, et al.
Published: (2025)
Disentangled representations of microscopy images
by: Dapueto, Jacopo, et al.
Published: (2025)
by: Dapueto, Jacopo, et al.
Published: (2025)
GRADEO: Towards Human-Like Evaluation for Text-to-Video Generation via Multi-Step Reasoning
by: Mou, Zhun, et al.
Published: (2025)
by: Mou, Zhun, et al.
Published: (2025)
DepthSeg: Depth prompting in remote sensing semantic segmentation
by: Zhou, Ning, et al.
Published: (2025)
by: Zhou, Ning, et al.
Published: (2025)
KAGE-Bench: Fast Known-Axis Visual Generalization Evaluation for Reinforcement Learning
by: Cherepanov, Egor, et al.
Published: (2026)
by: Cherepanov, Egor, et al.
Published: (2026)
Similar Items
-
Object-oriented backdoor attack against image captioning
by: Li, Meiling, et al.
Published: (2024) -
Semantic search for 100M+ galaxy images using AI-generated captions
by: Koblischke, Nolan, et al.
Published: (2025) -
Cognitive resilience: Unraveling the proficiency of image-captioning models to interpret masked visual content
by: Du, Zhicheng, et al.
Published: (2024) -
COIN: Counterfactual inpainting for weakly supervised semantic segmentation for medical images
by: Shvetsov, Dmytro, et al.
Published: (2024) -
Multi-Modal interpretable automatic video captioning
by: Hanna-Asaad, Antoine, et al.
Published: (2024)