Guardado en:
| Autor principal: | Dufera, Amanuel Tafese |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2509.17365 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
ConfidentSplat: Confidence-Weighted Depth Fusion for Accurate 3D Gaussian Splatting SLAM
por: Dufera, Amanuel T., et al.
Publicado: (2025)
por: Dufera, Amanuel T., et al.
Publicado: (2025)
SEMT: Static-Expansion-Mesh Transformer Network Architecture for Remote Sensing Image Captioning
por: Truong, Khang, et al.
Publicado: (2025)
por: Truong, Khang, et al.
Publicado: (2025)
Automated Image Captioning with CNNs and Transformers
por: Cahyono, Joshua Adrian, et al.
Publicado: (2024)
por: Cahyono, Joshua Adrian, et al.
Publicado: (2024)
GPT-NAS: Evolutionary Neural Architecture Search with the Generative Pre-Trained Model
por: Yu, Caiyang, et al.
Publicado: (2023)
por: Yu, Caiyang, et al.
Publicado: (2023)
MoColl: Agent-Based Specific and General Model Collaboration for Image Captioning
por: Yang, Pu, et al.
Publicado: (2025)
por: Yang, Pu, et al.
Publicado: (2025)
CaptionFool: Universal Image Captioning Model Attacks
por: Parekh, Swapnil
Publicado: (2026)
por: Parekh, Swapnil
Publicado: (2026)
An Ensemble Model with Attention Based Mechanism for Image Captioning
por: Badarneh, Israa Al, et al.
Publicado: (2025)
por: Badarneh, Israa Al, et al.
Publicado: (2025)
Disentangling Fine-Tuning from Pre-Training in Visual Captioning with Hybrid Markov Logic
por: Shah, Monika, et al.
Publicado: (2025)
por: Shah, Monika, et al.
Publicado: (2025)
Towards Retrieval-Augmented Architectures for Image Captioning
por: Sarto, Sara, et al.
Publicado: (2024)
por: Sarto, Sara, et al.
Publicado: (2024)
Explainable Image Captioning using CNN- CNN architecture and Hierarchical Attention
por: Mohan, Rishi Kesav, et al.
Publicado: (2024)
por: Mohan, Rishi Kesav, et al.
Publicado: (2024)
Masked Generative Story Transformer with Character Guidance and Caption Augmentation
por: Papadimitriou, Christos, et al.
Publicado: (2024)
por: Papadimitriou, Christos, et al.
Publicado: (2024)
A Comparative Study of Adversarial Robustness in CNN and CNN-ANFIS Architectures
por: Shankar, Kaaustaaub, et al.
Publicado: (2026)
por: Shankar, Kaaustaaub, et al.
Publicado: (2026)
Generating Accurate and Detailed Captions for High-Resolution Images
por: Lee, Hankyeol, et al.
Publicado: (2025)
por: Lee, Hankyeol, et al.
Publicado: (2025)
Development of CNN Architectures using Transfer Learning Methods for Medical Image Classification
por: Basyal, Ganga Prasad, et al.
Publicado: (2024)
por: Basyal, Ganga Prasad, et al.
Publicado: (2024)
Cross Modification Attention Based Deliberation Model for Image Captioning
por: Lian, Zheng, et al.
Publicado: (2021)
por: Lian, Zheng, et al.
Publicado: (2021)
AGIC: Attention-Guided Image Captioning to Improve Caption Relevance
por: Teja, L. D. M. S. Sai, et al.
Publicado: (2025)
por: Teja, L. D. M. S. Sai, et al.
Publicado: (2025)
Pre-Trained Video Generative Models as World Simulators
por: He, Haoran, et al.
Publicado: (2025)
por: He, Haoran, et al.
Publicado: (2025)
Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models
por: Lai, Zhengfeng, et al.
Publicado: (2024)
por: Lai, Zhengfeng, et al.
Publicado: (2024)
XMeCap: Meme Caption Generation with Sub-Image Adaptability
por: Chen, Yuyan, et al.
Publicado: (2024)
por: Chen, Yuyan, et al.
Publicado: (2024)
Low-hallucination Synthetic Captions for Large-Scale Vision-Language Model Pre-training
por: Zhang, Xinsong, et al.
Publicado: (2025)
por: Zhang, Xinsong, et al.
Publicado: (2025)
Evolving CNN Architectures: From Custom Designs to Deep Residual Models for Diverse Image Classification and Detection Tasks
por: Hasan, Mahmudul, et al.
Publicado: (2026)
por: Hasan, Mahmudul, et al.
Publicado: (2026)
When Better Eyes Lead to Blindness: A Diagnostic Study of the Information Bottleneck in CNN-LSTM Image Captioning Models
por: Gupta, Hitesh Kumar
Publicado: (2025)
por: Gupta, Hitesh Kumar
Publicado: (2025)
Is Your Text-to-Image Model Robust to Caption Noise?
por: Yu, Weichen, et al.
Publicado: (2024)
por: Yu, Weichen, et al.
Publicado: (2024)
How to Train your Text-to-Image Model: Evaluating Design Choices for Synthetic Training Captions
por: Brack, Manuel, et al.
Publicado: (2025)
por: Brack, Manuel, et al.
Publicado: (2025)
Optimizing Gastrointestinal Diagnostics: A CNN-Based Model for VCE Image Classification
por: Ahlawat, Vaneeta, et al.
Publicado: (2024)
por: Ahlawat, Vaneeta, et al.
Publicado: (2024)
MambaOutRS: A Hybrid CNN-Fourier Architecture for Remote Sensing Image Classification
por: Cheon, Minjong, et al.
Publicado: (2025)
por: Cheon, Minjong, et al.
Publicado: (2025)
Regeneration Based Training-free Attribution of Fake Images Generated by Text-to-Image Generative Models
por: Li, Meiling, et al.
Publicado: (2024)
por: Li, Meiling, et al.
Publicado: (2024)
AI-Powered Deepfake Detection Using CNN and Vision Transformer Architectures
por: Urmi, Sifatullah Sheikh, et al.
Publicado: (2026)
por: Urmi, Sifatullah Sheikh, et al.
Publicado: (2026)
Enhancing Image Caption Generation Using Reinforcement Learning with Human Feedback
por: L, Adarsh N, et al.
Publicado: (2024)
por: L, Adarsh N, et al.
Publicado: (2024)
Image Captioning in news report scenario
por: Liu, Tianrui, et al.
Publicado: (2024)
por: Liu, Tianrui, et al.
Publicado: (2024)
Team NYCU at Defactify4: Robust Detection and Source Identification of AI-Generated Images Using CNN and CLIP-Based Models
por: Yang, Tsan-Tsung, et al.
Publicado: (2025)
por: Yang, Tsan-Tsung, et al.
Publicado: (2025)
Any2Caption:Interpreting Any Condition to Caption for Controllable Video Generation
por: Wu, Shengqiong, et al.
Publicado: (2025)
por: Wu, Shengqiong, et al.
Publicado: (2025)
Mining Fine-Grained Image-Text Alignment for Zero-Shot Captioning via Text-Only Training
por: Qiu, Longtian, et al.
Publicado: (2024)
por: Qiu, Longtian, et al.
Publicado: (2024)
good4cir: Generating Detailed Synthetic Captions for Composed Image Retrieval
por: Kolouju, Pranavi, et al.
Publicado: (2025)
por: Kolouju, Pranavi, et al.
Publicado: (2025)
Exploring Annotation-free Image Captioning with Retrieval-augmented Pseudo Sentence Generation
por: Li, Zhiyuan, et al.
Publicado: (2023)
por: Li, Zhiyuan, et al.
Publicado: (2023)
Refining Pre-Trained Motion Models
por: Sun, Xinglong, et al.
Publicado: (2024)
por: Sun, Xinglong, et al.
Publicado: (2024)
Generative Pre-trained Autoregressive Diffusion Transformer
por: Zhang, Yuan, et al.
Publicado: (2025)
por: Zhang, Yuan, et al.
Publicado: (2025)
A Hybrid Fully Convolutional CNN-Transformer Model for Inherently Interpretable Disease Detection from Retinal Fundus Images
por: Djoumessi, Kerol, et al.
Publicado: (2025)
por: Djoumessi, Kerol, et al.
Publicado: (2025)
Image Embedding Sampling Method for Diverse Captioning
por: Waheed, Sania, et al.
Publicado: (2025)
por: Waheed, Sania, et al.
Publicado: (2025)
Top-Down Semantic Refinement for Image Captioning
por: Zhang, Jusheng, et al.
Publicado: (2025)
por: Zhang, Jusheng, et al.
Publicado: (2025)
Ejemplares similares
-
ConfidentSplat: Confidence-Weighted Depth Fusion for Accurate 3D Gaussian Splatting SLAM
por: Dufera, Amanuel T., et al.
Publicado: (2025) -
SEMT: Static-Expansion-Mesh Transformer Network Architecture for Remote Sensing Image Captioning
por: Truong, Khang, et al.
Publicado: (2025) -
Automated Image Captioning with CNNs and Transformers
por: Cahyono, Joshua Adrian, et al.
Publicado: (2024) -
GPT-NAS: Evolutionary Neural Architecture Search with the Generative Pre-Trained Model
por: Yu, Caiyang, et al.
Publicado: (2023) -
MoColl: Agent-Based Specific and General Model Collaboration for Image Captioning
por: Yang, Pu, et al.
Publicado: (2025)