Analyzing Transformer Models and Knowledge Distillation Approaches for Image Captioning on Edge AI
Fuente:
arXiv
Saved in:
| Main Authors: | Kwok, Wing Man Casca, Tung, Yip Chiu, Bhagchandani, Kunal |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
CAPEEN: Image Captioning with Early Exits and Knowledge Distillation
by: Bajpai, Divya Jyoti, et al.
Published: (2024)
by: Bajpai, Divya Jyoti, et al.
Published: (2024)
Edge-Efficient Image Restoration: Transformer Distillation into State-Space Models
by: Miriyala, Srinivas Soumitri, et al.
Published: (2026)
by: Miriyala, Srinivas Soumitri, et al.
Published: (2026)
A Novel Lightweight Transformer with Edge-Aware Fusion for Remote Sensing Image Captioning
by: Das, Swadhin, et al.
Published: (2025)
by: Das, Swadhin, et al.
Published: (2025)
Compositional Oil Spill Detection Based on Object Detector and Adapted Segment Anything Model from SAR Images
by: Wu, Wenhui, et al.
Published: (2024)
by: Wu, Wenhui, et al.
Published: (2024)
A Transformer-in-Transformer Network Utilizing Knowledge Distillation for Image Recognition
by: Rahman, Dewan Tauhid, et al.
Published: (2025)
by: Rahman, Dewan Tauhid, et al.
Published: (2025)
UIT-DarkCow team at ImageCLEFmedical Caption 2024: Diagnostic Captioning for Radiology Images Efficiency with Transformer Models
by: Van Nguyen, Quan, et al.
Published: (2024)
by: Van Nguyen, Quan, et al.
Published: (2024)
Token Compression Meets Compact Vision Transformers: A Survey and Comparative Evaluation for Edge AI
by: Nguyen, Phat, et al.
Published: (2025)
by: Nguyen, Phat, et al.
Published: (2025)
An Edge AI System Based on FPGA Platform for Railway Fault Detection
by: Li, Jiale, et al.
Published: (2024)
by: Li, Jiale, et al.
Published: (2024)
Efficient Few-Shot Learning for Edge AI via Knowledge Distillation on MobileViT
by: Tsuyuki, Shuhei, et al.
Published: (2026)
by: Tsuyuki, Shuhei, et al.
Published: (2026)
Where Do Images Come From? Analyzing Captions to Geographically Profile Datasets
by: Basu, Abhipsa, et al.
Published: (2026)
by: Basu, Abhipsa, et al.
Published: (2026)
Dual-Stream Collaborative Transformer for Image Captioning
by: Wan, Jun, et al.
Published: (2026)
by: Wan, Jun, et al.
Published: (2026)
Efficient Knowledge Distillation of SAM for Medical Image Segmentation
by: Patil, Kunal Dasharath, et al.
Published: (2025)
by: Patil, Kunal Dasharath, et al.
Published: (2025)
Analyzing Image Beyond Visual Aspect: Image Emotion Classification via Multiple-Affective Captioning
by: Zhou, Zibo, et al.
Published: (2025)
by: Zhou, Zibo, et al.
Published: (2025)
Image Generation from Image Captioning -- Invertible Approach
by: Menon, Nandakishore S, et al.
Published: (2024)
by: Menon, Nandakishore S, et al.
Published: (2024)
Beam-Guided Knowledge Replay for Knowledge-Rich Image Captioning using Vision-Language Model
by: AlJunaid, Reem, et al.
Published: (2025)
by: AlJunaid, Reem, et al.
Published: (2025)
CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era
by: Cheng, Kanzhi, et al.
Published: (2025)
by: Cheng, Kanzhi, et al.
Published: (2025)
Shifted Window Fourier Transform And Retention For Image Captioning
by: Hu, Jia Cheng, et al.
Published: (2024)
by: Hu, Jia Cheng, et al.
Published: (2024)
Towards Optimal Trade-offs in Knowledge Distillation for CNNs and Vision Transformers at the Edge
by: Violos, John, et al.
Published: (2024)
by: Violos, John, et al.
Published: (2024)
Automated Image Captioning with CNNs and Transformers
by: Cahyono, Joshua Adrian, et al.
Published: (2024)
by: Cahyono, Joshua Adrian, et al.
Published: (2024)
Knowledge Distillation via the Target-aware Transformer
by: Lin, Sihao, et al.
Published: (2022)
by: Lin, Sihao, et al.
Published: (2022)
CustomKD: Customizing Large Vision Foundation for Edge Model Improvement via Knowledge Distillation
by: Lee, Jungsoo, et al.
Published: (2025)
by: Lee, Jungsoo, et al.
Published: (2025)
EdgeGaussians -- 3D Edge Mapping via Gaussian Splatting
by: Chelani, Kunal, et al.
Published: (2024)
by: Chelani, Kunal, et al.
Published: (2024)
Context-aware Difference Distilling for Multi-change Captioning
by: Tu, Yunbin, et al.
Published: (2024)
by: Tu, Yunbin, et al.
Published: (2024)
CaptionQA: Is Your Caption as Useful as the Image Itself?
by: Yang, Shijia, et al.
Published: (2025)
by: Yang, Shijia, et al.
Published: (2025)
KD-DETR: Knowledge Distillation for Detection Transformer with Consistent Distillation Points Sampling
by: Wang, Yu, et al.
Published: (2022)
by: Wang, Yu, et al.
Published: (2022)
Point Clouds Are Specialized Images: A Knowledge Transfer Approach for 3D Understanding
by: Kang, Jiachen, et al.
Published: (2023)
by: Kang, Jiachen, et al.
Published: (2023)
Knowledge Distillation in Vision Transformers: A Critical Review
by: Habib, Gousia, et al.
Published: (2023)
by: Habib, Gousia, et al.
Published: (2023)
Knowledge Distillation via Query Selection for Detection Transformer
by: Liu, Yi, et al.
Published: (2024)
by: Liu, Yi, et al.
Published: (2024)
Transformer based Multitask Learning for Image Captioning and Object Detection
by: Basak, Debolena, et al.
Published: (2024)
by: Basak, Debolena, et al.
Published: (2024)
Caption-Matching: A Multimodal Approach for Cross-Domain Image Retrieval
by: Iijima, Lucas, et al.
Published: (2024)
by: Iijima, Lucas, et al.
Published: (2024)
CaptionFool: Universal Image Captioning Model Attacks
by: Parekh, Swapnil
Published: (2026)
by: Parekh, Swapnil
Published: (2026)
Visual-Language Model Knowledge Distillation Method for Image Quality Assessment
by: Hou, Yongkang, et al.
Published: (2025)
by: Hou, Yongkang, et al.
Published: (2025)
HiSem: Hierarchical Semantic Disentangling for Remote Sensing Image Change Captioning
by: Wang, Man, et al.
Published: (2026)
by: Wang, Man, et al.
Published: (2026)
CaptionSmiths: Flexibly Controlling Language Pattern in Image Captioning
by: Saito, Kuniaki, et al.
Published: (2025)
by: Saito, Kuniaki, et al.
Published: (2025)
A Lightweight Sparse Focus Transformer for Remote Sensing Image Change Captioning
by: Sun, Dongwei, et al.
Published: (2024)
by: Sun, Dongwei, et al.
Published: (2024)
DEVICE: Depth and Visual Concepts Aware Transformer for OCR-based Image Captioning
by: Xu, Dongsheng, et al.
Published: (2023)
by: Xu, Dongsheng, et al.
Published: (2023)
Embedded Heterogeneous Attention Transformer for Cross-lingual Image Captioning
by: Song, Zijie, et al.
Published: (2023)
by: Song, Zijie, et al.
Published: (2023)
AI-KD: Towards Alignment Invariant Face Image Quality Assessment Using Knowledge Distillation
by: Babnik, Žiga, et al.
Published: (2024)
by: Babnik, Žiga, et al.
Published: (2024)
m2mKD: Module-to-Module Knowledge Distillation for Modular Transformers
by: Lo, Ka Man, et al.
Published: (2024)
by: Lo, Ka Man, et al.
Published: (2024)
Visually-Aware Context Modeling for News Image Captioning
by: Qu, Tingyu, et al.
Published: (2023)
by: Qu, Tingyu, et al.
Published: (2023)
Similar Items
-
CAPEEN: Image Captioning with Early Exits and Knowledge Distillation
by: Bajpai, Divya Jyoti, et al.
Published: (2024) -
Edge-Efficient Image Restoration: Transformer Distillation into State-Space Models
by: Miriyala, Srinivas Soumitri, et al.
Published: (2026) -
A Novel Lightweight Transformer with Edge-Aware Fusion for Remote Sensing Image Captioning
by: Das, Swadhin, et al.
Published: (2025) -
Compositional Oil Spill Detection Based on Object Detector and Adapted Segment Anything Model from SAR Images
by: Wu, Wenhui, et al.
Published: (2024) -
A Transformer-in-Transformer Network Utilizing Knowledge Distillation for Image Recognition
by: Rahman, Dewan Tauhid, et al.
Published: (2025)