LLM as a Neural Architect: Controlled Generation of Image Captioning Models Under Strict API Contracts
Fuente:
arXiv
Saved in:
| Main Authors: | Jesani, Krunal, Ignatov, Dmitry, Timofte, Radu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Delta-Based Neural Architecture Search: LLM Fine-Tuning via Code Diffs
by: Adhikari, Santosh Premi, et al.
Published: (2026)
by: Adhikari, Santosh Premi, et al.
Published: (2026)
Enhancing LLM-Based Neural Network Generation: Few-Shot Prompting and Efficient Validation for Automated Architecture Design
by: Duvvuri, Raghuvir, et al.
Published: (2025)
by: Duvvuri, Raghuvir, et al.
Published: (2025)
VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs?
by: Gado, Mohamed, et al.
Published: (2025)
by: Gado, Mohamed, et al.
Published: (2025)
MobileAgeNet: Lightweight Facial Age Estimation for Mobile Deployment
by: Kumar, Arun, et al.
Published: (2026)
by: Kumar, Arun, et al.
Published: (2026)
From Memorization to Creativity: LLM as a Designer of Novel Neural Architectures
by: Khalid, Waleed, et al.
Published: (2026)
by: Khalid, Waleed, et al.
Published: (2026)
Closed-Loop LLM Discovery of Non-Standard Channel Priors in Vision Models
by: Uzun, Tolgay Atinc, et al.
Published: (2026)
by: Uzun, Tolgay Atinc, et al.
Published: (2026)
A Retrieval-Augmented Generation Approach to Extracting Algorithmic Logic from Neural Networks
by: Khalid, Waleed, et al.
Published: (2025)
by: Khalid, Waleed, et al.
Published: (2025)
Virtually Enriched NYU Depth V2 Dataset for Monocular Depth Estimation: Do We Need Artificial Augmentation?
by: Ignatov, Dmitry, et al.
Published: (2024)
by: Ignatov, Dmitry, et al.
Published: (2024)
From Code to Prediction: Fine-Tuning LLMs for Neural Network Performance Classification in NNGPT
by: Hanouneh, Mahmoud, et al.
Published: (2026)
by: Hanouneh, Mahmoud, et al.
Published: (2026)
AugmentGest: Can Random Data Cropping Augmentation Boost Gesture Recognition Performance?
by: Aboudeshish, Nada, et al.
Published: (2025)
by: Aboudeshish, Nada, et al.
Published: (2025)
Preparation of Fractal-Inspired Computational Architectures for Advanced Large Language Model Analysis
by: Mittal, Yash, et al.
Published: (2025)
by: Mittal, Yash, et al.
Published: (2025)
e5-omni: Explicit Cross-modal Alignment for Omni-modal Embeddings
by: Chen, Haonan, et al.
Published: (2026)
by: Chen, Haonan, et al.
Published: (2026)
MuDreamer: Learning Predictive World Models without Reconstruction
by: Burchi, Maxime, et al.
Published: (2024)
by: Burchi, Maxime, et al.
Published: (2024)
VELA: An LLM-Hybrid-as-a-Judge Approach for Evaluating Long Image Captions
by: Matsuda, Kazuki, et al.
Published: (2025)
by: Matsuda, Kazuki, et al.
Published: (2025)
From Brute Force to Semantic Insight: Performance-Guided Data Transformation Design with LLMs
by: Shrestha, Usha, et al.
Published: (2026)
by: Shrestha, Usha, et al.
Published: (2026)
Text-only Synthesis for Image Captioning
by: Zhou, Qing, et al.
Published: (2024)
by: Zhou, Qing, et al.
Published: (2024)
The Role of Data Curation in Image Captioning
by: Li, Wenyan, et al.
Published: (2023)
by: Li, Wenyan, et al.
Published: (2023)
Explaining Caption-Image Interactions in CLIP Models with Second-Order Attributions
by: Möller, Lucas, et al.
Published: (2024)
by: Möller, Lucas, et al.
Published: (2024)
The Power of Many: Multi-Agent Multimodal Models for Cultural Image Captioning
by: Bai, Longju, et al.
Published: (2024)
by: Bai, Longju, et al.
Published: (2024)
Learned Lightweight Smartphone ISP with Unpaired Data
by: Arhire, Andrei, et al.
Published: (2025)
by: Arhire, Andrei, et al.
Published: (2025)
WsiCaption: Multiple Instance Generation of Pathology Reports for Gigapixel Whole-Slide Images
by: Chen, Pingyi, et al.
Published: (2023)
by: Chen, Pingyi, et al.
Published: (2023)
CAPEEN: Image Captioning with Early Exits and Knowledge Distillation
by: Bajpai, Divya Jyoti, et al.
Published: (2024)
by: Bajpai, Divya Jyoti, et al.
Published: (2024)
CIC: A Framework for Culturally-Aware Image Captioning
by: Yun, Youngsik, et al.
Published: (2024)
by: Yun, Youngsik, et al.
Published: (2024)
LAViTeR: Learning Aligned Visual and Textual Representations Assisted by Image and Caption Generation
by: Hashemi, Mohammad Abuzar, et al.
Published: (2021)
by: Hashemi, Mohammad Abuzar, et al.
Published: (2021)
AC-Lite : A Lightweight Image Captioning Model for Low-Resource Assamese Language
by: Choudhury, Pankaj, et al.
Published: (2025)
by: Choudhury, Pankaj, et al.
Published: (2025)
EXPERT: An Explainable Image Captioning Evaluation Metric with Structured Explanations
by: Kim, Hyunjong, et al.
Published: (2025)
by: Kim, Hyunjong, et al.
Published: (2025)
Accurate and Efficient World Modeling with Masked Latent Transformers
by: Burchi, Maxime, et al.
Published: (2025)
by: Burchi, Maxime, et al.
Published: (2025)
Learning Transformer-based World Models with Contrastive Predictive Coding
by: Burchi, Maxime, et al.
Published: (2025)
by: Burchi, Maxime, et al.
Published: (2025)
Personalizing Multimodal Large Language Models for Image Captioning: An Experimental Analysis
by: Bucciarelli, Davide, et al.
Published: (2024)
by: Bucciarelli, Davide, et al.
Published: (2024)
Image Captioning Evaluation in the Age of Multimodal LLMs: Challenges and Future Perspectives
by: Sarto, Sara, et al.
Published: (2025)
by: Sarto, Sara, et al.
Published: (2025)
DENEB: A Hallucination-Robust Automatic Evaluation Metric for Image Captioning
by: Matsuda, Kazuki, et al.
Published: (2024)
by: Matsuda, Kazuki, et al.
Published: (2024)
Polos: Multimodal Metric Learning from Human Feedback for Image Captioning
by: Wada, Yuiga, et al.
Published: (2024)
by: Wada, Yuiga, et al.
Published: (2024)
FLEUR: An Explainable Reference-Free Evaluation Metric for Image Captioning Using a Large Multimodal Model
by: Lee, Yebin, et al.
Published: (2024)
by: Lee, Yebin, et al.
Published: (2024)
WaveHiT-SR: Hierarchical Wavelet Network for Efficient Image Super-Resolution
by: Ali, Fayaz, et al.
Published: (2025)
by: Ali, Fayaz, et al.
Published: (2025)
Towards Retrieval-Augmented Architectures for Image Captioning
by: Sarto, Sara, et al.
Published: (2024)
by: Sarto, Sara, et al.
Published: (2024)
CIC-BART-SSA: Controllable Image Captioning with Structured Semantic Augmentation
by: Basioti, Kalliopi, et al.
Published: (2024)
by: Basioti, Kalliopi, et al.
Published: (2024)
CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement Learning
by: Xing, Long, et al.
Published: (2025)
by: Xing, Long, et al.
Published: (2025)
Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M)
by: Merchant, Nicholas, et al.
Published: (2025)
by: Merchant, Nicholas, et al.
Published: (2025)
Improving Image Captioning Descriptiveness by Ranking and LLM-based Fusion
by: Celona, Luigi, et al.
Published: (2023)
by: Celona, Luigi, et al.
Published: (2023)
From Pixels to Posts: Retrieval-Augmented Fashion Captioning and Hashtag Generation
by: Gondal, Moazzam Umer, et al.
Published: (2025)
by: Gondal, Moazzam Umer, et al.
Published: (2025)
Similar Items
-
Delta-Based Neural Architecture Search: LLM Fine-Tuning via Code Diffs
by: Adhikari, Santosh Premi, et al.
Published: (2026) -
Enhancing LLM-Based Neural Network Generation: Few-Shot Prompting and Efficient Validation for Automated Architecture Design
by: Duvvuri, Raghuvir, et al.
Published: (2025) -
VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs?
by: Gado, Mohamed, et al.
Published: (2025) -
MobileAgeNet: Lightweight Facial Age Estimation for Mobile Deployment
by: Kumar, Arun, et al.
Published: (2026) -
From Memorization to Creativity: LLM as a Designer of Novel Neural Architectures
by: Khalid, Waleed, et al.
Published: (2026)