CycleCap: Improving VLMs Captioning Performance via Self-Supervised Cycle Consistency Fine-Tuning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Krestenitis, Marios, Tzelepis, Christos, Ioannidis, Konstantinos, Vrochidis, Stefanos, Kompatsiaris, Ioannis, Tzimiropoulos, Georgios, Gong, Shaogang, Patras, Ioannis |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MM2Latent: Text-to-facial image generation and editing in GANs with multimodal assistance
von: Meng, Debin, et al.
Veröffentlicht: (2024)
von: Meng, Debin, et al.
Veröffentlicht: (2024)
A Comprehensive Review of Deep Learning-Based Anomaly Detection Methods for Precision Agriculture
von: Gkountakos, Konstantinos, et al.
Veröffentlicht: (2024)
von: Gkountakos, Konstantinos, et al.
Veröffentlicht: (2024)
DiffusionAct: Controllable Diffusion Autoencoder for One-shot Face Reenactment
von: Bounareli, Stella, et al.
Veröffentlicht: (2024)
von: Bounareli, Stella, et al.
Veröffentlicht: (2024)
One-shot Neural Face Reenactment via Finding Directions in GAN's Latent Space
von: Bounareli, Stella, et al.
Veröffentlicht: (2024)
von: Bounareli, Stella, et al.
Veröffentlicht: (2024)
RAG Playground: A Framework for Systematic Evaluation of Retrieval Strategies and Prompt Engineering in RAG Systems
von: Papadimitriou, Ioannis, et al.
Veröffentlicht: (2024)
von: Papadimitriou, Ioannis, et al.
Veröffentlicht: (2024)
Efficient Unsupervised Visual Representation Learning with Explicit Cluster Balancing
von: Metaxas, Ioannis Maniadis, et al.
Veröffentlicht: (2024)
von: Metaxas, Ioannis Maniadis, et al.
Veröffentlicht: (2024)
CLIPCleaner: Cleaning Noisy Labels with CLIP
von: Feng, Chen, et al.
Veröffentlicht: (2024)
von: Feng, Chen, et al.
Veröffentlicht: (2024)
SSR: An Efficient and Robust Framework for Learning with Unknown Label Noise
von: Feng, Chen, et al.
Veröffentlicht: (2021)
von: Feng, Chen, et al.
Veröffentlicht: (2021)
A Multi-Task Text Classification Pipeline with Natural Language Explanations: A User-Centric Evaluation in Sentiment Analysis and Offensive Language Identification in Greek Tweets
von: Mylonas, Nikolaos, et al.
Veröffentlicht: (2024)
von: Mylonas, Nikolaos, et al.
Veröffentlicht: (2024)
Are CLIP features all you need for Universal Synthetic Image Origin Attribution?
von: Cioni, Dario, et al.
Veröffentlicht: (2024)
von: Cioni, Dario, et al.
Veröffentlicht: (2024)
AIM-Fair: Advancing Algorithmic Fairness via Selectively Fine-Tuning Biased Models with Contextual Synthetic Data
von: Zhao, Zengqun, et al.
Veröffentlicht: (2025)
von: Zhao, Zengqun, et al.
Veröffentlicht: (2025)
CemiFace: Center-based Semi-hard Synthetic Face Generation for Face Recognition
von: Sun, Zhonglin, et al.
Veröffentlicht: (2024)
von: Sun, Zhonglin, et al.
Veröffentlicht: (2024)
LAFS: Landmark-based Facial Self-supervised Learning for Face Recognition
von: Sun, Zhonglin, et al.
Veröffentlicht: (2024)
von: Sun, Zhonglin, et al.
Veröffentlicht: (2024)
Aligned Unsupervised Pretraining of Object Detectors with Self-training
von: Metaxas, Ioannis Maniadis, et al.
Veröffentlicht: (2023)
von: Metaxas, Ioannis Maniadis, et al.
Veröffentlicht: (2023)
Temporal Score Analysis for Understanding and Correcting Diffusion Artifacts
von: Cao, Yu, et al.
Veröffentlicht: (2025)
von: Cao, Yu, et al.
Veröffentlicht: (2025)
Enhancing Zero-Shot Facial Expression Recognition by LLM Knowledge Transfer
von: Zhao, Zengqun, et al.
Veröffentlicht: (2024)
von: Zhao, Zengqun, et al.
Veröffentlicht: (2024)
Assessment of Oil Spill Dispersion and Weathering Processes in Saronic Gulf
von: Papaioannou, Vassilios, et al.
Veröffentlicht: (2025)
von: Papaioannou, Vassilios, et al.
Veröffentlicht: (2025)
Motor Imagery Decoding Using Ensemble Curriculum Learning and Collaborative Training
von: Zoumpourlis, Georgios, et al.
Veröffentlicht: (2022)
von: Zoumpourlis, Georgios, et al.
Veröffentlicht: (2022)
VLLMs Provide Better Context for Emotion Understanding Through Common Sense Reasoning
von: Xenos, Alexandros, et al.
Veröffentlicht: (2024)
von: Xenos, Alexandros, et al.
Veröffentlicht: (2024)
Understanding the Performance Plateau in Text-to-Video Retrieval: A Comprehensive Empirical and Linguistic Analysis
von: Pegia, Maria-Eirini, et al.
Veröffentlicht: (2026)
von: Pegia, Maria-Eirini, et al.
Veröffentlicht: (2026)
Self-Supervised Facial Representation Learning with Facial Region Awareness
von: Gao, Zheng, et al.
Veröffentlicht: (2024)
von: Gao, Zheng, et al.
Veröffentlicht: (2024)
Training-Free Generation of Diverse and High-Fidelity Images via Prompt Semantic Space Optimization
von: Meng, Debin, et al.
Veröffentlicht: (2025)
von: Meng, Debin, et al.
Veröffentlicht: (2025)
Utilizing Large Language Models for Machine Learning Explainability
von: Vassiliades, Alexandros, et al.
Veröffentlicht: (2025)
von: Vassiliades, Alexandros, et al.
Veröffentlicht: (2025)
Multi-scale Image Super Resolution with a Single Auto-Regressive Model
von: Sanchez, Enrique, et al.
Veröffentlicht: (2025)
von: Sanchez, Enrique, et al.
Veröffentlicht: (2025)
InDistill: Information flow-preserving knowledge distillation for model compression
von: Sarridis, Ioannis, et al.
Veröffentlicht: (2022)
von: Sarridis, Ioannis, et al.
Veröffentlicht: (2022)
Prompting Visual-Language Models for Dynamic Facial Expression Recognition
von: Zhao, Zengqun, et al.
Veröffentlicht: (2023)
von: Zhao, Zengqun, et al.
Veröffentlicht: (2023)
Deconstructing the Failure of Ideal Noise Correction: A Three-Pillar Diagnosis
von: Feng, Chen, et al.
Veröffentlicht: (2026)
von: Feng, Chen, et al.
Veröffentlicht: (2026)
LightningNet: Distributed Graph-based Cellular Network Performance Forecasting for the Edge
von: Zacharopoulos, Konstantinos, et al.
Veröffentlicht: (2024)
von: Zacharopoulos, Konstantinos, et al.
Veröffentlicht: (2024)
Multilinear Mixture of Experts: Scalable Expert Specialization through Factorization
von: Oldfield, James, et al.
Veröffentlicht: (2024)
von: Oldfield, James, et al.
Veröffentlicht: (2024)
Towards Optimal Trade-offs in Knowledge Distillation for CNNs and Vision Transformers at the Edge
von: Violos, John, et al.
Veröffentlicht: (2024)
von: Violos, John, et al.
Veröffentlicht: (2024)
VladVA: Discriminative Fine-tuning of LVLMs
von: Ouali, Yassine, et al.
Veröffentlicht: (2024)
von: Ouali, Yassine, et al.
Veröffentlicht: (2024)
EmoCLIP: A Vision-Language Method for Zero-Shot Video Facial Expression Recognition
von: Foteinopoulou, Niki Maria, et al.
Veröffentlicht: (2023)
von: Foteinopoulou, Niki Maria, et al.
Veröffentlicht: (2023)
FashionSD-X: Multimodal Fashion Garment Synthesis using Latent Diffusion
von: Singh, Abhishek Kumar, et al.
Veröffentlicht: (2024)
von: Singh, Abhishek Kumar, et al.
Veröffentlicht: (2024)
A CLIP-based siamese approach for meme classification
von: Huertas-Tato, Javier, et al.
Veröffentlicht: (2024)
von: Huertas-Tato, Javier, et al.
Veröffentlicht: (2024)
Global Convergence of Multi-Agent Policy Gradient in Markov Potential Games
von: Leonardos, Stefanos, et al.
Veröffentlicht: (2021)
von: Leonardos, Stefanos, et al.
Veröffentlicht: (2021)
LatSearch: Latent Reward-Guided Search for Faster Inference-Time Scaling in Video Diffusion
von: Zhao, Zengqun, et al.
Veröffentlicht: (2026)
von: Zhao, Zengqun, et al.
Veröffentlicht: (2026)
Improving Lean4 Autoformalization via Cycle Consistency Fine-tuning
von: Shebzukhov, Arsen
Veröffentlicht: (2026)
von: Shebzukhov, Arsen
Veröffentlicht: (2026)
Py-GC/MS Dataset Loader: GeorgiosKirtsanis/pygcms_dataset_loader
von: Kirtsanis, Georgios, et al.
Veröffentlicht: (2026)
von: Kirtsanis, Georgios, et al.
Veröffentlicht: (2026)
VISion On Request: Enhanced VLLM efficiency with sparse, dynamically selected, vision-language interactions
von: Bulat, Adrian, et al.
Veröffentlicht: (2026)
von: Bulat, Adrian, et al.
Veröffentlicht: (2026)
White-Basilisk: A Hybrid Model for Code Vulnerability Detection
von: Lamprou, Ioannis, et al.
Veröffentlicht: (2025)
von: Lamprou, Ioannis, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MM2Latent: Text-to-facial image generation and editing in GANs with multimodal assistance
von: Meng, Debin, et al.
Veröffentlicht: (2024) -
A Comprehensive Review of Deep Learning-Based Anomaly Detection Methods for Precision Agriculture
von: Gkountakos, Konstantinos, et al.
Veröffentlicht: (2024) -
DiffusionAct: Controllable Diffusion Autoencoder for One-shot Face Reenactment
von: Bounareli, Stella, et al.
Veröffentlicht: (2024) -
One-shot Neural Face Reenactment via Finding Directions in GAN's Latent Space
von: Bounareli, Stella, et al.
Veröffentlicht: (2024) -
RAG Playground: A Framework for Systematic Evaluation of Retrieval Strategies and Prompt Engineering in RAG Systems
von: Papadimitriou, Ioannis, et al.
Veröffentlicht: (2024)