L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training
Fuente:
arXiv
Guardado en:
| Autores principales: | Li, Li, Peng, Yingzhe, Yang, Xu, Cheng, Ruoxi, Xu, Haiyang, Yan, Ming, Huang, Fei |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
HICEScore: A Hierarchical Metric for Image Captioning Evaluation
por: Zeng, Zequn, et al.
Publicado: (2024)
por: Zeng, Zequn, et al.
Publicado: (2024)
Adaptively Clustering Neighbor Elements for Image-Text Generation
por: Wang, Zihua, et al.
Publicado: (2023)
por: Wang, Zihua, et al.
Publicado: (2023)
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation
por: Zhang, Shi-Xue, et al.
Publicado: (2025)
por: Zhang, Shi-Xue, et al.
Publicado: (2025)
Evaluating Remote Sensing Image Captions Beyond Metric Biases
por: Chen, Ziyun, et al.
Publicado: (2026)
por: Chen, Ziyun, et al.
Publicado: (2026)
DEVICE: Depth and Visual Concepts Aware Transformer for OCR-based Image Captioning
por: Xu, Dongsheng, et al.
Publicado: (2023)
por: Xu, Dongsheng, et al.
Publicado: (2023)
CaptionQA: Is Your Caption as Useful as the Image Itself?
por: Yang, Shijia, et al.
Publicado: (2025)
por: Yang, Shijia, et al.
Publicado: (2025)
MIBench: Evaluating Multimodal Large Language Models over Multiple Images
por: Liu, Haowei, et al.
Publicado: (2024)
por: Liu, Haowei, et al.
Publicado: (2024)
TinyChart: Efficient Chart Understanding with Visual Token Merging and Program-of-Thoughts Learning
por: Zhang, Liang, et al.
Publicado: (2024)
por: Zhang, Liang, et al.
Publicado: (2024)
Cockatiel: Ensembling Synthetic and Human Preferenced Training for Detailed Video Caption
por: Qin, Luozheng, et al.
Publicado: (2025)
por: Qin, Luozheng, et al.
Publicado: (2025)
SimInversion: A Simple Framework for Inversion-Based Text-to-Image Editing
por: Qian, Qi, et al.
Publicado: (2024)
por: Qian, Qi, et al.
Publicado: (2024)
Efficient and Effective In-context Demonstration Selection with Coreset
por: Wang, Zihua, et al.
Publicado: (2025)
por: Wang, Zihua, et al.
Publicado: (2025)
G-VEval: A Versatile Metric for Evaluating Image and Video Captions Using GPT-4o
por: Tong, Tony Cheng, et al.
Publicado: (2024)
por: Tong, Tony Cheng, et al.
Publicado: (2024)
ICCV23 Visual-Dialog Emotion Explanation Challenge: SEU_309 Team Technical Report
por: Yuan, Yixiao, et al.
Publicado: (2024)
por: Yuan, Yixiao, et al.
Publicado: (2024)
Do Vision Encoders Truly Explain Object Hallucination?: Mitigating Object Hallucination via Simple Fine-Grained CLIPScore
por: Oh, Hongseok, et al.
Publicado: (2025)
por: Oh, Hongseok, et al.
Publicado: (2025)
Mobile-Agent-V: A Video-Guided Approach for Effortless and Efficient Operational Knowledge Injection in Mobile Automation
por: Wang, Junyang, et al.
Publicado: (2025)
por: Wang, Junyang, et al.
Publicado: (2025)
BUS:Efficient and Effective Vision-language Pre-training with Bottom-Up Patch Summarization
por: Jiang, Chaoya, et al.
Publicado: (2023)
por: Jiang, Chaoya, et al.
Publicado: (2023)
BTCChat: Advancing Remote Sensing Bi-temporal Change Captioning with Multimodal Large Language Model
por: Li, Yujie, et al.
Publicado: (2025)
por: Li, Yujie, et al.
Publicado: (2025)
Precision or Recall? An Analysis of Image Captions for Training Text-to-Image Generation Model
por: Cheng, Sheng, et al.
Publicado: (2024)
por: Cheng, Sheng, et al.
Publicado: (2024)
Context-aware Difference Distilling for Multi-change Captioning
por: Tu, Yunbin, et al.
Publicado: (2024)
por: Tu, Yunbin, et al.
Publicado: (2024)
AeroLite: Tag-Guided Lightweight Generation of Aerial Image Captions
por: Zi, Xing, et al.
Publicado: (2025)
por: Zi, Xing, et al.
Publicado: (2025)
Set Prediction Guided by Semantic Concepts for Diverse Video Captioning
por: Lu, Yifan, et al.
Publicado: (2023)
por: Lu, Yifan, et al.
Publicado: (2023)
TOD3Cap: Towards 3D Dense Captioning in Outdoor Scenes
por: Jin, Bu, et al.
Publicado: (2024)
por: Jin, Bu, et al.
Publicado: (2024)
Scene Graph-guided SegCaptioning Transformer with Fine-grained Alignment for Controllable Video Segmentation and Captioning
por: Zhang, Xu, et al.
Publicado: (2026)
por: Zhang, Xu, et al.
Publicado: (2026)
OmniCaptioner: One Captioner to Rule Them All
por: Lu, Yiting, et al.
Publicado: (2025)
por: Lu, Yiting, et al.
Publicado: (2025)
Dual-frequency Selected Knowledge Distillation with Statistical-based Sample Rectification for PolSAR Image Classification
por: Xin, Xinyue, et al.
Publicado: (2025)
por: Xin, Xinyue, et al.
Publicado: (2025)
Aesthetic Image Captioning with Saliency Enhanced MLLMs
por: Tao, Yilin, et al.
Publicado: (2025)
por: Tao, Yilin, et al.
Publicado: (2025)
Distractors-Immune Representation Learning with Cross-modal Contrastive Regularization for Change Captioning
por: Tu, Yunbin, et al.
Publicado: (2024)
por: Tu, Yunbin, et al.
Publicado: (2024)
CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language Models
por: Huang, Junming, et al.
Publicado: (2026)
por: Huang, Junming, et al.
Publicado: (2026)
A Conformal Risk Control Framework for Granular Word Assessment and Uncertainty Calibration of CLIPScore Quality Estimates
por: Gomes, Gonçalo, et al.
Publicado: (2025)
por: Gomes, Gonçalo, et al.
Publicado: (2025)
STAMP: Training Explicit Memory for Mobile GUI Agents in Controllable and Scalable Virtual Environments
por: Wang, Junyang, et al.
Publicado: (2026)
por: Wang, Junyang, et al.
Publicado: (2026)
ProxyImg: Towards Highly-Controllable Image Representation via Hierarchical Disentangled Proxy Embedding
por: Chen, Ye, et al.
Publicado: (2026)
por: Chen, Ye, et al.
Publicado: (2026)
Group Relative Policy Optimization for Image Captioning
por: Liang, Xu
Publicado: (2025)
por: Liang, Xu
Publicado: (2025)
Unifying Latent and Lexicon Representations for Effective Video-Text Retrieval
por: Liu, Haowei, et al.
Publicado: (2024)
por: Liu, Haowei, et al.
Publicado: (2024)
Semantics-enhanced Cross-modal Masked Image Modeling for Vision-Language Pre-training
por: Liu, Haowei, et al.
Publicado: (2024)
por: Liu, Haowei, et al.
Publicado: (2024)
Exploring Diverse In-Context Configurations for Image Captioning
por: Yang, Xu, et al.
Publicado: (2023)
por: Yang, Xu, et al.
Publicado: (2023)
Exploiting Auxiliary Caption for Video Grounding
por: Li, Hongxiang, et al.
Publicado: (2023)
por: Li, Hongxiang, et al.
Publicado: (2023)
EXPERT: An Explainable Image Captioning Evaluation Metric with Structured Explanations
por: Kim, Hyunjong, et al.
Publicado: (2025)
por: Kim, Hyunjong, et al.
Publicado: (2025)
Surveying the Landscape of Image Captioning Evaluation: A Comprehensive Taxonomy, Trends and Metrics Analysis
por: Berger, Uri, et al.
Publicado: (2024)
por: Berger, Uri, et al.
Publicado: (2024)
Culture-Aware Humorous Captioning: Multimodal Humor Generation across Cultural Contexts
por: Xu, Run, et al.
Publicado: (2026)
por: Xu, Run, et al.
Publicado: (2026)
MM-SeR: Multimodal Self-Refinement for Lightweight Image Captioning
por: Song, Junha, et al.
Publicado: (2025)
por: Song, Junha, et al.
Publicado: (2025)
Ejemplares similares
-
HICEScore: A Hierarchical Metric for Image Captioning Evaluation
por: Zeng, Zequn, et al.
Publicado: (2024) -
Adaptively Clustering Neighbor Elements for Image-Text Generation
por: Wang, Zihua, et al.
Publicado: (2023) -
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation
por: Zhang, Shi-Xue, et al.
Publicado: (2025) -
Evaluating Remote Sensing Image Captions Beyond Metric Biases
por: Chen, Ziyun, et al.
Publicado: (2026) -
DEVICE: Depth and Visual Concepts Aware Transformer for OCR-based Image Captioning
por: Xu, Dongsheng, et al.
Publicado: (2023)