DISCODE: Distribution-Aware Score Decoder for Robust Automatic Evaluation of Image Captioning
Fuente:
arXiv
Guardado en:
| Autores principales: | Inoue, Nakamasa, Goto, Kanoko, Oi, Masanari, Gruszka, Martyna, Ukai, Mahiro, Hirose, Takumi, Sekikawa, Yusuke |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Referring Expression Comprehension for Small Objects
por: Goto, Kanoko, et al.
Publicado: (2025)
por: Goto, Kanoko, et al.
Publicado: (2025)
DF-Mamba: Deformable State Space Modeling for 3D Hand Pose Estimation in Interactions
por: Zhou, Yifan, et al.
Publicado: (2025)
por: Zhou, Yifan, et al.
Publicado: (2025)
STATUS Bench: A Rigorous Benchmark for Evaluating Object State Understanding in Vision-Language Models
por: Ukai, Mahiro, et al.
Publicado: (2025)
por: Ukai, Mahiro, et al.
Publicado: (2025)
Multi-Point Positional Insertion Tuning for Small Object Detection
por: Goto, Kanoko, et al.
Publicado: (2024)
por: Goto, Kanoko, et al.
Publicado: (2024)
Autoregressive Direct Preference Optimization
por: Oi, Masanari, et al.
Publicado: (2026)
por: Oi, Masanari, et al.
Publicado: (2026)
AdaCoder: Adaptive Prompt Compression for Programmatic Visual Question Answering
por: Ukai, Mahiro, et al.
Publicado: (2024)
por: Ukai, Mahiro, et al.
Publicado: (2024)
EC-Bench: Enumeration and Counting Benchmark for Ultra-Long Videos
por: Tsuchiya, Fumihiko, et al.
Publicado: (2026)
por: Tsuchiya, Fumihiko, et al.
Publicado: (2026)
Multi-modal, Multi-task, Multi-criteria Automatic Evaluation with Vision Language Models
por: Ohi, Masanari, et al.
Publicado: (2024)
por: Ohi, Masanari, et al.
Publicado: (2024)
From Correspondence to Actions: Human-Like Multi-Image Spatial Reasoning in Multi-modal Large Language Models
por: Oi, Masanari, et al.
Publicado: (2026)
por: Oi, Masanari, et al.
Publicado: (2026)
HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis
por: Nishimura, Yuto, et al.
Publicado: (2024)
por: Nishimura, Yuto, et al.
Publicado: (2024)
ELP-Adapters: Parameter Efficient Adapter Tuning for Various Speech Processing Tasks
por: Inoue, Nakamasa, et al.
Publicado: (2024)
por: Inoue, Nakamasa, et al.
Publicado: (2024)
Masked Gated Linear Unit
por: Tajima, Yukito, et al.
Publicado: (2025)
por: Tajima, Yukito, et al.
Publicado: (2025)
MSD-Score: Multi-Scale Distributional Scoring for Reference-Free Image Caption Evaluation
por: Kan, Shichao, et al.
Publicado: (2026)
por: Kan, Shichao, et al.
Publicado: (2026)
DENEB: A Hallucination-Robust Automatic Evaluation Metric for Image Captioning
por: Matsuda, Kazuki, et al.
Publicado: (2024)
por: Matsuda, Kazuki, et al.
Publicado: (2024)
Pyramid Coder: Hierarchical Code Generator for Compositional Visual Question Answering
por: Shen, Ruoyue, et al.
Publicado: (2024)
por: Shen, Ruoyue, et al.
Publicado: (2024)
PhysQuantAgent: An Inference Pipeline of Mass Estimation for Vision-Language Models
por: Yokomizo, Hisayuki, et al.
Publicado: (2026)
por: Yokomizo, Hisayuki, et al.
Publicado: (2026)
SPECS: Specificity-Enhanced CLIP-Score for Long Image Caption Evaluation
por: Chen, Xiaofu, et al.
Publicado: (2025)
por: Chen, Xiaofu, et al.
Publicado: (2025)
Redemption Score: A Multi-Modal Evaluation Framework for Image Captioning via Distributional, Perceptual, and Linguistic Signal Triangulation
por: Dahal, Ashim, et al.
Publicado: (2025)
por: Dahal, Ashim, et al.
Publicado: (2025)
GUMBEL-NERF: Representing Unseen Objects as Part-Compositional Neural Radiance Fields
por: Sekikawa, Yusuke, et al.
Publicado: (2024)
por: Sekikawa, Yusuke, et al.
Publicado: (2024)
Neural Bloom: A Deep Learning Approach to Real-Time Lighting
por: Karp, Rafal, et al.
Publicado: (2025)
por: Karp, Rafal, et al.
Publicado: (2025)
AnimalClue: Recognizing Animals by their Traces
por: Shinoda, Risa, et al.
Publicado: (2025)
por: Shinoda, Risa, et al.
Publicado: (2025)
AgroBench: Vision-Language Model Benchmark in Agriculture
por: Shinoda, Risa, et al.
Publicado: (2025)
por: Shinoda, Risa, et al.
Publicado: (2025)
PowerCLIP: Powerset Alignment for Contrastive Pre-Training
por: Kawamura, Masaki, et al.
Publicado: (2025)
por: Kawamura, Masaki, et al.
Publicado: (2025)
An Examination of the Robustness of Reference-Free Image Captioning Evaluation Metrics
por: Ahmadi, Saba, et al.
Publicado: (2023)
por: Ahmadi, Saba, et al.
Publicado: (2023)
Compressed Image Captioning using CNN-based Encoder-Decoder Framework
por: Ridoy, Md Alif Rahman, et al.
Publicado: (2024)
por: Ridoy, Md Alif Rahman, et al.
Publicado: (2024)
Rethinking Evaluation Metrics for Grammatical Error Correction: Why Use a Different Evaluation Process than Human?
por: Goto, Takumi, et al.
Publicado: (2025)
por: Goto, Takumi, et al.
Publicado: (2025)
Visually-Aware Context Modeling for News Image Captioning
por: Qu, Tingyu, et al.
Publicado: (2023)
por: Qu, Tingyu, et al.
Publicado: (2023)
Synthesizing Instruction-Tuning Datasets with Contrastive Decoding
por: Ichinose, Tatsuya, et al.
Publicado: (2026)
por: Ichinose, Tatsuya, et al.
Publicado: (2026)
Divide and Restore: A Modular Task-Decoupled Framework for Universal Image Restoration
por: Wiekiera, Joanna, et al.
Publicado: (2026)
por: Wiekiera, Joanna, et al.
Publicado: (2026)
Grammatical Error Correction Evaluation by Optimally Transporting Edit Representation
por: Goto, Takumi, et al.
Publicado: (2026)
por: Goto, Takumi, et al.
Publicado: (2026)
Decoding Matters: Efficient Mamba-Based Decoder with Distribution-Aware Deep Supervision for Medical Image Segmentation
por: Bougourzi, Fares, et al.
Publicado: (2026)
por: Bougourzi, Fares, et al.
Publicado: (2026)
BioVITA: Biological Dataset, Model, and Benchmark for Visual-Textual-Acoustic Alignment
por: Shinoda, Risa, et al.
Publicado: (2026)
por: Shinoda, Risa, et al.
Publicado: (2026)
gec-metrics: A Unified Library for Grammatical Error Correction Evaluation
por: Goto, Takumi, et al.
Publicado: (2025)
por: Goto, Takumi, et al.
Publicado: (2025)
Understanding Retrieval Robustness for Retrieval-Augmented Image Captioning
por: Li, Wenyan, et al.
Publicado: (2024)
por: Li, Wenyan, et al.
Publicado: (2024)
Is Your Text-to-Image Model Robust to Caption Noise?
por: Yu, Weichen, et al.
Publicado: (2024)
por: Yu, Weichen, et al.
Publicado: (2024)
Removing Distributional Discrepancies in Captions Improves Image-Text Alignment
por: Li, Yuheng, et al.
Publicado: (2024)
por: Li, Yuheng, et al.
Publicado: (2024)
HICEScore: A Hierarchical Metric for Image Captioning Evaluation
por: Zeng, Zequn, et al.
Publicado: (2024)
por: Zeng, Zequn, et al.
Publicado: (2024)
On the Relationship Between Double Descent of CNNs and Shape/Texture Bias Under Learning Process
por: Iwase, Shun, et al.
Publicado: (2025)
por: Iwase, Shun, et al.
Publicado: (2025)
CIC: A Framework for Culturally-Aware Image Captioning
por: Yun, Youngsik, et al.
Publicado: (2024)
por: Yun, Youngsik, et al.
Publicado: (2024)
Evaluating Image Caption via Cycle-consistent Text-to-Image Generation
por: Cui, Tianyu, et al.
Publicado: (2025)
por: Cui, Tianyu, et al.
Publicado: (2025)
Ejemplares similares
-
Referring Expression Comprehension for Small Objects
por: Goto, Kanoko, et al.
Publicado: (2025) -
DF-Mamba: Deformable State Space Modeling for 3D Hand Pose Estimation in Interactions
por: Zhou, Yifan, et al.
Publicado: (2025) -
STATUS Bench: A Rigorous Benchmark for Evaluating Object State Understanding in Vision-Language Models
por: Ukai, Mahiro, et al.
Publicado: (2025) -
Multi-Point Positional Insertion Tuning for Small Object Detection
por: Goto, Kanoko, et al.
Publicado: (2024) -
Autoregressive Direct Preference Optimization
por: Oi, Masanari, et al.
Publicado: (2026)