LOTUS: A Leaderboard for Detailed Image Captioning from Quality to Societal Bias and User Preferences
Fuente:
arXiv
Saved in:
| Main Authors: | Hirota, Yusuke, Li, Boyi, Hachiuma, Ryo, Wu, Yueh-Hua, Ivanovic, Boris, Nakashima, Yuta, Pavone, Marco, Choi, Yejin, Wang, Yu-Chiang Frank, Yang, Chao-Han Huck |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Descriptive Richness to Bias: Unveiling the Dark Side of Generative Image Caption Enrichment
by: Hirota, Yusuke, et al.
Published: (2024)
by: Hirota, Yusuke, et al.
Published: (2024)
Bias in Gender Bias Benchmarks: How Spurious Features Distort Evaluation
by: Hirota, Yusuke, et al.
Published: (2025)
by: Hirota, Yusuke, et al.
Published: (2025)
SANER: Annotation-free Societal Attribute Neutralizer for Debiasing CLIP
by: Hirota, Yusuke, et al.
Published: (2024)
by: Hirota, Yusuke, et al.
Published: (2024)
Resampled Datasets Are Not Enough: Mitigating Societal Bias Beyond Single Attributes
by: Hirota, Yusuke, et al.
Published: (2024)
by: Hirota, Yusuke, et al.
Published: (2024)
From Global to Local: Social Bias Transfer in CLIP
by: Ramos, Ryan, et al.
Published: (2025)
by: Ramos, Ryan, et al.
Published: (2025)
Would Deep Generative Models Amplify Bias in Future Models?
by: Chen, Tianwei, et al.
Published: (2024)
by: Chen, Tianwei, et al.
Published: (2024)
Who Defines "Best"? Towards Interactive, User-Defined Evaluation of LLM Leaderboards
by: Jung, Minji, et al.
Published: (2026)
by: Jung, Minji, et al.
Published: (2026)
Human Preference-Aligned Concept Customization Benchmark via Decomposed Evaluation
by: Ishikawa, Reina, et al.
Published: (2025)
by: Ishikawa, Reina, et al.
Published: (2025)
Interpretable Debiasing of Vision-Language Models for Social Fairness
by: An, Na Min, et al.
Published: (2026)
by: An, Na Min, et al.
Published: (2026)
Describe Anything: Detailed Localized Image and Video Captioning
by: Lian, Long, et al.
Published: (2025)
by: Lian, Long, et al.
Published: (2025)
GenRecal: Generation after Recalibration from Large to Small Vision-Language Models
by: Lee, Byung-Kwan, et al.
Published: (2025)
by: Lee, Byung-Kwan, et al.
Published: (2025)
Towards Predicting Any Human Trajectory In Context
by: Fujii, Ryo, et al.
Published: (2025)
by: Fujii, Ryo, et al.
Published: (2025)
VLsI: Verbalized Layers-to-Interactions from Large to Small Vision Language Models
by: Lee, Byung-Kwan, et al.
Published: (2024)
by: Lee, Byung-Kwan, et al.
Published: (2024)
Unified Reinforcement and Imitation Learning for Vision-Language Models
by: Lee, Byung-Kwan, et al.
Published: (2025)
by: Lee, Byung-Kwan, et al.
Published: (2025)
Driving Everywhere with Large Language Model Policy Adaptation
by: Li, Boyi, et al.
Published: (2024)
by: Li, Boyi, et al.
Published: (2024)
LoRA3D: Low-Rank Self-Calibration of 3D Geometric Foundation Models
by: Lu, Ziqi, et al.
Published: (2024)
by: Lu, Ziqi, et al.
Published: (2024)
RealTraj: Towards Real-World Pedestrian Trajectory Forecasting
by: Fujii, Ryo, et al.
Published: (2024)
by: Fujii, Ryo, et al.
Published: (2024)
Gender Bias Evaluation in Text-to-image Generation: A Survey
by: Wu, Yankun, et al.
Published: (2024)
by: Wu, Yankun, et al.
Published: (2024)
Weakly Semi-supervised Tool Detection in Minimally Invasive Surgery Videos
by: Fujii, Ryo, et al.
Published: (2024)
by: Fujii, Ryo, et al.
Published: (2024)
VIOLA: Towards Video In-Context Learning with Minimal Annotations
by: Fujii, Ryo, et al.
Published: (2026)
by: Fujii, Ryo, et al.
Published: (2026)
Prompt-to-Leaderboard
by: Frick, Evan, et al.
Published: (2025)
by: Frick, Evan, et al.
Published: (2025)
Masking Teacher and Reinforcing Student for Distilling Vision-Language Models
by: Lee, Byung-Kwan, et al.
Published: (2025)
by: Lee, Byung-Kwan, et al.
Published: (2025)
CrowdMAC: Masked Crowd Density Completion for Robust Crowd Density Forecasting
by: Fujii, Ryo, et al.
Published: (2024)
by: Fujii, Ryo, et al.
Published: (2024)
Stable Diffusion Exposed: Gender Bias from Prompt to Image
by: Wu, Yankun, et al.
Published: (2023)
by: Wu, Yankun, et al.
Published: (2023)
Promptable Closed-loop Traffic Simulation
by: Tan, Shuhan, et al.
Published: (2024)
by: Tan, Shuhan, et al.
Published: (2024)
EMAG: Ego-motion Aware and Generalizable 2D Hand Forecasting from Egocentric Videos
by: Hatano, Masashi, et al.
Published: (2024)
by: Hatano, Masashi, et al.
Published: (2024)
Multimodal Cross-Domain Few-Shot Learning for Egocentric Action Recognition
by: Hatano, Masashi, et al.
Published: (2024)
by: Hatano, Masashi, et al.
Published: (2024)
Wolf: Dense Video Captioning with a World Summarization Framework
by: Li, Boyi, et al.
Published: (2024)
by: Li, Boyi, et al.
Published: (2024)
On The Conceptualization and Societal Impact of Cross-Cultural Bias
by: Bhandari, Vitthal
Published: (2025)
by: Bhandari, Vitthal
Published: (2025)
Enhancing Multimodal LLM for Detailed and Accurate Video Captioning using Multi-Round Preference Optimization
by: Tang, Changli, et al.
Published: (2024)
by: Tang, Changli, et al.
Published: (2024)
User-centric Subjective Leaderboard by Customizable Reward Modeling
by: Jia, Qi, et al.
Published: (2025)
by: Jia, Qi, et al.
Published: (2025)
SynPO: Synergizing Descriptiveness and Preference Optimization for Video Detailed Captioning
by: Dang, Jisheng, et al.
Published: (2025)
by: Dang, Jisheng, et al.
Published: (2025)
Producing and Leveraging Online Map Uncertainty in Trajectory Prediction
by: Gu, Xunjiang, et al.
Published: (2024)
by: Gu, Xunjiang, et al.
Published: (2024)
Accelerating Online Mapping and Behavior Prediction via Direct BEV Feature Attention
by: Gu, Xunjiang, et al.
Published: (2024)
by: Gu, Xunjiang, et al.
Published: (2024)
Towards Efficient and Effective Multi-Camera Encoding for End-to-End Driving
by: Yang, Jiawei, et al.
Published: (2025)
by: Yang, Jiawei, et al.
Published: (2025)
Zoom-Zero: Reinforced Coarse-to-Fine Video Understanding via Temporal Zoom-in
by: Shen, Xiaoqian, et al.
Published: (2025)
by: Shen, Xiaoqian, et al.
Published: (2025)
The Devil is in the EOS: Sequence Training for Detailed Image Captioning
by: Mohamed, Abdelrahman, et al.
Published: (2025)
by: Mohamed, Abdelrahman, et al.
Published: (2025)
Language-Image Models with 3D Understanding
by: Cho, Jang Hyun, et al.
Published: (2024)
by: Cho, Jang Hyun, et al.
Published: (2024)
Metropolis-Hastings Captioning Game: Knowledge Fusion of Vision Language Models via Decentralized Bayesian Inference
by: Matsui, Yuta, et al.
Published: (2025)
by: Matsui, Yuta, et al.
Published: (2025)
FoundationMotion: Auto-Labeling and Reasoning about Spatial Movement in Videos
by: Gan, Yulu, et al.
Published: (2025)
by: Gan, Yulu, et al.
Published: (2025)
Similar Items
-
From Descriptive Richness to Bias: Unveiling the Dark Side of Generative Image Caption Enrichment
by: Hirota, Yusuke, et al.
Published: (2024) -
Bias in Gender Bias Benchmarks: How Spurious Features Distort Evaluation
by: Hirota, Yusuke, et al.
Published: (2025) -
SANER: Annotation-free Societal Attribute Neutralizer for Debiasing CLIP
by: Hirota, Yusuke, et al.
Published: (2024) -
Resampled Datasets Are Not Enough: Mitigating Societal Bias Beyond Single Attributes
by: Hirota, Yusuke, et al.
Published: (2024) -
From Global to Local: Social Bias Transfer in CLIP
by: Ramos, Ryan, et al.
Published: (2025)