Vision-Language Models Generate More Homogeneous Stories for Phenotypically Black Individuals
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lee, Messi H. J., Jeon, Soyeon |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Visual Cues of Gender and Race are Associated with Stereotyping in Vision-Language Models
von: Lee, Messi H. J., et al.
Veröffentlicht: (2025)
von: Lee, Messi H. J., et al.
Veröffentlicht: (2025)
More Distinctively Black and Feminine Faces Lead to Increased Stereotyping in Vision-Language Models
von: Lee, Messi H. J., et al.
Veröffentlicht: (2024)
von: Lee, Messi H. J., et al.
Veröffentlicht: (2024)
Examining the Robustness of Homogeneity Bias to Hyperparameter Adjustments in GPT-4
von: Lee, Messi H. J.
Veröffentlicht: (2025)
von: Lee, Messi H. J.
Veröffentlicht: (2025)
Token-Level Entropy Reveals Demographic Disparities in Language Models
von: Lee, Messi H. J.
Veröffentlicht: (2025)
von: Lee, Messi H. J.
Veröffentlicht: (2025)
SEED-Story: Multimodal Long Story Generation with Large Language Model
von: Yang, Shuai, et al.
Veröffentlicht: (2024)
von: Yang, Shuai, et al.
Veröffentlicht: (2024)
Vision-aligned Latent Reasoning for Multi-modal Large Language Model
von: Jeon, Byungwoo, et al.
Veröffentlicht: (2026)
von: Jeon, Byungwoo, et al.
Veröffentlicht: (2026)
Diverse Rare Sample Generation with Pretrained GANs
von: Lee, Subeen, et al.
Veröffentlicht: (2024)
von: Lee, Subeen, et al.
Veröffentlicht: (2024)
Improving Calibration in Test-Time Prompt Tuning for Vision-Language Models via Data-Free Flatness-Aware Prompt Pretraining
von: Jang, Hyeonseo, et al.
Veröffentlicht: (2026)
von: Jang, Hyeonseo, et al.
Veröffentlicht: (2026)
uCLIP: Parameter-Efficient Multilingual Extension of Vision-Language Models with Unpaired Data
von: Chung, Dahyun, et al.
Veröffentlicht: (2025)
von: Chung, Dahyun, et al.
Veröffentlicht: (2025)
StoryGPT-V: Large Language Models as Consistent Story Visualizers
von: Shen, Xiaoqian, et al.
Veröffentlicht: (2023)
von: Shen, Xiaoqian, et al.
Veröffentlicht: (2023)
ForensicZip: More Tokens are Better but Not Necessary in Forensic Vision-Language Models
von: Lai, Yingxin, et al.
Veröffentlicht: (2026)
von: Lai, Yingxin, et al.
Veröffentlicht: (2026)
Why Far Looks Up: Probing Spatial Representation in Vision-Language Models
von: Min, Cheolhong, et al.
Veröffentlicht: (2026)
von: Min, Cheolhong, et al.
Veröffentlicht: (2026)
How Blind and Low-Vision Individuals Prefer Large Vision-Language Model-Generated Scene Descriptions
von: An, Na Min, et al.
Veröffentlicht: (2025)
von: An, Na Min, et al.
Veröffentlicht: (2025)
Test-Time Hinting for Black-Box Vision-Language Models
von: Hou, Kaihua, et al.
Veröffentlicht: (2026)
von: Hou, Kaihua, et al.
Veröffentlicht: (2026)
EmoStory: Emotion-Aware Story Generation
von: Yang, Jingyuan, et al.
Veröffentlicht: (2026)
von: Yang, Jingyuan, et al.
Veröffentlicht: (2026)
Infusing Environmental Captions for Long-Form Video Language Grounding
von: Lee, Hyogun, et al.
Veröffentlicht: (2024)
von: Lee, Hyogun, et al.
Veröffentlicht: (2024)
Language Models as Black-Box Optimizers for Vision-Language Models
von: Liu, Shihong, et al.
Veröffentlicht: (2023)
von: Liu, Shihong, et al.
Veröffentlicht: (2023)
More Than the Final Answer: Improving Visual Extraction and Logical Consistency in Vision-Language Models
von: Just, Hoang Anh, et al.
Veröffentlicht: (2025)
von: Just, Hoang Anh, et al.
Veröffentlicht: (2025)
ZIP: An Efficient Zeroth-order Prompt Tuning for Black-box Vision-Language Models
von: Park, Seonghwan, et al.
Veröffentlicht: (2025)
von: Park, Seonghwan, et al.
Veröffentlicht: (2025)
Cog3DMap: Multi-View Vision-Language Reasoning with 3D Cognitive Maps
von: Gwak, Chanyoung, et al.
Veröffentlicht: (2026)
von: Gwak, Chanyoung, et al.
Veröffentlicht: (2026)
Are Large Vision-Language Models Ready to Guide Blind and Low-Vision Individuals?
von: Kim, Eunki, et al.
Veröffentlicht: (2025)
von: Kim, Eunki, et al.
Veröffentlicht: (2025)
Generate, but Verify: Reducing Hallucination in Vision-Language Models with Retrospective Resampling
von: Wu, Tsung-Han, et al.
Veröffentlicht: (2025)
von: Wu, Tsung-Han, et al.
Veröffentlicht: (2025)
IdentityStory: Taming Your Identity-Preserving Generator for Human-Centric Story Generation
von: Zhou, Donghao, et al.
Veröffentlicht: (2025)
von: Zhou, Donghao, et al.
Veröffentlicht: (2025)
PhenoLIP: Integrating Phenotype Ontology Knowledge into Medical Vision-Language Pretraining
von: Liang, Cheng, et al.
Veröffentlicht: (2026)
von: Liang, Cheng, et al.
Veröffentlicht: (2026)
Active Prompt Learning in Vision Language Models
von: Bang, Jihwan, et al.
Veröffentlicht: (2023)
von: Bang, Jihwan, et al.
Veröffentlicht: (2023)
How Does Vision-Language Adaptation Impact the Safety of Vision Language Models?
von: Lee, Seongyun, et al.
Veröffentlicht: (2024)
von: Lee, Seongyun, et al.
Veröffentlicht: (2024)
How to Determine the Preferred Image Distribution of a Black-Box Vision-Language Model?
von: Taghanaki, Saeid Asgari, et al.
Veröffentlicht: (2024)
von: Taghanaki, Saeid Asgari, et al.
Veröffentlicht: (2024)
Multi-Modal Interpretability for Enhanced Localization in Vision-Language Models
von: Imran, Muhammad, et al.
Veröffentlicht: (2025)
von: Imran, Muhammad, et al.
Veröffentlicht: (2025)
Lay2Story: Extending Diffusion Transformers for Layout-Togglable Story Generation
von: Ma, Ao, et al.
Veröffentlicht: (2025)
von: Ma, Ao, et al.
Veröffentlicht: (2025)
ReDiStory: Region-Disentangled Diffusion for Consistent Visual Story Generation
von: Sarkar, Ayushman, et al.
Veröffentlicht: (2026)
von: Sarkar, Ayushman, et al.
Veröffentlicht: (2026)
Bridging Different Language Models and Generative Vision Models for Text-to-Image Generation
von: Zhao, Shihao, et al.
Veröffentlicht: (2024)
von: Zhao, Shihao, et al.
Veröffentlicht: (2024)
Quantized Prompt for Efficient Generalization of Vision-Language Models
von: Hao, Tianxiang, et al.
Veröffentlicht: (2024)
von: Hao, Tianxiang, et al.
Veröffentlicht: (2024)
Generalizing Vision-Language Models with Dedicated Prompt Guidance
von: Li, Xinyao, et al.
Veröffentlicht: (2025)
von: Li, Xinyao, et al.
Veröffentlicht: (2025)
Descriminative-Generative Custom Tokens for Vision-Language Models
von: Perera, Pramuditha, et al.
Veröffentlicht: (2025)
von: Perera, Pramuditha, et al.
Veröffentlicht: (2025)
Hypergraph Vision Transformers: Images are More than Nodes, More than Edges
von: Fixelle, Joshua
Veröffentlicht: (2025)
von: Fixelle, Joshua
Veröffentlicht: (2025)
A Vision Language Model for Generating Procedural Plant Architecture Representations from Simulated Images
von: Yun, Heesup, et al.
Veröffentlicht: (2026)
von: Yun, Heesup, et al.
Veröffentlicht: (2026)
Exploiting Domain Properties in Language-Driven Domain Generalization for Semantic Segmentation
von: Jeon, Seogkyu, et al.
Veröffentlicht: (2025)
von: Jeon, Seogkyu, et al.
Veröffentlicht: (2025)
More Thought, Less Accuracy? On the Dual Nature of Reasoning in Vision-Language Models
von: Tian, Xinyu, et al.
Veröffentlicht: (2025)
von: Tian, Xinyu, et al.
Veröffentlicht: (2025)
Learning More by Seeing Less: Structure First Learning for Efficient, Transferable, and Human-Aligned Vision
von: Li, Tianqin, et al.
Veröffentlicht: (2025)
von: Li, Tianqin, et al.
Veröffentlicht: (2025)
Focusing Where Vision Matters: Selective Training for Large Vision Language Models via Visual Information Gain
von: Lee, Seulbi, et al.
Veröffentlicht: (2026)
von: Lee, Seulbi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Visual Cues of Gender and Race are Associated with Stereotyping in Vision-Language Models
von: Lee, Messi H. J., et al.
Veröffentlicht: (2025) -
More Distinctively Black and Feminine Faces Lead to Increased Stereotyping in Vision-Language Models
von: Lee, Messi H. J., et al.
Veröffentlicht: (2024) -
Examining the Robustness of Homogeneity Bias to Hyperparameter Adjustments in GPT-4
von: Lee, Messi H. J.
Veröffentlicht: (2025) -
Token-Level Entropy Reveals Demographic Disparities in Language Models
von: Lee, Messi H. J.
Veröffentlicht: (2025) -
SEED-Story: Multimodal Long Story Generation with Large Language Model
von: Yang, Shuai, et al.
Veröffentlicht: (2024)