Debias your Large Multi-Modal Model at Test-Time via Non-Contrastive Visual Attribute Steering
Fuente:
arXiv
Guardado en:
| Autores principales: | Ratzlaff, Neale, Olson, Matthew Lyle, Hinck, Musashi, Aflalo, Estelle, Tseng, Shao-Yen, Lal, Vasudev, Howard, Phillip |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Debiasing Large Vision-Language Models by Ablating Protected Attribute Representations
por: Ratzlaff, Neale, et al.
Publicado: (2024)
por: Ratzlaff, Neale, et al.
Publicado: (2024)
Steering Large Language Models to Evaluate and Amplify Creativity
por: Olson, Matthew Lyle, et al.
Publicado: (2024)
por: Olson, Matthew Lyle, et al.
Publicado: (2024)
Probing the Representational Power of Sparse Autoencoders in Vision Models
por: Olson, Matthew Lyle, et al.
Publicado: (2025)
por: Olson, Matthew Lyle, et al.
Publicado: (2025)
Analyzing Hierarchical Structure in Vision Models with Sparse Autoencoders
por: Olson, Matthew Lyle, et al.
Publicado: (2025)
por: Olson, Matthew Lyle, et al.
Publicado: (2025)
LieCraft: A Multi-Agent Framework for Evaluating Deceptive Capabilities in Language Models
por: Olson, Matthew Lyle, et al.
Publicado: (2026)
por: Olson, Matthew Lyle, et al.
Publicado: (2026)
Probing Semantic Routing in Large Mixture-of-Expert Models
por: Olson, Matthew Lyle, et al.
Publicado: (2025)
por: Olson, Matthew Lyle, et al.
Publicado: (2025)
LLaVA-Gemma: Accelerating Multimodal Foundation Models with a Compact Language Model
por: Hinck, Musashi, et al.
Publicado: (2024)
por: Hinck, Musashi, et al.
Publicado: (2024)
Training-Free Mitigation of Language Reasoning Degradation After Multimodal Instruction Tuning
por: Ratzlaff, Neale, et al.
Publicado: (2024)
por: Ratzlaff, Neale, et al.
Publicado: (2024)
Learning from Reasoning Failures via Synthetic Data Generation
por: Stan, Gabriela Ben Melech, et al.
Publicado: (2025)
por: Stan, Gabriela Ben Melech, et al.
Publicado: (2025)
Why do LLaVA Vision-Language Models Reply to Images in English?
por: Hinck, Musashi, et al.
Publicado: (2024)
por: Hinck, Musashi, et al.
Publicado: (2024)
ClimDetect: A Benchmark Dataset for Climate Change Detection and Attribution
por: Yu, Sungduk, et al.
Publicado: (2024)
por: Yu, Sungduk, et al.
Publicado: (2024)
LVLM-Interpret: An Interpretability Tool for Large Vision-Language Models
por: Stan, Gabriela Ben Melech, et al.
Publicado: (2024)
por: Stan, Gabriela Ben Melech, et al.
Publicado: (2024)
FastRM: An efficient and automatic explainability framework for multimodal generative models
por: Stan, Gabriela Ben-Melech, et al.
Publicado: (2024)
por: Stan, Gabriela Ben-Melech, et al.
Publicado: (2024)
FiVL: A Framework for Improved Vision-Language Alignment through the Lens of Training, Evaluation and Explainability
por: Aflalo, Estelle, et al.
Publicado: (2024)
por: Aflalo, Estelle, et al.
Publicado: (2024)
A Causal World Model Underlying Next Token Prediction: Exploring GPT in a Controlled Environment
por: Rohekar, Raanan Y., et al.
Publicado: (2024)
por: Rohekar, Raanan Y., et al.
Publicado: (2024)
DPO Learning with LLMs-Judge Signal for Computer Use Agents
por: Luo, Man, et al.
Publicado: (2025)
por: Luo, Man, et al.
Publicado: (2025)
NeuroPrompts: An Adaptive Framework to Optimize Prompts for Text-to-Image Generation
por: Rosenman, Shachar, et al.
Publicado: (2023)
por: Rosenman, Shachar, et al.
Publicado: (2023)
Cultural Awareness in Vision-Language Models: A Cross-Country Exploration
por: Madasu, Avinash, et al.
Publicado: (2025)
por: Madasu, Avinash, et al.
Publicado: (2025)
Pruning the Paradox: How CLIP's Most Informative Heads Enhance Performance While Amplifying Bias
por: Madasu, Avinash, et al.
Publicado: (2025)
por: Madasu, Avinash, et al.
Publicado: (2025)
Quantifying and Enabling the Interpretability of CLIP-like Models
por: Madasu, Avinash, et al.
Publicado: (2024)
por: Madasu, Avinash, et al.
Publicado: (2024)
Is Your Paper Being Reviewed by an LLM? Benchmarking AI Text Detection in Peer Review
por: Yu, Sungduk, et al.
Publicado: (2025)
por: Yu, Sungduk, et al.
Publicado: (2025)
Is Your Paper Being Reviewed by an LLM? Investigating AI Text Detectability in Peer Review
por: Yu, Sungduk, et al.
Publicado: (2024)
por: Yu, Sungduk, et al.
Publicado: (2024)
Using Imperfect Surrogates for Downstream Inference: Design-based Supervised Learning for Social Science Applications of Large Language Models
por: Egami, Naoki, et al.
Publicado: (2023)
por: Egami, Naoki, et al.
Publicado: (2023)
AutoPersuade: A Framework for Evaluating and Explaining Persuasive Arguments
por: Saenger, Till Raphael, et al.
Publicado: (2024)
por: Saenger, Till Raphael, et al.
Publicado: (2024)
NeuroComparatives: Neuro-Symbolic Distillation of Comparative Knowledge
por: Howard, Phillip, et al.
Publicado: (2023)
por: Howard, Phillip, et al.
Publicado: (2023)
SocialCounterfactuals: Probing and Mitigating Intersectional Social Biases in Vision-Language Models with Counterfactual Examples
por: Howard, Phillip, et al.
Publicado: (2023)
por: Howard, Phillip, et al.
Publicado: (2023)
ICSVR: Investigating Compositional and Syntactic Understanding in Video Retrieval Models
por: Madasu, Avinash, et al.
Publicado: (2023)
por: Madasu, Avinash, et al.
Publicado: (2023)
SK-VQA: Synthetic Knowledge Generation at Scale for Training Context-Augmented Multimodal LLMs
por: Su, Xin, et al.
Publicado: (2024)
por: Su, Xin, et al.
Publicado: (2024)
Getting it Right: Improving Spatial Consistency in Text-to-Image Models
por: Chatterjee, Agneet, et al.
Publicado: (2024)
por: Chatterjee, Agneet, et al.
Publicado: (2024)
IssueBench: Millions of Realistic Prompts for Measuring Issue Bias in LLM Writing Assistance
por: Röttger, Paul, et al.
Publicado: (2025)
por: Röttger, Paul, et al.
Publicado: (2025)
Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models
por: Röttger, Paul, et al.
Publicado: (2024)
por: Röttger, Paul, et al.
Publicado: (2024)
L-MAGIC: Language Model Assisted Generation of Images with Coherence
por: Cai, Zhipeng, et al.
Publicado: (2024)
por: Cai, Zhipeng, et al.
Publicado: (2024)
Advocating with your story: Compelling our legislators to act
por: Zoe Tseng
Publicado: (2025)
por: Zoe Tseng
Publicado: (2025)
Utilizando Recursos Computacionais (Planilha) na Compreensão dos Números Racionais
por: Rosane Ratzlaff da Rosa
Publicado: (2008)
por: Rosane Ratzlaff da Rosa
Publicado: (2008)
LVLM-Compress-Bench: Benchmarking the Broader Impact of Large Vision-Language Model Compression
por: Kundu, Souvik, et al.
Publicado: (2025)
por: Kundu, Souvik, et al.
Publicado: (2025)
Controlling your Attributes in Voice
por: Li, Xuyuan, et al.
Publicado: (2025)
por: Li, Xuyuan, et al.
Publicado: (2025)
Geometry-Aware CLIP Retrieval via Local Cross-Modal Alignment and Steering
por: Prakash, Nirmalendu, et al.
Publicado: (2026)
por: Prakash, Nirmalendu, et al.
Publicado: (2026)
Debias-CLR: A Contrastive Learning Based Debiasing Method for Algorithmic Fairness in Healthcare Applications
por: Agarwal, Ankita, et al.
Publicado: (2024)
por: Agarwal, Ankita, et al.
Publicado: (2024)
Interaction-Enabled Two- and Three-Fold Exceptional Points
por: Kato, Musashi, et al.
Publicado: (2026)
por: Kato, Musashi, et al.
Publicado: (2026)
SegDebias: Test-Time Bias Mitigation for ViT-Based CLIP via Segmentation
por: Wu, Fangyu, et al.
Publicado: (2025)
por: Wu, Fangyu, et al.
Publicado: (2025)
Ejemplares similares
-
Debiasing Large Vision-Language Models by Ablating Protected Attribute Representations
por: Ratzlaff, Neale, et al.
Publicado: (2024) -
Steering Large Language Models to Evaluate and Amplify Creativity
por: Olson, Matthew Lyle, et al.
Publicado: (2024) -
Probing the Representational Power of Sparse Autoencoders in Vision Models
por: Olson, Matthew Lyle, et al.
Publicado: (2025) -
Analyzing Hierarchical Structure in Vision Models with Sparse Autoencoders
por: Olson, Matthew Lyle, et al.
Publicado: (2025) -
LieCraft: A Multi-Agent Framework for Evaluating Deceptive Capabilities in Language Models
por: Olson, Matthew Lyle, et al.
Publicado: (2026)