WISE: Weak-Supervision-Guided Step-by-Step Explanations for Multimodal LLMs in Image Classification
Fuente:
arXiv
Saved in:
| Main Authors: | Jiang, Yiwen, Mehta, Deval, Yan, Siyuan, Shen, Yaling, Wang, Zimu, Ge, Zongyuan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Enhancing Interpretable Image Classification Through LLM Agents and Conditional Concept Bottleneck Models
by: Jiang, Yiwen, et al.
Published: (2025)
by: Jiang, Yiwen, et al.
Published: (2025)
Neurosymbolic Framework for Concept-Driven Logical Reasoning in Skeleton-Based Human Action Recognition
by: Ilyas, Talha, et al.
Published: (2026)
by: Ilyas, Talha, et al.
Published: (2026)
Interpretable Few-Shot Retinal Disease Diagnosis with Concept-Guided Prompting of Vision-Language Models
by: Mehta, Deval, et al.
Published: (2025)
by: Mehta, Deval, et al.
Published: (2025)
Synthesize Step-by-Step: Tools, Templates and LLMs as Data Generators for Reasoning-Based Chart VQA
by: Li, Zhuowan, et al.
Published: (2024)
by: Li, Zhuowan, et al.
Published: (2024)
Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation
by: Lu, Jinghui, et al.
Published: (2026)
by: Lu, Jinghui, et al.
Published: (2026)
One-Step Diffusion for Perceptual Image Compression
by: Jia, Yiwen, et al.
Published: (2026)
by: Jia, Yiwen, et al.
Published: (2026)
MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
by: Zhang, Yi-Fan, et al.
Published: (2025)
by: Zhang, Yi-Fan, et al.
Published: (2025)
Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step
by: Guo, Ziyu, et al.
Published: (2025)
by: Guo, Ziyu, et al.
Published: (2025)
Weakly Supervised Gaussian Contrastive Grounding with Large Multimodal Models for Video Question Answering
by: Wang, Haibo, et al.
Published: (2024)
by: Wang, Haibo, et al.
Published: (2024)
MAKE: Multi-Aspect Knowledge-Enhanced Vision-Language Pretraining for Zero-shot Dermatological Assessment
by: Yan, Siyuan, et al.
Published: (2025)
by: Yan, Siyuan, et al.
Published: (2025)
CHiRPE: A Step Towards Real-World Clinical NLP with Clinician-Oriented Model Explanations
by: Fong, Stephanie, et al.
Published: (2026)
by: Fong, Stephanie, et al.
Published: (2026)
DiSA: Diffusion Step Annealing in Autoregressive Image Generation
by: Zhao, Qinyu, et al.
Published: (2025)
by: Zhao, Qinyu, et al.
Published: (2025)
Do No Harm: Exposing Hidden Vulnerabilities of LLMs via Persona-based Client Simulation Attack in Psychological Counseling
by: Xu, Qingyang, et al.
Published: (2026)
by: Xu, Qingyang, et al.
Published: (2026)
Whiteboard-of-Thought: Thinking Step-by-Step Across Modalities
by: Menon, Sachit, et al.
Published: (2024)
by: Menon, Sachit, et al.
Published: (2024)
EVALALIGN: Supervised Fine-Tuning Multimodal LLMs with Human-Aligned Data for Evaluating Text-to-Image Models
by: Tan, Zhiyu, et al.
Published: (2024)
by: Tan, Zhiyu, et al.
Published: (2024)
Universal Semi-Supervised Learning for Medical Image Classification
by: Ju, Lie, et al.
Published: (2023)
by: Ju, Lie, et al.
Published: (2023)
Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification
by: Bai, Tianyi, et al.
Published: (2025)
by: Bai, Tianyi, et al.
Published: (2025)
Harnessing Shared Relations via Multimodal Mixup Contrastive Learning for Multimodal Classification
by: Kumar, Raja, et al.
Published: (2024)
by: Kumar, Raja, et al.
Published: (2024)
Deep Intra-Image Contrastive Learning for Weakly Supervised One-Step Person Search
by: Wang, Jiabei, et al.
Published: (2023)
by: Wang, Jiabei, et al.
Published: (2023)
Guided Score identity Distillation for Data-Free One-Step Text-to-Image Generation
by: Zhou, Mingyuan, et al.
Published: (2024)
by: Zhou, Mingyuan, et al.
Published: (2024)
Chain-of-Jailbreak Attack for Image Generation Models via Editing Step by Step
by: Wang, Wenxuan, et al.
Published: (2024)
by: Wang, Wenxuan, et al.
Published: (2024)
Clinical Cognition Alignment for Gastrointestinal Diagnosis with Multimodal LLMs
by: Zheng, Huan, et al.
Published: (2026)
by: Zheng, Huan, et al.
Published: (2026)
Step-Video-TI2V Technical Report: A State-of-the-Art Text-Driven Image-to-Video Generation Model
by: Huang, Haoyang, et al.
Published: (2025)
by: Huang, Haoyang, et al.
Published: (2025)
ChestX-Reasoner: Advancing Radiology Foundation Models with Reasoning through Step-by-Step Verification
by: Fan, Ziqing, et al.
Published: (2025)
by: Fan, Ziqing, et al.
Published: (2025)
Memory-Augmented Multimodal LLMs for Surgical VQA via Self-Contained Inquiry
by: Hou, Wenjun, et al.
Published: (2024)
by: Hou, Wenjun, et al.
Published: (2024)
Enhancing Vision-Language Models Generalization via Diversity-Driven Novel Feature Synthesis
by: Yan, Siyuan, et al.
Published: (2024)
by: Yan, Siyuan, et al.
Published: (2024)
IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs
by: Ma, David, et al.
Published: (2025)
by: Ma, David, et al.
Published: (2025)
Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
by: Ma, Guoqing, et al.
Published: (2025)
by: Ma, Guoqing, et al.
Published: (2025)
Hunting Attributes: Context Prototype-Aware Learning for Weakly Supervised Semantic Segmentation
by: Tang, Feilong, et al.
Published: (2024)
by: Tang, Feilong, et al.
Published: (2024)
Derm1M: A Million-scale Vision-Language Dataset Aligned with Clinical Ontology Knowledge for Dermatology
by: Yan, Siyuan, et al.
Published: (2025)
by: Yan, Siyuan, et al.
Published: (2025)
MONICA: Benchmarking on Long-tailed Medical Image Classification
by: Ju, Lie, et al.
Published: (2024)
by: Ju, Lie, et al.
Published: (2024)
It Hears, It Sees too: Multi-Modal LLM for Depression Detection By Integrating Visual Understanding into Audio Language Models
by: Zhao, Xiangyu, et al.
Published: (2025)
by: Zhao, Xiangyu, et al.
Published: (2025)
Weakly-Supervised 3D Visual Grounding based on Visual Language Alignment
by: Xu, Xiaoxu, et al.
Published: (2023)
by: Xu, Xiaoxu, et al.
Published: (2023)
Beyond Image-Text Matching: Verb Understanding in Multimodal Transformers Using Guided Masking
by: Beňová, Ivana, et al.
Published: (2024)
by: Beňová, Ivana, et al.
Published: (2024)
ViCA: Efficient Multimodal LLMs with Vision-Only Cross-Attention
by: Liu, Wenjie, et al.
Published: (2026)
by: Liu, Wenjie, et al.
Published: (2026)
CHiP: Cross-modal Hierarchical Direct Preference Optimization for Multimodal LLMs
by: Fu, Jinlan, et al.
Published: (2025)
by: Fu, Jinlan, et al.
Published: (2025)
TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs
by: Zhang, Jun, et al.
Published: (2025)
by: Zhang, Jun, et al.
Published: (2025)
Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data Generation
by: Yang, Yue, et al.
Published: (2025)
by: Yang, Yue, et al.
Published: (2025)
Leveraging LLMs for On-the-Fly Instruction Guided Image Editing
by: Santos, Rodrigo, et al.
Published: (2024)
by: Santos, Rodrigo, et al.
Published: (2024)
UniVLR: Unifying Text and Vision in Visual Latent Reasoning for Multimodal LLMs
by: Jiang, Houcheng, et al.
Published: (2026)
by: Jiang, Houcheng, et al.
Published: (2026)
Similar Items
-
Enhancing Interpretable Image Classification Through LLM Agents and Conditional Concept Bottleneck Models
by: Jiang, Yiwen, et al.
Published: (2025) -
Neurosymbolic Framework for Concept-Driven Logical Reasoning in Skeleton-Based Human Action Recognition
by: Ilyas, Talha, et al.
Published: (2026) -
Interpretable Few-Shot Retinal Disease Diagnosis with Concept-Guided Prompting of Vision-Language Models
by: Mehta, Deval, et al.
Published: (2025) -
Synthesize Step-by-Step: Tools, Templates and LLMs as Data Generators for Reasoning-Based Chart VQA
by: Li, Zhuowan, et al.
Published: (2024) -
Xiaomi OneVL: One-Step Latent Reasoning and Planning with Vision-Language Explanation
by: Lu, Jinghui, et al.
Published: (2026)