Saved in:
| Main Authors: | Xie, Jianfei, Li, Ziyang |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2510.01990 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Secure Traffic Sign Recognition: An Attention-Enabled Universal Image Inpainting Mechanism against Light Patch Attacks
by: Cao, Hangcheng, et al.
Published: (2024)
by: Cao, Hangcheng, et al.
Published: (2024)
AlignVTOFF: Texture-Spatial Feature Alignment for High-Fidelity Virtual Try-Off
by: Zhu, Yihan, et al.
Published: (2026)
by: Zhu, Yihan, et al.
Published: (2026)
Synth-Align: Improving Trustworthiness in Vision-Language Model with Synthetic Preference Data Alignment
by: Wijaya, Robert, et al.
Published: (2024)
by: Wijaya, Robert, et al.
Published: (2024)
CARES: A Comprehensive Benchmark of Trustworthiness in Medical Vision Language Models
by: Xia, Peng, et al.
Published: (2024)
by: Xia, Peng, et al.
Published: (2024)
Generalized People Diversity: Learning a Human Perception-Aligned Diversity Representation for People Images
by: Srinivasan, Hansa, et al.
Published: (2024)
by: Srinivasan, Hansa, et al.
Published: (2024)
EnTri: Ensemble Learning with Tri-level Representations for Explainable Scene Recognition
by: Aminimehr, Amirhossein, et al.
Published: (2023)
by: Aminimehr, Amirhossein, et al.
Published: (2023)
Is Grad-CAM Explainable in Medical Images?
by: Suara, Subhashis, et al.
Published: (2023)
by: Suara, Subhashis, et al.
Published: (2023)
A Survey on Trustworthiness in Foundation Models for Medical Image Analysis
by: Shi, Congzhen, et al.
Published: (2024)
by: Shi, Congzhen, et al.
Published: (2024)
Embedding an Ethical Mind: Aligning Text-to-Image Synthesis via Lightweight Value Optimization
by: Wang, Xingqi, et al.
Published: (2024)
by: Wang, Xingqi, et al.
Published: (2024)
Position: Universal Aesthetic Alignment Narrows Artistic Expression
by: Guo, Wenqi Marshall, et al.
Published: (2025)
by: Guo, Wenqi Marshall, et al.
Published: (2025)
Aligning AI with Public Values: Deliberation and Decision-Making for Governing Multimodal LLMs in Political Video Analysis
by: Sharma, Tanusree, et al.
Published: (2024)
by: Sharma, Tanusree, et al.
Published: (2024)
Towards Energy-Efficiency by Navigating the Trilemma of Energy, Latency, and Accuracy
by: Tian, Boyuan, et al.
Published: (2024)
by: Tian, Boyuan, et al.
Published: (2024)
CORAL: Correspondence Alignment for Improved Virtual Try-On
by: Kim, Jiyoung, et al.
Published: (2026)
by: Kim, Jiyoung, et al.
Published: (2026)
SNIFFER: Multimodal Large Language Model for Explainable Out-of-Context Misinformation Detection
by: Qi, Peng, et al.
Published: (2024)
by: Qi, Peng, et al.
Published: (2024)
Detecting Visual Triggers in Cannabis Imagery: A CLIP-Based Multi-Labeling Framework with Local-Global Aggregation
by: Lu, Linqi, et al.
Published: (2024)
by: Lu, Linqi, et al.
Published: (2024)
A Tri-Dynamic Preprocessing Framework for UGC Video Compression
by: Zhao, Fei, et al.
Published: (2025)
by: Zhao, Fei, et al.
Published: (2025)
A Framework for Critical Evaluation of Text-to-Image Models: Integrating Art Historical Analysis, Artistic Exploration, and Critical Prompt Engineering
by: Foka, Amalia
Published: (2024)
by: Foka, Amalia
Published: (2024)
AgriCLIP: Adapting CLIP for Agriculture and Livestock via Domain-Specialized Cross-Model Alignment
by: Nawaz, Umair, et al.
Published: (2024)
by: Nawaz, Umair, et al.
Published: (2024)
Interdisciplinary Expertise to Advance Equitable Explainable AI
by: Bennett, Chloe R., et al.
Published: (2024)
by: Bennett, Chloe R., et al.
Published: (2024)
Multi-Modal Explainable Medical AI Assistant for Trustworthy Human-AI Collaboration
by: Yang, Honglong, et al.
Published: (2025)
by: Yang, Honglong, et al.
Published: (2025)
HyperAlign: Hypernetwork for Efficient Test-Time Alignment of Diffusion Models
by: Xie, Xin, et al.
Published: (2026)
by: Xie, Xin, et al.
Published: (2026)
LiOn-XA: Unsupervised Domain Adaptation via LiDAR-Only Cross-Modal Adversarial Training
by: Kreutz, Thomas, et al.
Published: (2024)
by: Kreutz, Thomas, et al.
Published: (2024)
Modality-Fair Preference Optimization for Trustworthy MLLM Alignment
by: Jiang, Songtao, et al.
Published: (2024)
by: Jiang, Songtao, et al.
Published: (2024)
MORALISE: A Structured Benchmark for Moral Alignment in Visual Language Models
by: Lin, Xiao, et al.
Published: (2025)
by: Lin, Xiao, et al.
Published: (2025)
EmoAssist: Emotional Assistant for Visual Impairment Community
by: Qi, Xingyu, et al.
Published: (2025)
by: Qi, Xingyu, et al.
Published: (2025)
Improved visual-information-driven model for crowd simulation and its modular application
by: Liang, Xuanwen, et al.
Published: (2025)
by: Liang, Xuanwen, et al.
Published: (2025)
Cycle-YOLO: A Efficient and Robust Framework for Pavement Damage Detection
by: Li, Zhengji, et al.
Published: (2024)
by: Li, Zhengji, et al.
Published: (2024)
Mitigating Hallucination in Large Vision-Language Models through Aligning Attention Distribution to Information Flow
by: Zhao, Jianfei, et al.
Published: (2025)
by: Zhao, Jianfei, et al.
Published: (2025)
A Unified Framework and Dataset for Assessing Societal Bias in Vision-Language Models
by: Sathe, Ashutosh, et al.
Published: (2024)
by: Sathe, Ashutosh, et al.
Published: (2024)
Agri-CPJ: A Training-Free Explainable Framework for Agricultural Pest Diagnosis Using Caption-Prompt-Judge and LLM-as-a-Judge
by: Zhang, Wentao, et al.
Published: (2026)
by: Zhang, Wentao, et al.
Published: (2026)
TrajSceneLLM: A Multimodal Perspective on Semantic GPS Trajectory Analysis
by: Ji, Chunhou, et al.
Published: (2025)
by: Ji, Chunhou, et al.
Published: (2025)
ViMU: Benchmarking Video Metaphorical Understanding
by: Li, Qi, et al.
Published: (2026)
by: Li, Qi, et al.
Published: (2026)
Enhancing Manufacturing Quality Prediction Models through the Integration of Explainability Methods
by: Gross, Dennis, et al.
Published: (2024)
by: Gross, Dennis, et al.
Published: (2024)
EgoPrivacy: What Your First-Person Camera Says About You?
by: Li, Yijiang, et al.
Published: (2025)
by: Li, Yijiang, et al.
Published: (2025)
Diagnosing Urban Street Vitality via a Visual-Semantic and Spatiotemporal Framework for Street-Level Economics
by: Zhuo, Xinxin, et al.
Published: (2026)
by: Zhuo, Xinxin, et al.
Published: (2026)
Automatic Teaching Platform on Vision Language Retrieval Augmented Generation
by: Gokhman, Ruslan, et al.
Published: (2025)
by: Gokhman, Ruslan, et al.
Published: (2025)
Towards Reliable Verification of Unauthorized Data Usage in Personalized Text-to-Image Diffusion Models
by: Li, Boheng, et al.
Published: (2024)
by: Li, Boheng, et al.
Published: (2024)
FairRAG: Fair Human Generation via Fair Retrieval Augmentation
by: Shrestha, Robik, et al.
Published: (2024)
by: Shrestha, Robik, et al.
Published: (2024)
STAMP: Multi-pattern Attention-aware Multiple Instance Learning for STAS Diagnosis in Multi-center Histopathology Images
by: Pan, Liangrui, et al.
Published: (2025)
by: Pan, Liangrui, et al.
Published: (2025)
UniFit: Towards Universal Virtual Try-on with MLLM-Guided Semantic Alignment
by: Zhang, Wei, et al.
Published: (2025)
by: Zhang, Wei, et al.
Published: (2025)
Similar Items
-
Secure Traffic Sign Recognition: An Attention-Enabled Universal Image Inpainting Mechanism against Light Patch Attacks
by: Cao, Hangcheng, et al.
Published: (2024) -
AlignVTOFF: Texture-Spatial Feature Alignment for High-Fidelity Virtual Try-Off
by: Zhu, Yihan, et al.
Published: (2026) -
Synth-Align: Improving Trustworthiness in Vision-Language Model with Synthetic Preference Data Alignment
by: Wijaya, Robert, et al.
Published: (2024) -
CARES: A Comprehensive Benchmark of Trustworthiness in Medical Vision Language Models
by: Xia, Peng, et al.
Published: (2024) -
Generalized People Diversity: Learning a Human Perception-Aligned Diversity Representation for People Images
by: Srinivasan, Hansa, et al.
Published: (2024)