Saved in:
| Main Authors: | Yang, Ziyuan, Yan, Ming, Chen, Yingyu, Wang, Hui, Lu, Zexin, Zhang, Yi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2409.13557 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Detecting and Understanding Hateful Contents in Memes Through Captioning and Visual Question-Answering
by: Anaissi, Ali, et al.
Published: (2025)
by: Anaissi, Ali, et al.
Published: (2025)
FedPalm: A General Federated Learning Framework for Closed- and Open-Set Palmprint Verification
by: Yang, Ziyuan, et al.
Published: (2025)
by: Yang, Ziyuan, et al.
Published: (2025)
Improving Medical Visual Reinforcement Fine-Tuning via Perception and Reasoning Augmentation
by: Yang, Guangjing, et al.
Published: (2026)
by: Yang, Guangjing, et al.
Published: (2026)
Reasoning-Aware Multimodal Fusion for Hateful Video Detection
by: Yang, Shuonan, et al.
Published: (2025)
by: Yang, Shuonan, et al.
Published: (2025)
One CT Unified Model Training Framework to Rule All Scanning Protocols
by: Xu, Fengzhi, et al.
Published: (2026)
by: Xu, Fengzhi, et al.
Published: (2026)
Hypernetwork-based Personalized Federated Learning for Multi-Institutional CT Imaging
by: Yang, Ziyuan, et al.
Published: (2022)
by: Yang, Ziyuan, et al.
Published: (2022)
Robust Adaptation of Large Multimodal Models for Retrieval Augmented Hateful Meme Detection
by: Mei, Jingbiao, et al.
Published: (2025)
by: Mei, Jingbiao, et al.
Published: (2025)
Targeted Downstream-Agnostic Attack
by: Lei, Zhuxin, et al.
Published: (2026)
by: Lei, Zhuxin, et al.
Published: (2026)
Augmentation Matters: A Mix-Paste Method for X-Ray Prohibited Item Detection under Noisy Annotations
by: Chen, Ruikang, et al.
Published: (2025)
by: Chen, Ruikang, et al.
Published: (2025)
Double Banking on Knowledge: Customized Modulation and Prototypes for Multi-Modality Semi-supervised Medical Image Segmentation
by: Chen, Yingyu, et al.
Published: (2024)
by: Chen, Yingyu, et al.
Published: (2024)
TSHA: A Benchmark for Visual Language Models in Trustworthy Safety Hazard Assessment Scenarios
by: Yu, Qiucheng, et al.
Published: (2026)
by: Yu, Qiucheng, et al.
Published: (2026)
Enhanced Multimodal Hate Video Detection via Channel-wise and Modality-wise Fusion
by: Zhang, Yinghui, et al.
Published: (2025)
by: Zhang, Yinghui, et al.
Published: (2025)
DiffusionTalker: Efficient and Compact Speech-Driven 3D Talking Head via Personalizer-Guided Distillation
by: Chen, Peng, et al.
Published: (2025)
by: Chen, Peng, et al.
Published: (2025)
Revealing Temporal Label Noise in Multimodal Hateful Video Classification
by: Yang, Shuonan, et al.
Published: (2025)
by: Yang, Shuonan, et al.
Published: (2025)
ViFP: A Framework for Visual False Positive Detection to Enhance Reasoning Reliability in VLMs
by: Zhang, Ben, et al.
Published: (2025)
by: Zhang, Ben, et al.
Published: (2025)
HOFAR: High-Order Augmentation of Flow Autoregressive Transformers
by: Liang, Yingyu, et al.
Published: (2025)
by: Liang, Yingyu, et al.
Published: (2025)
Universal Approximation of Visual Autoregressive Transformers
by: Chen, Yifang, et al.
Published: (2025)
by: Chen, Yifang, et al.
Published: (2025)
Seeing Through the Brain: New Insights from Decoding Visual Stimuli with fMRI
by: Huang, Zheng, et al.
Published: (2025)
by: Huang, Zheng, et al.
Published: (2025)
A Novel Trustworthy Video Summarization Algorithm Through a Mixture of LoRA Experts
by: Du, Wenzhuo, et al.
Published: (2025)
by: Du, Wenzhuo, et al.
Published: (2025)
MultiHateClip: A Multilingual Benchmark Dataset for Hateful Video Detection on YouTube and Bilibili
by: Wang, Han, et al.
Published: (2024)
by: Wang, Han, et al.
Published: (2024)
Detecting Visual Cues in the Intensive Care Unit and Association with Patient Clinical Status
by: Nerella, Subhash, et al.
Published: (2023)
by: Nerella, Subhash, et al.
Published: (2023)
Seeing is Not Reasoning: MVPBench for Graph-based Evaluation of Multi-path Visual Physical CoT
by: Dong, Zhuobai, et al.
Published: (2025)
by: Dong, Zhuobai, et al.
Published: (2025)
SeMOPO: Learning High-quality Model and Policy from Low-quality Offline Visual Datasets
by: Wan, Shenghua, et al.
Published: (2024)
by: Wan, Shenghua, et al.
Published: (2024)
A Comprehensive Augmentation Framework for Anomaly Detection
by: Lin, Jiang, et al.
Published: (2023)
by: Lin, Jiang, et al.
Published: (2023)
Visual Autoregressive Transformers Must Use $Ω(n^2 d)$ Memory
by: Cao, Yang, et al.
Published: (2025)
by: Cao, Yang, et al.
Published: (2025)
VCU-Bridge: Hierarchical Visual Connotation Understanding via Semantic Bridging
by: Zhong, Ming, et al.
Published: (2025)
by: Zhong, Ming, et al.
Published: (2025)
VG-CoT: Towards Trustworthy Visual Reasoning via Grounded Chain-of-Thought
by: Lim, Byeonggeuk, et al.
Published: (2026)
by: Lim, Byeonggeuk, et al.
Published: (2026)
VidLBEval: Benchmarking and Mitigating Language Bias in Video-Involved LVLMs
by: Yang, Yiming, et al.
Published: (2025)
by: Yang, Yiming, et al.
Published: (2025)
Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models
by: Van, Minh-Hao, et al.
Published: (2025)
by: Van, Minh-Hao, et al.
Published: (2025)
CLEAR-Mamba:Towards Accurate, Adaptive and Trustworthy Multi-Sequence Ophthalmic Angiography Classification
by: Wang, Zhuonan, et al.
Published: (2026)
by: Wang, Zhuonan, et al.
Published: (2026)
HieraTok: Multi-Scale Visual Tokenizer Improves Image Reconstruction and Generation
by: Chen, Cong, et al.
Published: (2025)
by: Chen, Cong, et al.
Published: (2025)
SCRA-VQA: Summarized Caption-Rerank for Augmented Large Language Models in Visual Question Answering
by: Zhang, Yan, et al.
Published: (2025)
by: Zhang, Yan, et al.
Published: (2025)
SkinGPT-X: A Self-Evolving Collaborative Multi-Agent System for Transparent and Trustworthy Dermatological Diagnosis
by: Chen, Zhangtianyi, et al.
Published: (2026)
by: Chen, Zhangtianyi, et al.
Published: (2026)
Deep Learning in Palmprint Recognition-A Comprehensive Survey
by: Gao, Chengrui, et al.
Published: (2025)
by: Gao, Chengrui, et al.
Published: (2025)
OVOD-Agent: A Markov-Bandit Framework for Proactive Visual Reasoning and Self-Evolving Detection
by: Wang, Chujie, et al.
Published: (2025)
by: Wang, Chujie, et al.
Published: (2025)
FBHM: Functional Benchmarking and Steering of VLMs for Hateful Meme Detection
by: Bhaskar, Paramananda, et al.
Published: (2026)
by: Bhaskar, Paramananda, et al.
Published: (2026)
VACoT: Rethinking Visual Data Augmentation with VLMs
by: Xu, Zhengzhuo, et al.
Published: (2025)
by: Xu, Zhengzhuo, et al.
Published: (2025)
AdaptCLIP: Adapting CLIP for Universal Visual Anomaly Detection
by: Gao, Bin-Bin, et al.
Published: (2025)
by: Gao, Bin-Bin, et al.
Published: (2025)
On the Black-box Explainability of Object Detection Models for Safe and Trustworthy Industrial Applications
by: Andres, Alain, et al.
Published: (2024)
by: Andres, Alain, et al.
Published: (2024)
KnowVal: A Knowledge-Augmented and Value-Guided Autonomous Driving System
by: Xia, Zhongyu, et al.
Published: (2025)
by: Xia, Zhongyu, et al.
Published: (2025)
Similar Items
-
Detecting and Understanding Hateful Contents in Memes Through Captioning and Visual Question-Answering
by: Anaissi, Ali, et al.
Published: (2025) -
FedPalm: A General Federated Learning Framework for Closed- and Open-Set Palmprint Verification
by: Yang, Ziyuan, et al.
Published: (2025) -
Improving Medical Visual Reinforcement Fine-Tuning via Perception and Reasoning Augmentation
by: Yang, Guangjing, et al.
Published: (2026) -
Reasoning-Aware Multimodal Fusion for Hateful Video Detection
by: Yang, Shuonan, et al.
Published: (2025) -
One CT Unified Model Training Framework to Rule All Scanning Protocols
by: Xu, Fengzhi, et al.
Published: (2026)