LogicalDefender: Discovering, Extracting, and Utilizing Common-Sense Knowledge
Fuente:
arXiv
Salvato in:
| Autori principali: | Liu, Yuhe, Kang, Mengxue, Qin, Zengchang, Chu, Xiangxiang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Integrating MedCLIP and Cross-Modal Fusion for Automatic Radiology Report Generation
di: Han, Qianhao, et al.
Pubblicazione: (2024)
di: Han, Qianhao, et al.
Pubblicazione: (2024)
TIE: Revolutionizing Text-based Image Editing for Complex-Prompt Following and High-Fidelity Editing
di: Zhang, Xinyu, et al.
Pubblicazione: (2024)
di: Zhang, Xinyu, et al.
Pubblicazione: (2024)
From Image to Video, what do we need in multimodal LLMs?
di: Huang, Suyuan, et al.
Pubblicazione: (2024)
di: Huang, Suyuan, et al.
Pubblicazione: (2024)
USP: Unified Self-Supervised Pretraining for Image Generation and Understanding
di: Chu, Xiangxiang, et al.
Pubblicazione: (2025)
di: Chu, Xiangxiang, et al.
Pubblicazione: (2025)
PLUG: Revisiting Amodal Segmentation with Foundation Model and Hierarchical Focus
di: Liu, Zhaochen, et al.
Pubblicazione: (2024)
di: Liu, Zhaochen, et al.
Pubblicazione: (2024)
ACTRESS: Active Retraining for Semi-supervised Visual Grounding
di: Kang, Weitai, et al.
Pubblicazione: (2024)
di: Kang, Weitai, et al.
Pubblicazione: (2024)
Discovering Intrinsic Spatial-Temporal Logic Rules to Explain Human Actions
di: Cao, Chengzhi, et al.
Pubblicazione: (2023)
di: Cao, Chengzhi, et al.
Pubblicazione: (2023)
FlowDreamer: Exploring High Fidelity Text-to-3D Generation via Rectified Flow
di: Li, Hangyu, et al.
Pubblicazione: (2024)
di: Li, Hangyu, et al.
Pubblicazione: (2024)
VisionLLaMA: A Unified LLaMA Backbone for Vision Tasks
di: Chu, Xiangxiang, et al.
Pubblicazione: (2024)
di: Chu, Xiangxiang, et al.
Pubblicazione: (2024)
Common-Sense Bias Modeling for Classification Tasks
di: Zhang, Miao, et al.
Pubblicazione: (2024)
di: Zhang, Miao, et al.
Pubblicazione: (2024)
Sparse Task Vector Mixup with Hypernetworks for Efficient Knowledge Transfer in Whole-Slide Image Prognosis
di: Liu, Pei, et al.
Pubblicazione: (2026)
di: Liu, Pei, et al.
Pubblicazione: (2026)
Learning Background Prompts to Discover Implicit Knowledge for Open Vocabulary Object Detection
di: Li, Jiaming, et al.
Pubblicazione: (2024)
di: Li, Jiaming, et al.
Pubblicazione: (2024)
Common Sense Reasoning for Deepfake Detection
di: Zhang, Yue, et al.
Pubblicazione: (2024)
di: Zhang, Yue, et al.
Pubblicazione: (2024)
Seeing the Unseen: Visual Common Sense for Semantic Placement
di: Ramrakhya, Ram, et al.
Pubblicazione: (2024)
di: Ramrakhya, Ram, et al.
Pubblicazione: (2024)
What Do You See in Common? Learning Hierarchical Prototypes over Tree-of-Life to Discover Evolutionary Traits
di: Manogaran, Harish Babu, et al.
Pubblicazione: (2024)
di: Manogaran, Harish Babu, et al.
Pubblicazione: (2024)
Revealing the Dark Secrets of Extremely Large Kernel ConvNets on Robustness
di: Chen, Honghao, et al.
Pubblicazione: (2024)
di: Chen, Honghao, et al.
Pubblicazione: (2024)
PeLK: Parameter-efficient Large Kernel ConvNets with Peripheral Convolution
di: Chen, Honghao, et al.
Pubblicazione: (2024)
di: Chen, Honghao, et al.
Pubblicazione: (2024)
Dyn-Adapter: Towards Disentangled Representation for Efficient Visual Recognition
di: Zhang, Yurong, et al.
Pubblicazione: (2024)
di: Zhang, Yurong, et al.
Pubblicazione: (2024)
AdaFedFR: Federated Face Recognition with Adaptive Inter-Class Representation Learning
di: Qiu, Di, et al.
Pubblicazione: (2024)
di: Qiu, Di, et al.
Pubblicazione: (2024)
Elucidating the SNR-t Bias of Diffusion Probabilistic Models
di: Yu, Meng, et al.
Pubblicazione: (2026)
di: Yu, Meng, et al.
Pubblicazione: (2026)
Next Token Is Enough: Realistic Image Quality and Aesthetic Scoring with Multimodal Large Language Model
di: Li, Mingxing, et al.
Pubblicazione: (2025)
di: Li, Mingxing, et al.
Pubblicazione: (2025)
MSA2-Net: Utilizing Self-Adaptive Convolution Module to Extract Multi-Scale Information in Medical Image Segmentation
di: Deng, Chao, et al.
Pubblicazione: (2025)
di: Deng, Chao, et al.
Pubblicazione: (2025)
Telling Stories for Common Sense Zero-Shot Action Recognition
di: Gowda, Shreyank N, et al.
Pubblicazione: (2023)
di: Gowda, Shreyank N, et al.
Pubblicazione: (2023)
One Shot Learning for Edge Detection on Point Clouds
di: Tu, Zhikun, et al.
Pubblicazione: (2026)
di: Tu, Zhikun, et al.
Pubblicazione: (2026)
FarSLIP: Discovering Effective CLIP Adaptation for Fine-Grained Remote Sensing Understanding
di: Li, Zhenshi, et al.
Pubblicazione: (2025)
di: Li, Zhenshi, et al.
Pubblicazione: (2025)
Exposing and Defending the Achilles' Heel of Video Mixture-of-Experts
di: Wang, Songping, et al.
Pubblicazione: (2026)
di: Wang, Songping, et al.
Pubblicazione: (2026)
Semantic Shield: Defending Vision-Language Models Against Backdooring and Poisoning via Fine-grained Knowledge Alignment
di: Ishmam, Alvi Md, et al.
Pubblicazione: (2024)
di: Ishmam, Alvi Md, et al.
Pubblicazione: (2024)
SecureGaze: Defending Gaze Estimation Against Backdoor Attacks
di: Du, Lingyu, et al.
Pubblicazione: (2025)
di: Du, Lingyu, et al.
Pubblicazione: (2025)
Unifying Visual and Vision-Language Tracking via Contrastive Learning
di: Ma, Yinchao, et al.
Pubblicazione: (2024)
di: Ma, Yinchao, et al.
Pubblicazione: (2024)
Intent3D: 3D Object Detection in RGB-D Scans Based on Human Intention
di: Kang, Weitai, et al.
Pubblicazione: (2024)
di: Kang, Weitai, et al.
Pubblicazione: (2024)
3DResT: A Strong Baseline for Semi-Supervised 3D Referring Expression Segmentation
di: Chen, Wenxin, et al.
Pubblicazione: (2025)
di: Chen, Wenxin, et al.
Pubblicazione: (2025)
StoryTailor:A Zero-Shot Pipeline for Action-Rich Multi-Subject Visual Narratives
di: Hu, Jinghao, et al.
Pubblicazione: (2026)
di: Hu, Jinghao, et al.
Pubblicazione: (2026)
Discovering Conceptual Knowledge with Analytic Ontology Templates for Articulated Objects
di: Sun, Jianhua, et al.
Pubblicazione: (2024)
di: Sun, Jianhua, et al.
Pubblicazione: (2024)
Cross-Cancer Knowledge Transfer in WSI-based Prognosis Prediction
di: Liu, Pei, et al.
Pubblicazione: (2025)
di: Liu, Pei, et al.
Pubblicazione: (2025)
Q-Hawkeye: Reliable Visual Policy Optimization for Image Quality Assessment
di: Xie, Wulin, et al.
Pubblicazione: (2026)
di: Xie, Wulin, et al.
Pubblicazione: (2026)
There is No VAE: End-to-End Pixel-Space Generative Modeling via Self-Supervised Pre-training
di: Lei, Jiachen, et al.
Pubblicazione: (2025)
di: Lei, Jiachen, et al.
Pubblicazione: (2025)
Understanding and Defending VLM Jailbreaks via Jailbreak-Related Representation Shift
di: Wei, Zhihua, et al.
Pubblicazione: (2026)
di: Wei, Zhihua, et al.
Pubblicazione: (2026)
Extracting Multimodal Learngene in CLIP: Unveiling the Multimodal Generalizable Knowledge
di: Chen, Ruiming, et al.
Pubblicazione: (2025)
di: Chen, Ruiming, et al.
Pubblicazione: (2025)
Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition
di: Hu, Xiaodan, et al.
Pubblicazione: (2025)
di: Hu, Xiaodan, et al.
Pubblicazione: (2025)
Remote Sensing Object Counting with Online Knowledge Learning
di: Jiang, Shengqin, et al.
Pubblicazione: (2023)
di: Jiang, Shengqin, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Integrating MedCLIP and Cross-Modal Fusion for Automatic Radiology Report Generation
di: Han, Qianhao, et al.
Pubblicazione: (2024) -
TIE: Revolutionizing Text-based Image Editing for Complex-Prompt Following and High-Fidelity Editing
di: Zhang, Xinyu, et al.
Pubblicazione: (2024) -
From Image to Video, what do we need in multimodal LLMs?
di: Huang, Suyuan, et al.
Pubblicazione: (2024) -
USP: Unified Self-Supervised Pretraining for Image Generation and Understanding
di: Chu, Xiangxiang, et al.
Pubblicazione: (2025) -
PLUG: Revisiting Amodal Segmentation with Foundation Model and Hierarchical Focus
di: Liu, Zhaochen, et al.
Pubblicazione: (2024)