CrisisViT: A Robust Vision Transformer for Crisis Image Classification
Fuente:
arXiv
Saved in:
| Main Authors: | Long, Zijun, McCreadie, Richard, Imran, Muhammad |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MultiWay-Adapater: Adapting large-scale multi-modal models for scalable image-text retrieval
by: Long, Zijun, et al.
Published: (2023)
by: Long, Zijun, et al.
Published: (2023)
LaCViT: A Label-aware Contrastive Fine-tuning Framework for Vision Transformers
by: Long, Zijun, et al.
Published: (2023)
by: Long, Zijun, et al.
Published: (2023)
SMP Challenge: An Overview and Analysis of Social Media Prediction Challenge
by: Wu, Bo, et al.
Published: (2024)
by: Wu, Bo, et al.
Published: (2024)
HSVLT: Hierarchical Scale-Aware Vision-Language Transformer for Multi-Label Image Classification
by: Ouyang, Shuyi, et al.
Published: (2024)
by: Ouyang, Shuyi, et al.
Published: (2024)
A Rate-Distortion-Classification Approach for Lossy Image Compression
by: Zhang, Yuefeng
Published: (2024)
by: Zhang, Yuefeng
Published: (2024)
Robust Latent Representation Tuning for Image-text Classification
by: Sun, Hao, et al.
Published: (2024)
by: Sun, Hao, et al.
Published: (2024)
Understanding and Mitigating Human-Labelling Errors in Supervised Contrastive Learning
by: Long, Zijun, et al.
Published: (2024)
by: Long, Zijun, et al.
Published: (2024)
A Matter of Time: Revealing the Structure of Time in Vision-Language Models
by: Tekaya, Nidham, et al.
Published: (2025)
by: Tekaya, Nidham, et al.
Published: (2025)
Epistemic-aware Vision-Language Foundation Model for Fetal Ultrasound Interpretation
by: He, Xiao, et al.
Published: (2025)
by: He, Xiao, et al.
Published: (2025)
RoboLLM: Robotic Vision Tasks Grounded on Multimodal Large Language Models
by: Long, Zijun, et al.
Published: (2023)
by: Long, Zijun, et al.
Published: (2023)
GMMFormer: Gaussian-Mixture-Model Based Transformer for Efficient Partially Relevant Video Retrieval
by: Wang, Yuting, et al.
Published: (2023)
by: Wang, Yuting, et al.
Published: (2023)
A Comprehensive Survey on Composed Image Retrieval
by: Song, Xuemeng, et al.
Published: (2025)
by: Song, Xuemeng, et al.
Published: (2025)
VQualA 2025 Challenge on Engagement Prediction for Short Videos: Methods and Results
by: Li, Dasong, et al.
Published: (2025)
by: Li, Dasong, et al.
Published: (2025)
Multi-modal Misinformation Detection: Approaches, Challenges and Opportunities
by: Abdali, Sara, et al.
Published: (2022)
by: Abdali, Sara, et al.
Published: (2022)
Image Complexity-Aware Adaptive Retrieval for Efficient Vision-Language Models
by: Williams-Lekuona, Mikel, et al.
Published: (2025)
by: Williams-Lekuona, Mikel, et al.
Published: (2025)
Advance Fake Video Detection via Vision Transformers
by: Battocchio, Joy, et al.
Published: (2025)
by: Battocchio, Joy, et al.
Published: (2025)
Feature CAM: Interpretable AI in Image Classification
by: Clement, Frincy, et al.
Published: (2024)
by: Clement, Frincy, et al.
Published: (2024)
Delving Deep into Engagement Prediction of Short Videos
by: Li, Dasong, et al.
Published: (2024)
by: Li, Dasong, et al.
Published: (2024)
LoViF 2026 The First Challenge on Weather Removal in Videos
by: Qian, Chenghao, et al.
Published: (2026)
by: Qian, Chenghao, et al.
Published: (2026)
VGA: Vision and Graph Fused Attention Network for Rumor Detection
by: Bai, Lin, et al.
Published: (2024)
by: Bai, Lin, et al.
Published: (2024)
Learning to Mask and Permute Visual Tokens for Vision Transformer Pre-Training
by: Baraldi, Lorenzo, et al.
Published: (2023)
by: Baraldi, Lorenzo, et al.
Published: (2023)
ViTMAlis: Towards Latency-Critical Mobile Video Analytics with Vision Transformers
by: Zhang, Miao, et al.
Published: (2026)
by: Zhang, Miao, et al.
Published: (2026)
Recurrence-Enhanced Vision-and-Language Transformers for Robust Multimodal Document Retrieval
by: Caffagni, Davide, et al.
Published: (2025)
by: Caffagni, Davide, et al.
Published: (2025)
TPC-ViT: Token Propagation Controller for Efficient Vision Transformer
by: Zhu, Wentao
Published: (2024)
by: Zhu, Wentao
Published: (2024)
End-to-End Optimized Image Compression with the Frequency-Oriented Transform
by: Zhang, Yuefeng, et al.
Published: (2024)
by: Zhang, Yuefeng, et al.
Published: (2024)
TIDE : Temporal-Aware Sparse Autoencoders for Interpretable Diffusion Transformers in Image Generation
by: Huang, Victor Shea-Jay, et al.
Published: (2025)
by: Huang, Victor Shea-Jay, et al.
Published: (2025)
Ges3ViG: Incorporating Pointing Gestures into Language-Based 3D Visual Grounding for Embodied Reference Understanding
by: Mane, Atharv Mahesh, et al.
Published: (2025)
by: Mane, Atharv Mahesh, et al.
Published: (2025)
MotionDreamer: One-to-Many Motion Synthesis with Localized Generative Masked Transformer
by: Wang, Yilin, et al.
Published: (2025)
by: Wang, Yilin, et al.
Published: (2025)
PhotoBench: Beyond Visual Matching Towards Personalized Intent-Driven Photo Retrieval
by: Xu, Tianyi, et al.
Published: (2026)
by: Xu, Tianyi, et al.
Published: (2026)
Cross-Modal Retrieval with Cauchy-Schwarz Divergence
by: Zhang, Jiahao, et al.
Published: (2025)
by: Zhang, Jiahao, et al.
Published: (2025)
The CASTLE 2024 Dataset: Advancing the Art of Multimodal Understanding
by: Rossetto, Luca, et al.
Published: (2025)
by: Rossetto, Luca, et al.
Published: (2025)
Multi-Modal Cross-Domain Alignment Network for Video Moment Retrieval
by: Fang, Xiang, et al.
Published: (2022)
by: Fang, Xiang, et al.
Published: (2022)
Beyond Final Answers: CRYSTAL Benchmark for Transparent Multimodal Reasoning Evaluation
by: Barrios, Wayner, et al.
Published: (2026)
by: Barrios, Wayner, et al.
Published: (2026)
VeriTaS: The First Dynamic Benchmark for Multimodal Automated Fact-Checking
by: Rothermel, Mark, et al.
Published: (2026)
by: Rothermel, Mark, et al.
Published: (2026)
Very Efficient Listwise Multimodal Reranking for Long Documents
by: Sun, Yiqun, et al.
Published: (2026)
by: Sun, Yiqun, et al.
Published: (2026)
An Empirical Study of Excitation and Aggregation Design Adaptions in CLIP4Clip for Video-Text Retrieval
by: Jing, Xiaolun, et al.
Published: (2024)
by: Jing, Xiaolun, et al.
Published: (2024)
PromptHash: Affinity-Prompted Collaborative Cross-Modal Learning for Adaptive Hashing Retrieval
by: Zou, Qiang, et al.
Published: (2025)
by: Zou, Qiang, et al.
Published: (2025)
MIRe: Enhancing Multimodal Queries Representation via Fusion-Free Modality Interaction for Multimodal Retrieval
by: Ju, Yeong-Joon, et al.
Published: (2024)
by: Ju, Yeong-Joon, et al.
Published: (2024)
A Survey of Multimodal Composite Editing and Retrieval
by: Li, Suyan, et al.
Published: (2024)
by: Li, Suyan, et al.
Published: (2024)
CLIP-PING: Boosting Lightweight Vision-Language Models with Proximus Intrinsic Neighbors Guidance
by: Thwal, Chu Myaet, et al.
Published: (2024)
by: Thwal, Chu Myaet, et al.
Published: (2024)
Similar Items
-
MultiWay-Adapater: Adapting large-scale multi-modal models for scalable image-text retrieval
by: Long, Zijun, et al.
Published: (2023) -
LaCViT: A Label-aware Contrastive Fine-tuning Framework for Vision Transformers
by: Long, Zijun, et al.
Published: (2023) -
SMP Challenge: An Overview and Analysis of Social Media Prediction Challenge
by: Wu, Bo, et al.
Published: (2024) -
HSVLT: Hierarchical Scale-Aware Vision-Language Transformer for Multi-Label Image Classification
by: Ouyang, Shuyi, et al.
Published: (2024) -
A Rate-Distortion-Classification Approach for Lossy Image Compression
by: Zhang, Yuefeng
Published: (2024)