Expanding Event Modality Applications through a Robust CLIP-Based Encoder
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Jeong, Sungheon, Chen, Hanning, Yun, Sanggeon, Cho, Suhyeon, Huang, Wenjun, Liu, Xiangjian, Imani, Mohsen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Draft and Refine with Visual Experts
von: Jeong, Sungheon, et al.
Veröffentlicht: (2025)
von: Jeong, Sungheon, et al.
Veröffentlicht: (2025)
EcoSense: Energy-Efficient Intelligent Sensing for In-Shore Ship Detection through Edge-Cloud Collaboration
von: Huang, Wenjun, et al.
Veröffentlicht: (2024)
von: Huang, Wenjun, et al.
Veröffentlicht: (2024)
Uncertainty-Weighted Image-Event Multimodal Fusion for Video Anomaly Detection
von: Jeong, Sungheon, et al.
Veröffentlicht: (2025)
von: Jeong, Sungheon, et al.
Veröffentlicht: (2025)
Fair Context Learning for Evidence-Balanced Test-Time Adaptation in Vision-Language Models
von: Yun, Sanggeon, et al.
Veröffentlicht: (2026)
von: Yun, Sanggeon, et al.
Veröffentlicht: (2026)
TaskCLIP: Extend Large Vision-Language Model for Task Oriented Object Detection
von: Chen, Hanning, et al.
Veröffentlicht: (2024)
von: Chen, Hanning, et al.
Veröffentlicht: (2024)
NeuroHash: A Hyperdimensional Neuro-Symbolic Framework for Spatially-Aware Image Hashing and Retrieval
von: Yun, Sanggeon, et al.
Veröffentlicht: (2024)
von: Yun, Sanggeon, et al.
Veröffentlicht: (2024)
QUILL: An Algorithm-Architecture Co-Design for Cache-Local Deformable Attention
von: Oh, Hyunwoo, et al.
Veröffentlicht: (2025)
von: Oh, Hyunwoo, et al.
Veröffentlicht: (2025)
PV-VTT: A Privacy-Centric Dataset for Mission-Specific Anomaly Detection and Natural Language Interpretation
von: Masukawa, Ryozo, et al.
Veröffentlicht: (2024)
von: Masukawa, Ryozo, et al.
Veröffentlicht: (2024)
Can Multimodal Large Language Models be Guided to Improve Industrial Anomaly Detection?
von: Chen, Zhiling, et al.
Veröffentlicht: (2025)
von: Chen, Zhiling, et al.
Veröffentlicht: (2025)
Internal Flow Signatures for Self-Checking and Refinement in LLMs
von: Jeong, Sungheon, et al.
Veröffentlicht: (2026)
von: Jeong, Sungheon, et al.
Veröffentlicht: (2026)
Tell Me What to Track: Infusing Robust Language Guidance for Enhanced Referring Multi-Object Tracking
von: Huang, Wenjun, et al.
Veröffentlicht: (2024)
von: Huang, Wenjun, et al.
Veröffentlicht: (2024)
Vision Language Model for Interpretable and Fine-grained Detection of Safety Compliance in Diverse Workplaces
von: Chen, Zhiling, et al.
Veröffentlicht: (2024)
von: Chen, Zhiling, et al.
Veröffentlicht: (2024)
LVLM_CSP: Accelerating Large Vision Language Models via Clustering, Scattering, and Pruning for Reasoning Segmentation
von: Chen, Hanning, et al.
Veröffentlicht: (2025)
von: Chen, Hanning, et al.
Veröffentlicht: (2025)
A Plug-in Tiny AI Module for Intelligent and Selective Sensor Data Transmission
von: Huang, Wenjun, et al.
Veröffentlicht: (2024)
von: Huang, Wenjun, et al.
Veröffentlicht: (2024)
MERIT: Multi-domain Efficient RAW Image Translation
von: Huang, Wenjun, et al.
Veröffentlicht: (2026)
von: Huang, Wenjun, et al.
Veröffentlicht: (2026)
PacketCLIP: Multi-Modal Embedding of Network Traffic and Language for Cybersecurity Reasoning
von: Masukawa, Ryozo, et al.
Veröffentlicht: (2025)
von: Masukawa, Ryozo, et al.
Veröffentlicht: (2025)
VLTP: Vision-Language Guided Token Pruning for Task-Oriented Segmentation
von: Chen, Hanning, et al.
Veröffentlicht: (2024)
von: Chen, Hanning, et al.
Veröffentlicht: (2024)
Recoverable Anonymization for Pose Estimation: A Privacy-Enhancing Approach
von: Huang, Wenjun, et al.
Veröffentlicht: (2024)
von: Huang, Wenjun, et al.
Veröffentlicht: (2024)
Leveraging CLIP Encoder for Multimodal Emotion Recognition
von: Song, Yehun, et al.
Veröffentlicht: (2025)
von: Song, Yehun, et al.
Veröffentlicht: (2025)
State-Centric Decision Process
von: Jeong, Sungheon, et al.
Veröffentlicht: (2026)
von: Jeong, Sungheon, et al.
Veröffentlicht: (2026)
Towards Robust Event-based Networks for Nighttime via Unpaired Day-to-Night Event Translation
von: Jeong, Yuhwan, et al.
Veröffentlicht: (2024)
von: Jeong, Yuhwan, et al.
Veröffentlicht: (2024)
Adversarial Robustness for Unified Multi-Modal Encoders via Efficient Calibration
von: Liao, Chih-Ting, et al.
Veröffentlicht: (2025)
von: Liao, Chih-Ting, et al.
Veröffentlicht: (2025)
HopFormer: Sparse Graph Transformers with Explicit Receptive Field Control
von: Yun, Sanggeon, et al.
Veröffentlicht: (2026)
von: Yun, Sanggeon, et al.
Veröffentlicht: (2026)
Robustness in Both Domains: CLIP Needs a Robust Text Encoder
von: Rocamora, Elias Abad, et al.
Veröffentlicht: (2025)
von: Rocamora, Elias Abad, et al.
Veröffentlicht: (2025)
Enhancing CLIP Robustness via Cross-Modality Alignment
von: Zhu, Xingyu, et al.
Veröffentlicht: (2025)
von: Zhu, Xingyu, et al.
Veröffentlicht: (2025)
ViewSplat: View-Adaptive Dynamic Gaussian Splatting for Feed-Forward Synthesis
von: Jeong, Moonyeon, et al.
Veröffentlicht: (2026)
von: Jeong, Moonyeon, et al.
Veröffentlicht: (2026)
DSERT-RoLL: Robust Multi-Modal Perception for Diverse Driving Conditions with Stereo Event-RGB-Thermal Cameras, 4D Radar, and Dual-LiDAR
von: Cho, Hoonhee, et al.
Veröffentlicht: (2026)
von: Cho, Hoonhee, et al.
Veröffentlicht: (2026)
Self-Attention Based Semantic Decomposition in Vector Symbolic Architectures
von: Yeung, Calvin, et al.
Veröffentlicht: (2024)
von: Yeung, Calvin, et al.
Veröffentlicht: (2024)
CLIP-KOA: Enhancing Knee Osteoarthritis Diagnosis with Multi-Modal Learning and Symmetry-Aware Loss Functions
von: Jeong, Yejin, et al.
Veröffentlicht: (2025)
von: Jeong, Yejin, et al.
Veröffentlicht: (2025)
Reangle-A-Video: 4D Video Generation as Video-to-Video Translation
von: Jeong, Hyeonho, et al.
Veröffentlicht: (2025)
von: Jeong, Hyeonho, et al.
Veröffentlicht: (2025)
CMTA: Cross-Modal Temporal Alignment for Event-guided Video Deblurring
von: Kim, Taewoo, et al.
Veröffentlicht: (2024)
von: Kim, Taewoo, et al.
Veröffentlicht: (2024)
Enhancing Temporal Understanding in Video-LLMs through Stacked Temporal Attention in Vision Encoders
von: Rasekh, Ali, et al.
Veröffentlicht: (2025)
von: Rasekh, Ali, et al.
Veröffentlicht: (2025)
Reevaluating the Intra-Modal Misalignment Hypothesis in CLIP
von: Herzog, Jonas, et al.
Veröffentlicht: (2026)
von: Herzog, Jonas, et al.
Veröffentlicht: (2026)
SPACE-CLIP: Spatial Perception via Adaptive CLIP Embeddings for Monocular Depth Estimation
von: Cho, Taewan, et al.
Veröffentlicht: (2026)
von: Cho, Taewan, et al.
Veröffentlicht: (2026)
CEIA: CLIP-Based Event-Image Alignment for Open-World Event-Based Understanding
von: Xu, Wenhao, et al.
Veröffentlicht: (2024)
von: Xu, Wenhao, et al.
Veröffentlicht: (2024)
MMeViT: Multi-Modal ensemble ViT for Post-Stroke Rehabilitation Action Recognition
von: Kim, Ye-eun, et al.
Veröffentlicht: (2025)
von: Kim, Ye-eun, et al.
Veröffentlicht: (2025)
Mind the Gap: Preserving and Compensating for the Modality Gap in CLIP-Based Continual Learning
von: Huang, Linlan, et al.
Veröffentlicht: (2025)
von: Huang, Linlan, et al.
Veröffentlicht: (2025)
Dual-Prompt CLIP with Hybrid Visual Encoders for Occluded Person Re-Identification
von: Ji, Zhangjian, et al.
Veröffentlicht: (2026)
von: Ji, Zhangjian, et al.
Veröffentlicht: (2026)
Towards Real-world Event-guided Low-light Video Enhancement and Deblurring
von: Kim, Taewoo, et al.
Veröffentlicht: (2024)
von: Kim, Taewoo, et al.
Veröffentlicht: (2024)
Quantization Robustness to Input Degradations for Object Detection
von: Karimov, Toghrul, et al.
Veröffentlicht: (2025)
von: Karimov, Toghrul, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Draft and Refine with Visual Experts
von: Jeong, Sungheon, et al.
Veröffentlicht: (2025) -
EcoSense: Energy-Efficient Intelligent Sensing for In-Shore Ship Detection through Edge-Cloud Collaboration
von: Huang, Wenjun, et al.
Veröffentlicht: (2024) -
Uncertainty-Weighted Image-Event Multimodal Fusion for Video Anomaly Detection
von: Jeong, Sungheon, et al.
Veröffentlicht: (2025) -
Fair Context Learning for Evidence-Balanced Test-Time Adaptation in Vision-Language Models
von: Yun, Sanggeon, et al.
Veröffentlicht: (2026) -
TaskCLIP: Extend Large Vision-Language Model for Task Oriented Object Detection
von: Chen, Hanning, et al.
Veröffentlicht: (2024)