COMMA: Co-Articulated Multi-Modal Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hu, Lianyu, Gao, Liqing, Liu, Zekang, Pun, Chi-Man, Feng, Wei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Dynamic Spatial-Temporal Aggregation for Skeleton-Aware Sign Language Recognition
von: Hu, Lianyu, et al.
Veröffentlicht: (2024)
von: Hu, Lianyu, et al.
Veröffentlicht: (2024)
Improving Continuous Sign Language Recognition with Adapted Image Models
von: Hu, Lianyu, et al.
Veröffentlicht: (2024)
von: Hu, Lianyu, et al.
Veröffentlicht: (2024)
CorrNet+: Sign Language Recognition and Translation via Spatial-Temporal Correlation
von: Hu, Lianyu, et al.
Veröffentlicht: (2024)
von: Hu, Lianyu, et al.
Veröffentlicht: (2024)
iLLaVA: An Image is Worth Fewer Than 1/3 Input Tokens in Large Multimodal Models
von: Hu, Lianyu, et al.
Veröffentlicht: (2024)
von: Hu, Lianyu, et al.
Veröffentlicht: (2024)
Deep Correlated Prompting for Visual Recognition with Missing Modalities
von: Hu, Lianyu, et al.
Veröffentlicht: (2024)
von: Hu, Lianyu, et al.
Veröffentlicht: (2024)
A Unified Spatial Alignment Framework for Highly Transferable Transformation-Based Attacks on Spatially Structured Tasks
von: Liang, Jiaming, et al.
Veröffentlicht: (2026)
von: Liang, Jiaming, et al.
Veröffentlicht: (2026)
Efficient Preemptive Robustification with Image Sharpening
von: Liang, Jiaming, et al.
Veröffentlicht: (2026)
von: Liang, Jiaming, et al.
Veröffentlicht: (2026)
Efficient Prototype Consistency Learning in Medical Image Segmentation via Joint Uncertainty and Data Augmentation
von: Li, Lijian, et al.
Veröffentlicht: (2025)
von: Li, Lijian, et al.
Veröffentlicht: (2025)
Dynamic Parameter Optimization for Highly Transferable Transformation-Based Attacks
von: Liang, Jiaming, et al.
Veröffentlicht: (2025)
von: Liang, Jiaming, et al.
Veröffentlicht: (2025)
MFDNet: Multi-Frequency Deflare Network for Efficient Nighttime Flare Removal
von: Jiang, Yiguo, et al.
Veröffentlicht: (2024)
von: Jiang, Yiguo, et al.
Veröffentlicht: (2024)
LightVLM: Acceleraing Large Multimodal Models with Pyramid Token Merging and KV Cache Compression
von: Hu, Lianyu, et al.
Veröffentlicht: (2025)
von: Hu, Lianyu, et al.
Veröffentlicht: (2025)
ForgeryTTT: Zero-Shot Image Manipulation Localization with Test-Time Training
von: Liu, Weihuang, et al.
Veröffentlicht: (2024)
von: Liu, Weihuang, et al.
Veröffentlicht: (2024)
OTI: A Model-free and Visually Interpretable Measure of Image Attackability
von: Liang, Jiaming, et al.
Veröffentlicht: (2026)
von: Liang, Jiaming, et al.
Veröffentlicht: (2026)
Co-Evidential Fusion with Information Volume for Medical Image Segmentation
von: He, Yuanpeng, et al.
Veröffentlicht: (2025)
von: He, Yuanpeng, et al.
Veröffentlicht: (2025)
Spatial-Aware Conformal Prediction for Trustworthy Hyperspectral Image Classification
von: Liu, Kangdao, et al.
Veröffentlicht: (2024)
von: Liu, Kangdao, et al.
Veröffentlicht: (2024)
Pose-Guided Fine-Grained Sign Language Video Generation
von: Shi, Tongkai, et al.
Veröffentlicht: (2024)
von: Shi, Tongkai, et al.
Veröffentlicht: (2024)
GaussianHead: High-fidelity Head Avatars with Learnable Gaussian Derivation
von: Wang, Jie, et al.
Veröffentlicht: (2023)
von: Wang, Jie, et al.
Veröffentlicht: (2023)
High-Resolution Document Shadow Removal via A Large-Scale Real-World Dataset and A Frequency-Aware Shadow Erasing Net
von: Li, Zinuo, et al.
Veröffentlicht: (2023)
von: Li, Zinuo, et al.
Veröffentlicht: (2023)
Underwater Image Restoration Through a Prior Guided Hybrid Sense Approach and Extensive Benchmark Analysis
von: Guo, Xiaojiao, et al.
Veröffentlicht: (2025)
von: Guo, Xiaojiao, et al.
Veröffentlicht: (2025)
Combating Pattern and Content Bias: Adversarial Feature Learning for Generalized AI-Generated Image Detection
von: Zhang, Haifeng, et al.
Veröffentlicht: (2026)
von: Zhang, Haifeng, et al.
Veröffentlicht: (2026)
UWFormer: Underwater Image Enhancement via a Semi-Supervised Multi-Scale Transformer
von: Chen, Weiwen, et al.
Veröffentlicht: (2023)
von: Chen, Weiwen, et al.
Veröffentlicht: (2023)
LES-Talker: Fine-Grained Emotion Editing for Talking Head Generation in Linear Emotion Space
von: Feng, Guanwen, et al.
Veröffentlicht: (2024)
von: Feng, Guanwen, et al.
Veröffentlicht: (2024)
Effective Gaussian Management for High-fidelity Object Reconstruction
von: Liu, Jiateng, et al.
Veröffentlicht: (2025)
von: Liu, Jiateng, et al.
Veröffentlicht: (2025)
SMAFormer: Synergistic Multi-Attention Transformer for Medical Image Segmentation
von: Zheng, Fuchen, et al.
Veröffentlicht: (2024)
von: Zheng, Fuchen, et al.
Veröffentlicht: (2024)
RS3Mamba: Visual State Space Model for Remote Sensing Images Semantic Segmentation
von: Ma, Xianping, et al.
Veröffentlicht: (2024)
von: Ma, Xianping, et al.
Veröffentlicht: (2024)
Multi-Modality Co-Learning for Efficient Skeleton-based Action Recognition
von: Liu, Jinfu, et al.
Veröffentlicht: (2024)
von: Liu, Jinfu, et al.
Veröffentlicht: (2024)
Rethinking Multi-Modal Object Detection from the Perspective of Mono-Modality Feature Learning
von: Zhao, Tianyi, et al.
Veröffentlicht: (2025)
von: Zhao, Tianyi, et al.
Veröffentlicht: (2025)
Generalized Uncertainty-Based Evidential Fusion with Hybrid Multi-Head Attention for Weak-Supervised Temporal Action Localization
von: He, Yuanpeng, et al.
Veröffentlicht: (2024)
von: He, Yuanpeng, et al.
Veröffentlicht: (2024)
MLLM-4D: Towards Visual-based Spatial-Temporal Intelligence
von: Yin, Xingyilang, et al.
Veröffentlicht: (2026)
von: Yin, Xingyilang, et al.
Veröffentlicht: (2026)
PersonaLive! Expressive Portrait Image Animation for Live Streaming
von: Li, Zhiyuan, et al.
Veröffentlicht: (2025)
von: Li, Zhiyuan, et al.
Veröffentlicht: (2025)
Dual-Hybrid Attention Network for Specular Highlight Removal
von: Guo, Xiaojiao, et al.
Veröffentlicht: (2024)
von: Guo, Xiaojiao, et al.
Veröffentlicht: (2024)
User-Friendly Customized Generation with Multi-Modal Prompts
von: Zhong, Linhao, et al.
Veröffentlicht: (2024)
von: Zhong, Linhao, et al.
Veröffentlicht: (2024)
Leveraging Arbitrary Data Sources for AI-Generated Image Detection Without Sacrificing Generalization
von: He, Qinghui, et al.
Veröffentlicht: (2026)
von: He, Qinghui, et al.
Veröffentlicht: (2026)
Kolmogorov-Arnold Network for Remote Sensing Image Semantic Segmentation
von: Ma, Xianping, et al.
Veröffentlicht: (2025)
von: Ma, Xianping, et al.
Veröffentlicht: (2025)
COMMA: Coordinate-aware Modulated Mamba Network for 3D Dispersed Vessel Segmentation
von: Shi, Gen, et al.
Veröffentlicht: (2025)
von: Shi, Gen, et al.
Veröffentlicht: (2025)
Mesoscopic Insights: Orchestrating Multi-scale & Hybrid Architecture for Image Manipulation Localization
von: Zhu, Xuekang, et al.
Veröffentlicht: (2024)
von: Zhu, Xuekang, et al.
Veröffentlicht: (2024)
Depth-aware Test-Time Training for Zero-shot Video Object Segmentation
von: Liu, Weihuang, et al.
Veröffentlicht: (2024)
von: Liu, Weihuang, et al.
Veröffentlicht: (2024)
ArticulatedGS: Self-supervised Digital Twin Modeling of Articulated Objects using 3D Gaussian Splatting
von: Guo, Junfu, et al.
Veröffentlicht: (2025)
von: Guo, Junfu, et al.
Veröffentlicht: (2025)
InfoGeo: Information-Theoretic Object-Centric Learning for Cross-View Generalizable UAV Geo-Localization
von: Zhang, Hongyang, et al.
Veröffentlicht: (2026)
von: Zhang, Hongyang, et al.
Veröffentlicht: (2026)
MDReID: Modality-Decoupled Learning for Any-to-Any Multi-Modal Object Re-Identification
von: Feng, Yingying, et al.
Veröffentlicht: (2025)
von: Feng, Yingying, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Dynamic Spatial-Temporal Aggregation for Skeleton-Aware Sign Language Recognition
von: Hu, Lianyu, et al.
Veröffentlicht: (2024) -
Improving Continuous Sign Language Recognition with Adapted Image Models
von: Hu, Lianyu, et al.
Veröffentlicht: (2024) -
CorrNet+: Sign Language Recognition and Translation via Spatial-Temporal Correlation
von: Hu, Lianyu, et al.
Veröffentlicht: (2024) -
iLLaVA: An Image is Worth Fewer Than 1/3 Input Tokens in Large Multimodal Models
von: Hu, Lianyu, et al.
Veröffentlicht: (2024) -
Deep Correlated Prompting for Visual Recognition with Missing Modalities
von: Hu, Lianyu, et al.
Veröffentlicht: (2024)