Semantic-Guided Natural Language and Visual Fusion for Cross-Modal Interaction Based on Tiny Object Detection
Fuente:
arXiv
Salvato in:
| Autori principali: | Huang, Xian-Hong, Su, Hui-Kai, Sun, Chi-Chia, Hsieh, Jun-Wei |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
RepSFNet : A Single Fusion Network with Structural Reparameterization for Crowd Counting
di: Achmadiah, Mas Nurul, et al.
Pubblicazione: (2026)
di: Achmadiah, Mas Nurul, et al.
Pubblicazione: (2026)
Fast-COS: A Fast One-Stage Object Detector Based on Reparameterized Attention Vision Transformer for Autonomous Driving
di: Setyawan, Novendra, et al.
Pubblicazione: (2025)
di: Setyawan, Novendra, et al.
Pubblicazione: (2025)
TinyFormer: Preserving Tiny Objects in YOLO-DETR Hybrid Real-time Detectors
di: Hsieh, Jun-Wei, et al.
Pubblicazione: (2026)
di: Hsieh, Jun-Wei, et al.
Pubblicazione: (2026)
RS-TinyNet: Stage-wise Feature Fusion Network for Detecting Tiny Objects in Remote Sensing Images
di: Jiang, Xiaozheng, et al.
Pubblicazione: (2025)
di: Jiang, Xiaozheng, et al.
Pubblicazione: (2025)
COMO: Cross-Mamba Interaction and Offset-Guided Fusion for Multimodal Object Detection
di: Liu, Chang, et al.
Pubblicazione: (2024)
di: Liu, Chang, et al.
Pubblicazione: (2024)
AUV-Fusion: Cross-Modal Adversarial Fusion of User Interactions and Visual Perturbations Against VARS
di: Ling, Hai, et al.
Pubblicazione: (2025)
di: Ling, Hai, et al.
Pubblicazione: (2025)
COXNet: Cross-Layer Fusion with Adaptive Alignment and Scale Integration for RGBT Tiny Object Detection
di: Peng, Peiran, et al.
Pubblicazione: (2025)
di: Peng, Peiran, et al.
Pubblicazione: (2025)
MOSA: Music Motion with Semantic Annotation Dataset for Cross-Modal Music Processing
di: Huang, Yu-Fen, et al.
Pubblicazione: (2024)
di: Huang, Yu-Fen, et al.
Pubblicazione: (2024)
DQ-DETR: DETR with Dynamic Query for Tiny Object Detection
di: Huang, Yi-Xin, et al.
Pubblicazione: (2024)
di: Huang, Yi-Xin, et al.
Pubblicazione: (2024)
Integrating Object Detection Modality into Visual Language Model for Enhanced Autonomous Driving Agent
di: He, Linfeng, et al.
Pubblicazione: (2024)
di: He, Linfeng, et al.
Pubblicazione: (2024)
Improving Audio-Visual Speech Recognition by Lip-Subword Correlation Based Visual Pre-training and Cross-Modal Fusion Encoder
di: Dai, Yusheng, et al.
Pubblicazione: (2023)
di: Dai, Yusheng, et al.
Pubblicazione: (2023)
Contrast-Guided Cross-Modal Distillation for Thermal Object Detection
di: Kim, SiWoo, et al.
Pubblicazione: (2025)
di: Kim, SiWoo, et al.
Pubblicazione: (2025)
SONAR: Semantic-Object Navigation with Aggregated Reasoning through a Cross-Modal Inference Paradigm
di: Wang, Yao, et al.
Pubblicazione: (2025)
di: Wang, Yao, et al.
Pubblicazione: (2025)
Dual-Domain Homogeneous Fusion with Cross-Modal Mamba and Progressive Decoder for 3D Object Detection
di: Hu, Xuzhong, et al.
Pubblicazione: (2025)
di: Hu, Xuzhong, et al.
Pubblicazione: (2025)
A DeNoising FPN With Transformer R-CNN for Tiny Object Detection
di: Liu, Hou-I, et al.
Pubblicazione: (2024)
di: Liu, Hou-I, et al.
Pubblicazione: (2024)
InterFusion: Text-Driven Generation of 3D Human-Object Interaction
di: Dai, Sisi, et al.
Pubblicazione: (2024)
di: Dai, Sisi, et al.
Pubblicazione: (2024)
Interacting Null Sources in Different Geometries
di: Hsieh, Chia-Li
Pubblicazione: (2024)
di: Hsieh, Chia-Li
Pubblicazione: (2024)
Thermal-Det: Language-Guided Cross-Modal Distillation for Open-Vocabulary Thermal Object Detection
di: Ranasinghe, Yasiru, et al.
Pubblicazione: (2026)
di: Ranasinghe, Yasiru, et al.
Pubblicazione: (2026)
Energy-Efficient Fast Object Detection on Edge Devices for IoT Systems
di: Achmadiah, Mas Nurul, et al.
Pubblicazione: (2026)
di: Achmadiah, Mas Nurul, et al.
Pubblicazione: (2026)
HGSFusion: Radar-Camera Fusion with Hybrid Generation and Synchronization for 3D Object Detection
di: Gu, Zijian, et al.
Pubblicazione: (2024)
di: Gu, Zijian, et al.
Pubblicazione: (2024)
RCTDistill: Cross-Modal Knowledge Distillation Framework for Radar-Camera 3D Object Detection with Temporal Fusion
di: Bang, Geonho, et al.
Pubblicazione: (2025)
di: Bang, Geonho, et al.
Pubblicazione: (2025)
MicroViTv2: Beyond the FLOPS for Edge Energy-Friendly Vision Transformers
di: Setyawan, Novendra, et al.
Pubblicazione: (2026)
di: Setyawan, Novendra, et al.
Pubblicazione: (2026)
FaceLiVTv2: An Improved Hybrid Architecture for Efficient Mobile Face Recognition
di: Setyawan, Novendra, et al.
Pubblicazione: (2026)
di: Setyawan, Novendra, et al.
Pubblicazione: (2026)
FaceLiVT: Face Recognition using Linear Vision Transformer with Structural Reparameterization For Mobile Device
di: Setyawan, Novendra, et al.
Pubblicazione: (2025)
di: Setyawan, Novendra, et al.
Pubblicazione: (2025)
MicroViT: A Vision Transformer with Low Complexity Self Attention for Edge Device
di: Setyawan, Novendra, et al.
Pubblicazione: (2025)
di: Setyawan, Novendra, et al.
Pubblicazione: (2025)
MANTA: A Large-Scale Multi-View and Visual-Text Anomaly Detection Dataset for Tiny Objects
di: Fan, Lei, et al.
Pubblicazione: (2024)
di: Fan, Lei, et al.
Pubblicazione: (2024)
High-Precision Transformer-Based Visual Servoing for Humanoid Robots in Aligning Tiny Objects
di: Xue, Jialong, et al.
Pubblicazione: (2025)
di: Xue, Jialong, et al.
Pubblicazione: (2025)
STMI: Segmentation-Guided Token Modulation with Cross-Modal Hypergraph Interaction for Multi-Modal Object Re-Identification
di: Xu, Xingguo, et al.
Pubblicazione: (2026)
di: Xu, Xingguo, et al.
Pubblicazione: (2026)
Cross-Modal Bottleneck Fusion For Noise Robust Audio-Visual Speech Recognition
di: Ok, Seaone, et al.
Pubblicazione: (2026)
di: Ok, Seaone, et al.
Pubblicazione: (2026)
Similarity Distance-Based Label Assignment for Tiny Object Detection
di: Shi, Shuohao, et al.
Pubblicazione: (2024)
di: Shi, Shuohao, et al.
Pubblicazione: (2024)
Cross-Modal Purification and Fusion for Small-Object RGB-D Transmission-Line Defect Detection
di: Cui, Jiaming, et al.
Pubblicazione: (2026)
di: Cui, Jiaming, et al.
Pubblicazione: (2026)
Bridging the Scale Gap: Balanced Tiny and General Object Detection in Remote Sensing Imagery
di: Zhao, Zhicheng, et al.
Pubblicazione: (2025)
di: Zhao, Zhicheng, et al.
Pubblicazione: (2025)
HiddenObject: Modality-Agnostic Fusion for Multimodal Hidden Object Detection
di: Song, Harris, et al.
Pubblicazione: (2025)
di: Song, Harris, et al.
Pubblicazione: (2025)
Visual Decision‐Making in Early Childhood Nutrition: Taiwanese Parents′ Infant Formula Choices via Eye‐Tracking and Hierarchical Decision Modeling
di: Chia-Yen Hsieh
Pubblicazione: (2026)
di: Chia-Yen Hsieh
Pubblicazione: (2026)
UFO-DETR: Frequency-Guided End-to-End Detector for UAV Tiny Objects
di: Chen, Yuankai, et al.
Pubblicazione: (2026)
di: Chen, Yuankai, et al.
Pubblicazione: (2026)
VIFO: Visual Feature Empowered Multivariate Time Series Forecasting with Cross-Modal Fusion
di: Wang, Yanlong, et al.
Pubblicazione: (2025)
di: Wang, Yanlong, et al.
Pubblicazione: (2025)
Seg the HAB: Language-Guided Geospatial Algae Bloom Reasoning and Segmentation
di: Hsieh, Patterson, et al.
Pubblicazione: (2025)
di: Hsieh, Patterson, et al.
Pubblicazione: (2025)
CFMD: Dynamic Cross-layer Feature Fusion for Salient Object Detection
di: Lian, Jin, et al.
Pubblicazione: (2025)
di: Lian, Jin, et al.
Pubblicazione: (2025)
Cross-modal Offset-guided Dynamic Alignment and Fusion for Weakly Aligned UAV Object Detection
di: Zongzhen, Liu, et al.
Pubblicazione: (2025)
di: Zongzhen, Liu, et al.
Pubblicazione: (2025)
Dynamic Translational Gains Manipulation for Tiny Object Interaction
di: Jiahui Dong, et al.
Pubblicazione: (2025)
di: Jiahui Dong, et al.
Pubblicazione: (2025)
Documenti analoghi
-
RepSFNet : A Single Fusion Network with Structural Reparameterization for Crowd Counting
di: Achmadiah, Mas Nurul, et al.
Pubblicazione: (2026) -
Fast-COS: A Fast One-Stage Object Detector Based on Reparameterized Attention Vision Transformer for Autonomous Driving
di: Setyawan, Novendra, et al.
Pubblicazione: (2025) -
TinyFormer: Preserving Tiny Objects in YOLO-DETR Hybrid Real-time Detectors
di: Hsieh, Jun-Wei, et al.
Pubblicazione: (2026) -
RS-TinyNet: Stage-wise Feature Fusion Network for Detecting Tiny Objects in Remote Sensing Images
di: Jiang, Xiaozheng, et al.
Pubblicazione: (2025) -
COMO: Cross-Mamba Interaction and Offset-Guided Fusion for Multimodal Object Detection
di: Liu, Chang, et al.
Pubblicazione: (2024)