Saved in:
| Main Authors: | Kiyokawa, Takuya, Shirakura, Naoki, Katayama, Hiroki, Tomochika, Keita, Takamatsu, Jun |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2304.04901 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Dual Mutual Learning Network with Global-local Awareness for RGB-D Salient Object Detection
by: Yi, Kang, et al.
Published: (2025)
by: Yi, Kang, et al.
Published: (2025)
HMPE:HeatMap Embedding for Efficient Transformer-Based Small Object Detection
by: Zeng, YangChen
Published: (2025)
by: Zeng, YangChen
Published: (2025)
TextRefiner: Internal Visual Feature as Efficient Refiner for Vision-Language Models Prompt Tuning
by: Xie, Jingjing, et al.
Published: (2024)
by: Xie, Jingjing, et al.
Published: (2024)
SFFNet: Synergistic Feature Fusion Network With Dual-Domain Edge Enhancement for UAV Image Object Detection
by: Zhang, Wenfeng, et al.
Published: (2026)
by: Zhang, Wenfeng, et al.
Published: (2026)
ACT360: An Efficient 360-Degree Action Detection and Summarization Framework for Mission-Critical Training and Debriefing
by: Tiwari, Aditi, et al.
Published: (2025)
by: Tiwari, Aditi, et al.
Published: (2025)
Spatial-Temporal Human-Object Interaction Detection
by: Sun, Xu, et al.
Published: (2025)
by: Sun, Xu, et al.
Published: (2025)
DPDETR: Decoupled Position Detection Transformer for Infrared-Visible Object Detection
by: Guo, Junjie, et al.
Published: (2024)
by: Guo, Junjie, et al.
Published: (2024)
Scaling Audio-Visual Quality Assessment Dataset via Crowdsourcing
by: Yang, Renyu, et al.
Published: (2026)
by: Yang, Renyu, et al.
Published: (2026)
Patch-level Sounding Object Tracking for Audio-Visual Question Answering
by: Li, Zhangbin, et al.
Published: (2024)
by: Li, Zhangbin, et al.
Published: (2024)
JPEG AI Image Compression Visual Artifacts: Detection Methods and Dataset
by: Tsereh, Daria, et al.
Published: (2024)
by: Tsereh, Daria, et al.
Published: (2024)
Joint Modeling of Big Five and HEXACO for Multimodal Apparent Personality-trait Recognition
by: Masumura, Ryo, et al.
Published: (2025)
by: Masumura, Ryo, et al.
Published: (2025)
Automatic Prompt Generation and Grounding Object Detection for Zero-Shot Image Anomaly Detection
by: Cheung, Tsun-Hin, et al.
Published: (2024)
by: Cheung, Tsun-Hin, et al.
Published: (2024)
3D2M Dataset: A 3-Dimension diverse Mesh Dataset
by: Dasgupta, Sankarshan
Published: (2024)
by: Dasgupta, Sankarshan
Published: (2024)
On the Robustness of Human-Object Interaction Detection against Distribution Shift
by: Xie, Chi, et al.
Published: (2025)
by: Xie, Chi, et al.
Published: (2025)
Efficient Object-centric Representation Learning with Pre-trained Geometric Prior
by: Khac, Phúc H. Le, et al.
Published: (2024)
by: Khac, Phúc H. Le, et al.
Published: (2024)
Probabilistic Temporal Masked Attention for Cross-view Online Action Detection
by: Xie, Liping, et al.
Published: (2025)
by: Xie, Liping, et al.
Published: (2025)
UVG-VPC: Voxelized Point Cloud Dataset for Visual Volumetric Video-based Coding
by: Gautier, Guillaume, et al.
Published: (2025)
by: Gautier, Guillaume, et al.
Published: (2025)
ManzaiSet: A Multimodal Dataset of Viewer Responses to Japanese Manzai Comedy
by: Kawamura, Kazuki, et al.
Published: (2025)
by: Kawamura, Kazuki, et al.
Published: (2025)
A Human-Annotated Video Dataset for Training and Evaluation of 360-Degree Video Summarization Methods
by: Kontostathis, Ioannis, et al.
Published: (2024)
by: Kontostathis, Ioannis, et al.
Published: (2024)
Multiverse Through Deepfakes: The MultiFakeVerse Dataset of Person-Centric Visual and Conceptual Manipulations
by: Gupta, Parul, et al.
Published: (2025)
by: Gupta, Parul, et al.
Published: (2025)
Towards Signboard-Oriented Visual Question Answering: ViSignVQA Dataset, Method and Benchmark
by: Nguyen, Hieu Minh, et al.
Published: (2025)
by: Nguyen, Hieu Minh, et al.
Published: (2025)
Divide-and-Conquer: Confluent Triple-Flow Network for RGB-T Salient Object Detection
by: Tang, Hao, et al.
Published: (2024)
by: Tang, Hao, et al.
Published: (2024)
Multi-Scale and Detail-Enhanced Segment Anything Model for Salient Object Detection
by: Gao, Shixuan, et al.
Published: (2024)
by: Gao, Shixuan, et al.
Published: (2024)
WaveMamba: Wavelet-Driven Mamba Fusion for RGB-Infrared Object Detection
by: Zhu, Haodong, et al.
Published: (2025)
by: Zhu, Haodong, et al.
Published: (2025)
Schrodinger Audio-Visual Editor: Object-Level Audiovisual Removal
by: Xu, Weihan, et al.
Published: (2025)
by: Xu, Weihan, et al.
Published: (2025)
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
by: Zhang, Zhenxing, et al.
Published: (2024)
by: Zhang, Zhenxing, et al.
Published: (2024)
FMNV: A Dataset of Media-Published News Videos for Fake News Detection
by: Wang, Yihao, et al.
Published: (2025)
by: Wang, Yihao, et al.
Published: (2025)
$\mathbf{C}^2$Former: Calibrated and Complementary Transformer for RGB-Infrared Object Detection
by: Yuan, Maoxun, et al.
Published: (2023)
by: Yuan, Maoxun, et al.
Published: (2023)
CUS3D :CLIP-based Unsupervised 3D Segmentation via Object-level Denoise
by: Yu, Fuyang, et al.
Published: (2024)
by: Yu, Fuyang, et al.
Published: (2024)
Learning Gaussian Data Augmentation in Feature Space for One-shot Object Detection in Manga
by: Taniguchi, Takara, et al.
Published: (2024)
by: Taniguchi, Takara, et al.
Published: (2024)
VP3D: Unleashing 2D Visual Prompt for Text-to-3D Generation
by: Chen, Yang, et al.
Published: (2024)
by: Chen, Yang, et al.
Published: (2024)
StableDub: Taming Diffusion Prior for Generalized and Efficient Visual Dubbing
by: Chen, Liyang, et al.
Published: (2025)
by: Chen, Liyang, et al.
Published: (2025)
Decoupled Audio-Visual Dataset Distillation
by: Li, Wenyuan, et al.
Published: (2025)
by: Li, Wenyuan, et al.
Published: (2025)
"Humor, Art, or Misinformation?": A Multimodal Dataset for Intent-Aware Synthetic Image Detection
by: Skoularikis, Anastasios, et al.
Published: (2025)
by: Skoularikis, Anastasios, et al.
Published: (2025)
SM3Det: A Unified Model for Multi-Modal Remote Sensing Object Detection
by: Li, Yuxuan, et al.
Published: (2024)
by: Li, Yuxuan, et al.
Published: (2024)
Segmentation-Based Attention Entropy: Detecting and Mitigating Object Hallucinations in Large Vision-Language Models
by: Song, Jiale, et al.
Published: (2026)
by: Song, Jiale, et al.
Published: (2026)
Ego3DT: Tracking Every 3D Object in Ego-centric Videos
by: Hao, Shengyu, et al.
Published: (2024)
by: Hao, Shengyu, et al.
Published: (2024)
KAN-SAM: Kolmogorov-Arnold Network Guided Segment Anything Model for RGB-T Salient Object Detection
by: Li, Xingyuan, et al.
Published: (2025)
by: Li, Xingyuan, et al.
Published: (2025)
Detection and Recovery of Adversarial Slow-Pose Drift in Offloaded Visual-Inertial Odometry
by: Saha, Soruya, et al.
Published: (2025)
by: Saha, Soruya, et al.
Published: (2025)
Multi-source Multimodal Progressive Domain Adaption for Audio-Visual Deception Detection
by: Lin, Ronghao, et al.
Published: (2025)
by: Lin, Ronghao, et al.
Published: (2025)
Similar Items
-
Dual Mutual Learning Network with Global-local Awareness for RGB-D Salient Object Detection
by: Yi, Kang, et al.
Published: (2025) -
HMPE:HeatMap Embedding for Efficient Transformer-Based Small Object Detection
by: Zeng, YangChen
Published: (2025) -
TextRefiner: Internal Visual Feature as Efficient Refiner for Vision-Language Models Prompt Tuning
by: Xie, Jingjing, et al.
Published: (2024) -
SFFNet: Synergistic Feature Fusion Network With Dual-Domain Edge Enhancement for UAV Image Object Detection
by: Zhang, Wenfeng, et al.
Published: (2026) -
ACT360: An Efficient 360-Degree Action Detection and Summarization Framework for Mission-Critical Training and Debriefing
by: Tiwari, Aditi, et al.
Published: (2025)