MV-DETR: Multi-modality indoor object detection by Multi-View DEtecton TRansformers
Fuente:
arXiv
Salvato in:
| Autori principali: | Dong, Zichao, Zhang, Yilin, Huang, Xufeng, Ji, Hang, Shi, Zhan, Zhan, Xin, Chen, Junbo |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
LVIC: Multi-modality segmentation by Lifting Visual Info as Cue
di: Dong, Zichao, et al.
Pubblicazione: (2024)
di: Dong, Zichao, et al.
Pubblicazione: (2024)
PeP: a Point enhanced Painting method for unified point cloud tasks
di: Dong, Zichao, et al.
Pubblicazione: (2023)
di: Dong, Zichao, et al.
Pubblicazione: (2023)
RoPETR: Improving Temporal Camera-Only 3D Detection by Integrating Enhanced Rotary Position Embedding
di: Ji, Hang, et al.
Pubblicazione: (2025)
di: Ji, Hang, et al.
Pubblicazione: (2025)
EW-DETR: Evolving World Object Detection via Incremental Low-Rank DEtection TRansformer
di: Monga, Munish, et al.
Pubblicazione: (2026)
di: Monga, Munish, et al.
Pubblicazione: (2026)
A Coarse-to-Fine Approach to Multi-Modality 3D Occupancy Grounding
di: Shi, Zhan, et al.
Pubblicazione: (2025)
di: Shi, Zhan, et al.
Pubblicazione: (2025)
LD-DETR: Loop Decoder DEtection TRansformer for Video Moment Retrieval and Highlight Detection
di: Zhao, Pengcheng, et al.
Pubblicazione: (2025)
di: Zhao, Pengcheng, et al.
Pubblicazione: (2025)
MGFs: Masked Gaussian Fields for Meshing Building based on Multi-View Images
di: Wang, Tengfei, et al.
Pubblicazione: (2024)
di: Wang, Tengfei, et al.
Pubblicazione: (2024)
MAG-VLAQ: Multi-modal Aerial-Ground Query Aggregation for Cross-View Place Recognition
di: Xu, Zhengyi, et al.
Pubblicazione: (2026)
di: Xu, Zhengyi, et al.
Pubblicazione: (2026)
MV-VTON: Multi-View Virtual Try-On with Diffusion Models
di: Wang, Haoyu, et al.
Pubblicazione: (2024)
di: Wang, Haoyu, et al.
Pubblicazione: (2024)
MV-TAP: Tracking Any Point in Multi-View Videos
di: Koo, Jahyeok, et al.
Pubblicazione: (2025)
di: Koo, Jahyeok, et al.
Pubblicazione: (2025)
MV2MAE: Multi-View Video Masked Autoencoders
di: Shah, Ketul, et al.
Pubblicazione: (2024)
di: Shah, Ketul, et al.
Pubblicazione: (2024)
RapidMV: Leveraging Spatio-Angular Representations for Efficient and Consistent Text-to-Multi-View Synthesis
di: Kim, Seungwook, et al.
Pubblicazione: (2025)
di: Kim, Seungwook, et al.
Pubblicazione: (2025)
MV-GMN: State Space Model for Multi-View Action Recognition
di: Lin, Yuhui, et al.
Pubblicazione: (2025)
di: Lin, Yuhui, et al.
Pubblicazione: (2025)
MV-SSM: Multi-View State Space Modeling for 3D Human Pose Estimation
di: Chharia, Aviral, et al.
Pubblicazione: (2025)
di: Chharia, Aviral, et al.
Pubblicazione: (2025)
DQ-DETR: DETR with Dynamic Query for Tiny Object Detection
di: Huang, Yi-Xin, et al.
Pubblicazione: (2024)
di: Huang, Yi-Xin, et al.
Pubblicazione: (2024)
Enhancing Incomplete Multi-modal Brain Tumor Segmentation with Intra-modal Asymmetry and Inter-modal Dependency
di: Liu, Weide, et al.
Pubblicazione: (2024)
di: Liu, Weide, et al.
Pubblicazione: (2024)
VideoMV: Consistent Multi-View Generation Based on Large Video Generative Model
di: Zuo, Qi, et al.
Pubblicazione: (2024)
di: Zuo, Qi, et al.
Pubblicazione: (2024)
COTR: Compact Occupancy TRansformer for Vision-based 3D Occupancy Prediction
di: Ma, Qihang, et al.
Pubblicazione: (2023)
di: Ma, Qihang, et al.
Pubblicazione: (2023)
MV-SAM3D: Adaptive Multi-View Fusion for Layout-Aware 3D Generation
di: Li, Baicheng, et al.
Pubblicazione: (2026)
di: Li, Baicheng, et al.
Pubblicazione: (2026)
MV-RoMa: From Pairwise Matching into Multi-View Track Reconstruction
di: Lee, Jongmin, et al.
Pubblicazione: (2026)
di: Lee, Jongmin, et al.
Pubblicazione: (2026)
DeCo-DETR: Decoupled Cognition DETR for efficient Open-Vocabulary Object Detection
di: Wang, Siheng, et al.
Pubblicazione: (2026)
di: Wang, Siheng, et al.
Pubblicazione: (2026)
FMC-DETR: Frequency-Decoupled Multi-Domain Coordination for Aerial-View Object Detection
di: Liang, Ben, et al.
Pubblicazione: (2025)
di: Liang, Ben, et al.
Pubblicazione: (2025)
WD-DETR: Wavelet Denoising-Enhanced Real-Time Object Detection Transformer for Robot Perception with Event Cameras
di: Cui, Yangjie, et al.
Pubblicazione: (2025)
di: Cui, Yangjie, et al.
Pubblicazione: (2025)
Multi-modal Generative AI: Multi-modal LLMs, Diffusions, and the Unification
di: Wang, Xin, et al.
Pubblicazione: (2024)
di: Wang, Xin, et al.
Pubblicazione: (2024)
YingVideo-MV: Music-Driven Multi-Stage Video Generation
di: Chen, Jiahui, et al.
Pubblicazione: (2025)
di: Chen, Jiahui, et al.
Pubblicazione: (2025)
MV-Match: Multi-View Matching for Domain-Adaptive Identification of Plant Nutrient Deficiencies
di: Yi, Jinhui, et al.
Pubblicazione: (2024)
di: Yi, Jinhui, et al.
Pubblicazione: (2024)
MV-CLIP: Multi-View CLIP for Zero-shot 3D Shape Recognition
di: Song, Dan, et al.
Pubblicazione: (2023)
di: Song, Dan, et al.
Pubblicazione: (2023)
AnomalyXFusion: Multi-modal Anomaly Synthesis with Diffusion
di: Hu, Jie, et al.
Pubblicazione: (2024)
di: Hu, Jie, et al.
Pubblicazione: (2024)
MV-MOS: Multi-View Feature Fusion for 3D Moving Object Segmentation
di: Cheng, Jintao, et al.
Pubblicazione: (2024)
di: Cheng, Jintao, et al.
Pubblicazione: (2024)
Siamese-DETR for Generic Multi-Object Tracking
di: Liu, Qiankun, et al.
Pubblicazione: (2023)
di: Liu, Qiankun, et al.
Pubblicazione: (2023)
MV-Fashion: Towards Enabling Virtual Try-On and Size Estimation with Multi-View Paired Data
di: Laczkó, Hunor, et al.
Pubblicazione: (2026)
di: Laczkó, Hunor, et al.
Pubblicazione: (2026)
MV2Cyl: Reconstructing 3D Extrusion Cylinders from Multi-View Images
di: Hong, Eunji, et al.
Pubblicazione: (2024)
di: Hong, Eunji, et al.
Pubblicazione: (2024)
Oracle Bone Inscriptions Multi-modal Dataset
di: Li, Bang, et al.
Pubblicazione: (2024)
di: Li, Bang, et al.
Pubblicazione: (2024)
PlanTRansformer: Unified Prediction and Planning with Goal-conditioned Transformer
di: Selzer, Constantin, et al.
Pubblicazione: (2026)
di: Selzer, Constantin, et al.
Pubblicazione: (2026)
Sim-DETR: Unlock DETR for Temporal Sentence Grounding
di: Tang, Jiajin, et al.
Pubblicazione: (2025)
di: Tang, Jiajin, et al.
Pubblicazione: (2025)
MV-S2V: Multi-View Subject-Consistent Video Generation
di: Song, Ziyang, et al.
Pubblicazione: (2026)
di: Song, Ziyang, et al.
Pubblicazione: (2026)
Advancing 3D Scene Understanding with MV-ScanQA Multi-View Reasoning Evaluation and TripAlign Pre-training Dataset
di: Mo, Wentao, et al.
Pubblicazione: (2025)
di: Mo, Wentao, et al.
Pubblicazione: (2025)
MV-Adapter: Multi-view Consistent Image Generation Made Easy
di: Huang, Zehuan, et al.
Pubblicazione: (2024)
di: Huang, Zehuan, et al.
Pubblicazione: (2024)
LUCES-MV: A Multi-View Dataset for Near-Field Point Light Source Photometric Stereo
di: Logothetis, Fotios, et al.
Pubblicazione: (2024)
di: Logothetis, Fotios, et al.
Pubblicazione: (2024)
SpecDETR: A transformer-based hyperspectral point object detection network
di: Li, Zhaoxu, et al.
Pubblicazione: (2024)
di: Li, Zhaoxu, et al.
Pubblicazione: (2024)
Documenti analoghi
-
LVIC: Multi-modality segmentation by Lifting Visual Info as Cue
di: Dong, Zichao, et al.
Pubblicazione: (2024) -
PeP: a Point enhanced Painting method for unified point cloud tasks
di: Dong, Zichao, et al.
Pubblicazione: (2023) -
RoPETR: Improving Temporal Camera-Only 3D Detection by Integrating Enhanced Rotary Position Embedding
di: Ji, Hang, et al.
Pubblicazione: (2025) -
EW-DETR: Evolving World Object Detection via Incremental Low-Rank DEtection TRansformer
di: Monga, Munish, et al.
Pubblicazione: (2026) -
A Coarse-to-Fine Approach to Multi-Modality 3D Occupancy Grounding
di: Shi, Zhan, et al.
Pubblicazione: (2025)