Saved in:
| Main Authors: | Dong, Zichao, Zhang, Yilin, Huang, Xufeng, Ji, Hang, Shi, Zhan, Zhan, Xin, Chen, Junbo |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2408.06604 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LVIC: Multi-modality segmentation by Lifting Visual Info as Cue
by: Dong, Zichao, et al.
Published: (2024)
by: Dong, Zichao, et al.
Published: (2024)
PeP: a Point enhanced Painting method for unified point cloud tasks
by: Dong, Zichao, et al.
Published: (2023)
by: Dong, Zichao, et al.
Published: (2023)
RoPETR: Improving Temporal Camera-Only 3D Detection by Integrating Enhanced Rotary Position Embedding
by: Ji, Hang, et al.
Published: (2025)
by: Ji, Hang, et al.
Published: (2025)
EW-DETR: Evolving World Object Detection via Incremental Low-Rank DEtection TRansformer
by: Monga, Munish, et al.
Published: (2026)
by: Monga, Munish, et al.
Published: (2026)
A Coarse-to-Fine Approach to Multi-Modality 3D Occupancy Grounding
by: Shi, Zhan, et al.
Published: (2025)
by: Shi, Zhan, et al.
Published: (2025)
LD-DETR: Loop Decoder DEtection TRansformer for Video Moment Retrieval and Highlight Detection
by: Zhao, Pengcheng, et al.
Published: (2025)
by: Zhao, Pengcheng, et al.
Published: (2025)
MGFs: Masked Gaussian Fields for Meshing Building based on Multi-View Images
by: Wang, Tengfei, et al.
Published: (2024)
by: Wang, Tengfei, et al.
Published: (2024)
MAG-VLAQ: Multi-modal Aerial-Ground Query Aggregation for Cross-View Place Recognition
by: Xu, Zhengyi, et al.
Published: (2026)
by: Xu, Zhengyi, et al.
Published: (2026)
MV-VTON: Multi-View Virtual Try-On with Diffusion Models
by: Wang, Haoyu, et al.
Published: (2024)
by: Wang, Haoyu, et al.
Published: (2024)
MV-TAP: Tracking Any Point in Multi-View Videos
by: Koo, Jahyeok, et al.
Published: (2025)
by: Koo, Jahyeok, et al.
Published: (2025)
MV2MAE: Multi-View Video Masked Autoencoders
by: Shah, Ketul, et al.
Published: (2024)
by: Shah, Ketul, et al.
Published: (2024)
RapidMV: Leveraging Spatio-Angular Representations for Efficient and Consistent Text-to-Multi-View Synthesis
by: Kim, Seungwook, et al.
Published: (2025)
by: Kim, Seungwook, et al.
Published: (2025)
DQ-DETR: DETR with Dynamic Query for Tiny Object Detection
by: Huang, Yi-Xin, et al.
Published: (2024)
by: Huang, Yi-Xin, et al.
Published: (2024)
Enhancing Incomplete Multi-modal Brain Tumor Segmentation with Intra-modal Asymmetry and Inter-modal Dependency
by: Liu, Weide, et al.
Published: (2024)
by: Liu, Weide, et al.
Published: (2024)
WD-DETR: Wavelet Denoising-Enhanced Real-Time Object Detection Transformer for Robot Perception with Event Cameras
by: Cui, Yangjie, et al.
Published: (2025)
by: Cui, Yangjie, et al.
Published: (2025)
VideoMV: Consistent Multi-View Generation Based on Large Video Generative Model
by: Zuo, Qi, et al.
Published: (2024)
by: Zuo, Qi, et al.
Published: (2024)
MV-SSM: Multi-View State Space Modeling for 3D Human Pose Estimation
by: Chharia, Aviral, et al.
Published: (2025)
by: Chharia, Aviral, et al.
Published: (2025)
DeCo-DETR: Decoupled Cognition DETR for efficient Open-Vocabulary Object Detection
by: Wang, Siheng, et al.
Published: (2026)
by: Wang, Siheng, et al.
Published: (2026)
COTR: Compact Occupancy TRansformer for Vision-based 3D Occupancy Prediction
by: Ma, Qihang, et al.
Published: (2023)
by: Ma, Qihang, et al.
Published: (2023)
MV-GMN: State Space Model for Multi-View Action Recognition
by: Lin, Yuhui, et al.
Published: (2025)
by: Lin, Yuhui, et al.
Published: (2025)
FMC-DETR: Frequency-Decoupled Multi-Domain Coordination for Aerial-View Object Detection
by: Liang, Ben, et al.
Published: (2025)
by: Liang, Ben, et al.
Published: (2025)
Multi-modal Generative AI: Multi-modal LLMs, Diffusions, and the Unification
by: Wang, Xin, et al.
Published: (2024)
by: Wang, Xin, et al.
Published: (2024)
MV-SAM3D: Adaptive Multi-View Fusion for Layout-Aware 3D Generation
by: Li, Baicheng, et al.
Published: (2026)
by: Li, Baicheng, et al.
Published: (2026)
MV-RoMa: From Pairwise Matching into Multi-View Track Reconstruction
by: Lee, Jongmin, et al.
Published: (2026)
by: Lee, Jongmin, et al.
Published: (2026)
AnomalyXFusion: Multi-modal Anomaly Synthesis with Diffusion
by: Hu, Jie, et al.
Published: (2024)
by: Hu, Jie, et al.
Published: (2024)
YingVideo-MV: Music-Driven Multi-Stage Video Generation
by: Chen, Jiahui, et al.
Published: (2025)
by: Chen, Jiahui, et al.
Published: (2025)
MV-MOS: Multi-View Feature Fusion for 3D Moving Object Segmentation
by: Cheng, Jintao, et al.
Published: (2024)
by: Cheng, Jintao, et al.
Published: (2024)
Siamese-DETR for Generic Multi-Object Tracking
by: Liu, Qiankun, et al.
Published: (2023)
by: Liu, Qiankun, et al.
Published: (2023)
Oracle Bone Inscriptions Multi-modal Dataset
by: Li, Bang, et al.
Published: (2024)
by: Li, Bang, et al.
Published: (2024)
MV-Match: Multi-View Matching for Domain-Adaptive Identification of Plant Nutrient Deficiencies
by: Yi, Jinhui, et al.
Published: (2024)
by: Yi, Jinhui, et al.
Published: (2024)
MV-CLIP: Multi-View CLIP for Zero-shot 3D Shape Recognition
by: Song, Dan, et al.
Published: (2023)
by: Song, Dan, et al.
Published: (2023)
PlanTRansformer: Unified Prediction and Planning with Goal-conditioned Transformer
by: Selzer, Constantin, et al.
Published: (2026)
by: Selzer, Constantin, et al.
Published: (2026)
Sim-DETR: Unlock DETR for Temporal Sentence Grounding
by: Tang, Jiajin, et al.
Published: (2025)
by: Tang, Jiajin, et al.
Published: (2025)
MV-S2V: Multi-View Subject-Consistent Video Generation
by: Song, Ziyang, et al.
Published: (2026)
by: Song, Ziyang, et al.
Published: (2026)
A comparison of extended object tracking with multi-modal sensors in indoor environment
by: Shuai, Jiangtao, et al.
Published: (2024)
by: Shuai, Jiangtao, et al.
Published: (2024)
MV-Fashion: Towards Enabling Virtual Try-On and Size Estimation with Multi-View Paired Data
by: Laczkó, Hunor, et al.
Published: (2026)
by: Laczkó, Hunor, et al.
Published: (2026)
MV2Cyl: Reconstructing 3D Extrusion Cylinders from Multi-View Images
by: Hong, Eunji, et al.
Published: (2024)
by: Hong, Eunji, et al.
Published: (2024)
Inst3D-LMM: Instance-Aware 3D Scene Understanding with Multi-modal Instruction Tuning
by: Yu, Hanxun, et al.
Published: (2025)
by: Yu, Hanxun, et al.
Published: (2025)
Efficient Multi-modal Large Language Models via Visual Token Grouping
by: Huang, Minbin, et al.
Published: (2024)
by: Huang, Minbin, et al.
Published: (2024)
GeoMM: On Geodesic Perspective for Multi-modal Learning
by: Mei, Shibin, et al.
Published: (2025)
by: Mei, Shibin, et al.
Published: (2025)
Similar Items
-
LVIC: Multi-modality segmentation by Lifting Visual Info as Cue
by: Dong, Zichao, et al.
Published: (2024) -
PeP: a Point enhanced Painting method for unified point cloud tasks
by: Dong, Zichao, et al.
Published: (2023) -
RoPETR: Improving Temporal Camera-Only 3D Detection by Integrating Enhanced Rotary Position Embedding
by: Ji, Hang, et al.
Published: (2025) -
EW-DETR: Evolving World Object Detection via Incremental Low-Rank DEtection TRansformer
by: Monga, Munish, et al.
Published: (2026) -
A Coarse-to-Fine Approach to Multi-Modality 3D Occupancy Grounding
by: Shi, Zhan, et al.
Published: (2025)