PD-APE: A Parallel Decoding Framework with Adaptive Position Encoding for 3D Visual Grounding
Fuente:
arXiv
Saved in:
| Main Authors: | Hou, Chenshu, Peng, Liang, Wu, Xiaopei, He, Xiaofei, Wang, Wenxiao |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Semi-supervised 3D Object Detection with PatchTeacher and PillarMix
by: Wu, Xiaopei, et al.
Published: (2024)
by: Wu, Xiaopei, et al.
Published: (2024)
An Efficient and Effective Transformer Decoder-Based Framework for Multi-Task Visual Grounding
by: Chen, Wei, et al.
Published: (2024)
by: Chen, Wei, et al.
Published: (2024)
APE: Agentic Prompt Enhancer for Image Generation and Editing
by: Huang, Zijian, et al.
Published: (2026)
by: Huang, Zijian, et al.
Published: (2026)
SelFLoc: Selective Feature Fusion for Large-scale Point Cloud-based Place Recognition
by: Qiu, Qibo, et al.
Published: (2023)
by: Qiu, Qibo, et al.
Published: (2023)
RAVE: Rate-Adaptive Visual Encoding for 3D Gaussian Splatting
by: Tran, Hoang-Nhat, et al.
Published: (2025)
by: Tran, Hoang-Nhat, et al.
Published: (2025)
SGCap: Decoding Semantic Group for Zero-shot Video Captioning
by: Pan, Zeyu, et al.
Published: (2025)
by: Pan, Zeyu, et al.
Published: (2025)
VL-UniTrack: A Unified Framework with Visual-Language Prompts for UAV-Ground Visual Tracking
by: Xu, Boyue, et al.
Published: (2026)
by: Xu, Boyue, et al.
Published: (2026)
DEGround: An Effective Baseline for Ego-centric 3D Visual Grounding with a Homogeneous Framework
by: Zhang, Yani, et al.
Published: (2025)
by: Zhang, Yani, et al.
Published: (2025)
LSVG: Language-Guided Scene Graphs with 2D-Assisted Multi-Modal Encoding for 3D Visual Grounding
by: Xiao, Feng, et al.
Published: (2025)
by: Xiao, Feng, et al.
Published: (2025)
PD-VLA: Accelerating Vision-Language-Action Model Integrated with Action Chunking via Parallel Decoding
by: Song, Wenxuan, et al.
Published: (2025)
by: Song, Wenxuan, et al.
Published: (2025)
Boosting Resolution Generalization of Diffusion Transformers with Randomized Positional Encodings
by: Hou, Liang, et al.
Published: (2025)
by: Hou, Liang, et al.
Published: (2025)
ParallelVLM: Lossless Video-LLM Acceleration with Visual Alignment Aware Parallel Speculative Decoding
by: Kong, Quan, et al.
Published: (2026)
by: Kong, Quan, et al.
Published: (2026)
ChangingGrounding: 3D Visual Grounding in Changing Scenes
by: Hu, Miao, et al.
Published: (2025)
by: Hu, Miao, et al.
Published: (2025)
NuGrounding: A Multi-View 3D Visual Grounding Framework in Autonomous Driving
by: Li, Fuhao, et al.
Published: (2025)
by: Li, Fuhao, et al.
Published: (2025)
Repurposing 2D Diffusion Models for 3D Shape Completion
by: He, Yao, et al.
Published: (2025)
by: He, Yao, et al.
Published: (2025)
Mono3DVG-EnSD: Enhanced Spatial-aware and Dimension-decoupled Text Encoding for Monocular 3D Visual Grounding
by: Li, Yuzhen, et al.
Published: (2025)
by: Li, Yuzhen, et al.
Published: (2025)
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding
by: Li, Rong, et al.
Published: (2024)
by: Li, Rong, et al.
Published: (2024)
Head-wise Adaptive Rotary Positional Encoding for Fine-Grained Image Generation
by: Li, Jiaye, et al.
Published: (2025)
by: Li, Jiaye, et al.
Published: (2025)
Joint Top-Down and Bottom-Up Frameworks for 3D Visual Grounding
by: Liu, Yang, et al.
Published: (2024)
by: Liu, Yang, et al.
Published: (2024)
OMEGA: Optimized Multimodal Position Encoding Index Derivation with Global Adaptive Scaling for Vision-Language Models
by: Huang, Ruoxiang, et al.
Published: (2025)
by: Huang, Ruoxiang, et al.
Published: (2025)
Accelerating Video Generation Inference with Sequential-Parallel 3D Positional Encoding Using a Global Time Index
by: Yuan, Chao, et al.
Published: (2026)
by: Yuan, Chao, et al.
Published: (2026)
Data-Efficient 3D Visual Grounding via Order-Aware Referring
by: Wu, Tung-Yu, et al.
Published: (2024)
by: Wu, Tung-Yu, et al.
Published: (2024)
MiKASA: Multi-Key-Anchor & Scene-Aware Transformer for 3D Visual Grounding
by: Chang, Chun-Peng, et al.
Published: (2024)
by: Chang, Chun-Peng, et al.
Published: (2024)
Neuro-3D: Towards 3D Visual Decoding from EEG Signals
by: Guo, Zhanqiang, et al.
Published: (2024)
by: Guo, Zhanqiang, et al.
Published: (2024)
Training-free Spatially Grounded Geometric Shape Encoding (Technical Report)
by: He, Yuhang
Published: (2026)
by: He, Yuhang
Published: (2026)
Revisiting Multimodal Positional Encoding in Vision-Language Models
by: Huang, Jie, et al.
Published: (2025)
by: Huang, Jie, et al.
Published: (2025)
Occlusion Handling in 3D Human Pose Estimation with Perturbed Positional Encoding
by: Azizi, Niloofar, et al.
Published: (2024)
by: Azizi, Niloofar, et al.
Published: (2024)
Lagrangian Motion Fields for Long-term Motion Generation
by: Yang, Yifei, et al.
Published: (2024)
by: Yang, Yifei, et al.
Published: (2024)
Sample-level Adaptive Knowledge Distillation for Action Recognition
by: Li, Ping, et al.
Published: (2025)
by: Li, Ping, et al.
Published: (2025)
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding
by: Zheng, Henry, et al.
Published: (2025)
by: Zheng, Henry, et al.
Published: (2025)
OBMO: One Bounding Box Multiple Objects for Monocular 3D Object Detection
by: Huang, Chenxi, et al.
Published: (2022)
by: Huang, Chenxi, et al.
Published: (2022)
Positional Encoding Field
by: Bai, Yunpeng, et al.
Published: (2025)
by: Bai, Yunpeng, et al.
Published: (2025)
Easy-Poly: An Easy Polyhedral Framework For 3D Multi-Object Tracking
by: Zhang, Peng, et al.
Published: (2025)
by: Zhang, Peng, et al.
Published: (2025)
Recursive Visual Imagination and Adaptive Linguistic Grounding for Vision Language Navigation
by: Chen, Bolei, et al.
Published: (2025)
by: Chen, Bolei, et al.
Published: (2025)
ProxyTransformation: Preshaping Point Cloud Manifold With Proxy Attention For 3D Visual Grounding
by: Peng, Qihang, et al.
Published: (2025)
by: Peng, Qihang, et al.
Published: (2025)
Multi-View Large Reconstruction Model via Geometry-Aware Positional Encoding and Attention
by: Li, Mengfei, et al.
Published: (2024)
by: Li, Mengfei, et al.
Published: (2024)
SUDO: Enhancing Text-to-Image Diffusion Models with Self-Supervised Direct Preference Optimization
by: Peng, Liang, et al.
Published: (2025)
by: Peng, Liang, et al.
Published: (2025)
Weierstrass Positional Encoding for Vision Transformers
by: Xin, Zhihang, et al.
Published: (2026)
by: Xin, Zhihang, et al.
Published: (2026)
LEDiT: Your Length-Extrapolatable Diffusion Transformer without Positional Encoding
by: Zhang, Shen, et al.
Published: (2025)
by: Zhang, Shen, et al.
Published: (2025)
Self-Prophetic Decoding to Unlock Visual Search in LVLMs
by: He, Zhendong, et al.
Published: (2026)
by: He, Zhendong, et al.
Published: (2026)
Similar Items
-
Semi-supervised 3D Object Detection with PatchTeacher and PillarMix
by: Wu, Xiaopei, et al.
Published: (2024) -
An Efficient and Effective Transformer Decoder-Based Framework for Multi-Task Visual Grounding
by: Chen, Wei, et al.
Published: (2024) -
APE: Agentic Prompt Enhancer for Image Generation and Editing
by: Huang, Zijian, et al.
Published: (2026) -
SelFLoc: Selective Feature Fusion for Large-scale Point Cloud-based Place Recognition
by: Qiu, Qibo, et al.
Published: (2023) -
RAVE: Rate-Adaptive Visual Encoding for 3D Gaussian Splatting
by: Tran, Hoang-Nhat, et al.
Published: (2025)