PD-APE: A Parallel Decoding Framework with Adaptive Position Encoding for 3D Visual Grounding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hou, Chenshu, Peng, Liang, Wu, Xiaopei, He, Xiaofei, Wang, Wenxiao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Semi-supervised 3D Object Detection with PatchTeacher and PillarMix
von: Wu, Xiaopei, et al.
Veröffentlicht: (2024)
von: Wu, Xiaopei, et al.
Veröffentlicht: (2024)
An Efficient and Effective Transformer Decoder-Based Framework for Multi-Task Visual Grounding
von: Chen, Wei, et al.
Veröffentlicht: (2024)
von: Chen, Wei, et al.
Veröffentlicht: (2024)
APE: Agentic Prompt Enhancer for Image Generation and Editing
von: Huang, Zijian, et al.
Veröffentlicht: (2026)
von: Huang, Zijian, et al.
Veröffentlicht: (2026)
SelFLoc: Selective Feature Fusion for Large-scale Point Cloud-based Place Recognition
von: Qiu, Qibo, et al.
Veröffentlicht: (2023)
von: Qiu, Qibo, et al.
Veröffentlicht: (2023)
RAVE: Rate-Adaptive Visual Encoding for 3D Gaussian Splatting
von: Tran, Hoang-Nhat, et al.
Veröffentlicht: (2025)
von: Tran, Hoang-Nhat, et al.
Veröffentlicht: (2025)
SGCap: Decoding Semantic Group for Zero-shot Video Captioning
von: Pan, Zeyu, et al.
Veröffentlicht: (2025)
von: Pan, Zeyu, et al.
Veröffentlicht: (2025)
VL-UniTrack: A Unified Framework with Visual-Language Prompts for UAV-Ground Visual Tracking
von: Xu, Boyue, et al.
Veröffentlicht: (2026)
von: Xu, Boyue, et al.
Veröffentlicht: (2026)
DEGround: An Effective Baseline for Ego-centric 3D Visual Grounding with a Homogeneous Framework
von: Zhang, Yani, et al.
Veröffentlicht: (2025)
von: Zhang, Yani, et al.
Veröffentlicht: (2025)
LSVG: Language-Guided Scene Graphs with 2D-Assisted Multi-Modal Encoding for 3D Visual Grounding
von: Xiao, Feng, et al.
Veröffentlicht: (2025)
von: Xiao, Feng, et al.
Veröffentlicht: (2025)
PD-VLA: Accelerating Vision-Language-Action Model Integrated with Action Chunking via Parallel Decoding
von: Song, Wenxuan, et al.
Veröffentlicht: (2025)
von: Song, Wenxuan, et al.
Veröffentlicht: (2025)
Boosting Resolution Generalization of Diffusion Transformers with Randomized Positional Encodings
von: Hou, Liang, et al.
Veröffentlicht: (2025)
von: Hou, Liang, et al.
Veröffentlicht: (2025)
ParallelVLM: Lossless Video-LLM Acceleration with Visual Alignment Aware Parallel Speculative Decoding
von: Kong, Quan, et al.
Veröffentlicht: (2026)
von: Kong, Quan, et al.
Veröffentlicht: (2026)
ChangingGrounding: 3D Visual Grounding in Changing Scenes
von: Hu, Miao, et al.
Veröffentlicht: (2025)
von: Hu, Miao, et al.
Veröffentlicht: (2025)
NuGrounding: A Multi-View 3D Visual Grounding Framework in Autonomous Driving
von: Li, Fuhao, et al.
Veröffentlicht: (2025)
von: Li, Fuhao, et al.
Veröffentlicht: (2025)
Repurposing 2D Diffusion Models for 3D Shape Completion
von: He, Yao, et al.
Veröffentlicht: (2025)
von: He, Yao, et al.
Veröffentlicht: (2025)
Mono3DVG-EnSD: Enhanced Spatial-aware and Dimension-decoupled Text Encoding for Monocular 3D Visual Grounding
von: Li, Yuzhen, et al.
Veröffentlicht: (2025)
von: Li, Yuzhen, et al.
Veröffentlicht: (2025)
SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding
von: Li, Rong, et al.
Veröffentlicht: (2024)
von: Li, Rong, et al.
Veröffentlicht: (2024)
Head-wise Adaptive Rotary Positional Encoding for Fine-Grained Image Generation
von: Li, Jiaye, et al.
Veröffentlicht: (2025)
von: Li, Jiaye, et al.
Veröffentlicht: (2025)
Joint Top-Down and Bottom-Up Frameworks for 3D Visual Grounding
von: Liu, Yang, et al.
Veröffentlicht: (2024)
von: Liu, Yang, et al.
Veröffentlicht: (2024)
OMEGA: Optimized Multimodal Position Encoding Index Derivation with Global Adaptive Scaling for Vision-Language Models
von: Huang, Ruoxiang, et al.
Veröffentlicht: (2025)
von: Huang, Ruoxiang, et al.
Veröffentlicht: (2025)
Accelerating Video Generation Inference with Sequential-Parallel 3D Positional Encoding Using a Global Time Index
von: Yuan, Chao, et al.
Veröffentlicht: (2026)
von: Yuan, Chao, et al.
Veröffentlicht: (2026)
Data-Efficient 3D Visual Grounding via Order-Aware Referring
von: Wu, Tung-Yu, et al.
Veröffentlicht: (2024)
von: Wu, Tung-Yu, et al.
Veröffentlicht: (2024)
MiKASA: Multi-Key-Anchor & Scene-Aware Transformer for 3D Visual Grounding
von: Chang, Chun-Peng, et al.
Veröffentlicht: (2024)
von: Chang, Chun-Peng, et al.
Veröffentlicht: (2024)
Neuro-3D: Towards 3D Visual Decoding from EEG Signals
von: Guo, Zhanqiang, et al.
Veröffentlicht: (2024)
von: Guo, Zhanqiang, et al.
Veröffentlicht: (2024)
Training-free Spatially Grounded Geometric Shape Encoding (Technical Report)
von: He, Yuhang
Veröffentlicht: (2026)
von: He, Yuhang
Veröffentlicht: (2026)
Revisiting Multimodal Positional Encoding in Vision-Language Models
von: Huang, Jie, et al.
Veröffentlicht: (2025)
von: Huang, Jie, et al.
Veröffentlicht: (2025)
Occlusion Handling in 3D Human Pose Estimation with Perturbed Positional Encoding
von: Azizi, Niloofar, et al.
Veröffentlicht: (2024)
von: Azizi, Niloofar, et al.
Veröffentlicht: (2024)
Lagrangian Motion Fields for Long-term Motion Generation
von: Yang, Yifei, et al.
Veröffentlicht: (2024)
von: Yang, Yifei, et al.
Veröffentlicht: (2024)
Sample-level Adaptive Knowledge Distillation for Action Recognition
von: Li, Ping, et al.
Veröffentlicht: (2025)
von: Li, Ping, et al.
Veröffentlicht: (2025)
DenseGrounding: Improving Dense Language-Vision Semantics for Ego-Centric 3D Visual Grounding
von: Zheng, Henry, et al.
Veröffentlicht: (2025)
von: Zheng, Henry, et al.
Veröffentlicht: (2025)
OBMO: One Bounding Box Multiple Objects for Monocular 3D Object Detection
von: Huang, Chenxi, et al.
Veröffentlicht: (2022)
von: Huang, Chenxi, et al.
Veröffentlicht: (2022)
Positional Encoding Field
von: Bai, Yunpeng, et al.
Veröffentlicht: (2025)
von: Bai, Yunpeng, et al.
Veröffentlicht: (2025)
Easy-Poly: An Easy Polyhedral Framework For 3D Multi-Object Tracking
von: Zhang, Peng, et al.
Veröffentlicht: (2025)
von: Zhang, Peng, et al.
Veröffentlicht: (2025)
Recursive Visual Imagination and Adaptive Linguistic Grounding for Vision Language Navigation
von: Chen, Bolei, et al.
Veröffentlicht: (2025)
von: Chen, Bolei, et al.
Veröffentlicht: (2025)
ProxyTransformation: Preshaping Point Cloud Manifold With Proxy Attention For 3D Visual Grounding
von: Peng, Qihang, et al.
Veröffentlicht: (2025)
von: Peng, Qihang, et al.
Veröffentlicht: (2025)
Multi-View Large Reconstruction Model via Geometry-Aware Positional Encoding and Attention
von: Li, Mengfei, et al.
Veröffentlicht: (2024)
von: Li, Mengfei, et al.
Veröffentlicht: (2024)
SUDO: Enhancing Text-to-Image Diffusion Models with Self-Supervised Direct Preference Optimization
von: Peng, Liang, et al.
Veröffentlicht: (2025)
von: Peng, Liang, et al.
Veröffentlicht: (2025)
Weierstrass Positional Encoding for Vision Transformers
von: Xin, Zhihang, et al.
Veröffentlicht: (2026)
von: Xin, Zhihang, et al.
Veröffentlicht: (2026)
LEDiT: Your Length-Extrapolatable Diffusion Transformer without Positional Encoding
von: Zhang, Shen, et al.
Veröffentlicht: (2025)
von: Zhang, Shen, et al.
Veröffentlicht: (2025)
Self-Prophetic Decoding to Unlock Visual Search in LVLMs
von: He, Zhendong, et al.
Veröffentlicht: (2026)
von: He, Zhendong, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Semi-supervised 3D Object Detection with PatchTeacher and PillarMix
von: Wu, Xiaopei, et al.
Veröffentlicht: (2024) -
An Efficient and Effective Transformer Decoder-Based Framework for Multi-Task Visual Grounding
von: Chen, Wei, et al.
Veröffentlicht: (2024) -
APE: Agentic Prompt Enhancer for Image Generation and Editing
von: Huang, Zijian, et al.
Veröffentlicht: (2026) -
SelFLoc: Selective Feature Fusion for Large-scale Point Cloud-based Place Recognition
von: Qiu, Qibo, et al.
Veröffentlicht: (2023) -
RAVE: Rate-Adaptive Visual Encoding for 3D Gaussian Splatting
von: Tran, Hoang-Nhat, et al.
Veröffentlicht: (2025)