An Investigation on The Position Encoding in Vision-Based Dynamics Prediction
Fuente:
arXiv
Saved in:
| Main Authors: | Zhu, Jiageng, Xie, Hanchen, Li, Jiazhi, Khayatkhoei, Mahyar, AbdAlmageed, Wael |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Look, Learn and Leverage (L$^3$): Mitigating Visual-Domain Shift and Discovering Intrinsic Relations via Symbolic Alignment
by: Xie, Hanchen, et al.
Published: (2024)
by: Xie, Hanchen, et al.
Published: (2024)
ManiFPT: Defining and Analyzing Fingerprints of Generative Models
by: Song, Hae Jin, et al.
Published: (2024)
by: Song, Hae Jin, et al.
Published: (2024)
Unsupervised Multimodal Deepfake Detection Using Intra- and Cross-Modal Inconsistencies
by: Tian, Mulin, et al.
Published: (2023)
by: Tian, Mulin, et al.
Published: (2023)
A Neuro-Symbolic Framework Combining Inductive and Deductive Reasoning for Autonomous Driving Planning
by: Wei, Hongyan, et al.
Published: (2026)
by: Wei, Hongyan, et al.
Published: (2026)
A Critical Review of Predominant Bias in Neural Networks
by: Li, Jiazhi, et al.
Published: (2025)
by: Li, Jiazhi, et al.
Published: (2025)
TRIGS: Trojan Identification from Gradient-based Signatures
by: Hussein, Mohamed E., et al.
Published: (2023)
by: Hussein, Mohamed E., et al.
Published: (2023)
AS2 -- Attention-Based Soft Answer Sets: An End-to-End Differentiable Neuro-Soft-Symbolic Reasoning Architecture
by: AbdAlmageed, Wael
Published: (2026)
by: AbdAlmageed, Wael
Published: (2026)
DiffusionCounterfactuals: Inferring High-dimensional Counterfactuals with Guidance of Causal Representations
by: Zhu, Jiageng, et al.
Published: (2024)
by: Zhu, Jiageng, et al.
Published: (2024)
Causal Representation Learning on High-Dimensional Data: Benchmarks, Reproducibility, and Evaluation Metrics
by: Sadeghi, Alireza, et al.
Published: (2026)
by: Sadeghi, Alireza, et al.
Published: (2026)
Towards Perceiving Small Visual Details in Zero-shot Visual Question Answering with Multimodal LLMs
by: Zhang, Jiarui, et al.
Published: (2023)
by: Zhang, Jiarui, et al.
Published: (2023)
MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs
by: Zhang, Jiarui, et al.
Published: (2025)
by: Zhang, Jiarui, et al.
Published: (2025)
Exploring Perceptual Limitation of Multimodal Large Language Models
by: Zhang, Jiarui, et al.
Published: (2024)
by: Zhang, Jiarui, et al.
Published: (2024)
Revisiting Multimodal Positional Encoding in Vision-Language Models
by: Huang, Jie, et al.
Published: (2025)
by: Huang, Jie, et al.
Published: (2025)
Positional Encoding Field
by: Bai, Yunpeng, et al.
Published: (2025)
by: Bai, Yunpeng, et al.
Published: (2025)
Weierstrass Positional Encoding for Vision Transformers
by: Xin, Zhihang, et al.
Published: (2026)
by: Xin, Zhihang, et al.
Published: (2026)
V2PE: Improving Multimodal Long-Context Capability of Vision-Language Models with Variable Visual Position Encoding
by: Ge, Junqi, et al.
Published: (2024)
by: Ge, Junqi, et al.
Published: (2024)
Head-wise Adaptive Rotary Positional Encoding for Fine-Grained Image Generation
by: Li, Jiaye, et al.
Published: (2025)
by: Li, Jiaye, et al.
Published: (2025)
A Weighted Vision Transformer-Based Multi-Task Learning Framework for Predicting ADAS-Cog Scores
by: Hamid, Nur Amirah Abd, et al.
Published: (2025)
by: Hamid, Nur Amirah Abd, et al.
Published: (2025)
A 2D Semantic-Aware Position Encoding for Vision Transformers
by: Chen, Xi, et al.
Published: (2025)
by: Chen, Xi, et al.
Published: (2025)
Parabolic Position Encoding: Vision-Centric, Principled, Extrapolatable, General
by: Øhrstrøm, Christoffer Koo, et al.
Published: (2026)
by: Øhrstrøm, Christoffer Koo, et al.
Published: (2026)
OMEGA: Optimized Multimodal Position Encoding Index Derivation with Global Adaptive Scaling for Vision-Language Models
by: Huang, Ruoxiang, et al.
Published: (2025)
by: Huang, Ruoxiang, et al.
Published: (2025)
Unified Camera Positional Encoding for Controlled Video Generation
by: Zhang, Cheng, et al.
Published: (2025)
by: Zhang, Cheng, et al.
Published: (2025)
Seeing the Unseen: Mask-Driven Positional Encoding and Strip-Convolution Context Modeling for Cross-View Object Geo-Localization
by: Hu, Shuhan, et al.
Published: (2025)
by: Hu, Shuhan, et al.
Published: (2025)
See&Say: Vision Language Guided Safe Zone Detection for Autonomous Package Delivery Drones
by: Ghazanfari, Mahyar, et al.
Published: (2026)
by: Ghazanfari, Mahyar, et al.
Published: (2026)
Cameras as Relative Positional Encoding
by: Li, Ruilong, et al.
Published: (2025)
by: Li, Ruilong, et al.
Published: (2025)
LEDiT: Your Length-Extrapolatable Diffusion Transformer without Positional Encoding
by: Zhang, Shen, et al.
Published: (2025)
by: Zhang, Shen, et al.
Published: (2025)
Anchor-free Cross-view Object Geo-localization with Gaussian Position Encoding and Cross-view Association
by: Ling, Xingtao, et al.
Published: (2025)
by: Ling, Xingtao, et al.
Published: (2025)
KeyPoint Relative Position Encoding for Face Recognition
by: Kim, Minchul, et al.
Published: (2024)
by: Kim, Minchul, et al.
Published: (2024)
Seamless Human Motion Composition with Blended Positional Encodings
by: Barquero, German, et al.
Published: (2024)
by: Barquero, German, et al.
Published: (2024)
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding
by: Chen, Zhanpeng, et al.
Published: (2025)
by: Chen, Zhanpeng, et al.
Published: (2025)
Mitigating Coordinate Prediction Bias from Positional Encoding Failures
by: Tao, Xingjian, et al.
Published: (2025)
by: Tao, Xingjian, et al.
Published: (2025)
Leveraging Positional Encoding for Robust Multi-Reference-Based Object 6D Pose Estimation
by: Park, Jaewoo, et al.
Published: (2024)
by: Park, Jaewoo, et al.
Published: (2024)
Dynamics Based Neural Encoding with Inter-Intra Region Connectivity
by: Gamal, Mai, et al.
Published: (2024)
by: Gamal, Mai, et al.
Published: (2024)
Predicting the Encoding Error of SIRENs
by: Vonderfecht, Jeremy, et al.
Published: (2024)
by: Vonderfecht, Jeremy, et al.
Published: (2024)
Multi-View Large Reconstruction Model via Geometry-Aware Positional Encoding and Attention
by: Li, Mengfei, et al.
Published: (2024)
by: Li, Mengfei, et al.
Published: (2024)
Positional Encodings Anchor Spatial Structure in Vision Transformers: A Geometric Perspective on Robustness
by: Mannes, Mahmoud
Published: (2026)
by: Mannes, Mahmoud
Published: (2026)
Uncovering Brain-Like Hierarchical Patterns in Vision-Language Models through fMRI-Based Neural Encoding
by: Ren, Yudan, et al.
Published: (2025)
by: Ren, Yudan, et al.
Published: (2025)
Boosting Resolution Generalization of Diffusion Transformers with Randomized Positional Encodings
by: Hou, Liang, et al.
Published: (2025)
by: Hou, Liang, et al.
Published: (2025)
Salient Temporal Encoding for Dynamic Scene Graph Generation
by: Zhu, Zhihao
Published: (2025)
by: Zhu, Zhihao
Published: (2025)
PuzzleBoard: A New Camera Calibration Pattern with Position Encoding
by: Stelldinger, Peer, et al.
Published: (2024)
by: Stelldinger, Peer, et al.
Published: (2024)
Similar Items
-
Look, Learn and Leverage (L$^3$): Mitigating Visual-Domain Shift and Discovering Intrinsic Relations via Symbolic Alignment
by: Xie, Hanchen, et al.
Published: (2024) -
ManiFPT: Defining and Analyzing Fingerprints of Generative Models
by: Song, Hae Jin, et al.
Published: (2024) -
Unsupervised Multimodal Deepfake Detection Using Intra- and Cross-Modal Inconsistencies
by: Tian, Mulin, et al.
Published: (2023) -
A Neuro-Symbolic Framework Combining Inductive and Deductive Reasoning for Autonomous Driving Planning
by: Wei, Hongyan, et al.
Published: (2026) -
A Critical Review of Predominant Bias in Neural Networks
by: Li, Jiazhi, et al.
Published: (2025)