Cameras as Relative Positional Encoding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Ruilong, Yi, Brent, Liu, Junchen, Gao, Hang, Ma, Yi, Kanazawa, Angjoo |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Estimating Body and Hand Motion in an Ego-sensed World
von: Yi, Brent, et al.
Veröffentlicht: (2024)
von: Yi, Brent, et al.
Veröffentlicht: (2024)
Reconstructing People, Places, and Cameras
von: Müller, Lea, et al.
Veröffentlicht: (2024)
von: Müller, Lea, et al.
Veröffentlicht: (2024)
Test-Time Training with KV Binding Is Secretly Linear Attention
von: Liu, Junchen, et al.
Veröffentlicht: (2026)
von: Liu, Junchen, et al.
Veröffentlicht: (2026)
SOAR: Self-Occluded Avatar Recovery from a Single Video In the Wild
von: Pan, Zhuoyang, et al.
Veröffentlicht: (2024)
von: Pan, Zhuoyang, et al.
Veröffentlicht: (2024)
NeRF-XL: Scaling NeRFs with Multiple GPUs
von: Li, Ruilong, et al.
Veröffentlicht: (2024)
von: Li, Ruilong, et al.
Veröffentlicht: (2024)
Weierstrass Positional Encoding for Vision Transformers
von: Xin, Zhihang, et al.
Veröffentlicht: (2026)
von: Xin, Zhihang, et al.
Veröffentlicht: (2026)
CRePE: Curved Ray Expectation Positional Encoding for Unified-Camera-Controlled Video Generation
von: Jin, Seonghyun, et al.
Veröffentlicht: (2026)
von: Jin, Seonghyun, et al.
Veröffentlicht: (2026)
gsplat: An Open-Source Library for Gaussian Splatting
von: Ye, Vickie, et al.
Veröffentlicht: (2024)
von: Ye, Vickie, et al.
Veröffentlicht: (2024)
Vision-as-Inverse-Graphics Agent via Interleaved Multimodal Reasoning
von: Yin, Shaofeng, et al.
Veröffentlicht: (2026)
von: Yin, Shaofeng, et al.
Veröffentlicht: (2026)
A 2D Semantic-Aware Position Encoding for Vision Transformers
von: Chen, Xi, et al.
Veröffentlicht: (2025)
von: Chen, Xi, et al.
Veröffentlicht: (2025)
Prompt When the Animal is: Temporal Animal Behavior Grounding with Positional Recovery Training
von: Yan, Sheng, et al.
Veröffentlicht: (2024)
von: Yan, Sheng, et al.
Veröffentlicht: (2024)
Equipping Sketch Patches with Context-Aware Positional Encoding for Graphic Sketch Representation
von: Zang, Sicong, et al.
Veröffentlicht: (2024)
von: Zang, Sicong, et al.
Veröffentlicht: (2024)
Predict-Optimize-Distill: A Self-Improving Cycle for 4D Object Understanding
von: Wu, Mingxuan, et al.
Veröffentlicht: (2025)
von: Wu, Mingxuan, et al.
Veröffentlicht: (2025)
MCOO-SLAM: A Multi-Camera Omnidirectional Object SLAM System
von: Pan, Miaoxin, et al.
Veröffentlicht: (2025)
von: Pan, Miaoxin, et al.
Veröffentlicht: (2025)
I2VControl-Camera: Precise Video Camera Control with Adjustable Motion Strength
von: Feng, Wanquan, et al.
Veröffentlicht: (2024)
von: Feng, Wanquan, et al.
Veröffentlicht: (2024)
GeLoc3r: Enhancing Relative Camera Pose Regression with Geometric Consistency Regularization
von: Li, Jingxing, et al.
Veröffentlicht: (2025)
von: Li, Jingxing, et al.
Veröffentlicht: (2025)
CamReasoner: Reinforcing Camera Movement Understanding via Structured Spatial Reasoning
von: Wu, Hang, et al.
Veröffentlicht: (2026)
von: Wu, Hang, et al.
Veröffentlicht: (2026)
Accelerating Video Generation Inference with Sequential-Parallel 3D Positional Encoding Using a Global Time Index
von: Yuan, Chao, et al.
Veröffentlicht: (2026)
von: Yuan, Chao, et al.
Veröffentlicht: (2026)
Splatfacto-W: A Nerfstudio Implementation of Gaussian Splatting for Unconstrained Photo Collections
von: Xu, Congrong, et al.
Veröffentlicht: (2024)
von: Xu, Congrong, et al.
Veröffentlicht: (2024)
Human-level 3D shape perception emerges from multi-view learning
von: Bonnen, Tyler, et al.
Veröffentlicht: (2026)
von: Bonnen, Tyler, et al.
Veröffentlicht: (2026)
SeqPE: Transformer with Sequential Position Encoding
von: Li, Huayang, et al.
Veröffentlicht: (2025)
von: Li, Huayang, et al.
Veröffentlicht: (2025)
Vanishing Depth: A Depth Adapter with Positional Depth Encoding for Generalized Image Encoders
von: Koch, Paul, et al.
Veröffentlicht: (2025)
von: Koch, Paul, et al.
Veröffentlicht: (2025)
Efficient Point Cloud Processing with High-Dimensional Positional Encoding and Non-Local MLPs
von: Zou, Yanmei, et al.
Veröffentlicht: (2026)
von: Zou, Yanmei, et al.
Veröffentlicht: (2026)
PAS: A Training-Free Stabilizer for Temporal Encoding in Video LLMs
von: Sun, Bowen, et al.
Veröffentlicht: (2025)
von: Sun, Bowen, et al.
Veröffentlicht: (2025)
Mitigating Coordinate Prediction Bias from Positional Encoding Failures
von: Tao, Xingjian, et al.
Veröffentlicht: (2025)
von: Tao, Xingjian, et al.
Veröffentlicht: (2025)
PyramidStyler: Transformer-Based Neural Style Transfer with Pyramidal Positional Encoding and Reinforcement Learning
von: Durairaju, Raahul Krishna, et al.
Veröffentlicht: (2025)
von: Durairaju, Raahul Krishna, et al.
Veröffentlicht: (2025)
Tracking spatial temporal details in ultrasound long video via wavelet analysis and memory bank
von: Zhang, Chenxiao, et al.
Veröffentlicht: (2025)
von: Zhang, Chenxiao, et al.
Veröffentlicht: (2025)
InCaRPose: In-Cabin Relative Camera Pose Estimation Model and Dataset
von: Stillger, Felix, et al.
Veröffentlicht: (2026)
von: Stillger, Felix, et al.
Veröffentlicht: (2026)
Finite Difference Flow Optimization for RL Post-Training of Text-to-Image Models
von: McAllister, David, et al.
Veröffentlicht: (2026)
von: McAllister, David, et al.
Veröffentlicht: (2026)
Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding
von: Chen, Zhanpeng, et al.
Veröffentlicht: (2025)
von: Chen, Zhanpeng, et al.
Veröffentlicht: (2025)
MAP: Unleashing Hybrid Mamba-Transformer Vision Backbone's Potential with Masked Autoregressive Pretraining
von: Liu, Yunze, et al.
Veröffentlicht: (2024)
von: Liu, Yunze, et al.
Veröffentlicht: (2024)
AerialVG: A Challenging Benchmark for Aerial Visual Grounding by Exploring Positional Relations
von: Liu, Junli, et al.
Veröffentlicht: (2025)
von: Liu, Junli, et al.
Veröffentlicht: (2025)
Robot See Robot Do: Imitating Articulated Object Manipulation with Monocular 4D Reconstruction
von: Kerr, Justin, et al.
Veröffentlicht: (2024)
von: Kerr, Justin, et al.
Veröffentlicht: (2024)
MemCam: Memory-Augmented Camera Control for Consistent Video Generation
von: Gao, Xinhang, et al.
Veröffentlicht: (2026)
von: Gao, Xinhang, et al.
Veröffentlicht: (2026)
First RAG, Second SEG: A Training-Free Paradigm for Camouflaged Object Detection
von: Liu, Wutao, et al.
Veröffentlicht: (2025)
von: Liu, Wutao, et al.
Veröffentlicht: (2025)
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization
von: Wang, Xiaoqi, et al.
Veröffentlicht: (2025)
von: Wang, Xiaoqi, et al.
Veröffentlicht: (2025)
VidCRAFT3: Camera, Object, and Lighting Control for Image-to-Video Generation
von: Zheng, Sixiao, et al.
Veröffentlicht: (2025)
von: Zheng, Sixiao, et al.
Veröffentlicht: (2025)
C^2ROPE: Causal Continuous Rotary Positional Encoding for 3D Large Multimodal-Models Reasoning
von: Ye, Guanting, et al.
Veröffentlicht: (2026)
von: Ye, Guanting, et al.
Veröffentlicht: (2026)
Shape of Motion: 4D Reconstruction from a Single Video
von: Wang, Qianqian, et al.
Veröffentlicht: (2024)
von: Wang, Qianqian, et al.
Veröffentlicht: (2024)
ERetinex: Event Camera Meets Retinex Theory for Low-Light Image Enhancement
von: Guo, Xuejian, et al.
Veröffentlicht: (2025)
von: Guo, Xuejian, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Estimating Body and Hand Motion in an Ego-sensed World
von: Yi, Brent, et al.
Veröffentlicht: (2024) -
Reconstructing People, Places, and Cameras
von: Müller, Lea, et al.
Veröffentlicht: (2024) -
Test-Time Training with KV Binding Is Secretly Linear Attention
von: Liu, Junchen, et al.
Veröffentlicht: (2026) -
SOAR: Self-Occluded Avatar Recovery from a Single Video In the Wild
von: Pan, Zhuoyang, et al.
Veröffentlicht: (2024) -
NeRF-XL: Scaling NeRFs with Multiple GPUs
von: Li, Ruilong, et al.
Veröffentlicht: (2024)