GeoPE:A Unified Geometric Positional Embedding for Structured Tensors
Fuente:
arXiv
Guardado en:
| Autores principales: | Yao, Yupu, Yang, Bowen |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
FishRoPE: Projective Rotary Position Embeddings for Omnidirectional Visual Perception
por: Ahuja, Rahul, et al.
Publicado: (2026)
por: Ahuja, Rahul, et al.
Publicado: (2026)
ComRoPE: Scalable and Robust Rotary Position Embedding Parameterized by Trainable Commuting Angle Matrices
por: Yu, Hao, et al.
Publicado: (2025)
por: Yu, Hao, et al.
Publicado: (2025)
SoPE: Spherical Coordinate-Based Positional Embedding for Enhancing Spatial Perception of 3D LVLMs
por: Ye, Guanting, et al.
Publicado: (2026)
por: Ye, Guanting, et al.
Publicado: (2026)
Renormalization Group Guided Tensor Network Structure Search
por: Wang, Maolin, et al.
Publicado: (2025)
por: Wang, Maolin, et al.
Publicado: (2025)
APLA: Additional Perturbation for Latent Noise with Adversarial Training Enables Consistency
por: Yao, Yupu, et al.
Publicado: (2023)
por: Yao, Yupu, et al.
Publicado: (2023)
CRePE: Curved Ray Expectation Positional Encoding for Unified-Camera-Controlled Video Generation
por: Jin, Seonghyun, et al.
Publicado: (2026)
por: Jin, Seonghyun, et al.
Publicado: (2026)
GERA: Geometric Embedding for Efficient Point Registration Analysis
por: Li, Geng, et al.
Publicado: (2024)
por: Li, Geng, et al.
Publicado: (2024)
SeqPE: Transformer with Sequential Position Encoding
por: Li, Huayang, et al.
Publicado: (2025)
por: Li, Huayang, et al.
Publicado: (2025)
PixelBytes: Catching Unified Embedding for Multimodal Generation
por: Furfaro, Fabien
Publicado: (2024)
por: Furfaro, Fabien
Publicado: (2024)
Enhancing Angular Resolution via Directionality Encoding and Geometric Constraints in Brain Diffusion Tensor Imaging
por: Chen, Sheng, et al.
Publicado: (2024)
por: Chen, Sheng, et al.
Publicado: (2024)
URoPE: Universal Relative Position Embedding across Geometric Spaces
por: Xie, Yichen, et al.
Publicado: (2026)
por: Xie, Yichen, et al.
Publicado: (2026)
Cross-Axis Transformer with 3D Rotary Positional Embeddings
por: Erickson, Lily
Publicado: (2023)
por: Erickson, Lily
Publicado: (2023)
FoundationPose: Unified 6D Pose Estimation and Tracking of Novel Objects
por: Wen, Bowen, et al.
Publicado: (2023)
por: Wen, Bowen, et al.
Publicado: (2023)
LOOPE: Learnable Optimal Patch Order in Positional Embeddings for Vision Transformers
por: Chowdhury, Md Abtahi Majeed, et al.
Publicado: (2025)
por: Chowdhury, Md Abtahi Majeed, et al.
Publicado: (2025)
FQ-PETR: Fully Quantized Position Embedding Transformation for Multi-View 3D Object Detection
por: Yu, Jiangyong, et al.
Publicado: (2025)
por: Yu, Jiangyong, et al.
Publicado: (2025)
Geometric Decoupling: Diagnosing the Structural Instability of Latent
por: Liang, Yuanbang, et al.
Publicado: (2026)
por: Liang, Yuanbang, et al.
Publicado: (2026)
Geometrical Properties of Text Token Embeddings for Strong Semantic Binding in Text-to-Image Generation
por: Seo, Hoigi, et al.
Publicado: (2025)
por: Seo, Hoigi, et al.
Publicado: (2025)
ICDAR 2025 Competition on End-to-End Document Image Machine Translation Towards Complex Layouts
por: Zhang, Yaping, et al.
Publicado: (2026)
por: Zhang, Yaping, et al.
Publicado: (2026)
GeoSURGE: Geo-localization using Semantic Fusion with Hierarchy of Geographic Embeddings
por: Daruna, Angel, et al.
Publicado: (2025)
por: Daruna, Angel, et al.
Publicado: (2025)
From Noisy Labels to Intrinsic Structure: A Geometric-Structural Dual-Guided Framework for Noise-Robust Medical Image Segmentation
por: Wang, Tao, et al.
Publicado: (2025)
por: Wang, Tao, et al.
Publicado: (2025)
GeoGuide: Hierarchical Geometric Guidance for Open-Vocabulary 3D Semantic Segmentation
por: Tao, Xujing, et al.
Publicado: (2026)
por: Tao, Xujing, et al.
Publicado: (2026)
GeoSym127K: Scalable Symbolically-verifiable Synthesis for Multimodal Geometric Reasoning
por: Jing, Jinhao, et al.
Publicado: (2026)
por: Jing, Jinhao, et al.
Publicado: (2026)
A Circular Argument : Does RoPE need to be Equivariant for Vision?
por: van de Geijn, Chase, et al.
Publicado: (2025)
por: van de Geijn, Chase, et al.
Publicado: (2025)
Single-to-mix Modality Alignment with Multimodal Large Language Model for Document Image Machine Translation
por: Liang, Yupu, et al.
Publicado: (2025)
por: Liang, Yupu, et al.
Publicado: (2025)
GeoDiffusion: Text-Prompted Geometric Control for Object Detection Data Generation
por: Chen, Kai, et al.
Publicado: (2023)
por: Chen, Kai, et al.
Publicado: (2023)
CapGeo: A Caption-Assisted Approach to Geometric Reasoning
por: Li, Yuying, et al.
Publicado: (2025)
por: Li, Yuying, et al.
Publicado: (2025)
EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization
por: Wang, Xiaoqi, et al.
Publicado: (2025)
por: Wang, Xiaoqi, et al.
Publicado: (2025)
HiSciBench: A Hierarchical Multi-disciplinary Benchmark for Scientific Intelligence from Reading to Discovery
por: Zhang, Yaping, et al.
Publicado: (2025)
por: Zhang, Yaping, et al.
Publicado: (2025)
Q-Ponder: A Unified Training Pipeline for Reasoning-based Visual Quality Assessment
por: Cai, Zhuoxuan, et al.
Publicado: (2025)
por: Cai, Zhuoxuan, et al.
Publicado: (2025)
Socratic-Geo: Synthetic Data Generation and Geometric Reasoning via Multi-Agent Interaction
por: Jiao, Zhengbo, et al.
Publicado: (2026)
por: Jiao, Zhengbo, et al.
Publicado: (2026)
Nexus-Gen: Unified Image Understanding, Generation, and Editing via Prefilled Autoregression in Shared Embedding Space
por: Zhang, Hong, et al.
Publicado: (2025)
por: Zhang, Hong, et al.
Publicado: (2025)
Security Tensors as a Cross-Modal Bridge: Extending Text-Aligned Safety to Vision in LVLM
por: Li, Shen, et al.
Publicado: (2025)
por: Li, Shen, et al.
Publicado: (2025)
Aether: Geometric-Aware Unified World Modeling
por: Aether Team, et al.
Publicado: (2025)
por: Aether Team, et al.
Publicado: (2025)
GeoBiked: A Dataset with Geometric Features and Automated Labeling Techniques to Enable Deep Generative Models in Engineering Design
por: Mueller, Phillip, et al.
Publicado: (2024)
por: Mueller, Phillip, et al.
Publicado: (2024)
Improving MLLM's Document Image Machine Translation via Synchronously Self-reviewing Its OCR Proficiency
por: Liang, Yupu, et al.
Publicado: (2025)
por: Liang, Yupu, et al.
Publicado: (2025)
ParkFormer: A Transformer-Based Parking Policy with Goal Embedding and Pedestrian-Aware Control
por: Fu, Jun, et al.
Publicado: (2025)
por: Fu, Jun, et al.
Publicado: (2025)
MoPE-CLIP: Structured Pruning for Efficient Vision-Language Models with Module-wise Pruning Error Metric
por: Lin, Haokun, et al.
Publicado: (2024)
por: Lin, Haokun, et al.
Publicado: (2024)
EarthSpatialBench: Benchmarking Spatial Reasoning Capabilities of Multimodal LLMs on Earth Imagery
por: Xu, Zelin, et al.
Publicado: (2026)
por: Xu, Zelin, et al.
Publicado: (2026)
VERDI: VLM-Embedded Reasoning for Autonomous Driving
por: Feng, Bowen, et al.
Publicado: (2025)
por: Feng, Bowen, et al.
Publicado: (2025)
Hunyuan3D 1.0: A Unified Framework for Text-to-3D and Image-to-3D Generation
por: Yang, Xianghui, et al.
Publicado: (2024)
por: Yang, Xianghui, et al.
Publicado: (2024)
Ejemplares similares
-
FishRoPE: Projective Rotary Position Embeddings for Omnidirectional Visual Perception
por: Ahuja, Rahul, et al.
Publicado: (2026) -
ComRoPE: Scalable and Robust Rotary Position Embedding Parameterized by Trainable Commuting Angle Matrices
por: Yu, Hao, et al.
Publicado: (2025) -
SoPE: Spherical Coordinate-Based Positional Embedding for Enhancing Spatial Perception of 3D LVLMs
por: Ye, Guanting, et al.
Publicado: (2026) -
Renormalization Group Guided Tensor Network Structure Search
por: Wang, Maolin, et al.
Publicado: (2025) -
APLA: Additional Perturbation for Latent Noise with Adversarial Training Enables Consistency
por: Yao, Yupu, et al.
Publicado: (2023)