Galaxy Walker: Geometry-aware VLMs For Galaxy-scale Understanding
Fuente:
arXiv
Guardado en:
| Autores principales: | Chen, Tianyu, Fu, Xingcheng, Gao, Yisen, Qian, Haodong, Wei, Yuecen, Yan, Kun, Zhou, Haoyi, Li, Jianxin |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
ETC: Extreme Token Compression via Task-aware Visual Information Distillation in VLMs
por: Gao, Yiling, et al.
Publicado: (2026)
por: Gao, Yiling, et al.
Publicado: (2026)
Towards Long-window Anchoring in Vision-Language Model Distillation
por: Zhou, Haoyi, et al.
Publicado: (2025)
por: Zhou, Haoyi, et al.
Publicado: (2025)
Reinforcing 3D Understanding in Point-VLMs via Geometric Reward Credit Assignment
por: Chen, Jingkun, et al.
Publicado: (2026)
por: Chen, Jingkun, et al.
Publicado: (2026)
DocRefine: An Intelligent Framework for Scientific Document Understanding and Content Optimization based on Multimodal Large Model Agents
por: Qian, Kun, et al.
Publicado: (2025)
por: Qian, Kun, et al.
Publicado: (2025)
MMTok: Multimodal Coverage Maximization for Efficient Inference of VLMs
por: Dong, Sixun, et al.
Publicado: (2025)
por: Dong, Sixun, et al.
Publicado: (2025)
Are VLMs Ready for Lane Topology Awareness in Autonomous Driving?
por: Chen, Xin, et al.
Publicado: (2025)
por: Chen, Xin, et al.
Publicado: (2025)
Deep Pre-Alignment for VLMs
por: Yu, Tianyu, et al.
Publicado: (2026)
por: Yu, Tianyu, et al.
Publicado: (2026)
ChartGalaxy: A Dataset for Infographic Chart Understanding and Generation
por: Li, Zhen, et al.
Publicado: (2025)
por: Li, Zhen, et al.
Publicado: (2025)
SGTA: Scene-Graph Based Multi-Modal Traffic Agent for Video Understanding
por: Zhou, Xingcheng, et al.
Publicado: (2026)
por: Zhou, Xingcheng, et al.
Publicado: (2026)
DualCamCtrl: Dual-Branch Diffusion Model for Geometry-Aware Camera-Controlled Video Generation
por: Zhang, Hongfei, et al.
Publicado: (2025)
por: Zhang, Hongfei, et al.
Publicado: (2025)
Exploring Large Vision-Language Models for Robust and Efficient Industrial Anomaly Detection
por: Qian, Kun, et al.
Publicado: (2024)
por: Qian, Kun, et al.
Publicado: (2024)
GPT-4V as Traffic Assistant: An In-depth Look at Vision Language Model on Complex Traffic Events
por: Zhou, Xingcheng, et al.
Publicado: (2024)
por: Zhou, Xingcheng, et al.
Publicado: (2024)
Rectify the Regression Bias in Long-Tailed Object Detection
por: Zhu, Ke, et al.
Publicado: (2024)
por: Zhu, Ke, et al.
Publicado: (2024)
Learning Fine-Grained Geometry for Sparse-View Splatting via Cascade Depth Loss
por: Lu, Wenjun, et al.
Publicado: (2025)
por: Lu, Wenjun, et al.
Publicado: (2025)
Gaze-Regularized VLMs for Ego-Centric Behavior Understanding
por: Pani, Anupam, et al.
Publicado: (2026)
por: Pani, Anupam, et al.
Publicado: (2026)
Linear Scaling Video VLMs for Long Video Understanding
por: Eyzaguirre, Cristobal, et al.
Publicado: (2026)
por: Eyzaguirre, Cristobal, et al.
Publicado: (2026)
Data Factory with Minimal Human Effort Using VLMs
por: Ye, Jiaojiao, et al.
Publicado: (2025)
por: Ye, Jiaojiao, et al.
Publicado: (2025)
Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs
por: Qiao, Yuxuan, et al.
Publicado: (2024)
por: Qiao, Yuxuan, et al.
Publicado: (2024)
Geometry-aware Distance Measure for Diverse Hierarchical Structures in Hyperbolic Spaces
por: Li, Pengxiang, et al.
Publicado: (2025)
por: Li, Pengxiang, et al.
Publicado: (2025)
Observe Less, Understand More: Cost-aware Cross-scale Observation for Remote Sensing Understanding
por: Xie, Zhenghao, et al.
Publicado: (2026)
por: Xie, Zhenghao, et al.
Publicado: (2026)
WM-MoE: Weather-aware Multi-scale Mixture-of-Experts for Blind Adverse Weather Removal
por: Luo, Yulin, et al.
Publicado: (2023)
por: Luo, Yulin, et al.
Publicado: (2023)
$π^3$: Permutation-Equivariant Visual Geometry Learning
por: Wang, Yifan, et al.
Publicado: (2025)
por: Wang, Yifan, et al.
Publicado: (2025)
On the Perception Bottleneck of VLMs for Chart Understanding
por: Liu, Junteng, et al.
Publicado: (2025)
por: Liu, Junteng, et al.
Publicado: (2025)
CIVET: Systematic Evaluation of Understanding in VLMs
por: Rizzoli, Massimo, et al.
Publicado: (2025)
por: Rizzoli, Massimo, et al.
Publicado: (2025)
Hypergraph Multi-modal Large Language Model: Exploiting EEG and Eye-tracking Modalities to Evaluate Heterogeneous Responses for Video Understanding
por: Wu, Minghui, et al.
Publicado: (2024)
por: Wu, Minghui, et al.
Publicado: (2024)
LMHaze: Intensity-aware Image Dehazing with a Large-scale Multi-intensity Real Haze Dataset
por: Zhang, Ruikun, et al.
Publicado: (2024)
por: Zhang, Ruikun, et al.
Publicado: (2024)
Does Your 3D Encoder Really Work? When Pretrain-SFT from 2D VLMs Meets 3D VLMs
por: Li, Haoyuan, et al.
Publicado: (2025)
por: Li, Haoyuan, et al.
Publicado: (2025)
GeoGaussian: Geometry-aware Gaussian Splatting for Scene Rendering
por: Li, Yanyan, et al.
Publicado: (2024)
por: Li, Yanyan, et al.
Publicado: (2024)
Stepping VLMs onto the Court: Benchmarking Spatial Intelligence in Sports
por: Yang, Yuchen, et al.
Publicado: (2026)
por: Yang, Yuchen, et al.
Publicado: (2026)
GEARS: Local Geometry-aware Hand-object Interaction Synthesis
por: Zhou, Keyang, et al.
Publicado: (2024)
por: Zhou, Keyang, et al.
Publicado: (2024)
Enhancing Underwater Light Field Images via Global Geometry-aware Diffusion Process
por: Lin, Yuji, et al.
Publicado: (2026)
por: Lin, Yuji, et al.
Publicado: (2026)
GLS: Geometry-aware 3D Language Gaussian Splatting
por: Qiu, Jiaxiong, et al.
Publicado: (2024)
por: Qiu, Jiaxiong, et al.
Publicado: (2024)
Coordinative Learning with Ordinal and Relational Priors for Volumetric Medical Image Segmentation
por: Wang, Haoyi
Publicado: (2025)
por: Wang, Haoyi
Publicado: (2025)
Real-time 3D-aware Portrait Video Relighting
por: Cai, Ziqi, et al.
Publicado: (2024)
por: Cai, Ziqi, et al.
Publicado: (2024)
FlashGS: Efficient 3D Gaussian Splatting for Large-scale and High-resolution Rendering
por: Feng, Guofeng, et al.
Publicado: (2024)
por: Feng, Guofeng, et al.
Publicado: (2024)
Direction-aware multi-scale gradient loss for infrared and visible image fusion
por: Yang, Kaixuan, et al.
Publicado: (2025)
por: Yang, Kaixuan, et al.
Publicado: (2025)
Long Story Short: Disentangling Compositionality and Long-Caption Understanding in Contrastive VLMs
por: Salazar, Israfel, et al.
Publicado: (2025)
por: Salazar, Israfel, et al.
Publicado: (2025)
LVC: A Lightweight Compression Framework for Enhancing VLMs in Long Video Understanding
por: Wang, Ziyi, et al.
Publicado: (2025)
por: Wang, Ziyi, et al.
Publicado: (2025)
Decompose and Compare Consistency: Measuring VLMs' Answer Reliability via Task-Decomposition Consistency Comparison
por: Yang, Qian, et al.
Publicado: (2024)
por: Yang, Qian, et al.
Publicado: (2024)
OmniScience: A Large-scale Multi-modal Dataset for Scientific Image Understanding
por: Tao, Haoyi, et al.
Publicado: (2026)
por: Tao, Haoyi, et al.
Publicado: (2026)
Ejemplares similares
-
ETC: Extreme Token Compression via Task-aware Visual Information Distillation in VLMs
por: Gao, Yiling, et al.
Publicado: (2026) -
Towards Long-window Anchoring in Vision-Language Model Distillation
por: Zhou, Haoyi, et al.
Publicado: (2025) -
Reinforcing 3D Understanding in Point-VLMs via Geometric Reward Credit Assignment
por: Chen, Jingkun, et al.
Publicado: (2026) -
DocRefine: An Intelligent Framework for Scientific Document Understanding and Content Optimization based on Multimodal Large Model Agents
por: Qian, Kun, et al.
Publicado: (2025) -
MMTok: Multimodal Coverage Maximization for Efficient Inference of VLMs
por: Dong, Sixun, et al.
Publicado: (2025)