VG3T: Visual Geometry Grounded Gaussian Transformer
Fuente:
arXiv
Salvato in:
| Autori principali: | Kim, Junho, Lee, Seongwon |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
ExpVG: Investigating the Design Space of Visual Grounding in Multimodal Large Language Model
di: Kang, Weitai, et al.
Pubblicazione: (2025)
di: Kang, Weitai, et al.
Pubblicazione: (2025)
VG-CoT: Towards Trustworthy Visual Reasoning via Grounded Chain-of-Thought
di: Lim, Byeonggeuk, et al.
Pubblicazione: (2026)
di: Lim, Byeonggeuk, et al.
Pubblicazione: (2026)
FedVG: Gradient-Guided Aggregation for Enhanced Federated Learning
di: Devkota, Alina, et al.
Pubblicazione: (2026)
di: Devkota, Alina, et al.
Pubblicazione: (2026)
Streaming 4D Visual Geometry Transformer
di: Zhuo, Dong, et al.
Pubblicazione: (2025)
di: Zhuo, Dong, et al.
Pubblicazione: (2025)
TTA-DAME: Test-Time Adaptation with Domain Augmentation and Model Ensemble for Dynamic Driving Conditions
di: Jeon, Dongjae, et al.
Pubblicazione: (2025)
di: Jeon, Dongjae, et al.
Pubblicazione: (2025)
Look Every Frame All at Once: Video-Ma$^2$mba for Efficient Long-form Video Understanding with Multi-Axis Gradient Checkpointing
di: Lee, Hosu, et al.
Pubblicazione: (2024)
di: Lee, Hosu, et al.
Pubblicazione: (2024)
Causal Unsupervised Semantic Segmentation
di: Kim, Junho, et al.
Pubblicazione: (2023)
di: Kim, Junho, et al.
Pubblicazione: (2023)
GenOL: Generating Diverse Examples for Name-only Online Learning
di: Seo, Minhyuk, et al.
Pubblicazione: (2024)
di: Seo, Minhyuk, et al.
Pubblicazione: (2024)
Toward an Artificial General Teacher: Procedural Geometry Data Generation and Visual Grounding with Vision-Language Models
di: Nguyen-Truong, Hai, et al.
Pubblicazione: (2026)
di: Nguyen-Truong, Hai, et al.
Pubblicazione: (2026)
Visual Test-time Scaling for GUI Agent Grounding
di: Luo, Tiange, et al.
Pubblicazione: (2025)
di: Luo, Tiange, et al.
Pubblicazione: (2025)
GSplatLoc: Grounding Keypoint Descriptors into 3D Gaussian Splatting for Improved Visual Localization
di: Sidorov, Gennady, et al.
Pubblicazione: (2024)
di: Sidorov, Gennady, et al.
Pubblicazione: (2024)
Grounding Continuous Representations in Geometry: Equivariant Neural Fields
di: Wessels, David R, et al.
Pubblicazione: (2024)
di: Wessels, David R, et al.
Pubblicazione: (2024)
AerialVG: A Challenging Benchmark for Aerial Visual Grounding by Exploring Positional Relations
di: Liu, Junli, et al.
Pubblicazione: (2025)
di: Liu, Junli, et al.
Pubblicazione: (2025)
PropVG: End-to-End Proposal-Driven Visual Grounding with Multi-Granularity Discrimination
di: Dai, Ming, et al.
Pubblicazione: (2025)
di: Dai, Ming, et al.
Pubblicazione: (2025)
AgroVG: A Large-Scale Multi-Source Benchmark for Agricultural Visual Grounding
di: Li, Haocheng, et al.
Pubblicazione: (2026)
di: Li, Haocheng, et al.
Pubblicazione: (2026)
SimVG: A Simple Framework for Visual Grounding with Decoupled Multi-modal Fusion
di: Dai, Ming, et al.
Pubblicazione: (2024)
di: Dai, Ming, et al.
Pubblicazione: (2024)
Emergent Visual Grounding in Large Multimodal Models Without Grounding Supervision
di: Cao, Shengcao, et al.
Pubblicazione: (2024)
di: Cao, Shengcao, et al.
Pubblicazione: (2024)
$\text{VG}^2$GT: Voxel-Gaussian Splatting Visual Geometry Grounded Transformer
di: Zhao, Yibin, et al.
Pubblicazione: (2026)
di: Zhao, Yibin, et al.
Pubblicazione: (2026)
Scratching Visual Transformer's Back with Uniform Attention
di: Hyeon-Woo, Nam, et al.
Pubblicazione: (2022)
di: Hyeon-Woo, Nam, et al.
Pubblicazione: (2022)
SwiftVGGT: A Scalable Visual Geometry Grounded Transformer for Large-Scale Scenes
di: Lee, Jungho, et al.
Pubblicazione: (2025)
di: Lee, Jungho, et al.
Pubblicazione: (2025)
Facial Wrinkle Segmentation for Cosmetic Dermatology: Pretraining with Texture Map-Based Weak Supervision
di: Moon, Junho, et al.
Pubblicazione: (2024)
di: Moon, Junho, et al.
Pubblicazione: (2024)
ReSpec: Relevance and Specificity Grounded Online Filtering for Learning on Video-Text Data Streams
di: Kim, Chris Dongjoo, et al.
Pubblicazione: (2025)
di: Kim, Chris Dongjoo, et al.
Pubblicazione: (2025)
GeM-VG: Towards Generalized Multi-image Visual Grounding with Multimodal Large Language Models
di: Zheng, Shurong, et al.
Pubblicazione: (2026)
di: Zheng, Shurong, et al.
Pubblicazione: (2026)
Weakly Supervised Pretraining and Multi-Annotator Supervised Finetuning for Facial Wrinkle Detection
di: Moon, Ik Jun, et al.
Pubblicazione: (2024)
di: Moon, Ik Jun, et al.
Pubblicazione: (2024)
GTA: A Geometry-Aware Attention Mechanism for Multi-View Transformers
di: Miyato, Takeru, et al.
Pubblicazione: (2023)
di: Miyato, Takeru, et al.
Pubblicazione: (2023)
By My Eyes: Grounding Multimodal Large Language Models with Sensor Data via Visual Prompting
di: Yoon, Hyungjun, et al.
Pubblicazione: (2024)
di: Yoon, Hyungjun, et al.
Pubblicazione: (2024)
Paint Outside the Box: Synthesizing and Selecting Training Data for Visual Grounding
di: Du, Zilin, et al.
Pubblicazione: (2024)
di: Du, Zilin, et al.
Pubblicazione: (2024)
Aligning Forest and Trees in Images & Long Captions for Visually Grounded Understanding
di: Woo, Byeongju, et al.
Pubblicazione: (2026)
di: Woo, Byeongju, et al.
Pubblicazione: (2026)
VG3S: Visual Geometry Grounded Gaussian Splatting for Semantic Occupancy Prediction
di: Yan, Xiaoyang, et al.
Pubblicazione: (2026)
di: Yan, Xiaoyang, et al.
Pubblicazione: (2026)
Domain-Invariant Per-Frame Feature Extraction for Cross-Domain Imitation Learning with Visual Observations
di: Kim, Minung, et al.
Pubblicazione: (2025)
di: Kim, Minung, et al.
Pubblicazione: (2025)
SHeaP: Self-Supervised Head Geometry Predictor Learned via 2D Gaussians
di: Schoneveld, Liam, et al.
Pubblicazione: (2025)
di: Schoneveld, Liam, et al.
Pubblicazione: (2025)
Audio-3DVG: Unified Audio -- Point Cloud Fusion for 3D Visual Grounding
di: Cao-Dinh, Duc, et al.
Pubblicazione: (2025)
di: Cao-Dinh, Duc, et al.
Pubblicazione: (2025)
Good Token Hunting: A Hitchhiker's Guide to Token Selection for Visual Geometry Transformers
di: Zheng, Shuhong, et al.
Pubblicazione: (2026)
di: Zheng, Shuhong, et al.
Pubblicazione: (2026)
Naturally Supervised 3D Visual Grounding with Language-Regularized Concept Learners
di: Feng, Chun, et al.
Pubblicazione: (2024)
di: Feng, Chun, et al.
Pubblicazione: (2024)
Geometry-Correct Diffusion Posterior Sampling with Denoiser-Pullback Curvature Guidance and Manifold-Aligned Damping
di: Shin, Seunghyeok, et al.
Pubblicazione: (2026)
di: Shin, Seunghyeok, et al.
Pubblicazione: (2026)
Self-supervised Transformation Learning for Equivariant Representations
di: Yu, Jaemyung, et al.
Pubblicazione: (2025)
di: Yu, Jaemyung, et al.
Pubblicazione: (2025)
Pygmalion Effect in Vision: Image-to-Clay Translation for Reflective Geometry Reconstruction
di: Lee, Gayoung, et al.
Pubblicazione: (2025)
di: Lee, Gayoung, et al.
Pubblicazione: (2025)
HART: Efficient Visual Generation with Hybrid Autoregressive Transformer
di: Tang, Haotian, et al.
Pubblicazione: (2024)
di: Tang, Haotian, et al.
Pubblicazione: (2024)
IVR-R1: Refining Trajectories through Iterative Visual-Grounded Reasoning in Reinforcement Learning
di: Li, Chenghao, et al.
Pubblicazione: (2026)
di: Li, Chenghao, et al.
Pubblicazione: (2026)
ViPCap: Retrieval Text-Based Visual Prompts for Lightweight Image Captioning
di: Kim, Taewhan, et al.
Pubblicazione: (2024)
di: Kim, Taewhan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
ExpVG: Investigating the Design Space of Visual Grounding in Multimodal Large Language Model
di: Kang, Weitai, et al.
Pubblicazione: (2025) -
VG-CoT: Towards Trustworthy Visual Reasoning via Grounded Chain-of-Thought
di: Lim, Byeonggeuk, et al.
Pubblicazione: (2026) -
FedVG: Gradient-Guided Aggregation for Enhanced Federated Learning
di: Devkota, Alina, et al.
Pubblicazione: (2026) -
Streaming 4D Visual Geometry Transformer
di: Zhuo, Dong, et al.
Pubblicazione: (2025) -
TTA-DAME: Test-Time Adaptation with Domain Augmentation and Model Ensemble for Dynamic Driving Conditions
di: Jeon, Dongjae, et al.
Pubblicazione: (2025)