InstAP: Instance-Aware Vision-Language Pre-Train for Spatial-Temporal Understanding
Fuente:
arXiv
Guardado en:
| Autores principales: | Kumar, Ashutosh, Saini, Rajat, Pan, Jingjing, Erdogan, Mustafa, Zhang, Mingfang, Dem, Betty Le, Kobori, Norimasa, Kong, Quan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
WTS: A Pedestrian-Centric Traffic Video Dataset for Fine-grained Spatial-Temporal Understanding
por: Kong, Quan, et al.
Publicado: (2024)
por: Kong, Quan, et al.
Publicado: (2024)
CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering
por: Zhang, Mingfang, et al.
Publicado: (2026)
por: Zhang, Mingfang, et al.
Publicado: (2026)
E-VLC: A Real-World Dataset for Event-based Visible Light Communication And Localization
por: Shiba, Shintaro, et al.
Publicado: (2025)
por: Shiba, Shintaro, et al.
Publicado: (2025)
GA3CE: Unconstrained 3D Gaze Estimation with Gaze-Aware 3D Context Encoding
por: Kawana, Yuki, et al.
Publicado: (2025)
por: Kawana, Yuki, et al.
Publicado: (2025)
Evaluation of Mobile Environment for Vehicular Visible Light Communication Using Multiple LEDs and Event Cameras
por: Soga, Ryota, et al.
Publicado: (2025)
por: Soga, Ryota, et al.
Publicado: (2025)
Distance Estimation in Outdoor Driving Environments Using Phase-only Correlation Method with Event Cameras
por: Kobayashi, Masataka, et al.
Publicado: (2025)
por: Kobayashi, Masataka, et al.
Publicado: (2025)
Inst3D-LMM: Instance-Aware 3D Scene Understanding with Multi-modal Instruction Tuning
por: Yu, Hanxun, et al.
Publicado: (2025)
por: Yu, Hanxun, et al.
Publicado: (2025)
InstDrive: Instance-Aware 3D Gaussian Splatting for Driving Scenes
por: Liu, Hongyuan, et al.
Publicado: (2025)
por: Liu, Hongyuan, et al.
Publicado: (2025)
One Trajectory, One Token: Grounded Video Tokenization via Panoptic Sub-object Trajectory
por: Zheng, Chenhao, et al.
Publicado: (2025)
por: Zheng, Chenhao, et al.
Publicado: (2025)
UDA4Inst: Unsupervised Domain Adaptation for Instance Segmentation
por: Guo, Yachan, et al.
Publicado: (2024)
por: Guo, Yachan, et al.
Publicado: (2024)
Books in the Field: Physical Anthropology
por: Polacheck, Dem
Publicado: (1969)
por: Polacheck, Dem
Publicado: (1969)
Trainable Quantum Neural Network for Multiclass Image Classification with the Power of Pre-trained Tree Tensor Networks
por: Murota, Keisuke, et al.
Publicado: (2025)
por: Murota, Keisuke, et al.
Publicado: (2025)
AutoInst: Automatic Instance-Based Segmentation of LiDAR 3D Scans
por: Perauer, Cedric, et al.
Publicado: (2024)
por: Perauer, Cedric, et al.
Publicado: (2024)
FastInstShadow: A Simple Query-Based Model for Instance Shadow Detection
por: Inoue, Takeru, et al.
Publicado: (2025)
por: Inoue, Takeru, et al.
Publicado: (2025)
Enhanced BoxInst for Weakly Supervised Liver Tumor Instance Segmentation in CT Images
por: Shanshan Li, et al.
Publicado: (2025)
por: Shanshan Li, et al.
Publicado: (2025)
Entre littérature et anthropologie : la "raison littéraire" d’Éric Chauvier
por: Laurent DemAnze
Publicado: (2022)
por: Laurent DemAnze
Publicado: (2022)
Leveraging GNSS and Onboard Visual Data from Consumer Vehicles for Robust Road Network Estimation
por: Opra, Balázs, et al.
Publicado: (2024)
por: Opra, Balázs, et al.
Publicado: (2024)
Adapting Pre-Trained Vision Models for Novel Instance Detection and Segmentation
por: Lu, Yangxiao, et al.
Publicado: (2024)
por: Lu, Yangxiao, et al.
Publicado: (2024)
A Temporal Modeling Framework for Video Pre-Training on Video Instance Segmentation
por: Zhong, Qing, et al.
Publicado: (2025)
por: Zhong, Qing, et al.
Publicado: (2025)
BLO-Inst: Bi-Level Optimization Based Alignment of YOLO and SAM for Robust Instance Segmentation
por: Zhang, Li, et al.
Publicado: (2026)
por: Zhang, Li, et al.
Publicado: (2026)
InstCache: A Predictive Cache for LLM Serving
por: Zou, Longwei, et al.
Publicado: (2024)
por: Zou, Longwei, et al.
Publicado: (2024)
Reprojection Errors as Prompts for Efficient Scene Coordinate Regression
por: Liu, Ting-Ru, et al.
Publicado: (2024)
por: Liu, Ting-Ru, et al.
Publicado: (2024)
An Embodied AR Navigation Agent: Integrating BIM with Retrieval-Augmented Generation for Language Guidance
por: Yang, Hsuan-Kung, et al.
Publicado: (2025)
por: Yang, Hsuan-Kung, et al.
Publicado: (2025)
Inst4DGS: Instance-Decomposed 4D Gaussian Splatting with Multi-Video Label Permutation Learning
por: Lee, Yonghan, et al.
Publicado: (2026)
por: Lee, Yonghan, et al.
Publicado: (2026)
Decoupled Spatial and Temporal Processing for Resource Efficient Multichannel Speech Enhancement
por: Pandey, Ashutosh, et al.
Publicado: (2024)
por: Pandey, Ashutosh, et al.
Publicado: (2024)
Impact of Scalar NSI on Spatial and Temporal Correlations in Neutrino Oscillations
por: Yadav, Bhavna, et al.
Publicado: (2024)
por: Yadav, Bhavna, et al.
Publicado: (2024)
Spatially Constrained Transformer with Efficient Global Relation Modelling for Spatio-Temporal Prediction
por: Sao, Ashutosh, et al.
Publicado: (2024)
por: Sao, Ashutosh, et al.
Publicado: (2024)
Preschool Students' Understanding of a Geometric Shape, the Square
por: Erdogan Halat
Publicado: (2016)
por: Erdogan Halat
Publicado: (2016)
Score Broadcast and Decorrelation: A General Framework for Broadcast-Based Credit Assignment
por: Uzun, Mustafa, et al.
Publicado: (2026)
por: Uzun, Mustafa, et al.
Publicado: (2026)
O INDIVÍDUO MODERNO E O ROMANCE: UM POSSÍVEL DIÁLOGO ENTRE O JOVEM LUKÁCS E MARTHE ROBERT
por: Eduardo Toshio Kobori
Publicado: (2015)
por: Eduardo Toshio Kobori
Publicado: (2015)
Synthetic Visual Genome
por: Park, Jae Sung, et al.
Publicado: (2025)
por: Park, Jae Sung, et al.
Publicado: (2025)
3DVLA: Enhancing Vision-Language-Action Models via 3D Spatial and Instance Understanding
por: Xia, Zhongyu, et al.
Publicado: (2026)
por: Xia, Zhongyu, et al.
Publicado: (2026)
Instance-Aware Group Quantization for Vision Transformers
por: Moon, Jaehyeon, et al.
Publicado: (2024)
por: Moon, Jaehyeon, et al.
Publicado: (2024)
Leveraging RGB Images for Pre-Training of Event-Based Hand Pose Estimation
por: Liu, Ruicong, et al.
Publicado: (2025)
por: Liu, Ruicong, et al.
Publicado: (2025)
Unsupervised Pre-Training for 3D Leaf Instance Segmentation
por: Roggiolani, Gianmarco, et al.
Publicado: (2024)
por: Roggiolani, Gianmarco, et al.
Publicado: (2024)
LeafInst - Unified Instance Segmentation Network for Fine-Grained Forestry Leaf Phenotype Analysis: A New UAV based Benchmark
por: Luo, Taige, et al.
Publicado: (2026)
por: Luo, Taige, et al.
Publicado: (2026)
Evaluation of bioactive potential of Edirne beyaz peyniri (Edirne white cheese) produced from different milk types during maturation time
por: Fatmagül Halici Demİr, et al.
Publicado: (2025)
por: Fatmagül Halici Demİr, et al.
Publicado: (2025)
Spatial-Temporal Pre-Training for Embryo Viability Prediction Using Time-Lapse Videos
por: Shi, Zhiyi, et al.
Publicado: (2025)
por: Shi, Zhiyi, et al.
Publicado: (2025)
TraceVision: Trajectory-Aware Vision-Language Model for Human-Like Spatial Understanding
por: Yang, Fan, et al.
Publicado: (2026)
por: Yang, Fan, et al.
Publicado: (2026)
Vision-TTT: Efficient and Expressive Visual Representation Learning with Test-Time Training
por: Kong, Quan, et al.
Publicado: (2026)
por: Kong, Quan, et al.
Publicado: (2026)
Ejemplares similares
-
WTS: A Pedestrian-Centric Traffic Video Dataset for Fine-grained Spatial-Temporal Understanding
por: Kong, Quan, et al.
Publicado: (2024) -
CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering
por: Zhang, Mingfang, et al.
Publicado: (2026) -
E-VLC: A Real-World Dataset for Event-based Visible Light Communication And Localization
por: Shiba, Shintaro, et al.
Publicado: (2025) -
GA3CE: Unconstrained 3D Gaze Estimation with Gaze-Aware 3D Context Encoding
por: Kawana, Yuki, et al.
Publicado: (2025) -
Evaluation of Mobile Environment for Vehicular Visible Light Communication Using Multiple LEDs and Event Cameras
por: Soga, Ryota, et al.
Publicado: (2025)