InstAP: Instance-Aware Vision-Language Pre-Train for Spatial-Temporal Understanding
Fuente:
arXiv
Salvato in:
| Autori principali: | Kumar, Ashutosh, Saini, Rajat, Pan, Jingjing, Erdogan, Mustafa, Zhang, Mingfang, Dem, Betty Le, Kobori, Norimasa, Kong, Quan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
WTS: A Pedestrian-Centric Traffic Video Dataset for Fine-grained Spatial-Temporal Understanding
di: Kong, Quan, et al.
Pubblicazione: (2024)
di: Kong, Quan, et al.
Pubblicazione: (2024)
CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering
di: Zhang, Mingfang, et al.
Pubblicazione: (2026)
di: Zhang, Mingfang, et al.
Pubblicazione: (2026)
E-VLC: A Real-World Dataset for Event-based Visible Light Communication And Localization
di: Shiba, Shintaro, et al.
Pubblicazione: (2025)
di: Shiba, Shintaro, et al.
Pubblicazione: (2025)
GA3CE: Unconstrained 3D Gaze Estimation with Gaze-Aware 3D Context Encoding
di: Kawana, Yuki, et al.
Pubblicazione: (2025)
di: Kawana, Yuki, et al.
Pubblicazione: (2025)
Evaluation of Mobile Environment for Vehicular Visible Light Communication Using Multiple LEDs and Event Cameras
di: Soga, Ryota, et al.
Pubblicazione: (2025)
di: Soga, Ryota, et al.
Pubblicazione: (2025)
Distance Estimation in Outdoor Driving Environments Using Phase-only Correlation Method with Event Cameras
di: Kobayashi, Masataka, et al.
Pubblicazione: (2025)
di: Kobayashi, Masataka, et al.
Pubblicazione: (2025)
Inst3D-LMM: Instance-Aware 3D Scene Understanding with Multi-modal Instruction Tuning
di: Yu, Hanxun, et al.
Pubblicazione: (2025)
di: Yu, Hanxun, et al.
Pubblicazione: (2025)
InstDrive: Instance-Aware 3D Gaussian Splatting for Driving Scenes
di: Liu, Hongyuan, et al.
Pubblicazione: (2025)
di: Liu, Hongyuan, et al.
Pubblicazione: (2025)
One Trajectory, One Token: Grounded Video Tokenization via Panoptic Sub-object Trajectory
di: Zheng, Chenhao, et al.
Pubblicazione: (2025)
di: Zheng, Chenhao, et al.
Pubblicazione: (2025)
UDA4Inst: Unsupervised Domain Adaptation for Instance Segmentation
di: Guo, Yachan, et al.
Pubblicazione: (2024)
di: Guo, Yachan, et al.
Pubblicazione: (2024)
Books in the Field: Physical Anthropology
di: Polacheck, Dem
Pubblicazione: (1969)
di: Polacheck, Dem
Pubblicazione: (1969)
Trainable Quantum Neural Network for Multiclass Image Classification with the Power of Pre-trained Tree Tensor Networks
di: Murota, Keisuke, et al.
Pubblicazione: (2025)
di: Murota, Keisuke, et al.
Pubblicazione: (2025)
AutoInst: Automatic Instance-Based Segmentation of LiDAR 3D Scans
di: Perauer, Cedric, et al.
Pubblicazione: (2024)
di: Perauer, Cedric, et al.
Pubblicazione: (2024)
FastInstShadow: A Simple Query-Based Model for Instance Shadow Detection
di: Inoue, Takeru, et al.
Pubblicazione: (2025)
di: Inoue, Takeru, et al.
Pubblicazione: (2025)
Enhanced BoxInst for Weakly Supervised Liver Tumor Instance Segmentation in CT Images
di: Shanshan Li, et al.
Pubblicazione: (2025)
di: Shanshan Li, et al.
Pubblicazione: (2025)
Entre littérature et anthropologie : la "raison littéraire" d’Éric Chauvier
di: Laurent DemAnze
Pubblicazione: (2022)
di: Laurent DemAnze
Pubblicazione: (2022)
Leveraging GNSS and Onboard Visual Data from Consumer Vehicles for Robust Road Network Estimation
di: Opra, Balázs, et al.
Pubblicazione: (2024)
di: Opra, Balázs, et al.
Pubblicazione: (2024)
Adapting Pre-Trained Vision Models for Novel Instance Detection and Segmentation
di: Lu, Yangxiao, et al.
Pubblicazione: (2024)
di: Lu, Yangxiao, et al.
Pubblicazione: (2024)
A Temporal Modeling Framework for Video Pre-Training on Video Instance Segmentation
di: Zhong, Qing, et al.
Pubblicazione: (2025)
di: Zhong, Qing, et al.
Pubblicazione: (2025)
InstCache: A Predictive Cache for LLM Serving
di: Zou, Longwei, et al.
Pubblicazione: (2024)
di: Zou, Longwei, et al.
Pubblicazione: (2024)
BLO-Inst: Bi-Level Optimization Based Alignment of YOLO and SAM for Robust Instance Segmentation
di: Zhang, Li, et al.
Pubblicazione: (2026)
di: Zhang, Li, et al.
Pubblicazione: (2026)
Reprojection Errors as Prompts for Efficient Scene Coordinate Regression
di: Liu, Ting-Ru, et al.
Pubblicazione: (2024)
di: Liu, Ting-Ru, et al.
Pubblicazione: (2024)
An Embodied AR Navigation Agent: Integrating BIM with Retrieval-Augmented Generation for Language Guidance
di: Yang, Hsuan-Kung, et al.
Pubblicazione: (2025)
di: Yang, Hsuan-Kung, et al.
Pubblicazione: (2025)
Inst4DGS: Instance-Decomposed 4D Gaussian Splatting with Multi-Video Label Permutation Learning
di: Lee, Yonghan, et al.
Pubblicazione: (2026)
di: Lee, Yonghan, et al.
Pubblicazione: (2026)
Decoupled Spatial and Temporal Processing for Resource Efficient Multichannel Speech Enhancement
di: Pandey, Ashutosh, et al.
Pubblicazione: (2024)
di: Pandey, Ashutosh, et al.
Pubblicazione: (2024)
Impact of Scalar NSI on Spatial and Temporal Correlations in Neutrino Oscillations
di: Yadav, Bhavna, et al.
Pubblicazione: (2024)
di: Yadav, Bhavna, et al.
Pubblicazione: (2024)
Spatially Constrained Transformer with Efficient Global Relation Modelling for Spatio-Temporal Prediction
di: Sao, Ashutosh, et al.
Pubblicazione: (2024)
di: Sao, Ashutosh, et al.
Pubblicazione: (2024)
Preschool Students' Understanding of a Geometric Shape, the Square
di: Erdogan Halat
Pubblicazione: (2016)
di: Erdogan Halat
Pubblicazione: (2016)
Score Broadcast and Decorrelation: A General Framework for Broadcast-Based Credit Assignment
di: Uzun, Mustafa, et al.
Pubblicazione: (2026)
di: Uzun, Mustafa, et al.
Pubblicazione: (2026)
O INDIVÍDUO MODERNO E O ROMANCE: UM POSSÍVEL DIÁLOGO ENTRE O JOVEM LUKÁCS E MARTHE ROBERT
di: Eduardo Toshio Kobori
Pubblicazione: (2015)
di: Eduardo Toshio Kobori
Pubblicazione: (2015)
Synthetic Visual Genome
di: Park, Jae Sung, et al.
Pubblicazione: (2025)
di: Park, Jae Sung, et al.
Pubblicazione: (2025)
3DVLA: Enhancing Vision-Language-Action Models via 3D Spatial and Instance Understanding
di: Xia, Zhongyu, et al.
Pubblicazione: (2026)
di: Xia, Zhongyu, et al.
Pubblicazione: (2026)
Instance-Aware Group Quantization for Vision Transformers
di: Moon, Jaehyeon, et al.
Pubblicazione: (2024)
di: Moon, Jaehyeon, et al.
Pubblicazione: (2024)
LeafInst - Unified Instance Segmentation Network for Fine-Grained Forestry Leaf Phenotype Analysis: A New UAV based Benchmark
di: Luo, Taige, et al.
Pubblicazione: (2026)
di: Luo, Taige, et al.
Pubblicazione: (2026)
Leveraging RGB Images for Pre-Training of Event-Based Hand Pose Estimation
di: Liu, Ruicong, et al.
Pubblicazione: (2025)
di: Liu, Ruicong, et al.
Pubblicazione: (2025)
Unsupervised Pre-Training for 3D Leaf Instance Segmentation
di: Roggiolani, Gianmarco, et al.
Pubblicazione: (2024)
di: Roggiolani, Gianmarco, et al.
Pubblicazione: (2024)
Evaluation of bioactive potential of Edirne beyaz peyniri (Edirne white cheese) produced from different milk types during maturation time
di: Fatmagül Halici Demİr, et al.
Pubblicazione: (2025)
di: Fatmagül Halici Demİr, et al.
Pubblicazione: (2025)
Spatial-Temporal Pre-Training for Embryo Viability Prediction Using Time-Lapse Videos
di: Shi, Zhiyi, et al.
Pubblicazione: (2025)
di: Shi, Zhiyi, et al.
Pubblicazione: (2025)
TraceVision: Trajectory-Aware Vision-Language Model for Human-Like Spatial Understanding
di: Yang, Fan, et al.
Pubblicazione: (2026)
di: Yang, Fan, et al.
Pubblicazione: (2026)
Vision-TTT: Efficient and Expressive Visual Representation Learning with Test-Time Training
di: Kong, Quan, et al.
Pubblicazione: (2026)
di: Kong, Quan, et al.
Pubblicazione: (2026)
Documenti analoghi
-
WTS: A Pedestrian-Centric Traffic Video Dataset for Fine-grained Spatial-Temporal Understanding
di: Kong, Quan, et al.
Pubblicazione: (2024) -
CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering
di: Zhang, Mingfang, et al.
Pubblicazione: (2026) -
E-VLC: A Real-World Dataset for Event-based Visible Light Communication And Localization
di: Shiba, Shintaro, et al.
Pubblicazione: (2025) -
GA3CE: Unconstrained 3D Gaze Estimation with Gaze-Aware 3D Context Encoding
di: Kawana, Yuki, et al.
Pubblicazione: (2025) -
Evaluation of Mobile Environment for Vehicular Visible Light Communication Using Multiple LEDs and Event Cameras
di: Soga, Ryota, et al.
Pubblicazione: (2025)