GTASA: Ground Truth Annotations for Spatiotemporal Analysis, Evaluation and Training of Video Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Cudlenco, Nicolae, Masala, Mihai, Leordeanu, Marius |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Agentic Video Generation: From Text to Executable Event Graphs via Tool-Constrained LLM Planning
von: Cudlenco, Nicolae, et al.
Veröffentlicht: (2026)
von: Cudlenco, Nicolae, et al.
Veröffentlicht: (2026)
Towards Zero-Shot & Explainable Video Description by Reasoning over Graphs of Events in Space and Time
von: Masala, Mihai, et al.
Veröffentlicht: (2025)
von: Masala, Mihai, et al.
Veröffentlicht: (2025)
From Vision To Language through Graph of Events in Space and Time: An Explainable Self-supervised Approach
von: Masala, Mihai, et al.
Veröffentlicht: (2025)
von: Masala, Mihai, et al.
Veröffentlicht: (2025)
Closer to Ground Truth: Realistic Shape and Appearance Labeled Data Generation for Unsupervised Underwater Image Segmentation
von: Jelea, Andrei, et al.
Veröffentlicht: (2025)
von: Jelea, Andrei, et al.
Veröffentlicht: (2025)
Probabilistic Hyper-Graphs using Multiple Randomly Masked Autoencoders for Semi-supervised Multi-modal Multi-task Learning
von: Mihai-Cristian, Pîrvu, et al.
Veröffentlicht: (2025)
von: Mihai-Cristian, Pîrvu, et al.
Veröffentlicht: (2025)
Multi-modal video data-pipelines for machine learning with minimal human supervision
von: Pîrvu, Mihai-Cristian, et al.
Veröffentlicht: (2025)
von: Pîrvu, Mihai-Cristian, et al.
Veröffentlicht: (2025)
With Great Context Comes Great Prediction Power: Classifying Objects via Geo-Semantic Scene Graphs
von: Constantinescu, Ciprian, et al.
Veröffentlicht: (2025)
von: Constantinescu, Ciprian, et al.
Veröffentlicht: (2025)
Multiple Random Masking Autoencoder Ensembles for Robust Multimodal Semi-supervised Learning
von: Todoran, Alexandru-Raul, et al.
Veröffentlicht: (2024)
von: Todoran, Alexandru-Raul, et al.
Veröffentlicht: (2024)
Learning on the Fly: Replay-Based Continual Object Perception for Indoor Drones
von: Nae, Sebastian-Ion, et al.
Veröffentlicht: (2026)
von: Nae, Sebastian-Ion, et al.
Veröffentlicht: (2026)
Efficient Self-Supervised Neuro-Analytic Visual Servoing for Real-time Quadrotor Control
von: Mocanu, Sebastian, et al.
Veröffentlicht: (2025)
von: Mocanu, Sebastian, et al.
Veröffentlicht: (2025)
Self-Supervised Learning to Fly using Efficient Semantic Segmentation and Metric Depth Estimation for Low-Cost Autonomous UAVs
von: Mocanu, Sebastian, et al.
Veröffentlicht: (2025)
von: Mocanu, Sebastian, et al.
Veröffentlicht: (2025)
Learning from Random Subspace Exploration: Generalized Test-Time Augmentation with Self-supervised Distillation
von: Jelea, Andrei, et al.
Veröffentlicht: (2025)
von: Jelea, Andrei, et al.
Veröffentlicht: (2025)
Inside Knowledge: Graph-based Path Generation with Explainable Data Augmentation and Curriculum Learning for Visual Indoor Navigation
von: Airinei, Daniel, et al.
Veröffentlicht: (2025)
von: Airinei, Daniel, et al.
Veröffentlicht: (2025)
Maia: A Real-time Non-Verbal Chat for Human-AI Interaction
von: Costea, Dragos, et al.
Veröffentlicht: (2024)
von: Costea, Dragos, et al.
Veröffentlicht: (2024)
A self-supervised cyclic neural-analytic approach for novel view synthesis and 3D reconstruction
von: Costea, Dragos, et al.
Veröffentlicht: (2025)
von: Costea, Dragos, et al.
Veröffentlicht: (2025)
3D Ground Truth Reconstruction from Multi-Camera Annotations Using UKF
von: Van Ma, Linh, et al.
Veröffentlicht: (2025)
von: Van Ma, Linh, et al.
Veröffentlicht: (2025)
Non-verbal Real-time Human-AI Interaction in Constrained Robotic Environments
von: Costea, Dragos, et al.
Veröffentlicht: (2026)
von: Costea, Dragos, et al.
Veröffentlicht: (2026)
Fine-grained Spatiotemporal Grounding on Egocentric Videos
von: Liang, Shuo, et al.
Veröffentlicht: (2025)
von: Liang, Shuo, et al.
Veröffentlicht: (2025)
Iterative Explainability for Weakly Supervised Segmentation in Medical PE Detection
von: Condrea, Florin, et al.
Veröffentlicht: (2024)
von: Condrea, Florin, et al.
Veröffentlicht: (2024)
"Înţelegi Româneşte?'' A Recipe for Romanian Vision-Language Models
von: Masala, Mihai, et al.
Veröffentlicht: (2026)
von: Masala, Mihai, et al.
Veröffentlicht: (2026)
An Evaluation of Hybrid Annotation Workflows on High-Ambiguity Spatiotemporal Video Footage
von: Gutiérrez, Juan, et al.
Veröffentlicht: (2025)
von: Gutiérrez, Juan, et al.
Veröffentlicht: (2025)
Amodal Ground Truth and Completion in the Wild
von: Zhan, Guanqi, et al.
Veröffentlicht: (2023)
von: Zhan, Guanqi, et al.
Veröffentlicht: (2023)
Minority Reports: Balancing Cost and Quality in Ground Truth Data Annotation
von: Liao, Hsuan Wei, et al.
Veröffentlicht: (2025)
von: Liao, Hsuan Wei, et al.
Veröffentlicht: (2025)
IKIWISI: An Interactive Visual Pattern Generator for Evaluating the Reliability of Vision-Language Models Without Ground Truth
von: Islam, Md Touhidul, et al.
Veröffentlicht: (2025)
von: Islam, Md Touhidul, et al.
Veröffentlicht: (2025)
A Human-Annotated Video Dataset for Training and Evaluation of 360-Degree Video Summarization Methods
von: Kontostathis, Ioannis, et al.
Veröffentlicht: (2024)
von: Kontostathis, Ioannis, et al.
Veröffentlicht: (2024)
Look Ma, No Ground Truth! Ground-Truth-Free Tuning of Structure from Motion and Visual SLAM
von: Fontan, Alejandro, et al.
Veröffentlicht: (2024)
von: Fontan, Alejandro, et al.
Veröffentlicht: (2024)
Training-free Online Video Step Grounding
von: Zanella, Luca, et al.
Veröffentlicht: (2025)
von: Zanella, Luca, et al.
Veröffentlicht: (2025)
Beyond the Ground Truth: Enhanced Supervision for Image Restoration
von: Ryou, Donghun, et al.
Veröffentlicht: (2025)
von: Ryou, Donghun, et al.
Veröffentlicht: (2025)
MIDAS: Modeling Ground-Truth Distributions with Dark Knowledge for Domain Generalized Stereo Matching
von: Xu, Peng, et al.
Veröffentlicht: (2025)
von: Xu, Peng, et al.
Veröffentlicht: (2025)
Not Just Streaks: Towards Ground Truth for Single Image Deraining
von: Ba, Yunhao, et al.
Veröffentlicht: (2022)
von: Ba, Yunhao, et al.
Veröffentlicht: (2022)
LLM4VG: Large Language Models Evaluation for Video Grounding
von: Feng, Wei, et al.
Veröffentlicht: (2023)
von: Feng, Wei, et al.
Veröffentlicht: (2023)
Training-free Video Temporal Grounding using Large-scale Pre-trained Models
von: Zheng, Minghang, et al.
Veröffentlicht: (2024)
von: Zheng, Minghang, et al.
Veröffentlicht: (2024)
VideoTG-R1: Boosting Video Temporal Grounding via Curriculum Reinforcement Learning on Reflected Boundary Annotations
von: Dong, Lu, et al.
Veröffentlicht: (2025)
von: Dong, Lu, et al.
Veröffentlicht: (2025)
Perspective-Equivariant Fine-tuning for Multispectral Demosaicing without Ground Truth
von: Wang, Andrew, et al.
Veröffentlicht: (2026)
von: Wang, Andrew, et al.
Veröffentlicht: (2026)
Learning Cross-view Visual Geo-localization without Ground Truth
von: Li, Haoyuan, et al.
Veröffentlicht: (2024)
von: Li, Haoyuan, et al.
Veröffentlicht: (2024)
Marine Snow Removal Using Internally Generated Pseudo Ground Truth
von: Malyugina, Alexandra, et al.
Veröffentlicht: (2025)
von: Malyugina, Alexandra, et al.
Veröffentlicht: (2025)
CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models
von: Huang, Hsiang-Wei, et al.
Veröffentlicht: (2026)
von: Huang, Hsiang-Wei, et al.
Veröffentlicht: (2026)
UniCtrl: Improving the Spatiotemporal Consistency of Text-to-Video Diffusion Models via Training-Free Unified Attention Control
von: Xia, Tian, et al.
Veröffentlicht: (2024)
von: Xia, Tian, et al.
Veröffentlicht: (2024)
Scene-R1: Video-Grounded Large Language Models for 3D Scene Reasoning without 3D Annotations
von: Yuan, Zhihao, et al.
Veröffentlicht: (2025)
von: Yuan, Zhihao, et al.
Veröffentlicht: (2025)
Reducing Annotation Burden: Exploiting Image Knowledge for Few-Shot Medical Video Object Segmentation via Spatiotemporal Consistency Relearning
von: Zheng, Zixuan, et al.
Veröffentlicht: (2025)
von: Zheng, Zixuan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Agentic Video Generation: From Text to Executable Event Graphs via Tool-Constrained LLM Planning
von: Cudlenco, Nicolae, et al.
Veröffentlicht: (2026) -
Towards Zero-Shot & Explainable Video Description by Reasoning over Graphs of Events in Space and Time
von: Masala, Mihai, et al.
Veröffentlicht: (2025) -
From Vision To Language through Graph of Events in Space and Time: An Explainable Self-supervised Approach
von: Masala, Mihai, et al.
Veröffentlicht: (2025) -
Closer to Ground Truth: Realistic Shape and Appearance Labeled Data Generation for Unsupervised Underwater Image Segmentation
von: Jelea, Andrei, et al.
Veröffentlicht: (2025) -
Probabilistic Hyper-Graphs using Multiple Randomly Masked Autoencoders for Semi-supervised Multi-modal Multi-task Learning
von: Mihai-Cristian, Pîrvu, et al.
Veröffentlicht: (2025)