Synthetic Visual Genome 2: Extracting Large-scale Spatio-Temporal Scene Graphs from Videos
Fuente:
arXiv
Salvato in:
| Autori principali: | Gao, Ziqi, Zhang, Jieyu, Ikezogwo, Wisdom Oluchi, Park, Jae Sung, You, Tario G., Ogbu, Daniel, Zheng, Chenhao, Huang, Weikai, Yang, Yinuo, Han, Winson, Kong, Quan, Saini, Rajat, Krishna, Ranjay |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Synthetic Object Compositions for Scalable and Accurate Learning in Detection, Segmentation, and Grounding
di: Huang, Weikai, et al.
Pubblicazione: (2025)
di: Huang, Weikai, et al.
Pubblicazione: (2025)
Generate Any Scene: Scene Graph Driven Data Synthesis for Visual Generation Training
di: Gao, Ziqi, et al.
Pubblicazione: (2024)
di: Gao, Ziqi, et al.
Pubblicazione: (2024)
When Rubrics Fail: Error Enumeration as Reward in Reference-Free RL Post-Training for Virtual Try-On
di: Ikezogwo, Wisdom, et al.
Pubblicazione: (2026)
di: Ikezogwo, Wisdom, et al.
Pubblicazione: (2026)
Quilt-LLaVA: Visual Instruction Tuning by Extracting Localized Narratives from Open-Source Histopathology Videos
di: Seyfioglu, Mehmet Saygin, et al.
Pubblicazione: (2023)
di: Seyfioglu, Mehmet Saygin, et al.
Pubblicazione: (2023)
Synthetic Visual Genome
di: Park, Jae Sung, et al.
Pubblicazione: (2025)
di: Park, Jae Sung, et al.
Pubblicazione: (2025)
One Trajectory, One Token: Grounded Video Tokenization via Panoptic Sub-object Trajectory
di: Zheng, Chenhao, et al.
Pubblicazione: (2025)
di: Zheng, Chenhao, et al.
Pubblicazione: (2025)
Quilt-1M: One Million Image-Text Pairs for Histopathology
di: Ikezogwo, Wisdom Oluchi, et al.
Pubblicazione: (2023)
di: Ikezogwo, Wisdom Oluchi, et al.
Pubblicazione: (2023)
MedicalNarratives: Connecting Medical Vision and Language with Localized Narratives
di: Ikezogwo, Wisdom O., et al.
Pubblicazione: (2025)
di: Ikezogwo, Wisdom O., et al.
Pubblicazione: (2025)
TrajTok: Learning Trajectory Tokens enables better Video Understanding
di: Zheng, Chenhao, et al.
Pubblicazione: (2026)
di: Zheng, Chenhao, et al.
Pubblicazione: (2026)
Iterated Learning Improves Compositionality in Large Vision-Language Models
di: Zheng, Chenhao, et al.
Pubblicazione: (2024)
di: Zheng, Chenhao, et al.
Pubblicazione: (2024)
m&m's: A Benchmark to Evaluate Tool-Use for multi-step multi-modal Tasks
di: Ma, Zixian, et al.
Pubblicazione: (2024)
di: Ma, Zixian, et al.
Pubblicazione: (2024)
MedBLINK: Probing Basic Perception in Multimodal Language Models for Medicine
di: Bigverdi, Mahtab, et al.
Pubblicazione: (2025)
di: Bigverdi, Mahtab, et al.
Pubblicazione: (2025)
You Only Judge Once: Multi-response Reward Modeling in a Single Forward Pass
di: Yang, Yinuo, et al.
Pubblicazione: (2026)
di: Yang, Yinuo, et al.
Pubblicazione: (2026)
Unified Spatio-Temporal Token Scoring for Efficient Video VLMs
di: Zhang, Jianrui, et al.
Pubblicazione: (2026)
di: Zhang, Jianrui, et al.
Pubblicazione: (2026)
MolmoPoint: Better Pointing for VLMs with Grounding Tokens
di: Clark, Christopher, et al.
Pubblicazione: (2026)
di: Clark, Christopher, et al.
Pubblicazione: (2026)
Jud, el mediocre
di: Francisco Tario
Pubblicazione: (2011)
di: Francisco Tario
Pubblicazione: (2011)
PathFinder: A Multi-Modal Multi-Agent System for Medical Diagnostic Decision-Making Applied to Histopathology
di: Ghezloo, Fatemeh, et al.
Pubblicazione: (2025)
di: Ghezloo, Fatemeh, et al.
Pubblicazione: (2025)
Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding
di: Clark, Christopher, et al.
Pubblicazione: (2026)
di: Clark, Christopher, et al.
Pubblicazione: (2026)
Visual Representations inside the Language Model
di: Liu, Benlin, et al.
Pubblicazione: (2025)
di: Liu, Benlin, et al.
Pubblicazione: (2025)
GraspMolmo: Generalizable Task-Oriented Grasping via Large-Scale Synthetic Data Generation
di: Deshpande, Abhay, et al.
Pubblicazione: (2025)
di: Deshpande, Abhay, et al.
Pubblicazione: (2025)
CaST-Bench: Benchmarking Causal Chain-Grounded Spatio-Temporal Reasoning for Video Question Answering
di: Zhang, Mingfang, et al.
Pubblicazione: (2026)
di: Zhang, Mingfang, et al.
Pubblicazione: (2026)
SOLID: a Framework of Synergizing Optimization and LLMs for Intelligent Decision-Making
di: Wang, Yinsheng, et al.
Pubblicazione: (2025)
di: Wang, Yinsheng, et al.
Pubblicazione: (2025)
Task Me Anything
di: Zhang, Jieyu, et al.
Pubblicazione: (2024)
di: Zhang, Jieyu, et al.
Pubblicazione: (2024)
Towards Acyclic Preference Evaluation of Language Models via Multiple Evaluators
di: Hu, Zhengyu, et al.
Pubblicazione: (2024)
di: Hu, Zhengyu, et al.
Pubblicazione: (2024)
CREATIVE DESIGN: AN INTEGRAL ASPECT OF INNOVATION IN INDUSTRIAL DESIGN AND TECHNOLOGY
di: Yohanna Ogbu Egiri
Pubblicazione: (2015)
di: Yohanna Ogbu Egiri
Pubblicazione: (2015)
Video-Based Reward Modeling for Computer-Use Agents
di: Song, Linxin, et al.
Pubblicazione: (2026)
di: Song, Linxin, et al.
Pubblicazione: (2026)
Offline Training of Language Model Agents with Functions as Learnable Weights
di: Zhang, Shaokun, et al.
Pubblicazione: (2024)
di: Zhang, Shaokun, et al.
Pubblicazione: (2024)
WildDet3D: Scaling Promptable 3D Detection in the Wild
di: Huang, Weikai, et al.
Pubblicazione: (2026)
di: Huang, Weikai, et al.
Pubblicazione: (2026)
Posterior Augmented Flow Matching
di: Stoica, George, et al.
Pubblicazione: (2026)
di: Stoica, George, et al.
Pubblicazione: (2026)
Videoshop: Localized Semantic Video Editing with Noise-Extrapolated Diffusion Inversion
di: Fan, Xiang, et al.
Pubblicazione: (2024)
di: Fan, Xiang, et al.
Pubblicazione: (2024)
Ablate-to-Validate: Are Vision-Language Models Really Using Continuous Thought Tokens?
di: Zhang, Tianyi, et al.
Pubblicazione: (2026)
di: Zhang, Tianyi, et al.
Pubblicazione: (2026)
ActionAtlas: A VideoQA Benchmark for Domain-specialized Action Recognition
di: Salehi, Mohammadreza, et al.
Pubblicazione: (2024)
di: Salehi, Mohammadreza, et al.
Pubblicazione: (2024)
Unfolding Spatial Cognition: Evaluating Multimodal Models on Visual Simulations
di: Li, Linjie, et al.
Pubblicazione: (2025)
di: Li, Linjie, et al.
Pubblicazione: (2025)
InstAP: Instance-Aware Vision-Language Pre-Train for Spatial-Temporal Understanding
di: Kumar, Ashutosh, et al.
Pubblicazione: (2026)
di: Kumar, Ashutosh, et al.
Pubblicazione: (2026)
Effects of Sesamum indicum L. Seed Extract on Male Reproductive Parameters and In Silico Anti‐Infertility Insights
di: John I. Ogbu, et al.
Pubblicazione: (2026)
di: John I. Ogbu, et al.
Pubblicazione: (2026)
Properties and Examples of $A$-Landweber Exact Spectra
di: Wisdom, Noah
Pubblicazione: (2024)
di: Wisdom, Noah
Pubblicazione: (2024)
The subgroup stratification of Nakaoka spectra
di: Wisdom, Noah
Pubblicazione: (2025)
di: Wisdom, Noah
Pubblicazione: (2025)
A classification of $C_{p^n}$-Tambara fields
di: Wisdom, Noah
Pubblicazione: (2024)
di: Wisdom, Noah
Pubblicazione: (2024)
The $K$-theory of finite Tambara fields: away from $p$
di: Wisdom, Noah
Pubblicazione: (2026)
di: Wisdom, Noah
Pubblicazione: (2026)
Clarification and Coinduction of Tambara Functors
di: Wisdom, Noah
Pubblicazione: (2025)
di: Wisdom, Noah
Pubblicazione: (2025)
Documenti analoghi
-
Synthetic Object Compositions for Scalable and Accurate Learning in Detection, Segmentation, and Grounding
di: Huang, Weikai, et al.
Pubblicazione: (2025) -
Generate Any Scene: Scene Graph Driven Data Synthesis for Visual Generation Training
di: Gao, Ziqi, et al.
Pubblicazione: (2024) -
When Rubrics Fail: Error Enumeration as Reward in Reference-Free RL Post-Training for Virtual Try-On
di: Ikezogwo, Wisdom, et al.
Pubblicazione: (2026) -
Quilt-LLaVA: Visual Instruction Tuning by Extracting Localized Narratives from Open-Source Histopathology Videos
di: Seyfioglu, Mehmet Saygin, et al.
Pubblicazione: (2023) -
Synthetic Visual Genome
di: Park, Jae Sung, et al.
Pubblicazione: (2025)