Synthetic Visual Genome
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Park, Jae Sung, Ma, Zixian, Li, Linjie, Zheng, Chenhao, Hsieh, Cheng-Yu, Lu, Ximing, Chandu, Khyathi, Kong, Quan, Kobori, Norimasa, Farhadi, Ali, Choi, Yejin, Krishna, Ranjay |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Certainly Uncertain: A Benchmark and Metric for Multimodal Epistemic and Aleatoric Awareness
von: Chandu, Khyathi Raghavi, et al.
Veröffentlicht: (2024)
von: Chandu, Khyathi Raghavi, et al.
Veröffentlicht: (2024)
One Trajectory, One Token: Grounded Video Tokenization via Panoptic Sub-object Trajectory
von: Zheng, Chenhao, et al.
Veröffentlicht: (2025)
von: Zheng, Chenhao, et al.
Veröffentlicht: (2025)
E-VLC: A Real-World Dataset for Event-based Visible Light Communication And Localization
von: Shiba, Shintaro, et al.
Veröffentlicht: (2025)
von: Shiba, Shintaro, et al.
Veröffentlicht: (2025)
ActionAtlas: A VideoQA Benchmark for Domain-specialized Action Recognition
von: Salehi, Mohammadreza, et al.
Veröffentlicht: (2024)
von: Salehi, Mohammadreza, et al.
Veröffentlicht: (2024)
Unfolding Spatial Cognition: Evaluating Multimodal Models on Visual Simulations
von: Li, Linjie, et al.
Veröffentlicht: (2025)
von: Li, Linjie, et al.
Veröffentlicht: (2025)
Synthetic Visual Genome 2: Extracting Large-scale Spatio-Temporal Scene Graphs from Videos
von: Gao, Ziqi, et al.
Veröffentlicht: (2026)
von: Gao, Ziqi, et al.
Veröffentlicht: (2026)
GA3CE: Unconstrained 3D Gaze Estimation with Gaze-Aware 3D Context Encoding
von: Kawana, Yuki, et al.
Veröffentlicht: (2025)
von: Kawana, Yuki, et al.
Veröffentlicht: (2025)
VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognition
von: Yadav, Tanush, et al.
Veröffentlicht: (2026)
von: Yadav, Tanush, et al.
Veröffentlicht: (2026)
Synthetic Object Compositions for Scalable and Accurate Learning in Detection, Segmentation, and Grounding
von: Huang, Weikai, et al.
Veröffentlicht: (2025)
von: Huang, Weikai, et al.
Veröffentlicht: (2025)
RESTOR: Knowledge Recovery in Machine Unlearning
von: Rezaei, Keivan, et al.
Veröffentlicht: (2024)
von: Rezaei, Keivan, et al.
Veröffentlicht: (2024)
L3GO: Language Agents with Chain-of-3D-Thoughts for Generating Unconventional Objects
von: Yamada, Yutaro, et al.
Veröffentlicht: (2024)
von: Yamada, Yutaro, et al.
Veröffentlicht: (2024)
Scale Can't Overcome Pragmatics: The Impact of Reporting Bias on Vision-Language Reasoning
von: Kamath, Amita, et al.
Veröffentlicht: (2026)
von: Kamath, Amita, et al.
Veröffentlicht: (2026)
AI as Humanity's Salieri: Quantifying Linguistic Creativity of Language Models via Systematic Attribution of Machine Text against Web Text
von: Lu, Ximing, et al.
Veröffentlicht: (2024)
von: Lu, Ximing, et al.
Veröffentlicht: (2024)
Selective Visual Representations Improve Convergence and Generalization for Embodied AI
von: Eftekhar, Ainaz, et al.
Veröffentlicht: (2023)
von: Eftekhar, Ainaz, et al.
Veröffentlicht: (2023)
Agent Lumos: Unified and Modular Training for Open-Source Language Agents
von: Yin, Da, et al.
Veröffentlicht: (2023)
von: Yin, Da, et al.
Veröffentlicht: (2023)
Selective "Selective Prediction": Reducing Unnecessary Abstention in Vision-Language Reasoning
von: Srinivasan, Tejas, et al.
Veröffentlicht: (2024)
von: Srinivasan, Tejas, et al.
Veröffentlicht: (2024)
ThinkMorph: Emergent Properties in Multimodal Interleaved Chain-of-Thought Reasoning
von: Gu, Jiawei, et al.
Veröffentlicht: (2025)
von: Gu, Jiawei, et al.
Veröffentlicht: (2025)
Evaluation of Mobile Environment for Vehicular Visible Light Communication Using Multiple LEDs and Event Cameras
von: Soga, Ryota, et al.
Veröffentlicht: (2025)
von: Soga, Ryota, et al.
Veröffentlicht: (2025)
Distance Estimation in Outdoor Driving Environments Using Phase-only Correlation Method with Event Cameras
von: Kobayashi, Masataka, et al.
Veröffentlicht: (2025)
von: Kobayashi, Masataka, et al.
Veröffentlicht: (2025)
Iterated Learning Improves Compositionality in Large Vision-Language Models
von: Zheng, Chenhao, et al.
Veröffentlicht: (2024)
von: Zheng, Chenhao, et al.
Veröffentlicht: (2024)
AdaReasoner: Dynamic Tool Orchestration for Iterative Visual Reasoning
von: Song, Mingyang, et al.
Veröffentlicht: (2026)
von: Song, Mingyang, et al.
Veröffentlicht: (2026)
Contrastive Flow Matching
von: Stoica, George, et al.
Veröffentlicht: (2025)
von: Stoica, George, et al.
Veröffentlicht: (2025)
MutaGReP: Execution-Free Repository-Grounded Plan Search for Code-Use
von: Khan, Zaid, et al.
Veröffentlicht: (2025)
von: Khan, Zaid, et al.
Veröffentlicht: (2025)
Superposed Decoding: Multiple Generations from a Single Autoregressive Inference Pass
von: Shen, Ethan, et al.
Veröffentlicht: (2024)
von: Shen, Ethan, et al.
Veröffentlicht: (2024)
InstAP: Instance-Aware Vision-Language Pre-Train for Spatial-Temporal Understanding
von: Kumar, Ashutosh, et al.
Veröffentlicht: (2026)
von: Kumar, Ashutosh, et al.
Veröffentlicht: (2026)
WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
von: Lin, Bill Yuchen, et al.
Veröffentlicht: (2024)
von: Lin, Bill Yuchen, et al.
Veröffentlicht: (2024)
m&m's: A Benchmark to Evaluate Tool-Use for multi-step multi-modal Tasks
von: Ma, Zixian, et al.
Veröffentlicht: (2024)
von: Ma, Zixian, et al.
Veröffentlicht: (2024)
You Only Judge Once: Multi-response Reward Modeling in a Single Forward Pass
von: Yang, Yinuo, et al.
Veröffentlicht: (2026)
von: Yang, Yinuo, et al.
Veröffentlicht: (2026)
Task Me Anything
von: Zhang, Jieyu, et al.
Veröffentlicht: (2024)
von: Zhang, Jieyu, et al.
Veröffentlicht: (2024)
Convergent Functions, Divergent Forms
von: Jeon, Hyeonseong, et al.
Veröffentlicht: (2025)
von: Jeon, Hyeonseong, et al.
Veröffentlicht: (2025)
Lookback Lens: Detecting and Mitigating Contextual Hallucinations in Large Language Models Using Only Attention Maps
von: Chuang, Yung-Sung, et al.
Veröffentlicht: (2024)
von: Chuang, Yung-Sung, et al.
Veröffentlicht: (2024)
The Hard Positive Truth about Vision-Language Compositionality
von: Kamath, Amita, et al.
Veröffentlicht: (2024)
von: Kamath, Amita, et al.
Veröffentlicht: (2024)
Deal, or no deal (or who knows)? Forecasting Uncertainty in Conversations using Large Language Models
von: Sicilia, Anthony, et al.
Veröffentlicht: (2024)
von: Sicilia, Anthony, et al.
Veröffentlicht: (2024)
TrajTok: Learning Trajectory Tokens enables better Video Understanding
von: Zheng, Chenhao, et al.
Veröffentlicht: (2026)
von: Zheng, Chenhao, et al.
Veröffentlicht: (2026)
The Unmet Promise of Synthetic Training Images: Using Retrieved Real Images Performs Better
von: Geng, Scott, et al.
Veröffentlicht: (2024)
von: Geng, Scott, et al.
Veröffentlicht: (2024)
MolmoPoint: Better Pointing for VLMs with Grounding Tokens
von: Clark, Christopher, et al.
Veröffentlicht: (2026)
von: Clark, Christopher, et al.
Veröffentlicht: (2026)
Reinforced Visual Perception with Tools
von: Zhou, Zetong, et al.
Veröffentlicht: (2025)
von: Zhou, Zetong, et al.
Veröffentlicht: (2025)
MolmoWeb: Open Visual Web Agent and Open Data for the Open Web
von: Gupta, Tanmay, et al.
Veröffentlicht: (2026)
von: Gupta, Tanmay, et al.
Veröffentlicht: (2026)
Perception Tokens Enhance Visual Reasoning in Multimodal Language Models
von: Bigverdi, Mahtab, et al.
Veröffentlicht: (2024)
von: Bigverdi, Mahtab, et al.
Veröffentlicht: (2024)
Rethinking Human Preference Evaluation of LLM Rationales
von: Li, Ziang, et al.
Veröffentlicht: (2025)
von: Li, Ziang, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Certainly Uncertain: A Benchmark and Metric for Multimodal Epistemic and Aleatoric Awareness
von: Chandu, Khyathi Raghavi, et al.
Veröffentlicht: (2024) -
One Trajectory, One Token: Grounded Video Tokenization via Panoptic Sub-object Trajectory
von: Zheng, Chenhao, et al.
Veröffentlicht: (2025) -
E-VLC: A Real-World Dataset for Event-based Visible Light Communication And Localization
von: Shiba, Shintaro, et al.
Veröffentlicht: (2025) -
ActionAtlas: A VideoQA Benchmark for Domain-specialized Action Recognition
von: Salehi, Mohammadreza, et al.
Veröffentlicht: (2024) -
Unfolding Spatial Cognition: Evaluating Multimodal Models on Visual Simulations
von: Li, Linjie, et al.
Veröffentlicht: (2025)