What can Off-the-Shelves Large Multi-Modal Models do for Dynamic Scene Graph Generation?
Fuente:
arXiv
Saved in:
| Main Authors: | Cui, Xuanming, Bhoi, Jaiminkumar Ashokbhai, Peng, Chionh Wei, Kuek, Adriel, Lim, Ser Nam |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AirSketch: Generative Motion to Sketch
by: Lim, Hui Xian Grace, et al.
Published: (2024)
by: Lim, Hui Xian Grace, et al.
Published: (2024)
LASER: A Neuro-Symbolic Framework for Learning Spatial-Temporal Scene Graphs with Weak Supervision
by: Huang, Jiani, et al.
Published: (2023)
by: Huang, Jiani, et al.
Published: (2023)
Scene Co-pilot: Procedural Text to Video Generation with Human in the Loop
by: Qian, Zhaofang, et al.
Published: (2024)
by: Qian, Zhaofang, et al.
Published: (2024)
Off the Shelves
by: McCrea, Bridget
Published: (2013)
by: McCrea, Bridget
Published: (2013)
Towards Chunk-Wise Generation for Long Videos
by: Zhang, Siyang, et al.
Published: (2024)
by: Zhang, Siyang, et al.
Published: (2024)
Booktalking Them Off the Shelves.
by: Rochman, Hazel
Published: (1984)
by: Rochman, Hazel
Published: (1984)
ESCA: Contextualizing Embodied Agents via Scene-Graph Generation
by: Huang, Jiani, et al.
Published: (2025)
by: Huang, Jiani, et al.
Published: (2025)
Visual Delta Generator with Large Multi-modal Models for Semi-supervised Composed Image Retrieval
by: Jang, Young Kyun, et al.
Published: (2024)
by: Jang, Young Kyun, et al.
Published: (2024)
VideoMerge: Towards Training-free Long Video Generation
by: Zhang, Siyang, et al.
Published: (2025)
by: Zhang, Siyang, et al.
Published: (2025)
Generalization or Memorization: Dynamic Decoding for Mode Steering
by: Zhang, Xuanming
Published: (2025)
by: Zhang, Xuanming
Published: (2025)
Beyond Generation: Unlocking Universal Editing via Self-Supervised Fine-Tuning
by: Chen, Harold Haodong, et al.
Published: (2024)
by: Chen, Harold Haodong, et al.
Published: (2024)
An Ethical Literary Criticism of Han Suyin’s Autobiography
by: Kuek, Florence
Published: (2025)
by: Kuek, Florence
Published: (2025)
Delta Activations: A Representation for Finetuned Large Language Models
by: Xu, Zhiqiu, et al.
Published: (2025)
by: Xu, Zhiqiu, et al.
Published: (2025)
Self-Captioning Multimodal Interaction Tuning: Amplifying Exploitable Redundancies for Robust Vision Language Models
by: Ryan, Yuriel, et al.
Published: (2026)
by: Ryan, Yuriel, et al.
Published: (2026)
Mitigating Dialogue Hallucination for Large Vision Language Models via Adversarial Instruction Tuning
by: Park, Dongmin, et al.
Published: (2024)
by: Park, Dongmin, et al.
Published: (2024)
Video Decomposition Prior: A Methodology to Decompose Videos into Layers
by: Shrivastava, Gaurav, et al.
Published: (2024)
by: Shrivastava, Gaurav, et al.
Published: (2024)
Niagara: Normal-Integrated Geometric Affine Field for Scene Reconstruction from a Single View
by: Wu, Xianzu, et al.
Published: (2025)
by: Wu, Xianzu, et al.
Published: (2025)
Users’ quality expectations and their correspondence with the realistic features of translation applications
by: Jing Rou Kuek
Published: (2024)
by: Jing Rou Kuek
Published: (2024)
Library Off-Site Shelving: Guide for High-Density Facilities.
by: Nitecki, Danuta A., Ed., et al.
Published: (2001)
by: Nitecki, Danuta A., Ed., et al.
Published: (2001)
Approaching Clairvoyance: Notes Toward Selection for Off-Site Shelving.
by: Powell, Margaret K.
Published: (1998)
by: Powell, Margaret K.
Published: (1998)
DiReCT: Disentangled Regularization of Contrastive Trajectories for Physics-Refined Video Generation
by: Meyarian, Abolfazl, et al.
Published: (2026)
by: Meyarian, Abolfazl, et al.
Published: (2026)
LARM: Large Auto-Regressive Model for Long-Horizon Embodied Intelligence
by: Li, Zhuoling, et al.
Published: (2024)
by: Li, Zhuoling, et al.
Published: (2024)
Distilling Vision-Language Pretraining for Efficient Cross-Modal Retrieval
by: Jang, Young Kyun, et al.
Published: (2024)
by: Jang, Young Kyun, et al.
Published: (2024)
FSViewFusion: Few-Shots View Generation of Novel Objects
by: Hussain, Rukhshanda, et al.
Published: (2024)
by: Hussain, Rukhshanda, et al.
Published: (2024)
InEx: Hallucination Mitigation via Introspection and Cross-Modal Multi-Agent Collaboration
by: Yang, Zhongyu, et al.
Published: (2025)
by: Yang, Zhongyu, et al.
Published: (2025)
Deep Graph Learning for Industrial Carbon Emission Analysis and Policy Impact
by: Zhang, Xuanming
Published: (2025)
by: Zhang, Xuanming
Published: (2025)
BOOKAGENT: Orchestrating Safety-Aware Visual Narratives via Multi-Agent Cognitive Calibration
by: Gao, Bo, et al.
Published: (2026)
by: Gao, Bo, et al.
Published: (2026)
Metric Compatible Training for Online Backfilling in Large-Scale Retrieval
by: Seo, Seonguk, et al.
Published: (2023)
by: Seo, Seonguk, et al.
Published: (2023)
Multi-Modal 3D Scene Graph Updater for Shared and Dynamic Environments
by: Olivastri, Emilio, et al.
Published: (2024)
by: Olivastri, Emilio, et al.
Published: (2024)
When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding
by: Shu, Yan, et al.
Published: (2025)
by: Shu, Yan, et al.
Published: (2025)
Generation mean analysis of important yield traits in Bitter gourd (Momordica charantia)
by: Bhoi Swamini
Published: (2021)
by: Bhoi Swamini
Published: (2021)
DreamMask: Boosting Open-vocabulary Panoptic Segmentation with Synthetic Data
by: Tu, Yuanpeng, et al.
Published: (2025)
by: Tu, Yuanpeng, et al.
Published: (2025)
Towards Unified 3D Object Detection via Algorithm and Data Unification
by: Li, Zhuoling, et al.
Published: (2024)
by: Li, Zhuoling, et al.
Published: (2024)
Composing Object Relations and Attributes for Image-Text Matching
by: Pham, Khoi, et al.
Published: (2024)
by: Pham, Khoi, et al.
Published: (2024)
Fast Encoding and Decoding for Implicit Video Representation
by: Chen, Hao, et al.
Published: (2024)
by: Chen, Hao, et al.
Published: (2024)
Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation
by: Chen, Harold Haodong, et al.
Published: (2025)
by: Chen, Harold Haodong, et al.
Published: (2025)
Open Shelves/Closed Shelves in Research Libraries
by: Rovelstad, Mathilde V.
Published: (1976)
by: Rovelstad, Mathilde V.
Published: (1976)
BCTR: Bidirectional Conditioning Transformer for Scene Graph Generation
by: Hao, Peng, et al.
Published: (2024)
by: Hao, Peng, et al.
Published: (2024)
Graph-based Integrated Gradients for Explaining Graph Neural Networks
by: Simpson, Lachlan, et al.
Published: (2025)
by: Simpson, Lachlan, et al.
Published: (2025)
OP-LoRA: The Blessing of Dimensionality
by: Teterwak, Piotr, et al.
Published: (2024)
by: Teterwak, Piotr, et al.
Published: (2024)
Similar Items
-
AirSketch: Generative Motion to Sketch
by: Lim, Hui Xian Grace, et al.
Published: (2024) -
LASER: A Neuro-Symbolic Framework for Learning Spatial-Temporal Scene Graphs with Weak Supervision
by: Huang, Jiani, et al.
Published: (2023) -
Scene Co-pilot: Procedural Text to Video Generation with Human in the Loop
by: Qian, Zhaofang, et al.
Published: (2024) -
Off the Shelves
by: McCrea, Bridget
Published: (2013) -
Towards Chunk-Wise Generation for Long Videos
by: Zhang, Siyang, et al.
Published: (2024)