Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhou, Honglu, Peng, Xiangyu, Kendre, Shrikant, Ryoo, Michael S., Savarese, Silvio, Xiong, Caiming, Niebles, Juan Carlos |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs
di: Ryoo, Michael S., et al.
Pubblicazione: (2024)
di: Ryoo, Michael S., et al.
Pubblicazione: (2024)
SMILE: A Composite Lexical-Semantic Metric for Question-Answering Evaluation
di: Kendre, Shrikant, et al.
Pubblicazione: (2025)
di: Kendre, Shrikant, et al.
Pubblicazione: (2025)
HIVE: Harnessing Human Feedback for Instructional Visual Editing
di: Zhang, Shu, et al.
Pubblicazione: (2023)
di: Zhang, Shu, et al.
Pubblicazione: (2023)
Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding
di: Wang, Ziyang, et al.
Pubblicazione: (2025)
di: Wang, Ziyang, et al.
Pubblicazione: (2025)
ReSpark: Leveraging Previous Data Reports as References to Generate New Reports with LLMs
di: Tian, Yuan, et al.
Pubblicazione: (2025)
di: Tian, Yuan, et al.
Pubblicazione: (2025)
Jupybara: Operationalizing a Design Space for Actionable Data Analysis and Storytelling with LLMs
di: Wang, Huichen Will, et al.
Pubblicazione: (2025)
di: Wang, Huichen Will, et al.
Pubblicazione: (2025)
A Survey of LLM Alignment: Instruction Understanding, Intention Reasoning, and Reliable Generation
di: Chang, Zongyu, et al.
Pubblicazione: (2025)
di: Chang, Zongyu, et al.
Pubblicazione: (2025)
IP-Dialog: Evaluating Implicit Personalization in Dialogue Systems with Synthetic Data
di: Peng, Bo, et al.
Pubblicazione: (2025)
di: Peng, Bo, et al.
Pubblicazione: (2025)
Diverse and Fine-Grained Instruction-Following Ability Exploration with Synthetic Data
di: Gu, Zihui, et al.
Pubblicazione: (2024)
di: Gu, Zihui, et al.
Pubblicazione: (2024)
LLMs as Academic Reading Companions: Extending HCI Through Synthetic Personae
di: Chen, Celia, et al.
Pubblicazione: (2024)
di: Chen, Celia, et al.
Pubblicazione: (2024)
When LLMs Enter Everyday Feminism on Chinese Social Media: Opportunities and Risks for Women's Empowerment
di: Zhang, Runhua, et al.
Pubblicazione: (2026)
di: Zhang, Runhua, et al.
Pubblicazione: (2026)
Gig2Gether: Data-sharing to Empower, Unify and Demystify Gig Work
di: Hsieh, Jane, et al.
Pubblicazione: (2025)
di: Hsieh, Jane, et al.
Pubblicazione: (2025)
DesignFromX: Empowering Consumer-Driven Design Space Exploration through Feature Composition of Referenced Products
di: Duan, Runlin, et al.
Pubblicazione: (2025)
di: Duan, Runlin, et al.
Pubblicazione: (2025)
Substantial, Decomposable, and Invisible: Visual Context Misalignment in Instructional Videos for Physical Tasks
di: Li, Yayuan, et al.
Pubblicazione: (2026)
di: Li, Yayuan, et al.
Pubblicazione: (2026)
Empowering Vocabulary Learning Through Teaching AI: Using LLMs as a Student to Perform Learning by Teaching in Vocabulary Acquisition
di: Uchida, Tokio, et al.
Pubblicazione: (2026)
di: Uchida, Tokio, et al.
Pubblicazione: (2026)
Exploring the Role of Interaction Data to Empower End-User Decision-Making In UI Personalization
di: Alves, Sérgio, et al.
Pubblicazione: (2026)
di: Alves, Sérgio, et al.
Pubblicazione: (2026)
A Dynamic and High-Precision Method for Scenario-Based HRA Synthetic Data Collection in Multi-Agent Collaborative Environments Driven by LLMs
di: Xiao, Xingyu, et al.
Pubblicazione: (2025)
di: Xiao, Xingyu, et al.
Pubblicazione: (2025)
MIND: Empowering Mental Health Clinicians with Multimodal Data Insights through a Narrative Dashboard
di: Zou, Ruishi, et al.
Pubblicazione: (2026)
di: Zou, Ruishi, et al.
Pubblicazione: (2026)
Video2MR: Automatically Generating Mixed Reality 3D Instructions by Augmenting Extracted Motion from 2D Videos
di: Ihara, Keiichi, et al.
Pubblicazione: (2024)
di: Ihara, Keiichi, et al.
Pubblicazione: (2024)
Examining the Expanding Role of Synthetic Data Throughout the AI Development Pipeline
di: Kapania, Shivani, et al.
Pubblicazione: (2025)
di: Kapania, Shivani, et al.
Pubblicazione: (2025)
Understanding the Impact of Referent Design on Scale Perception in Immersive Data Visualization
di: Hou, Yihan, et al.
Pubblicazione: (2024)
di: Hou, Yihan, et al.
Pubblicazione: (2024)
Data Playwright: Authoring Data Videos with Annotated Narration
di: Shen, Leixian, et al.
Pubblicazione: (2024)
di: Shen, Leixian, et al.
Pubblicazione: (2024)
XR and Hybrid Data Visualization Spaces for Enhanced Data Analytics
di: Lombeyda, Santiago, et al.
Pubblicazione: (2026)
di: Lombeyda, Santiago, et al.
Pubblicazione: (2026)
Generative Lecture: Making Lecture Videos Interactive with LLMs and AI Clone Instructors
di: Jo, Hye-Young, et al.
Pubblicazione: (2025)
di: Jo, Hye-Young, et al.
Pubblicazione: (2025)
Volume-Based Space-Time Cube for Large-Scale Continuous Spatial Time Series
di: Deng, Zikun, et al.
Pubblicazione: (2025)
di: Deng, Zikun, et al.
Pubblicazione: (2025)
Enhancing EEG Signal-Based Emotion Recognition with Synthetic Data: Diffusion Model Approach
di: Siddhad, Gourav, et al.
Pubblicazione: (2024)
di: Siddhad, Gourav, et al.
Pubblicazione: (2024)
Leveraging LLMs for Persona-Based Visualization of Election Data
di: Panda, Swaroop, et al.
Pubblicazione: (2025)
di: Panda, Swaroop, et al.
Pubblicazione: (2025)
Offscript: Automated Auditing of Instruction Adherence in LLMs
di: Clark, Nicholas, et al.
Pubblicazione: (2025)
di: Clark, Nicholas, et al.
Pubblicazione: (2025)
VisTR: Visualizations as Representations for Time-series Table Reasoning
di: Hao, Jianing, et al.
Pubblicazione: (2024)
di: Hao, Jianing, et al.
Pubblicazione: (2024)
Enhancing Computational Notebooks with Code+Data Space Versioning
di: Fang, Hanxi, et al.
Pubblicazione: (2025)
di: Fang, Hanxi, et al.
Pubblicazione: (2025)
Nonvisual Support for Understanding and Reasoning about Data Structures
di: Wimer, Brianna L., et al.
Pubblicazione: (2026)
di: Wimer, Brianna L., et al.
Pubblicazione: (2026)
AI-Empowered Human Research Integrating Brain Science and Social Sciences Insights
di: Xiong, Feng, et al.
Pubblicazione: (2024)
di: Xiong, Feng, et al.
Pubblicazione: (2024)
Aesthetics of Connectivity: Envisioning Empowerment Through Smart Clothing
di: Mulundule, Yannick Kibolwe, et al.
Pubblicazione: (2025)
di: Mulundule, Yannick Kibolwe, et al.
Pubblicazione: (2025)
God's Innovation Project -- Empowering The Player With Generative AI
di: Nair, Ritvik, et al.
Pubblicazione: (2025)
di: Nair, Ritvik, et al.
Pubblicazione: (2025)
Reflecting on Design Paradigms of Animated Data Video Tools
di: Shen, Leixian, et al.
Pubblicazione: (2025)
di: Shen, Leixian, et al.
Pubblicazione: (2025)
Visualization in Motion in Video Games for Different Types of Data
di: Bucchieri, Federica, et al.
Pubblicazione: (2024)
di: Bucchieri, Federica, et al.
Pubblicazione: (2024)
RealUserSim: Bridging the Reality Gap in Agent Benchmarking via Grounded User Simulation
di: Zhu, Ming, et al.
Pubblicazione: (2026)
di: Zhu, Ming, et al.
Pubblicazione: (2026)
Video-Conferencing Beyond Screen-Sharing and Thumbnail Webcam Videos: Gesture-Aware Augmented Reality Video for Data-Rich Remote Presentations
di: Brehmer, Matthew
Pubblicazione: (2025)
di: Brehmer, Matthew
Pubblicazione: (2025)
AnimAlte:Designing AI-Infused Cartoon Videos to Improve Preschoolers' Language Learning with Family Engagement at Home
di: Tsang, Shiya, et al.
Pubblicazione: (2025)
di: Tsang, Shiya, et al.
Pubblicazione: (2025)
A Decentralized Frontier AI Architecture Based on Personal Instances, Synthetic Data, and Collective Context Synchronization
di: Małecki, Jacek, et al.
Pubblicazione: (2026)
di: Małecki, Jacek, et al.
Pubblicazione: (2026)
Documenti analoghi
-
xGen-MM-Vid (BLIP-3-Video): You Only Need 32 Tokens to Represent a Video Even in VLMs
di: Ryoo, Michael S., et al.
Pubblicazione: (2024) -
SMILE: A Composite Lexical-Semantic Metric for Question-Answering Evaluation
di: Kendre, Shrikant, et al.
Pubblicazione: (2025) -
HIVE: Harnessing Human Feedback for Instructional Visual Editing
di: Zhang, Shu, et al.
Pubblicazione: (2023) -
Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding
di: Wang, Ziyang, et al.
Pubblicazione: (2025) -
ReSpark: Leveraging Previous Data Reports as References to Generate New Reports with LLMs
di: Tian, Yuan, et al.
Pubblicazione: (2025)