ViSpeak: Visual Instruction Feedback in Streaming Videos
Fuente:
arXiv
Saved in:
| Main Authors: | Fu, Shenghao, Yang, Qize, Li, Yuan-Ming, Peng, Yi-Xing, Lin, Kun-Yu, Wei, Xihan, Hu, Jian-Fang, Xie, Xiaohua, Zheng, Wei-Shi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LOVE-R1: Advancing Long Video Understanding with an Adaptive Zoom-in Mechanism via Multi-Step Reasoning
by: Fu, Shenghao, et al.
Published: (2025)
by: Fu, Shenghao, et al.
Published: (2025)
Frozen-DETR: Enhancing DETR with Image Understanding from Frozen Foundation Models
by: Fu, Shenghao, et al.
Published: (2024)
by: Fu, Shenghao, et al.
Published: (2024)
A Hierarchical Semantic Distillation Framework for Open-Vocabulary Object Detection
by: Fu, Shenghao, et al.
Published: (2025)
by: Fu, Shenghao, et al.
Published: (2025)
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding
by: Peng, Yi-Xing, et al.
Published: (2025)
by: Peng, Yi-Xing, et al.
Published: (2025)
LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models
by: Fu, Shenghao, et al.
Published: (2025)
by: Fu, Shenghao, et al.
Published: (2025)
IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation
by: Li, Yuan-Ming, et al.
Published: (2025)
by: Li, Yuan-Ming, et al.
Published: (2025)
Omni-Emotion: Extending Video MLLM with Detailed Face and Audio Modeling for Multimodal Emotion Analysis
by: Yang, Qize, et al.
Published: (2025)
by: Yang, Qize, et al.
Published: (2025)
HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding
by: Zhao, Jiaxing, et al.
Published: (2025)
by: Zhao, Jiaxing, et al.
Published: (2025)
DreamView: Injecting View-specific Text Guidance into Text-to-3D Generation
by: Yan, Junkai, et al.
Published: (2024)
by: Yan, Junkai, et al.
Published: (2024)
ObjEmbed: Towards Universal Multimodal Object Embeddings
by: Fu, Shenghao, et al.
Published: (2026)
by: Fu, Shenghao, et al.
Published: (2026)
WeDetect: Fast Open-Vocabulary Object Detection as Retrieval
by: Fu, Shenghao, et al.
Published: (2025)
by: Fu, Shenghao, et al.
Published: (2025)
HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context
by: Yang, Qize, et al.
Published: (2025)
by: Yang, Qize, et al.
Published: (2025)
Facial Dynamics in Video: Instruction Tuning for Improved Facial Expression Perception and Contextual Awareness
by: Zhao, Jiaxing, et al.
Published: (2025)
by: Zhao, Jiaxing, et al.
Published: (2025)
Streaming Video Instruction Tuning
by: Xia, Jiaer, et al.
Published: (2025)
by: Xia, Jiaer, et al.
Published: (2025)
ReferDINO: Referring Video Object Segmentation with Visual Grounding Foundations
by: Liang, Tianming, et al.
Published: (2025)
by: Liang, Tianming, et al.
Published: (2025)
ViScratch: Using Large Language Models and Gameplay Videos for Automated Feedback in Scratch
by: Si, Yuan, et al.
Published: (2025)
by: Si, Yuan, et al.
Published: (2025)
Let ViT Speak: Generative Language-Image Pre-training
by: Fang, Yan, et al.
Published: (2026)
by: Fang, Yan, et al.
Published: (2026)
LLaVA-Octopus: Unlocking Instruction-Driven Adaptive Projector Fusion for Video Understanding
by: Sun, Boyuan, et al.
Published: (2025)
by: Sun, Boyuan, et al.
Published: (2025)
ReferDINO-Plus: 2nd Solution for 4th PVUW MeViS Challenge at CVPR 2025
by: Liang, Tianming, et al.
Published: (2025)
by: Liang, Tianming, et al.
Published: (2025)
Towards Sampling Data Structures for Tensor Products in Turnstile Streams
by: Song, Zhao, et al.
Published: (2025)
by: Song, Zhao, et al.
Published: (2025)
ViViD: Video Virtual Try-on using Diffusion Models
by: Fang, Zixun, et al.
Published: (2024)
by: Fang, Zixun, et al.
Published: (2024)
Perfect Sampling in Turnstile Streams Beyond Small Moments
by: Woodruff, David P., et al.
Published: (2025)
by: Woodruff, David P., et al.
Published: (2025)
InstructionBench: An Instructional Video Understanding Benchmark
by: Wei, Haiwan, et al.
Published: (2025)
by: Wei, Haiwan, et al.
Published: (2025)
Foreign Language Instruction in Chinese‐Speaking Neurotypical and Neurodivergent Children
by: Liming Zhou, et al.
Published: (2026)
by: Liming Zhou, et al.
Published: (2026)
Liminal Nostalgia: Internet Cafés, Immobile Millennials, and Suspended Identity in Hebei, China
by: Shenghao Xing
Published: (2026)
by: Shenghao Xing
Published: (2026)
ViLLa: Video Reasoning Segmentation with Large Language Model
by: Zheng, Rongkun, et al.
Published: (2024)
by: Zheng, Rongkun, et al.
Published: (2024)
Prediction, Communication, and Computing Duration Optimization for VR Video Streaming
by: Wei, Xing, et al.
Published: (2019)
by: Wei, Xing, et al.
Published: (2019)
Teaching Time Series to See and Speak: Forecasting with Aligned Visual and Textual Perspectives
by: Dong, Sixun, et al.
Published: (2025)
by: Dong, Sixun, et al.
Published: (2025)
Rethinking CLIP-based Video Learners in Cross-Domain Open-Vocabulary Action Recognition
by: Lin, Kun-Yu, et al.
Published: (2024)
by: Lin, Kun-Yu, et al.
Published: (2024)
Modeling Multiple Normal Action Representations for Error Detection in Procedural Tasks
by: Huang, Wei-Jin, et al.
Published: (2025)
by: Huang, Wei-Jin, et al.
Published: (2025)
Rethinking Few-shot Class-incremental Learning: Learning from Yourself
by: Tang, Yu-Ming, et al.
Published: (2024)
by: Tang, Yu-Ming, et al.
Published: (2024)
ReViP: Mitigating False Completion in Vision-Language-Action Models with Vision-Proprioception Rebalance
by: Li, Zhuohao, et al.
Published: (2026)
by: Li, Zhuohao, et al.
Published: (2026)
A Versatile Framework for Multi-scene Person Re-identification
by: Zheng, Wei-Shi, et al.
Published: (2024)
by: Zheng, Wei-Shi, et al.
Published: (2024)
Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding
by: Jin, Peng, et al.
Published: (2023)
by: Jin, Peng, et al.
Published: (2023)
In-Video Instructions: Visual Signals as Generative Control
by: Fang, Gongfan, et al.
Published: (2025)
by: Fang, Gongfan, et al.
Published: (2025)
Weight Distribution of Repeated-Root Cyclic Codes with Prime Power Lengths
by: Zhao, Wei, et al.
Published: (2023)
by: Zhao, Wei, et al.
Published: (2023)
STRIDE: When to Speak Meets Sequence Denoising for Streaming Video Understanding
by: Kim, Junho, et al.
Published: (2026)
by: Kim, Junho, et al.
Published: (2026)
Siamese Learning with Joint Alignment and Regression for Weakly-Supervised Video Paragraph Grounding
by: Tan, Chaolei, et al.
Published: (2024)
by: Tan, Chaolei, et al.
Published: (2024)
LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs
by: Sun, Boyuan, et al.
Published: (2025)
by: Sun, Boyuan, et al.
Published: (2025)
Tumor budding is an optimal indictor of occult cervical metastasis in clinical early‐stage buccal mucosa squamous cell carcinoma
by: Zhi Zheng, et al.
Published: (2024)
by: Zhi Zheng, et al.
Published: (2024)
Similar Items
-
LOVE-R1: Advancing Long Video Understanding with an Adaptive Zoom-in Mechanism via Multi-Step Reasoning
by: Fu, Shenghao, et al.
Published: (2025) -
Frozen-DETR: Enhancing DETR with Image Understanding from Frozen Foundation Models
by: Fu, Shenghao, et al.
Published: (2024) -
A Hierarchical Semantic Distillation Framework for Open-Vocabulary Object Detection
by: Fu, Shenghao, et al.
Published: (2025) -
ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding
by: Peng, Yi-Xing, et al.
Published: (2025) -
LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models
by: Fu, Shenghao, et al.
Published: (2025)