StreamSense: Streaming Social Task Detection with Selective Vision-Language Model Routing
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Han, Ji, Deyi, Zhu, Lanyun, Luo, Jiebo, Lee, Roy Ka-Wei |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A Simple Baseline for Streaming Video Understanding
by: Shen, Yujiao, et al.
Published: (2026)
by: Shen, Yujiao, et al.
Published: (2026)
Mahalanobis PatchCore: Covariance-Aware and Streaming-Compatible Industrial Anomaly Detection
by: Ferrari, Niccolò, et al.
Published: (2026)
by: Ferrari, Niccolò, et al.
Published: (2026)
Enhancing Eye Feature Estimation from Event Data Streams through Adaptive Inference State Space Modeling
by: Nguyen, Viet Dung, et al.
Published: (2026)
by: Nguyen, Viet Dung, et al.
Published: (2026)
Prompt Sensitivity in Vision-Language Grounding: How Small Changes in Wording Affect Object Detection
by: Deka, Dawar Jyoti, et al.
Published: (2026)
by: Deka, Dawar Jyoti, et al.
Published: (2026)
LLM-empowered Dynamic Prompt Routing for Vision-Language Models Tuning under Long-Tailed Distributions
by: Jia, Yongju, et al.
Published: (2025)
by: Jia, Yongju, et al.
Published: (2025)
SVGS-DSGAT: An IoT-Enabled Innovation in Underwater Robotic Object Detection Technology
by: Wu, Dongli, et al.
Published: (2025)
by: Wu, Dongli, et al.
Published: (2025)
Saliency-Aware Multi-Route Thinking: Revisiting Vision-Language Reasoning
by: Shi, Mingjia, et al.
Published: (2026)
by: Shi, Mingjia, et al.
Published: (2026)
DisasterM3: A Remote Sensing Vision-Language Dataset for Disaster Damage Assessment and Response
by: Wang, Junjue, et al.
Published: (2025)
by: Wang, Junjue, et al.
Published: (2025)
DNRSelect: Active Best View Selection for Deferred Neural Rendering
by: Wu, Dongli, et al.
Published: (2025)
by: Wu, Dongli, et al.
Published: (2025)
StreamCacheVGGT: Streaming Visual Geometry Transformers with Robust Scoring and Hybrid Cache Compression
by: Liu, Xuanyi, et al.
Published: (2026)
by: Liu, Xuanyi, et al.
Published: (2026)
FusionSense: Bridging Common Sense, Vision, and Touch for Robust Sparse-View Reconstruction
by: Fang, Irving, et al.
Published: (2024)
by: Fang, Irving, et al.
Published: (2024)
A Vision-Language Model for Focal Liver Lesion Classification
by: Jian, Song, et al.
Published: (2025)
by: Jian, Song, et al.
Published: (2025)
Grounding Synthetic Data Generation With Vision and Language Models
by: Çağlar, Ümit Mert, et al.
Published: (2026)
by: Çağlar, Ümit Mert, et al.
Published: (2026)
Decoupled Sensitivity-Consistency Learning for Weakly Supervised Video Anomaly Detection
by: Zheng, Hantao, et al.
Published: (2026)
by: Zheng, Hantao, et al.
Published: (2026)
Supervised Embedded Methods for Hyperspectral Band Selection
by: Zimmer, Yaniv, et al.
Published: (2024)
by: Zimmer, Yaniv, et al.
Published: (2024)
Streaming Anchor Loss: Augmenting Supervision with Temporal Significance
by: Sarawgi, Utkarsh Oggy, et al.
Published: (2023)
by: Sarawgi, Utkarsh Oggy, et al.
Published: (2023)
Multispectral Remote Sensing for Weed Detection in West Australian Agricultural Lands
by: Wang, Haitian, et al.
Published: (2025)
by: Wang, Haitian, et al.
Published: (2025)
DVLA-RL: Dual-Level Vision-Language Alignment with Reinforcement Learning Gating for Few-Shot Learning
by: Li, Wenhao, et al.
Published: (2026)
by: Li, Wenhao, et al.
Published: (2026)
Cross-Modal Transfer from Memes to Videos: Addressing Data Scarcity in Hateful Video Detection
by: Wang, Han, et al.
Published: (2025)
by: Wang, Han, et al.
Published: (2025)
Quantized Vision-Language Models for Damage Assessment: A Comparative Study of LLaVA-1.5-7B Quantization Levels
by: Yasuno, Takato
Published: (2026)
by: Yasuno, Takato
Published: (2026)
TD3Net: A temporal densely connected multi-dilated convolutional network for lipreading
by: Lee, Byung Hoon, et al.
Published: (2025)
by: Lee, Byung Hoon, et al.
Published: (2025)
High-Throughput Phenotyping using Computer Vision and Machine Learning
by: Singhvi, Vivaan, et al.
Published: (2024)
by: Singhvi, Vivaan, et al.
Published: (2024)
Pan-Arctic Permafrost Landform and Human-built Infrastructure Feature Detection with Vision Transformers and Location Embeddings
by: Perera, Amal S., et al.
Published: (2025)
by: Perera, Amal S., et al.
Published: (2025)
Streetscape Analysis with Generative AI (SAGAI): Vision-Language Assessment and Mapping of Urban Scenes
by: Perez, Joan, et al.
Published: (2025)
by: Perez, Joan, et al.
Published: (2025)
MSPCaps: A Multi-Scale Patchify Capsule Network with Cross-Agreement Routing for Visual Recognition
by: Hu, Yudong, et al.
Published: (2025)
by: Hu, Yudong, et al.
Published: (2025)
MdaIF: Robust One-Stop Multi-Degradation-Aware Image Fusion with Language-Driven Semantics
by: Li, Jing, et al.
Published: (2025)
by: Li, Jing, et al.
Published: (2025)
DVGBench: Implicit-to-Explicit Visual Grounding Benchmark in UAV Imagery with Large Vision-Language Models
by: Zhou, Yue, et al.
Published: (2026)
by: Zhou, Yue, et al.
Published: (2026)
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
by: Raoufi, Behnam, et al.
Published: (2025)
by: Raoufi, Behnam, et al.
Published: (2025)
Evaluating the Impact of Synthetic Data on Object Detection Tasks in Autonomous Driving
by: Özeren, Enes, et al.
Published: (2025)
by: Özeren, Enes, et al.
Published: (2025)
VLMaxxing through FrameMogging Training-Free Anti-Recomputation for Video Vision-Language Models
by: Bastien, JF, et al.
Published: (2026)
by: Bastien, JF, et al.
Published: (2026)
VLM-NCD:Novel Class Discovery with Vision-Based Large Language Models
by: Su, Yuetong, et al.
Published: (2025)
by: Su, Yuetong, et al.
Published: (2025)
Point, Detect, Count: Multi-Task Medical Image Understanding with Instruction-Tuned Vision-Language Models
by: Gautam, Sushant, et al.
Published: (2025)
by: Gautam, Sushant, et al.
Published: (2025)
Cytoarchitecture in Words: Weakly Supervised Vision-Language Modeling for Human Brain Microscopy
by: Sutton, Matthew, et al.
Published: (2026)
by: Sutton, Matthew, et al.
Published: (2026)
A Genealogy of Foundation Models in Remote Sensing
by: Lane, Kevin, et al.
Published: (2025)
by: Lane, Kevin, et al.
Published: (2025)
Interpretable Modeling of Driver Attention Shifts with a Vision--Language Model
by: Hamid, Kaiser, et al.
Published: (2025)
by: Hamid, Kaiser, et al.
Published: (2025)
DeltaVLM: Interactive Remote Sensing Image Change Analysis via Instruction-guided Difference Perception
by: Deng, Pei, et al.
Published: (2025)
by: Deng, Pei, et al.
Published: (2025)
Traffic Scene Small Target Detection Method Based on YOLOv8n-SPTS Model for Autonomous Driving
by: Wu, Songhan
Published: (2025)
by: Wu, Songhan
Published: (2025)
Global-Local Similarity for Efficient Fine-Grained Image Recognition with Vision Transformers
by: Rios, Edwin Arkel, et al.
Published: (2024)
by: Rios, Edwin Arkel, et al.
Published: (2024)
Fusion and Grouping Strategies in Deep Learning for Local Climate Zone Classification of Multimodal Remote Sensing Data
by: Thomas, Ancymol, et al.
Published: (2026)
by: Thomas, Ancymol, et al.
Published: (2026)
Model Agnostic Defense against Adversarial Patch Attacks on Object Detection in Unmanned Aerial Vehicles
by: Pathak, Saurabh, et al.
Published: (2024)
by: Pathak, Saurabh, et al.
Published: (2024)
Similar Items
-
A Simple Baseline for Streaming Video Understanding
by: Shen, Yujiao, et al.
Published: (2026) -
Mahalanobis PatchCore: Covariance-Aware and Streaming-Compatible Industrial Anomaly Detection
by: Ferrari, Niccolò, et al.
Published: (2026) -
Enhancing Eye Feature Estimation from Event Data Streams through Adaptive Inference State Space Modeling
by: Nguyen, Viet Dung, et al.
Published: (2026) -
Prompt Sensitivity in Vision-Language Grounding: How Small Changes in Wording Affect Object Detection
by: Deka, Dawar Jyoti, et al.
Published: (2026) -
LLM-empowered Dynamic Prompt Routing for Vision-Language Models Tuning under Long-Tailed Distributions
by: Jia, Yongju, et al.
Published: (2025)