StreamSense: Streaming Social Task Detection with Selective Vision-Language Model Routing
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Han, Ji, Deyi, Zhu, Lanyun, Luo, Jiebo, Lee, Roy Ka-Wei |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
A Simple Baseline for Streaming Video Understanding
por: Shen, Yujiao, et al.
Publicado: (2026)
por: Shen, Yujiao, et al.
Publicado: (2026)
Mahalanobis PatchCore: Covariance-Aware and Streaming-Compatible Industrial Anomaly Detection
por: Ferrari, Niccolò, et al.
Publicado: (2026)
por: Ferrari, Niccolò, et al.
Publicado: (2026)
Enhancing Eye Feature Estimation from Event Data Streams through Adaptive Inference State Space Modeling
por: Nguyen, Viet Dung, et al.
Publicado: (2026)
por: Nguyen, Viet Dung, et al.
Publicado: (2026)
Prompt Sensitivity in Vision-Language Grounding: How Small Changes in Wording Affect Object Detection
por: Deka, Dawar Jyoti, et al.
Publicado: (2026)
por: Deka, Dawar Jyoti, et al.
Publicado: (2026)
LLM-empowered Dynamic Prompt Routing for Vision-Language Models Tuning under Long-Tailed Distributions
por: Jia, Yongju, et al.
Publicado: (2025)
por: Jia, Yongju, et al.
Publicado: (2025)
SVGS-DSGAT: An IoT-Enabled Innovation in Underwater Robotic Object Detection Technology
por: Wu, Dongli, et al.
Publicado: (2025)
por: Wu, Dongli, et al.
Publicado: (2025)
Saliency-Aware Multi-Route Thinking: Revisiting Vision-Language Reasoning
por: Shi, Mingjia, et al.
Publicado: (2026)
por: Shi, Mingjia, et al.
Publicado: (2026)
DisasterM3: A Remote Sensing Vision-Language Dataset for Disaster Damage Assessment and Response
por: Wang, Junjue, et al.
Publicado: (2025)
por: Wang, Junjue, et al.
Publicado: (2025)
DNRSelect: Active Best View Selection for Deferred Neural Rendering
por: Wu, Dongli, et al.
Publicado: (2025)
por: Wu, Dongli, et al.
Publicado: (2025)
StreamCacheVGGT: Streaming Visual Geometry Transformers with Robust Scoring and Hybrid Cache Compression
por: Liu, Xuanyi, et al.
Publicado: (2026)
por: Liu, Xuanyi, et al.
Publicado: (2026)
FusionSense: Bridging Common Sense, Vision, and Touch for Robust Sparse-View Reconstruction
por: Fang, Irving, et al.
Publicado: (2024)
por: Fang, Irving, et al.
Publicado: (2024)
A Vision-Language Model for Focal Liver Lesion Classification
por: Jian, Song, et al.
Publicado: (2025)
por: Jian, Song, et al.
Publicado: (2025)
Grounding Synthetic Data Generation With Vision and Language Models
por: Çağlar, Ümit Mert, et al.
Publicado: (2026)
por: Çağlar, Ümit Mert, et al.
Publicado: (2026)
Decoupled Sensitivity-Consistency Learning for Weakly Supervised Video Anomaly Detection
por: Zheng, Hantao, et al.
Publicado: (2026)
por: Zheng, Hantao, et al.
Publicado: (2026)
Supervised Embedded Methods for Hyperspectral Band Selection
por: Zimmer, Yaniv, et al.
Publicado: (2024)
por: Zimmer, Yaniv, et al.
Publicado: (2024)
Streaming Anchor Loss: Augmenting Supervision with Temporal Significance
por: Sarawgi, Utkarsh Oggy, et al.
Publicado: (2023)
por: Sarawgi, Utkarsh Oggy, et al.
Publicado: (2023)
Multispectral Remote Sensing for Weed Detection in West Australian Agricultural Lands
por: Wang, Haitian, et al.
Publicado: (2025)
por: Wang, Haitian, et al.
Publicado: (2025)
DVLA-RL: Dual-Level Vision-Language Alignment with Reinforcement Learning Gating for Few-Shot Learning
por: Li, Wenhao, et al.
Publicado: (2026)
por: Li, Wenhao, et al.
Publicado: (2026)
Cross-Modal Transfer from Memes to Videos: Addressing Data Scarcity in Hateful Video Detection
por: Wang, Han, et al.
Publicado: (2025)
por: Wang, Han, et al.
Publicado: (2025)
Quantized Vision-Language Models for Damage Assessment: A Comparative Study of LLaVA-1.5-7B Quantization Levels
por: Yasuno, Takato
Publicado: (2026)
por: Yasuno, Takato
Publicado: (2026)
TD3Net: A temporal densely connected multi-dilated convolutional network for lipreading
por: Lee, Byung Hoon, et al.
Publicado: (2025)
por: Lee, Byung Hoon, et al.
Publicado: (2025)
High-Throughput Phenotyping using Computer Vision and Machine Learning
por: Singhvi, Vivaan, et al.
Publicado: (2024)
por: Singhvi, Vivaan, et al.
Publicado: (2024)
Pan-Arctic Permafrost Landform and Human-built Infrastructure Feature Detection with Vision Transformers and Location Embeddings
por: Perera, Amal S., et al.
Publicado: (2025)
por: Perera, Amal S., et al.
Publicado: (2025)
Streetscape Analysis with Generative AI (SAGAI): Vision-Language Assessment and Mapping of Urban Scenes
por: Perez, Joan, et al.
Publicado: (2025)
por: Perez, Joan, et al.
Publicado: (2025)
MSPCaps: A Multi-Scale Patchify Capsule Network with Cross-Agreement Routing for Visual Recognition
por: Hu, Yudong, et al.
Publicado: (2025)
por: Hu, Yudong, et al.
Publicado: (2025)
MdaIF: Robust One-Stop Multi-Degradation-Aware Image Fusion with Language-Driven Semantics
por: Li, Jing, et al.
Publicado: (2025)
por: Li, Jing, et al.
Publicado: (2025)
DVGBench: Implicit-to-Explicit Visual Grounding Benchmark in UAV Imagery with Large Vision-Language Models
por: Zhou, Yue, et al.
Publicado: (2026)
por: Zhou, Yue, et al.
Publicado: (2026)
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
por: Raoufi, Behnam, et al.
Publicado: (2025)
por: Raoufi, Behnam, et al.
Publicado: (2025)
Evaluating the Impact of Synthetic Data on Object Detection Tasks in Autonomous Driving
por: Özeren, Enes, et al.
Publicado: (2025)
por: Özeren, Enes, et al.
Publicado: (2025)
VLMaxxing through FrameMogging Training-Free Anti-Recomputation for Video Vision-Language Models
por: Bastien, JF, et al.
Publicado: (2026)
por: Bastien, JF, et al.
Publicado: (2026)
VLM-NCD:Novel Class Discovery with Vision-Based Large Language Models
por: Su, Yuetong, et al.
Publicado: (2025)
por: Su, Yuetong, et al.
Publicado: (2025)
Point, Detect, Count: Multi-Task Medical Image Understanding with Instruction-Tuned Vision-Language Models
por: Gautam, Sushant, et al.
Publicado: (2025)
por: Gautam, Sushant, et al.
Publicado: (2025)
Cytoarchitecture in Words: Weakly Supervised Vision-Language Modeling for Human Brain Microscopy
por: Sutton, Matthew, et al.
Publicado: (2026)
por: Sutton, Matthew, et al.
Publicado: (2026)
A Genealogy of Foundation Models in Remote Sensing
por: Lane, Kevin, et al.
Publicado: (2025)
por: Lane, Kevin, et al.
Publicado: (2025)
Interpretable Modeling of Driver Attention Shifts with a Vision--Language Model
por: Hamid, Kaiser, et al.
Publicado: (2025)
por: Hamid, Kaiser, et al.
Publicado: (2025)
DeltaVLM: Interactive Remote Sensing Image Change Analysis via Instruction-guided Difference Perception
por: Deng, Pei, et al.
Publicado: (2025)
por: Deng, Pei, et al.
Publicado: (2025)
Traffic Scene Small Target Detection Method Based on YOLOv8n-SPTS Model for Autonomous Driving
por: Wu, Songhan
Publicado: (2025)
por: Wu, Songhan
Publicado: (2025)
Global-Local Similarity for Efficient Fine-Grained Image Recognition with Vision Transformers
por: Rios, Edwin Arkel, et al.
Publicado: (2024)
por: Rios, Edwin Arkel, et al.
Publicado: (2024)
Fusion and Grouping Strategies in Deep Learning for Local Climate Zone Classification of Multimodal Remote Sensing Data
por: Thomas, Ancymol, et al.
Publicado: (2026)
por: Thomas, Ancymol, et al.
Publicado: (2026)
Model Agnostic Defense against Adversarial Patch Attacks on Object Detection in Unmanned Aerial Vehicles
por: Pathak, Saurabh, et al.
Publicado: (2024)
por: Pathak, Saurabh, et al.
Publicado: (2024)
Ejemplares similares
-
A Simple Baseline for Streaming Video Understanding
por: Shen, Yujiao, et al.
Publicado: (2026) -
Mahalanobis PatchCore: Covariance-Aware and Streaming-Compatible Industrial Anomaly Detection
por: Ferrari, Niccolò, et al.
Publicado: (2026) -
Enhancing Eye Feature Estimation from Event Data Streams through Adaptive Inference State Space Modeling
por: Nguyen, Viet Dung, et al.
Publicado: (2026) -
Prompt Sensitivity in Vision-Language Grounding: How Small Changes in Wording Affect Object Detection
por: Deka, Dawar Jyoti, et al.
Publicado: (2026) -
LLM-empowered Dynamic Prompt Routing for Vision-Language Models Tuning under Long-Tailed Distributions
por: Jia, Yongju, et al.
Publicado: (2025)