Perceive What Matters: Relevance-Driven Scheduling for Multimodal Streaming Perception
Fuente:
arXiv
Saved in:
| Main Authors: | Huang, Dingcheng, Zhang, Xiaotong, Youcef-Toumi, Kamal |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Relevance-driven Decision Making for Safer and More Efficient Human Robot Collaboration
by: Zhang, Xiaotong, et al.
Published: (2024)
by: Zhang, Xiaotong, et al.
Published: (2024)
Teller: Real-Time Streaming Audio-Driven Portrait Animation with Autoregressive Motion Generation
by: Zhen, Dingcheng, et al.
Published: (2025)
by: Zhen, Dingcheng, et al.
Published: (2025)
Relevance for Human Robot Collaboration
by: Zhang, Xiaotong, et al.
Published: (2024)
by: Zhang, Xiaotong, et al.
Published: (2024)
VideoPerceiver: Enhancing Fine-Grained Temporal Perception in Video Multimodal Large Language Models
by: Zhao, Fufangchen, et al.
Published: (2025)
by: Zhao, Fufangchen, et al.
Published: (2025)
Data Augmentation through Background Removal for Apple Leaf Disease Classification Using the MobileNetV2 Model
by: Ferdi, Youcef
Published: (2024)
by: Ferdi, Youcef
Published: (2024)
Move What Matters: Parameter-Efficient Domain Adaptation via Optimal Transport Flow for Collaborative Perception
by: Jia, Zesheng, et al.
Published: (2026)
by: Jia, Zesheng, et al.
Published: (2026)
Continuous Perception Matters: Diagnosing Temporal Integration Failures in Multimodal Models
by: Wang, Zeyu, et al.
Published: (2024)
by: Wang, Zeyu, et al.
Published: (2024)
What Matters When Repurposing Diffusion Models for General Dense Perception Tasks?
by: Xu, Guangkai, et al.
Published: (2024)
by: Xu, Guangkai, et al.
Published: (2024)
StixelNExT: Toward Monocular Low-Weight Perception for Object Segmentation and Free Space Detection
by: Vosshans, Marcel, et al.
Published: (2024)
by: Vosshans, Marcel, et al.
Published: (2024)
Slow Perception: Let's Perceive Geometric Figures Step-by-step
by: Wei, Haoran, et al.
Published: (2024)
by: Wei, Haoran, et al.
Published: (2024)
DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception
by: Li, Xiaotong, et al.
Published: (2024)
by: Li, Xiaotong, et al.
Published: (2024)
OmniStream: Mastering Perception, Reconstruction and Action in Continuous Streams
by: Yan, Yibin, et al.
Published: (2026)
by: Yan, Yibin, et al.
Published: (2026)
StixelNExT++: Lightweight Monocular Scene Segmentation and Representation for Collective Perception
by: Vosshans, Marcel, et al.
Published: (2025)
by: Vosshans, Marcel, et al.
Published: (2025)
Real-time Stereo-based 3D Object Detection for Streaming Perception
by: Li, Changcai, et al.
Published: (2024)
by: Li, Changcai, et al.
Published: (2024)
LLAVADI: What Matters For Multimodal Large Language Models Distillation
by: Xu, Shilin, et al.
Published: (2024)
by: Xu, Shilin, et al.
Published: (2024)
Label What Matters: Modality-Balanced and Difficulty-Aware Multimodal Active Learning
by: Zeng, Yuqiao, et al.
Published: (2026)
by: Zeng, Yuqiao, et al.
Published: (2026)
Multimodal Label Relevance Ranking via Reinforcement Learning
by: Guo, Taian, et al.
Published: (2024)
by: Guo, Taian, et al.
Published: (2024)
Semantic-aware Next-Best-View for Multi-DoFs Mobile System in Search-and-Acquisition based Visual Perception
by: Yu, Xiaotong, et al.
Published: (2024)
by: Yu, Xiaotong, et al.
Published: (2024)
Aligning What Vision-Language Models See and Perceive with Adaptive Information Flow
by: Liu, Chengxin, et al.
Published: (2026)
by: Liu, Chengxin, et al.
Published: (2026)
StreamReady: Learning What to Answer and When in Long Streaming Videos
by: Azad, Shehreen, et al.
Published: (2026)
by: Azad, Shehreen, et al.
Published: (2026)
QUEST: Query Stream for Practical Cooperative Perception
by: Fan, Siqi, et al.
Published: (2023)
by: Fan, Siqi, et al.
Published: (2023)
Multimodal Prompt Perceiver: Empower Adaptiveness, Generalizability and Fidelity for All-in-One Image Restoration
by: Ai, Yuang, et al.
Published: (2023)
by: Ai, Yuang, et al.
Published: (2023)
What You Perceive Is What You Conceive: A Cognition-Inspired Framework for Open Vocabulary Image Segmentation
by: Lin, Jianghang, et al.
Published: (2025)
by: Lin, Jianghang, et al.
Published: (2025)
Starve to Perceive: Taming Lazy Perception in VLMs with Constrained Visual Bandwidth
by: Wu, Yuhuan, et al.
Published: (2026)
by: Wu, Yuhuan, et al.
Published: (2026)
Seeing What Matters: Visual Preference Policy Optimization for Visual Generation
by: Ni, Ziqi, et al.
Published: (2025)
by: Ni, Ziqi, et al.
Published: (2025)
Spotlight on Token Perception for Multimodal Reinforcement Learning
by: Huang, Siyuan, et al.
Published: (2025)
by: Huang, Siyuan, et al.
Published: (2025)
Transtreaming: Adaptive Delay-aware Transformer for Real-time Streaming Perception
by: Zhang, Xiang, et al.
Published: (2024)
by: Zhang, Xiang, et al.
Published: (2024)
What Really Matters for Learning-based LiDAR-Camera Calibration
by: Huang, Shujuan, et al.
Published: (2025)
by: Huang, Shujuan, et al.
Published: (2025)
Do VLMs Perceive or Recall? Probing Visual Perception vs. Memory with Classic Visual Illusions
by: Sun, Xiaoxiao, et al.
Published: (2026)
by: Sun, Xiaoxiao, et al.
Published: (2026)
Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length
by: Huang, Yubo, et al.
Published: (2025)
by: Huang, Yubo, et al.
Published: (2025)
Conservative Estimation of Perception Relevance of Dynamic Objects for Safe Trajectories in Automotive Scenarios
by: Mori, Ken, et al.
Published: (2023)
by: Mori, Ken, et al.
Published: (2023)
Perceive, Verify and Understand Long Video: Multi-Granular Perception and Active Verification via Interactive Agents
by: Li, Jiahua, et al.
Published: (2025)
by: Li, Jiahua, et al.
Published: (2025)
SMSP: A Plug-and-Play Strategy of Multi-Scale Perception for MLLMs to Perceive Visual Illusions
by: Tu, Jinzhe, et al.
Published: (2026)
by: Tu, Jinzhe, et al.
Published: (2026)
Perceive, Understand and Restore: Real-World Image Super-Resolution with Autoregressive Multimodal Generative Models
by: Wei, Hongyang, et al.
Published: (2025)
by: Wei, Hongyang, et al.
Published: (2025)
Perceive-IR: Learning to Perceive Degradation Better for All-in-One Image Restoration
by: Zhang, Xu, et al.
Published: (2024)
by: Zhang, Xu, et al.
Published: (2024)
MR-MLLM: Mutual Reinforcement of Multimodal Comprehension and Vision Perception
by: Wang, Guanqun, et al.
Published: (2024)
by: Wang, Guanqun, et al.
Published: (2024)
Who Walks With You Matters: Perceiving Social Interactions with Groups for Pedestrian Trajectory Prediction
by: Zou, Ziqian, et al.
Published: (2024)
by: Zou, Ziqian, et al.
Published: (2024)
SiMO: Single-Modality-Operable Multimodal Collaborative Perception
by: Wen, Jiageng, et al.
Published: (2026)
by: Wen, Jiageng, et al.
Published: (2026)
CSDNet: Detect Salient Object in Depth-Thermal via A Lightweight Cross Shallow and Deep Perception Network
by: Yu, Xiaotong, et al.
Published: (2024)
by: Yu, Xiaotong, et al.
Published: (2024)
PlantSAM: An Object Detection-Driven Segmentation Pipeline for Herbarium Specimens
by: Sklab, Youcef, et al.
Published: (2025)
by: Sklab, Youcef, et al.
Published: (2025)
Similar Items
-
Relevance-driven Decision Making for Safer and More Efficient Human Robot Collaboration
by: Zhang, Xiaotong, et al.
Published: (2024) -
Teller: Real-Time Streaming Audio-Driven Portrait Animation with Autoregressive Motion Generation
by: Zhen, Dingcheng, et al.
Published: (2025) -
Relevance for Human Robot Collaboration
by: Zhang, Xiaotong, et al.
Published: (2024) -
VideoPerceiver: Enhancing Fine-Grained Temporal Perception in Video Multimodal Large Language Models
by: Zhao, Fufangchen, et al.
Published: (2025) -
Data Augmentation through Background Removal for Apple Leaf Disease Classification Using the MobileNetV2 Model
by: Ferdi, Youcef
Published: (2024)