Adaptive Guidance Semantically Enhanced via Multimodal LLM for Edge-Cloud Object Detection
Fuente:
arXiv
Saved in:
| Main Authors: | Hu, Yunqing, Yang, Zheming, Zhao, Chang, Ji, Wen |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SAEC: Scene-Aware Enhanced Edge-Cloud Collaborative Industrial Vision Inspection with Multimodal LLM
by: Tian, Yuhao, et al.
Published: (2025)
by: Tian, Yuhao, et al.
Published: (2025)
AIVD: Adaptive Edge-Cloud Collaboration for Accurate and Efficient Industrial Visual Detection
by: Hu, Yunqing, et al.
Published: (2026)
by: Hu, Yunqing, et al.
Published: (2026)
RSGen: Enhancing Layout-Driven Remote Sensing Image Generation with Diverse Edge Guidance
by: Hou, Xianbao, et al.
Published: (2026)
by: Hou, Xianbao, et al.
Published: (2026)
Frequency-Adaptive Low-Latency Object Detection Using Events and Frames
by: Zhang, Haitian, et al.
Published: (2024)
by: Zhang, Haitian, et al.
Published: (2024)
Diffusion as Reasoning: Enhancing Object Navigation via Diffusion Model Conditioned on LLM-based Object-Room Knowledge
by: Ji, Yiming, et al.
Published: (2024)
by: Ji, Yiming, et al.
Published: (2024)
Enhancing Weakly Supervised Multimodal Video Anomaly Detection through Text Guidance
by: Sun, Shengyang, et al.
Published: (2026)
by: Sun, Shengyang, et al.
Published: (2026)
LED: LLM Enhanced Open-Vocabulary Object Detection without Human Curated Data Generation
by: Zhou, Yang, et al.
Published: (2025)
by: Zhou, Yang, et al.
Published: (2025)
COMO: Cross-Mamba Interaction and Offset-Guided Fusion for Multimodal Object Detection
by: Liu, Chang, et al.
Published: (2024)
by: Liu, Chang, et al.
Published: (2024)
A new method for optical steel rope non-destructive damage detection
by: Bao, Yunqing, et al.
Published: (2024)
by: Bao, Yunqing, et al.
Published: (2024)
Unlocking Textual and Visual Wisdom: Open-Vocabulary 3D Object Detection Enhanced by Comprehensive Guidance from Text and Image
by: Jiao, Pengkun, et al.
Published: (2024)
by: Jiao, Pengkun, et al.
Published: (2024)
Towards Adaptive Open-Set Object Detection via Category-Level Collaboration Knowledge Mining
by: Ji, Yuqi, et al.
Published: (2026)
by: Ji, Yuqi, et al.
Published: (2026)
FlexiTex: Enhancing Texture Generation via Visual Guidance
by: Jiang, DaDong, et al.
Published: (2024)
by: Jiang, DaDong, et al.
Published: (2024)
Utilizing Graph Generation for Enhanced Domain Adaptive Object Detection
by: Wang, Mu
Published: (2024)
by: Wang, Mu
Published: (2024)
Multimodal-Enhanced Objectness Learner for Corner Case Detection in Autonomous Driving
by: Xiao, Lixing, et al.
Published: (2024)
by: Xiao, Lixing, et al.
Published: (2024)
From Street to Orbit: Training-Free Cross-View Retrieval via Location Semantics and LLM Guidance
by: Min, Jeongho, et al.
Published: (2025)
by: Min, Jeongho, et al.
Published: (2025)
EdgeSync: Faster Edge-model Updating via Adaptive Continuous Learning for Video Data Drift
by: Zhao, Peng, et al.
Published: (2024)
by: Zhao, Peng, et al.
Published: (2024)
Latent Distillation for Continual Object Detection at the Edge
by: Pasti, Francesco, et al.
Published: (2024)
by: Pasti, Francesco, et al.
Published: (2024)
Object Navigation with Structure-Semantic Reasoning-Based Multi-level Map and Multimodal Decision-Making LLM
by: Yan, Chongshang, et al.
Published: (2025)
by: Yan, Chongshang, et al.
Published: (2025)
Is 'Right' Right? Enhancing Object Orientation Understanding in Multimodal Large Language Models through Egocentric Instruction Tuning
by: Jung, Ji Hyeok, et al.
Published: (2024)
by: Jung, Ji Hyeok, et al.
Published: (2024)
Proto-OOD: Enhancing OOD Object Detection with Prototype Feature Similarity
by: Chen, Junkun, et al.
Published: (2024)
by: Chen, Junkun, et al.
Published: (2024)
Multimodal Prototype-Enhanced Network for Few-Shot Action Recognition
by: Ni, Xinzhe, et al.
Published: (2022)
by: Ni, Xinzhe, et al.
Published: (2022)
Implementing Edge Based Object Detection For Microplastic Debris
by: Singh, Amardeep, et al.
Published: (2023)
by: Singh, Amardeep, et al.
Published: (2023)
MCIE: Multimodal LLM-Driven Complex Instruction Image Editing with Spatial Guidance
by: Bai, Xuehai, et al.
Published: (2026)
by: Bai, Xuehai, et al.
Published: (2026)
Enhancing Alignment for Unified Multimodal Models via Semantically-Grounded Supervision
by: Kim, Jiyeong, et al.
Published: (2026)
by: Kim, Jiyeong, et al.
Published: (2026)
Monocular Normal Estimation via Shading Sequence Estimation
by: Li, Zongrui, et al.
Published: (2026)
by: Li, Zongrui, et al.
Published: (2026)
Weakly Supervised Camouflaged Object Detection Based on the SAM Model and Mask Guidance
by: Li, Xia, et al.
Published: (2026)
by: Li, Xia, et al.
Published: (2026)
Modelling Visual Semantics via Image Captioning to extract Enhanced Multi-Level Cross-Modal Semantic Incongruity Representation with Attention for Multimodal Sarcasm Detection
by: Aggarwal, Sajal, et al.
Published: (2024)
by: Aggarwal, Sajal, et al.
Published: (2024)
Tell Me What to Track: Infusing Robust Language Guidance for Enhanced Referring Multi-Object Tracking
by: Huang, Wenjun, et al.
Published: (2024)
by: Huang, Wenjun, et al.
Published: (2024)
Feature Recalibration Based Olfactory-Visual Multimodal Model for Enhanced Rice Deterioration Detection
by: Zhao, Rongqiang, et al.
Published: (2026)
by: Zhao, Rongqiang, et al.
Published: (2026)
SemVideo: Reconstructs What You Watch from Brain Activity via Hierarchical Semantic Guidance
by: Yang, Minghan, et al.
Published: (2026)
by: Yang, Minghan, et al.
Published: (2026)
AdaTok: Adaptive Token Compression with Object-Aware Representations for Efficient Multimodal LLMs
by: Zhang, Xinliang, et al.
Published: (2025)
by: Zhang, Xinliang, et al.
Published: (2025)
Spectral Discrepancy and Cross-modal Semantic Consistency Learning for Object Detection in Hyperspectral Image
by: He, Xiao, et al.
Published: (2025)
by: He, Xiao, et al.
Published: (2025)
Leveraging Unlabeled Data from Unknown Sources via Dual-Path Guidance for Deepfake Face Detection
by: Yang, Zhiqiang, et al.
Published: (2025)
by: Yang, Zhiqiang, et al.
Published: (2025)
HiddenObject: Modality-Agnostic Fusion for Multimodal Hidden Object Detection
by: Song, Harris, et al.
Published: (2025)
by: Song, Harris, et al.
Published: (2025)
Semantic Robustness Probing via Inpainting: An Interactive Tool for Safety-Critical Object Detection
by: Steckhan, Nico, et al.
Published: (2026)
by: Steckhan, Nico, et al.
Published: (2026)
Dual-Model Distillation for Efficient Action Classification with Hybrid Edge-Cloud Solution
by: Wei, Timothy, et al.
Published: (2024)
by: Wei, Timothy, et al.
Published: (2024)
Q-Zoom: Query-Aware Adaptive Perception for Efficient Multimodal Large Language Models
by: Shi, Yuheng, et al.
Published: (2026)
by: Shi, Yuheng, et al.
Published: (2026)
Domain Adaptive SAR Wake Detection: Leveraging Similarity Filtering and Memory Guidance
by: Gao, He, et al.
Published: (2025)
by: Gao, He, et al.
Published: (2025)
Enhancing Source-Free Domain Adaptive Object Detection with Low-confidence Pseudo Label Distillation
by: Yoon, Ilhoon, et al.
Published: (2024)
by: Yoon, Ilhoon, et al.
Published: (2024)
Complementary Pseudo Multimodal Feature for Point Cloud Anomaly Detection
by: Cao, Yunkang, et al.
Published: (2023)
by: Cao, Yunkang, et al.
Published: (2023)
Similar Items
-
SAEC: Scene-Aware Enhanced Edge-Cloud Collaborative Industrial Vision Inspection with Multimodal LLM
by: Tian, Yuhao, et al.
Published: (2025) -
AIVD: Adaptive Edge-Cloud Collaboration for Accurate and Efficient Industrial Visual Detection
by: Hu, Yunqing, et al.
Published: (2026) -
RSGen: Enhancing Layout-Driven Remote Sensing Image Generation with Diverse Edge Guidance
by: Hou, Xianbao, et al.
Published: (2026) -
Frequency-Adaptive Low-Latency Object Detection Using Events and Frames
by: Zhang, Haitian, et al.
Published: (2024) -
Diffusion as Reasoning: Enhancing Object Navigation via Diffusion Model Conditioned on LLM-based Object-Room Knowledge
by: Ji, Yiming, et al.
Published: (2024)