OCC-MLLM-Alpha:Empowering Multi-modal Large Language Model for the Understanding of Occluded Objects with Self-Supervised Test-Time Learning
Fuente:
arXiv
Salvato in:
| Autori principali: | Yang, Shuxin, Di, Xinhan |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
OCC-MLLM:Empowering Multimodal Large Language Model For the Understanding of Occluded Objects
di: Qiu, Wenmo, et al.
Pubblicazione: (2024)
di: Qiu, Wenmo, et al.
Pubblicazione: (2024)
OCC-MLLM-CoT-Alpha: Towards Multi-stage Occlusion Recognition Based on Large Language Models via 3D-Aware Supervision and Chain-of-Thoughts Guidance
di: Wang, Chaoyi, et al.
Pubblicazione: (2025)
di: Wang, Chaoyi, et al.
Pubblicazione: (2025)
Empowering Segmentation Ability to Multi-modal Large Language Models
di: Yang, Yuqi, et al.
Pubblicazione: (2024)
di: Yang, Yuqi, et al.
Pubblicazione: (2024)
dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models
di: Xin, Yi, et al.
Pubblicazione: (2025)
di: Xin, Yi, et al.
Pubblicazione: (2025)
LLMTrack: Semantic Multi-Object Tracking with Multi-modal Large Language Models
di: Liao, Pan, et al.
Pubblicazione: (2026)
di: Liao, Pan, et al.
Pubblicazione: (2026)
CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language Models
di: Huang, Junming, et al.
Pubblicazione: (2026)
di: Huang, Junming, et al.
Pubblicazione: (2026)
Test-Time Temporal Sampling for Efficient MLLM Video Understanding
di: Wang, Kaibin, et al.
Pubblicazione: (2025)
di: Wang, Kaibin, et al.
Pubblicazione: (2025)
4D-Bench: Benchmarking Multi-modal Large Language Models for 4D Object Understanding
di: Zhu, Wenxuan, et al.
Pubblicazione: (2025)
di: Zhu, Wenxuan, et al.
Pubblicazione: (2025)
Self-Supervised Learning of Deviation in Latent Representation for Co-speech Gesture Video Generation
di: Yang, Huan, et al.
Pubblicazione: (2024)
di: Yang, Huan, et al.
Pubblicazione: (2024)
Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models
di: Xu, Runsen, et al.
Pubblicazione: (2025)
di: Xu, Runsen, et al.
Pubblicazione: (2025)
EALD-MLLM: Emotion Analysis in Long-sequential and De-identity videos with Multi-modal Large Language Model
di: Li, Deng, et al.
Pubblicazione: (2024)
di: Li, Deng, et al.
Pubblicazione: (2024)
Understanding Information Storage and Transfer in Multi-modal Large Language Models
di: Basu, Samyadeep, et al.
Pubblicazione: (2024)
di: Basu, Samyadeep, et al.
Pubblicazione: (2024)
Multiple Consistency-guided Test-Time Adaptation for Contrastive Audio-Language Models with Unlabeled Audio
di: Chen, Gongyu, et al.
Pubblicazione: (2024)
di: Chen, Gongyu, et al.
Pubblicazione: (2024)
Bootstrapping MLLM for Weakly-Supervised Class-Agnostic Object Counting
di: Zhang, Xiaowen, et al.
Pubblicazione: (2026)
di: Zhang, Xiaowen, et al.
Pubblicazione: (2026)
GEM: Empowering MLLM for Grounded ECG Understanding with Time Series and Images
di: Lan, Xiang, et al.
Pubblicazione: (2025)
di: Lan, Xiang, et al.
Pubblicazione: (2025)
PUMA: Empowering Unified MLLM with Multi-granular Visual Generation
di: Fang, Rongyao, et al.
Pubblicazione: (2024)
di: Fang, Rongyao, et al.
Pubblicazione: (2024)
CoF: Coarse to Fine-Grained Image Understanding for Multi-modal Large Language Models
di: Wang, Yeyuan, et al.
Pubblicazione: (2024)
di: Wang, Yeyuan, et al.
Pubblicazione: (2024)
Explorations in Self-Supervised Learning: Dataset Composition Testing for Object Classification
di: Chavez, Raynor Kirkson E., et al.
Pubblicazione: (2024)
di: Chavez, Raynor Kirkson E., et al.
Pubblicazione: (2024)
ROODI: Reconstructing Occluded Objects with Denoising Inpainters
di: Chang, Yeonjin, et al.
Pubblicazione: (2025)
di: Chang, Yeonjin, et al.
Pubblicazione: (2025)
3DCity-LLM: Empowering Multi-modality Large Language Models for 3D City-scale Perception and Understanding
di: Chen, Yiping, et al.
Pubblicazione: (2026)
di: Chen, Yiping, et al.
Pubblicazione: (2026)
MLLM-HWSI: A Multimodal Large Language Model for Hierarchical Whole Slide Image Understanding
di: Alawode, Basit, et al.
Pubblicazione: (2026)
di: Alawode, Basit, et al.
Pubblicazione: (2026)
MLLM-SUL: Multimodal Large Language Model for Semantic Scene Understanding and Localization in Traffic Scenarios
di: Fan, Jiaqi, et al.
Pubblicazione: (2024)
di: Fan, Jiaqi, et al.
Pubblicazione: (2024)
Universal Video Temporal Grounding with Generative Multi-modal Large Language Models
di: Li, Zeqian, et al.
Pubblicazione: (2025)
di: Li, Zeqian, et al.
Pubblicazione: (2025)
SSAM: Self-Supervised Association Modeling for Test-Time Adaption
di: Wang, Yaxiong, et al.
Pubblicazione: (2025)
di: Wang, Yaxiong, et al.
Pubblicazione: (2025)
MLLM-Tool: A Multimodal Large Language Model For Tool Agent Learning
di: Wang, Chenyu, et al.
Pubblicazione: (2024)
di: Wang, Chenyu, et al.
Pubblicazione: (2024)
LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models
di: Fu, Shenghao, et al.
Pubblicazione: (2025)
di: Fu, Shenghao, et al.
Pubblicazione: (2025)
Improving Classification of Occluded Objects through Scene Context
di: King, Courtney M., et al.
Pubblicazione: (2025)
di: King, Courtney M., et al.
Pubblicazione: (2025)
A Diffusion-Based Framework for Occluded Object Movement
di: Duan, Zheng-Peng, et al.
Pubblicazione: (2025)
di: Duan, Zheng-Peng, et al.
Pubblicazione: (2025)
Hypergraph Multi-modal Large Language Model: Exploiting EEG and Eye-tracking Modalities to Evaluate Heterogeneous Responses for Video Understanding
di: Wu, Minghui, et al.
Pubblicazione: (2024)
di: Wu, Minghui, et al.
Pubblicazione: (2024)
LLMRA: Multi-modal Large Language Model based Restoration Assistant
di: Jin, Xiaoyu, et al.
Pubblicazione: (2024)
di: Jin, Xiaoyu, et al.
Pubblicazione: (2024)
Test-Time Adaptive Object Detection with Foundation Model
di: Gao, Yingjie, et al.
Pubblicazione: (2025)
di: Gao, Yingjie, et al.
Pubblicazione: (2025)
RGB-T Object Detection via Group Shuffled Multi-receptive Attention and Multi-modal Supervision
di: Wang, Jinzhong, et al.
Pubblicazione: (2024)
di: Wang, Jinzhong, et al.
Pubblicazione: (2024)
Understanding the Multi-modal Prompts of the Pre-trained Vision-Language Model
di: Ma, Shuailei, et al.
Pubblicazione: (2023)
di: Ma, Shuailei, et al.
Pubblicazione: (2023)
Real-Time Object Detection in Occluded Environment with Background Cluttering Effects Using Deep Learning
di: Aamir, Syed Muhammad, et al.
Pubblicazione: (2024)
di: Aamir, Syed Muhammad, et al.
Pubblicazione: (2024)
DocGenome: An Open Large-scale Scientific Document Benchmark for Training and Testing Multi-modal Large Language Models
di: Xia, Renqiu, et al.
Pubblicazione: (2024)
di: Xia, Renqiu, et al.
Pubblicazione: (2024)
ControlMLLM: Training-Free Visual Prompt Learning for Multimodal Large Language Models
di: Wu, Mingrui, et al.
Pubblicazione: (2024)
di: Wu, Mingrui, et al.
Pubblicazione: (2024)
ANTS: Adaptive Negative Textual Space Shaping for OOD Detection via Test-Time MLLM Understanding and Reasoning
di: Zhu, Wenjie, et al.
Pubblicazione: (2025)
di: Zhu, Wenjie, et al.
Pubblicazione: (2025)
Weakly Supervised Test-Time Domain Adaptation for Object Detection
di: Doan, Anh-Dzung, et al.
Pubblicazione: (2024)
di: Doan, Anh-Dzung, et al.
Pubblicazione: (2024)
Low-Rank Adaptation with Task-Relevant Feature Enhancement for Fine-tuning Language Models
di: Li, Changqun, et al.
Pubblicazione: (2024)
di: Li, Changqun, et al.
Pubblicazione: (2024)
HAIC: Improving Human Action Understanding and Generation with Better Captions for Multi-modal Large Language Models
di: Wang, Xiao, et al.
Pubblicazione: (2025)
di: Wang, Xiao, et al.
Pubblicazione: (2025)
Documenti analoghi
-
OCC-MLLM:Empowering Multimodal Large Language Model For the Understanding of Occluded Objects
di: Qiu, Wenmo, et al.
Pubblicazione: (2024) -
OCC-MLLM-CoT-Alpha: Towards Multi-stage Occlusion Recognition Based on Large Language Models via 3D-Aware Supervision and Chain-of-Thoughts Guidance
di: Wang, Chaoyi, et al.
Pubblicazione: (2025) -
Empowering Segmentation Ability to Multi-modal Large Language Models
di: Yang, Yuqi, et al.
Pubblicazione: (2024) -
dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models
di: Xin, Yi, et al.
Pubblicazione: (2025) -
LLMTrack: Semantic Multi-Object Tracking with Multi-modal Large Language Models
di: Liao, Pan, et al.
Pubblicazione: (2026)