An analysis of HOI: using a training-free method with multimodal visual foundation models when only the test set is available, without the training set
Fuente:
arXiv
Saved in:
| Main Author: | Ai, Chaoyi |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
A multimodal vision foundation model for generalizable knee pathology
by: Yu, Kang, et al.
Published: (2026)
by: Yu, Kang, et al.
Published: (2026)
Transformers self-organize like newborn visual systems when trained in prenatal worlds
by: Pandey, Lalit, et al.
Published: (2026)
by: Pandey, Lalit, et al.
Published: (2026)
Avoid Wasted Annotation Costs in Open-set Active Learning with Pre-trained Vision-Language Model
by: Heo, Jaehyuk, et al.
Published: (2024)
by: Heo, Jaehyuk, et al.
Published: (2024)
EyeCLIP: A visual-language foundation model for multi-modal ophthalmic image analysis
by: Shi, Danli, et al.
Published: (2024)
by: Shi, Danli, et al.
Published: (2024)
ActiveMark: on watermarking of visual foundation models via massive activations
by: Chistyakova, Anna, et al.
Published: (2025)
by: Chistyakova, Anna, et al.
Published: (2025)
PB-IAD: Utilizing multimodal foundation models for semantic industrial anomaly detection in dynamic manufacturing environments
by: Hofmann, Bernd, et al.
Published: (2025)
by: Hofmann, Bernd, et al.
Published: (2025)
Leveraging AI multimodal geospatial foundation models for improved near-real-time flood mapping at a global scale
by: Tulbure, Mirela G., et al.
Published: (2025)
by: Tulbure, Mirela G., et al.
Published: (2025)
DiTraj: training-free trajectory control for video diffusion transformer
by: Lei, Cheng, et al.
Published: (2025)
by: Lei, Cheng, et al.
Published: (2025)
IMPACT-HOI: Supervisory Control for Onset-Anchored Partial HOI Event Construction
by: Zhang, Haoshen, et al.
Published: (2026)
by: Zhang, Haoshen, et al.
Published: (2026)
Enabling clinical use of foundation models for computational pathology
by: Henriksen, Audun L, et al.
Published: (2026)
by: Henriksen, Audun L, et al.
Published: (2026)
An interpretable framework using foundation models for fish sex identification
by: Miao, Zheng, et al.
Published: (2026)
by: Miao, Zheng, et al.
Published: (2026)
SHOE: Semantic HOI Open-Vocabulary Evaluation Metric
by: Noack, Maja, et al.
Published: (2026)
by: Noack, Maja, et al.
Published: (2026)
NERVE: Neighbourhood & Entropy-guided Random-walk for training free open-Vocabulary sEgmentation
by: Mahatha, Kunal, et al.
Published: (2025)
by: Mahatha, Kunal, et al.
Published: (2025)
Improving Pre-trained Segmentation Models using Post-Processing
by: Parida, Abhijeet, et al.
Published: (2025)
by: Parida, Abhijeet, et al.
Published: (2025)
Recall and Refine: A Simple but Effective Source-free Open-set Domain Adaptation Framework
by: Nejjar, Ismail, et al.
Published: (2024)
by: Nejjar, Ismail, et al.
Published: (2024)
Pixel Seal: Adversarial-only training for invisible image and video watermarking
by: Souček, Tomáš, et al.
Published: (2025)
by: Souček, Tomáš, et al.
Published: (2025)
Generalized Denoising Diffusion Codebook Models (gDDCM): Tokenizing images using a pre-trained diffusion model
by: Kong, Fei
Published: (2025)
by: Kong, Fei
Published: (2025)
DynaHOI: Benchmarking Hand-Object Interaction for Dynamic Target
by: Hu, BoCheng, et al.
Published: (2026)
by: Hu, BoCheng, et al.
Published: (2026)
Referential communication in heterogeneous communities of pre-trained visual deep networks
by: Mahaut, Matéo, et al.
Published: (2023)
by: Mahaut, Matéo, et al.
Published: (2023)
Driving scenario generation and evaluation using a structured layer representation and foundational models
by: Hubert, Arthur, et al.
Published: (2025)
by: Hubert, Arthur, et al.
Published: (2025)
LLMs can see and hear without any training
by: Ashutosh, Kumar, et al.
Published: (2025)
by: Ashutosh, Kumar, et al.
Published: (2025)
Exploring Annotation-free Image Captioning with Retrieval-augmented Pseudo Sentence Generation
by: Li, Zhiyuan, et al.
Published: (2023)
by: Li, Zhiyuan, et al.
Published: (2023)
Geospatial foundation models for image analysis: evaluating and enhancing NASA-IBM Prithvi's domain adaptability
by: Hsu, Chia-Yu, et al.
Published: (2024)
by: Hsu, Chia-Yu, et al.
Published: (2024)
VLM-HOI: Vision Language Models for Interpretable Human-Object Interaction Analysis
by: Kang, Donggoo, et al.
Published: (2024)
by: Kang, Donggoo, et al.
Published: (2024)
Mitigating Long-Tail Bias in HOI Detection via Adaptive Diversity Cache
by: Jiang, Yuqiu, et al.
Published: (2025)
by: Jiang, Yuqiu, et al.
Published: (2025)
StyleAutoEncoder for manipulating image attributes using pre-trained StyleGAN
by: Bedychaj, Andrzej, et al.
Published: (2024)
by: Bedychaj, Andrzej, et al.
Published: (2024)
Generative Pre-trained Autoregressive Diffusion Transformer
by: Zhang, Yuan, et al.
Published: (2025)
by: Zhang, Yuan, et al.
Published: (2025)
UniHOI: Unified Human-Object Interaction Understanding via Unified Token Space
by: Yang, Panqi, et al.
Published: (2025)
by: Yang, Panqi, et al.
Published: (2025)
ReCorD: Reasoning and Correcting Diffusion for HOI Generation
by: Jiang-Lin, Jian-Yu, et al.
Published: (2024)
by: Jiang-Lin, Jian-Yu, et al.
Published: (2024)
Precise localization of corneal reflections in eye images using deep learning trained on synthetic data
by: Byrne, Sean Anthony, et al.
Published: (2023)
by: Byrne, Sean Anthony, et al.
Published: (2023)
One-to-many Reconstruction of 3D Geometry of cultural Artifacts using a synthetically trained Generative Model
by: Pöllabauer, Thomas, et al.
Published: (2024)
by: Pöllabauer, Thomas, et al.
Published: (2024)
PhyGenHOI: Physically-Aware 4D Generation of Dynamic Human-Object Interactions
by: Benishu, Omer, et al.
Published: (2026)
by: Benishu, Omer, et al.
Published: (2026)
Lightweight, Pre-trained Transformers for Remote Sensing Timeseries
by: Tseng, Gabriel, et al.
Published: (2023)
by: Tseng, Gabriel, et al.
Published: (2023)
Towards Accurate Post-training Quantization for Reparameterized Models
by: Zhang, Luoming, et al.
Published: (2024)
by: Zhang, Luoming, et al.
Published: (2024)
QVD: Post-training Quantization for Video Diffusion Models
by: Tian, Shilong, et al.
Published: (2024)
by: Tian, Shilong, et al.
Published: (2024)
An Empirical Study of Autoregressive Pre-training from Videos
by: Rajasegaran, Jathushan, et al.
Published: (2025)
by: Rajasegaran, Jathushan, et al.
Published: (2025)
Thinker: A vision-language foundation model for embodied intelligence
by: Pan, Baiyu, et al.
Published: (2026)
by: Pan, Baiyu, et al.
Published: (2026)
Self-supervised pre-training with diffusion model for few-shot landmark detection in x-ray images
by: Di Via, Roberto, et al.
Published: (2024)
by: Di Via, Roberto, et al.
Published: (2024)
MoMBS: Mixed-order minibatch sampling enhances model training from diverse-quality images
by: Li, Han, et al.
Published: (2025)
by: Li, Han, et al.
Published: (2025)
FFF: Fixing Flawed Foundations in contrastive pre-training results in very strong Vision-Language models
by: Bulat, Adrian, et al.
Published: (2024)
by: Bulat, Adrian, et al.
Published: (2024)
Similar Items
-
A multimodal vision foundation model for generalizable knee pathology
by: Yu, Kang, et al.
Published: (2026) -
Transformers self-organize like newborn visual systems when trained in prenatal worlds
by: Pandey, Lalit, et al.
Published: (2026) -
Avoid Wasted Annotation Costs in Open-set Active Learning with Pre-trained Vision-Language Model
by: Heo, Jaehyuk, et al.
Published: (2024) -
EyeCLIP: A visual-language foundation model for multi-modal ophthalmic image analysis
by: Shi, Danli, et al.
Published: (2024) -
ActiveMark: on watermarking of visual foundation models via massive activations
by: Chistyakova, Anna, et al.
Published: (2025)