Efficient Object-centric Representation Learning with Pre-trained Geometric Prior
Fuente:
arXiv
Saved in:
| Main Authors: | Khac, Phúc H. Le, Healy, Graham, Smeaton, Alan F. |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Using Saliency and Cropping to Improve Video Memorability
by: Mudgal, Vaibhav, et al.
Published: (2023)
by: Mudgal, Vaibhav, et al.
Published: (2023)
Dynamic Self-adaptive Multiscale Distillation from Pre-trained Multimodal Large Model for Efficient Cross-modal Representation Learning
by: Liang, Zhengyang, et al.
Published: (2024)
by: Liang, Zhengyang, et al.
Published: (2024)
Ego3DT: Tracking Every 3D Object in Ego-centric Videos
by: Hao, Shengyu, et al.
Published: (2024)
by: Hao, Shengyu, et al.
Published: (2024)
Generalized Face Forgery Detection via Adaptive Learning for Pre-trained Vision Transformer
by: Luo, Anwei, et al.
Published: (2023)
by: Luo, Anwei, et al.
Published: (2023)
Reinforcing Pre-trained Models Using Counterfactual Images
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
Modularized Zero-shot VQA with Pre-trained Models
by: Cao, Rui, et al.
Published: (2023)
by: Cao, Rui, et al.
Published: (2023)
Scaling up Multimodal Pre-training for Sign Language Understanding
by: Zhou, Wengang, et al.
Published: (2024)
by: Zhou, Wengang, et al.
Published: (2024)
Improving Text-guided Object Inpainting with Semantic Pre-inpainting
by: Chen, Yifu, et al.
Published: (2024)
by: Chen, Yifu, et al.
Published: (2024)
Pre-training CLIP against Data Poisoning with Optimal Transport-based Matching and Alignment
by: Zhang, Tong, et al.
Published: (2025)
by: Zhang, Tong, et al.
Published: (2025)
Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training
by: Han, Junlin, et al.
Published: (2025)
by: Han, Junlin, et al.
Published: (2025)
StableDub: Taming Diffusion Prior for Generalized and Efficient Visual Dubbing
by: Chen, Liyang, et al.
Published: (2025)
by: Chen, Liyang, et al.
Published: (2025)
SNP-S3: Shared Network Pre-training and Significant Semantic Strengthening for Various Video-Text Tasks
by: Dong, Xingning, et al.
Published: (2024)
by: Dong, Xingning, et al.
Published: (2024)
REArtGS: Reconstructing and Generating Articulated Objects via 3D Gaussian Splatting with Geometric and Motion Constraints
by: Wu, Di, et al.
Published: (2025)
by: Wu, Di, et al.
Published: (2025)
Advancing 3D Scene Understanding with MV-ScanQA Multi-View Reasoning Evaluation and TripAlign Pre-training Dataset
by: Mo, Wentao, et al.
Published: (2025)
by: Mo, Wentao, et al.
Published: (2025)
CoPRS: Learning Positional Prior from Chain-of-Thought for Reasoning Segmentation
by: Lu, Zhenyu, et al.
Published: (2025)
by: Lu, Zhenyu, et al.
Published: (2025)
TRIP: Temporal Residual Learning with Image Noise Prior for Image-to-Video Diffusion Models
by: Zhang, Zhongwei, et al.
Published: (2024)
by: Zhang, Zhongwei, et al.
Published: (2024)
3D-LMVIC: Learning-based Multi-View Image Coding with 3D Gaussian Geometric Priors
by: Huang, Yujun, et al.
Published: (2024)
by: Huang, Yujun, et al.
Published: (2024)
HMPE:HeatMap Embedding for Efficient Transformer-Based Small Object Detection
by: Zeng, YangChen
Published: (2025)
by: Zeng, YangChen
Published: (2025)
Efficiently Collecting Training Dataset for 2D Object Detection by Online Visual Feedback
by: Kiyokawa, Takuya, et al.
Published: (2023)
by: Kiyokawa, Takuya, et al.
Published: (2023)
Learning Brain Representation with Hierarchical Visual Embeddings
by: Zheng, Jiawen, et al.
Published: (2026)
by: Zheng, Jiawen, et al.
Published: (2026)
GSVC: Efficient Video Representation and Compression Through 2D Gaussian Splatting
by: Wang, Longan, et al.
Published: (2025)
by: Wang, Longan, et al.
Published: (2025)
Sketch and Patch: Efficient 3D Gaussian Representation for Man-Made Scenes
by: Shi, Yuang, et al.
Published: (2025)
by: Shi, Yuang, et al.
Published: (2025)
DeepSPG: Exploring Deep Semantic Prior Guidance for Low-light Image Enhancement with Multimodal Learning
by: Lu, Jialang, et al.
Published: (2025)
by: Lu, Jialang, et al.
Published: (2025)
Self-supervised Photographic Image Layout Representation Learning
by: Zhao, Zhaoran, et al.
Published: (2024)
by: Zhao, Zhaoran, et al.
Published: (2024)
Creatively Upscaling Images with Global-Regional Priors
by: Qian, Yurui, et al.
Published: (2025)
by: Qian, Yurui, et al.
Published: (2025)
UniScene: Multi-Camera Unified Pre-training via 3D Scene Reconstruction for Autonomous Driving
by: Min, Chen, et al.
Published: (2023)
by: Min, Chen, et al.
Published: (2023)
Rethinking Multi-view Representation Learning via Distilled Disentangling
by: Ke, Guanzhou, et al.
Published: (2024)
by: Ke, Guanzhou, et al.
Published: (2024)
MSLIQA: Enhancing Learning Representations for Image Quality Assessment through Multi-Scale Learning
by: Avanaki, Nasim Jamshidi, et al.
Published: (2024)
by: Avanaki, Nasim Jamshidi, et al.
Published: (2024)
RETRO: REthinking Tactile Representation Learning with Material PriOrs
by: Xia, Weihao, et al.
Published: (2025)
by: Xia, Weihao, et al.
Published: (2025)
Holistic Visual-Textual Sentiment Analysis with Prior Models
by: Chen, Junyu, et al.
Published: (2022)
by: Chen, Junyu, et al.
Published: (2022)
Self-similarity Prior Distillation for Unsupervised Remote Physiological Measurement
by: Zhang, Xinyu, et al.
Published: (2023)
by: Zhang, Xinyu, et al.
Published: (2023)
EVE: Efficient Vision-Language Pre-training with Masked Prediction and Modality-Aware MoE
by: Chen, Junyi, et al.
Published: (2023)
by: Chen, Junyi, et al.
Published: (2023)
Learning Gaussian Data Augmentation in Feature Space for One-shot Object Detection in Manga
by: Taniguchi, Takara, et al.
Published: (2024)
by: Taniguchi, Takara, et al.
Published: (2024)
Large-scale Multi-Modal Pre-trained Models: A Comprehensive Survey
by: Wang, Xiao, et al.
Published: (2023)
by: Wang, Xiao, et al.
Published: (2023)
Towards Unbiased Cross-Modal Representation Learning for Food Image-to-Recipe Retrieval
by: Wang, Qing, et al.
Published: (2025)
by: Wang, Qing, et al.
Published: (2025)
Disparity-based Stereo Image Compression with Aligned Cross-View Priors
by: Zhai, Yongqi, et al.
Published: (2022)
by: Zhai, Yongqi, et al.
Published: (2022)
Dual Mutual Learning Network with Global-local Awareness for RGB-D Salient Object Detection
by: Yi, Kang, et al.
Published: (2025)
by: Yi, Kang, et al.
Published: (2025)
UAU-Net: Uncertainty-aware Representation Learning and Evidential Classification for Facial Action Unit Detection
by: Li, Yuze, et al.
Published: (2026)
by: Li, Yuze, et al.
Published: (2026)
MISS: Memory-efficient Instance Segmentation Framework By Visual Inductive Priors Flow Propagation
by: Hsu, Chih-Chung, et al.
Published: (2024)
by: Hsu, Chih-Chung, et al.
Published: (2024)
Improving Adversarial Transferability of Vision-Language Pre-training Models through Collaborative Multimodal Interaction
by: Fu, Jiyuan, et al.
Published: (2024)
by: Fu, Jiyuan, et al.
Published: (2024)
Similar Items
-
Using Saliency and Cropping to Improve Video Memorability
by: Mudgal, Vaibhav, et al.
Published: (2023) -
Dynamic Self-adaptive Multiscale Distillation from Pre-trained Multimodal Large Model for Efficient Cross-modal Representation Learning
by: Liang, Zhengyang, et al.
Published: (2024) -
Ego3DT: Tracking Every 3D Object in Ego-centric Videos
by: Hao, Shengyu, et al.
Published: (2024) -
Generalized Face Forgery Detection via Adaptive Learning for Pre-trained Vision Transformer
by: Luo, Anwei, et al.
Published: (2023) -
Reinforcing Pre-trained Models Using Counterfactual Images
by: Li, Xiang, et al.
Published: (2024)