Efficient Image Pre-Training with Siamese Cropped Masked Autoencoders
Fuente:
arXiv
Guardado en:
| Autores principales: | Eymaël, Alexandre, Vandeghen, Renaud, Cioppa, Anthony, Giancola, Silvio, Ghanem, Bernard, Van Droogenbroeck, Marc |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Adaptive Self-Training for Object Detection
por: Vandeghen, Renaud, et al.
Publicado: (2022)
por: Vandeghen, Renaud, et al.
Publicado: (2022)
Action Anticipation from SoccerNet Football Video Broadcasts
por: Dalal, Mohamad, et al.
Publicado: (2025)
por: Dalal, Mohamad, et al.
Publicado: (2025)
AGOP as Explanation: From Feature Learning to Per-Sample Attribution in Image Classifiers
por: Katakam, Raj Kiran Gupta
Publicado: (2026)
por: Katakam, Raj Kiran Gupta
Publicado: (2026)
Dream to Fly: Model-Based Reinforcement Learning for Vision-Based Drone Flight
por: Romero, Angel, et al.
Publicado: (2025)
por: Romero, Angel, et al.
Publicado: (2025)
WayFASTER: a Self-Supervised Traversability Prediction for Increased Navigation Awareness
por: Gasparino, Mateus Valverde, et al.
Publicado: (2024)
por: Gasparino, Mateus Valverde, et al.
Publicado: (2024)
SoccerLens: Grounded Soccer Video Understanding Beyond Accuracy
por: Elsharkawi, Ismael, et al.
Publicado: (2026)
por: Elsharkawi, Ismael, et al.
Publicado: (2026)
EduFlow: Advancing MLLMs' Problem-Solving Proficiency through Multi-Stage, Multi-Perspective Critique
por: Zhu, Chenglin, et al.
Publicado: (2025)
por: Zhu, Chenglin, et al.
Publicado: (2025)
Lifelong Learning in Vision-Language Models: Enhanced EWC with Cross-Modal Knowledge Retention
por: Durrani, Hamza Ahmed, et al.
Publicado: (2026)
por: Durrani, Hamza Ahmed, et al.
Publicado: (2026)
SUN Team's Contribution to ABAW 2024 Competition: Audio-visual Valence-Arousal Estimation and Expression Recognition
por: Dresvyanskiy, Denis, et al.
Publicado: (2024)
por: Dresvyanskiy, Denis, et al.
Publicado: (2024)
Learning the meanings of function words from grounded language using a visual question answering model
por: Portelance, Eva, et al.
Publicado: (2023)
por: Portelance, Eva, et al.
Publicado: (2023)
Closed-Loop Neural Activation Control in Vision-Language-Action Models
por: Babu, Abhijith, et al.
Publicado: (2026)
por: Babu, Abhijith, et al.
Publicado: (2026)
Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning
por: Yang, Shan
Publicado: (2026)
por: Yang, Shan
Publicado: (2026)
Mitigating Catastrophic Forgetting in the Incremental Learning of Medical Images
por: Yavari, Sara, et al.
Publicado: (2025)
por: Yavari, Sara, et al.
Publicado: (2025)
Cortex 2.0: Grounding World Models in Real-World Industrial Deployment
por: Aida, Adriana, et al.
Publicado: (2026)
por: Aida, Adriana, et al.
Publicado: (2026)
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
por: Raoufi, Behnam, et al.
Publicado: (2025)
por: Raoufi, Behnam, et al.
Publicado: (2025)
Better Schedules for Low Precision Training of Deep Neural Networks
por: Wolfe, Cameron R., et al.
Publicado: (2024)
por: Wolfe, Cameron R., et al.
Publicado: (2024)
MaSC: A Masked Similarity Metric for Evaluating Concept-Driven Generation
por: Bartkowiak, Patryk, et al.
Publicado: (2026)
por: Bartkowiak, Patryk, et al.
Publicado: (2026)
CLARE: Continual Learning for Vision-Language-Action Models via Autonomous Adapter Routing and Expansion
por: Römer, Ralf, et al.
Publicado: (2026)
por: Römer, Ralf, et al.
Publicado: (2026)
ExpReS-VLA: Specializing Vision-Language-Action Models Through Experience Replay and Retrieval
por: Syed, Shahram Najam, et al.
Publicado: (2025)
por: Syed, Shahram Najam, et al.
Publicado: (2025)
Once-For-All: A Train-Once and Select-Anytime Framework for Multimodal Instruction Tuning
por: Dong, Mingkang, et al.
Publicado: (2026)
por: Dong, Mingkang, et al.
Publicado: (2026)
Exposing Blindspots: Cultural Bias Evaluation in Generative Image Models
por: Seo, Huichan, et al.
Publicado: (2025)
por: Seo, Huichan, et al.
Publicado: (2025)
Physics-Informed Spectral Modeling for Hyperspectral Imaging
por: Gawrysiak, Zuzanna, et al.
Publicado: (2025)
por: Gawrysiak, Zuzanna, et al.
Publicado: (2025)
From Demonstrations to Safe Deployment: Path-Consistent Safety Filtering for Diffusion Policies
por: Römer, Ralf, et al.
Publicado: (2025)
por: Römer, Ralf, et al.
Publicado: (2025)
CrossVLA: Cross-Paradigm Post-Training and Inference Optimization for Vision-Language-Action Models
por: Liu, Zhi
Publicado: (2026)
por: Liu, Zhi
Publicado: (2026)
Curb Your Attention: Causal Attention Gating for Robust Trajectory Prediction in Autonomous Driving
por: Ahmadi, Ehsan, et al.
Publicado: (2024)
por: Ahmadi, Ehsan, et al.
Publicado: (2024)
ReLKD: Inter-Class Relation Learning with Knowledge Distillation for Generalized Category Discovery
por: Zhou, Fang, et al.
Publicado: (2025)
por: Zhou, Fang, et al.
Publicado: (2025)
Reducing the Sensitivity of Neural Physics Simulators to Mesh Topology via Pretraining
por: Vaska, Nathan, et al.
Publicado: (2025)
por: Vaska, Nathan, et al.
Publicado: (2025)
Data Organization Matters in Multimodal Instruction Tuning: A Controlled Study of Capability Trade-offs
por: Tang, Guowei
Publicado: (2026)
por: Tang, Guowei
Publicado: (2026)
RDPO: Real Data Preference Optimization for Physics Consistency Video Generation
por: Qian, Wenxu, et al.
Publicado: (2025)
por: Qian, Wenxu, et al.
Publicado: (2025)
CrystalDiT: A Diffusion Transformer for Crystal Generation
por: Yi, Xiaohan, et al.
Publicado: (2025)
por: Yi, Xiaohan, et al.
Publicado: (2025)
Survey Transfer Learning: Recycling Data with Silicon Responses
por: Amini, Ali
Publicado: (2025)
por: Amini, Ali
Publicado: (2025)
Reframing linguistic bootstrapping as joint inference using visually-grounded grammar induction models
por: Portelance, Eva, et al.
Publicado: (2024)
por: Portelance, Eva, et al.
Publicado: (2024)
eStonefish-Scenes: A Sim-to-Real Validated and Robot-Centric Event-based Optical Flow Dataset for Underwater Vehicles
por: Mansour, Jad, et al.
Publicado: (2025)
por: Mansour, Jad, et al.
Publicado: (2025)
eCARLA-scenes: A synthetically generated dataset for event-based optical flow prediction
por: Mansour, Jad, et al.
Publicado: (2024)
por: Mansour, Jad, et al.
Publicado: (2024)
Engineering an Efficient Object Tracker for Non-Linear Motion
por: Adžemović, Momir, et al.
Publicado: (2024)
por: Adžemović, Momir, et al.
Publicado: (2024)
Physics-informed Variational Autoencoders for Improved Robustness to Environmental Factors of Variation
por: Thoreau, Romain, et al.
Publicado: (2022)
por: Thoreau, Romain, et al.
Publicado: (2022)
EvoPrune: Early-Stage Visual Token Pruning for Efficient MLLMs
por: Chen, Yuhao, et al.
Publicado: (2026)
por: Chen, Yuhao, et al.
Publicado: (2026)
Invariant Representation via Decoupling Style and Spurious Features from Images
por: Li, Ruimeng, et al.
Publicado: (2023)
por: Li, Ruimeng, et al.
Publicado: (2023)
FrescoDiffusion: 4K Image-to-Video with Prior-Regularized Tiled Diffusion
por: Caselles-Dupré, Hugo, et al.
Publicado: (2026)
por: Caselles-Dupré, Hugo, et al.
Publicado: (2026)
Training a Student Expert via Semi-Supervised Foundation Model Distillation
por: Taghavi, Pardis, et al.
Publicado: (2026)
por: Taghavi, Pardis, et al.
Publicado: (2026)
Ejemplares similares
-
Adaptive Self-Training for Object Detection
por: Vandeghen, Renaud, et al.
Publicado: (2022) -
Action Anticipation from SoccerNet Football Video Broadcasts
por: Dalal, Mohamad, et al.
Publicado: (2025) -
AGOP as Explanation: From Feature Learning to Per-Sample Attribution in Image Classifiers
por: Katakam, Raj Kiran Gupta
Publicado: (2026) -
Dream to Fly: Model-Based Reinforcement Learning for Vision-Based Drone Flight
por: Romero, Angel, et al.
Publicado: (2025) -
WayFASTER: a Self-Supervised Traversability Prediction for Increased Navigation Awareness
por: Gasparino, Mateus Valverde, et al.
Publicado: (2024)