Memory-Modular Classification: Learning to Generalize with Memory Replacement
Fuente:
arXiv
Salvato in:
| Autori principali: | Kang, Dahyun, Iscen, Ahmet, Jo, Eunchan, Choi, Sua, Cho, Minsu, Schmid, Cordelia |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Contrastive Mean-Shift Learning for Generalized Category Discovery
di: Choi, Sua, et al.
Pubblicazione: (2024)
di: Choi, Sua, et al.
Pubblicazione: (2024)
Few-Shot Pattern Detection via Template Matching and Regression
di: Jo, Eunchan, et al.
Pubblicazione: (2025)
di: Jo, Eunchan, et al.
Pubblicazione: (2025)
A Generative Approach for Wikipedia-Scale Visual Entity Recognition
di: Caron, Mathilde, et al.
Pubblicazione: (2024)
di: Caron, Mathilde, et al.
Pubblicazione: (2024)
Web-Scale Visual Entity Recognition: An LLM-Driven Data Approach
di: Caron, Mathilde, et al.
Pubblicazione: (2024)
di: Caron, Mathilde, et al.
Pubblicazione: (2024)
Retrieval-Enhanced Contrastive Vision-Text Models
di: Iscen, Ahmet, et al.
Pubblicazione: (2023)
di: Iscen, Ahmet, et al.
Pubblicazione: (2023)
In Defense of Lazy Visual Grounding for Open-Vocabulary Semantic Segmentation
di: Kang, Dahyun, et al.
Pubblicazione: (2024)
di: Kang, Dahyun, et al.
Pubblicazione: (2024)
Learning Correlation Structures for Vision Transformers
di: Kim, Manjin, et al.
Pubblicazione: (2024)
di: Kim, Manjin, et al.
Pubblicazione: (2024)
Affogato: Learning Open-Vocabulary Affordance Grounding with Automated Data Generation at Scale
di: Lee, Junha, et al.
Pubblicazione: (2025)
di: Lee, Junha, et al.
Pubblicazione: (2025)
MoReVQA: Exploring Modular Reasoning Models for Video Question Answering
di: Min, Juhong, et al.
Pubblicazione: (2024)
di: Min, Juhong, et al.
Pubblicazione: (2024)
AMES: Asymmetric and Memory-Efficient Similarity Estimation for Instance-level Retrieval
di: Suma, Pavel, et al.
Pubblicazione: (2024)
di: Suma, Pavel, et al.
Pubblicazione: (2024)
CAViAR: Critic-Augmented Video Agentic Reasoning
di: Menon, Sachit, et al.
Pubblicazione: (2025)
di: Menon, Sachit, et al.
Pubblicazione: (2025)
Continual Learning in Vision-Language Models via Aligned Model Merging
di: Sokar, Ghada, et al.
Pubblicazione: (2025)
di: Sokar, Ghada, et al.
Pubblicazione: (2025)
Time-, Memory- and Parameter-Efficient Visual Adaptation
di: Mercea, Otniel-Bogdan, et al.
Pubblicazione: (2024)
di: Mercea, Otniel-Bogdan, et al.
Pubblicazione: (2024)
Online Temporal Action Localization with Memory-Augmented Transformer
di: Song, Youngkil, et al.
Pubblicazione: (2024)
di: Song, Youngkil, et al.
Pubblicazione: (2024)
BrickNet: Graph-Backed Generative Brick Assembly
di: Kulits, Peter, et al.
Pubblicazione: (2026)
di: Kulits, Peter, et al.
Pubblicazione: (2026)
Grounded Video Caption Generation
di: Kazakos, Evangelos, et al.
Pubblicazione: (2024)
di: Kazakos, Evangelos, et al.
Pubblicazione: (2024)
ComposeAnything: Composite Object Priors for Text-to-Image Generation
di: Khan, Zeeshan, et al.
Pubblicazione: (2025)
di: Khan, Zeeshan, et al.
Pubblicazione: (2025)
Large-scale Pre-training for Grounded Video Caption Generation
di: Kazakos, Evangelos, et al.
Pubblicazione: (2025)
di: Kazakos, Evangelos, et al.
Pubblicazione: (2025)
Learning text-to-video retrieval from image captioning
di: Ventura, Lucas, et al.
Pubblicazione: (2024)
di: Ventura, Lucas, et al.
Pubblicazione: (2024)
SceneCraft: An LLM Agent for Synthesizing 3D Scene as Blender Code
di: Hu, Ziniu, et al.
Pubblicazione: (2024)
di: Hu, Ziniu, et al.
Pubblicazione: (2024)
VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation
di: Bousselham, Walid, et al.
Pubblicazione: (2025)
di: Bousselham, Walid, et al.
Pubblicazione: (2025)
Space-Time Forecasting of Dynamic Scenes with Motion-aware Gaussian Grouping
di: Lee, Junmyeong, et al.
Pubblicazione: (2026)
di: Lee, Junmyeong, et al.
Pubblicazione: (2026)
LoFT: LoRA-fused Training Dataset Generation with Few-shot Guidance
di: Kim, Jae Myung, et al.
Pubblicazione: (2025)
di: Kim, Jae Myung, et al.
Pubblicazione: (2025)
Dense Optical Tracking: Connecting the Dots
di: Moing, Guillaume Le, et al.
Pubblicazione: (2023)
di: Moing, Guillaume Le, et al.
Pubblicazione: (2023)
PointACT: Vision-Language-Action Models with Multi-Scale Point-Action Interaction
di: Chen, Shizhe, et al.
Pubblicazione: (2026)
di: Chen, Shizhe, et al.
Pubblicazione: (2026)
Towards Generalizable Vision-Language Robotic Manipulation: A Benchmark and LLM-guided 3D Policy
di: Garcia, Ricardo, et al.
Pubblicazione: (2024)
di: Garcia, Ricardo, et al.
Pubblicazione: (2024)
Harnessing the Power of Training-Free Techniques in Text-to-2D Generation for Text-to-3D Generation via Score Distillation Sampling
di: Lee, Junhong, et al.
Pubblicazione: (2025)
di: Lee, Junhong, et al.
Pubblicazione: (2025)
Affostruction: 3D Affordance Grounding with Generative Reconstruction
di: Park, Chunghyun, et al.
Pubblicazione: (2026)
di: Park, Chunghyun, et al.
Pubblicazione: (2026)
Tempered Self-Similarity Alignment for Physically Plausible Video Generation
di: Kim, Manjin, et al.
Pubblicazione: (2026)
di: Kim, Manjin, et al.
Pubblicazione: (2026)
MINERVA: Evaluating Complex Video Reasoning
di: Nagrani, Arsha, et al.
Pubblicazione: (2025)
di: Nagrani, Arsha, et al.
Pubblicazione: (2025)
Improving Text-to-Image Generation with Intrinsic Self-Confidence Rewards
di: Kim, Seungwook, et al.
Pubblicazione: (2026)
di: Kim, Seungwook, et al.
Pubblicazione: (2026)
RoDyGS: Robust Dynamic Gaussian Splatting for Casual Videos
di: Lee, Junmyeong, et al.
Pubblicazione: (2024)
di: Lee, Junmyeong, et al.
Pubblicazione: (2024)
Similarity-Aware Selective State-Space Modeling for Semantic Correspondence
di: Kim, Seungwook, et al.
Pubblicazione: (2025)
di: Kim, Seungwook, et al.
Pubblicazione: (2025)
Leveraging 3D Geometric Priors in 2D Rotation Symmetry Detection
di: Seo, Ahyun, et al.
Pubblicazione: (2025)
di: Seo, Ahyun, et al.
Pubblicazione: (2025)
RAVEN: Rethinking Adversarial Video Generation with Efficient Tri-plane Networks
di: Ghosh, Partha, et al.
Pubblicazione: (2024)
di: Ghosh, Partha, et al.
Pubblicazione: (2024)
Chapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMs
di: Ventura, Lucas, et al.
Pubblicazione: (2025)
di: Ventura, Lucas, et al.
Pubblicazione: (2025)
OVFact: Measuring and Improving Open-Vocabulary Factuality for Long Caption Models
di: Wysoczańska, Monika, et al.
Pubblicazione: (2025)
di: Wysoczańska, Monika, et al.
Pubblicazione: (2025)
SUGAR: Pre-training 3D Visual Representations for Robotics
di: Chen, Shizhe, et al.
Pubblicazione: (2024)
di: Chen, Shizhe, et al.
Pubblicazione: (2024)
Dense Video Object Captioning from Disjoint Supervision
di: Zhou, Xingyi, et al.
Pubblicazione: (2023)
di: Zhou, Xingyi, et al.
Pubblicazione: (2023)
CoVR-2: Automatic Data Construction for Composed Video Retrieval
di: Ventura, Lucas, et al.
Pubblicazione: (2023)
di: Ventura, Lucas, et al.
Pubblicazione: (2023)
Documenti analoghi
-
Contrastive Mean-Shift Learning for Generalized Category Discovery
di: Choi, Sua, et al.
Pubblicazione: (2024) -
Few-Shot Pattern Detection via Template Matching and Regression
di: Jo, Eunchan, et al.
Pubblicazione: (2025) -
A Generative Approach for Wikipedia-Scale Visual Entity Recognition
di: Caron, Mathilde, et al.
Pubblicazione: (2024) -
Web-Scale Visual Entity Recognition: An LLM-Driven Data Approach
di: Caron, Mathilde, et al.
Pubblicazione: (2024) -
Retrieval-Enhanced Contrastive Vision-Text Models
di: Iscen, Ahmet, et al.
Pubblicazione: (2023)