Ego: Embedding-Guided Personalization of Vision-Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Seifi, Soroush, Gardier, Simon, Dorovatas, Vaggelis, Reino, Daniel Olmeda, Aljundi, Rahaf |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Personalization Toolkit: Training Free Personalization of Large Vision Language Models
von: Seifi, Soroush, et al.
Veröffentlicht: (2025)
von: Seifi, Soroush, et al.
Veröffentlicht: (2025)
Recurrent Attention-based Token Selection for Efficient Streaming Video-LLMs
von: Dorovatas, Vaggelis, et al.
Veröffentlicht: (2025)
von: Dorovatas, Vaggelis, et al.
Veröffentlicht: (2025)
Annotation Free Semantic Segmentation with Vision Foundation Models
von: Seifi, Soroush, et al.
Veröffentlicht: (2024)
von: Seifi, Soroush, et al.
Veröffentlicht: (2024)
Efficient Few-Shot Continual Learning in Vision-Language Models
von: Panos, Aristeidis, et al.
Veröffentlicht: (2025)
von: Panos, Aristeidis, et al.
Veröffentlicht: (2025)
Online In-Context Distillation for Low-Resource Vision Language Models
von: Kang, Zhiqi, et al.
Veröffentlicht: (2025)
von: Kang, Zhiqi, et al.
Veröffentlicht: (2025)
Imperfect Vision Encoders: Efficient and Robust Tuning for Vision-Language Models
von: Panos, Aristeidis, et al.
Veröffentlicht: (2024)
von: Panos, Aristeidis, et al.
Veröffentlicht: (2024)
SEM: Sparse Embedding Modulation for Post-Hoc Debiasing of Vision-Language Models
von: Guimard, Quentin, et al.
Veröffentlicht: (2026)
von: Guimard, Quentin, et al.
Veröffentlicht: (2026)
First Session Adaptation: A Strong Replay-Free Baseline for Class-Incremental Learning
von: Panos, Aristeidis, et al.
Veröffentlicht: (2023)
von: Panos, Aristeidis, et al.
Veröffentlicht: (2023)
Memo: Training Memory-Efficient Embodied Agents with Reinforcement Learning
von: Gupta, Gunshi, et al.
Veröffentlicht: (2025)
von: Gupta, Gunshi, et al.
Veröffentlicht: (2025)
The Phantom Menace: Unmasking Privacy Leakages in Vision-Language Models
von: Caldarella, Simone, et al.
Veröffentlicht: (2024)
von: Caldarella, Simone, et al.
Veröffentlicht: (2024)
Evaluating the Impact of Post-Training Quantization on Reliable VQA with Multimodal LLMs
von: Kurz, Paul Jonas, et al.
Veröffentlicht: (2026)
von: Kurz, Paul Jonas, et al.
Veröffentlicht: (2026)
Fusion Embedding for Pose-Guided Person Image Synthesis with Diffusion Model
von: Lee, Donghwna, et al.
Veröffentlicht: (2024)
von: Lee, Donghwna, et al.
Veröffentlicht: (2024)
EgoSelf: From Memory to Personalized Egocentric Assistant
von: Wang, Yanshuo, et al.
Veröffentlicht: (2026)
von: Wang, Yanshuo, et al.
Veröffentlicht: (2026)
Language-Guided Invariance Probing of Vision-Language Models
von: Lee, Jae Joong
Veröffentlicht: (2025)
von: Lee, Jae Joong
Veröffentlicht: (2025)
Incremental Object-Based Novelty Detection with Feedback Loop
von: Caldarella, Simone, et al.
Veröffentlicht: (2023)
von: Caldarella, Simone, et al.
Veröffentlicht: (2023)
Semantic Compositions Enhance Vision-Language Contrastive Learning
von: Aladago, Maxwell, et al.
Veröffentlicht: (2024)
von: Aladago, Maxwell, et al.
Veröffentlicht: (2024)
Ego-Grounding for Personalized Question-Answering in Egocentric Videos
von: Xiao, Junbin, et al.
Veröffentlicht: (2026)
von: Xiao, Junbin, et al.
Veröffentlicht: (2026)
EgoPro-Bench: Benchmarking Personalized Proactive Interaction in Egocentric Video Streams
von: Ran, Dongchuan, et al.
Veröffentlicht: (2026)
von: Ran, Dongchuan, et al.
Veröffentlicht: (2026)
Are Large Vision-Language Models Ready to Guide Blind and Low-Vision Individuals?
von: Kim, Eunki, et al.
Veröffentlicht: (2025)
von: Kim, Eunki, et al.
Veröffentlicht: (2025)
Gate-and-Merge: Zero-shot Compositional Personalization of Vision Language Models
von: Ding, Guodong, et al.
Veröffentlicht: (2026)
von: Ding, Guodong, et al.
Veröffentlicht: (2026)
MC-LLaVA: Multi-Concept Personalized Vision-Language Model
von: An, Ruichuan, et al.
Veröffentlicht: (2024)
von: An, Ruichuan, et al.
Veröffentlicht: (2024)
MC-LLaVA: Multi-Concept Personalized Vision-Language Model
von: An, Ruichuan, et al.
Veröffentlicht: (2025)
von: An, Ruichuan, et al.
Veröffentlicht: (2025)
TrajMamba: An Ego-Motion-Guided Mamba Model for Pedestrian Trajectory Prediction from an Egocentric Perspective
von: Peng, Yusheng, et al.
Veröffentlicht: (2026)
von: Peng, Yusheng, et al.
Veröffentlicht: (2026)
YoChameleon: Personalized Vision and Language Generation
von: Nguyen, Thao, et al.
Veröffentlicht: (2025)
von: Nguyen, Thao, et al.
Veröffentlicht: (2025)
EgoNight: Towards Egocentric Vision Understanding at Night with a Challenging Benchmark
von: Zhang, Deheng, et al.
Veröffentlicht: (2025)
von: Zhang, Deheng, et al.
Veröffentlicht: (2025)
HCQA @ Ego4D EgoSchema Challenge 2024
von: Zhang, Haoyu, et al.
Veröffentlicht: (2024)
von: Zhang, Haoyu, et al.
Veröffentlicht: (2024)
Cure or Poison? Embedding Instructions Visually Alters Hallucination in Vision-Language Models
von: Wang, Zhaochen, et al.
Veröffentlicht: (2025)
von: Wang, Zhaochen, et al.
Veröffentlicht: (2025)
Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives
von: Grauman, Kristen, et al.
Veröffentlicht: (2023)
von: Grauman, Kristen, et al.
Veröffentlicht: (2023)
PPU-Bench:Real World Benchmark for Personalized Partial Unlearning in Vision Language Models
von: Guang, Jiahui, et al.
Veröffentlicht: (2026)
von: Guang, Jiahui, et al.
Veröffentlicht: (2026)
HCQA-1.5 @ Ego4D EgoSchema Challenge 2025
von: Zhang, Haoyu, et al.
Veröffentlicht: (2025)
von: Zhang, Haoyu, et al.
Veröffentlicht: (2025)
EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos
von: Yang, Ruihan, et al.
Veröffentlicht: (2025)
von: Yang, Ruihan, et al.
Veröffentlicht: (2025)
ATA: Bridging Implicit Reasoning with Attention-Guided and Action-Guided Inference for Vision-Language Action Models
von: Yang, Cheng, et al.
Veröffentlicht: (2026)
von: Yang, Cheng, et al.
Veröffentlicht: (2026)
Non-Contrastive Vision-Language Learning with Predictive Embedding Alignment
von: Kuhn, Lukas, et al.
Veröffentlicht: (2026)
von: Kuhn, Lukas, et al.
Veröffentlicht: (2026)
Locatability-Guided Adaptive Reasoning for Image Geo-Localization with Vision-Language Models
von: Yu, Bo, et al.
Veröffentlicht: (2026)
von: Yu, Bo, et al.
Veröffentlicht: (2026)
On the Utility of Foundation Models for Fast MRI: Vision-Language-Guided Image Reconstruction
von: Feng, Ruimin, et al.
Veröffentlicht: (2025)
von: Feng, Ruimin, et al.
Veröffentlicht: (2025)
Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning
von: Zhan, Yufei, et al.
Veröffentlicht: (2025)
von: Zhan, Yufei, et al.
Veröffentlicht: (2025)
EgoCross: Benchmarking Multimodal Large Language Models for Cross-Domain Egocentric Video Question Answering
von: Li, Yanjun, et al.
Veröffentlicht: (2025)
von: Li, Yanjun, et al.
Veröffentlicht: (2025)
V-DPO: Mitigating Hallucination in Large Vision Language Models via Vision-Guided Direct Preference Optimization
von: Xie, Yuxi, et al.
Veröffentlicht: (2024)
von: Xie, Yuxi, et al.
Veröffentlicht: (2024)
CLAY: Conditional Visual Similarity Modulation in Vision-Language Embedding Space
von: Lim, Sohwi, et al.
Veröffentlicht: (2026)
von: Lim, Sohwi, et al.
Veröffentlicht: (2026)
Procedural Knowledge Extraction from Industrial Troubleshooting Guides Using Vision Language Models
von: de Avalle, Guillermo Gil, et al.
Veröffentlicht: (2026)
von: de Avalle, Guillermo Gil, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Personalization Toolkit: Training Free Personalization of Large Vision Language Models
von: Seifi, Soroush, et al.
Veröffentlicht: (2025) -
Recurrent Attention-based Token Selection for Efficient Streaming Video-LLMs
von: Dorovatas, Vaggelis, et al.
Veröffentlicht: (2025) -
Annotation Free Semantic Segmentation with Vision Foundation Models
von: Seifi, Soroush, et al.
Veröffentlicht: (2024) -
Efficient Few-Shot Continual Learning in Vision-Language Models
von: Panos, Aristeidis, et al.
Veröffentlicht: (2025) -
Online In-Context Distillation for Low-Resource Vision Language Models
von: Kang, Zhiqi, et al.
Veröffentlicht: (2025)