According to Me: Long-Term Personalized Referential Memory QA
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Mei, Jingbiao, Chen, Jinghong, Yang, Guangyu, Hou, Xinyu, Li, Margaret, Byrne, Bill |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Robust Adaptation of Large Multimodal Models for Retrieval Augmented Hateful Meme Detection
von: Mei, Jingbiao, et al.
Veröffentlicht: (2025)
von: Mei, Jingbiao, et al.
Veröffentlicht: (2025)
Improving Hateful Meme Detection through Retrieval-Guided Contrastive Learning
von: Mei, Jingbiao, et al.
Veröffentlicht: (2023)
von: Mei, Jingbiao, et al.
Veröffentlicht: (2023)
BERAG: Bayesian Ensemble Retrieval-Augmented Generation for Knowledge-based Visual Question Answering
von: Chen, Jinghong, et al.
Veröffentlicht: (2026)
von: Chen, Jinghong, et al.
Veröffentlicht: (2026)
Few-Shot VQA with Frozen LLMs: A Tale of Two Approaches
von: Sterner, Igor, et al.
Veröffentlicht: (2024)
von: Sterner, Igor, et al.
Veröffentlicht: (2024)
On Extending Direct Preference Optimization to Accommodate Ties
von: Chen, Jinghong, et al.
Veröffentlicht: (2024)
von: Chen, Jinghong, et al.
Veröffentlicht: (2024)
Retrieval-Augmented Defense: Adaptive and Controllable Jailbreak Prevention for Large Language Models
von: Yang, Guangyu, et al.
Veröffentlicht: (2025)
von: Yang, Guangyu, et al.
Veröffentlicht: (2025)
TeleMem: Building Long-Term and Multimodal Memory for Agentic AI
von: Chen, Chunliang, et al.
Veröffentlicht: (2025)
von: Chen, Chunliang, et al.
Veröffentlicht: (2025)
PreFLMR: Scaling Up Fine-Grained Late-Interaction Multi-modal Retrievers
von: Lin, Weizhe, et al.
Veröffentlicht: (2024)
von: Lin, Weizhe, et al.
Veröffentlicht: (2024)
Control-DAG: Constrained Decoding for Non-Autoregressive Directed Acyclic T5 using Weighted Finite State Automata
von: Chen, Jinghong, et al.
Veröffentlicht: (2024)
von: Chen, Jinghong, et al.
Veröffentlicht: (2024)
Moment Sampling in Video LLMs for Long-Form Video QA
von: Chasmai, Mustafa, et al.
Veröffentlicht: (2025)
von: Chasmai, Mustafa, et al.
Veröffentlicht: (2025)
Grounding Language in Multi-Perspective Referential Communication
von: Tang, Zineng, et al.
Veröffentlicht: (2024)
von: Tang, Zineng, et al.
Veröffentlicht: (2024)
PersonaVLM: Long-Term Personalized Multimodal LLMs
von: Nie, Chang, et al.
Veröffentlicht: (2026)
von: Nie, Chang, et al.
Veröffentlicht: (2026)
Benchmarking Deflection and Hallucination in Large Vision-Language Models
von: Moratelli, Nicholas, et al.
Veröffentlicht: (2026)
von: Moratelli, Nicholas, et al.
Veröffentlicht: (2026)
Referencing Where to Focus: Improving VisualGrounding with Referential Query
von: Wang, Yabing, et al.
Veröffentlicht: (2024)
von: Wang, Yabing, et al.
Veröffentlicht: (2024)
FunQA: Towards Surprising Video Comprehension
von: Xie, Binzhu, et al.
Veröffentlicht: (2023)
von: Xie, Binzhu, et al.
Veröffentlicht: (2023)
SRPO: Self-Referential Policy Optimization for Vision-Language-Action Models
von: Fei, Senyu, et al.
Veröffentlicht: (2025)
von: Fei, Senyu, et al.
Veröffentlicht: (2025)
3DLLM-Mem: Long-Term Spatial-Temporal Memory for Embodied 3D Large Language Model
von: Hu, Wenbo, et al.
Veröffentlicht: (2025)
von: Hu, Wenbo, et al.
Veröffentlicht: (2025)
PASK: Toward Intent-Aware Proactive Agents with Long-Term Memory
von: Xie, Zhifei, et al.
Veröffentlicht: (2026)
von: Xie, Zhifei, et al.
Veröffentlicht: (2026)
ReMeREC: Relation-aware and Multi-entity Referring Expression Comprehension
von: Hu, Yizhi, et al.
Veröffentlicht: (2025)
von: Hu, Yizhi, et al.
Veröffentlicht: (2025)
Catch Me If You Can Describe Me: Open-Vocabulary Camouflaged Instance Segmentation with Diffusion
von: Vu, Tuan-Anh, et al.
Veröffentlicht: (2023)
von: Vu, Tuan-Anh, et al.
Veröffentlicht: (2023)
EviMem: Evidence-Gap-Driven Iterative Retrieval for Long-Term Conversational Memory
von: Li, Yuyang, et al.
Veröffentlicht: (2026)
von: Li, Yuyang, et al.
Veröffentlicht: (2026)
ESG Accountability Made Easy: DocQA at Your Service
von: Mishra, Lokesh, et al.
Veröffentlicht: (2023)
von: Mishra, Lokesh, et al.
Veröffentlicht: (2023)
MimeQA: Towards Socially-Intelligent Nonverbal Foundation Models
von: Li, Hengzhi, et al.
Veröffentlicht: (2025)
von: Li, Hengzhi, et al.
Veröffentlicht: (2025)
DIAGRAMS: A Review Framework for Reasoning-Level Attribution in Diagram QA
von: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Veröffentlicht: (2026)
von: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Veröffentlicht: (2026)
Mem4Nav: Boosting Vision-and-Language Navigation in Urban Environments with a Hierarchical Spatial-Cognition Long-Short Memory System
von: He, Lixuan, et al.
Veröffentlicht: (2025)
von: He, Lixuan, et al.
Veröffentlicht: (2025)
Common Objects Out of Context (COOCo): Investigating Multimodal Context and Semantic Scene Violations in Referential Communication
von: Merlo, Filippo, et al.
Veröffentlicht: (2025)
von: Merlo, Filippo, et al.
Veröffentlicht: (2025)
WildVision: Evaluating Vision-Language Models in the Wild with Human Preferences
von: Lu, Yujie, et al.
Veröffentlicht: (2024)
von: Lu, Yujie, et al.
Veröffentlicht: (2024)
EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding
von: Wang, Ziyang, et al.
Veröffentlicht: (2026)
von: Wang, Ziyang, et al.
Veröffentlicht: (2026)
QA-Dragon: Query-Aware Dynamic RAG System for Knowledge-Intensive Visual Question Answering
von: Jiang, Zhuohang, et al.
Veröffentlicht: (2025)
von: Jiang, Zhuohang, et al.
Veröffentlicht: (2025)
Disentangled Representations for Short-Term and Long-Term Person Re-Identification
von: Eom, Chanho, et al.
Veröffentlicht: (2024)
von: Eom, Chanho, et al.
Veröffentlicht: (2024)
LongWriter-V: Enabling Ultra-Long and High-Fidelity Generation in Vision-Language Models
von: Tu, Shangqing, et al.
Veröffentlicht: (2025)
von: Tu, Shangqing, et al.
Veröffentlicht: (2025)
TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding
von: Ren, Shuhuai, et al.
Veröffentlicht: (2023)
von: Ren, Shuhuai, et al.
Veröffentlicht: (2023)
MGM-Omni: Scaling Omni LLMs to Personalized Long-Horizon Speech
von: Wang, Chengyao, et al.
Veröffentlicht: (2025)
von: Wang, Chengyao, et al.
Veröffentlicht: (2025)
SpatialMem: Metric-Aligned Long-Horizon Video Memory for Language Grounding and QA
von: Zheng, Xinyi, et al.
Veröffentlicht: (2026)
von: Zheng, Xinyi, et al.
Veröffentlicht: (2026)
MLLMReID: Multimodal Large Language Model-based Person Re-identification
von: Yang, Shan, et al.
Veröffentlicht: (2024)
von: Yang, Shan, et al.
Veröffentlicht: (2024)
LLMs Meet Long Video: Advancing Long Video Question Answering with An Interactive Visual Adapter in LLMs
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
von: Li, Yunxin, et al.
Veröffentlicht: (2024)
MeMOTR: Long-Term Memory-Augmented Transformer for Multi-Object Tracking
von: Gao, Ruopeng, et al.
Veröffentlicht: (2023)
von: Gao, Ruopeng, et al.
Veröffentlicht: (2023)
GameplayQA: A Benchmarking Framework for Decision-Dense POV-Synced Multi-Video Understanding of 3D Virtual Agents
von: Wang, Yunzhe, et al.
Veröffentlicht: (2026)
von: Wang, Yunzhe, et al.
Veröffentlicht: (2026)
MedXpertQA: Benchmarking Expert-Level Medical Reasoning and Understanding
von: Zuo, Yuxin, et al.
Veröffentlicht: (2025)
von: Zuo, Yuxin, et al.
Veröffentlicht: (2025)
Hierarchical Memory for Long Video QA
von: Wang, Yiqin, et al.
Veröffentlicht: (2024)
von: Wang, Yiqin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Robust Adaptation of Large Multimodal Models for Retrieval Augmented Hateful Meme Detection
von: Mei, Jingbiao, et al.
Veröffentlicht: (2025) -
Improving Hateful Meme Detection through Retrieval-Guided Contrastive Learning
von: Mei, Jingbiao, et al.
Veröffentlicht: (2023) -
BERAG: Bayesian Ensemble Retrieval-Augmented Generation for Knowledge-based Visual Question Answering
von: Chen, Jinghong, et al.
Veröffentlicht: (2026) -
Few-Shot VQA with Frozen LLMs: A Tale of Two Approaches
von: Sterner, Igor, et al.
Veröffentlicht: (2024) -
On Extending Direct Preference Optimization to Accommodate Ties
von: Chen, Jinghong, et al.
Veröffentlicht: (2024)