VisualLens: Personalization through Task-Agnostic Visual History
Fuente:
arXiv
Saved in:
| Main Authors: | Zhu, Wang Bill, Fu, Deqing, Sun, Kai, Lu, Yi, Lin, Zhaojiang, Moon, Seungwhan, Narang, Kanika, Canim, Mustafa, Liu, Yue, Kumar, Anuj, Dong, Xin Luna |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Proactive Assistant Dialogue Generation from Streaming Egocentric Videos
by: Zhang, Yichi, et al.
Published: (2025)
by: Zhang, Yichi, et al.
Published: (2025)
Pixel-Grounded Retrieval for Knowledgeable Large Multimodal Models
by: Kim, Jeonghwan, et al.
Published: (2026)
by: Kim, Jeonghwan, et al.
Published: (2026)
SnapNTell: Enhancing Entity-Centric Visual Question Answering with Retrieval Augmented Multimodal LLM
by: Qiu, Jielin, et al.
Published: (2024)
by: Qiu, Jielin, et al.
Published: (2024)
Large Language Models as Zero-shot Dialogue State Tracker through Function Calling
by: Li, Zekun, et al.
Published: (2024)
by: Li, Zekun, et al.
Published: (2024)
Toward Architecture-Agnostic Local Control of Posterior Collapse in VAEs
by: Song, Hyunsoo, et al.
Published: (2025)
by: Song, Hyunsoo, et al.
Published: (2025)
Stream RAG: Instant and Accurate Spoken Dialogue Systems with Streaming Tool Usage
by: Arora, Siddhant, et al.
Published: (2025)
by: Arora, Siddhant, et al.
Published: (2025)
SCRIBES: Web-Scale Script-Based Semi-Structured Data Extraction with Reinforcement Learning
by: Liu, Shicheng, et al.
Published: (2025)
by: Liu, Shicheng, et al.
Published: (2025)
Doppelgänger's Watch: A Split Objective Approach to Large Language Models
by: Ghasemlou, Shervin, et al.
Published: (2024)
by: Ghasemlou, Shervin, et al.
Published: (2024)
FocalLens: Visualizing Narratives through Focalization
by: Alam, S M Raihanul, et al.
Published: (2026)
by: Alam, S M Raihanul, et al.
Published: (2026)
LayLens: Improving Deepfake Understanding through Simplified Explanations
by: Narang, Abhijeet, et al.
Published: (2025)
by: Narang, Abhijeet, et al.
Published: (2025)
Revisiting MLLM Token Technology through the Lens of Classical Visual Coding
by: Liu, Jinming, et al.
Published: (2025)
by: Liu, Jinming, et al.
Published: (2025)
Understanding Visual Feature Reliance through the Lens of Complexity
by: Fel, Thomas, et al.
Published: (2024)
by: Fel, Thomas, et al.
Published: (2024)
AssoMem: Scalable Memory QA with Multi-Signal Associative Retrieval
by: Zhang, Kai, et al.
Published: (2025)
by: Zhang, Kai, et al.
Published: (2025)
TRACE: A Framework for Analyzing and Enhancing Stepwise Reasoning in Vision-Language Models
by: Imani, Shima, et al.
Published: (2025)
by: Imani, Shima, et al.
Published: (2025)
Visual Prompt-Agnostic Evolution
by: Wang, Junze, et al.
Published: (2026)
by: Wang, Junze, et al.
Published: (2026)
Knowledge Extraction on Semi-Structured Content: Does It Remain Relevant for Question Answering in the Era of LLMs?
by: Sun, Kai, et al.
Published: (2025)
by: Sun, Kai, et al.
Published: (2025)
Contextualized Visual Personalization in Vision-Language Models
by: Oh, Yeongtak, et al.
Published: (2026)
by: Oh, Yeongtak, et al.
Published: (2026)
Dimension Agnostic Testing of Survey Data Credibility through the Lens of Regression
by: Basu, Debabrota, et al.
Published: (2025)
by: Basu, Debabrota, et al.
Published: (2025)
Language-Agnostic Visual Embeddings for Cross-Script Handwriting Retrieval
by: Chen, Fangke, et al.
Published: (2026)
by: Chen, Fangke, et al.
Published: (2026)
Synthesis and Characterization of Coated CoFe2O4 Nanoparticles with Biocompatible Compounds and In Vitro Toxicity Assessment on Glioma Cell Lines
by: Sevil Ozer, et al.
Published: (2025)
by: Sevil Ozer, et al.
Published: (2025)
Understanding Tourists' Environmental Behavior Through the Lens of Willingness to Change and Motivational Factors
by: Millo Yaja, et al.
Published: (2026)
by: Millo Yaja, et al.
Published: (2026)
ChartCitor: Multi-Agent Framework for Fine-Grained Chart Visual Attribution
by: Goswami, Kanika, et al.
Published: (2025)
by: Goswami, Kanika, et al.
Published: (2025)
DeformAr: Rethinking NER Evaluation through Component Analysis and Visual Analytics
by: Younes, Ahmed Mustafa
Published: (2025)
by: Younes, Ahmed Mustafa
Published: (2025)
Task-Driven Lens Design
by: Yang, Xinge, et al.
Published: (2023)
by: Yang, Xinge, et al.
Published: (2023)
How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding
by: Yu, Zhuoran, et al.
Published: (2025)
by: Yu, Zhuoran, et al.
Published: (2025)
Generative Visual Foresight Meets Task-Agnostic Pose Estimation in Robotic Table-Top Manipulation
by: Zhang, Chuye, et al.
Published: (2025)
by: Zhang, Chuye, et al.
Published: (2025)
LossLens: Diagnostics for Machine Learning through Loss Landscape Visual Analytics
by: Xie, Tiankai, et al.
Published: (2024)
by: Xie, Tiankai, et al.
Published: (2024)
Revisiting Visual Understanding in Multimodal Reasoning through a Lens of Image Perturbation
by: Li, Yuting, et al.
Published: (2025)
by: Li, Yuting, et al.
Published: (2025)
ConfRAG: Confidence-Guided Retrieval-Augmenting Generation
by: Huang, Yin, et al.
Published: (2025)
by: Huang, Yin, et al.
Published: (2025)
PlotGen: Multi-Agent LLM-based Scientific Data Visualization via Multimodal Feedback
by: Goswami, Kanika, et al.
Published: (2025)
by: Goswami, Kanika, et al.
Published: (2025)
SymPyBench: A Dynamic Benchmark for Scientific Reasoning with Executable Python Code
by: Imani, Shima, et al.
Published: (2025)
by: Imani, Shima, et al.
Published: (2025)
PRiSM: An Agentic Multimodal Benchmark for Scientific Reasoning via Python-Grounded Evaluation
by: Imani, Shima, et al.
Published: (2025)
by: Imani, Shima, et al.
Published: (2025)
Data Type Agnostic Visual Sensitivity Analysis
by: Piccolotto, Nikolaus, et al.
Published: (2023)
by: Piccolotto, Nikolaus, et al.
Published: (2023)
Visual Attention Graph
by: Yang, Kai-Fu, et al.
Published: (2025)
by: Yang, Kai-Fu, et al.
Published: (2025)
WearVox: An Egocentric Multichannel Voice Assistant Benchmark for Wearables
by: Lin, Zhaojiang, et al.
Published: (2025)
by: Lin, Zhaojiang, et al.
Published: (2025)
Visual Histories of Occupation
Published: (2022)
Published: (2022)
Toward Real-Time Edge AI: Model-Agnostic Task-Oriented Communication with Visual Feature Alignment
by: Xie, Songjie, et al.
Published: (2024)
by: Xie, Songjie, et al.
Published: (2024)
Embodiment-Agnostic Navigation Policy Trained with Visual Demonstrations
by: Curtis, Nimrod, et al.
Published: (2024)
by: Curtis, Nimrod, et al.
Published: (2024)
HoLens: A Visual Analytics Design for Higher-order Movement Modeling and Visualization
by: Feng, Zezheng, et al.
Published: (2024)
by: Feng, Zezheng, et al.
Published: (2024)
Sharing Black History Month Art through Xerography and Visual Literacy.
by: Demery, Marie
Published: (1985)
by: Demery, Marie
Published: (1985)
Similar Items
-
Proactive Assistant Dialogue Generation from Streaming Egocentric Videos
by: Zhang, Yichi, et al.
Published: (2025) -
Pixel-Grounded Retrieval for Knowledgeable Large Multimodal Models
by: Kim, Jeonghwan, et al.
Published: (2026) -
SnapNTell: Enhancing Entity-Centric Visual Question Answering with Retrieval Augmented Multimodal LLM
by: Qiu, Jielin, et al.
Published: (2024) -
Large Language Models as Zero-shot Dialogue State Tracker through Function Calling
by: Li, Zekun, et al.
Published: (2024) -
Toward Architecture-Agnostic Local Control of Posterior Collapse in VAEs
by: Song, Hyunsoo, et al.
Published: (2025)