Unsupervised Memorability Modeling from Tip-of-the-Tongue Retrieval Queries
Fuente:
arXiv
Saved in:
| Main Authors: | Bhattacharyya, Sree, Singla, Yaman Kumar, Yarram, Sudhir, Singh, Somesh Kumar, S I, Harini, Wang, James Z. |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Measuring and Improving Persuasiveness of Large Language Models
by: Singh, Somesh, et al.
Published: (2024)
by: Singh, Somesh, et al.
Published: (2024)
Long-Term Ad Memorability: Understanding & Generating Memorable Ads
by: SI, Harini, et al.
Published: (2023)
by: SI, Harini, et al.
Published: (2023)
Teaching Human Behavior Improves Content Understanding Abilities Of LLMs
by: Singh, Somesh, et al.
Published: (2024)
by: Singh, Somesh, et al.
Published: (2024)
Evaluating Vision-Language Models for Emotion Recognition
by: Bhattacharyya, Sree, et al.
Published: (2025)
by: Bhattacharyya, Sree, et al.
Published: (2025)
Large Content And Behavior Models To Understand, Simulate, And Optimize Content And Behavior
by: Khandelwal, Ashmit, et al.
Published: (2023)
by: Khandelwal, Ashmit, et al.
Published: (2023)
Forecasting Future Videos from Novel Views via Disentangled 3D Scene Representation
by: Yarram, Sudhir, et al.
Published: (2024)
by: Yarram, Sudhir, et al.
Published: (2024)
A Heterogeneous Multimodal Graph Learning Framework for Recognizing User Emotions in Social Networks
by: Bhattacharyya, Sree, et al.
Published: (2025)
by: Bhattacharyya, Sree, et al.
Published: (2025)
Enhancing Image Retrieval : A Comprehensive Study on Photo Search using the CLIP Mode
by: Lahajal, Naresh Kumar, et al.
Published: (2024)
by: Lahajal, Naresh Kumar, et al.
Published: (2024)
Decomposing Queries into Tool Calls for Long-Video Keyframe Retrieval
by: Shlapentokh-Rothman, Michal, et al.
Published: (2026)
by: Shlapentokh-Rothman, Michal, et al.
Published: (2026)
Towards Adversarial Robustness And Backdoor Mitigation in SSL
by: Satpathy, Aryan, et al.
Published: (2024)
by: Satpathy, Aryan, et al.
Published: (2024)
GC-KBVQA: A New Four-Stage Framework for Enhancing Knowledge Based Visual Question Answering Performance
by: Moradi, Mohammad Mahdi, et al.
Published: (2025)
by: Moradi, Mohammad Mahdi, et al.
Published: (2025)
Watch Out Your Album! On the Inadvertent Privacy Memorization in Multi-Modal Large Language Models
by: Ju, Tianjie, et al.
Published: (2025)
by: Ju, Tianjie, et al.
Published: (2025)
MERIT: Multilingual Semantic Retrieval with Interleaved Multi-Condition Query
by: Chow, Wei, et al.
Published: (2025)
by: Chow, Wei, et al.
Published: (2025)
Generative Cross-Modal Retrieval: Memorizing Images in Multimodal Language Models for Retrieval and Beyond
by: Li, Yongqi, et al.
Published: (2024)
by: Li, Yongqi, et al.
Published: (2024)
Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries
by: Wu, Yin, et al.
Published: (2025)
by: Wu, Yin, et al.
Published: (2025)
Q2E: Query-to-Event Decomposition for Zero-Shot Multilingual Text-to-Video Retrieval
by: Dipta, Shubhashis Roy, et al.
Published: (2025)
by: Dipta, Shubhashis Roy, et al.
Published: (2025)
Transformer-based Clipped Contrastive Quantization Learning for Unsupervised Image Retrieval
by: Dubey, Ayush, et al.
Published: (2024)
by: Dubey, Ayush, et al.
Published: (2024)
RoundTripOCR: A Data Generation Technique for Enhancing Post-OCR Error Correction in Low-Resource Devanagari Languages
by: Kashid, Harshvivek, et al.
Published: (2024)
by: Kashid, Harshvivek, et al.
Published: (2024)
Frame-Voyager: Learning to Query Frames for Video Large Language Models
by: Yu, Sicheng, et al.
Published: (2024)
by: Yu, Sicheng, et al.
Published: (2024)
AdaMMS: Model Merging for Heterogeneous Multimodal Large Language Models with Unsupervised Coefficient Optimization
by: Du, Yiyang, et al.
Published: (2025)
by: Du, Yiyang, et al.
Published: (2025)
Text Change Detection in Multilingual Documents Using Image Comparison
by: Park, Doyoung, et al.
Published: (2024)
by: Park, Doyoung, et al.
Published: (2024)
Give me a hint: Can LLMs take a hint to solve math problems?
by: Agrawal, Vansh, et al.
Published: (2024)
by: Agrawal, Vansh, et al.
Published: (2024)
The Percept-V Challenge: Can Multimodal LLMs Crack Simple Perception Problems?
by: Ghosh, Samrajnee, et al.
Published: (2025)
by: Ghosh, Samrajnee, et al.
Published: (2025)
Investigating Spatial Attention Bias in Vision-Language Models
by: Chaudhary, Aryan, et al.
Published: (2025)
by: Chaudhary, Aryan, et al.
Published: (2025)
Query-centric Audio-Visual Cognition Network for Moment Retrieval, Segmentation and Step-Captioning
by: Tu, Yunbin, et al.
Published: (2024)
by: Tu, Yunbin, et al.
Published: (2024)
VisCon-100K: Leveraging Contextual Web Data for Fine-tuning Vision Language Models
by: Kumar, Gokul Karthik, et al.
Published: (2025)
by: Kumar, Gokul Karthik, et al.
Published: (2025)
Do Vision-Language Models Understand Compound Nouns?
by: Kumar, Sonal, et al.
Published: (2024)
by: Kumar, Sonal, et al.
Published: (2024)
A Multimodal, Multitask System for Generating E Commerce Text Listings from Images
by: Singh, Nayan Kumar
Published: (2025)
by: Singh, Nayan Kumar
Published: (2025)
ALN-P3: Unified Language Alignment for Perception, Prediction, and Planning in Autonomous Driving
by: Ma, Yunsheng, et al.
Published: (2025)
by: Ma, Yunsheng, et al.
Published: (2025)
Seeing Eye to AI: Comparing Human Gaze and Model Attention in Video Memorability
by: Kumar, Prajneya, et al.
Published: (2023)
by: Kumar, Prajneya, et al.
Published: (2023)
Understanding Retrieval Robustness for Retrieval-Augmented Image Captioning
by: Li, Wenyan, et al.
Published: (2024)
by: Li, Wenyan, et al.
Published: (2024)
StegaVision: Enhancing Steganography with Attention Mechanism
by: Kumar, Abhinav, et al.
Published: (2024)
by: Kumar, Abhinav, et al.
Published: (2024)
Scaling Concept With Text-Guided Diffusion Models
by: Huang, Chao, et al.
Published: (2024)
by: Huang, Chao, et al.
Published: (2024)
Referencing Where to Focus: Improving VisualGrounding with Referential Query
by: Wang, Yabing, et al.
Published: (2024)
by: Wang, Yabing, et al.
Published: (2024)
UniVS: Unified and Universal Video Segmentation with Prompts as Queries
by: Li, Minghan, et al.
Published: (2024)
by: Li, Minghan, et al.
Published: (2024)
Enhancing Adverse Drug Event Detection with Multimodal Dataset: Corpus Creation and Model Development
by: Sahoo, Pranab, et al.
Published: (2024)
by: Sahoo, Pranab, et al.
Published: (2024)
brat: Aligned Multi-View Embeddings for Brain MRI Analysis
by: Kayser, Maxime, et al.
Published: (2025)
by: Kayser, Maxime, et al.
Published: (2025)
PatientVLM Meets DocVLM: Pre-Consultation Dialogue Between Vision-Language Models for Efficient Diagnosis
by: Lokesh, K, et al.
Published: (2026)
by: Lokesh, K, et al.
Published: (2026)
Entropy-Gradient Grounding: Training-Free Evidence Retrieval in Vision-Language Models
by: Gröpl, Marcel, et al.
Published: (2026)
by: Gröpl, Marcel, et al.
Published: (2026)
Reminding Multimodal Large Language Models of Object-aware Knowledge with Retrieved Tags
by: Qi, Daiqing, et al.
Published: (2024)
by: Qi, Daiqing, et al.
Published: (2024)
Similar Items
-
Measuring and Improving Persuasiveness of Large Language Models
by: Singh, Somesh, et al.
Published: (2024) -
Long-Term Ad Memorability: Understanding & Generating Memorable Ads
by: SI, Harini, et al.
Published: (2023) -
Teaching Human Behavior Improves Content Understanding Abilities Of LLMs
by: Singh, Somesh, et al.
Published: (2024) -
Evaluating Vision-Language Models for Emotion Recognition
by: Bhattacharyya, Sree, et al.
Published: (2025) -
Large Content And Behavior Models To Understand, Simulate, And Optimize Content And Behavior
by: Khandelwal, Ashmit, et al.
Published: (2023)