MERLIN: Multimodal Embedding Refinement via LLM-based Iterative Navigation for Text-Video Retrieval-Rerank Pipeline
Fuente:
arXiv
Guardado en:
| Autores principales: | Han, Donghoon, Park, Eunhwan, Lee, Gisang, Lee, Adam, Kwak, Nojun |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
CacheFocus: Dynamic Cache Re-Positioning for Efficient Retrieval-Augmented Generation
por: Lee, Kun-Hui, et al.
Publicado: (2025)
por: Lee, Kun-Hui, et al.
Publicado: (2025)
Unleash the Potential of CLIP for Video Highlight Detection
por: Han, Donghoon, et al.
Publicado: (2024)
por: Han, Donghoon, et al.
Publicado: (2024)
Retrieval-Augmented Generation Based Nurse Observation Extraction
por: Hwang, Kyomin, et al.
Publicado: (2026)
por: Hwang, Kyomin, et al.
Publicado: (2026)
Visual Words Meet BM25: Sparse Auto-Encoder Visual Word Scoring for Image Retrieval
por: Han, Donghoon, et al.
Publicado: (2026)
por: Han, Donghoon, et al.
Publicado: (2026)
Unlocking the Potential of Diffusion Language Models through Template Infilling
por: Lee, Junhoo, et al.
Publicado: (2025)
por: Lee, Junhoo, et al.
Publicado: (2025)
KRETA: A Benchmark for Korean Reading and Reasoning in Text-Rich VQA Attuned to Diverse Visual Contexts
por: Hwang, Taebaek, et al.
Publicado: (2025)
por: Hwang, Taebaek, et al.
Publicado: (2025)
FeRG-LLM : Feature Engineering by Reason Generation Large Language Models
por: Ko, Jeonghyun, et al.
Publicado: (2025)
por: Ko, Jeonghyun, et al.
Publicado: (2025)
Expanding Search Space with Diverse Prompting Agents: An Efficient Sampling Approach for LLM Mathematical Reasoning
por: Lee, Gisang, et al.
Publicado: (2024)
por: Lee, Gisang, et al.
Publicado: (2024)
Do not think about pink elephant!
por: Hwang, Kyomin, et al.
Publicado: (2024)
por: Hwang, Kyomin, et al.
Publicado: (2024)
Let LLMs Speak Embedding Languages: Generative Text Embeddings via Iterative Contrastive Refinement
por: Tsai, Yu-Che, et al.
Publicado: (2025)
por: Tsai, Yu-Che, et al.
Publicado: (2025)
Cross-Genre Authorship Attribution via LLM-Based Retrieve-and-Rerank
por: Agarwal, Shantanu, et al.
Publicado: (2025)
por: Agarwal, Shantanu, et al.
Publicado: (2025)
When Confidence Misleads: Suffix Anchoring and Anchor-Proximity Confidence Modulation for Diffusion Language Models
por: Park, Jungwon, et al.
Publicado: (2026)
por: Park, Jungwon, et al.
Publicado: (2026)
Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models
por: Zhang, Yanzhao, et al.
Publicado: (2025)
por: Zhang, Yanzhao, et al.
Publicado: (2025)
Omni-Embed-Audio: Leveraging Multimodal LLMs for Robust Audio-Text Retrieval
por: Yoo, HaeJun, et al.
Publicado: (2026)
por: Yoo, HaeJun, et al.
Publicado: (2026)
How Good are LLM-based Rerankers? An Empirical Analysis of State-of-the-Art Reranking Models
por: Abdallah, Abdelrahman, et al.
Publicado: (2025)
por: Abdallah, Abdelrahman, et al.
Publicado: (2025)
MERLIN: Multi-Stage Curriculum Alignment for Multilingual Encoder-LLM Integration in Cross-Lingual Reasoning
por: Uemura, Kosei, et al.
Publicado: (2025)
por: Uemura, Kosei, et al.
Publicado: (2025)
CSF: Black-box Fingerprinting via Compositional Semantics for Text-to-Image Models
por: Lee, Junhoo, et al.
Publicado: (2026)
por: Lee, Junhoo, et al.
Publicado: (2026)
Omni-Embed-Nemotron: A Unified Multimodal Retrieval Model for Text, Image, Audio, and Video
por: Xu, Mengyao, et al.
Publicado: (2025)
por: Xu, Mengyao, et al.
Publicado: (2025)
Knowledge Beyond Language: Bridging the Gap in Multilingual Machine Unlearning Evaluation
por: Hwang, Kyomin, et al.
Publicado: (2026)
por: Hwang, Kyomin, et al.
Publicado: (2026)
MemReranker: Reasoning-Aware Reranking for Agent Memory Retrieval
por: Li, Chunyu, et al.
Publicado: (2026)
por: Li, Chunyu, et al.
Publicado: (2026)
Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking
por: Li, Mingxin, et al.
Publicado: (2026)
por: Li, Mingxin, et al.
Publicado: (2026)
CoRank: LLM-Based Compact Reranking with Document Features for Scientific Retrieval
por: Tian, Runchu, et al.
Publicado: (2025)
por: Tian, Runchu, et al.
Publicado: (2025)
MERLIN: A Testbed for Multilingual Multimodal Entity Recognition and Linking
por: Ramamoorthy, Sathyanarayanan, et al.
Publicado: (2025)
por: Ramamoorthy, Sathyanarayanan, et al.
Publicado: (2025)
Mira-Embeddings-V1: Domain-Adapted Semantic Reranking for Recruitment via LLM-Synthesized Data
por: Liang, Zhaohua, et al.
Publicado: (2026)
por: Liang, Zhaohua, et al.
Publicado: (2026)
Embedding-Based Context-Aware Reranker
por: Yuan, Ye, et al.
Publicado: (2025)
por: Yuan, Ye, et al.
Publicado: (2025)
FLAIRR-TS -- Forecasting LLM-Agents with Iterative Refinement and Retrieval for Time Series
por: Jalori, Gunjan, et al.
Publicado: (2025)
por: Jalori, Gunjan, et al.
Publicado: (2025)
Efficiency-Effectiveness Reranking FLOPs for LLM-based Rerankers
por: Peng, Zhiyuan, et al.
Publicado: (2025)
por: Peng, Zhiyuan, et al.
Publicado: (2025)
RAIR: Retrieval-Augmented Iterative Refinement for Chinese Spelling Correction
por: Liang, Junhong, et al.
Publicado: (2025)
por: Liang, Junhong, et al.
Publicado: (2025)
DS@GT at CheckThat! 2025: Exploring Retrieval and Reranking Pipelines for Scientific Claim Source Retrieval on Social Media Discourse
por: Schofield, Jeanette, et al.
Publicado: (2025)
por: Schofield, Jeanette, et al.
Publicado: (2025)
DEO: Training-Free Direct Embedding Optimization for Negation-Aware Retrieval
por: Lee, Taegyeong, et al.
Publicado: (2026)
por: Lee, Taegyeong, et al.
Publicado: (2026)
Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement
por: Lee, Daeun, et al.
Publicado: (2024)
por: Lee, Daeun, et al.
Publicado: (2024)
Advancing Beyond Identification: Multi-bit Watermark for Large Language Models
por: Yoo, KiYoon, et al.
Publicado: (2023)
por: Yoo, KiYoon, et al.
Publicado: (2023)
QCG-Rerank: Chunks Graph Rerank with Query Expansion in Retrieval-Augmented LLMs for Tourism Domain
por: Wei, Qikai, et al.
Publicado: (2024)
por: Wei, Qikai, et al.
Publicado: (2024)
MLLM Is a Strong Reranker: Advancing Multimodal Retrieval-augmented Generation via Knowledge-enhanced Reranking and Noise-injected Training
por: Chen, Zhanpeng, et al.
Publicado: (2024)
por: Chen, Zhanpeng, et al.
Publicado: (2024)
LLM driven Text-to-Table Generation through Sub-Tasks Guidance and Iterative Refinement
por: C, Rajmohan, et al.
Publicado: (2025)
por: C, Rajmohan, et al.
Publicado: (2025)
mGTE: Generalized Long-Context Text Representation and Reranking Models for Multilingual Text Retrieval
por: Zhang, Xin, et al.
Publicado: (2024)
por: Zhang, Xin, et al.
Publicado: (2024)
DeAR: Dual-Stage Document Reranking with Reasoning Agents via LLM Distillation
por: Abdallah, Abdelrahman, et al.
Publicado: (2025)
por: Abdallah, Abdelrahman, et al.
Publicado: (2025)
What's Making That Sound Right Now? Video-centric Audio-Visual Localization
por: Choi, Hahyeon, et al.
Publicado: (2025)
por: Choi, Hahyeon, et al.
Publicado: (2025)
ELITE: Embedding-Less retrieval with Iterative Text Exploration
por: Wang, Zhangyu, et al.
Publicado: (2025)
por: Wang, Zhangyu, et al.
Publicado: (2025)
Decompose-and-Refine: Structured Legal Question Answering with Parametric Retrieval
por: lee, Jihyung, et al.
Publicado: (2026)
por: lee, Jihyung, et al.
Publicado: (2026)
Ejemplares similares
-
CacheFocus: Dynamic Cache Re-Positioning for Efficient Retrieval-Augmented Generation
por: Lee, Kun-Hui, et al.
Publicado: (2025) -
Unleash the Potential of CLIP for Video Highlight Detection
por: Han, Donghoon, et al.
Publicado: (2024) -
Retrieval-Augmented Generation Based Nurse Observation Extraction
por: Hwang, Kyomin, et al.
Publicado: (2026) -
Visual Words Meet BM25: Sparse Auto-Encoder Visual Word Scoring for Image Retrieval
por: Han, Donghoon, et al.
Publicado: (2026) -
Unlocking the Potential of Diffusion Language Models through Template Infilling
por: Lee, Junhoo, et al.
Publicado: (2025)