SODIUM: From Open Web Data to Queryable Databases
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Hu, Chuxuan, Li, Philip, Yang, Maxwell, Kang, Daniel |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DRAMA: Unifying Data Retrieval and Analysis for Open-Domain Analytic Queries
von: Hu, Chuxuan, et al.
Veröffentlicht: (2025)
von: Hu, Chuxuan, et al.
Veröffentlicht: (2025)
Multimodal Neural Databases
von: Trappolini, Giovanni, et al.
Veröffentlicht: (2023)
von: Trappolini, Giovanni, et al.
Veröffentlicht: (2023)
GPT-4V(ision) is a Generalist Web Agent, if Grounded
von: Zheng, Boyuan, et al.
Veröffentlicht: (2024)
von: Zheng, Boyuan, et al.
Veröffentlicht: (2024)
GaussMaster: An LLM-based Database Copilot System
von: Zhou, Wei, et al.
Veröffentlicht: (2025)
von: Zhou, Wei, et al.
Veröffentlicht: (2025)
Introduction of a tree-based technique for efficient and real-time label retrieval in the object tracking system
von: Benrazek, Ala-Eddine, et al.
Veröffentlicht: (2022)
von: Benrazek, Ala-Eddine, et al.
Veröffentlicht: (2022)
LazyVLM: Neuro-Symbolic Approach to Video Analytics
von: Jian, Xiangru, et al.
Veröffentlicht: (2025)
von: Jian, Xiangru, et al.
Veröffentlicht: (2025)
MagicLens: Self-Supervised Image Retrieval with Open-Ended Instructions
von: Zhang, Kai, et al.
Veröffentlicht: (2024)
von: Zhang, Kai, et al.
Veröffentlicht: (2024)
From Videos to Indexed Knowledge Graphs -- Framework to Marry Methods for Multimodal Content Analysis and Understanding
von: Rizk, Basem, et al.
Veröffentlicht: (2025)
von: Rizk, Basem, et al.
Veröffentlicht: (2025)
$M^3EL$: A Multi-task Multi-topic Dataset for Multi-modal Entity Linking
von: Wang, Fang, et al.
Veröffentlicht: (2024)
von: Wang, Fang, et al.
Veröffentlicht: (2024)
ImplicitAVE: An Open-Source Dataset and Multimodal LLMs Benchmark for Implicit Attribute Value Extraction
von: Zou, Henry Peng, et al.
Veröffentlicht: (2024)
von: Zou, Henry Peng, et al.
Veröffentlicht: (2024)
Personalized Multimodal Large Language Models: A Survey
von: Wu, Junda, et al.
Veröffentlicht: (2024)
von: Wu, Junda, et al.
Veröffentlicht: (2024)
MMSearch: Benchmarking the Potential of Large Models as Multi-modal Search Engines
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2024)
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2024)
MMDocIR: Benchmarking Multimodal Retrieval for Long Documents
von: Dong, Kuicai, et al.
Veröffentlicht: (2025)
von: Dong, Kuicai, et al.
Veröffentlicht: (2025)
Self Knowledge Re-expression: A Fully Local Method for Adapting LLMs to Tasks Using Intrinsic Knowledge
von: Wang, Mengyu, et al.
Veröffentlicht: (2026)
von: Wang, Mengyu, et al.
Veröffentlicht: (2026)
Agent-UniRAG: A Trainable Open-Source LLM Agent Framework for Unified Retrieval-Augmented Generation Systems
von: Pham, Hoang, et al.
Veröffentlicht: (2025)
von: Pham, Hoang, et al.
Veröffentlicht: (2025)
SPARQL Generation: an analysis on fine-tuning OpenLLaMA for Question Answering over a Life Science Knowledge Graph
von: Rangel, Julio C., et al.
Veröffentlicht: (2024)
von: Rangel, Julio C., et al.
Veröffentlicht: (2024)
Automating Database-Native Function Code Synthesis with LLMs
von: Zhou, Wei, et al.
Veröffentlicht: (2026)
von: Zhou, Wei, et al.
Veröffentlicht: (2026)
From Verbatim to Gist: Distilling Pyramidal Multimodal Memory via Semantic Information Bottleneck for Long-Horizon Video Agents
von: Lian, Niu, et al.
Veröffentlicht: (2026)
von: Lian, Niu, et al.
Veröffentlicht: (2026)
Taxonomy Inference for Tabular Data Using Large Language Models
von: Wu, Zhenyu, et al.
Veröffentlicht: (2025)
von: Wu, Zhenyu, et al.
Veröffentlicht: (2025)
DSRAG: A Domain-Specific Retrieval Framework Based on Document-derived Multimodal Knowledge Graph
von: Yang, Mengzheng, et al.
Veröffentlicht: (2025)
von: Yang, Mengzheng, et al.
Veröffentlicht: (2025)
VideoComp: Advancing Fine-Grained Compositional and Temporal Alignment in Video-Text Models
von: Kim, Dahun, et al.
Veröffentlicht: (2025)
von: Kim, Dahun, et al.
Veröffentlicht: (2025)
VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents
von: Yu, Shi, et al.
Veröffentlicht: (2024)
von: Yu, Shi, et al.
Veröffentlicht: (2024)
LITTA: Late-Interaction and Test-Time Alignment for Visually-Grounded Multimodal Retrieval
von: Kim, Seonok
Veröffentlicht: (2026)
von: Kim, Seonok
Veröffentlicht: (2026)
Smart Multi-Modal Search: Contextual Sparse and Dense Embedding Integration in Adobe Express
von: Aroraa, Cherag, et al.
Veröffentlicht: (2024)
von: Aroraa, Cherag, et al.
Veröffentlicht: (2024)
INQUIRE: A Natural World Text-to-Image Retrieval Benchmark
von: Vendrow, Edward, et al.
Veröffentlicht: (2024)
von: Vendrow, Edward, et al.
Veröffentlicht: (2024)
ViBERTgrid BiLSTM-CRF: Multimodal Key Information Extraction from Unstructured Financial Documents
von: Pala, Furkan, et al.
Veröffentlicht: (2024)
von: Pala, Furkan, et al.
Veröffentlicht: (2024)
Leveraging Customer Feedback for Multi-modal Insight Extraction
von: Mukku, Sandeep Sricharan, et al.
Veröffentlicht: (2024)
von: Mukku, Sandeep Sricharan, et al.
Veröffentlicht: (2024)
Learning Visual Composition through Improved Semantic Guidance
von: Stone, Austin, et al.
Veröffentlicht: (2024)
von: Stone, Austin, et al.
Veröffentlicht: (2024)
Progressive Multimodal Reasoning via Active Retrieval
von: Dong, Guanting, et al.
Veröffentlicht: (2024)
von: Dong, Guanting, et al.
Veröffentlicht: (2024)
NTSEBENCH: Cognitive Reasoning Benchmark for Vision Language Models
von: Pandya, Pranshu, et al.
Veröffentlicht: (2024)
von: Pandya, Pranshu, et al.
Veröffentlicht: (2024)
M3DR: Towards Universal Multilingual Multimodal Document Retrieval
von: Kolavi, Adithya S, et al.
Veröffentlicht: (2025)
von: Kolavi, Adithya S, et al.
Veröffentlicht: (2025)
Naiad: Novel Agentic Intelligent Autonomous System for Inland Water Monitoring
von: Baltzi, Eirini, et al.
Veröffentlicht: (2025)
von: Baltzi, Eirini, et al.
Veröffentlicht: (2025)
VDocRAG: Retrieval-Augmented Generation over Visually-Rich Documents
von: Tanaka, Ryota, et al.
Veröffentlicht: (2025)
von: Tanaka, Ryota, et al.
Veröffentlicht: (2025)
MegaRAG: Multimodal Knowledge Graph-Based Retrieval Augmented Generation
von: Hsiao, Chi-Hsiang, et al.
Veröffentlicht: (2025)
von: Hsiao, Chi-Hsiang, et al.
Veröffentlicht: (2025)
Multi-Grained Query-Guided Set Prediction Network for Grounded Multimodal Named Entity Recognition
von: Tang, Jielong, et al.
Veröffentlicht: (2024)
von: Tang, Jielong, et al.
Veröffentlicht: (2024)
Beyond Unimodal Boundaries: Generative Recommendation with Multimodal Semantics
von: Zhu, Jing, et al.
Veröffentlicht: (2025)
von: Zhu, Jing, et al.
Veröffentlicht: (2025)
ViDoRAG: Visual Document Retrieval-Augmented Generation via Dynamic Iterative Reasoning Agents
von: Wang, Qiuchen, et al.
Veröffentlicht: (2025)
von: Wang, Qiuchen, et al.
Veröffentlicht: (2025)
Chain of Evidence: Pixel-Level Visual Attribution for Iterative Retrieval-Augmented Generation
von: Liu, Peiyang, et al.
Veröffentlicht: (2026)
von: Liu, Peiyang, et al.
Veröffentlicht: (2026)
VideoAgent: Long-form Video Understanding with Large Language Model as Agent
von: Wang, Xiaohan, et al.
Veröffentlicht: (2024)
von: Wang, Xiaohan, et al.
Veröffentlicht: (2024)
XL-HeadTags: Leveraging Multimodal Retrieval Augmentation for the Multilingual Generation of News Headlines and Tags
von: Shohan, Faisal Tareque, et al.
Veröffentlicht: (2024)
von: Shohan, Faisal Tareque, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
DRAMA: Unifying Data Retrieval and Analysis for Open-Domain Analytic Queries
von: Hu, Chuxuan, et al.
Veröffentlicht: (2025) -
Multimodal Neural Databases
von: Trappolini, Giovanni, et al.
Veröffentlicht: (2023) -
GPT-4V(ision) is a Generalist Web Agent, if Grounded
von: Zheng, Boyuan, et al.
Veröffentlicht: (2024) -
GaussMaster: An LLM-based Database Copilot System
von: Zhou, Wei, et al.
Veröffentlicht: (2025) -
Introduction of a tree-based technique for efficient and real-time label retrieval in the object tracking system
von: Benrazek, Ala-Eddine, et al.
Veröffentlicht: (2022)