Proximity QA: Unleashing the Power of Multi-Modal Large Language Models for Spatial Proximity Analysis
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Li, Jianing, Nan, Xi, Lu, Ming, Du, Li, Zhang, Shanghang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Automating Clinical Information Retrieval from Finnish Electronic Health Records Using Large Language Models
von: Saukkoriipi, Mikko, et al.
Veröffentlicht: (2026)
von: Saukkoriipi, Mikko, et al.
Veröffentlicht: (2026)
Pipeline and Dataset Generation for Automated Fact-checking in Almost Any Language
von: Drchal, Jan, et al.
Veröffentlicht: (2023)
von: Drchal, Jan, et al.
Veröffentlicht: (2023)
Language Predicts Identity Fusion Across Cultures and Reveals Divergent Pathways to Violence
von: Wright, Devin R., et al.
Veröffentlicht: (2026)
von: Wright, Devin R., et al.
Veröffentlicht: (2026)
CG-TTRL: Context-Guided Test-Time Reinforcement Learning for On-Device Large Language Models
von: Hosseini, Peyman, et al.
Veröffentlicht: (2025)
von: Hosseini, Peyman, et al.
Veröffentlicht: (2025)
Cytoarchitecture in Words: Weakly Supervised Vision-Language Modeling for Human Brain Microscopy
von: Sutton, Matthew, et al.
Veröffentlicht: (2026)
von: Sutton, Matthew, et al.
Veröffentlicht: (2026)
Scaling Laws for State Dynamics in Large Language Models
von: Li, Jacob X, et al.
Veröffentlicht: (2025)
von: Li, Jacob X, et al.
Veröffentlicht: (2025)
Fine-Tuning Vision-Language Models for Understanding Current Damage and Scoring Priority with Quality Guard Agent
von: Yasuno, Takato
Veröffentlicht: (2026)
von: Yasuno, Takato
Veröffentlicht: (2026)
propella-1: Multi-Property Document Annotation for LLM Data Curation at Scale
von: Idahl, Maximilian, et al.
Veröffentlicht: (2026)
von: Idahl, Maximilian, et al.
Veröffentlicht: (2026)
Multilingual Multi-Label Emotion Classification at Scale with Synthetic Data
von: Borisov, Vadim
Veröffentlicht: (2026)
von: Borisov, Vadim
Veröffentlicht: (2026)
Detection of Personal Data in Structured Datasets Using a Large Language Model
von: Ntwali, Albert Agisha, et al.
Veröffentlicht: (2025)
von: Ntwali, Albert Agisha, et al.
Veröffentlicht: (2025)
PRISMA: Preference-Reinforced Self-Training Approach for Interpretable Emotionally Intelligent Negotiation Dialogues
von: Kajare, Prajwal Vijay, et al.
Veröffentlicht: (2026)
von: Kajare, Prajwal Vijay, et al.
Veröffentlicht: (2026)
LongSumEval: Question-Answering Based Evaluation and Feedback-Driven Refinement for Long Document Summarization
von: Nguyen, Huyen, et al.
Veröffentlicht: (2026)
von: Nguyen, Huyen, et al.
Veröffentlicht: (2026)
Logits-Constrained Framework with RoBERTa for Ancient Chinese NER
von: Hua, Wenjie, et al.
Veröffentlicht: (2025)
von: Hua, Wenjie, et al.
Veröffentlicht: (2025)
Enhancing Spatial Reasoning in Vision-Language Models via Chain-of-Thought Prompting and Reinforcement Learning
von: Ji, Binbin, et al.
Veröffentlicht: (2025)
von: Ji, Binbin, et al.
Veröffentlicht: (2025)
Distilled HuBERT for Mobile Speech Emotion Recognition: A Cross-Corpus Validation Study
von: Ismail, Saifelden M.
Veröffentlicht: (2025)
von: Ismail, Saifelden M.
Veröffentlicht: (2025)
Chronic pain patient narratives allow for the estimation of current pain intensity
von: Nunes, Diogo A. P., et al.
Veröffentlicht: (2022)
von: Nunes, Diogo A. P., et al.
Veröffentlicht: (2022)
Measuring Robustness of Speech Recognition from MEG Signals Under Distribution Shift
von: Chien, Sheng-You, et al.
Veröffentlicht: (2026)
von: Chien, Sheng-You, et al.
Veröffentlicht: (2026)
Real-World En Call Center Transcripts Dataset with PII Redaction
von: Dao, Ha, et al.
Veröffentlicht: (2025)
von: Dao, Ha, et al.
Veröffentlicht: (2025)
ParaNoise-SV: Integrated Approach for Noise-Robust Speaker Verification with Parallel Joint Learning of Speech Enhancement and Noise Extraction
von: Kim, Minu, et al.
Veröffentlicht: (2025)
von: Kim, Minu, et al.
Veröffentlicht: (2025)
Cognitive Linguistic Identity Fusion Score (CLIFS): A Scalable Cognition-Informed Approach to Quantifying Identity Fusion from Text
von: Wright, Devin R., et al.
Veröffentlicht: (2025)
von: Wright, Devin R., et al.
Veröffentlicht: (2025)
Improving Cross-Lingual Phonetic Representation of Low-Resource Languages Through Language Similarity Analysis
von: Kim, Minu, et al.
Veröffentlicht: (2025)
von: Kim, Minu, et al.
Veröffentlicht: (2025)
Talking Tennis: Language Feedback from 3D Biomechanical Action Recognition
von: Dashore, Arushi, et al.
Veröffentlicht: (2025)
von: Dashore, Arushi, et al.
Veröffentlicht: (2025)
Transfer-learning for video classification: Video Swin Transformer on multiple domains
von: Oliveira, Daniel A. P., et al.
Veröffentlicht: (2022)
von: Oliveira, Daniel A. P., et al.
Veröffentlicht: (2022)
Thaka at KSAA-2026 Task 2: Regularized Fine-Tuning for Arabic Speech Diacritization
von: Alamr, Meshal, et al.
Veröffentlicht: (2026)
von: Alamr, Meshal, et al.
Veröffentlicht: (2026)
Generative AI for Strategic Plan Development
von: Ponnock, Jesse
Veröffentlicht: (2025)
von: Ponnock, Jesse
Veröffentlicht: (2025)
Predicting Upcoming Stuttering Events from Three-Second Audio: Stratified Evaluation Reveals Severity-Selective Precursors, and the Model Deploys Fully On-Device
von: Kozak, Nazar
Veröffentlicht: (2026)
von: Kozak, Nazar
Veröffentlicht: (2026)
VocSim: A Training-free Benchmark for Zero-shot Content Identity in Single-source Audio
von: Basha, Maris, et al.
Veröffentlicht: (2025)
von: Basha, Maris, et al.
Veröffentlicht: (2025)
Delayed Fusion: Integrating Large Language Models into First-Pass Decoding in End-to-end Speech Recognition
von: Hori, Takaaki, et al.
Veröffentlicht: (2025)
von: Hori, Takaaki, et al.
Veröffentlicht: (2025)
Modular Deep Learning Framework for Assistive Perception: Gaze, Affect, and Speaker Identification
von: Anchan, Akshit Pramod, et al.
Veröffentlicht: (2025)
von: Anchan, Akshit Pramod, et al.
Veröffentlicht: (2025)
Motion-Based Sign Language Video Summarization using Curvature and Torsion
von: Sartinas, Evangelos G., et al.
Veröffentlicht: (2023)
von: Sartinas, Evangelos G., et al.
Veröffentlicht: (2023)
Structural Stress and Learned Helplessness in Afghanistan: A Multi-Layer Analysis of the AFSTRESS Dari Corpus
von: Baktash, Jawid Ahmad, et al.
Veröffentlicht: (2026)
von: Baktash, Jawid Ahmad, et al.
Veröffentlicht: (2026)
Leveraging large multimodal models for audio-video deepfake detection: a pilot study
von: Cao, Songjun, et al.
Veröffentlicht: (2026)
von: Cao, Songjun, et al.
Veröffentlicht: (2026)
From Capabilities to Performance: Evaluating Key Functional Properties of LLM Architectures in Penetration Testing
von: Huang, Lanxiao, et al.
Veröffentlicht: (2025)
von: Huang, Lanxiao, et al.
Veröffentlicht: (2025)
FS-DAG: Few Shot Domain Adapting Graph Networks for Visually Rich Document Understanding
von: Agarwal, Amit, et al.
Veröffentlicht: (2025)
von: Agarwal, Amit, et al.
Veröffentlicht: (2025)
Deep Interest Mining for Intent-Enriched Semantic IDs in Multimodal Generative Recommendation
von: Zeng, Yangchen, et al.
Veröffentlicht: (2026)
von: Zeng, Yangchen, et al.
Veröffentlicht: (2026)
Machine Learning for Coding Retail Product Names to Consumer-Price Categories: A Rule-plus-Bag-of-Words Pipeline with Reliability-Weighted Human-in-the-Loop Labeling
von: Beskorovainyi, Vladimir
Veröffentlicht: (2026)
von: Beskorovainyi, Vladimir
Veröffentlicht: (2026)
Emotional Voice Messages (EMOVOME) database: emotion recognition in spontaneous voice messages
von: Zaragozá, Lucía Gómez, et al.
Veröffentlicht: (2024)
von: Zaragozá, Lucía Gómez, et al.
Veröffentlicht: (2024)
Making a Pipeline Production-Ready: Challenges and Lessons Learned in the Healthcare Domain
von: Lawand, Daniel Angelo Esteves, et al.
Veröffentlicht: (2025)
von: Lawand, Daniel Angelo Esteves, et al.
Veröffentlicht: (2025)
SPIRA: Building an Intelligent System for Respiratory Insufficiency Detection
von: Ferreira, Renato Cordeiro, et al.
Veröffentlicht: (2025)
von: Ferreira, Renato Cordeiro, et al.
Veröffentlicht: (2025)
DeepC4: Deep Conditional Census-Constrained Clustering for Large-scale Multitask Spatial Disaggregation of Urban Morphology
von: Dimasaka, Joshua, et al.
Veröffentlicht: (2025)
von: Dimasaka, Joshua, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Automating Clinical Information Retrieval from Finnish Electronic Health Records Using Large Language Models
von: Saukkoriipi, Mikko, et al.
Veröffentlicht: (2026) -
Pipeline and Dataset Generation for Automated Fact-checking in Almost Any Language
von: Drchal, Jan, et al.
Veröffentlicht: (2023) -
Language Predicts Identity Fusion Across Cultures and Reveals Divergent Pathways to Violence
von: Wright, Devin R., et al.
Veröffentlicht: (2026) -
CG-TTRL: Context-Guided Test-Time Reinforcement Learning for On-Device Large Language Models
von: Hosseini, Peyman, et al.
Veröffentlicht: (2025) -
Cytoarchitecture in Words: Weakly Supervised Vision-Language Modeling for Human Brain Microscopy
von: Sutton, Matthew, et al.
Veröffentlicht: (2026)