Chat-Driven Text Generation and Interaction for Person Retrieval
Fuente:
arXiv
Saved in:
| Main Authors: | Xie, Zequn, Wang, Chuxin, Cai, Sihang, Wang, Yeqiang, Wang, Shulei, Jin, Tao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation
by: Niu, Yuwei, et al.
Published: (2025)
by: Niu, Yuwei, et al.
Published: (2025)
Deep Learning based Key Information Extraction from Business Documents: Systematic Literature Review
by: Rombach, Alexander Michael, et al.
Published: (2024)
by: Rombach, Alexander Michael, et al.
Published: (2024)
Training-Free Diffusion Priors for Text-to-Image Generation via Optimization-based Visual Inversion
by: Dell'Erba, Samuele, et al.
Published: (2025)
by: Dell'Erba, Samuele, et al.
Published: (2025)
FlexDoc: Parameterized Sampling for Diverse Multilingual Synthetic Documents for Training Document Understanding Models
by: Dua, Karan, et al.
Published: (2025)
by: Dua, Karan, et al.
Published: (2025)
Semantic Similarity is a Spurious Measure of Comic Understanding: Lessons Learned from Hallucinations in a Benchmarking Experiment
by: Driggers-Ellis, Christopher, et al.
Published: (2026)
by: Driggers-Ellis, Christopher, et al.
Published: (2026)
From Rule-Based Models to Deep Learning Transformers Architectures for Natural Language Processing and Sign Language Translation Systems: Survey, Taxonomy and Performance Evaluation
by: Shahin, Nada, et al.
Published: (2024)
by: Shahin, Nada, et al.
Published: (2024)
GLoT: A Novel Gated-Logarithmic Transformer for Efficient Sign Language Translation
by: Shahin, Nada, et al.
Published: (2025)
by: Shahin, Nada, et al.
Published: (2025)
Scaling Large Vision-Language Models for Enhanced Multimodal Comprehension In Biomedical Image Analysis
by: Umeike, Robinson, et al.
Published: (2025)
by: Umeike, Robinson, et al.
Published: (2025)
Cytoarchitecture in Words: Weakly Supervised Vision-Language Modeling for Human Brain Microscopy
by: Sutton, Matthew, et al.
Published: (2026)
by: Sutton, Matthew, et al.
Published: (2026)
ADAT: Time-Series-Aware Adaptive Transformer Architecture for Sign Language Translation
by: Shahin, Nada, et al.
Published: (2025)
by: Shahin, Nada, et al.
Published: (2025)
Motion-Based Sign Language Video Summarization using Curvature and Torsion
by: Sartinas, Evangelos G., et al.
Published: (2023)
by: Sartinas, Evangelos G., et al.
Published: (2023)
Unified Modeling Language Code Generation from Diagram Images Using Multimodal Large Language Models
by: Bates, Averi, et al.
Published: (2025)
by: Bates, Averi, et al.
Published: (2025)
KidsNanny: A Two-Stage Multimodal Content Moderation Pipeline Integrating Visual Classification, Object Detection, OCR, and Contextual Reasoning for Child Safety
by: Panchal, Viraj, et al.
Published: (2026)
by: Panchal, Viraj, et al.
Published: (2026)
MOVE: A Mixture-of-Vision-Encoders Approach for Domain-Focused Vision-Language Processing
by: Skripkin, Matvey, et al.
Published: (2025)
by: Skripkin, Matvey, et al.
Published: (2025)
DocShield: Towards AI Document Safety via Evidence-Grounded Agentic Reasoning
by: Zeng, Fanwei, et al.
Published: (2026)
by: Zeng, Fanwei, et al.
Published: (2026)
Spatially-Aware Speaker for Vision-and-Language Navigation Instruction Generation
by: Gopinathan, Muraleekrishna, et al.
Published: (2024)
by: Gopinathan, Muraleekrishna, et al.
Published: (2024)
HATL: Hierarchical Adaptive-Transfer Learning Framework for Sign Language Machine Translation
by: Shahin, Nada, et al.
Published: (2026)
by: Shahin, Nada, et al.
Published: (2026)
Visual Aesthetic Benchmark: Can Frontier Models Judge Beauty?
by: Feng, Yichen, et al.
Published: (2026)
by: Feng, Yichen, et al.
Published: (2026)
Robustness of Large Language Models to Perturbations in Text
by: Singh, Ayush, et al.
Published: (2024)
by: Singh, Ayush, et al.
Published: (2024)
Optimizing Multi-Scale Representations to Detect Effect Heterogeneity Using Earth Observation and Computer Vision: Applications to Two Anti-Poverty RCTs
by: Zhu, Fucheng Warren, et al.
Published: (2024)
by: Zhu, Fucheng Warren, et al.
Published: (2024)
Cross-Lingual Generalization and Compression: From Language-Specific to Shared Neurons
by: Riemenschneider, Frederick, et al.
Published: (2025)
by: Riemenschneider, Frederick, et al.
Published: (2025)
MIRAGE: Scaling Test-Time Inference with Parallel Graph-Retrieval-Augmented Reasoning Chains
by: Wei, Kaiwen, et al.
Published: (2025)
by: Wei, Kaiwen, et al.
Published: (2025)
HyperPersona: A Multi-Level Hypergraph Framework for Text-Based Automatic Personality Prediction
by: Heydari, Sina, et al.
Published: (2026)
by: Heydari, Sina, et al.
Published: (2026)
Automating Clinical Information Retrieval from Finnish Electronic Health Records Using Large Language Models
by: Saukkoriipi, Mikko, et al.
Published: (2026)
by: Saukkoriipi, Mikko, et al.
Published: (2026)
Overcoming the Generalization Limits of SLM Finetuning for Shape-Based Extraction of Datatype and Object Properties
by: Ringwald, Célian, et al.
Published: (2025)
by: Ringwald, Célian, et al.
Published: (2025)
From Knowledge Generation to Knowledge Verification: Examining the BioMedical Generative Capabilities of ChatGPT
by: Hamed, Ahmed Abdeen, et al.
Published: (2025)
by: Hamed, Ahmed Abdeen, et al.
Published: (2025)
Pipeline and Dataset Generation for Automated Fact-checking in Almost Any Language
by: Drchal, Jan, et al.
Published: (2023)
by: Drchal, Jan, et al.
Published: (2023)
LongSumEval: Question-Answering Based Evaluation and Feedback-Driven Refinement for Long Document Summarization
by: Nguyen, Huyen, et al.
Published: (2026)
by: Nguyen, Huyen, et al.
Published: (2026)
Bridge Diffusion Model: Bridge Chinese Text-to-Image Diffusion Model with English Communities
by: Liu, Shanyuan, et al.
Published: (2023)
by: Liu, Shanyuan, et al.
Published: (2023)
SAGE: A Strategy-Aware Graph-Enhanced Generation Framework For Online Counseling
by: Aharon, Eliya Naomi, et al.
Published: (2026)
by: Aharon, Eliya Naomi, et al.
Published: (2026)
UAV-assisted Visual SLAM Generating Reconstructed 3D Scene Graphs in GPS-denied Environments
by: Radwan, Ahmed, et al.
Published: (2024)
by: Radwan, Ahmed, et al.
Published: (2024)
ICG: Improving Cover Image Generation via MLLM-based Prompting and Personalized Preference Alignment
by: Bian, Zhipeng, et al.
Published: (2026)
by: Bian, Zhipeng, et al.
Published: (2026)
Improving LLMs with a knowledge from databases
by: Máša, Petr
Published: (2025)
by: Máša, Petr
Published: (2025)
CR-LT-KGQA: A Knowledge Graph Question Answering Dataset Requiring Commonsense Reasoning and Long-Tail Knowledge
by: Guo, Willis, et al.
Published: (2024)
by: Guo, Willis, et al.
Published: (2024)
Kastor: Fine-tuned Small Language Models for Shape-based Active Relation Extraction
by: Celian, Ringwald, et al.
Published: (2025)
by: Celian, Ringwald, et al.
Published: (2025)
Multi-Hierarchical Feature Detection for Large Language Model Generated Text
by: Zhang, Luyan, et al.
Published: (2025)
by: Zhang, Luyan, et al.
Published: (2025)
YETI (YET to Intervene) Proactive Interventions by Multimodal AI Agents in Augmented Reality Tasks
by: Bandyopadhyay, Saptarashmi, et al.
Published: (2025)
by: Bandyopadhyay, Saptarashmi, et al.
Published: (2025)
Technical Report of TeleChat2, TeleChat2.5 and T1
by: Wang, Zihan, et al.
Published: (2025)
by: Wang, Zihan, et al.
Published: (2025)
Retrieval-Based Multi-Label Legal Annotation: Extensible, Data-Efficient and Hallucination-Free
by: Zhang, Li, et al.
Published: (2026)
by: Zhang, Li, et al.
Published: (2026)
Fixed-Budget Parameter-Efficient Training with Frozen Encoders Improves Multimodal Chest X-Ray Classification
by: Khan, Md Ashik, et al.
Published: (2025)
by: Khan, Md Ashik, et al.
Published: (2025)
Similar Items
-
WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation
by: Niu, Yuwei, et al.
Published: (2025) -
Deep Learning based Key Information Extraction from Business Documents: Systematic Literature Review
by: Rombach, Alexander Michael, et al.
Published: (2024) -
Training-Free Diffusion Priors for Text-to-Image Generation via Optimization-based Visual Inversion
by: Dell'Erba, Samuele, et al.
Published: (2025) -
FlexDoc: Parameterized Sampling for Diverse Multilingual Synthetic Documents for Training Document Understanding Models
by: Dua, Karan, et al.
Published: (2025) -
Semantic Similarity is a Spurious Measure of Comic Understanding: Lessons Learned from Hallucinations in a Benchmarking Experiment
by: Driggers-Ellis, Christopher, et al.
Published: (2026)