Improving Region Representation Learning from Urban Imagery with Noisy Long-Caption Supervision
Fuente:
arXiv
Salvato in:
| Autori principali: | Zhang, Yimei, Shen, Guojiang, Ning, Kaili, Ren, Tongwei, Qiu, Xuebo, Wang, Mengmeng, Kong, Xiangjie |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
VISAFF: Speaker-Centered Visual Affective Feature Learning for Emotion Recognition in Conversation
di: ZHU, Linan, et al.
Pubblicazione: (2026)
di: ZHU, Linan, et al.
Pubblicazione: (2026)
GARLIC: GPT-Augmented Reinforcement Learning with Intelligent Control for Vehicle Dispatching
di: Han, Xiao, et al.
Pubblicazione: (2024)
di: Han, Xiao, et al.
Pubblicazione: (2024)
Deep Learning for Time Series Forecasting: A Survey
di: Kong, Xiangjie, et al.
Pubblicazione: (2025)
di: Kong, Xiangjie, et al.
Pubblicazione: (2025)
UrbanVerse: Learning Urban Region Representation Across Cities and Tasks
di: Sun, Fengze, et al.
Pubblicazione: (2026)
di: Sun, Fengze, et al.
Pubblicazione: (2026)
Boundary Prompting: Elastic Urban Region Representation via Graph-based Spatial Tokenization
di: Zhu, Haojia, et al.
Pubblicazione: (2025)
di: Zhu, Haojia, et al.
Pubblicazione: (2025)
Learning with Imbalanced Noisy Data by Preventing Bias in Sample Selection
di: Liu, Huafeng, et al.
Pubblicazione: (2024)
di: Liu, Huafeng, et al.
Pubblicazione: (2024)
WISTERIA: Learning Clinical Representations from Noisy Supervision via Multi-View Consistency in Electronic Health Records
di: Dong, Ruan, et al.
Pubblicazione: (2026)
di: Dong, Ruan, et al.
Pubblicazione: (2026)
ToPT: Task-Oriented Prompt Tuning for Urban Region Representation Learning
di: Guo, Zitao, et al.
Pubblicazione: (2026)
di: Guo, Zitao, et al.
Pubblicazione: (2026)
Rule-driven News Captioning
di: Xu, Ning, et al.
Pubblicazione: (2024)
di: Xu, Ning, et al.
Pubblicazione: (2024)
MobiCLR: Mobility Time Series Contrastive Learning for Urban Region Representations
di: Kim, Namwoo, et al.
Pubblicazione: (2025)
di: Kim, Namwoo, et al.
Pubblicazione: (2025)
Controllable Hybrid Captioner for Improved Long-form Video Understanding
di: Sasse, Kuleen, et al.
Pubblicazione: (2025)
di: Sasse, Kuleen, et al.
Pubblicazione: (2025)
Retrieval Augmented Learning: A Retrial-based Large Language Model Self-Supervised Learning and Autonomous Knowledge Generation
di: Li, Zongyuan, et al.
Pubblicazione: (2025)
di: Li, Zongyuan, et al.
Pubblicazione: (2025)
Can LLMs Learn to Reason Robustly under Noisy Supervision?
di: Yang, Shenzhi, et al.
Pubblicazione: (2026)
di: Yang, Shenzhi, et al.
Pubblicazione: (2026)
Weakly Supervised Point Clouds Transformer for 3D Object Detection
di: Tang, Zuojin, et al.
Pubblicazione: (2023)
di: Tang, Zuojin, et al.
Pubblicazione: (2023)
Mining Fine-Grained Image-Text Alignment for Zero-Shot Captioning via Text-Only Training
di: Qiu, Longtian, et al.
Pubblicazione: (2024)
di: Qiu, Longtian, et al.
Pubblicazione: (2024)
Pattern-Matching Dynamic Memory Network for Dual-Mode Traffic Prediction
di: Weng, Wenchao, et al.
Pubblicazione: (2024)
di: Weng, Wenchao, et al.
Pubblicazione: (2024)
FedRG: Unleashing the Representation Geometry for Federated Learning with Noisy Clients
di: Wen, Tian, et al.
Pubblicazione: (2026)
di: Wen, Tian, et al.
Pubblicazione: (2026)
SAIL: Similarity-Aware Guidance and Inter-Caption Augmentation-based Learning for Weakly-Supervised Dense Video Captioning
di: Kim, Ye-Chan, et al.
Pubblicazione: (2026)
di: Kim, Ye-Chan, et al.
Pubblicazione: (2026)
Multimodal Trajectory Representation Learning for Travel Time Estimation
di: Liu, Zhi, et al.
Pubblicazione: (2025)
di: Liu, Zhi, et al.
Pubblicazione: (2025)
CLAP: Learning Transferable Binary Code Representations with Natural Language Supervision
di: Wang, Hao, et al.
Pubblicazione: (2024)
di: Wang, Hao, et al.
Pubblicazione: (2024)
Improving Attributed Text Generation of Large Language Models via Preference Learning
di: Li, Dongfang, et al.
Pubblicazione: (2024)
di: Li, Dongfang, et al.
Pubblicazione: (2024)
Noisy-Pair Robust Representation Alignment for Positive-Unlabeled Learning
di: Zhao, Hengwei, et al.
Pubblicazione: (2025)
di: Zhao, Hengwei, et al.
Pubblicazione: (2025)
Augment, Drop & Swap: Improving Diversity in LLM Captions for Efficient Music-Text Representation Learning
di: Manco, Ilaria, et al.
Pubblicazione: (2024)
di: Manco, Ilaria, et al.
Pubblicazione: (2024)
AGIC: Attention-Guided Image Captioning to Improve Caption Relevance
di: Teja, L. D. M. S. Sai, et al.
Pubblicazione: (2025)
di: Teja, L. D. M. S. Sai, et al.
Pubblicazione: (2025)
LooGLE: Can Long-Context Language Models Understand Long Contexts?
di: Li, Jiaqi, et al.
Pubblicazione: (2023)
di: Li, Jiaqi, et al.
Pubblicazione: (2023)
Denoising-Aware Contrastive Learning for Noisy Time Series
di: Zhou, Shuang, et al.
Pubblicazione: (2024)
di: Zhou, Shuang, et al.
Pubblicazione: (2024)
URECA: Unique Region Caption Anything
di: Lim, Sangbeom, et al.
Pubblicazione: (2025)
di: Lim, Sangbeom, et al.
Pubblicazione: (2025)
Impact of Noisy Supervision in Foundation Model Learning
di: Chen, Hao, et al.
Pubblicazione: (2024)
di: Chen, Hao, et al.
Pubblicazione: (2024)
LLM Agent Framework for Intelligent Change Analysis in Urban Environment using Remote Sensing Imagery
di: Xiao, Zixuan, et al.
Pubblicazione: (2026)
di: Xiao, Zixuan, et al.
Pubblicazione: (2026)
Position Prediction Self-Supervised Learning for Multimodal Satellite Imagery Semantic Segmentation
di: Waithaka, John, et al.
Pubblicazione: (2025)
di: Waithaka, John, et al.
Pubblicazione: (2025)
Learning to Model Graph Structural Information on MLPs via Graph Structure Self-Contrasting
di: Wu, Lirong, et al.
Pubblicazione: (2024)
di: Wu, Lirong, et al.
Pubblicazione: (2024)
Relation Modeling and Distillation for Learning with Noisy Labels
di: Che, Xiaming, et al.
Pubblicazione: (2024)
di: Che, Xiaming, et al.
Pubblicazione: (2024)
UrbanKGent: A Unified Large Language Model Agent Framework for Urban Knowledge Graph Construction
di: Ning, Yansong, et al.
Pubblicazione: (2024)
di: Ning, Yansong, et al.
Pubblicazione: (2024)
Multimodal Contrastive Learning of Urban Space Representations from POI Data
di: Wang, Xinglei, et al.
Pubblicazione: (2024)
di: Wang, Xinglei, et al.
Pubblicazione: (2024)
TrackAny3D: Transferring Pretrained 3D Models for Category-unified 3D Point Cloud Tracking
di: Wang, Mengmeng, et al.
Pubblicazione: (2025)
di: Wang, Mengmeng, et al.
Pubblicazione: (2025)
Graph-Based Captioning: Enhancing Visual Descriptions by Interconnecting Region Captions
di: Hsieh, Yu-Guan, et al.
Pubblicazione: (2024)
di: Hsieh, Yu-Guan, et al.
Pubblicazione: (2024)
SRA 2: Variational Autoencoder Self-Representation Alignment for Efficient Diffusion Training
di: Wang, Mengmeng, et al.
Pubblicazione: (2026)
di: Wang, Mengmeng, et al.
Pubblicazione: (2026)
Learning from Noisy Labels for Long-tailed Data via Optimal Transport
di: Li, Mengting, et al.
Pubblicazione: (2024)
di: Li, Mengting, et al.
Pubblicazione: (2024)
Seeing vs. Believing: Evaluating the Language Bias of Open-Source MLLMs in Counter-Intuitive Scenes
di: Ling, Chen, et al.
Pubblicazione: (2026)
di: Ling, Chen, et al.
Pubblicazione: (2026)
Mitigating Object Dependencies: Improving Point Cloud Self-Supervised Learning through Object Exchange
di: Wu, Yanhao, et al.
Pubblicazione: (2024)
di: Wu, Yanhao, et al.
Pubblicazione: (2024)
Documenti analoghi
-
VISAFF: Speaker-Centered Visual Affective Feature Learning for Emotion Recognition in Conversation
di: ZHU, Linan, et al.
Pubblicazione: (2026) -
GARLIC: GPT-Augmented Reinforcement Learning with Intelligent Control for Vehicle Dispatching
di: Han, Xiao, et al.
Pubblicazione: (2024) -
Deep Learning for Time Series Forecasting: A Survey
di: Kong, Xiangjie, et al.
Pubblicazione: (2025) -
UrbanVerse: Learning Urban Region Representation Across Cities and Tasks
di: Sun, Fengze, et al.
Pubblicazione: (2026) -
Boundary Prompting: Elastic Urban Region Representation via Graph-based Spatial Tokenization
di: Zhu, Haojia, et al.
Pubblicazione: (2025)