WalkCLIP: Multimodal Learning for Urban Walkability Prediction
Fuente:
arXiv
Saved in:
| Main Authors: | Xiang, Shilong, Lee, JangHyeon, Namgung, Min, Chiang, Yao-Yi |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Less is More: Multimodal Region Representation via Pairwise Inter-view Learning
by: Namgung, Min, et al.
Published: (2025)
by: Namgung, Min, et al.
Published: (2025)
Transit for All: Mapping Equitable Bike2Subway Connection using Region Representation Learning
by: Namgung, Min, et al.
Published: (2025)
by: Namgung, Min, et al.
Published: (2025)
SceneAware: Scene-Constrained Pedestrian Trajectory Prediction with LLM-Guided Walkability
by: Bai, Juho, et al.
Published: (2025)
by: Bai, Juho, et al.
Published: (2025)
Learning Generalizable Prompt for CLIP with Class Similarity Knowledge
by: Jung, Sehun, et al.
Published: (2025)
by: Jung, Sehun, et al.
Published: (2025)
DesignCLIP: Multimodal Learning with CLIP for Design Patent Understanding
by: Wang, Zhu, et al.
Published: (2025)
by: Wang, Zhu, et al.
Published: (2025)
CLIP Can Understand Depth
by: Kim, Sohee, et al.
Published: (2024)
by: Kim, Sohee, et al.
Published: (2024)
CellCLIP -- Learning Perturbation Effects in Cell Painting via Text-Guided Contrastive Learning
by: Lu, Mingyu, et al.
Published: (2025)
by: Lu, Mingyu, et al.
Published: (2025)
DiffCLIP: Differential Attention Meets CLIP
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2025)
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2025)
CountCLIP -- [Re] Teaching CLIP to Count to Ten
by: Mestha, Harshvardhan, et al.
Published: (2024)
by: Mestha, Harshvardhan, et al.
Published: (2024)
SeMoBridge: Semantic Modality Bridge for Efficient Few-Shot Adaptation of CLIP
by: Timmermann, Christoph, et al.
Published: (2025)
by: Timmermann, Christoph, et al.
Published: (2025)
Cross-modal Affinity-aligned Multimodal Learning Analytics for Predicting Student Collaboration Satisfaction in Game-Based Learning
by: Tsai, Wen-Hsin, et al.
Published: (2026)
by: Tsai, Wen-Hsin, et al.
Published: (2026)
Exploring a Multimodal Fusion-based Deep Learning Network for Detecting Facial Palsy
by: Oo, Heng Yim Nicole, et al.
Published: (2024)
by: Oo, Heng Yim Nicole, et al.
Published: (2024)
Transformer-Based Classification Outcome Prediction for Multimodal Stroke Treatment
by: Ma, Danqing, et al.
Published: (2024)
by: Ma, Danqing, et al.
Published: (2024)
Anchors Aweigh! Sail for Optimal Unified Multi-Modal Representations
by: Jeong, Minoh, et al.
Published: (2024)
by: Jeong, Minoh, et al.
Published: (2024)
Learning from Oblivion: Predicting Knowledge Overflowed Weights via Retrodiction of Forgetting
by: Jang, Jinhyeok, et al.
Published: (2025)
by: Jang, Jinhyeok, et al.
Published: (2025)
VeCLIP: Improving CLIP Training via Visual-enriched Captions
by: Lai, Zhengfeng, et al.
Published: (2023)
by: Lai, Zhengfeng, et al.
Published: (2023)
SynthCLIP: Are We Ready for a Fully Synthetic CLIP Training?
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2024)
by: Hammoud, Hasan Abed Al Kader, et al.
Published: (2024)
On the Reproducibility of "FairCLIP: Harnessing Fairness in Vision-Language Learning''
by: Bakker, Hua Chang, et al.
Published: (2025)
by: Bakker, Hua Chang, et al.
Published: (2025)
Distilled Prompt Learning for Incomplete Multimodal Survival Prediction
by: Xu, Yingxue, et al.
Published: (2025)
by: Xu, Yingxue, et al.
Published: (2025)
A Multimodal Fusion Model Leveraging MLP Mixer and Handcrafted Features-based Deep Learning Networks for Facial Palsy Detection
by: Oo, Heng Yim Nicole, et al.
Published: (2025)
by: Oo, Heng Yim Nicole, et al.
Published: (2025)
CLIP with Generative Latent Replay: a Strong Baseline for Incremental Learning
by: Frascaroli, Emanuele, et al.
Published: (2024)
by: Frascaroli, Emanuele, et al.
Published: (2024)
GoldiCLIP: The Goldilocks Approach for Balancing Explicit Supervision for Language-Image Pretraining
by: Mohan, Deen Dayal, et al.
Published: (2026)
by: Mohan, Deen Dayal, et al.
Published: (2026)
DenseMLLM: Standard Multimodal LLMs for Dense Prediction
by: Li, Yi, et al.
Published: (2026)
by: Li, Yi, et al.
Published: (2026)
Can Multimodal Large Language Models Truly Perform Multimodal In-Context Learning?
by: Chen, Shuo, et al.
Published: (2023)
by: Chen, Shuo, et al.
Published: (2023)
Label Distribution Shift-Aware Prediction Refinement for Test-Time Adaptation
by: Jang, Minguk, et al.
Published: (2024)
by: Jang, Minguk, et al.
Published: (2024)
Soft Task-Aware Routing of Experts for Equivariant Representation Learning
by: Jeon, Jaebyeong, et al.
Published: (2025)
by: Jeon, Jaebyeong, et al.
Published: (2025)
ECOR: Explainable CLIP for Object Recognition
by: Rasekh, Ali, et al.
Published: (2024)
by: Rasekh, Ali, et al.
Published: (2024)
Interpreting CLIP with Hierarchical Sparse Autoencoders
by: Zaigrajew, Vladimir, et al.
Published: (2025)
by: Zaigrajew, Vladimir, et al.
Published: (2025)
EXAONEPath 1.0 Patch-level Foundation Model for Pathology
by: Yun, Juseung, et al.
Published: (2024)
by: Yun, Juseung, et al.
Published: (2024)
Possibilistic Predictive Uncertainty for Deep Learning
by: Ni, Yao, et al.
Published: (2026)
by: Ni, Yao, et al.
Published: (2026)
Reproducibility Study of CDUL: CLIP-Driven Unsupervised Learning for Multi-Label Image Classification
by: Shah, Manan, et al.
Published: (2024)
by: Shah, Manan, et al.
Published: (2024)
ClipTBP: Clip-Pair based Temporal Boundary Prediction with Boundary-Aware Learning for Moment Retrieval
by: Kim, Ji-Hyeon, et al.
Published: (2026)
by: Kim, Ji-Hyeon, et al.
Published: (2026)
Zoom-shot: Fast and Efficient Unsupervised Zero-Shot Transfer of CLIP to Vision Encoders with Multimodal Loss
by: Shipard, Jordan, et al.
Published: (2024)
by: Shipard, Jordan, et al.
Published: (2024)
Implicit Inversion turns CLIP into a Decoder
by: D'Orazio, Antonio, et al.
Published: (2025)
by: D'Orazio, Antonio, et al.
Published: (2025)
Steering CLIP's vision transformer with sparse autoencoders
by: Joseph, Sonia, et al.
Published: (2025)
by: Joseph, Sonia, et al.
Published: (2025)
IDEA: Image Description Enhanced CLIP-Adapter
by: Ye, Zhipeng, et al.
Published: (2025)
by: Ye, Zhipeng, et al.
Published: (2025)
Strong but simple: A Baseline for Domain Generalized Dense Perception by CLIP-based Transfer Learning
by: Hümmer, Christoph, et al.
Published: (2023)
by: Hümmer, Christoph, et al.
Published: (2023)
Visual Representation Learning with Stochastic Frame Prediction
by: Jang, Huiwon, et al.
Published: (2024)
by: Jang, Huiwon, et al.
Published: (2024)
UrbanGraphEmbeddings: Learning and Evaluating Spatially Grounded Multimodal Embeddings for Urban Science
by: Zhang, Jie, et al.
Published: (2026)
by: Zhang, Jie, et al.
Published: (2026)
CLIP-based Camera-Agnostic Feature Learning for Intra-camera Person Re-Identification
by: Tan, Xuan, et al.
Published: (2024)
by: Tan, Xuan, et al.
Published: (2024)
Similar Items
-
Less is More: Multimodal Region Representation via Pairwise Inter-view Learning
by: Namgung, Min, et al.
Published: (2025) -
Transit for All: Mapping Equitable Bike2Subway Connection using Region Representation Learning
by: Namgung, Min, et al.
Published: (2025) -
SceneAware: Scene-Constrained Pedestrian Trajectory Prediction with LLM-Guided Walkability
by: Bai, Juho, et al.
Published: (2025) -
Learning Generalizable Prompt for CLIP with Class Similarity Knowledge
by: Jung, Sehun, et al.
Published: (2025) -
DesignCLIP: Multimodal Learning with CLIP for Design Patent Understanding
by: Wang, Zhu, et al.
Published: (2025)