UrbanAlign: Post-hoc Semantic Calibration for VLM-Human Preference Alignment
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhang, Yecheng, Zhao, Rong, Sha, Zhizhou, Li, Yong, Wang, Lei, Hou, Ce, Ji, Wen, Huang, Hao, Wan, Yunshan, Yu, Jian, Xia, Junhao, Zhang, Yuru, Shi, Chunlei |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence
von: Zeng, Zhitao, et al.
Veröffentlicht: (2025)
von: Zeng, Zhitao, et al.
Veröffentlicht: (2025)
ClustViT: Clustering-based Token Merging for Semantic Segmentation
von: Montello, Fabio, et al.
Veröffentlicht: (2025)
von: Montello, Fabio, et al.
Veröffentlicht: (2025)
VIAFormer: Voxel-Image Alignment Transformer for High-Fidelity Voxel Refinement
von: Fang, Tiancheng, et al.
Veröffentlicht: (2026)
von: Fang, Tiancheng, et al.
Veröffentlicht: (2026)
Semantic Prioritization in Visual Counterfactual Explanations with Weighted Segmentation and Auto-Adaptive Region Selection
von: Zhang, Lintong, et al.
Veröffentlicht: (2025)
von: Zhang, Lintong, et al.
Veröffentlicht: (2025)
DiffYOLO: Object Detection for Anti-Noise via YOLO and Diffusion Models
von: Liu, Yichen, et al.
Veröffentlicht: (2024)
von: Liu, Yichen, et al.
Veröffentlicht: (2024)
Surg$Σ$: A Spectrum of Large-Scale Multimodal Data and Foundation Models for Surgical Intelligence
von: Zeng, Zhitao, et al.
Veröffentlicht: (2026)
von: Zeng, Zhitao, et al.
Veröffentlicht: (2026)
VLM-NCD:Novel Class Discovery with Vision-Based Large Language Models
von: Su, Yuetong, et al.
Veröffentlicht: (2025)
von: Su, Yuetong, et al.
Veröffentlicht: (2025)
One-to-Normal: Anomaly Personalization for Few-shot Anomaly Detection
von: Li, Yiyue, et al.
Veröffentlicht: (2025)
von: Li, Yiyue, et al.
Veröffentlicht: (2025)
VDPP: Video Depth Post-Processing for Speed and Scalability
von: Yoon, Daewon, et al.
Veröffentlicht: (2026)
von: Yoon, Daewon, et al.
Veröffentlicht: (2026)
VA-$π$: Variational Policy Alignment for Pixel-Aware Autoregressive Generation
von: Liao, Xinyao, et al.
Veröffentlicht: (2025)
von: Liao, Xinyao, et al.
Veröffentlicht: (2025)
Semi-supervised Latent Disentangled Diffusion Model for Textile Pattern Generation
von: Hu, Chenggong, et al.
Veröffentlicht: (2026)
von: Hu, Chenggong, et al.
Veröffentlicht: (2026)
Meaning over Motion: A Semantic-First Approach to 360° Viewport Prediction
von: Khah, Arman Nik, et al.
Veröffentlicht: (2026)
von: Khah, Arman Nik, et al.
Veröffentlicht: (2026)
Predicting Pedestrian Crossing Behavior in Germany and Japan: Insights into Model Transferability
von: Zhang, Chi, et al.
Veröffentlicht: (2024)
von: Zhang, Chi, et al.
Veröffentlicht: (2024)
Predicting and Analyzing Pedestrian Crossing Behavior at Unsignalized Crossings
von: Zhang, Chi, et al.
Veröffentlicht: (2024)
von: Zhang, Chi, et al.
Veröffentlicht: (2024)
Generalization performance of neural mapping schemes for the space-time interpolation of satellite-derived ocean colour datasets
von: Nguyen, Thi Thuy Nga, et al.
Veröffentlicht: (2025)
von: Nguyen, Thi Thuy Nga, et al.
Veröffentlicht: (2025)
Deep Learning Approaches for Human Action Recognition in Video Data
von: Xie, Yufei
Veröffentlicht: (2024)
von: Xie, Yufei
Veröffentlicht: (2024)
LatentForensics: Towards frugal deepfake detection in the StyleGAN latent space
von: Delmas, Matthieu, et al.
Veröffentlicht: (2023)
von: Delmas, Matthieu, et al.
Veröffentlicht: (2023)
Synthetic Industrial Object Detection: GenAI vs. Feature-Based Methods
von: Araya-Martinez, Jose Moises, et al.
Veröffentlicht: (2025)
von: Araya-Martinez, Jose Moises, et al.
Veröffentlicht: (2025)
Zero-Shot Multi-Criteria Visual Quality Inspection for Semi-Controlled Industrial Environments via Real-Time 3D Digital Twin Simulation
von: Araya-Martinez, Jose Moises, et al.
Veröffentlicht: (2025)
von: Araya-Martinez, Jose Moises, et al.
Veröffentlicht: (2025)
Multi-scale Temporal Prediction via Incremental Generation and Multi-agent Collaboration
von: Zeng, Zhitao, et al.
Veröffentlicht: (2025)
von: Zeng, Zhitao, et al.
Veröffentlicht: (2025)
Observation-only learning of neural mapping schemes for gappy satellite-derived ocean colour parameters
von: Dorffer, Clément, et al.
Veröffentlicht: (2025)
von: Dorffer, Clément, et al.
Veröffentlicht: (2025)
PlaneSAM: Multimodal Plane Instance Segmentation Using the Segment Anything Model
von: Deng, Zhongchen, et al.
Veröffentlicht: (2024)
von: Deng, Zhongchen, et al.
Veröffentlicht: (2024)
Extrapolating and Decoupling Image-to-Video Generation Models: Motion Modeling is Easier Than You Think
von: Tian, Jie, et al.
Veröffentlicht: (2025)
von: Tian, Jie, et al.
Veröffentlicht: (2025)
SynthRender and IRIS: Open-Source Framework and Dataset for Bidirectional Sim-Real Transfer in Industrial Object Perception
von: Araya-Martinez, Jose Moises, et al.
Veröffentlicht: (2026)
von: Araya-Martinez, Jose Moises, et al.
Veröffentlicht: (2026)
Video-CoE: Reinforcing Video Event Prediction via Chain of Events
von: Su, Qile, et al.
Veröffentlicht: (2026)
von: Su, Qile, et al.
Veröffentlicht: (2026)
Learning Discriminative Spatio-temporal Representations for Semi-supervised Action Recognition
von: Wang, Yu, et al.
Veröffentlicht: (2024)
von: Wang, Yu, et al.
Veröffentlicht: (2024)
Force-Aware 3D Contact Modeling for Stable Grasp Generation
von: Chen, Zhuo, et al.
Veröffentlicht: (2025)
von: Chen, Zhuo, et al.
Veröffentlicht: (2025)
Lightweight Low-SNR-Robust Semantic Communication System for Autonomous Driving
von: Ren, Ruixing, et al.
Veröffentlicht: (2026)
von: Ren, Ruixing, et al.
Veröffentlicht: (2026)
Web-Scale Collection of Video Data for 4D Animal Reconstruction
von: Zhao, Brian Nlong, et al.
Veröffentlicht: (2025)
von: Zhao, Brian Nlong, et al.
Veröffentlicht: (2025)
LESV: Language Embedded Sparse Voxel Fusion for Open-Vocabulary 3D Scene Understanding
von: Wang, Fusang, et al.
Veröffentlicht: (2026)
von: Wang, Fusang, et al.
Veröffentlicht: (2026)
Flexible-weighted Chamfer Distance: Enhanced Objective Function for Point Cloud Completion
von: Li, Jie, et al.
Veröffentlicht: (2025)
von: Li, Jie, et al.
Veröffentlicht: (2025)
RGB-Only Gaussian Splatting SLAM for Unbounded Outdoor Scenes
von: Yu, Sicheng, et al.
Veröffentlicht: (2025)
von: Yu, Sicheng, et al.
Veröffentlicht: (2025)
J-NeuS: Joint field optimization for Neural Surface reconstruction in urban scenes with limited image overlap
von: Wang, Fusang, et al.
Veröffentlicht: (2025)
von: Wang, Fusang, et al.
Veröffentlicht: (2025)
BUFF: Bayesian Uncertainty Guided Diffusion Probabilistic Model for Single Image Super-Resolution
von: He, Zihao, et al.
Veröffentlicht: (2025)
von: He, Zihao, et al.
Veröffentlicht: (2025)
From Gaze to Insight: Bridging Human Visual Attention and Vision Language Model Explanation for Weakly-Supervised Medical Image Segmentation
von: Chen, Jingkun, et al.
Veröffentlicht: (2025)
von: Chen, Jingkun, et al.
Veröffentlicht: (2025)
Quantifying and Narrowing the Unknown: Interactive Text-to-Video Retrieval via Uncertainty Minimization
von: Zhang, Bingqing, et al.
Veröffentlicht: (2025)
von: Zhang, Bingqing, et al.
Veröffentlicht: (2025)
Learning Association via Track-Detection Matching for Multi-Object Tracking
von: Adžemović, Momir
Veröffentlicht: (2025)
von: Adžemović, Momir
Veröffentlicht: (2025)
AMANet: Advancing SAR Ship Detection with Adaptive Multi-Hierarchical Attention Network
von: Ma, Xiaolin, et al.
Veröffentlicht: (2024)
von: Ma, Xiaolin, et al.
Veröffentlicht: (2024)
3D Reconstruction from Sketches
von: Talwar, Abhimanyu, et al.
Veröffentlicht: (2025)
von: Talwar, Abhimanyu, et al.
Veröffentlicht: (2025)
AI-Enhanced Precision in Sport Taekwondo: Increasing Fairness, Speed, and Trust in Competition (FST.ai)
von: Shariatmadar, Keivan, et al.
Veröffentlicht: (2025)
von: Shariatmadar, Keivan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
SurgVLM: A Large Vision-Language Model and Systematic Evaluation Benchmark for Surgical Intelligence
von: Zeng, Zhitao, et al.
Veröffentlicht: (2025) -
ClustViT: Clustering-based Token Merging for Semantic Segmentation
von: Montello, Fabio, et al.
Veröffentlicht: (2025) -
VIAFormer: Voxel-Image Alignment Transformer for High-Fidelity Voxel Refinement
von: Fang, Tiancheng, et al.
Veröffentlicht: (2026) -
Semantic Prioritization in Visual Counterfactual Explanations with Weighted Segmentation and Auto-Adaptive Region Selection
von: Zhang, Lintong, et al.
Veröffentlicht: (2025) -
DiffYOLO: Object Detection for Anti-Noise via YOLO and Diffusion Models
von: Liu, Yichen, et al.
Veröffentlicht: (2024)