Saved in:
| Main Authors: | Geng, Yuang, Zhou, Zhuoyang, Zhang, Zhongzheng, Pan, Siyuan, Tran, Hoang-Dung, Ruchkin, Ivan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2512.08991 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MLE-UVAD: Minimal Latent Entropy Autoencoder for Fully Unsupervised Video Anomaly Detection
by: Geng, Yuang, et al.
Published: (2026)
by: Geng, Yuang, et al.
Published: (2026)
Four Principles for Physically Interpretable World Models
by: Peper, Jordan, et al.
Published: (2025)
by: Peper, Jordan, et al.
Published: (2025)
Evaluating Precise Geolocation Inference Capabilities of Vision Language Models
by: Jay, Neel, et al.
Published: (2025)
by: Jay, Neel, et al.
Published: (2025)
Ultrasound Vision-Language Alignment via Contrastive Learning
by: Lyu, Zhuoyang, et al.
Published: (2026)
by: Lyu, Zhuoyang, et al.
Published: (2026)
Language-Enhanced Latent Representations for Out-of-Distribution Detection in Autonomous Driving
by: Mao, Zhenjiang, et al.
Published: (2024)
by: Mao, Zhenjiang, et al.
Published: (2024)
Hypergraph Laplacian Eigenmaps and Face Recognition Problems
by: Tran, Loc Hoang
Published: (2024)
by: Tran, Loc Hoang
Published: (2024)
Space Syntax-guided Post-training for Residential Floor Plan Generation
by: Jiang, Zhuoyang, et al.
Published: (2026)
by: Jiang, Zhuoyang, et al.
Published: (2026)
KOPPA: Improving Prompt-based Continual Learning with Key-Query Orthogonal Projection and Prototype-based One-Versus-All
by: Tran, Quyen, et al.
Published: (2023)
by: Tran, Quyen, et al.
Published: (2023)
Learned Image Compression with Text Quality Enhancement
by: Lai, Chih-Yu, et al.
Published: (2024)
by: Lai, Chih-Yu, et al.
Published: (2024)
Gradient-Aligned Calibration for Post-Training Quantization of Diffusion Models
by: Hoang, Dung Anh, et al.
Published: (2026)
by: Hoang, Dung Anh, et al.
Published: (2026)
Enhancing Domain Adaptation through Prompt Gradient Alignment
by: Phan, Hoang, et al.
Published: (2024)
by: Phan, Hoang, et al.
Published: (2024)
Provably Improving Generalization of Few-Shot Models with Synthetic Data
by: Nguyen, Lan-Cuong, et al.
Published: (2025)
by: Nguyen, Lan-Cuong, et al.
Published: (2025)
EfficientViT-SAM: Accelerated Segment Anything Model Without Accuracy Loss
by: Zhang, Zhuoyang, et al.
Published: (2024)
by: Zhang, Zhuoyang, et al.
Published: (2024)
Progressive Multi-granular Alignments for Grounded Reasoning in Large Vision-Language Models
by: Le, Quang-Hung, et al.
Published: (2024)
by: Le, Quang-Hung, et al.
Published: (2024)
Zero-shot Safety Prediction for Autonomous Robots with Foundation World Models
by: Mao, Zhenjiang, et al.
Published: (2024)
by: Mao, Zhenjiang, et al.
Published: (2024)
AdaWorld: Learning Adaptable World Models with Latent Actions
by: Gao, Shenyuan, et al.
Published: (2025)
by: Gao, Shenyuan, et al.
Published: (2025)
Continual Visual Reinforcement Learning with A Life-Long World Model
by: Pan, Minting, et al.
Published: (2023)
by: Pan, Minting, et al.
Published: (2023)
Active Intelligence in Video Avatars via Closed-loop World Modeling
by: He, Xuanhua, et al.
Published: (2025)
by: He, Xuanhua, et al.
Published: (2025)
Cloning Deterministic Worlds: The Critical Role of Latent Geometry in Long-Horizon World Models
by: Xia, Zaishuo, et al.
Published: (2025)
by: Xia, Zaishuo, et al.
Published: (2025)
Geometric Regularity in Deterministic Sampling Dynamics of Diffusion-based Generative Models
by: Chen, Defang, et al.
Published: (2025)
by: Chen, Defang, et al.
Published: (2025)
Robust Self-Training with Closed-loop Label Correction for Learning from Noisy Labels
by: Lin, Zhanhui, et al.
Published: (2026)
by: Lin, Zhanhui, et al.
Published: (2026)
PVI: Plug-in Visual Injection for Vision-Language-Action Models
by: Zhang, Zezhou, et al.
Published: (2026)
by: Zhang, Zezhou, et al.
Published: (2026)
Maximising the Utility of Validation Sets for Imbalanced Noisy-label Meta-learning
by: Hoang, Dung Anh, et al.
Published: (2022)
by: Hoang, Dung Anh, et al.
Published: (2022)
Sharpness-Aware Data Generation for Zero-shot Quantization
by: Hoang-Anh, Dung, et al.
Published: (2025)
by: Hoang-Anh, Dung, et al.
Published: (2025)
MOOSE: Pay Attention to Temporal Dynamics for Video Understanding via Optical Flows
by: Nguyen, Hong, et al.
Published: (2025)
by: Nguyen, Hong, et al.
Published: (2025)
CAKE: Real-time Action Detection via Motion Distillation and Background-aware Contrastive Learning
by: Hoang, Hieu, et al.
Published: (2026)
by: Hoang, Hieu, et al.
Published: (2026)
Navigation World Models
by: Bar, Amir, et al.
Published: (2024)
by: Bar, Amir, et al.
Published: (2024)
Consistent3D: Towards Consistent High-Fidelity Text-to-3D Generation with Deterministic Sampling Prior
by: Wu, Zike, et al.
Published: (2024)
by: Wu, Zike, et al.
Published: (2024)
Doe-1: Closed-Loop Autonomous Driving with Large World Model
by: Zheng, Wenzhao, et al.
Published: (2024)
by: Zheng, Wenzhao, et al.
Published: (2024)
Attention Guided Alignment in Efficient Vision-Language Models
by: Mahajan, Shweta, et al.
Published: (2025)
by: Mahajan, Shweta, et al.
Published: (2025)
LLM-attacker: Enhancing Closed-loop Adversarial Scenario Generation for Autonomous Driving with Large Language Models
by: Mei, Yuewen, et al.
Published: (2025)
by: Mei, Yuewen, et al.
Published: (2025)
Pix2Fact: When Vision Is Not Enough -- Benchmarking Fine-Grained VQA with Web Verification on High-Resolution Real-World Scenes
by: Jiang, Yifan, et al.
Published: (2026)
by: Jiang, Yifan, et al.
Published: (2026)
RefDecoder: Enhancing Visual Generation with Conditional Video Decoding
by: Fan, Xiang, et al.
Published: (2026)
by: Fan, Xiang, et al.
Published: (2026)
Continual Retinal Vision-Language Pre-training upon Incremental Imaging Modalities
by: Yao, Yuang, et al.
Published: (2025)
by: Yao, Yuang, et al.
Published: (2025)
Vision Verification Enhanced Fusion of VLMs for Efficient Visual Reasoning
by: Tekin, Selim Furkan, et al.
Published: (2026)
by: Tekin, Selim Furkan, et al.
Published: (2026)
Universal Camouflage Attack on Vision-Language Models for Autonomous Driving
by: Kong, Dehong, et al.
Published: (2025)
by: Kong, Dehong, et al.
Published: (2025)
Stochastic Sampling from Deterministic Flow Models
by: Singh, Saurabh, et al.
Published: (2024)
by: Singh, Saurabh, et al.
Published: (2024)
Generalization Bounds for Robust Contrastive Learning: From Theory to Practice
by: Tran, Ngoc N., et al.
Published: (2023)
by: Tran, Ngoc N., et al.
Published: (2023)
Vision-Language Model Based Handwriting Verification
by: Chauhan, Mihir, et al.
Published: (2024)
by: Chauhan, Mihir, et al.
Published: (2024)
Reading $\neq$ Seeing: Diagnosing and Closing the Typography Gap in Vision-Language Models
by: Zhou, Heng, et al.
Published: (2026)
by: Zhou, Heng, et al.
Published: (2026)
Similar Items
-
MLE-UVAD: Minimal Latent Entropy Autoencoder for Fully Unsupervised Video Anomaly Detection
by: Geng, Yuang, et al.
Published: (2026) -
Four Principles for Physically Interpretable World Models
by: Peper, Jordan, et al.
Published: (2025) -
Evaluating Precise Geolocation Inference Capabilities of Vision Language Models
by: Jay, Neel, et al.
Published: (2025) -
Ultrasound Vision-Language Alignment via Contrastive Learning
by: Lyu, Zhuoyang, et al.
Published: (2026) -
Language-Enhanced Latent Representations for Out-of-Distribution Detection in Autonomous Driving
by: Mao, Zhenjiang, et al.
Published: (2024)