Saved in:
| Main Authors: | Kong, Lingdong, Sun, Xian, Chow, Wei, Li, Linfeng, Lin, Kevin Qinghong, Zhang, Xuan Billy, Wang, Song, Li, Rong, Wu, Qing, Gao, Wei, Wang, Yingshuo, Xie, Shaoyuan, Liu, Jiachen, Qu, Leigang, Li, Shijie, Ng, Lai Xing, Cottereau, Benoit R., Liu, Ziwei, Chua, Tat-Seng, Ooi, Wei Tsang |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2605.18661 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
OpenESS: Event-based Semantic Scene Understanding with Open Vocabularies
by: Kong, Lingdong, et al.
Published: (2024)
by: Kong, Lingdong, et al.
Published: (2024)
EventFly: Event Camera Perception from Ground to the Sky
by: Kong, Lingdong, et al.
Published: (2025)
by: Kong, Lingdong, et al.
Published: (2025)
Visual Grounding from Event Cameras
by: Kong, Lingdong, et al.
Published: (2025)
by: Kong, Lingdong, et al.
Published: (2025)
Talk2Event: Grounded Understanding of Dynamic Scenes from Event Cameras
by: Kong, Lingdong, et al.
Published: (2025)
by: Kong, Lingdong, et al.
Published: (2025)
Composed Image Retrieval with Text Feedback via Multi-grained Uncertainty Regularization
by: Chen, Yiyang, et al.
Published: (2022)
by: Chen, Yiyang, et al.
Published: (2022)
NExT-GPT: Any-to-Any Multimodal LLM
by: Wu, Shengqiong, et al.
Published: (2023)
by: Wu, Shengqiong, et al.
Published: (2023)
Learning to Remove Lens Flare in Event Camera
by: Han, Haiqian, et al.
Published: (2025)
by: Han, Haiqian, et al.
Published: (2025)
Generative Cross-Modal Retrieval: Memorizing Images in Multimodal Language Models for Retrieval and Beyond
by: Li, Yongqi, et al.
Published: (2024)
by: Li, Yongqi, et al.
Published: (2024)
Revolutionizing Text-to-Image Retrieval as Autoregressive Token-to-Voken Generation
by: Li, Yongqi, et al.
Published: (2024)
by: Li, Yongqi, et al.
Published: (2024)
SILMM: Self-Improving Large Multimodal Models for Compositional Text-to-Image Generation
by: Qu, Leigang, et al.
Published: (2024)
by: Qu, Leigang, et al.
Published: (2024)
Discriminative Probing and Tuning for Text-to-Image Generation
by: Qu, Leigang, et al.
Published: (2024)
by: Qu, Leigang, et al.
Published: (2024)
Learning to Generate 4D LiDAR Sequences
by: Liang, Ao, et al.
Published: (2025)
by: Liang, Ao, et al.
Published: (2025)
LiDARCrafter: Dynamic 4D World Modeling from LiDAR Sequences
by: Liang, Ao, et al.
Published: (2025)
by: Liang, Ao, et al.
Published: (2025)
TIGeR: Unifying Text-to-Image Generation and Retrieval with Large Multimodal Models
by: Qu, Leigang, et al.
Published: (2024)
by: Qu, Leigang, et al.
Published: (2024)
Auto-Encoding Morph-Tokens for Multimodal LLM
by: Pan, Kaihang, et al.
Published: (2024)
by: Pan, Kaihang, et al.
Published: (2024)
S ee 4D: Pose‐Free 4D Generation via Auto‐Regressive Video Inpainting
by: Dongyue Lu, et al.
Published: (2026)
by: Dongyue Lu, et al.
Published: (2026)
See4D: Pose-Free 4D Generation via Auto-Regressive Video Inpainting
by: Lu, Dongyue, et al.
Published: (2025)
by: Lu, Dongyue, et al.
Published: (2025)
TTOM: Test-Time Optimization and Memorization for Compositional Video Generation
by: Qu, Leigang, et al.
Published: (2025)
by: Qu, Leigang, et al.
Published: (2025)
Is Your Driving World Model an All-Around Player?
by: Kong, Lingdong, et al.
Published: (2026)
by: Kong, Lingdong, et al.
Published: (2026)
Multi-Modal Data-Efficient 3D Scene Understanding for Autonomous Driving
by: Kong, Lingdong, et al.
Published: (2024)
by: Kong, Lingdong, et al.
Published: (2024)
Dynamic Multimodal Fusion via Meta-Learning Towards Micro-Video Recommendation
by: Liu, Han, et al.
Published: (2025)
by: Liu, Han, et al.
Published: (2025)
Perspective-Invariant 3D Object Detection
by: Liang, Ao, et al.
Published: (2025)
by: Liang, Ao, et al.
Published: (2025)
Principled Multimodal Representation Learning
by: Liu, Xiaohao, et al.
Published: (2025)
by: Liu, Xiaohao, et al.
Published: (2025)
Continual Multimodal Contrastive Learning
by: Liu, Xiaohao, et al.
Published: (2025)
by: Liu, Xiaohao, et al.
Published: (2025)
FlexEvent: Towards Flexible Event-Frame Object Detection at Varying Operational Frequencies
by: Lu, Dongyue, et al.
Published: (2024)
by: Lu, Dongyue, et al.
Published: (2024)
Understanding Long Videos via LLM-Powered Entity Relation Graphs
by: Chu, Meng, et al.
Published: (2025)
by: Chu, Meng, et al.
Published: (2025)
Learning to Ask Critical Questions for Assisting Product Search
by: Li, Zixuan, et al.
Published: (2024)
by: Li, Zixuan, et al.
Published: (2024)
Turing Patterns for Multimedia: Reaction-Diffusion Multi-Modal Fusion for Language-Guided Video Moment Retrieval
by: Fang, Xiang, et al.
Published: (2026)
by: Fang, Xiang, et al.
Published: (2026)
Generative Ghost: Investigating Ranking Bias Hidden in AI-Generated Videos
by: Gao, Haowen, et al.
Published: (2025)
by: Gao, Haowen, et al.
Published: (2025)
P-GSVC: Layered Progressive 2D Gaussian Splatting for Scalable Image and Video
by: Wang, Longan, et al.
Published: (2026)
by: Wang, Longan, et al.
Published: (2026)
GSVC: Efficient Video Representation and Compression Through 2D Gaussian Splatting
by: Wang, Longan, et al.
Published: (2025)
by: Wang, Longan, et al.
Published: (2025)
Towards Natural Language-Guided Drones: GeoText-1652 Benchmark with Spatial Relation Matching
by: Chu, Meng, et al.
Published: (2023)
by: Chu, Meng, et al.
Published: (2023)
Calibrated Multimodal Representation Learning with Missing Modalities
by: Liu, Xiaohao, et al.
Published: (2025)
by: Liu, Xiaohao, et al.
Published: (2025)
U4D: Uncertainty-Aware 4D World Modeling from LiDAR Sequences
by: Xu, Xiang, et al.
Published: (2025)
by: Xu, Xiang, et al.
Published: (2025)
Don't Just Say "I don't know"! Self-aligning Large Language Models for Responding to Unknown Questions with Explanations
by: Deng, Yang, et al.
Published: (2024)
by: Deng, Yang, et al.
Published: (2024)
Compose Your Aesthetics: Empowering Text-to-Image Models with the Principles of Art
by: Jin, Zhe, et al.
Published: (2025)
by: Jin, Zhe, et al.
Published: (2025)
One-Stage Top-$k$ Learning-to-Defer: Score-Based Surrogates with Theoretical Guarantees
by: Montreuil, Yannis, et al.
Published: (2025)
by: Montreuil, Yannis, et al.
Published: (2025)
Adversarial Robustness in Two-Stage Learning-to-Defer: Algorithms and Guarantees
by: Montreuil, Yannis, et al.
Published: (2025)
by: Montreuil, Yannis, et al.
Published: (2025)
Beyond Augmented-Action Surrogates for Multi-Expert Learning-to-Defer
by: Montreuil, Yannis, et al.
Published: (2026)
by: Montreuil, Yannis, et al.
Published: (2026)
Why Ask One When You Can Ask $k$? Learning-to-Defer to the Top-$k$ Experts
by: Montreuil, Yannis, et al.
Published: (2025)
by: Montreuil, Yannis, et al.
Published: (2025)
Similar Items
-
OpenESS: Event-based Semantic Scene Understanding with Open Vocabularies
by: Kong, Lingdong, et al.
Published: (2024) -
EventFly: Event Camera Perception from Ground to the Sky
by: Kong, Lingdong, et al.
Published: (2025) -
Visual Grounding from Event Cameras
by: Kong, Lingdong, et al.
Published: (2025) -
Talk2Event: Grounded Understanding of Dynamic Scenes from Event Cameras
by: Kong, Lingdong, et al.
Published: (2025) -
Composed Image Retrieval with Text Feedback via Multi-grained Uncertainty Regularization
by: Chen, Yiyang, et al.
Published: (2022)