ProCap: Projection-Aware Captioning for Spatial Augmented Reality
Fuente:
arXiv
Guardado en:
| Autores principales: | Cao, Zimo, Deng, Yuchen, Ling, Haibin, Huang, Bingyao |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
LAPIG: Language Guided Projector Image Generation with Surface Adaptation and Stylization
por: Deng, Yuchen, et al.
Publicado: (2025)
por: Deng, Yuchen, et al.
Publicado: (2025)
GS-ProCams: Gaussian Splatting-based Projector-Camera Systems
por: Deng, Qingyue, et al.
Publicado: (2024)
por: Deng, Qingyue, et al.
Publicado: (2024)
See or Guess: Counterfactually Regularized Image Captioning
por: Cao, Qian, et al.
Publicado: (2024)
por: Cao, Qian, et al.
Publicado: (2024)
FinCap: Topic-Aligned Captions for Short-Form Financial YouTube Videos
por: Sukhani, Siddhant, et al.
Publicado: (2025)
por: Sukhani, Siddhant, et al.
Publicado: (2025)
ESIQA: Perceptual Quality Assessment of Vision-Pro-based Egocentric Spatial Images
por: Zhu, Xilei, et al.
Publicado: (2024)
por: Zhu, Xilei, et al.
Publicado: (2024)
PolySmart @ TRECVid 2024 Video Captioning (VTT)
por: Wu, Jiaxin, et al.
Publicado: (2024)
por: Wu, Jiaxin, et al.
Publicado: (2024)
Mitigating Image Captioning Hallucinations in Vision-Language Models
por: Zhao, Fei, et al.
Publicado: (2025)
por: Zhao, Fei, et al.
Publicado: (2025)
Towards Retrieval-Augmented Architectures for Image Captioning
por: Sarto, Sara, et al.
Publicado: (2024)
por: Sarto, Sara, et al.
Publicado: (2024)
ProFD: Prompt-Guided Feature Disentangling for Occluded Person Re-Identification
por: Cui, Can, et al.
Publicado: (2024)
por: Cui, Can, et al.
Publicado: (2024)
OneDiff: A Generalist Model for Image Difference Captioning
por: Hu, Erdong, et al.
Publicado: (2024)
por: Hu, Erdong, et al.
Publicado: (2024)
Embedded Heterogeneous Attention Transformer for Cross-lingual Image Captioning
por: Song, Zijie, et al.
Publicado: (2023)
por: Song, Zijie, et al.
Publicado: (2023)
AeroLite: Tag-Guided Lightweight Generation of Aerial Image Captions
por: Zi, Xing, et al.
Publicado: (2025)
por: Zi, Xing, et al.
Publicado: (2025)
Group-based Distinctive Image Captioning with Memory Difference Encoding and Attention
por: Wang, Jiuniu, et al.
Publicado: (2025)
por: Wang, Jiuniu, et al.
Publicado: (2025)
Spatial-Aware Efficient Projector for MLLMs via Multi-Layer Feature Aggregation
por: Qian, Shun, et al.
Publicado: (2024)
por: Qian, Shun, et al.
Publicado: (2024)
Edit As You Wish: Video Caption Editing with Multi-grained User Control
por: Yao, Linli, et al.
Publicado: (2023)
por: Yao, Linli, et al.
Publicado: (2023)
Video Summarization: Towards Entity-Aware Captions
por: Ayyubi, Hammad A., et al.
Publicado: (2023)
por: Ayyubi, Hammad A., et al.
Publicado: (2023)
SmartFreeEdit: Mask-Free Spatial-Aware Image Editing with Complex Instruction Understanding
por: Sun, Qianqian, et al.
Publicado: (2025)
por: Sun, Qianqian, et al.
Publicado: (2025)
MotionPro: A Precise Motion Controller for Image-to-Video Generation
por: Zhang, Zhongwei, et al.
Publicado: (2025)
por: Zhang, Zhongwei, et al.
Publicado: (2025)
Mesquite MoCap: Democratizing Real-Time Motion Capture with Affordable, Bodyworn IoT Sensors and WebXR SLAM
por: Vanani, Poojan, et al.
Publicado: (2025)
por: Vanani, Poojan, et al.
Publicado: (2025)
Generating Attribute-Aware Human Motions from Textual Prompt
por: Wang, Xinghan, et al.
Publicado: (2025)
por: Wang, Xinghan, et al.
Publicado: (2025)
MambaPro: Multi-Modal Object Re-Identification with Mamba Aggregation and Synergistic Prompt
por: Wang, Yuhao, et al.
Publicado: (2024)
por: Wang, Yuhao, et al.
Publicado: (2024)
Autoregressive Image Generation with Linear Complexity: A Spatial-Aware Decay Perspective
por: Mao, Yuxin, et al.
Publicado: (2025)
por: Mao, Yuxin, et al.
Publicado: (2025)
IG Captioner: Information Gain Captioners are Strong Zero-shot Classifiers
por: Yang, Chenglin, et al.
Publicado: (2023)
por: Yang, Chenglin, et al.
Publicado: (2023)
ChronusOmni: Improving Time Awareness of Omni Large Language Models
por: Chen, Yijing, et al.
Publicado: (2025)
por: Chen, Yijing, et al.
Publicado: (2025)
2D or 3D: Who Governs Salience in VLA Models? -- Tri-Stage Token Pruning Framework with Modality Salience Awareness
por: Zheng, Zihao, et al.
Publicado: (2026)
por: Zheng, Zihao, et al.
Publicado: (2026)
Spatial-Temporal Human-Object Interaction Detection
por: Sun, Xu, et al.
Publicado: (2025)
por: Sun, Xu, et al.
Publicado: (2025)
IWN: Image Watermarking Based on Idempotency
por: Deng, Kaixin
Publicado: (2024)
por: Deng, Kaixin
Publicado: (2024)
DetailVerifyBench: A Benchmark for Dense Hallucination Localization in Long Image Captions
por: Wang, Xinran, et al.
Publicado: (2026)
por: Wang, Xinran, et al.
Publicado: (2026)
Memory-enhanced Retrieval Augmentation for Long Video Understanding
por: Yuan, Huaying, et al.
Publicado: (2025)
por: Yuan, Huaying, et al.
Publicado: (2025)
Selective Vision-Language Subspace Projection for Few-shot CLIP
por: Zhu, Xingyu, et al.
Publicado: (2024)
por: Zhu, Xingyu, et al.
Publicado: (2024)
LiftProj: Space Lifting and Projection-Based Panorama Stitching
por: Jia, Yuan, et al.
Publicado: (2025)
por: Jia, Yuan, et al.
Publicado: (2025)
ESVQA: Perceptual Quality Assessment of Egocentric Spatial Videos
por: Zhu, Xilei, et al.
Publicado: (2024)
por: Zhu, Xilei, et al.
Publicado: (2024)
Visual Lifelog Retrieval through Captioning-Enhanced Interpretation
por: Shih, Yu-Fei, et al.
Publicado: (2025)
por: Shih, Yu-Fei, et al.
Publicado: (2025)
Efficient Token Compression for Vision Transformer with Spatial Information Preserved
por: Mao, Junzhu, et al.
Publicado: (2025)
por: Mao, Junzhu, et al.
Publicado: (2025)
Can Multimodal Large Language Models Understand Spatial Relations?
por: Liu, Jingping, et al.
Publicado: (2025)
por: Liu, Jingping, et al.
Publicado: (2025)
TextRefiner: Internal Visual Feature as Efficient Refiner for Vision-Language Models Prompt Tuning
por: Xie, Jingjing, et al.
Publicado: (2024)
por: Xie, Jingjing, et al.
Publicado: (2024)
Cross Modification Attention Based Deliberation Model for Image Captioning
por: Lian, Zheng, et al.
Publicado: (2021)
por: Lian, Zheng, et al.
Publicado: (2021)
MoRAG -- Multi-Fusion Retrieval Augmented Generation for Human Motion
por: Kalakonda, Sai Shashank, et al.
Publicado: (2024)
por: Kalakonda, Sai Shashank, et al.
Publicado: (2024)
Interactive Spatial-Frequency Fusion Mamba for Multi-Modal Image Fusion
por: Zhu, Yixin, et al.
Publicado: (2026)
por: Zhu, Yixin, et al.
Publicado: (2026)
Dual Attribute-Spatial Relation Alignment for 3D Visual Grounding
por: Xu, Yue, et al.
Publicado: (2024)
por: Xu, Yue, et al.
Publicado: (2024)
Ejemplares similares
-
LAPIG: Language Guided Projector Image Generation with Surface Adaptation and Stylization
por: Deng, Yuchen, et al.
Publicado: (2025) -
GS-ProCams: Gaussian Splatting-based Projector-Camera Systems
por: Deng, Qingyue, et al.
Publicado: (2024) -
See or Guess: Counterfactually Regularized Image Captioning
por: Cao, Qian, et al.
Publicado: (2024) -
FinCap: Topic-Aligned Captions for Short-Form Financial YouTube Videos
por: Sukhani, Siddhant, et al.
Publicado: (2025) -
ESIQA: Perceptual Quality Assessment of Vision-Pro-based Egocentric Spatial Images
por: Zhu, Xilei, et al.
Publicado: (2024)