Point Cloud Understanding via Attention-Driven Contrastive Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Wang, Yi, Wang, Jiaze, Guo, Ziyu, Zhang, Renrui, Zhou, Donghao, Chen, Guangyong, Liu, Anfeng, Heng, Pheng-Ann |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MM-Mixing: Multi-Modal Mixing Alignment for 3D Understanding
von: Wang, Jiaze, et al.
Veröffentlicht: (2024)
von: Wang, Jiaze, et al.
Veröffentlicht: (2024)
SignVTCL: Multi-Modal Continuous Sign Language Recognition Enhanced by Visual-Textual Contrastive Learning
von: Chen, Hao, et al.
Veröffentlicht: (2024)
von: Chen, Hao, et al.
Veröffentlicht: (2024)
MagicTailor: Component-Controllable Personalization in Text-to-Image Diffusion Models
von: Zhou, Donghao, et al.
Veröffentlicht: (2024)
von: Zhou, Donghao, et al.
Veröffentlicht: (2024)
SFANet: Spatial-Frequency Attention Network for Weather Forecasting
von: Wang, Jiaze, et al.
Veröffentlicht: (2024)
von: Wang, Jiaze, et al.
Veröffentlicht: (2024)
Beyond Binary Contrast: Modeling Continuous Skeleton Action Spaces with Transitional Anchors
von: Feng, Yingjie, et al.
Veröffentlicht: (2026)
von: Feng, Yingjie, et al.
Veröffentlicht: (2026)
SceneDecorator: Towards Scene-Oriented Story Generation with Scene Planning and Scene Consistency
von: Song, Quanjian, et al.
Veröffentlicht: (2025)
von: Song, Quanjian, et al.
Veröffentlicht: (2025)
SAM2Point: Segment Any 3D as Videos in Zero-shot and Promptable Manners
von: Guo, Ziyu, et al.
Veröffentlicht: (2024)
von: Guo, Ziyu, et al.
Veröffentlicht: (2024)
SciVerse: Unveiling the Knowledge Comprehension and Visual Reasoning of LMMs on Multi-modal Scientific Problems
von: Guo, Ziyu, et al.
Veröffentlicht: (2025)
von: Guo, Ziyu, et al.
Veröffentlicht: (2025)
Distribution-Aware Calibration for Object Detection with Noisy Bounding Boxes
von: Zhou, Donghao, et al.
Veröffentlicht: (2023)
von: Zhou, Donghao, et al.
Veröffentlicht: (2023)
CrossVideo: Self-supervised Cross-modal Contrastive Learning for Point Cloud Video Understanding
von: Liu, Yunze, et al.
Veröffentlicht: (2024)
von: Liu, Yunze, et al.
Veröffentlicht: (2024)
ATLAS: Agentic or Latent Visual Reasoning? One Word is Enough for Both
von: Guo, Ziyu, et al.
Veröffentlicht: (2026)
von: Guo, Ziyu, et al.
Veröffentlicht: (2026)
Cross-modality Guidance-aided Multi-modal Learning with Dual Attention for MRI Brain Tumor Grading
von: Xu, Dunyuan, et al.
Veröffentlicht: (2024)
von: Xu, Dunyuan, et al.
Veröffentlicht: (2024)
Thinking-while-Generating: Interleaving Textual Reasoning throughout Visual Generation
von: Guo, Ziyu, et al.
Veröffentlicht: (2025)
von: Guo, Ziyu, et al.
Veröffentlicht: (2025)
GATS: Gaussian Aware Temporal Scaling Transformer for Invariant 4D Spatio-Temporal Point Cloud Representation
von: Tian, Jiayi, et al.
Veröffentlicht: (2026)
von: Tian, Jiayi, et al.
Veröffentlicht: (2026)
LaST-R1: Reinforcing Robotic Manipulation via Adaptive Physical Latent Reasoning
von: Chen, Hao, et al.
Veröffentlicht: (2026)
von: Chen, Hao, et al.
Veröffentlicht: (2026)
Adaptive Negative Evidential Deep Learning for Open-set Semi-supervised Learning
von: Yu, Yang, et al.
Veröffentlicht: (2023)
von: Yu, Yang, et al.
Veröffentlicht: (2023)
Synergistic Bleeding Region and Point Detection in Laparoscopic Surgical Videos
von: Pei, Jialun, et al.
Veröffentlicht: (2025)
von: Pei, Jialun, et al.
Veröffentlicht: (2025)
Uni-Synergy: Bridging Understanding and Generation for Personalized Reasoning via Co-operative Reinforcement Learning
von: Shen, Zijun, et al.
Veröffentlicht: (2026)
von: Shen, Zijun, et al.
Veröffentlicht: (2026)
PMA: Towards Parameter-Efficient Point Cloud Understanding via Point Mamba Adapter
von: Zha, Yaohua, et al.
Veröffentlicht: (2025)
von: Zha, Yaohua, et al.
Veröffentlicht: (2025)
Point-In-Context: Understanding Point Cloud via In-Context Learning
von: Liu, Mengyuan, et al.
Veröffentlicht: (2024)
von: Liu, Mengyuan, et al.
Veröffentlicht: (2024)
EPContrast: Effective Point-level Contrastive Learning for Large-scale Point Cloud Understanding
von: Pan, Zhiyi, et al.
Veröffentlicht: (2024)
von: Pan, Zhiyi, et al.
Veröffentlicht: (2024)
Medical Large Vision Language Models with Multi-Image Visual Ability
von: Yang, Xikai, et al.
Veröffentlicht: (2025)
von: Yang, Xikai, et al.
Veröffentlicht: (2025)
Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO
von: Tong, Chengzhuo, et al.
Veröffentlicht: (2025)
von: Tong, Chengzhuo, et al.
Veröffentlicht: (2025)
Depth-Driven Geometric Prompt Learning for Laparoscopic Liver Landmark Detection
von: Pei, Jialun, et al.
Veröffentlicht: (2024)
von: Pei, Jialun, et al.
Veröffentlicht: (2024)
Are Video Models Ready as Zero-Shot Reasoners? An Empirical Study with the MME-CoF Benchmark
von: Guo, Ziyu, et al.
Veröffentlicht: (2025)
von: Guo, Ziyu, et al.
Veröffentlicht: (2025)
DisCo-Layout: Disentangling and Coordinating Semantic and Physical Refinement in a Multi-Agent Framework for 3D Indoor Layout Synthesis
von: Gao, Jialin, et al.
Veröffentlicht: (2025)
von: Gao, Jialin, et al.
Veröffentlicht: (2025)
Joint Learning for Scattered Point Cloud Understanding with Hierarchical Self-Distillation
von: Zhou, Kaiyue, et al.
Veröffentlicht: (2023)
von: Zhou, Kaiyue, et al.
Veröffentlicht: (2023)
Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding
von: Wang, Yan, et al.
Veröffentlicht: (2025)
von: Wang, Yan, et al.
Veröffentlicht: (2025)
Parameter-efficient Prompt Learning for 3D Point Cloud Understanding
von: Sun, Hongyu, et al.
Veröffentlicht: (2024)
von: Sun, Hongyu, et al.
Veröffentlicht: (2024)
Towards Synchronous Memorizability and Generalizability with Site-Modulated Diffusion Replay for Cross-Site Continual Segmentation
von: Xu, Dunyuan, et al.
Veröffentlicht: (2024)
von: Xu, Dunyuan, et al.
Veröffentlicht: (2024)
Does Engram Do Memory Retrieval in Autoregressive Image Generation?
von: Wang, Jinghao, et al.
Veröffentlicht: (2026)
von: Wang, Jinghao, et al.
Veröffentlicht: (2026)
DG-PIC: Domain Generalized Point-In-Context Learning for Point Cloud Understanding
von: Jiang, Jincen, et al.
Veröffentlicht: (2024)
von: Jiang, Jincen, et al.
Veröffentlicht: (2024)
Unlocking Positive Transfer in Incrementally Learning Surgical Instruments: A Self-reflection Hierarchical Prompt Framework
von: Zhu, Yu, et al.
Veröffentlicht: (2026)
von: Zhu, Yu, et al.
Veröffentlicht: (2026)
MME-CoF-Pro: Evaluating Reasoning Coherence in Video Generative Models with Text and Visual Hints
von: Qi, Yu, et al.
Veröffentlicht: (2026)
von: Qi, Yu, et al.
Veröffentlicht: (2026)
Physics-Driven Local-Whole Elastic Deformation Modeling for Point Cloud Representation Learning
von: Chen, Zhongyu, et al.
Veröffentlicht: (2025)
von: Chen, Zhongyu, et al.
Veröffentlicht: (2025)
PCF-Lift: Panoptic Lifting by Probabilistic Contrastive Fusion
von: Zhu, Runsong, et al.
Veröffentlicht: (2024)
von: Zhu, Runsong, et al.
Veröffentlicht: (2024)
Deep Omni-supervised Learning for Rib Fracture Detection from Chest Radiology Images
von: Chai, Zhizhong, et al.
Veröffentlicht: (2023)
von: Chai, Zhizhong, et al.
Veröffentlicht: (2023)
HiFi-Inpaint: Towards High-Fidelity Reference-Based Inpainting for Generating Detail-Preserving Human-Product Images
von: Liu, Yichen, et al.
Veröffentlicht: (2026)
von: Liu, Yichen, et al.
Veröffentlicht: (2026)
T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2025)
von: Jiang, Dongzhi, et al.
Veröffentlicht: (2025)
UniHOPE: A Unified Approach for Hand-Only and Hand-Object Pose Estimation
von: Wang, Yinqiao, et al.
Veröffentlicht: (2025)
von: Wang, Yinqiao, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MM-Mixing: Multi-Modal Mixing Alignment for 3D Understanding
von: Wang, Jiaze, et al.
Veröffentlicht: (2024) -
SignVTCL: Multi-Modal Continuous Sign Language Recognition Enhanced by Visual-Textual Contrastive Learning
von: Chen, Hao, et al.
Veröffentlicht: (2024) -
MagicTailor: Component-Controllable Personalization in Text-to-Image Diffusion Models
von: Zhou, Donghao, et al.
Veröffentlicht: (2024) -
SFANet: Spatial-Frequency Attention Network for Weather Forecasting
von: Wang, Jiaze, et al.
Veröffentlicht: (2024) -
Beyond Binary Contrast: Modeling Continuous Skeleton Action Spaces with Transitional Anchors
von: Feng, Yingjie, et al.
Veröffentlicht: (2026)