Abstract:
In recent years, the field of 3D vision has garnered substantial attention. However, acquiring 3D data, particularly annotated data, poses a significant challenge that impedes the progress of large- scale pre- trained models in the realm of 3D vision. In this context, the inherent abundant knowledge present in 2D pre- trained models presents a promising approach for understanding 3D point clouds by directly projecting them onto 2D images, demonstrating both feasibility and potential. However, it is still challenging to narrow the domain gap between 3D and 2D and eliminate the noise introduced during the 3D- to- 2D projection process. Therefore, this paper proposes a novel method dubbed Diff3DU that utilizes a pre- trained diffusion model to alleviate the noise problem and improve the performance of the 2D pre- trained models on 3D point cloud understanding. Comprehensive comparison and analysis of the experimental results against widely- used benchmarks across diverse settings demonstrate the effectiveness of the proposed method.